Supercomputing center emergency response method and system based on data fusion analysis
Through multi-source data fusion and real-time anomaly detection, an emergency response strategy collection of supercomputer centers is generated, which solves the problem of mismatch between policy singularity and abnormal complexity in the existing technology, and realizes adaptive recovery and stability improvement of supercomputer centers.
Patent Information
- Application Number
- CN202510863795.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The emergency response methods of existing supercomputing centers are difficult to accurately identify composite abnormal characteristics, resulting in mismatch between strategies and abnormal types, low resource scheduling efficiency, and easy to cause recovery operation overload or resource conflict in the operating environment, lack of adaptive parameter adjustment mechanism with real-time feedback, resulting in increased risk of system recovery delay and stability in complex abnormal scenarios.
By obtaining multi-source real-time monitoring data, data fusion processing is carried out to generate a multi-dimensional fusion feature set, real-time abnormality detection is performed based on preset anomaly detection strategies, an emergency response strategy set is generated, and the matching parameters of the policy template are updated in real time, to realize global perception and adaptive optimization of the supercomputer center.
It improves the adaptive recovery ability and system stability of the supercomputer center for complex abnormal events, reduces the risk of misjudgment, improves resource release efficiency and environmental stability recovery capabilities, and optimizes the strategy adaptation accuracy and execution priority.
Smart Images

Figure CN120353635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a supercomputer center emergency response method and system based on data fusion analysis. Background Art
[0002] With the large-scale and intelligent development of supercomputer centers, system emergency response technology has become a key link to ensure the reliable operation of high-performance computing. The current mainstream emergency response methods usually trigger preset response rules based on device operation indicators or service request status, and generate static disposal strategies through threshold comparison. However, abnormal events in supercomputer centers are often caused by the coupling of multiple factors such as devices, environment, and service requests. Traditional data-driven response mechanisms are difficult to accurately identify complex abnormal features, resulting in inaccurate matching between strategies and abnormal types, low resource scheduling efficiency, and easy occurrence of recovery operation overload or resource conflicts due to fixed policy parameters in the operating environment. Although multi-source data monitoring has been attempted in the prior art, the lack of fusion modeling between device operation correlation features and environmental impact features makes the abnormal detection results unable to be accurately mapped to a multi-dimensional emergency strategy library. At the same time, the policy execution process lacks an adaptive parameter adjustment mechanism based on real-time feedback, further exacerbating the system recovery delay and stability risk in complex abnormal scenarios. Summary of the Invention
[0003] The present invention provides a supercomputer center emergency response method and system based on data fusion analysis.
[0004] In a first aspect, an embodiment of the present invention provides a supercomputer center emergency response method based on data fusion analysis, including: Obtaining a multi-source real-time monitoring data set of the supercomputer center, where the multi-source real-time monitoring data set includes device operation monitoring data, environmental status monitoring data, and service request monitoring data; Performing data fusion processing on the multi-source real-time monitoring data set to generate a multi-dimensional fusion feature set corresponding to the supercomputer center, where the multi-dimensional fusion feature set includes device operation correlation features, environmental impact features, and service request fluctuation features; Based on a preset abnormal detection strategy, performing real-time abnormal detection analysis on the multi-dimensional fusion feature set to generate an abnormal event set of the supercomputer center, where the abnormal event set includes multiple abnormal events and an abnormal type identifier for each abnormal event; According to the abnormal type identifier of each abnormal event in the abnormal event set, calling an emergency response strategy template matching the abnormal type identifier to generate an emergency response strategy set of the supercomputer center; Execute the target emergency response strategy in the set of emergency response strategies, and obtain the execution feedback data of the target emergency response strategy in real time. Update the matching parameters of the emergency response strategy template according to the execution feedback data.
[0005] In a second aspect, an embodiment of the present invention provides a computer system, including: A memory in which a computer program is stored; A processor for loading the computer program to implement the above-mentioned supercomputer center emergency response method based on data fusion analysis.
[0006] The supercomputer center emergency response method based on data fusion analysis provided by the present invention integrates multi-source real-time monitoring data of equipment operation, environmental status, and service requests, constructs a multi-dimensional collaborative monitoring system, and generates device operation correlation features, environmental impact features, and service request fluctuation features based on data fusion to achieve a global perception of the operation status of the supercomputer center; through the precise mapping mechanism between the anomaly type identifier and the emergency response strategy template, match the multi-dimensional anomaly detection results with the strategy execution conditions and priority parameters in the preset strategy library to generate an emergency response strategy set highly adaptable to complex anomaly events, effectively solving the problem of response lag caused by the mismatch between the single nature of the strategy and the complexity of anomalies in traditional methods; further based on the parameter update mechanism of the strategy execution feedback data, by real-time evaluating indicators such as resource release efficiency, environmental stability recovery, and service response delay improvement, adaptively optimize the matching range and execution priority of the strategy template, continuously improve the anomaly recovery efficiency and strategy adaptation accuracy without manual intervention, and at the same time, through multi-dimensional feature fusion and anomaly correlation analysis, avoid the misjudgment risk caused by noise in a single data source or local environmental fluctuations, significantly enhancing the adaptive recovery ability and system stability of the supercomputer center to cope with complex anomaly events. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 is a flowchart of a supercomputer center emergency response method based on data fusion analysis provided by an embodiment of the present invention.
[0007] Figure 2 is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0008] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0009] Please refer toFigure 1 , which is a flowchart of an emergency response method for a supercomputer center based on data fusion analysis provided by an embodiment of the present invention. This method can be executed by a computer system and may include the following steps: Step S100: Obtain a multi-source real-time monitoring data set of the supercomputer center. The multi-source real-time monitoring data set includes device operation monitoring data, environmental status monitoring data, and service request monitoring data.
[0010] The multi-source real-time monitoring data set of the supercomputer center is a series of data obtained in real time from multiple different sources and used to reflect the operation status of the supercomputer center. Among them, the device operation monitoring data is data describing the operation of various devices in the supercomputer center, such as the operation status information of servers, storage devices, etc.; the environmental status monitoring data is data on the environmental conditions of the supercomputer center, such as data on environmental factors such as power supply, network, temperature, and humidity; the service request monitoring data is request data related to the services provided by the supercomputer center, including information such as the concurrency volume and response delay of service requests.
[0011] The process of obtaining the multi-source real-time monitoring data set of the supercomputer center is as follows: First, use various sensors and monitoring modules deployed in the supercomputer center to collect different types of data. For device operation monitoring data, it is collected through a sensor network deployed on the devices. These sensors can obtain information such as the temperature, load rate, and energy consumption of the devices in real time and convert them into corresponding sequence data. For example, temperature sensors are installed at key parts of the server, and the temperature value is recorded every set time to form a device temperature sequence. For environmental status monitoring data, it is collected through environmental monitoring nodes. These nodes can monitor information such as power supply voltage fluctuations, network bandwidth occupancy rates, and environmental temperature and humidity and generate corresponding sequence data. For example, voltage sensors are installed on the power supply lines of the supercomputer center to monitor the voltage fluctuation situation in real time and obtain a power supply voltage fluctuation sequence. For service request monitoring data, it is collected through a service request monitoring module. This module can record information such as the concurrency volume, response delay, and resource occupancy rate of service requests and form corresponding sequence data. For example, monitoring points are set at the service request entrance to count the concurrency volume of service requests in each time period and obtain a service request concurrency volume sequence.
[0012] After these data are collected, since the collection times of different data sources may vary, it is necessary to perform timestamp alignment processing on the device operation monitoring data, environmental status monitoring data, and service request monitoring data to generate a time-synchronized multi-source real-time monitoring data set.
[0013] As an implementation method, step S100, obtaining the multi-source real-time monitoring data set of the supercomputer center, may specifically include the following steps S110~S140: Step S110: Collect device operation monitoring data through the device sensor network deployed in the supercomputing center. The device operation monitoring data includes the device temperature sequence, the device load rate sequence, and the device energy consumption sequence.
[0014] The device sensor network is a network system composed of multiple sensors distributed on various devices in the supercomputing center, which is used to collect the operation status information of the devices in real time. The device operation monitoring data is a series of data reflecting the operation status of the devices. Among them, the device temperature sequence is a set of device temperature values recorded in chronological order, reflecting the temperature changes of the devices at different time points; the device load rate sequence is a record of the device load rate changing with time, reflecting the working load degree of the devices at each moment; the device energy consumption sequence is the change of the device energy consumption with time, reflecting the energy consumption of the devices during operation.
[0015] In the actual collection process, the temperature sensors in the device sensor network will regularly measure the temperature of the devices and record the measured values in chronological order to form the device temperature sequence. For example, for a server, the temperature sensor measures the temperature every 5 minutes, and arranging these measured values in sequence will obtain the device temperature sequence of this server. The device load rate can be calculated by monitoring the usage of resources such as the CPU and memory of the device. Recording the load rate values at different time points will form the device load rate sequence. For example, the CPU usage rate of the server is obtained every 10 minutes through the system monitoring tool, which is used as an index of the load rate, and arranging these values in chronological order will obtain the device load rate sequence. The device energy consumption can be measured by the energy consumption monitoring device installed on the power line of the device. Recording the energy consumption values at different time points will form the device energy consumption sequence. For example, an intelligent electricity meter is installed on the power line of the server, and the energy consumption value is recorded every 15 minutes to obtain the device energy consumption sequence.
[0016] Step S120: Collect environmental status monitoring data through the environmental monitoring nodes deployed in the supercomputing center. The environmental status monitoring data includes the power supply voltage fluctuation sequence, the network bandwidth occupancy rate sequence, and the environmental temperature and humidity sequence.
[0017] The environmental monitoring nodes are devices or modules specifically used to monitor the environmental conditions of the supercomputing center. They are distributed at various key positions in the supercomputing center and can collect environmental-related data in real time. The environmental status monitoring data is a series of data reflecting the environmental conditions of the supercomputing center. Among them, the power supply voltage fluctuation sequence is a set of power supply voltage fluctuation conditions recorded in chronological order, reflecting the stability of the power supply voltage at different time points; the network bandwidth occupancy rate sequence is a record of the network bandwidth occupancy rate changing with time, reflecting the usage of the network at each moment; the environmental temperature and humidity sequence is the change of the environmental temperature and humidity with time, reflecting the temperature and humidity conditions of the internal environment of the supercomputing center.
[0018] When collecting the power supply voltage fluctuation sequence, the voltage sensor in the environmental monitoring node will monitor the voltage value of the power supply line in real time, calculate the difference between the voltage values at adjacent time points, and arrange these differences in chronological order to obtain the power supply voltage fluctuation sequence. For example, the voltage sensor measures the power supply voltage every 2 minutes, calculates the difference between each measurement value and the previous one, and records these differences in sequence to form the power supply voltage fluctuation sequence. For the network bandwidth occupancy rate sequence, the environmental monitoring node can collect the data traffic in the network through a network traffic monitoring tool, calculate the bandwidth occupancy rate at each time point, and record it in chronological order. For instance, the in and out data traffic of the network interface is counted every 3 minutes, the bandwidth occupancy rate is calculated, and the network bandwidth occupancy rate sequence is obtained. The environmental temperature and humidity sequence can be collected by a temperature and humidity sensor. The temperature and humidity sensor will regularly measure the temperature and humidity values of the environment and record these values in chronological order.
[0019] Step S130: Collect service request monitoring data through the service request monitoring module deployed in the supercomputing center. The service request monitoring data includes the service request concurrency sequence, the service response latency sequence, and the service resource occupancy rate sequence.
[0020] The service request monitoring module is a module used to monitor the request situation of the services provided by the supercomputing center and can record the relevant information of service requests in real time. The service request monitoring data is a series of data reflecting the service request situation. Among them, the service request concurrency sequence is a set of the number of service requests within the same time recorded in chronological order, reflecting the concurrency degree of service requests at different time points; the service response latency sequence is a record of the time delay from the issuance of a service request to its response changing over time, reflecting the response speed of the service at each moment; the service resource occupancy rate sequence is the change of the proportion of system resources occupied by the service during operation over time, reflecting the usage of system resources by the service.
[0021] When collecting the service request concurrency volume sequence, the service request monitoring module will count the number of requests within the same time in real time at the service request entry, and record these numbers in chronological order. For example, the number of service requests entering the supercomputing center is counted every 1 minute, and these numbers are arranged in sequence to obtain the service request concurrency volume sequence. The service response delay sequence can be obtained by recording the sending time and response time of the service request and calculating the difference between the two. The service request monitoring module will record the sending time and response time for each service request, calculate the delay time, and record these delay times in chronological order. For example, for each service request, record its sending time and the time when the response is received, calculate the difference between the two, and arrange these differences in sequence to obtain the service response delay sequence. The service resource occupancy rate sequence can be calculated by monitoring the usage of system resources such as CPU, memory, and disk during the service operation. The service request monitoring module will regularly obtain the usage ratios of the service for various resources and record these ratios in chronological order. For example, the usage ratio of the service for the CPU is obtained every 2 minutes, and these ratios are arranged in sequence to obtain the service resource occupancy rate sequence.
[0022] Step S140: Perform timestamp alignment processing on the device operation monitoring data, environmental status monitoring data, and service request monitoring data to generate a time-synchronized multi-source real-time monitoring data set.
[0023] Timestamp alignment processing is an operation to uniformly align the data collected from different data sources in terms of time. The purpose is to ensure the consistency of different types of data in time for subsequent data fusion and analysis. After the timestamp alignment processing of the multi-source real-time monitoring data set, the data points in each data sequence are corresponding in time, forming a time-synchronized data set.
[0024] The specific process of timestamp alignment processing is as follows: First, extract the first timestamp sequence of device operation monitoring data, the second timestamp sequence of environmental status monitoring data, and the third timestamp sequence of service request monitoring data. These timestamp sequences record the collection time of each data point. For example, each temperature value in the device temperature sequence corresponds to a collection time. Arranging these collection times in order gives the first timestamp sequence of device operation monitoring data. Then, according to the preset time window length, divide the first timestamp sequence, the second timestamp sequence, and the third timestamp sequence into multiple time window intervals. The preset time window length can be set according to actual needs, for example, set to 10 minutes. Divide each timestamp sequence at 10 - minute time intervals to obtain multiple time window intervals. Within each time window interval, perform data interpolation processing on the device operation monitoring data, environmental status monitoring data, and service request monitoring data to generate an interpolation data set with time - window alignment. Data interpolation processing is a method of estimating unknown data points between known data points. For example, use a linear interpolation algorithm. If there are two adjacent data points in the device temperature sequence within a time window interval, namely the temperature value T1 corresponding to time t1 and the temperature value T2 corresponding to time t2, and it is necessary to estimate the temperature value T corresponding to time t (t1 < t < t2), the linear interpolation formula T = T1+(T2 - T1)×(t - t1) / (t2 - t1) can be used for calculation. Finally, perform data smoothing processing on the interpolation data set to eliminate data acquisition noise and generate a multi - source real - time monitoring data set with time synchronization and noise filtering. Data smoothing processing can use the moving average method. For example, calculate the average value of each data point and several data points before and after it, and use this average value to replace the value of this data point, thereby reducing the impact of noise.
[0025] As an implementation manner, in step S140, performing timestamp alignment processing on the device operation monitoring data, environmental status monitoring data, and service request monitoring data may specifically include the following steps S141 to S144: Step S141: Extract the first timestamp sequence of device operation monitoring data, the second timestamp sequence of environmental status monitoring data, and the third timestamp sequence of service request monitoring data.
[0026] A timestamp sequence is a sequence that records the data collection time, providing a time mark for each data point, enabling different types of data to be associated and compared in the time dimension. The first timestamp sequence is a set of timestamps corresponding to the device operation monitoring data, recording the collection time of each data point in the device operation monitoring data; the second timestamp sequence is a set of timestamps corresponding to the environmental status monitoring data, recording the collection time of the environmental status monitoring data; the third timestamp sequence is a set of timestamps corresponding to the service request monitoring data, recording the collection time of the service request monitoring data.
[0027] In actual operation, for the device operation monitoring data, the collection time of each data point is recorded simultaneously during collection. By extracting these collection times in sequence, the first timestamp sequence is obtained. For example, each temperature value in the device temperature sequence has a corresponding collection time. Arranging these collection times in sequence forms the first timestamp sequence. Similarly, for the environmental status monitoring data and service request monitoring data, the collection times of each data point are extracted respectively to form the second timestamp sequence and the third timestamp sequence.
[0028] Step S142: Divide the first timestamp sequence, the second timestamp sequence, and the third timestamp sequence into multiple time window intervals according to a preset time window length.
[0029] The preset time window length is a preset time interval used to divide the timestamp sequence into different intervals. A time window interval is a period of time obtained by dividing the timestamp sequence in the time dimension. By dividing the time window intervals, different types of data can be aligned and grouped in time, facilitating subsequent processing and analysis.
[0030] When dividing the time window intervals, first determine the preset time window length, for example, set to 15 minutes. Then, for the first timestamp sequence, starting from the first timestamp, divide it at 15-minute intervals. If the first timestamp is t1, then the first time window interval is [t1, t1 + 15 minutes), and so on, dividing the entire first timestamp sequence into multiple time window intervals. The same method is applied to the second timestamp sequence and the third timestamp sequence to obtain the corresponding time window intervals.
[0031] Step S143: Perform data interpolation processing on the device operation monitoring data, environmental status monitoring data, and service request monitoring data within each time window interval to generate an interpolation data set with time window alignment.
[0032] Data interpolation processing is a method for estimating unknown data points between known data points. Its purpose is to supplement and improve different types of data within the time window interval, making the data more continuous and complete in time, thereby generating an interpolation data set with time window alignment.
[0033] Within each time window interval, for the device operation monitoring data, environmental status monitoring data, and service request monitoring data, there may be cases where data is not collected at certain time points. At this time, data interpolation processing is required to estimate the data values at these unknown time points. For example, using the linear interpolation algorithm, if in a time window interval, there are two adjacent data points in the device temperature sequence, which are the temperature value T1 corresponding to time t1 and the temperature value T2 corresponding to time t2, and it is necessary to estimate the temperature value T corresponding to time t (t1 < t < t2), the linear interpolation formula T = T1 + (T2 - T1) × (t - t1) / (t2 - t1) can be used for calculation. For the environmental status monitoring data and service request monitoring data, the linear interpolation algorithm can also be used for processing. By performing interpolation processing on different types of data within each time window interval, an interpolation data set aligned in time is obtained, making the different types of data more consistent in time.
[0034] Step S144: Perform data smoothing processing on the interpolation data set to eliminate data acquisition noise and generate a multi-source real-time monitoring data set that is time-synchronized and noise-filtered.
[0035] Data smoothing processing is an operation of filtering and noise reduction on data, and its purpose is to eliminate the noise generated during data acquisition, making the data smoother and more stable. After the multi-source real-time monitoring data set undergoes data smoothing processing, it not only achieves synchronization in time, but also the noise is effectively filtered, making it more suitable for subsequent data analysis and processing.
[0036] When performing data smoothing processing on the interpolation data set, the moving average method can be adopted. The moving average method calculates the average value of each data point and several data points before and after it, and uses this average value to replace the value of this data point, thereby reducing the influence of noise. For example, for the interpolation data set of the device temperature sequence, select a moving window size, such as 3. For each data point, calculate the average value of it and one data point before and after it, and use this average value to replace the value of this data point. For the interpolation data sets of the environmental status monitoring data and service request monitoring data, the moving average method is also adopted for processing. By performing data smoothing processing on the entire interpolation data set, a multi-source real-time monitoring data set that is time-synchronized and noise-filtered is generated, improving the quality and usability of the data.
[0037] Step S200: Perform data fusion processing on the multi-source real-time monitoring data set to generate a multi-dimensional fusion feature set corresponding to the supercomputer center. The multi-dimensional fusion feature set includes device operation correlation features, environmental impact features, and service request fluctuation features.
[0038] Data fusion processing is the process of integrating and analyzing data from multiple different sources to extract more valuable information and features. The multi-dimensional fusion feature set is a set containing features of multiple dimensions obtained through data fusion processing, which can more comprehensively reflect the operating status of the supercomputer center. Among them, the device operation correlation features are the features related to the operation of the supercomputer center devices, reflecting the correlation and mutual influence between devices; the environmental impact features are the features related to the environmental conditions of the supercomputer center, reflecting the impact of environmental factors on the operation of the supercomputer center; the service request fluctuation features are the features related to the request situation of the services provided by the supercomputer center, reflecting the fluctuating changes in service requests.
[0039] The process of data fusion processing for the multi-source real-time monitoring data set is as follows: First, perform feature extraction processing on the device operation monitoring data, environmental status monitoring data, and service request monitoring data respectively to obtain the device operation correlation features, environmental impact features, and service request fluctuation features. Then, according to the preset fusion weight parameters, perform weighted fusion processing on these features to generate a multi-dimensional fusion feature set. The preset fusion weight parameters can be adjusted according to the historical emergency event data of the supercomputer center to ensure that the fused features can more accurately reflect the actual situation of the supercomputer center.
[0040] As an implementation method, in step S200, perform data fusion processing on the multi-source real-time monitoring data set to generate a multi-dimensional fusion feature set corresponding to the supercomputer center, which may specifically include the following steps S210~S240: Step S210: Perform first feature extraction processing on the device operation monitoring data to obtain the device operation correlation features. The first feature extraction processing includes the correlation analysis of the device load volatility, device temperature change gradient, and device energy consumption offset.
[0041] The first feature extraction processing is the process of extracting key features from the device operation monitoring data that can reflect the device operation correlation situation. The device operation correlation features are a set of features used to describe the correlation and mutual influence between devices. Among them, the device load volatility is the change rate of the device load rate within a certain period of time, reflecting the fluctuation of the device load; the device temperature change gradient is the change rate of the device temperature in space or time, reflecting the change trend of the device temperature; the device energy consumption offset is the difference between the actual energy consumption of the device and the expected energy consumption, reflecting the abnormal situation of the device energy consumption.
[0042] When performing the first feature extraction process, first calculate the device load volatility. The load volatility for each time interval can be obtained by calculating the ratio of the difference in device load rates at adjacent time points to the load rate at the previous time point. For example, if the device load rates are L1 and L2 at times t1 and t2 (t2 > t1) respectively, the device load volatility for this time interval is (L2 - L1) / L1. Then, calculate the device temperature change gradient. The device temperature change gradient can be obtained by numerically differentiating the device temperature sequence and calculating the ratio of the difference in temperature values at adjacent time points to the time interval. For example, for two adjacent temperature values T1 and T2 in the device temperature sequence, with the corresponding time interval being Δt, the device temperature change gradient is (T2 - T1) / Δt. Finally, calculate the device energy consumption offset. First, a prediction model of the device energy consumption needs to be established based on the device's historical operation data and performance parameters to predict the expected energy consumption of the device in the current operating state. Then, subtract the actual energy consumption of the device from the expected energy consumption to obtain the device energy consumption offset. By performing a correlation analysis on the device load volatility, device temperature change gradient, and device energy consumption offset, the device operation correlation features are obtained. For example, a correlation analysis method can be used to calculate the correlation coefficients between these three features and analyze the degree of their association.
[0043] Step S220: Perform a second feature extraction process on the environmental status monitoring data to obtain environmental impact features. The second feature extraction process includes a matching analysis of the environmental temperature and humidity correlation, power supply stability volatility, and network delay offset.
[0044] The second feature extraction process is a process of extracting key features from the environmental status monitoring data that can reflect the impact of the environment on the operation of the supercomputer center. The environmental impact features are a set of features used to describe the impact of environmental factors on the operation of the supercomputer center. Among them, the environmental temperature and humidity correlation is the degree of association between the environmental temperature and humidity, reflecting the mutual relationship between the environmental temperature and humidity; the power supply stability volatility is the change rate of the power supply voltage fluctuation within a set time, reflecting the fluctuation of the power supply stability; the network delay offset is the difference between the actual network delay and the expected delay, reflecting the abnormal situation of the network delay.
[0045] When performing the second feature extraction process, first calculate the environmental temperature and humidity correlation. The correlation analysis method can be used to calculate the correlation coefficient between the environmental temperature sequence and the humidity sequence to measure the environmental temperature and humidity correlation. For example, using the Pearson correlation coefficient calculation method, for the environmental temperature sequence T = [T1, T2,..., Tn] and the humidity sequence H = [H1, H2,..., Hn], calculate their Pearson correlation coefficient r. The closer the value of r is to 1 or -1, the stronger the correlation between the environmental temperature and humidity. Then, calculate the power supply stability volatility. The power supply stability volatility for each time interval can be obtained by calculating the ratio of the difference between the power supply voltage fluctuation values at adjacent time points to the fluctuation value at the previous time point. For example, if the power supply voltage fluctuation values V1 and V2 are collected at times t1 and t2 (t2 > t1) respectively, the power supply stability volatility for this time interval is (V2 - V1) / V1. Finally, calculate the network delay offset. First, a prediction model of network delay needs to be established based on the historical performance data and topological structure of the network to predict the expected delay of the network in the current operating state. Then, subtract the actual delay of the network from the expected delay to obtain the network delay offset. By performing matching analysis on the environmental temperature and humidity correlation, power supply stability volatility, and network delay offset, environmental impact features are obtained. For example, the clustering analysis method can be used to cluster these three features and analyze their matching relationship.
[0046] Step S230: Perform a third feature extraction process on the service request monitoring data to obtain service request fluctuation features. The third feature extraction process includes trend analysis of the service request concurrency change rate, service response delay volatility, and service resource occupancy offset.
[0047] The third feature extraction process is a process of extracting key features from the service request monitoring data that can reflect the service request fluctuation situation. The service request fluctuation features are a set of features used to describe the service request fluctuation change situation. Among them, the service request concurrency change rate is the change rate of the service request concurrency within a set time, reflecting the fluctuation of the service request concurrency; the service response delay volatility is the change rate of the service response delay within a set time, reflecting the fluctuation of the service response delay; the service resource occupancy offset is the difference between the actual occupied resources of the service and the expected occupied resources, reflecting the abnormal situation of the service resource occupancy.
[0048] When performing the third feature extraction process, first calculate the change rate of service request concurrency. The change rate of service request concurrency within each time interval can be obtained by calculating the ratio of the difference in service request concurrency between adjacent time points to the concurrency at the previous time point. For example, if the service request concurrency is collected as C1 and C2 at time t1 and t2 (t2 > t1) respectively, then the change rate of service request concurrency within this time interval is (C2 - C1) / C1. Then, calculate the volatility of service response latency. The volatility of service response latency within each time interval can be obtained by calculating the ratio of the difference in service response latency values between adjacent time points to the latency value at the previous time point. For example, for two adjacent latency values D1 and D2 in the service response latency sequence, with the corresponding time interval of Δt, the volatility of service response latency is (D2 - D1) / D1. Finally, calculate the offset of service resource occupancy. First, establish a prediction model of service resource occupancy based on the historical operation data and business requirements of the service, and predict the expected resource occupancy of the service in the current operating state. Then, subtract the expected resource occupancy from the actual resource occupancy of the service to obtain the offset of service resource occupancy. By performing trend analysis on the change rate of service request concurrency, the volatility of service response latency, and the offset of service resource occupancy, the service request fluctuation characteristics are obtained. For example, time series analysis methods can be used to analyze the changing trends of these three characteristics over time.
[0049] Step S240: According to the preset fusion weight parameters, perform weighted fusion processing on the device operation correlation characteristics, environmental impact characteristics, and service request fluctuation characteristics to generate a multi-dimensional fusion feature set; wherein, the fusion weight parameters are adjusted according to the historical emergency event data of the supercomputing center.
[0050] The preset fusion weight parameters are parameters preset for weighting different characteristics, which determine the importance of each characteristic in the fusion process. The weighted fusion processing is a process of performing weighted summation on the device operation correlation characteristics, environmental impact characteristics, and service request fluctuation characteristics according to the preset fusion weight parameters to generate a comprehensive multi-dimensional fusion feature set. The fusion weight parameters can be adjusted according to the historical emergency event data of the supercomputing center to ensure that the fused characteristics can more accurately reflect the actual situation of the supercomputing center.
[0051] When performing weighted fusion processing, first obtain the historical operation data set of the supercomputer center, including historical device operation data, historical environmental fluctuation data, and historical service request data. Then, perform a weight impact analysis on these historical data to determine the first weight coefficient of the device operation association feature, the second weight coefficient of the environmental impact feature, and the third weight coefficient of the service request fluctuation feature. Machine learning algorithms, such as linear regression algorithms, can be used to establish a regression model based on historical emergency event data and feature data, and solve for the weight coefficients of each feature. For example, taking the occurrence of historical emergency events as the dependent variable, and the device operation association feature, environmental impact feature, and service request fluctuation feature as independent variables, establish a linear regression model, and solve for the first weight coefficient, the second weight coefficient, and the third weight coefficient through the least squares method. Next, according to these weight coefficients, construct a weight adjustment function. The weight adjustment function can be a simple linear function, such as F(x)=w1×x1+w2×x2+w3×x3, where w1, w2, and w3 are the first weight coefficient, the second weight coefficient, and the third weight coefficient respectively, and x1, x2, and x3 are the device operation association feature, environmental impact feature, and service request fluctuation feature respectively. Finally, perform real-time weight allocation on the device operation association feature, environmental impact feature, and service request fluctuation feature through the weight adjustment function to generate a weighted multi-dimensional fusion feature set.
[0052] As an implementation manner, in step S240, perform weighted fusion processing on the device operation association feature, environmental impact feature, and service request fluctuation feature according to the preset fusion weight parameters, which may specifically include the following steps S241 to S244: Step S241: Obtain the historical operation data set of the supercomputer center, and the historical operation data set includes historical device operation data, historical environmental fluctuation data, and historical service request data.
[0053] The historical operation data set is a set of operation data of the supercomputer center in the past period of time, including historical device operation data, historical environmental fluctuation data, and historical service request data. The historical device operation data is the operation status information of the devices in the supercomputer center in the past, such as data on device temperature, load rate, energy consumption, etc.; the historical environmental fluctuation data is the fluctuation situation of the environment where the supercomputer center is located in the past, such as power supply voltage fluctuation, network bandwidth occupancy rate, environmental temperature and humidity, etc.; the historical service request data is the request situation of the services provided by the supercomputer center in the past, such as service request concurrency, service response latency, service resource occupancy rate, etc.
[0054] The process of obtaining the historical operation data set can be achieved through the log system and database of the supercomputing center. The supercomputing center usually records the operation status of devices, environmental parameters, and service request information, and stores this data in the database. Query statements can be written to extract historical device operation data, historical environmental fluctuation data, and historical service request data within a specified time period from the database to form a historical operation data set. For example, using SQL query statements, extract the device temperature, load rate, and energy consumption data for the past year from the device operation log table, extract the power supply voltage fluctuation, network bandwidth occupancy rate, and environmental temperature and humidity data for the past year from the environmental monitoring log table, and extract the service request concurrency, service response latency, and service resource occupancy rate data for the past year from the service request log table.
[0055] Step S242: Conduct a weight impact analysis on the historical device operation data, historical environmental fluctuation data, and historical service request data to determine the first weight coefficient of the device operation correlation feature, the second weight coefficient of the environmental impact feature, and the third weight coefficient of the service request fluctuation feature.
[0056] The weight impact analysis is to analyze the historical device operation data, historical environmental fluctuation data, and historical service request data to determine the importance of each feature in the fusion process, so as to obtain the first weight coefficient of the device operation correlation feature, the second weight coefficient of the environmental impact feature, and the third weight coefficient of the service request fluctuation feature.
[0057] When conducting the weight impact analysis, machine learning algorithms such as the linear regression algorithm can be used. First, take the occurrence of historical emergency events as the dependent variable, and the device operation correlation feature, environmental impact feature, and service request fluctuation feature as independent variables to establish a linear regression model. For example, if the occurrence of historical emergency events is represented by variable Y, the device operation correlation feature is represented by variable X1, the environmental impact feature is represented by variable X2, and the service request fluctuation feature is represented by variable X3, then the linear regression model can be expressed as Y = β0 + β1×X1 + β2×X2 + β3×X3 + ε, where β0 is the intercept, β1, β2, and β3 are the first weight coefficient, the second weight coefficient, and the third weight coefficient respectively, and ε is the error term. Then, use the historical operation data set to train the linear regression model, and solve the values of β1, β2, and β3 through the least squares method, which are the weight coefficients of the device operation correlation feature, environmental impact feature, and service request fluctuation feature. During the training process, divide the historical operation data set into a training set and a test set, use the training set to train the model, and use the test set to evaluate and optimize the model to improve the accuracy and generalization ability of the model.
[0058] Step S243: Construct a weight adjustment function according to the first weight coefficient, the second weight coefficient, and the third weight coefficient.
[0059] The weight adjustment function is a function constructed based on the first weight coefficient, the second weight coefficient, and the third weight coefficient, and is used to perform real-time weight allocation on the device operation-related features, environmental impact features, and service request fluctuation features.
[0060] When constructing the weight adjustment function, a simple linear combination method can be adopted. If the first weight coefficient is w1, the second weight coefficient is w2, the third weight coefficient is w3, the device operation-related feature is x1, the environmental impact feature is x2, and the service request fluctuation feature is x3, then the weight adjustment function can be expressed as F(x)=w1×x1+w2×x2+w3×x3. This function performs weighted summation on the device operation-related features, environmental impact features, and service request fluctuation features according to their respective weight coefficients to obtain a comprehensive feature value. The role of the weight adjustment function is to perform reasonable weighted fusion on them according to the importance of different features, so that the fused features can more accurately reflect the actual operation status of the supercomputing center.
[0061] Step S244: Perform real-time weight allocation on the device operation-related features, environmental impact features, and service request fluctuation features through the weight adjustment function to generate a weighted multi-dimensional fusion feature set.
[0062] Real-time weight allocation is to immediately use the weight adjustment function to perform weighted processing on these features after obtaining the latest device operation-related features, environmental impact features, and service request fluctuation features, in order to obtain a weighted multi-dimensional fusion feature set.
[0063] When performing real-time weight allocation, substitute the current device operation-related features, environmental impact features, and service request fluctuation features into the weight adjustment function and calculate according to the calculation rules of the function. For example, substitute the device operation-related feature x1, the environmental impact feature x2, and the service request fluctuation feature x3 into the weight adjustment function F(x)=w1×x1+w2×x2+w3×x3 to calculate the weighted comprehensive feature value. Perform such processing on the feature data at each time point to obtain a series of weighted feature values, and combine these feature values to generate a weighted multi-dimensional fusion feature set. Through real-time weight allocation, the weights of different features in the fusion process can be dynamically adjusted according to their importance, so that the multi-dimensional fusion feature set can more accurately reflect the real-time operation status of the supercomputing center.
[0064] Step S300: Based on a preset anomaly detection strategy, perform real-time anomaly detection analysis on the multi-dimensional fusion feature set to generate an anomaly event set of the supercomputing center. The anomaly event set includes multiple anomaly events and the anomaly type identifier of each anomaly event.
[0065] The preset anomaly detection strategy is a set of rules and methods pre-established for detecting anomalies in the multi-dimensional fusion feature set. Real-time anomaly detection analysis is to monitor and analyze the multi-dimensional fusion feature set in real time to discover anomalies therein. The anomaly event set is a set containing multiple anomaly events and their anomaly type identifiers, recording the anomalies that occur during the operation of the supercomputer center. When performing real-time anomaly detection analysis, first, extract the first feature vector corresponding to the device operation correlation feature, the second feature vector corresponding to the environmental impact feature, and the third feature vector corresponding to the service request fluctuation feature from the multi-dimensional fusion feature set. Then, input these feature vectors into the preset anomaly detection model to output the real-time anomaly detection results of the supercomputer center, including the anomaly event identifier, the anomaly event trigger time, and the anomaly type identifier. Next, according to the anomaly event trigger time in the real-time anomaly detection results, divide the multiple anomaly events into time windows to generate multiple subsets of anomaly events. Finally, perform type aggregation analysis on the anomaly events in each subset of anomaly events to determine the dominant anomaly type identifier for each subset of anomaly events and update the anomaly type identifier of each anomaly event in the anomaly event set with the dominant anomaly type identifier.
[0066] As an implementation manner, in step S300, based on the preset anomaly detection strategy, perform real-time anomaly detection analysis on the multi-dimensional fusion feature set to generate the anomaly event set of the supercomputer center, which may specifically include the following steps S310 to S340: Step S310: Extract the first feature vector corresponding to the device operation correlation feature, the second feature vector corresponding to the environmental impact feature, and the third feature vector corresponding to the service request fluctuation feature from the multi-dimensional fusion feature set.
[0067] A feature vector represents feature data in the form of a vector, facilitating subsequent analysis and processing. The first feature vector is a vector corresponding to the device operation correlation feature, containing the feature values of each dimension of the device operation correlation feature; the second feature vector is a vector corresponding to the environmental impact feature, containing the feature values of each dimension of the environmental impact feature; the third feature vector is a vector corresponding to the service request fluctuation feature, containing the feature values of each dimension of the service request fluctuation feature.
[0068] When extracting feature vectors from the multi-dimensional fusion feature set, they can be extracted according to the dimensions and order of the features. For example, the device operation correlation feature contains the feature values of three dimensions: device load volatility, device temperature change gradient, and device energy consumption offset. Arrange these three feature values in order to form a three-dimensional first feature vector. Similarly, for the environmental impact feature and the service request fluctuation feature, extract the feature values of their respective dimensions to form the second feature vector and the third feature vector.
[0069] Step S320: Input the first feature vector, the second feature vector, and the third feature vector into a preset anomaly detection model to output the real-time anomaly detection result of the supercomputing center. The real-time anomaly detection result includes an anomaly event identifier, an anomaly event trigger time, and an anomaly type identifier.
[0070] The preset anomaly detection model is a pre-trained model for detecting anomalies. It can determine whether there are anomalies based on the input feature vectors and output the corresponding anomaly detection results. The real-time anomaly detection result is the result of detecting anomalies in the supercomputing center at the current moment, including an anomaly event identifier, an anomaly event trigger time, and an anomaly type identifier. The anomaly event identifier is a number used to uniquely identify an anomaly event; the anomaly event trigger time is the specific time when the anomaly event occurs; the anomaly type identifier is a number or label used to represent the type of the anomaly event.
[0071] When inputting the first feature vector, the second feature vector, and the third feature vector into a preset anomaly detection model, first perform standardization processing on these feature vectors to generate a first standardized vector corresponding to the device operation-related features, a second standardized vector corresponding to the environmental impact features, and a third standardized vector corresponding to the service request fluctuation features. Standardization processing can make the feature vectors have the same scale and range, improving the training effect and accuracy of the model. For example, using the Z-score standardization method, for each feature vector, calculate its mean and standard deviation, and then subtract the mean from each eigenvalue and divide by the standard deviation to obtain the standardized feature vector. Then, input the first standardized vector, the second standardized vector, and the third standardized vector into the first-level feature fusion processing unit of the anomaly detection model to perform first-level feature fusion processing and generate a fused first-level fusion feature vector. The first-level feature fusion processing unit can adopt the fully connected layer of a neural network to perform linear combination and non-linear transformation on the input feature vectors to obtain the fused feature vector. Next, input the first-level fusion feature vector into the second-level anomaly correlation analysis unit of the anomaly detection model to perform second-level anomaly correlation analysis and generate a first anomaly correlation degree between the device operation-related features and the environmental impact features, a second anomaly correlation degree between the environmental impact features and the service request fluctuation features, and a third anomaly correlation degree between the service request fluctuation features and the device operation-related features. The second-level anomaly correlation analysis unit can adopt a correlation analysis algorithm, such as the Pearson correlation coefficient calculation method, to calculate the correlation degree between different features. Based on the first anomaly correlation degree, the second anomaly correlation degree, and the third anomaly correlation degree, determine the comprehensive anomaly probability distribution of the supercomputer center, and generate the anomaly event identifier and the anomaly type identifier in the real-time anomaly detection result according to the anomaly type identifier corresponding to the probability peak in the comprehensive anomaly probability distribution. Finally, extract the timestamp sequences of the first standardized vector, the second standardized vector, and the third standardized vector, determine the anomaly event trigger time according to the timestamp corresponding to the probability peak in the timestamp sequence, and associate and integrate the anomaly event trigger time with the anomaly event identifier and the anomaly type identifier to generate a real-time anomaly detection result including the anomaly event identifier, the anomaly event trigger time, and the anomaly type identifier.
[0072] As an implementation, in step S320, input the first feature vector, the second feature vector, and the third feature vector into a preset anomaly detection model to output the real-time anomaly detection result of the supercomputer center, which may specifically include the following steps S321 to S325: Step S321: Perform standardization processing on the first feature vector, the second feature vector, and the third feature vector to generate a first standardized vector corresponding to the device operation-related features, a second standardized vector corresponding to the environmental impact features, and a third standardized vector corresponding to the service request fluctuation features.
[0073] Normalization is to uniformly process feature vectors of different scales and ranges so that they have the same mean and standard deviation, thereby improving the training effect and accuracy of the model. The first normalized vector is the vector obtained by normalizing the first feature vector and corresponds to the device operation-related features; the second normalized vector is the vector obtained by normalizing the second feature vector and corresponds to the environmental impact features; the third normalized vector is the vector obtained by normalizing the third feature vector and corresponds to the service request fluctuation features.
[0074] When performing normalization, the Z-score normalization method can be used. For the first feature vector x1 = [x11, x12,..., x1n], first calculate its mean μ1 and standard deviation σ1, and the calculation formulas are respectively and , where n is the dimension of the feature vector. Then, subtract the mean from each element in the first feature vector and divide by the standard deviation to obtain the first normalized vector y1 = [(x11 - μ1) / σ1, (x12 - μ1) / σ1,..., (x1n - μ1) / σ1]. Similarly, perform the same normalization process on the second and third feature vectors to obtain the second normalized vector y2 and the third normalized vector y3.
[0075] Step S322: Input the first normalized vector, the second normalized vector, and the third normalized vector into the first-level feature fusion processing unit of the anomaly detection model to perform first-level feature fusion processing and generate a fused first-level fusion feature vector.
[0076] The first-level feature fusion processing unit is a module in the anomaly detection model, which is used to fuse the input first normalized vector, second normalized vector, and third normalized vector to generate a comprehensive feature vector. The first-level fusion feature vector is the vector obtained after the first-level feature fusion processing and contains the comprehensive information of the device operation-related features, environmental impact features, and service request fluctuation features.
[0077] The first-level feature fusion processing unit can adopt the fully connected layer of a neural network. The fully connected layer is a common neural network layer that connects each input neuron to each output neuron. When inputting the first normalized vector, the second normalized vector, and the third normalized vector into the first-level feature fusion processing unit, the fully connected layer will perform linear combination and non-linear transformation on the input vectors. Specifically, the fully connected layer multiplies the input vectors by a weight matrix, then adds a bias vector, and then performs non-linear transformation through an activation function to obtain the output first-level fusion feature vector. For example, if the first normalized vector is y1, the second normalized vector is y2, the third normalized vector is y3, the weight matrix is W, the bias vector is b, and the activation function is f, then the first-level fusion feature vector z can be expressed as z = f(W × [y1; y2; y3] + b), where [y1; y2; y3] represents concatenating the three vectors column-wise into a new vector.
[0078] Step S323: Input the first-level fusion feature vector into the second-level anomaly correlation analysis unit of the anomaly detection model to perform second-level anomaly correlation analysis, and generate the first anomaly correlation degree between the device operation correlation feature and the environmental impact feature, the second anomaly correlation degree between the environmental impact feature and the service request fluctuation feature, and the third anomaly correlation degree between the service request fluctuation feature and the device operation correlation feature.
[0079] The second-level anomaly correlation analysis unit is a module in the anomaly detection model, which is used to analyze the anomaly correlation degree between different features. The first anomaly correlation degree is the anomaly correlation degree between the device operation correlation feature and the environmental impact feature, reflecting the mutual influence between the device operation situation and environmental factors; the second anomaly correlation degree is the anomaly correlation degree between the environmental impact feature and the service request fluctuation feature, reflecting the relationship between environmental factors and service request situations; the third anomaly correlation degree is the anomaly correlation degree between the service request fluctuation feature and the device operation correlation feature, reflecting the interaction between service request situations and device operation situations.
[0080] When performing second-level anomaly correlation analysis, a correlation analysis algorithm such as the Pearson correlation coefficient calculation method can be used. For the device operation correlation feature part and the environmental impact feature part in the first-level fusion feature vector, calculate the Pearson correlation coefficient between them to obtain the first anomaly correlation degree. The value obtained by calculating the Pearson correlation coefficient is closer to 1 or -1, indicating a stronger anomaly correlation degree between the device operation correlation feature and the environmental impact feature; the closer it is to 0, the weaker the correlation degree. Similarly, perform the same calculation on the environmental impact feature part and the service request fluctuation feature part, and the service request fluctuation feature part and the device operation correlation feature part respectively to obtain the second anomaly correlation degree and the third anomaly correlation degree.
[0081] Step S324: Based on the first anomaly correlation degree, the second anomaly correlation degree, and the third anomaly correlation degree, determine the comprehensive anomaly probability distribution of the supercomputer center, and generate the anomaly event identifier and the anomaly type identifier in the real-time anomaly detection result according to the anomaly type identifier corresponding to the probability peak in the comprehensive anomaly probability distribution.
[0082] The comprehensive anomaly probability distribution is the distribution of the probability of anomalies occurring in the supercomputer center over different anomaly types by combining the first anomaly correlation degree, the second anomaly correlation degree, and the third anomaly correlation degree. It can be determined by establishing a probability model, such as using the Gaussian Mixture Model (GMM). The Gaussian Mixture Model is a probability model that can model complex data distributions, if the data is composed of multiple Gaussian distributions mixed together.
[0083] First, input the first anomaly correlation degree, the second anomaly correlation degree, and the third anomaly correlation degree as features into the trained Gaussian Mixture Model. When training the Gaussian Mixture Model, use historical data as the training set. These historical data contain the anomaly correlation degree features under different anomaly types and the corresponding anomaly type identifiers. Determine the parameters of the Gaussian Mixture Model, such as the means, covariance matrices, and weights of each Gaussian distribution, through methods such as maximum likelihood estimation.
[0084] After obtaining the comprehensive anomaly probability distribution of the supercomputer center, find the probability peak among them. The probability peak indicates the anomaly type that the supercomputer center is most likely to have under the current situation. According to the anomaly type identifier corresponding to the probability peak, combined with the pre-set anomaly event numbering rule, generate the anomaly event identifier and the anomaly type identifier in the real-time anomaly detection result. For example, if the anomaly type identifier corresponding to the probability peak is "equipment overheating anomaly", assign a unique anomaly event identifier, such as "AE001", to this anomaly event according to the numbering rule, and record "equipment overheating anomaly" as the anomaly type identifier in the real-time anomaly detection result.
[0085] Step S325: Extract the timestamp sequences of the first standardized vector, the second standardized vector, and the third standardized vector. Determine the anomaly event trigger time according to the timestamp corresponding to the probability peak in the timestamp sequence, and associate and integrate the anomaly event trigger time with the anomaly event identifier and the anomaly type identifier to generate a real-time anomaly detection result including the anomaly event identifier, the anomaly event trigger time, and the anomaly type identifier.
[0086] The timestamp sequence records the acquisition time of each data point in the first normalized vector, the second normalized vector, and the third normalized vector. When determining the abnormal event trigger time, first extract the respective timestamp sequences from these normalized vectors. Then, find the timestamp corresponding to the probability peak in the comprehensive abnormal probability distribution. This timestamp is the time when the abnormal event is most likely to occur, and it is determined as the abnormal event trigger time.
[0087] Next, associate and integrate the abnormal event trigger time with the previously generated abnormal event identifier and abnormal type identifier. These information can be stored in a data structure, such as a dictionary, where the keys are "abnormal event identifier", "abnormal event trigger time", and "abnormal type identifier", and the values are the corresponding specific information. In this way, a real-time abnormal detection result including the abnormal event identifier, the abnormal event trigger time, and the abnormal type identifier is generated.
[0088] Step S330: According to the abnormal event trigger time in the real-time abnormal detection result, perform time window partitioning on multiple abnormal events to generate multiple abnormal event subsets.
[0089] Time window partitioning is an operation of grouping multiple abnormal events at a preset time interval according to the abnormal event trigger time. Its purpose is to group the abnormal events that occur in a similar time for more detailed analysis of each group of abnormal events later.
[0090] First, determine the preset time window length, for example, set it to 30 minutes. Then, sort all abnormal events according to the abnormal event trigger time. Starting from the first abnormal event, use its trigger time as the starting point to divide a time window with a length of 30 minutes, and group all the abnormal events triggered within this time window into an abnormal event subset. Then, use the end time of this time window as the new starting point to continue dividing the next 30-minute time window, and repeat the above operation until all abnormal events are divided into the corresponding abnormal event subsets.
[0091] For example, if the trigger time of abnormal event A is "2024-01-01 12:00:00", the trigger time of abnormal event B is "2024-01-01 12:15:00", the trigger time of abnormal event C is "2024-01-01 12:30:00", and the trigger time of abnormal event D is "2024-01-01 13:05:00". With a 30-minute time window length, abnormal events A, B, and C will be grouped into one abnormal event subset, and abnormal event D will be grouped into another abnormal event subset.
[0092] Step S340: Perform type aggregation analysis on the abnormal events in each subset of abnormal events, determine the dominant abnormal type identifier for each subset of abnormal events, and update the dominant abnormal type identifier to the abnormal type identifier of each abnormal event in the set of abnormal events.
[0093] Type aggregation analysis is to count and analyze the abnormal events in each subset of abnormal events according to the abnormal type identifier to determine the dominant abnormal type identifier in the subset. The dominant abnormal type identifier represents the most common or main abnormal type in the subset of abnormal events.
[0094] As an implementation, in step S340, performing type aggregation analysis on the abnormal events in each subset of abnormal events, determining the dominant abnormal type identifier for each subset of abnormal events, and updating the dominant abnormal type identifier to the abnormal type identifier of each abnormal event in the set of abnormal events may specifically include the following steps S341 - S345: Step S341: Traverse each abnormal event in the subset of abnormal events, extract the abnormal type identifier of the abnormal event, and count the number of occurrences of each abnormal type identifier in the subset of abnormal events.
[0095] Traverse each abnormal event in the subset of abnormal events, and extract its abnormal type identifier from the information of each abnormal event. A dictionary can be used to count the number of occurrences of each abnormal type identifier. The key of the dictionary is the abnormal type identifier, and the value is the number of times the abnormal type identifier appears.
[0096] Step S342: Compare the number of occurrences of the abnormal type identifier with a preset type occurrence threshold. If the number of occurrences of the abnormal type identifier exceeds the type occurrence threshold, mark the abnormal type identifier as a candidate dominant abnormal type identifier.
[0097] The preset type occurrence threshold is a pre - set number of times value used to screen out the identifiers that may become the dominant abnormal type. Compare the number of occurrences of each abnormal type identifier with this threshold. If the number of occurrences of an abnormal type identifier exceeds the threshold, mark it as a candidate dominant abnormal type identifier.
[0098] For example, the preset type occurrence threshold is 2. For the above statistical results, "device overheating anomaly" appears 2 times and "network latency anomaly" appears 1 time. Then, "device overheating anomaly" will be marked as a candidate dominant abnormal type identifier.
[0099] Step S343: According to the timestamp sequence of the candidate dominant abnormal type identifier in the subset of abnormal events, screen out the candidate dominant abnormal type identifier corresponding to the latest timestamp in the timestamp sequence, and determine the candidate dominant abnormal type identifier as the dominant abnormal type identifier of the subset of abnormal events.
[0100] For each candidate dominant anomaly type identifier, extract the timestamp sequence in which it appears in the subset of anomaly events. The timestamp sequence records the triggering time of the anomaly events corresponding to the anomaly type identifier. Then, find the latest timestamp from these timestamp sequences, and the candidate dominant anomaly type identifier corresponding to this timestamp is the dominant anomaly type identifier of this subset of anomaly events.
[0101] For example, if "device overheating anomaly" has two occurrence timestamps in the subset of anomaly events, which are "2024-01-01 12:00:00" and "2024-01-01 12:15:00" respectively, the occurrence timestamp of "network latency anomaly" is "2024-01-01 12:05:00", and "device overheating anomaly" is the candidate dominant anomaly type identifier, then the "device overheating anomaly" corresponding to the latest timestamp "2024-01-01 12:15:00" will be determined as the dominant anomaly type identifier of this subset of anomaly events.
[0102] Step S344: Traverse each anomaly event in the anomaly event set. If the anomaly event belongs to the subset of anomaly events, then associate and match the dominant anomaly type identifier of the subset of anomaly events with the original anomaly type identifier of the anomaly event.
[0103] Traverse each anomaly event in the anomaly event set and determine whether the anomaly event belongs to the current subset of anomaly events. If it belongs, then associate and match the dominant anomaly type identifier of the subset of anomaly events with the original anomaly type identifier of the anomaly event. The association and matching can be achieved by comparing whether the two identifiers are the same.
[0104] For example, for an anomaly event in the anomaly event set, its original anomaly type identifier is "network latency anomaly", while the dominant anomaly type identifier of the current subset of anomaly events is "device overheating anomaly", then this association and matching situation needs to be recorded.
[0105] Step S345: When the dominant anomaly type identifier is inconsistent with the original anomaly type identifier, update the anomaly type identifier of the anomaly event to the dominant anomaly type identifier, and bind the updated anomaly type identifier with the event identifier of the anomaly event to generate the updated anomaly type identifier for each anomaly event in the anomaly event set.
[0106] If the dominant exception type identifier is inconsistent with the original exception type identifier, it indicates that the exception event may have a closer association with the dominant exception situation in the subset of this exception event. It is necessary to update the exception type identifier of this exception event to the dominant exception type identifier. Then, bind the updated exception type identifier with the event identifier of this exception event, store it in the exception event set, and generate the updated exception type identifier for each exception event.
[0107] For example, for the above exception event with the original exception type identifier of "network latency exception", update its exception type identifier to "device overheating exception", bind it with the event identifier of this exception event, and update the corresponding information in the exception event set.
[0108] Step S400: According to the exception type identifier of each exception event in the exception event set, call the emergency response policy template that matches the exception type identifier to generate the emergency response policy set of the supercomputer center.
[0109] The emergency response policy template is a pre-established template for response strategies for different exception types, which contains the specific steps and measures for handling exception events. The emergency response policy set is a set of a series of emergency response policies generated by calling the corresponding emergency response policy template according to the exception type identifier of each exception event in the exception event set.
[0110] As an implementation method, in step S400, according to the exception type identifier of each exception event in the exception event set, call the emergency response policy template that matches the exception type identifier to generate the emergency response policy set of the supercomputer center, which can specifically include the following steps S410~S440: Step S410: Obtain the preset emergency response policy template library. Multiple emergency response policy templates are stored in the emergency response policy template library, and each emergency response policy template is associated with at least one exception type identifier and a priority parameter.
[0111] The preset emergency response policy template library is a database or data structure that stores multiple emergency response policy templates. Each emergency response policy template is associated with at least one exception type identifier and a priority parameter. The exception type identifier is used to identify the exception type applicable to this emergency response policy template, and the priority parameter is used to determine the execution order of this emergency response policy template. The emergency response policy template library can be obtained through database query or file reading, etc.
[0112] Step S420: According to the exception type identifier of each exception event in the exception event set, screen out the candidate emergency response policy templates that match the exception type identifier from the emergency response policy template library.
[0113] Traverse each exception event in the set of exception events and extract its exception type identifier. Then, based on these exception type identifiers, screen out the corresponding emergency response policy templates from the emergency response policy template library. Screening can be performed by querying the database or traversing the data structure. The screened candidate emergency response policy templates will serve as the basis for generating the emergency response policy set subsequently.
[0114] Step S430: Based on the priority parameters of the candidate emergency response policy templates and the occurrence timestamps of the exception events in the set of exception events, sort the execution order of multiple candidate emergency response policy templates to generate a policy execution sequence.
[0115] The priority parameters and the occurrence timestamps of the exception events are important bases for determining the execution order of the candidate emergency response policy templates. First, perform a preliminary sort on the candidate emergency response policy templates according to their priority parameters, and the templates with higher priorities are ranked in the front. Then, for the templates with the same priority, sort them according to the occurrence timestamps of the exception events, and the templates corresponding to the exception events that occurred earlier are ranked in the front.
[0116] Step S440: According to the execution order in the policy execution sequence, perform parameter adaptation processing on multiple candidate emergency response policy templates to generate an emergency response policy set that matches the current operating state of the supercomputer center.
[0117] Parameter adaptation processing is to adjust and optimize the parameters in the candidate emergency response policy templates according to the current operating state of the supercomputer center, so that the generated emergency response policies can better adapt to the actual situation of the supercomputer center.
[0118] According to the execution order in the policy execution sequence, perform parameter adaptation processing on each candidate emergency response policy template in turn. For example, for an emergency response policy template for equipment overheating exception, it may contain parameters for adjusting the rotation speed of the equipment fan. When performing parameter adaptation processing, dynamically adjust the parameter value of the fan rotation speed according to the current operating state of the equipment such as temperature and load in the supercomputer center. By performing parameter adaptation processing on all candidate emergency response policy templates, an emergency response policy set that matches the current operating state of the supercomputer center is generated.
[0119] Step S500: Execute the target emergency response policy in the emergency response policy set, and obtain the execution feedback data of the target emergency response policy in real time. Update the matching parameters of the emergency response policy template according to the execution feedback data.
[0120] The target emergency response strategy is a specific strategy selected from the set of emergency response strategies for execution. The execution feedback data is the data reflecting the execution effect of the strategy collected in real time during the execution of the target emergency response strategy. The matching parameter is a parameter in the emergency response strategy template used to match the abnormal type and determine the execution order, etc.
[0121] As an implementation manner, in step S500, execute the target emergency response strategy in the set of emergency response strategies, and obtain the execution feedback data of the target emergency response strategy in real time, which may specifically include the following steps S510 to S550: Step S510: According to the strategy execution conditions of each emergency response strategy in the set of emergency response strategies, detect whether the current device operating status, the current environmental stability index, and the current service request response index of the supercomputer center meet the strategy execution conditions.
[0122] The strategy execution condition is the condition required to execute the strategy specified in the emergency response strategy, which is usually related to the device operating status, environmental stability index, and service request response index of the supercomputer center. The current device operating status includes information such as the temperature, load rate, and energy consumption of the device; the current environmental stability index includes information such as the stability of the power supply voltage, network bandwidth stability, and environmental temperature and humidity stability; the current service request response index includes information such as the service request concurrency and service response delay.
[0123] These index information are collected in real time through various sensors and monitoring modules deployed in the supercomputer center, and then compared with the strategy execution conditions of the emergency response strategy. For example, for an emergency response strategy, its strategy execution condition is "the device temperature exceeds 80°C", and the device temperature is monitored in real time through a temperature sensor to determine whether this condition is met.
[0124] Step S520: If the current device operating status, the current environmental stability index, and the current service request response index meet the strategy execution conditions of the target emergency response strategy, allocate strategy execution resources for the target emergency response strategy in the resource scheduling module of the supercomputer center. The strategy execution resources include device resource scheduling permissions, environmental regulation permissions, and service response priority adjustment permissions.
[0125] If the current device operating status, the current environmental stability index, and the current service request response index meet the strategy execution conditions of the target emergency response strategy, it indicates that the strategy can be executed. At this time, allocate the required strategy execution resources for the target emergency response strategy in the resource scheduling module of the supercomputer center.
[0126] The device resource scheduling permission allows the policy to schedule the devices in the supercomputing center, such as adjusting the operating parameters of the devices, allocating the computing resources of the devices, etc. The environmental control permission allows the policy to control the environment of the supercomputing center, such as adjusting the power supply voltage, regulating the network bandwidth, controlling the environmental temperature and humidity, etc. The service response priority adjustment permission enables the policy to adjust the response priority of service requests to ensure that important service requests can be processed in a timely manner.
[0127] For example, for an emergency response policy for device overheating abnormality, if the policy execution conditions are met, the resource scheduling module can allocate the device resource scheduling permission to it, allowing the policy to increase the rotation speed of the device fan to reduce the device temperature.
[0128] Step S530: Trigger the execution operation of the target emergency response policy through the resource scheduling module, and in the process of execution, collect the device resource occupancy data, environmental control operation record data, and service request response optimization data corresponding to the target emergency response policy in real time.
[0129] The resource scheduling module is responsible for triggering the execution operation of the target emergency response policy. During the execution process, collect the relevant data corresponding to the target emergency response policy in real time.
[0130] The device resource occupancy data includes information such as the CPU usage rate, memory usage rate, and disk I / O of the device during the execution of the policy, which is collected through device monitoring tools. The environmental control operation record data includes the specific operation records of environmental control, such as the adjusted value of the power supply voltage, the allocated situation of the network bandwidth, and the adjusted parameters of the environmental temperature and humidity, which are recorded through environmental monitoring devices. The service request response optimization data includes the change in the concurrency of service requests, the improvement of service response latency, etc., which are collected through the service request monitoring module.
[0131] For example, when executing an emergency response policy for adjusting the device load, collect the change in the CPU usage rate of the device in real time as the device resource occupancy data, record the specific operation of adjusting the device load as the environmental control operation record data, and count the improvement of the service request response latency as the service request response optimization data.
[0132] Step S540: Divide the device resource occupancy data, environmental control operation record data, and service request response optimization data into time windows to generate multiple execution phase data sets, and perform abnormal recovery index extraction processing on each execution phase data set to obtain the resource release efficiency parameter, environmental stability recovery parameter, and service response latency improvement parameter of each execution phase data set.
[0133] Time window division is an operation of grouping the collected device resource occupancy data, environmental regulation operation record data, and service request response optimization data at a preset time interval. Each time window corresponds to an execution phase, generating multiple execution phase data sets.
[0134] Anomaly recovery index extraction and processing is to extract indicators from each execution phase data set that can reflect the anomaly recovery situation. The resource release efficiency parameter is a measure of the efficiency of the device in releasing resources during the execution of the policy. For example, the rate at which the device CPU usage drops after the execution of the policy. The environmental stability recovery parameter is an evaluation of the degree to which the environment returns to stability after the execution of the policy, such as the proportion by which the fluctuation range of the power supply voltage shrinks after the execution of the policy. The service response delay improvement parameter reflects the improvement of the service request response delay after the execution of the policy, such as the average reduction value of the service response delay.
[0135] For example, divide the collected data into time windows of 10 minutes. For each 10-minute execution phase data set, calculate the rate of decrease in the device CPU usage during this phase as the resource release efficiency parameter, calculate the proportion of the reduction in the fluctuation range of the power supply voltage during this phase as the environmental stability recovery parameter, and calculate the average reduction value of the service response delay during this phase as the service response delay improvement parameter.
[0136] Step S550: Generate an evaluation result of the policy execution efficiency of the target emergency response policy based on the resource release efficiency parameter, the environmental stability recovery parameter, and the service response delay improvement parameter, and associate and integrate the policy execution efficiency evaluation result with the time stamp of the execution phase data set to generate execution feedback data containing the policy execution efficiency parameter and the anomaly recovery success rate parameter.
[0137] The policy execution efficiency evaluation result is the result of evaluating the execution efficiency of the target emergency response policy by comprehensively considering the resource release efficiency parameter, the environmental stability recovery parameter, and the service response delay improvement parameter. An evaluation model can be established to generate the policy execution efficiency evaluation result. For example, use the weighted summation method to assign different weights to the resource release efficiency parameter, the environmental stability recovery parameter, and the service response delay improvement parameter, and then sum them up weighted to obtain the policy execution efficiency evaluation result.
[0138] The anomaly recovery success rate parameter is a measure of the probability that the target emergency response policy successfully recovers abnormal situations during execution. The anomaly recovery success rate can be determined based on whether the indicators in the execution phase data set meet the preset recovery criteria.
[0139] Associate and integrate the policy execution efficiency evaluation result with the time stamp of the execution phase data set to generate execution feedback data containing the policy execution efficiency parameter and the anomaly recovery success rate parameter.
[0140] As an implementation manner, in step S500, the matching parameters of the emergency response policy template are updated according to the execution feedback data, which may specifically include the following steps S560 to S5100: Step S560: Extract the policy execution efficiency parameter and the abnormal recovery success rate parameter from the execution feedback data, and obtain the original priority parameter and the original abnormal type identifier set corresponding to the target emergency response policy in the emergency response policy template library.
[0141] Extract the policy execution efficiency parameter and the abnormal recovery success rate parameter from the execution feedback data. These two parameters reflect the execution effect of the target emergency response policy. At the same time, obtain the original priority parameter and the original abnormal type identifier set corresponding to the target emergency response policy from the emergency response policy template library. The original priority parameter is used to determine the position of the policy in the execution order, and the original abnormal type identifier set is used to identify the abnormal types applicable to the policy.
[0142] The original priority parameter and the original abnormal type identifier set can be obtained by querying the records of the emergency response policy template library.
[0143] Step S570: Calculate the policy execution efficiency deviation degree of the target emergency response policy according to the difference between the policy execution efficiency parameter and the preset efficiency reference value, and calculate the abnormal recovery success rate deviation degree of the target emergency response policy according to the difference between the abnormal recovery success rate parameter and the preset success rate reference value.
[0144] The preset efficiency reference value and the preset success rate reference value are preset standard values for evaluating the policy execution effect. The policy execution efficiency deviation degree is the difference between the policy execution efficiency parameter and the preset efficiency reference value, reflecting the deviation degree of the policy execution efficiency from the standard. The abnormal recovery success rate deviation degree is the difference between the abnormal recovery success rate parameter and the preset success rate reference value, reflecting the difference between the abnormal recovery success rate and the standard.
[0145] For example, if the preset efficiency reference value is 0.9 and the policy execution efficiency parameter is 0.8, then the policy execution efficiency deviation degree is 0.8 - 0.9 = -0.1. If the preset success rate reference value is 0.95 and the abnormal recovery success rate parameter is 0.9, then the abnormal recovery success rate deviation degree is 0.9 - 0.95 = -0.05.
[0146] Step S580: Determine the policy optimization weight of the target emergency response policy based on the policy execution efficiency deviation degree and the abnormal recovery success rate deviation degree, and adjust the original priority parameter according to the policy optimization weight to generate the updated priority parameter.
[0147] The policy optimization weight is a weight value determined based on the deviation degree of policy execution efficiency and the deviation degree of abnormal recovery success rate, and is used to adjust the original priority parameter of the target emergency response policy. The policy optimization weight can be determined by establishing a weight calculation model. For example, using the method of linear combination, different weight coefficients are assigned to the deviation degree of policy execution efficiency and the deviation degree of abnormal recovery success rate, and then they are weighted and summed to obtain the policy optimization weight.
[0148] Step S590: If the deviation degree of policy execution efficiency is lower than the preset efficiency deviation threshold, incrementally adjust the device resource scheduling permission parameter in the resource scheduling module of the target emergency response policy, and update the adjusted device resource scheduling permission parameter to the execution condition of the emergency response policy template.
[0149] The preset efficiency deviation threshold is a deviation value set in advance, used to determine whether the policy execution efficiency is too low. If the deviation degree of policy execution efficiency is lower than the preset efficiency deviation threshold, it indicates that the policy execution efficiency is low, and it is necessary to incrementally adjust the device resource scheduling permission parameter in the resource scheduling module of the target emergency response policy.
[0150] For example, the preset efficiency deviation threshold is -0.05, and the deviation degree of policy execution efficiency is -0.1, which is lower than this threshold. At this time, incrementally adjust the device resource scheduling permission parameter, such as increasing the allocation ratio of the device CPU resource by 10%. Then, update the adjusted device resource scheduling permission parameter to the execution condition of the emergency response policy template so that the adjusted parameter can be used when the policy is executed next time.
[0151] Step S5100: If the deviation degree of abnormal recovery success rate is higher than the preset success rate deviation threshold, expand the original abnormal type identification set of the target emergency response policy to the associated abnormal type identifications that are not matched in the abnormal event set, and bind the expanded associated abnormal type identifications with the updated priority parameter to generate the updated matching parameter of the emergency response policy template.
[0152] The preset success rate deviation threshold is a deviation value set in advance, used to determine whether the abnormal recovery success rate is too low. If the deviation degree of abnormal recovery success rate is higher than the preset success rate deviation threshold, it indicates that the abnormal recovery success rate is low, and it is necessary to expand the original abnormal type identification set of the target emergency response policy.
[0153] For example, the preset success rate deviation threshold is -0.03, and the abnormal recovery success rate deviation is -0.05, which is higher than this threshold. At this time, find the associated abnormal type identifiers in the abnormal event set that do not match the target emergency response strategy, and add these identifiers to the original abnormal type identifier set. Then, bind the extended associated abnormal type identifiers to the updated priority parameters to generate the updated matching parameters of the emergency response strategy template.
[0154] As an implementation manner, after executing the target emergency response strategy in the emergency response strategy set in step S500, the embodiments of the present invention may further include the following steps S600 to S900: Step S600: Real-time monitor the device operation status, environmental stability indicators, and service request response indicators of the supercomputer center to generate an emergency recovery monitoring data set.
[0155] Real-time monitor the device operation status, environmental stability indicators, and service request response indicators of the supercomputer center, and collect relevant data through various sensors and monitoring modules deployed in the supercomputer center. The device operation status monitoring includes information such as the temperature, load rate, and energy consumption of the device; the environmental stability indicator monitoring includes information such as the power supply voltage stability, network bandwidth stability, and environmental temperature and humidity stability; the service request response indicator monitoring includes information such as the service request concurrency and service response latency.
[0156] Organize and store the collected data to generate an emergency recovery monitoring data set. This data set records the changes in various indicators of the supercomputer center after executing the emergency response strategy, and is used to evaluate the emergency recovery effect subsequently.
[0157] Step S700: Compare and analyze the emergency recovery monitoring data set with the preset emergency recovery reference data to determine the emergency recovery effect score of the target emergency response strategy.
[0158] The preset emergency recovery reference data is a set of standard data set in advance, representing the values of various indicators of the supercomputer center in the normal operation state or the ideal emergency recovery state. Compare and analyze each indicator in the emergency recovery monitoring data set with the preset emergency recovery reference data.
[0159] For example, for the device temperature indicator, the device temperature recorded in the emergency recovery monitoring data set is 70°C, and the standard value of the device temperature in the preset emergency recovery reference data is 60°C. Then, the deviation value of this indicator can be calculated. By performing similar comparative analyses on all indicators, comprehensively considering the deviation situations of each indicator, an evaluation model is used to determine the emergency recovery effect score of the target emergency response strategy. The evaluation model can adopt the method of weighted summation, assign different weights to each indicator, and then sum the weighted deviation values of each indicator to obtain the emergency recovery effect score.
[0160] Step S800: If the emergency recovery effect score is lower than the preset score threshold, trigger the iterative optimization instruction for the emergency response strategy, and re-match the strategy parameters in the emergency response strategy set according to the iterative optimization instruction.
[0161] The preset score threshold is a pre-set scoring criterion used to determine whether the emergency recovery effect meets the requirements. If the emergency recovery effect score is lower than the preset score threshold, it indicates that the emergency recovery effect is not good, and the iterative optimization instruction for the emergency response strategy needs to be triggered.
[0162] According to the iterative optimization instruction, re-match the strategy parameters in the emergency response strategy set. The re-matching process can refer to the previous parameter adaptation processing method, and combine the data in the emergency recovery monitoring data set to adjust and optimize the strategy parameters to improve the execution effect of the emergency response strategy.
[0163] Step S900: Update the re-matched strategy parameters to the emergency response strategy template library and generate an emergency response strategy optimization report.
[0164] Update the re-matched strategy parameters to the emergency response strategy template library so that the optimized parameters can be used when the emergency response strategy is executed next time. At the same time, generate an emergency response strategy optimization report, which should include the analysis results of the emergency recovery monitoring data set, the emergency recovery effect score, the re-matching situation of the strategy parameters, etc. This report can provide a reference for the subsequent optimization of the emergency response strategy and help continuously improve the emergency response ability of the supercomputer center.
[0165] It can be understood that in the above introductions of the embodiments of the present invention, various algorithms involved, such as the Pearson correlation coefficient algorithm, the linear interpolation algorithm, etc., can be obtained from relevant content in the prior art. To save space, they will not be elaborated in the embodiments of the present invention. In addition, those skilled in the art can make detailed supplements according to the common general knowledge in the art. For example, according to the general knowledge in the art, normalization can be used to eliminate the dimensional conflict before feature fusion, interpolation can be used to eliminate the dimensional difference, historical data, experience or business scenario requirements can be combined to reasonably set the threshold, the model can be trained based on the general model training method, the number of layers in the model structure can be set based on actual needs, the activation function can be selected, etc. The present invention will no longer give redundant introductions to the overly detailed implementation process here.
[0166] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a computer system provided by an embodiment of the present invention. The computer system at least includes a processor 101, a communication interface 102, and a memory 103. Among them, the processor 101, the communication interface 102, and the memory 103 can be connected through a bus or other means. Among them, the processor 101 (or the Central Processing Unit (CPU)) is the computing core and control core of the computer system, which can parse various instructions in the computer system and process various data in the computer system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for the transmission and interaction of internal data in the computer system. The memory 103 (Memory) is a memory device in the computer system, used to store programs and data. It can be understood that the memory 103 here can include both the built-in memory of the computer system and, of course, the extended memory supported by the computer system. The memory 103 provides a storage space, and the operating system of the computer system is stored in this storage space, which is not limited in the present invention. In one embodiment, the processor 101 executes the supercomputer center emergency response method based on data fusion analysis provided above in the embodiments of the present invention by running the computer program in the memory 103.
Claims
1. A supercomputer center emergency response method based on data fusion analysis, characterized in that, Including: Obtain a multi-source real-time monitoring data set of the supercomputing center, where the multi-source real-time monitoring data set includes device operation monitoring data, environmental status monitoring data, and service request monitoring data; Perform data fusion processing on the multi-source real-time monitoring data set to generate a multi-dimensional fusion feature set corresponding to the supercomputing center, where the multi-dimensional fusion feature set includes device operation correlation features, environmental impact features, and service request fluctuation features; Based on a preset anomaly detection strategy, perform real-time anomaly detection analysis on the multi-dimensional fusion feature set to generate an anomaly event set of the supercomputing center, where the anomaly event set includes multiple anomaly events and the anomaly type identifier of each anomaly event; According to the anomaly type identifier of each anomaly event in the anomaly event set, call an emergency response strategy template that matches the anomaly type identifier to generate an emergency response strategy set of the supercomputing center; Execute the target emergency response strategy in the emergency response strategy set, and obtain the execution feedback data of the target emergency response strategy in real time, and update the matching parameters of the emergency response strategy template according to the execution feedback data.
2. The method according to claim 1, characterized in that, The performing data fusion processing on the multi-source real-time monitoring data set to generate a multi-dimensional fusion feature set corresponding to the supercomputing center includes: Perform first feature extraction processing on the device operation monitoring data to obtain device operation correlation features, where the first feature extraction processing includes correlation analysis of device load volatility, device temperature change gradient, and device energy consumption offset; Perform second feature extraction processing on the environmental status monitoring data to obtain environmental impact features, where the second feature extraction processing includes matching analysis of environmental temperature and humidity correlation, power supply stability volatility, and network delay offset; Perform third feature extraction processing on the service request monitoring data to obtain service request fluctuation features, where the third feature extraction processing includes trend analysis of service request concurrency change rate, service response delay volatility, and service resource occupancy offset; According to preset fusion weight parameters, perform weighted fusion processing on the device operation correlation features, environmental impact features, and service request fluctuation features to generate the multi-dimensional fusion feature set; where the fusion weight parameters are adjusted according to the historical emergency event data of the supercomputing center.
3. The method according to claim 2, characterized in that, The performing real-time anomaly detection analysis on the multi-dimensional fusion feature set based on a preset anomaly detection strategy to generate an anomaly event set of the supercomputing center includes: Extract a first feature vector corresponding to the device operation correlation feature, a second feature vector corresponding to the environmental impact feature, and a third feature vector corresponding to the service request fluctuation feature from the multi-dimensional fusion feature set; Input the first feature vector, the second feature vector, and the third feature vector into a preset anomaly detection model, and output the real-time anomaly detection result of the supercomputing center, where the real-time anomaly detection result includes an anomaly event identifier, an anomaly event trigger time, and an anomaly type identifier; According to the trigger time of the abnormal events in the real-time abnormal detection results, perform time window partitioning on multiple abnormal events to generate multiple subsets of abnormal events; Perform type aggregation analysis on the abnormal events in each subset of abnormal events, determine the dominant abnormal type identifier for each subset of abnormal events, and update the abnormal type identifier of each abnormal event in the abnormal event set with the dominant abnormal type identifier.
4. The method according to claim 3, wherein According to the abnormal type identifier of each abnormal event in the abnormal event set, call the emergency response policy template matching the abnormal type identifier to generate the emergency response policy set of the supercomputing center, including: Obtain a preset emergency response policy template library, which stores multiple emergency response policy templates, and each emergency response policy template is associated with at least one abnormal type identifier and a priority parameter; According to the abnormal type identifier of each abnormal event in the abnormal event set, screen out the candidate emergency response policy templates that match the abnormal type identifier from the emergency response policy template library; Based on the priority parameter of the candidate emergency response policy template and the occurrence timestamp of the abnormal events in the abnormal event set, sort the execution order of multiple candidate emergency response policy templates to generate a policy execution sequence; According to the execution order in the policy execution sequence, perform parameter adaptation processing on multiple candidate emergency response policy templates to generate the emergency response policy set that matches the current operating state of the supercomputing center.
5. The method according to claim 4, wherein Execute the target emergency response policy in the emergency response policy set and obtain the execution feedback data of the target emergency response policy in real time, including: According to the policy execution conditions of each emergency response policy in the emergency response policy set, detect whether the current device operating state, the current environmental stability index, and the current service request response index of the supercomputing center meet the policy execution conditions; If the current device operating state, the current environmental stability index, and the current service request response index meet the policy execution conditions of the target emergency response policy, allocate policy execution resources for the target emergency response policy in the resource scheduling module of the supercomputing center, and the policy execution resources include device resource scheduling permissions, environmental regulation permissions, and service response priority adjustment permissions; Trigger the execution operation of the target emergency response policy through the resource scheduling module, and collect the device resource occupancy data, environmental regulation operation record data, and service request response optimization data corresponding to the target emergency response policy in real time during the execution process; Perform time window partitioning on the device resource occupancy data, environmental regulation operation record data, and service request response optimization data to generate multiple execution phase data sets, and perform abnormal recovery index extraction processing on each execution phase data set to obtain the resource release efficiency parameter, environmental stability recovery parameter, and service response delay improvement parameter of each execution phase data set; Generate the policy execution efficiency evaluation result of the target emergency response policy according to the resource release efficiency parameter, the environmental stability recovery parameter, and the service response delay improvement parameter, and associate and integrate the policy execution efficiency evaluation result with the time stamp of the execution phase data set to generate the execution feedback data including the policy execution efficiency parameter and the abnormal recovery success rate parameter.
6. The method according to claim 5, wherein Updating the matching parameters of the emergency response policy template according to the execution feedback data includes: Extract the policy execution efficiency parameter and the abnormal recovery success rate parameter from the execution feedback data, and obtain the original priority parameter and the original abnormal type identification set corresponding to the target emergency response policy in the emergency response policy template library; Calculate the policy execution efficiency deviation degree of the target emergency response policy according to the difference between the policy execution efficiency parameter and the preset efficiency benchmark value, and calculate the abnormal recovery success rate deviation degree of the target emergency response policy according to the difference between the abnormal recovery success rate parameter and the preset success rate benchmark value; Based on the policy execution efficiency deviation degree and the abnormal recovery success rate deviation degree, determine the policy optimization weight of the target emergency response policy, and adjust the original priority parameter according to the policy optimization weight to generate an updated priority parameter; If the policy execution efficiency deviation degree is lower than the preset efficiency deviation threshold, incrementally adjust the device resource scheduling permission parameter in the resource scheduling module of the target emergency response policy, and update the adjusted device resource scheduling permission parameter to the execution condition of the emergency response policy template; If the abnormal recovery success rate deviation degree is higher than the preset success rate deviation threshold, expand the original abnormal type identification set of the target emergency response policy to the associated abnormal type identifications that are not matched in the abnormal event set, and bind the expanded associated abnormal type identifications to the updated priority parameter to generate the updated matching parameters of the emergency response policy template.
7. The method according to claim 2, wherein The weighted fusion process of the device operation association feature, the environmental impact feature, and the service request fluctuation feature according to the preset fusion weight parameter includes: Obtain the historical operation data set of the supercomputer center, and the historical operation data set includes historical device operation data, historical environmental fluctuation data, and historical service request data; Conduct a weight impact analysis on the historical device operation data, historical environmental fluctuation data, and historical service request data to determine the first weight coefficient of the device operation association feature, the second weight coefficient of the environmental impact feature, and the third weight coefficient of the service request fluctuation feature; Construct a weight adjustment function according to the first weight coefficient, the second weight coefficient, and the third weight coefficient; Perform real-time weight allocation on the device operation association feature, the environmental impact feature, and the service request fluctuation feature through the weight adjustment function to generate the weighted multi-dimensional fusion feature set.
8. The method according to claim 3, characterized in that Inputting the first feature vector, the second feature vector, and the third feature vector into a preset abnormal detection model to output the real-time abnormal detection result of the supercomputer center includes: Normalize the first feature vector, the second feature vector, and the third feature vector to generate a first normalized vector corresponding to the device operation related feature, a second normalized vector corresponding to the environmental impact feature, and a third normalized vector corresponding to the service request fluctuation feature; Input the first normalized vector, the second normalized vector, and the third normalized vector into the first-level feature fusion processing unit of the anomaly detection model to perform first-level feature fusion processing and generate a fused first-level fusion feature vector; Input the first-level fusion feature vector into the second-level anomaly correlation analysis unit of the anomaly detection model to perform second-level anomaly correlation analysis and generate a first anomaly correlation degree between the device operation related feature and the environmental impact feature, a second anomaly correlation degree between the environmental impact feature and the service request fluctuation feature, and a third anomaly correlation degree between the service request fluctuation feature and the device operation related feature; Based on the first anomaly correlation degree, the second anomaly correlation degree, and the third anomaly correlation degree, determine the comprehensive anomaly probability distribution of the supercomputer center, and generate the anomaly event identifier and the anomaly type identifier in the real-time anomaly detection result according to the anomaly type identifier corresponding to the probability peak in the comprehensive anomaly probability distribution; Extract the timestamp sequences of the first normalized vector, the second normalized vector, and the third normalized vector, determine the anomaly event trigger time according to the timestamp corresponding to the probability peak in the timestamp sequence, and associate and integrate the anomaly event trigger time with the anomaly event identifier and the anomaly type identifier to generate the real-time anomaly detection result including the anomaly event identifier, the anomaly event trigger time, and the anomaly type identifier; 9. The method according to claim 3, wherein The type aggregation analysis of the anomaly events in each anomaly event subset to determine the dominant anomaly type identifier of each anomaly event subset and update the dominant anomaly type identifier to the anomaly type identifier of each anomaly event in the anomaly event set includes: Traverse each anomaly event in the anomaly event subset, extract the anomaly type identifier of the anomaly event, and count the number of occurrences of each anomaly type identifier in the anomaly event subset; Compare the number of occurrences of the anomaly type identifier with a preset type occurrence threshold. If the number of occurrences of the anomaly type identifier exceeds the type occurrence threshold, mark the anomaly type identifier as a candidate dominant anomaly type identifier; According to the timestamp sequence of the appearance of the candidate dominant anomaly type identifier in the anomaly event subset, filter out the candidate dominant anomaly type identifier corresponding to the latest timestamp in the timestamp sequence, and determine the candidate dominant anomaly type identifier as the dominant anomaly type identifier of the anomaly event subset; Traverse each anomaly event in the anomaly event set. If the anomaly event belongs to the anomaly event subset, perform an association match between the dominant anomaly type identifier of the anomaly event subset and the original anomaly type identifier of the anomaly event; When the dominant exception type identifier is inconsistent with the original exception type identifier, update the exception type identifier of the exception event to the dominant exception type identifier, and bind the updated exception type identifier with the event identifier of the exception event to generate the updated exception type identifier of each exception event in the exception event set.
10. A computer system, characterized in that, Including: A memory, in which a computer program is stored; A processor, configured to load the computer program to implement the supercomputer center emergency response method based on data fusion analysis according to any one of claims 1-9.
Citation Information
Patent Citations
Reactive voltage control method and system, terminal equipment and storage medium
CN116914766A
Police data anomaly detection method and system based on deep learning
CN118427815A
Multi-source heterogeneous data fusion method and system
CN118503915A
Abnormal event emergency scheduling method and device, equipment and storage medium
CN119228022A
Multi-source heterogeneous data fusion communication system and method
CN119697046A
Cited By
Hyper-converged server multi-resource integration system and scheduling method
CN120561343A
Emergency event identification method based on neural network model and computer equipment
CN120744858A
Emergency event identification method based on neural network model and computer device
CN120744858B
Abnormal mobile application determination method and device, equipment and medium
CN120930142A
Document conversion service stability optimization method and system based on artificial intelligence
CN121144274A