An AI-based data center operation and maintenance security system
By constructing an operational intent fingerprint set and counterfactual security twin modeling, the problem of identifying weak anomalies and locating vulnerable points in the data center operation and maintenance process is solved, achieving highly accurate and low-interference anomaly identification and handling, and ensuring the security and stability of the data center.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YIYUN TECHNOLOGY CO LTD
- Filing Date
- 2026-04-17
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies cannot effectively identify subtle anomalies and locate weak security points during data center operation and maintenance, leading to accidental shutdowns and security risks. In particular, during single-cabinet PDU hot-swapping, cold aisle baffle reset, or battery inspection, extremely short-term temperature rises, current fluctuations, and access control opening signals are often misjudged as normal samples, causing abnormal backflow or thermal runaway signs to be masked.
The system constructs an operation and maintenance intent fingerprint set, combines the disturbed closed domain determination module, counterfactual security twin modeling, and multi-source data mapping, obtains task information through the operation and maintenance intent fingerprint construction module, establishes a normal disturbance model through the counterfactual security twin modeling module, separates abnormal disturbances through the disturbance decomposition module, tracks the propagation of abnormalities through the propagation tracking module, and finally locates the security vulnerability location and outputs the handling instructions through the security vulnerability location determination module.
It enables precise differentiation between normal and abnormal disturbances caused by operation and maintenance, improves the accuracy of identifying weak, hidden, and rare anomalies, reduces the risk of false alarms and missed alarms, and can accurately locate security vulnerabilities, output adaptive handling instructions, and ensure the security and robustness of continuous data center operation.
Smart Images

Figure CN122286586A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data center operation and maintenance security technology, and specifically to an AI-based data center operation and maintenance security system. Background Technology
[0002] When data centers perform single-cabinet PDU hot-swap verification, cold aisle baffle reset, or battery inspection at night, extremely short-term temperature rises, current fluctuations, and access control opening signals often overlap. Existing threshold alarms or global AI baselines can easily write such rare, small-scale maintenance disturbances into normal samples, causing abnormal return current in a single-row cabinet or early signs of thermal runaway in the backup power supply to be masked, leading to unintended shutdowns. Therefore, there is an urgent need for a method that can identify weak anomalies and locate security vulnerabilities in operation and maintenance scenarios. Summary of the Invention
[0003] The purpose of this invention is to provide an AI-based data center operation and maintenance security system to address the shortcomings of the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: an AI-based data center operation and maintenance security system, comprising: Operation and maintenance intent fingerprint construction module: Obtain the work order information, operation and maintenance personnel identity information, entry path information, operation terminal records and operation device list of the target operation and maintenance task, and construct the operation and maintenance intent fingerprint set F corresponding to the target operation and maintenance task; Disturbed Closed Domain Determination Module: Based on the equipment topology, power connection, and airflow organization of F and the target data center, determine the disturbed closed domain C related to the target operation and maintenance task and the corresponding permissible disturbance boundary B; Counterfactual security twin modeling module: Select a group of reference devices G outside the disturbed closed domain C that have the same load characteristics as C and have not participated in this operation and maintenance, and combine C and G to establish a counterfactual security twin model W to characterize normal operation and maintenance disturbances; Disturbance decomposition module: Collects multi-source real-time data sequence A within the target operation and maintenance period, including current sequence, temperature rise sequence, access control sequence, operation log sequence and network session sequence, and maps A to W to obtain the interpretable disturbance component M triggered by operation and maintenance intention; Propagation tracing module: Based on A, M and the permissible disturbance boundary B, the residual abnormal sequence D that exceeds the scope of the operation and maintenance intention is separated, and the reverse propagation tracing of D is performed based on the time delay coupling relationship between devices to determine the initial source point P of the abnormality and the abnormal propagation chain L; Security Vulnerability Location Determination Module: Analyzes the deviation entropy of operation and maintenance intentions and the cross-domain coupling hysteresis gain coefficient of each node in P, L and the disturbed closed domain C to determine the security vulnerability location Q corresponding to the target operation and maintenance task, and outputs the blocking, rollback or isolation handling instructions corresponding to Q.
[0005] Preferably, determining the disturbed closed domain C and the corresponding permissible disturbance boundary B related to the target operation and maintenance task includes: Map the operated device objects in the operation and maintenance intention fingerprint set to the device topology relationship, extract the set of devices that have a direct connection or second-order adjacency relationship with the operated device, and form the initial topology association domain; Based on the initial topology association domain, and combined with the power connection relationship, the device nodes that share a power supply branch or have a current coupling relationship with the operated device are identified, and the initial topology association domain is expanded to obtain the power coupling extended domain. Based on the aforementioned power coupling extended domain, the set of devices with cold and hot channel associations or airflow return paths is determined according to the airflow organization relationship, and the power coupling extended domain is further extended to form a disturbed closed domain. Based on the historical operational fluctuation range of each device within the disturbed closed domain and the disturbance characteristics of the corresponding operation type in the operation and maintenance intention fingerprint set, differentiated disturbance thresholds are set for each device to construct an allowable disturbance boundary.
[0006] Preferably, the step of establishing a counterfactual security twin model W using C and G to characterize normal operational disturbances includes: Based on the historical operating data of each device within the disturbed closed domain, a set of load characteristic parameters is extracted. The set of load characteristic parameters includes at least periodic load fluctuation characteristics, peak power distribution characteristics, and load change rate characteristics. A load characteristic vector is then constructed based on the set of load characteristic parameters. Based on the load feature vector, a set of devices with an Euclidean distance less than a preset similarity threshold and no operation record within the target maintenance time window is selected outside the disturbed closed domain to determine the reference device group; The devices within the disturbed closed domain are mapped and paired one by one with the reference device group according to the device type and connection relationship. A disturbance response function is established based on historical synchronous operation data. The disturbance response function is used to characterize the natural change relationship between devices under the condition of no operation and maintenance intervention. Based on the disturbance response function and the mapping pairing relationship, a counterfactual security twin model is constructed to generate operational status prediction results under the condition of no operation and maintenance intervention.
[0007] Preferably, obtaining the interpretable disturbance component M includes: Time alignment processing is performed on current sequence, temperature rise sequence, access control sequence, operation log sequence and network session sequence to construct a multi-source synchronous data matrix under a unified time axis, and interpolation method is used to fill in missing data to obtain a standardized multi-source data sequence; The standardized multi-source data sequence is input into the counterfactual secure twin model to obtain the counterfactual prediction sequence at the corresponding time point, and the difference between the actual data sequence and the counterfactual prediction sequence is calculated to form the initial perturbation difference sequence. Based on the perturbation pattern corresponding to the operation type in the operation and maintenance intent fingerprint, the initial perturbation difference sequence is subjected to pattern matching and component decomposition to extract the perturbation component consistent with the operation and maintenance operation, and candidate perturbation components are obtained. The candidate disturbance components are constrained and verified to meet the allowable disturbance boundary range, and an interpretable disturbance component that conforms to the operation and maintenance intention is output.
[0008] Preferably, the step of separating the residual abnormal sequence D that exceeds the scope of the operation and maintenance intention based on A, M, and the permissible disturbance boundary B includes: The multi-source real-time data sequence is aligned with the interpretable perturbation component on a time-by-time basis. The difference between the two is calculated and compared with the allowable perturbation boundary. The difference data exceeding the allowable perturbation boundary range is extracted to form a residual abnormal sequence.
[0009] Preferably, determining the initial source point P of the anomaly and the anomaly propagation chain L includes: constructing an anomaly triggering time matrix between devices based on the anomaly occurrence time of each device in the residual anomaly sequence, and calculating the time delay difference between any two devices to obtain a time delay coupling relationship matrix between devices; according to the time delay coupling relationship matrix and the device topology connection relationship, using the minimum time delay path search method to backtrack the residual anomaly sequence to determine the device node where the anomaly signal first appears, as the initial source point of the anomaly; starting from the initial source point of the anomaly, and combining the propagation order relationship between each node in the time delay coupling relationship matrix, constructing an anomaly propagation path sequence, thereby determining the anomaly propagation chain.
[0010] Preferably, the method for calculating the deviation of the operation and maintenance intention from the entropy includes: Based on the entry path information and operation step sequence in the operation and maintenance intent fingerprint set, a standard operation and maintenance path state transition sequence is constructed, and the transition probability between each adjacent state is calculated to form a standard path transition probability matrix. Based on the access control sequence, operation log sequence, and network session sequence in the multi-source real-time data sequence, the actual operation and maintenance behavior path sequence is extracted, and an actual path transition probability matrix is constructed according to the same state division rule. The actual path transition probability matrix and the standard path transition probability matrix are subjected to difference mapping, the corresponding state transition probability deviation is calculated, and a weighted path probability distribution is constructed based on the deviation. The path information entropy is calculated based on the weighted path probability distribution and used as the operation and maintenance intent deviation entropy.
[0011] Preferably, the calculation of the cross-domain coupling hysteresis gain coefficient includes: Time series data corresponding to different physical domains within the disturbed closed domain are acquired separately, and the time series data are uniformly sampled and detrended. Each time series is converted into its corresponding frequency domain representation using Fast Fourier Transform to obtain the power spectral density at each frequency component. The cross-power spectral density between any two physical domains is calculated based on the frequency domain representation, and a frequency domain coherence function is constructed based on the cross-power spectral density and their respective power spectral densities. The dominant frequency component with the largest coherence value is selected from the frequency domain coherence function, and the phase difference at the corresponding frequency is calculated. The phase difference is converted into a time lag. Based on the time lag and the coherence strength of the corresponding frequency component, the amplitude transfer ratio of the disturbance between different physical domains is calculated, and the cross-domain coupling lag gain coefficient is constructed.
[0012] Preferably, determining the security vulnerability location Q includes: Based on the initial source of the anomaly and the anomaly propagation chain, the temporal position of each node in the propagation path within the disturbed closed domain is extracted, and a node propagation sequence is constructed. At the same time, the deviation entropy of the operation and maintenance intention and the cross-domain coupling hysteresis gain coefficient of each node are obtained to form a multi-dimensional feature set. The multidimensional feature set is normalized, and the deviation entropy of operation and maintenance intention and the cross-domain coupling lag gain coefficient are weighted and fused based on preset weights to construct a node anomaly contribution index, wherein the weights are determined by the statistical results of the impact of the two types of features on the fault in historical anomaly samples. Based on the node's abnormal contribution index and the node's position in the propagation sequence, the node's cumulative risk value is calculated. The cumulative risk value is obtained by superimposing the node's own abnormal contribution value and the contribution value of its upstream node according to the time delay decay function. The node with the highest cumulative risk value is identified as the vulnerable location Q.
[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention constructs an operational intent fingerprint set and combines it with disturbed closed-domain delineation, counterfactual security twin modeling, and multi-source data mapping to achieve a fine distinction between "normal disturbances caused by operations and maintenance" and "abnormal disturbances deviating from operational intent." Compared to traditional anomaly detection methods based on thresholds or single models, this invention introduces path information entropy to characterize the temporal deviation of operational behavior and utilizes frequency domain coherence functions to quantify the propagation coupling relationship across physical domains. This elevates anomaly identification from single-point feature judgment to a multi-dimensional joint analysis of "behavioral semantics + physical propagation," thereby significantly improving the accuracy of identifying weak, hidden, and rare anomalies and reducing the risk of false alarms and missed alarms.
[0014] 2. This invention achieves precise location of security vulnerabilities by integrating the initial source of anomalies, the anomaly propagation chain, and the node risk accumulation mechanism. It also adaptively outputs blocking, rollback, or isolation commands based on differences in risk sources. This method not only enables proactive intervention before anomalies spread widely but also avoids the interference caused by traditional "one-size-fits-all" handling strategies to normal operations and maintenance. Therefore, it improves overall operational security and system robustness while ensuring continuous data center operation. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0016] Figure 1 This is a flowchart of the system modules of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] For examples, please refer to Figure 1 As shown in this embodiment, an AI-based data center operation and maintenance security system includes: Operation and maintenance intent fingerprint construction module: Obtain the work order information, operation and maintenance personnel identity information, entry path information, operation terminal records and operation device list of the target operation and maintenance task, and construct the operation and maintenance intent fingerprint set F corresponding to the target operation and maintenance task.
[0019] In some embodiments of the present invention, the operation and maintenance intent fingerprint construction module is used to integrate multi-source heterogeneous information and perform structured expression before the target operation and maintenance task is executed, and construct an operation and maintenance intent fingerprint set F that can characterize the true intent of the operation and maintenance task, so as to serve as the basic input for subsequent determination of disturbed closed domain and anomaly identification.
[0020] Specifically, the work order information obtained by the operation and maintenance intent fingerprint construction module includes at least the work order number, task type, planned execution time window, target device object, and standard operation steps; the operation and maintenance personnel identity information obtained includes at least the personnel identifier, permission level, work group, and historical operation qualifications; the entry path information obtained includes at least the access control record, path node sequence, and dwell time at each node; the operation terminal record obtained includes at least the terminal identifier, login time, operation command sequence, and session association; and the list of operated devices obtained includes at least the device identifier, device type, physical location, and associated connection relationship.
[0021] After obtaining the above information, each type of information is encoded to form a work order encoding subset f1, an identity encoding subset f2, a path encoding subset f3, a terminal behavior encoding subset f4, and a device object encoding subset f5. Each encoding subset is represented by a unified data structure as a multi-dimensional feature vector or structured data with attribute labels to ensure consistent processing in subsequent steps.
[0022] Furthermore, to avoid semantic loss due to simple concatenation, the operation and maintenance intent fingerprint construction module performs constraint fusion on f1 to f5 based on preset association rules. The preset association rules are a set of rules used to establish consistency, correspondence, and constraint relationships between different subsets of information, and include at least: (1) Time consistency rules are used to constrain the work order planning time, path arrival time and terminal login time to meet the same time window; (2) Spatial mapping rules, used to establish the correspondence between the end point of the path and the physical location of the operated device; (3) Permission matching rules, used to constrain the matching relationship between the permission level of operation and maintenance personnel and the operation level of work orders; (4) Operation correspondence rules, used to constrain the inclusion or sequential consistency relationship between the terminal command sequence and the standard operation steps of the work order; (5) Device range constraint rules, which are used to limit the actual access range of devices to not exceed the list of devices to be operated or their topological adjacency range.
[0023] Under the constraints of the above rules, each encoded subset is associated, bound, and fused to generate an operation and maintenance intent fingerprint set F. Preferably, F represents a constrained structured feature set, which not only contains the features of each subset but also the association information between the subsets, thereby forming a multi-dimensional coupled expression between "task-person-path-terminal-device".
[0024] In some implementations, the maintenance intent fingerprint set F can be further constructed as graph structure data, where nodes correspond to work orders, personnel, path nodes, terminals and devices, and edges are used to represent time association, spatial association and operational association, thereby improving the ability to express complex maintenance behaviors.
[0025] The operation and maintenance intent fingerprint set F constructed in the above manner, compared with traditional single logs or simple feature splicing, can express the expected behavioral boundaries of operation and maintenance tasks under a unified framework. This provides a calculable basis for separating "disturbances that conform to operation and maintenance intent" and "abnormalities that deviate from operation and maintenance intent" from actual operation data, thereby improving the accuracy and interpretability of anomaly identification.
[0026] Disturbed Closed Domain Determination Module: Based on the equipment topology, power connection, and airflow organization of F and the target data center, determine the disturbed closed domain C related to the target operation and maintenance task and the corresponding permissible disturbance boundary B.
[0027] In some embodiments of the present invention, firstly, the operated device objects in the operation and maintenance intention fingerprint set are mapped to device topology relationships. Specifically, a data center device topology graph is pre-constructed, where each device is a node and network connection links or physical connection relationships are edges, forming a directed or undirected graph structure. Based on the operated device identifier in the operation and maintenance intention fingerprint set, the corresponding node is located in the device topology graph, and a graph traversal algorithm is used to extract first-order adjacent nodes that have direct connections to the node, as well as second-order adjacent nodes that are indirectly connected through first-order nodes. This set of nodes is defined as the initial topology association domain. Adjacency relationships are determined by edge connection relationships, and the traversal process uses a breadth-first search algorithm, with the traversal depth limited to two levels to control the scope of influence.
[0028] Based on the initial topology association domain, power-coupled device nodes are identified by combining power connection relationships. Specifically, a data center power connection model is constructed, which models the connection relationships between devices and power distribution units, uninterruptible power supplies (UPS), and power distribution branches based on power supply paths. By querying the power supply branch identifiers corresponding to each device within the initial topology association domain, device nodes located on the same power supply branch as the operated device or sharing a common UPS output terminal are identified, and the current coupling coefficient between devices is calculated. The current coupling coefficient is defined as the correlation coefficient of the current change sequences of two devices during historical operation. It is calculated by aligning the current sequences within a time window and calculating the Pearson correlation coefficient. When the current coupling coefficient is greater than a preset coupling threshold, the corresponding device is included in the extended range, thus obtaining the power coupling extended domain.
[0029] Based on the aforementioned power coupling extended domain, further expansion is made according to airflow organization relationships. Specifically, a data center airflow organization model is constructed, spatially modeling rack locations, cold aisles, hot aisles, and air conditioning supply and return air paths to form an airflow path diagram. Based on the rack location of the equipment in the power coupling extended domain, other equipment located in the same cold aisle or with an airflow return path is identified, and an airflow influence coefficient is calculated. This airflow influence coefficient is defined as a time-series response relationship based on temperature sensor data, obtained by calculating the time lag correlation between the temperature rise change of one device and the temperature rise change of another device. When the airflow influence coefficient exceeds a preset threshold, the corresponding device is included in the extended range, ultimately forming a disturbed closed domain.
[0030] Based on the disturbed closed domain, a permissible disturbance boundary is constructed. Specifically, multi-dimensional operational data of each device within the disturbed closed domain under historical normal operating conditions are extracted, including current, temperature, and power data. Statistical characteristic values of each parameter within a sliding time window are calculated, including the mean, standard deviation, and rate of change. These statistical characteristics are then weighted and corrected based on historical disturbance samples corresponding to the operation type in the maintenance intention fingerprint set. The permissible disturbance boundary is defined as the upper limit range of variation for each parameter at a given confidence level. It is calculated by adding or subtracting a certain number of standard deviations centered on the historical mean and then adding an operational disturbance compensation amount. This operational disturbance compensation amount is obtained through cluster analysis of historical data from similar maintenance tasks. Through this method, differentiated disturbance tolerances are generated for different devices within the disturbed closed domain, thus forming the permissible disturbance boundary.
[0031] Counterfactual security twin modeling module: Select a group of reference devices G outside the disturbed closed domain C that have the same load characteristics as C and have not participated in this operation and maintenance, and combine C and G to establish a counterfactual security twin model W to characterize normal operation and maintenance disturbances.
[0032] In some embodiments of the present invention, firstly, a set of load characteristic parameters is extracted based on the historical operating data of each device within the disturbed closed domain. Specifically, historical operating data for at least seven consecutive days is selected, and a time series is constructed with a sampling interval of five minutes. Subsequently, the time series is segmented using a sliding time window of one hour, with an overlap ratio of 50% between adjacent windows. Within each sliding time window, a fast Fourier transform is first performed on the power sequence to extract the three main frequency components with the largest amplitudes as periodic load fluctuation features. Secondly, extreme value detection is performed on the power peaks, and the number of peaks exceeding the mean plus two standard deviations within each window is counted, and a peak occurrence frequency distribution is constructed as the peak power distribution feature. Thirdly, the power difference between adjacent sampling points is calculated using the first-order difference, and the average of their absolute values is taken as the load change rate feature. Finally, the above features are subjected to min-max normalization processing according to a unified dimension, mapped to the range of zero to one, and concatenated in a fixed order to form a load feature vector.
[0033] Secondly, reference devices are selected based on the load feature vector. Specifically, for each device within the disturbed closed domain, all candidate devices outside the disturbed closed domain are traversed, and the Euclidean distance between their load feature vectors is calculated. This Euclidean distance is calculated by summing the squares of the differences in each feature dimension and taking the square root. Subsequently, the Euclidean distances between devices of the same type are statistically analyzed, a distance distribution histogram is constructed, and the 15th percentile of this distribution is selected as the similarity threshold. Only when the Euclidean distance between a candidate device and the target device is less than this similarity threshold, and there are no login records, command execution records, or configuration change records in its operation log within the target maintenance time window, is the candidate device included in the reference device group. This ensures that the selected reference devices are highly similar in load behavior and have not been subject to human interference.
[0034] Then, deterministic mapping pairing is performed between the devices within the disturbed closed domain and the reference device group, and a disturbance response function is constructed. Specifically, the devices are first grouped by type. Within each group, a one-to-one matching is performed based on the principle of minimum Euclidean distance, ensuring that each reference device is used only once. After pairing, the operating data of each pair of devices within the same historical time period is extracted and time-axis aligned. Subsequently, an input matrix and output vector are constructed. The input matrix consists of the current, power, and temperature sequences of the reference device, and the output vector is the parameter sequence of the corresponding disturbed device. Based on this data, a multiple linear regression method with a regularization term is used to solve for the regression coefficients. The regularization term uses a squared penalty term, and its coefficient is selected through cross-validation to avoid overfitting, thus obtaining the disturbance response function. This function characterizes the stable mapping relationship between the two types of devices under conditions without maintenance intervention.
[0035] Finally, a counterfactual security twin model is constructed based on the disturbance response function. Specifically, during the target maintenance period, operational data of the reference device group is collected in real time, and this data is input into the corresponding disturbance response function. The predicted operational state of each device within the disturbed closed domain under conditions without maintenance operations is calculated time-by-time. The predicted results of all devices are then concatenated according to a time series to form a complete set of counterfactual operational state sequences. Furthermore, to ensure the continuity of the prediction results, the predicted values at adjacent time points are smoothed using a weighted moving average method, where the weight of the current time point is higher than that of historical time points. Ultimately, the mapping relationship, disturbance response function parameters, and counterfactual operational state sequences together constitute the counterfactual security twin model, used to characterize the baseline operational state of the data center under conditions without maintenance intervention.
[0036] Disturbance decomposition module: Collects multi-source real-time data sequence A during the target operation and maintenance period, including current sequence, temperature rise sequence, access control sequence, operation log sequence and network session sequence, and maps A to W to obtain the interpretable disturbance component M triggered by operation and maintenance intention.
[0037] In some embodiments of the present invention, firstly, time alignment processing is performed on the current sequence, temperature rise sequence, access control sequence, operation log sequence, and network session sequence. Specifically, a unified time axis is constructed, with a time granularity set to 1 minute, and a continuous time scale is established with the start time of the target maintenance period as the zero point. For the current sequence and temperature rise sequence, if the sampling frequency is higher than 1 minute, the arithmetic mean within each 1-minute interval is taken as the data at that time point. For the access control sequence and operation log sequence, an event expansion method is used to map discrete events to state values within a time interval, where an event is recorded as 1 and no event as 0. For the network session sequence, the number of sessions and data packets within each 1-minute interval are counted and normalized. For missing data points, a piecewise linear interpolation method is used to fill in the missing data, with the interpolation function determined by the values of two adjacent known time points. After completion, the minimum and maximum values are calculated for all sequences, and normalization is performed according to the following formula: Normalized value = (current value - minimum value) / (maximum value - minimum value), thereby obtaining a multi-source synchronized data matrix of a unified scale.
[0038] The multi-source synchronous data matrix is input into the counterfactual security twin model to obtain the counterfactual prediction sequence. Specifically, for each time point t, the current, temperature, and power data corresponding to the reference device are extracted from the multi-source synchronous data matrix and used as input vectors. These vectors are substituted into the trained disturbance response function to calculate the predicted value of each device within the disturbed closed domain at time point t. This calculation is performed for all time points to form a complete counterfactual prediction sequence. Subsequently, the actually collected multi-source synchronous data matrix and the counterfactual prediction sequence are aligned element-wise to calculate the difference sequence. The difference is calculated point-by-point according to the method of "actual value minus predicted value," forming an initial disturbance difference sequence consistent with the time axis.
[0039] The initial perturbation difference sequence is decomposed based on operational intent. Specifically, an operational perturbation pattern library is first constructed by extracting difference sequence samples from historical similar operational tasks and training them using a K-means clustering method. The number of clusters is determined using the elbow method, with the cluster number corresponding to the point with the minimum rate of change of the sum of squared errors within each cluster serving as the final number of categories. Each cluster center represents a typical perturbation pattern. Subsequently, the current initial perturbation difference sequence is matched with each typical perturbation pattern. Similarity calculation uses dynamic time warping distance, specifically by constructing time alignment paths to minimize the cumulative distance between two sequences. When the minimum distance is less than a preset matching threshold, the perturbation in that time period is determined to be a perturbation triggered by operational intent, and that portion of the sequence is extracted. Parts that do not meet the matching conditions are retained as unexplained perturbations, thus completing the perturbation component decomposition and obtaining candidate perturbation components.
[0040] Candidate disturbance components are subject to permissible disturbance boundary constraint verification. Specifically, for each device within the disturbed closed domain, the mean and standard deviation are calculated based on its historical normal operation data. An upper limit for the disturbance is set as the mean plus twice the standard deviation, and a lower limit is set as the mean minus twice the standard deviation. Simultaneously, an operation compensation amount is calculated based on the historical disturbance amplitude corresponding to the operation type in the maintenance intention fingerprint. This compensation amount is the average of the disturbance peak values in historical operations of the same type. The compensation amount is then superimposed on the aforementioned upper and lower limits to form the final permissible disturbance boundary. Subsequently, the candidate disturbance component is compared with the permissible disturbance boundary of the corresponding device at each time point. When all data points fall within the boundary range, the disturbance component is retained as an interpretable disturbance component. When some data points exceed the boundary, the excess parts are discarded or marked, thus obtaining the final interpretable disturbance component M.
[0041] Propagation tracing module: Based on A, M and the permissible disturbance boundary B, the residual abnormal sequence D that exceeds the scope of the operation and maintenance intention is separated, and the reverse propagation tracing of D is performed based on the time delay coupling relationship between devices to determine the initial source point P of the abnormality and the abnormal propagation chain L.
[0042] In some embodiments of the present invention, firstly, the multi-source real-time data sequence is aligned with the interpretable disturbance component on a time-by-time basis, and residual abnormal sequences are separated. Specifically, under a unified time axis, for each device at each time point t, the difference between the multi-source real-time data sequence and the interpretable disturbance component is calculated. The difference is calculated point-by-point according to the method of "actual observation value minus interpretable disturbance value". Subsequently, the difference is compared with the allowable disturbance boundary of the corresponding device, wherein the allowable disturbance boundary is composed of an upper boundary and a lower boundary. When the difference is greater than the upper boundary or less than the lower boundary, the time point is determined to be an anomaly. Further, all anomalies are connected in chronological order, and consecutive anomalies are merged to form residual abnormal sequences, wherein each residual abnormal sequence includes a start time, an end time, and an abnormal amplitude.
[0043] The time delay coupling relationship between devices is constructed based on residual abnormal sequences. Specifically, for any two devices within the disturbed closed domain, the set of abnormal occurrence times in their residual abnormal sequences is extracted, and the abnormal triggering delay between the two devices is determined using a time difference calculation method. Specifically, for device i and device j, each abnormal start time of device i is traversed, and the time point with the smallest positive time difference in the abnormal time set of device j is found. This time difference is used as a first-order propagation delay sample; the average of all samples is calculated to obtain the average time delay value from device i to device j. Subsequently, the average time delay values between all pairs of devices are used to construct a time delay coupling relationship matrix, where the matrix elements represent the propagation delay strength between devices. When the average time delay is less than a preset time delay threshold, an effective coupling relationship is considered to exist between the two devices. This time delay threshold is determined by statistically analyzing the time delay distribution of all device pairs and taking the 20th percentile.
[0044] The initial source of anomalies is located based on the time-delay coupling matrix and the device topology connection relationships. Specifically, the device topology relationships and the time-delay coupling matrix are superimposed with constraints, retaining only device pairs that have a connection in the topology and meet the time-delay threshold condition. Based on this, for each device, its "input delay sum" (the sum of all delay values pointing to that device) is calculated; simultaneously, its "output delay sum" (the sum of delay values pointing from that device to other devices) is calculated. Then, a source point determination index is constructed, defined as the input delay sum minus the output delay sum. When this index is at its minimum and the corresponding device has the earliest anomaly occurrence time, that device is determined as the initial source of the anomaly.
[0045] An anomaly propagation chain is constructed starting from the initial source of the anomaly. Specifically, starting from the initial source, the system expands layer by layer along the delay order in the delay coupling matrix, prioritizing devices with effective coupling relationships. In each expansion layer, only devices whose anomaly occurrence time is later than the current node and which satisfy the minimum delay difference are selected as the next propagation node. This process is repeated recursively, layer by layer, until all participating device nodes are covered, thus forming a sequence of anomaly propagation paths arranged in chronological order. This sequence of paths constitutes the anomaly propagation chain, used to characterize the propagation process of the anomaly within the data center.
[0046] Security Vulnerability Location Determination Module: Analyzes the deviation entropy of operation and maintenance intentions and the cross-domain coupling hysteresis gain coefficient of each node in P, L and the disturbed closed domain C to determine the security vulnerability location Q corresponding to the target operation and maintenance task, and outputs the blocking, rollback or isolation handling instructions corresponding to Q.
[0047] In one embodiment of the present invention, the deviation entropy of the operation and maintenance intention corresponding to each node within the disturbed closed domain is first calculated. Specifically, a standard operation and maintenance path state transition sequence is constructed based on the entry path information, operation step sequence, operation terminal records, and list of operated devices from the operation and maintenance intention fingerprint set. This state transition sequence does not merely record spatial paths; rather, it combines "spatial location, operation action, and target object" into a unified state unit. Specifically, access control nodes, server room areas, rack rows, and rack numbers are used as spatial location sub-states; login, verification, switching, writing parameters, issuing commands, and exiting the session are used as operation action sub-states; and the identifier of the operated device is used as the target object sub-state. These three are combined into a composite state in a fixed order. Based on the standard operation steps of the work order and the planned entry path, the composite states are arranged in chronological order to obtain the standard operation and maintenance path state transition sequence. Subsequently, the number of transitions between any two adjacent composite states is counted to construct a standard path transition count matrix. Then, the standard path transition probability matrix is calculated using Laplace smoothing. The formula is: the transition probability from a preceding state to a subsequent state equals the number of transitions plus 1, divided by the sum of all outgoing transitions from the preceding state plus the total number of states. This approach aims to avoid situations where certain paths do not occur, resulting in a transition probability of 0, which could affect subsequent entropy calculations.
[0048] After obtaining the standard path transition probability matrix, the actual operation and maintenance behavior path sequence is further extracted based on the access control sequence, operation log sequence, and network session sequence from the multi-source real-time data sequence. Specifically, using a 1-minute time granularity, access control events are mapped to the actual spatial location state, command categories in the operation logs are mapped to operation action states, and access target devices in network session records are mapped to target object states. Then, following the same composite state encoding rules as the standard operation and maintenance path state transition sequence, the spatial location state, operation action state, and target object state within the same time granularity are combined into a composite state of actual operation and maintenance behavior. For composite states that remain unchanged across multiple consecutive time granularities, only the first occurrence is retained to eliminate duplicate counting bias caused by long-term dwell time. All composite states of actual operation and maintenance behavior are then arranged in chronological order to form the actual operation and maintenance behavior path sequence, and the actual path transition count matrix and actual path transition probability matrix are constructed using the same statistical methods as described above. To ensure direct comparison between the standard path transition probability matrix and the actual path transition probability matrix, both use the exact same set of states and state order.
[0049] After obtaining the standard path transition probability matrix and the actual path transition probability matrix, a difference mapping is performed between the two, and a weighted path probability distribution is further constructed. Specifically, for each identical preceding and subsequent state, the absolute difference between the actual path transition probability and the standard path transition probability is calculated, and this absolute difference is defined as the probability deviation of that state transition pair. To increase the weight of critical state transition pairs in the operational intent deviation analysis, a path weight is also set for each state transition pair. The path weight consists of a step position weight and a device sensitivity weight. The step position weight is calculated according to the position of the state transition pair in the standard operational path state transition sequence; the closer the state transition pair is to the core operation area, the greater the step position weight. The device sensitivity weight is set according to the influence level of the target device in the data center; the device sensitivity weight of power supply nodes, cooling control nodes, and backbone network nodes is higher than that of ordinary computing nodes. The probability deviation is multiplied by the corresponding path weight to obtain the weighted deviation value. Subsequently, the sum of the weighted deviation values of all state transition pairs is used as the normalized denominator, and the weighted deviation value of each state transition pair is divided by this normalized denominator to form the weighted path probability distribution. In this way, the deviation between actual behavior and expected behavior is reflected not only in the difference in probability, but also in the intensity of the deviation in key steps and key equipment.
[0050] After obtaining the weighted path probability distribution, the path information entropy is calculated and used as the deviation entropy of the operational intent. Specifically, all state transition pairs are traversed, and the weighted probability value of each pair is multiplied by the negative of its natural logarithm. The sum of all results is then obtained to obtain the path information entropy. If the weighted probability value of a state transition pair is 0, that term is treated as 0. The larger the calculated path information entropy, the stronger the dispersion of the actual operational behavior in state transitions and the higher the degree of deviation from the standard operational path state transition sequence; the smaller the calculated path information entropy, the closer the actual operational behavior is to the expected path described by the operational intent fingerprint set. To make the deviation entropy of operational intent comparable between different operational tasks, the path information entropy is also normalized. Specifically, the path information entropy of the current task is subtracted from the minimum value of the path information entropy of historical tasks of the same type, then divided by the difference between the maximum and minimum values plus 0.000001, thus obtaining the normalized deviation entropy of operational intent in the range of 0 to 1.
[0051] In one embodiment of the present invention, after calculating the deviation entropy of the operation and maintenance intention, the cross-domain coupling hysteresis gain coefficient of each node within the disturbed closed domain is further calculated. Specifically, time-series data of each node in different physical domains are first extracted from the disturbed closed domain. These different physical domains include at least a power domain, a thermal domain, and a network domain. The power domain time-series data consists of current, power, or power supply branch load; the thermal domain time-series data consists of temperature rise, return air temperature, or internal cabinet temperature difference; and the network domain time-series data consists of network session counts, port traffic, or control message counts. To ensure that time series from different physical domains can be analyzed in a unified frequency domain, all time series are first uniformly resampled at 1-minute sampling intervals. High-frequency sampling sequences are downsampled using interval averaging, and low-frequency sampling sequences are filled with missing points using linear interpolation. Subsequently, each time series is detrended. Specifically, a linear trend term is fitted to the original time series using the least squares method, and then the linear trend term is subtracted from the original time series to obtain the detrended time series, thereby reducing the impact of long-term slow drift on the frequency domain analysis results.
[0052] After completing unified sampling and detrending processing, frequency domain transformation is performed on the time series of different physical domains. Specifically, each segment consists of 64 sampling points with a 50% overlap. A Hanning window is applied to each segment, followed by a Fast Fourier Transform (FFT) to obtain the complex spectrum for each frequency component. Based on the complex spectrum, the power spectral density (PSD) of each physical domain's time series is calculated. The PSD is defined as the square of the magnitude of the complex spectrum of that frequency component after length normalization. Furthermore, the cross-power spectral density (CPSD) between any two physical domains is calculated. The CPSD is defined as the length normalized product of the complex spectrum of one physical domain and the conjugate of the complex spectrum of the other physical domain. Subsequently, based on the CPSD and the power spectral densities of the two physical domains, the frequency domain coherence function is calculated. The frequency domain coherence function is defined as the ratio of the square of the magnitude of the CPSD to the product of the power spectral densities of the two physical domains, ranging from 0 to 1. The closer the value is to 1, the stronger the coupling between the two physical domains at that frequency component.
[0053] After obtaining the frequency domain coherence function, the dominant frequency component is selected, and the time lag is calculated. Specifically, the frequency point with the largest frequency domain coherence function among all frequency components is first found, and the frequency corresponding to this point is defined as the dominant frequency. To avoid weak coupling noise being misjudged as effective coupling, a coherence strength threshold of 0.6 is set. When the frequency domain coherence function corresponding to the dominant frequency is less than 0.6, it is determined that there is no effective cross-domain coupling between the two physical domains within the current time window, and its cross-domain coupling lag gain coefficient is directly recorded as 0. When the frequency domain coherence function corresponding to the dominant frequency is not less than 0.6, the phase angle of the cross-power spectral density at the dominant frequency is further calculated, and the phase angle is adjusted to a continuous interval using phase expansion. Then, the phase angle is divided by the product of twice pi and the dominant frequency to obtain the time lag. The time lag represents the time delay experienced by a disturbance in one physical domain propagating to another physical domain. The smaller the absolute value of the time lag, the more direct the cross-domain propagation.
[0054] After obtaining the time lag, the amplitude transfer ratio is further calculated, and a cross-domain coupling lag gain coefficient is constructed. Specifically, the amplitude transfer ratio is obtained by dividing the spectral amplitude of the target physical domain at the dominant frequency by the spectral amplitude of the upstream physical domain. To ensure that the cross-domain coupling lag gain coefficient simultaneously reflects coupling strength, amplitude amplification effect, and propagation delay, the following construction method is adopted: multiply the frequency domain coherence function corresponding to the dominant frequency by the amplitude transfer ratio, and then multiply by a time decay term. The time decay term uses an exponential decay function, expressed as a negative exponent of the natural constant, where the exponent is obtained by dividing the absolute value of the time lag by the reference lag time. The reference lag time is taken as the median of the forward propagation time lag between similar devices over the past 30 days. According to this construction method, if the coherence strength between two physical domains is high, the amplitude transfer ratio is large, and the time lag is small, the cross-domain coupling lag gain coefficient is large; if the coherence strength is low, the amplitude transfer ratio is low, or the time lag is large, the cross-domain coupling lag gain coefficient is small. For cases where the same node involves multiple physical domain pairs, the cross-domain coupling lag gain coefficients of each physical domain pair directly related to the node are summed up by arithmetic average to obtain the node-level cross-domain coupling lag gain coefficient of the node.
[0055] In one embodiment of the present invention, after obtaining the operational intention deviation entropy and cross-domain coupling hysteresis gain coefficient of each node, a node propagation sequence is constructed based on the initial source point of the anomaly, the anomaly propagation chain, and the propagation position relationships of each node within the disturbed closed domain. Specifically, the initial source point of the anomaly in the aforementioned residual anomaly reverse propagation tracing results is first read and used as the first node in the propagation sequence. Then, subsequent nodes are read sequentially along the time order of the anomaly propagation chain, and the time of the first occurrence of the anomaly, the cumulative propagation time between the node and the initial source point, and the hierarchical position in the anomaly propagation chain are recorded for each node. If a node appears in multiple branch propagation paths simultaneously, the earliest time of the first occurrence of the anomaly is used as the propagation time of that node, and the minimum hierarchical position of that node from the initial source point is used as the propagation level of that node. After completing the above processing, a node propagation sequence containing "node identifier, propagation time, propagation level, and upstream node set" is obtained. This node propagation sequence is used to characterize the actual propagation order of the anomaly within the disturbed closed domain.
[0056] After forming the node propagation sequence, the node propagation sequence is combined with the operational intention deviation entropy and cross-domain coupling hysteresis gain coefficient of each node to form a multi-dimensional feature set of the nodes. To ensure the comparability of the two types of features, the operational intention deviation entropy of all nodes within the disturbed closed domain is first normalized, and then the cross-domain coupling hysteresis gain coefficient of all nodes is normalized. The normalization method uniformly adopts linear normalization from minimum to maximum value, specifically: the normalized feature value of a node is equal to the original feature value of the node minus the minimum value of the same type of feature within the disturbed closed domain, then divided by the difference between the maximum and minimum values plus 0.000001. Then, the fusion weight is determined based on the statistical results of the impact of the two types of features on the fault in historical anomaly samples. In specific implementation, historical anomaly samples occurring in the same type of data center in the last 180 days are selected, and a fault impact score is calculated for each historical anomaly sample. The fault impact score is composed of recovery time, the number of affected nodes, and the magnitude of the abnormal peak. The recovery time is weighted at 0.5, the number of affected nodes at 0.3, and the magnitude of the abnormal peak at 0.2. These three factors are normalized and then weighted and summed to obtain the fault impact score. Subsequently, the Pearson correlation coefficients between the operational intention deviation entropy and the fault impact score in historical samples, as well as the cross-domain coupling lag gain coefficient and the fault impact score, are calculated. The absolute values of these two coefficients are then normalized to obtain the operational intention deviation entropy weight and the cross-domain coupling lag gain coefficient weight. The weights determined in this way are not arbitrarily assigned but driven by the statistical relationships of historical abnormal samples, reflecting the true contribution of these two types of features to the abnormal consequences.
[0057] After obtaining the normalized features and fusion weights, the node anomaly contribution index for each node is calculated. Specifically, the normalized operational intention deviation entropy of a node is multiplied by its operational intention deviation entropy weight, and the normalized cross-domain coupling lag gain coefficient of that node is multiplied by its cross-domain coupling lag gain coefficient weight. The sum of these two values yields the node anomaly contribution index. A larger node anomaly contribution index indicates that the node possesses both stronger operational intention deviation characteristics and a higher cross-domain propagation amplification capability, resulting in a more significant driving effect on subsequent anomaly propagation.
[0058] After obtaining the node anomaly contribution index for each node, the cumulative node risk value is further calculated based on the node's positional relationship in the propagation sequence. Specifically, for each node in the propagation sequence, its own node anomaly contribution index is first used as the base risk value. Then, all upstream nodes in the anomaly propagation chain are traversed, and the time difference between the first occurrence time of an anomaly at an upstream node and the first occurrence time of an anomaly at the current node is calculated. Next, the node anomaly contribution index of the upstream node is multiplied by a time delay decay function, and the decay results of all upstream nodes are summed. This sum is then added to the base risk value of the current node to obtain the cumulative node risk value. The time delay decay function adopts an exponential decay form, expressed as a negative exponential form of the natural constant, where the exponent is obtained by dividing the time difference from the upstream node to the current node by the propagation reference time. The propagation reference time is the average of the propagation time differences of all adjacent nodes in the current anomaly propagation chain. When the propagation time difference between the upstream node and the current node is small, its risk contribution to the current node is attenuated less; when the propagation time difference is large, its risk contribution to the current node is attenuated more. In this way, the node risk accumulation value not only takes into account the node's own characteristics, but also the risk impact superimposed when the anomaly propagates from upstream to the node, thus better conforming to the risk accumulation law in the real propagation process.
[0059] After calculating the cumulative node risk value for all nodes, the vulnerable locations are determined. Specifically, the cumulative node risk values of all nodes within the disturbed closed domain are sorted, and the node with the highest cumulative node risk value is identified as the vulnerable location. If two or more nodes have the same cumulative node risk value, the node with the earlier anomaly occurrence time is prioritized; if the anomaly occurrence time is also the same, the node with the larger cross-domain coupling hysteresis gain coefficient is prioritized. This determination method ensures the uniqueness and repeatability of the vulnerable locations. The identified vulnerable locations are not simply the initial source of the anomaly, nor are they simply the propagation endpoints. Rather, they are the determination results made based on a comprehensive consideration of the degree of deviation from operational behavior, cross-domain propagation capability, and risk accumulation effect, identifying the critical locations most likely to trigger anomaly propagation within the entire target operational task.
[0060] After identifying the security vulnerability, the system further outputs blocking, rollback, or isolation instructions corresponding to the vulnerability. Specifically, the cumulative risk value of the node at the vulnerability is first normalized again to obtain a normalized risk value within the range of 0 to 1. Then, based on the distribution of the cumulative risk value of nodes in the most recent 180 days of historical anomaly samples, the 75th percentile is taken as the primary risk threshold, and the 90th percentile as the secondary risk threshold. When the normalized risk value is not less than the secondary risk threshold, and the node corresponding to the security vulnerability is an account entry point, remote login terminal, bastion access node, or network session entry point node, a blocking instruction is output. The blocking instruction includes terminating the current session, freezing the corresponding account, closing the corresponding network port, and rejecting subsequent same-origin connection requests. The blocking instruction is applicable when the operational intent deviates significantly from the entropy and the anomaly's starting position is close to the access entry point, to prevent the anomaly from continuing to inject into the disturbed closed domain.
[0061] When the normalized risk value is not less than the first-level risk threshold, and the node corresponding to the security vulnerability is a parameter configuration node, control command issuance node, or policy switching node, and the deviation entropy of the node's operation and maintenance intention is greater than its cross-domain coupling hysteresis gain coefficient, a rollback handling instruction is output. The rollback handling instruction includes restoring the configuration snapshot saved 5 minutes before the operation and maintenance, canceling parameters added during the current work order period, and restoring to the control table entry of the previous stable version. The configuration snapshot is obtained by performing a full backup of the target node's parameter file, control register values, and routing policy table before each high-privilege operation and maintenance task begins. The rollback handling instruction is applicable to situations where the anomaly is mainly triggered by deviation from the operation path, rather than having formed a strong cross-domain spread, in order to restore the node state to the pre-operation and maintenance baseline as quickly as possible.
[0062] When the normalized risk value is not less than the secondary risk threshold, and the node corresponding to the security vulnerability is a power supply node, cooling linkage node, column-level control node, or a critical node that has shown significant cross-domain propagation, and the cross-domain coupling hysteresis gain coefficient of the node is greater than its operational intention deviation entropy, an isolation and handling instruction is output. The isolation and handling instruction includes cutting off the unnecessary control links corresponding to the node, closing the linkage control interface connected to it, removing the cabinet column where it is located from the automatic linkage strategy, and migrating the relevant load to a backup node or backup branch. The isolation and handling instruction is applicable to situations where the anomaly has already generated a significant propagation amplification effect between the power domain, thermal domain, and network domain, to prevent cross-domain coupling from causing the anomaly to expand from a local node into a regional security event.
[0063] Through the above processing, this invention performs a unified analysis of the deviation entropy of operation and maintenance intention, the cross-domain coupling lag gain coefficient, the initial source of the anomaly, the anomaly propagation chain, and the propagation time sequence of nodes within the disturbed closed domain. This not only accurately identifies the security vulnerability corresponding to the target operation and maintenance task, but also automatically provides targeted handling instructions based on the risk source and propagation characteristics of the security vulnerability. This avoids the problems of false blocking, false rollback, or delayed handling caused by existing methods that only rely on a single alarm threshold or a single anomaly score.
[0064] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. An AI-based data center operation and maintenance security system, characterized in that: include: Operation and maintenance intent fingerprint construction module: Obtain the work order information, operation and maintenance personnel identity information, entry path information, operation terminal records and operation device list of the target operation and maintenance task, and construct the operation and maintenance intent fingerprint set F corresponding to the target operation and maintenance task; Disturbed Closed Domain Determination Module: Based on the equipment topology, power connection, and airflow organization of F and the target data center, determine the disturbed closed domain C related to the target operation and maintenance task and the corresponding permissible disturbance boundary B; Counterfactual security twin modeling module: Select a group of reference devices G outside the disturbed closed domain C that have the same load characteristics as C and have not participated in this operation and maintenance, and combine C and G to establish a counterfactual security twin model W to characterize normal operation and maintenance disturbances; Disturbance decomposition module: Collects multi-source real-time data sequence A within the target operation and maintenance period, including current sequence, temperature rise sequence, access control sequence, operation log sequence and network session sequence, and maps A to W to obtain the interpretable disturbance component M triggered by operation and maintenance intention; Propagation tracing module: Based on A, M and the permissible disturbance boundary B, the residual abnormal sequence D that exceeds the scope of the operation and maintenance intention is separated, and the reverse propagation tracing of D is performed based on the time delay coupling relationship between devices to determine the initial source point P of the abnormality and the abnormal propagation chain L; Security Vulnerability Location Determination Module: Analyzes the deviation entropy of operation and maintenance intentions and the cross-domain coupling hysteresis gain coefficient of each node in P, L and the disturbed closed domain C to determine the security vulnerability location Q corresponding to the target operation and maintenance task, and outputs the blocking, rollback or isolation handling instructions corresponding to Q.
2. The AI-based data center operation and maintenance security system according to claim 1, characterized in that: The determination of the disturbed closed domain C and the corresponding permissible disturbance boundary B related to the target operation and maintenance task includes: Map the operated device objects in the operation and maintenance intention fingerprint set to the device topology relationship, extract the set of devices that have a direct connection or second-order adjacency relationship with the operated device, and form the initial topology association domain; Based on the initial topology association domain, and combined with the power connection relationship, the device nodes that share a power supply branch or have a current coupling relationship with the operated device are identified, and the initial topology association domain is expanded to obtain the power coupling extended domain. Based on the aforementioned power coupling extended domain, the set of devices with cold and hot channel associations or airflow return paths is determined according to the airflow organization relationship, and the power coupling extended domain is further extended to form a disturbed closed domain. Based on the historical operational fluctuation range of each device within the disturbed closed domain and the disturbance characteristics of the corresponding operation type in the operation and maintenance intention fingerprint set, differentiated disturbance thresholds are set for each device to construct an allowable disturbance boundary.
3. The AI-based data center operation and maintenance security system according to claim 1, characterized in that: The aforementioned combination of C and G to establish a counterfactual security twin model W for characterizing normal operational disturbances includes: Based on the historical operating data of each device within the disturbed closed domain, a set of load characteristic parameters is extracted. The set of load characteristic parameters includes at least periodic load fluctuation characteristics, peak power distribution characteristics, and load change rate characteristics. A load characteristic vector is then constructed based on the set of load characteristic parameters. Based on the load feature vector, a set of devices with an Euclidean distance less than a preset similarity threshold and no operation record within the target maintenance time window is selected outside the disturbed closed domain to determine the reference device group; The devices within the disturbed closed domain are mapped and paired one by one with the reference device group according to the device type and connection relationship. A disturbance response function is established based on historical synchronous operation data. The disturbance response function is used to characterize the natural change relationship between devices under the condition of no operation and maintenance intervention. Based on the disturbance response function and the mapping pairing relationship, a counterfactual security twin model is constructed to generate operational status prediction results under the condition of no operation and maintenance intervention.
4. The AI-based data center operation and maintenance security system according to claim 1, characterized in that: The obtained interpretable perturbation component M includes: Time alignment processing is performed on current sequence, temperature rise sequence, access control sequence, operation log sequence and network session sequence to construct a multi-source synchronous data matrix under a unified time axis, and interpolation method is used to fill in missing data to obtain a standardized multi-source data sequence; The standardized multi-source data sequence is input into the counterfactual secure twin model to obtain the counterfactual prediction sequence at the corresponding time point, and the difference between the actual data sequence and the counterfactual prediction sequence is calculated to form the initial perturbation difference sequence. Based on the perturbation pattern corresponding to the operation type in the operation and maintenance intent fingerprint, the initial perturbation difference sequence is subjected to pattern matching and component decomposition to extract the perturbation component consistent with the operation and maintenance operation, and candidate perturbation components are obtained. The candidate disturbance components are constrained and verified to meet the allowable disturbance boundary range, and an interpretable disturbance component that conforms to the operation and maintenance intention is output.
5. The AI-based data center operation and maintenance security system according to claim 1, characterized in that: The residual abnormal sequence D, which exceeds the scope of the operational intention, is obtained by separating it based on A, M, and the permissible disturbance boundary B, including: The multi-source real-time data sequence is aligned with the interpretable perturbation component on a time-by-time basis. The difference between the two is calculated and compared with the allowable perturbation boundary. The difference data exceeding the allowable perturbation boundary range is extracted to form a residual abnormal sequence.
6. The AI-based data center operation and maintenance security system according to claim 5, characterized in that: Determining the initial source point P of the anomaly and the anomaly propagation chain L includes: constructing an anomaly triggering time matrix between devices based on the anomaly occurrence time of each device in the residual anomaly sequence, and calculating the time delay difference between any two devices to obtain the time delay coupling relationship matrix between devices; according to the time delay coupling relationship matrix and the device topology connection relationship, using the minimum time delay path search method to backtrack the residual anomaly sequence to determine the device node where the anomaly signal first appears, as the initial source point of the anomaly; starting from the initial source point of the anomaly, and combining the propagation order relationship between each node in the time delay coupling relationship matrix, constructing an anomaly propagation path sequence, thereby determining the anomaly propagation chain.
7. The AI-based data center operation and maintenance security system according to claim 1, characterized in that: The method for calculating the deviation entropy of the operational intention includes: Based on the entry path information and operation step sequence in the operation and maintenance intent fingerprint set, a standard operation and maintenance path state transition sequence is constructed, and the transition probability between each adjacent state is calculated to form a standard path transition probability matrix. Based on the access control sequence, operation log sequence, and network session sequence in the multi-source real-time data sequence, the actual operation and maintenance behavior path sequence is extracted, and an actual path transition probability matrix is constructed according to the same state division rule. The actual path transition probability matrix and the standard path transition probability matrix are subjected to difference mapping, the corresponding state transition probability deviation is calculated, and a weighted path probability distribution is constructed based on the deviation. The path information entropy is calculated based on the weighted path probability distribution and used as the operation and maintenance intent deviation entropy.
8. The AI-based data center operation and maintenance security system according to claim 1, characterized in that: The calculation of the cross-domain coupling hysteresis gain coefficient includes: Time series data corresponding to different physical domains within the disturbed closed domain are acquired separately, and the time series data are uniformly sampled and detrended. Each time series is converted into its corresponding frequency domain representation using Fast Fourier Transform to obtain the power spectral density at each frequency component. The cross-power spectral density between any two physical domains is calculated based on the frequency domain representation, and a frequency domain coherence function is constructed based on the cross-power spectral density and their respective power spectral densities. The dominant frequency component with the largest coherence value is selected from the frequency domain coherence function, and the phase difference at the corresponding frequency is calculated. The phase difference is converted into a time lag. Based on the time lag and the coherence strength of the corresponding frequency component, the amplitude transfer ratio of the disturbance between different physical domains is calculated, and the cross-domain coupling lag gain coefficient is constructed.
9. The AI-based data center operation and maintenance security system according to claim 8, characterized in that: The determination of the security vulnerability location Q includes: Based on the initial source of the anomaly and the anomaly propagation chain, the temporal position of each node in the propagation path within the disturbed closed domain is extracted, and a node propagation sequence is constructed. At the same time, the deviation entropy of the operation and maintenance intention and the cross-domain coupling hysteresis gain coefficient of each node are obtained to form a multi-dimensional feature set. The multidimensional feature set is normalized, and the deviation entropy of operation and maintenance intention and the cross-domain coupling lag gain coefficient are weighted and fused based on preset weights to construct a node anomaly contribution index, wherein the weights are determined by the statistical results of the impact of the two types of features on the fault in historical anomaly samples. Based on the node's abnormal contribution index and the node's position in the propagation sequence, the node's cumulative risk value is calculated. The cumulative risk value is obtained by superimposing the node's own abnormal contribution value and the contribution value of its upstream node according to the time delay decay function. The node with the highest cumulative risk value is identified as the vulnerable location Q.