A running fault prediction method based on computing power server health management

By performing dual-stream data sequence calibration and matching on the operational data and thermal environment data of the computing server, unique prototype identifiers and component identifiers are generated. A resource relationship diagram is constructed and a set of maintenance windows is generated, which solves the coupling problem between operation and thermal environment in the health management of computing servers, achieves the accuracy and consistency of fault prediction, and supports the closed loop of health management.

CN121433953BActive Publication Date: 2026-07-14BEIJING AEROSPACE STAR BRIDGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing health management methods for computing servers lack site signature constraints, resulting in incomparability across devices and difficulty in characterizing the operation-thermal environment coupling problem due to single-stream triggering based solely on threshold scoring. Furthermore, predictions often remain at the level of probability and alarms, failing to form a set of constraints for handling operational failures upon expiration, leading to a disconnect between prediction and handling.

Method used

By collecting and aligning operational and thermal environment data of the computing server according to the unique resource identifier, calibrating the dual-stream data sequence and site signature, constructing the data coupling trajectory and matching it with the computing server prototype library under the site signature, generating unique prototype identifiers and component identifiers, mapping them to disposal scripts, constructing resource relationship diagrams and generating maintenance window sets, and finally converting them into a set of operational fault expiration disposal constraints.

Benefits of technology

It achieves precise location of potential failure points and certainty of overdue handling paths, improves the certainty and consistency of operational failure overdue identification, forms a health management closed loop, and ensures a unified expression of action sets, sequential parallel relationships, and time bandwidth constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121433953B_ABST
    Figure CN121433953B_ABST
Patent Text Reader

Abstract

The application discloses a kind of running fault prediction methods based on computing power server health management, it is related to data center operation and maintenance technical field, including, to computing power server is according to resource unique identification collection and aligns running data and thermal environment data, calibrates running data and thermal environment data to obtain double-flow data sequence and site signature;Based on double-flow data sequence, construct data coupling track, match under site signature with computing power server prototype library, obtain unique prototype identification, and map as component identification and disposal scenario;By being attached to resource unique prototype identification, component identification and disposal scenario, obtain resource entry, according to disposal scenario, construct resource relationship diagram to obtain strong relationship diagram, obtain resource total order arrangement and scan generation maintenance window set;Through maintenance window set, construct partial order lattice, convert maintenance window and disposal scenario into running fault due date disposal constraint set.The application is accurately positioned fault and forms disposal boundary by double-flow prototype mapping and partial order constraint set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center operation and maintenance technology, and in particular to a method for predicting operational failures based on the health management of computing servers. Background Technology

[0002] Computing infrastructure is evolving towards higher density, heterogeneity, and greener practices, with server operating status and data center thermal field exhibiting strong coupling characteristics. The industry has established a health management system based on monitoring, modeling, and operation and maintenance orchestration, covering aspects such as log and telemetry data collection, equipment and environmental monitoring, fault mode database construction, time series analysis, and predictive maintenance, playing a role in ensuring availability and improving energy efficiency.

[0003] However, existing methods still have two limitations: First, they lack a dual-flow approach to operation and thermal environment constrained by site signatures, relying mainly on single-flow indicators or threshold scoring fusion. This results in non-closed dimensions across cabinets and cooling zones, template migration distortion, and difficulty in stabilizing grayscale degradation and expiration risks. Second, predictions are mostly based on probability and alarms, failing to structure maintenance windows and handling scripts into sequential parallel relationships and time and bandwidth constraints. This makes it impossible to form a set of constraints for handling operational failures upon expiration, leading to a disconnect between prediction and handling. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a method for predicting operational failures based on the health management of computing servers to solve the problems of cross-device incomparability caused by the lack of site signature constraints and the difficulty in characterizing the operation-thermal environment coupling problem by single-stream triggering based solely on threshold scoring.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] This invention provides a method for predicting operational failures based on the health management of computing servers, comprising:

[0008] The computing server collects and aligns operational data and thermal environment data according to the unique resource identifier, and calibrates the operational data and thermal environment data to obtain a dual-stream data sequence and site signature;

[0009] A data coupling trajectory is constructed based on the dual-stream data sequence, and matched with the computing server prototype library under the site signature to obtain a unique prototype identifier, which is then mapped to a component identifier and a disposal script.

[0010] By attaching a unique prototype identifier, component identifier, and disposal script to the resource, resource entries are obtained. A strong relationship graph is obtained by constructing a resource relationship graph based on the disposal script. The total order of resources is obtained and scanned to generate a set of maintenance windows.

[0011] By constructing a partially ordered lattice through a set of maintenance windows, the maintenance windows and disposal scripts are transformed into a set of constraints for handling expired faults.

[0012] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the specific steps for collecting and aligning operational data and thermal environment data of the computing server according to unique resource identifiers are as follows:

[0013] By using hierarchical naming to connect the known physical and logical levels of computing power servers, the physical uniqueness of computing power servers and computing power server components is solidified to generate unique resource identifiers.

[0014] A data acquisition channel is established using the unique resource identifier as the primary key. The data from the computing server is time-aligned using the sampling periods for operational data and thermal environment data, respectively, to obtain operational data and thermal environment data within a unified time grid.

[0015] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the calibration operation data and thermal environment data are used to obtain a dual-stream data sequence and a site signature. The specific steps are as follows:

[0016] Normalize the operational data and thermal environment data within a unified time grid to obtain normalized operational data sequences and thermal environment data sequences;

[0017] Within each unified time grid, using the normalized operational data sequence as the parent time axis, causal time warping is performed on the normalized thermal environment data sequence. The normalized operational data and thermal environment data at the matching points of each unified time grid are statistically analyzed and merged to obtain a dual-stream data sequence.

[0018] Robust covariance is obtained by performing a block diagonal matrix on the dual-stream data sequence to obtain the site covariance.

[0019] Read the known cooling zone identifiers within the computing server and combine them with the site covariance and resource unique identifier to obtain the site signature.

[0020] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the specific steps of constructing a data coupling trajectory based on dual-stream data sequences are as follows:

[0021] Based on the station covariance of the normalized operating data sequence and the thermal environment data sequence, the whitening matrix of the normalized operating data sequence and the thermal environment data sequence is calculated.

[0022] The increment is obtained by calculating the difference increment between adjacent time steps of the normalized operating data sequence and the thermal environment data sequence. The increment is then transformed by the whitening matrix of the normalized operating data sequence and the thermal environment data sequence, and the data coupling trajectory is obtained by connecting them according to time.

[0023] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the step of matching the site signature with the computing server prototype library to obtain a unique prototype identifier specifically involves the following steps:

[0024] Bind component identifiers and handling scripts to each data prototype trajectory in the computing server prototype library;

[0025] The data coupling trajectory is aligned temporally with the data prototype trajectory within the same site signature in the computing server prototype library to obtain a unique prototype identifier.

[0026] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the mapping of component identifiers and disposal scripts includes, within the same site signature, determining the component identifier and disposal script corresponding to the unique prototype identifier of the data coupling trajectory based on the component identifiers and disposal scripts in the computing server prototype library.

[0027] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the specific steps of obtaining resource entries by attaching unique prototype identifiers, component identifiers, and disposal scripts to resources are as follows:

[0028] The unique prototype identifier, component identifier, disposal script, cooldown section identifier, and resource unique identifier are bound together to obtain the initial resource entries, and time alignment is performed based on a unified time grid.

[0029] The initial resource entry is supplemented with enhanced fields for tenant, project, host, and machine location to form the resource entry.

[0030] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the steps of constructing a resource relationship graph based on the disposal script to obtain a strong relationship graph, obtaining the total order of resources, and scanning to generate a maintenance window set are as follows:

[0031] Within the statistics window, a resource relationship graph is constructed based on the action set of resource items extracted from the handling script, and a strong relationship graph is output after resource relationship judgment.

[0032] Strong relation graphs calculate the order variables of all resource pairs using mixed-integer linear programming to obtain the total order permutation of resources;

[0033] Set maintenance window rules and obtain the maintenance window set based on the total order of resources.

[0034] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the specific steps of constructing a partial-order lattice through a maintenance window set are as follows:

[0035] Within the maintenance window set, obtain the complete set of actions for each maintenance window, and filter the order of maintenance window actions, the parallel relationship of actions, and the priority order of actions;

[0036] By summarizing the sequence of actions, the parallel relationships of actions, and the priority of actions, a partial order grid for maintaining the window is obtained.

[0037] As a preferred embodiment of the operational fault prediction method based on computing server health management described in this invention, the specific steps of converting the maintenance window and handling script into an operational fault expiration handling constraint set are as follows:

[0038] Apply time and bandwidth constraints to the entire set of actions for the maintenance window to obtain the basis of the constraint set for handling operational faults upon expiration.

[0039] The set of constraints for handling expired operational faults is obtained by intersecting the base set of constraints for handling expired operational faults with the partially ordered lattice of the maintenance window.

[0040] The beneficial effects of this invention are as follows: By constructing a data coupling trajectory based on dual-stream data sequences and matching it with the computing server prototype library under the site signature, the invention achieves a precise correspondence between the unique prototype identifier, component identifier, and disposal script, which is used to locate potential failure locations and due disposal paths, ultimately improving the certainty and consistency of operational failure due identification; by constructing a partial order lattice through a maintenance window set and converting it into an operational failure due disposal constraint set, the invention achieves a unified expression of action sets, sequential parallel relationships, and time bandwidth limitations, which is used to form a directly referenceable disposal boundary, ultimately supporting a health management closed loop. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of a method for predicting operational failures based on the health management of computing servers.

[0043] Figure 2 Generate a flowchart that is aligned with the data time to uniquely identify the resource.

[0044] Figure 3A flowchart for generating dual-stream data sequences and site signatures.

[0045] Figure 4 A flowchart for constructing a data coupling trajectory that matches a unique prototype identifier. Detailed Implementation

[0046] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0047] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0048] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0049] Reference Figures 1-4 This is one embodiment of the present invention, which provides a method for predicting operational failures based on the health management of computing servers, including the following steps:

[0050] S1. Collect and align the operating data and thermal environment data of the computing server according to the unique resource identifier, and calibrate the operating data and thermal environment data to obtain a dual-stream data sequence and site signature.

[0051] The computing power servers are connected by hierarchical naming, which is known to be at the physical and logical levels. The physical uniqueness of the computing power servers and computing power server components is solidified based on composite hardware fingerprints, and one-way hashing is used to generate unique resource identifiers for the computing power servers and computing power server components.

[0052] Hierarchical naming arranges the known physical and logical levels of the computing server in descending spatial order.

[0053] Composite hardware fingerprints generate a unique identifier for a device using the characteristic information of a computing server.

[0054] One-way hashing is an irreversible hashing algorithm that uses a normalized string to represent the physical uniqueness of a computing server and its components.

[0055] A data acquisition channel is established using the unique resource identifier as the primary key. Through the network device time synchronization protocol, the data of the computing server is time-aligned according to the sampling period of the running data and the sampling period of the thermal environment data, so as to obtain the running data and thermal environment data within a unified time grid.

[0056] The sampling periods for operational data and thermal environment data are sampled separately. The operational data sampling period targets rapidly changing quantities and captures second-level abrupt changes, such as 1 second. The thermal environment data sampling period targets slowly changing quantities with high thermal inertia and performs low-frequency sampling, such as 5 seconds.

[0057] Calculate robust statistics for the data within the sampling period. Normalize the operational data and thermal environment data within the unified time grid based on the station scale parameters within the unified time grid to obtain the normalized operational data sequence and thermal environment data sequence.

[0058] The site scale parameter is constructed by taking the median of the data in each sampling period as the center vector and the median absolute deviation as the scale vector. Based on the center vector and scale vector in the robust statistics of the data, a diagonal matrix is ​​constructed based on the dimension and the center vector is calculated. This constitutes the site scale parameter of the operational data and thermal environment data as the normalization base.

[0059] Within each unified time grid, using the normalized operational data sequence as the parent time axis, causal time warping is performed on the normalized thermal environment data sequence. The normalized operational data and thermal environment data at the matching points of each unified time grid are statistically analyzed and merged to obtain a dual-stream data sequence.

[0060] Causal time warping refers to performing forward-only time alignment on normalized thermal environment data using normalized operational data as the parent time base on a unified time grid. It only uses the current moment and historical samples of the current moment, with the starting point fixed as the first valid sample of the sequence and the ending point fixed as the current moment of the parent time base. The path is monotonous and continuous, and the path is restricted to a local window with a fixed width centered on the current moment. Only one forward step is allowed for horizontal and diagonal movement. The window width is determined based on the historical average configuration of the computing server, resulting in a dual-stream data sequence.

[0061] Furthermore, robust covariance is obtained by performing analysis on the normalized operational data sequence and thermal environment data sequence within a unified time grid, and a block diagonal matrix is ​​constructed to obtain the site covariance.

[0062] Robust covariance is achieved by filtering normalized operational data sequences and thermal environment data sequences with a fixed truncation threshold, retaining data sequences whose vectors in every dimension are below the truncation threshold to form a clean subset, and calculating the covariance of the clean subset as the robust covariance.

[0063] The truncation threshold is an integer multiple of the site scale parameter to satisfy the requirements of extremely low false deletion probability and strong resistance to outliers, such as 6 times the site scale parameter.

[0064] The block diagonal matrix is ​​constructed based on the robust covariance of the normalized operational data sequence and the thermal environment data sequence to construct the site covariance, expressed as:

[0065] ;

[0066] in, It is the current moment. It is the robust covariance of the normalized running data sequence at the current moment. It is a sequence of running data. It is the robust covariance of the normalized thermal environment data sequence at the current moment. It is a thermal environment data sequence. It is the station covariance at the current moment.

[0067] Furthermore, the known cooling zone identifiers within the computing server are read and encapsulated with the operating data site scale parameters, thermal environment data site scale parameters, site covariance, and resource unique identifier to obtain the site signature.

[0068] S2. Construct a data coupling trajectory based on the dual-stream data sequence, match it with the computing server prototype library under the site signature to obtain a unique prototype identifier, and map it to a component identifier and disposal script.

[0069] Based on the station covariance of the normalized operational data sequence and the thermal environment data sequence, the whitening matrix expression of the normalized operational data sequence and the thermal environment data sequence is calculated as follows:

[0070] ;

[0071] in, It is a whitening matrix. It is a block diagonal splicing operator. yes The inverse square root matrix of the site covariance, It is the inverse square root matrix of the site covariance of thermal environment data.

[0072] The increment is obtained by calculating the difference increment between adjacent time steps of the normalized operating data sequence and the thermal environment data sequence. The increment is then transformed using the whitening matrix of the normalized operating data sequence and the thermal environment data sequence, and the data coupling trajectory is obtained by concatenating them according to time. The expression is as follows:

[0073] ;

[0074] ;

[0075] in, It is a discrete time index in a unified time grid. It's an increment. It is the difference increment of normalized running data at adjacent time points. It is the difference increment between adjacent time points in the normalized thermal environment data. It is a moment The data coupling trajectory vectors are concatenated to obtain the data coupling trajectory.

[0076] Furthermore, the data coupling trajectories are aggregated through covariance alignment to obtain the data prototype trajectories. The data prototype trajectories that exhibit time jumps are split into multiple data prototype trajectories. Discretization rules are applied to the data prototype trajectories to obtain the computing server prototype library.

[0077] Time jump means that the discrete time index in the unified time grid is less than the previous sample.

[0078] Covariance alignment aggregation refers to the process of measuring the distance between data coupling trajectories within the same site signature based on the site covariance, and taking the average of the covariances of time-aligned sites to form the prototype data trajectory.

[0079] The distance metric used is Mahalanobis distance, which is the covariance of the stations.

[0080] Discrete rules are used to obtain a prototype library of computing servers by filtering data prototype trajectories that meet the minimum sample size and minimum support.

[0081] The minimum sample size is determined by the median of the minimum sample size based on the prototype trajectory of historical data.

[0082] Minimum support is determined by the median of the minimum number of times the data prototype trajectory is mapped by the data coupling trajectory in the historical local window.

[0083] The computing power server prototype library is a data prototype trajectory after the trajectory within the same site signature is subjected to discrete rules.

[0084] Furthermore, component identifiers and handling scripts are bound to each data prototype trajectory in the computing server prototype library.

[0085] Component identification is based on hardware error events, replacement records, and physical topology within the time period of the data prototype trajectory, and is explicitly pointed to a specific component through a deterministic mapping table.

[0086] The handling script generates a partial order description based on the action set in historical work orders and execution receipts with the same data prototype trajectory, the sequential and parallel relationship of the action set, and the upper limit of the parallel capacity of the action, and then solidifies it through discrete rules.

[0087] The data coupling trajectory is aligned with the data prototype trajectory within the same site signature in the computing server prototype library in terms of time sequence shape. The data prototype trajectory with the minimum covariance alignment and aggregation value within the same site signature is selected as the unique prototype identifier.

[0088] Within the same site signature, based on the component identifiers and disposal scripts in the computing power server prototype library, read-only component identifier mapping tables and disposal script mapping tables are constructed respectively.

[0089] Within the same site signature, a unique prototype identifier for the data coupling trajectory is mapped. Based on the component identifier mapping table and the disposal script mapping table, the component identifier and disposal script corresponding to the unique prototype identifier are determined.

[0090] S3. By attaching a unique prototype identifier, component identifier, and disposal script to the resource, resource entries are obtained. A strong relationship graph is obtained by constructing a resource relationship graph based on the disposal script. The total order of resources is obtained and scanned to generate a set of maintenance windows.

[0091] The unique prototype identifier, component identifier, disposal script, cooling zone identifier, and resource unique identifier are bound together to obtain the initial resource entries, and time alignment is performed based on a unified time grid.

[0092] The initial resource entry is supplemented with enhanced fields for tenant, project, host, rack number, and rack position to form a resource entry.

[0093] Furthermore, within the statistics window, a resource relationship graph is constructed based on the sequential and parallel relationships of the action sets of resource items extracted from the disposal script. For any two resources, a resource relationship judgment is made. If two resources have a relationship, they are considered as a strong relationship pair. The strong relationship pairs are counted to obtain a strong relationship graph.

[0094] The statistics window is the time period used for map creation and cost statistics.

[0095] Resource relationships include ownership relationships, host relationships, and cooling zone relationships.

[0096] Ownership association refers to a direct management relationship within the same tenant and the same project.

[0097] Same-host association refers to the deployment relationship between the same host and the same physical machine.

[0098] Cooling section association is a thermal coupling relationship between the same cooling section and the same cooling circuit.

[0099] Strong relation graphs compute the total order permutation of resources using mixed-integer linear programming, expressed as:

[0100] ;

[0101] ;

[0102] in, It is a ranking variable. It is any strong relation pair within the observation window. It refers to any resource entry in a strong relation graph within the observation window. It is the ability to observe the window. Another resource item that constitutes a strong relationship pair It is a strong relation pair within the observation window. It is the price of severing strong ties. It is a sequential variable, only when Prior to Take 1 at time. yes The engineering coupling cost, This means that only pairs of different resource entries are counted. It is a total order arrangement of resources.

[0103] The observation window takes the time of the shortest executable action in the disposal script.

[0104] The cost of splitting strong ties is the median of the incremental cost of splitting a pair of strong ties into different adjacent locations.

[0105] Incremental cost is obtained by multiplying the migration unit price by the amount of migration data.

[0106] Engineering coupling cost is the measurable engineering cost incurred when any strong relation pair within the statistical window of the resource pair is processed concurrently within the same site signature. It is obtained by multiplying the known data unit price with the amount of migrated data.

[0107] Furthermore, for each resource, a set of fixed actions, the order and parallel relationships of the action sets, and the upper limit of the parallel capacity of the actions are defined, and maintenance window rules are set.

[0108] The maintenance window rule refers to the allocation of different maintenance windows when any two resources do not meet the general parallelizable rule, and the number of resources in a single maintenance window must not exceed the maintenance window parallel limit.

[0109] The general parallelizability rule is based on the parallel relationship of the action set of resources within the observation window, and the common part of the statistical action set is used as the general parallelizability rule within the observation window.

[0110] The maintenance window is a set of time periods and resources that can be processed simultaneously, arranged in full order of resources within the observation window.

[0111] The maintenance window parallel limit is the upper limit of the parallel capacity of resource actions within the statistical maintenance window. The minimum value of the action parallel capacity upper limit is taken as the maintenance window parallel limit.

[0112] The resources are scanned in total order and arranged in order to obtain a set of maintenance windows.

[0113] S4. Construct a partially ordered lattice through the maintenance window set, and convert the maintenance windows and handling scripts into a set of constraints for handling expired faults.

[0114] Within the set of maintenance windows, obtain the resource list for each maintenance window, and extract the action list, action prerequisite relationship set, and parallel action pair set from the disposal script of the resource list.

[0115] Merge the action lists in the resource list to obtain the complete action set of the maintenance window. Merge the action prerequisite relationship set in the resource list to obtain the maintenance window's prerequisite relationship master table. Take the intersection of the set of parallelizable action pairs in the resource list to obtain the window's common parallel list. Determine whether action pairs remain incomparable in the partial order relationship.

[0116] Based on the order of actions in the preceding relationship table of the maintenance window, the initial preceding relationship of the maintenance window is constructed. For action pairs that are not in the window's common parallel list, they are arranged according to action priority to determine a fixed priority order.

[0117] Action priority is arranged according to the logical setting of software taking precedence over hardware. For example, migration takes precedence over isolation, isolation takes precedence over load reduction, load reduction takes precedence over disk check, disk check takes precedence over fan adjustment, and fan adjustment takes precedence over cooling settings.

[0118] Furthermore, the sequence of actions is solidified, and action pairs that do not belong to the action parallel relationship are arranged according to the action priority. Action pairs within the window's common parallel list are kept incomparable. By satisfying the reflexivity, antisymmetry, and transitivity in the partial order relation, the partial order relation of the maintenance window is obtained. All lower closed subsets and inclusion relations of the partial order relation of the maintenance window constitute the partial order lattice of the maintenance window.

[0119] Furthermore, based on the start and end times of the maintenance window and the resource list, constraints are imposed on the entire set of actions for the maintenance window. The actions must fall within the start and end times of the maintenance window and the bandwidth occupied by the actions must not exceed the available bandwidth margin of the maintenance window, thus obtaining the basis of the constraint set for handling operational faults upon expiration.

[0120] The available bandwidth margin for the maintenance window is the minimum of the known available bandwidth margins of the computing server in the network transmission links involved in the maintenance window.

[0121] The set of constraints for handling expired operational faults is obtained by intersecting the base set of constraints for handling expired operational faults with the partially ordered lattice of the maintenance window.

[0122] In summary, this invention achieves precise correspondence between unique prototype identifiers, component identifiers, and disposal scripts by constructing data coupling trajectories based on dual-stream data sequences and matching them with the computing server prototype library under site signatures. This is used to locate potential failure locations and due disposal paths, ultimately improving the certainty and consistency of operational failure due identification. Furthermore, by constructing a partial-order lattice through a maintenance window set and converting it into an operational failure due disposal constraint set, a unified expression of action sets, sequential parallel relationships, and time and bandwidth limitations is achieved. This forms directly referable disposal boundaries, ultimately supporting a closed-loop health management system.

[0123] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for predicting operational failures based on the health management of computing servers, characterized in that: include, The computing server collects and aligns operational data and thermal environment data according to the unique resource identifier, and calibrates the operational data and thermal environment data to obtain a dual-stream data sequence and site signature; A data coupling trajectory is constructed based on the dual-stream data sequence, and matched with the computing server prototype library under the site signature to obtain a unique prototype identifier, which is then mapped to a component identifier and a disposal script. By attaching a unique prototype identifier, component identifier, and disposal script to the resource, resource entries are obtained. A strong relationship graph is obtained by constructing a resource relationship graph based on the disposal script. The total order of resources is obtained and scanned to generate a set of maintenance windows. A partial order lattice is constructed by maintaining a set of windows, and the maintenance windows and disposal scripts are converted into a set of constraints for handling expired faults. The specific steps for constructing the data coupling trajectory based on the dual-stream data sequence are as follows: based on the station covariance of the normalized operating data sequence and the thermal environment data sequence, the whitening matrix of the normalized operating data sequence and the thermal environment data sequence is calculated. The increment is obtained by calculating the difference increment between adjacent time steps of the normalized operating data sequence and the thermal environment data sequence, and the increment is transformed by the whitening matrix of the normalized operating data sequence and the thermal environment data sequence. The data coupling trajectory is obtained by connecting them according to time. The steps for constructing a resource relationship graph based on the disposal script to obtain a strong relationship graph, obtaining the total order of resources and scanning to generate a maintenance window set are as follows: within the statistics window, constructing a resource relationship graph based on the action set of resource items extracted from the disposal script, and performing resource relationship judgment to output a strong relationship graph. Strong relation graphs calculate the order variables of all resource pairs using mixed-integer linear programming to obtain the total order permutation of resources; Set maintenance window rules and obtain the maintenance window set based on the total order of resources; The specific steps for constructing a partial order lattice through a set of maintenance windows are as follows: within the set of maintenance windows, obtain the complete set of actions for each maintenance window, and filter the order of actions, the parallel relationship of actions, and the priority order of actions. By summarizing the sequence of actions, the parallel relationships of actions, and the priority of actions, a partial order grid for maintaining the window is obtained.

2. The operational fault prediction method based on computing server health management as described in claim 1, characterized in that: The specific steps for collecting and aligning operational data and thermal environment data of the computing server using unique resource identifiers are as follows: By using hierarchical naming to connect the known physical and logical levels of computing power servers, the physical uniqueness of computing power servers and computing power server components is solidified to generate unique resource identifiers. A data acquisition channel is established using the unique resource identifier as the primary key. The data from the computing server is time-aligned using the sampling periods for operational data and thermal environment data, respectively, to obtain operational data and thermal environment data within a unified time grid.

3. The operational fault prediction method based on computing server health management as described in claim 2, characterized in that: The calibration operation data and thermal environment data are combined to obtain a dual-stream data sequence and site signature. The specific steps are as follows: Normalize the operational data and thermal environment data within a unified time grid to obtain normalized operational data sequences and thermal environment data sequences; Within each unified time grid, using the normalized operational data sequence as the parent time axis, causal time warping is performed on the normalized thermal environment data sequence. The normalized operational data and thermal environment data at the matching points of each unified time grid are statistically analyzed and merged to obtain a dual-stream data sequence. Robust covariance is obtained by performing a block diagonal matrix on the dual-stream data sequence to obtain the site covariance. Read the known cooling zone identifiers within the computing server and combine them with the site covariance and resource unique identifier to obtain the site signature.

4. The operational fault prediction method based on computing server health management as described in claim 3, characterized in that: The process of matching the site signature with the computing server prototype library to obtain a unique prototype identifier involves the following steps: Bind component identifiers and handling scripts to each data prototype trajectory in the computing server prototype library; The data coupling trajectory is aligned temporally with the data prototype trajectory within the same site signature in the computing server prototype library to obtain a unique prototype identifier.

5. The operational fault prediction method based on computing server health management as described in claim 4, characterized in that: The mapping of component identifiers and disposal scripts includes, within the same site signature, determining the component identifier and disposal script corresponding to the unique prototype identifier of the data coupling trajectory based on the component identifiers and disposal scripts in the computing power server prototype library.

6. The operational fault prediction method based on computing server health management as described in claim 5, characterized in that: The process of obtaining resource entries by attaching unique prototype identifiers, component identifiers, and disposal scripts to resources involves the following steps: The unique prototype identifier, component identifier, disposal script, cooldown section identifier, and resource unique identifier are bound together to obtain the initial resource entries, and time alignment is performed based on a unified time grid. The initial resource entry is supplemented with enhanced fields for tenant, project, host, and machine location to form the resource entry.

7. The operational fault prediction method based on computing server health management as described in claim 6, characterized in that: The specific steps for converting the maintenance window and handling script into a set of runtime fault expiration handling constraints are as follows: Apply time and bandwidth constraints to the entire set of actions for the maintenance window to obtain the basis of the constraint set for handling operational faults upon expiration. The set of constraints for handling expired operational faults is obtained by intersecting the base set of constraints for handling expired operational faults with the partially ordered lattice of the maintenance window.

Citation Information

Patent Citations

  • Server alarm processing method and related equipment

    CN113821408A

  • Server fault prediction method and system based on BMC module edge calculation

    CN118733317A