A multi-modal based server failure prediction method, apparatus and device

By combining multimodal data fusion and deep learning models with fault knowledge graphs, a fault risk map is generated, which solves the problem of difficulty in discovering hidden faults in GPU servers in traditional methods, and achieves more accurate and efficient fault detection and location.

CN120610874BActive Publication Date: 2025-11-18ZHEJIANG DETACENT DATA TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511121467.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Traditional fault prediction methods based on a single data source are unable to detect hidden faults in GPU servers, resulting in the inability to predict and intervene in potential thermal runaway or equipment damage in a timely manner.

Method used

Multimodal data fusion technology is adopted, including thermal images, system logs, GPU utilization, power consumption, temperature, fan speed and fan audio. Modal features are extracted and fused through deep learning models. Fault knowledge graphs are used to verify the accuracy of predictions and generate fault risk maps to display fault information.

Benefits of technology

It improves the accuracy and reliability of fault detection, reduces false alarms and missed alarms, helps maintenance personnel quickly locate fault areas, improves maintenance efficiency and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610874B_ABST
    Figure CN120610874B_ABST
Patent Text Reader

Abstract

The application provides a kind of multi-modal based server fault prediction method, device and equipment, applied to fault prediction technical field, the method includes the multi-modal data of acquisition node server. Multi-modal data includes at least two of thermal image, system log, GPU utilization, power consumption, temperature, fan speed and fan audio. From the modal feature corresponding to each mode of multi-modal data is extracted respectively. The modal feature of all modes in multi-modal data is fused, and the fault information of node server is predicted. The fault information is mapped to the thermal image to obtain the fault risk map of node server. Fault risk map is used to superimpose and display fault information in thermal image. It can solve the problem that traditional fault prediction method based on single data source is difficult to find implicit problems generated by server, so as to predict server failure in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fault prediction technology, and in particular to a method, apparatus and device for server fault prediction based on multimodality. Background Technology

[0002] With the rapid development of artificial intelligence, big data, and high-performance computing, GPU servers, due to their superior parallel computing capabilities, have been widely used in data centers and high-performance computing clusters to handle high-load tasks such as large model training, image rendering, and scientific computing. During continuous high-load operation, GPUs easily generate a large amount of heat. If this heat is not dissipated effectively and promptly, the GPU core temperature may become too high, leading to serious consequences such as system throttling, performance degradation, hardware damage, or even system crashes, affecting task continuity and system stability.

[0003] To ensure the reliable operation of GPU servers, related technologies commonly employ monitoring methods based on hardware sensors and system logs to monitor the GPU's operational status. These methods typically rely on a single-modal data source, such as temperature values ​​from temperature sensors or error messages recorded in system logs, to determine if the GPU is malfunctioning. However, in practical applications, GPU failures are often complex and insidious, such as gradual fan performance degradation, partial failures in the cooling system, or abnormal changes in heat conduction paths. These problems may not immediately trigger obvious temperature spikes or system alarms, but they accumulate gradually, eventually leading to thermal runaway or equipment damage. Therefore, traditional fault prediction methods based on a single data source often lack sufficient early warning capabilities when facing these slowly evolving, multi-factor coupled, and latent faults, making early detection and timely intervention difficult. Summary of the Invention

[0004] The purpose of this application is to provide a multimodal server fault prediction method, apparatus and equipment to solve the problem that traditional fault prediction methods based on a single data source are difficult to discover hidden problems caused by servers, and thus cannot predict server faults in a timely manner.

[0005] Firstly, embodiments of this application provide a multimodal server fault prediction method applied to a server platform, which includes at least one node server. The method includes: collecting multimodal data from the node server. The multimodal data includes at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio. Modal features corresponding to each modality are extracted from the multimodal data. The modal features of all modalities in the multimodal data are fused to predict fault information of the node server. The fault information is mapped onto the thermal image to obtain a fault risk map of the node server. The fault risk map is used to overlay and display fault information on the thermal image.

[0006] The multimodal server fault prediction method provided in this application collects multimodal data from node servers, including at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio. It extracts modal features corresponding to each modality, fuses these features to predict fault information of the node server, and maps the fault information onto the thermal image to generate a fault risk map to display the fault information. This method can effectively improve the accuracy and reliability of fault detection, reduce false alarms and missed alarms, help maintenance personnel quickly locate fault areas, improve maintenance efficiency, and reduce maintenance costs.

[0007] One possible implementation involves incorporating fault information including all fault types and the corresponding fault probability for each type. This involves fusing modal features from all modalities in the multimodal data to predict fault information for the node server, including: temporally aligning and fusing the modal features of all modalities to construct a feature vector. The trained fault prediction model is then used to classify and predict faults on the feature vector, determining all fault types of the node server and the corresponding fault probability for each fault type.

[0008] One possible implementation method further includes: acquiring a pre-built fault knowledge graph on the server platform. The fault knowledge graph includes multiple nodes. Nodes include modal feature nodes, fault type nodes, device type nodes, and feedback information nodes. Modal feature nodes correspond to multimodal data. Fault type nodes correspond to fault types. Device type nodes represent the device models of devices in the node servers. Devices include GPUs and fans. Feedback information nodes represent feedback from operations and maintenance personnel to fault information. The fault knowledge graph is used to determine the confidence level of the fault information. The confidence level is used to verify the authenticity of the fault information.

[0009] One possible implementation involves temporally aligning and fusing the modal features of all modalities to construct a feature vector. This includes: for any target modality among all modalities, determining the fluctuation intensity of the target modality within the current acquisition window. The fluctuation intensity represents the degree of change of the modal features corresponding to the target modality within the current acquisition window. The attention weights of the target modality are adjusted using the fluctuation intensity to obtain the fluctuation weights corresponding to the target modality. Based on the modal features of all modalities and the corresponding fluctuation weights, the feature vector is determined.

[0010] One possible implementation involves mapping fault information to a thermal image to obtain a fault risk map for the node server. This includes: determining the corresponding hotspot activation map based on the thermal image. The hotspot activation map represents the image region in the thermal image related to the fault information. The fault information is then mapped to the thermal image using the hotspot activation map to determine the fault risk map for the node server.

[0011] One possible implementation method further includes: receiving feedback information from operations and maintenance personnel regarding fault information; binding the feedback information with multimodal data to determine the labeled samples for node servers; determining the node weight for each node server using the device types of multiple node servers on the server platform and their corresponding labeled samples; and adjusting the global parameters of the global fault prediction model of the server platform based on the node weight and the model parameters of the fault prediction model for each node server. The global fault prediction model is used to indicate the adjustments made to the fault prediction model for each node server.

[0012] One possible implementation involves, after receiving feedback from maintenance personnel regarding fault information, updating the server platform's fault knowledge graph using the fault information of the node server, the modal features corresponding to each modality in the multimodal data, the device type of the node server, and the feedback from maintenance personnel regarding fault information.

[0013] One possible implementation method further includes: when any modality data is detected to be missing in the multimodal data, reasoning is performed using the relationships in a pre-built fault knowledge graph to determine the missing modality data. The determined missing modality data is then filled into the multimodal data accordingly.

[0014] Secondly, embodiments of this application provide a multimodal server fault prediction device applied to a server platform, the server platform including at least one node server, and the device including: a collection module, an extraction module, a prediction module and a mapping module.

[0015] The acquisition module is used to collect multimodal data from the node servers. This multimodal data includes at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio.

[0016] The extraction module is used to extract modal features corresponding to each modality from multimodal data.

[0017] The prediction module is used to fuse the modal features of all modalities in multimodal data to predict fault information of node servers.

[0018] The mapping module maps fault information onto thermal images to obtain a fault risk map of the node server. This fault risk map is then used to overlay fault information onto the thermal images.

[0019] Thirdly, embodiments of this application provide a multimodal server fault prediction device. This multimodal server fault prediction device has the function of implementing the multimodal server fault prediction method of the first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the multimodal server fault prediction method described in the first aspect or any possible implementation thereof.

[0021] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, enable the computer to execute the multimodal server fault prediction method described in the first aspect or any possible implementation thereof.

[0022] The technical effects of any of the design methods in aspects two through five can be found in aspect one or in different possible implementations of aspect one, and will not be repeated here. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a multimodal server fault prediction method provided in this application embodiment;

[0025] Figure 2 A specific example diagram of a thermal image provided in an embodiment of this application;

[0026] Figure 3 A specific example diagram of a fault knowledge graph provided in an embodiment of this application;

[0027] Figure 4 A specific example diagram of a fault risk diagram provided in the embodiments of this application;

[0028] Figure 5 A schematic diagram of a multimodal server fault prediction device provided in this application embodiment;

[0029] Figure 6 This is a system architecture diagram of a multimodal server fault prediction system provided in an embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0031] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0032] To ensure the reliable operation of GPU servers, related technologies commonly employ monitoring methods based on hardware sensors and system logs to monitor the GPU's operational status. These methods typically rely on a single-modal data source, such as temperature values ​​from temperature sensors or error messages recorded in system logs, to determine if the GPU is malfunctioning. However, in practical applications, GPU failures are often complex and insidious, such as gradual fan performance degradation, partial failures in the cooling system, or abnormal changes in heat conduction paths. These problems may not immediately trigger obvious temperature spikes or system alarms, but they accumulate gradually, eventually leading to thermal runaway or equipment damage. Therefore, traditional fault prediction methods based on a single data source often lack sufficient early warning capabilities when facing these slowly evolving, multi-factor coupled, and latent faults, making early detection and timely intervention difficult.

[0033] Based on this, this application provides a multimodal server fault prediction method applied to a server platform, which includes at least one node server. The method includes: collecting multimodal data from the node server. The multimodal data includes at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio. Modal features corresponding to each modality are extracted from the multimodal data. The modal features of all modalities in the multimodal data are fused to predict fault information of the node server. The fault information is mapped onto the thermal image to obtain a fault risk map of the node server. The fault risk map is used to overlay and display fault information on the thermal image.

[0034] The multimodal server fault prediction method provided in this application collects multimodal data from node servers, including at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio. It extracts modal features corresponding to each modality, fuses these features to predict fault information of the node server, and maps the fault information onto the thermal image to generate a fault risk map to display the fault information. This method can effectively improve the accuracy and reliability of fault detection, reduce false alarms and missed alarms, help maintenance personnel quickly locate fault areas, improve maintenance efficiency, and reduce maintenance costs.

[0035] The methods provided in the embodiments of this application will now be described in conjunction with the specific accompanying drawings.

[0036] On the one hand, embodiments of this application provide a multimodal server fault prediction method, which may include the following steps.

[0037] S101 collects multimodal data from the node server.

[0038] The multimodal data includes at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio.

[0039] One possible implementation is to use a thermal infrared camera to capture thermal images of the GPU chip or the entire server.

[0040] For example, a thermal imager such as the FLIR A700 or Seek ShotPro can be used to perform high-frequency sampling on the GPU chip or the entire server to capture its thermal distribution and obtain a thermal image. Figure 2 As shown, Figure 2 This is a specific example diagram of a thermal image provided in an embodiment of this application. The image resolution of this thermal image is 640×480@30Hz. The pixel values ​​in the thermal image represent specific temperature values ​​in degrees Celsius. Brighter areas in the thermal image indicate higher temperatures in those areas.

[0041] Another possible approach is to collect log information from multiple levels of the server.

[0042] For example, log information may include GPU driver logs (such as logs generated by NVIDIA's nvidia-smi tool), Kubernetes container logs, and system syslogs. The log information records various warnings and error messages generated by the server during operation.

[0043] For example, one log entry shows "2025-06-10 10:12:44 [WARN] GPU0 Core Temp 95°C (overheating)," clearly indicating that the GPU0 core temperature is too high, reaching 95°C, posing a risk of overheating; another log entry, "2025-06-10 10:13:01 [ERROR] GPU0 Fan 0 RPM (stopped)," indicates that the GPU0 fan speed is 0, meaning the fan has stopped running.

[0044] Another possible implementation involves collecting data on the server's GPU utilization, power consumption, temperature, and fan speed, and converting these into independent time-series vectors.

[0045] For example, taking power consumption as an example, the collected power consumption can be converted into a time-series vector: P =( P 1, P 2, ..., Pt ,...).in, Pt This represents the power value at time t.

[0046] Another possible implementation is to use a microphone to capture the server's fan audio.

[0047] For example, miniature stationary microphones (such as MEMS arrays) are placed on the GPU server or rack side to capture audio signals at a sampling rate of 44.1 kHz, with the option of single-channel or dual-channel acquisition. Each audio segment is captured for 3 to 5 seconds and updated using a sliding window to continuously monitor fan audio generated during server operation.

[0048] S102, extract the modal features corresponding to each modality from the multimodal data.

[0049] Specifically, spatial features of thermal distribution are extracted from thermal images. Semantic features related to faults are extracted from system logs. Trend change features are extracted from GPU utilization, power consumption, temperature, and fan speed. Abnormal noise frequency band features are extracted from fan audio.

[0050] One possible implementation involves using a pre-trained deep learning model, such as a convolutional neural network (CNN) or a Vision Transformer, to extract spatial features of heat distribution from thermal images. This deep learning model can automatically learn complex patterns and structures in the image, capturing subtle changes in heat distribution and thus providing rich spatial information for fault detection. For example, after processing the thermal image using a CNN or Vision Transformer, we obtain an image feature vector. feature = [0.23, -0.12, ..., 0.09], with a dimension of 512. These feature values ​​can be used to represent the feature intensity of an image at different spatial locations.

[0051] Another possible approach is to use a BERT encoder to extract fault-related semantic features from system logs. The BERT encoder can understand the context and semantic information within the system log text, transforming it into a high-dimensional semantic feature vector. For example, after BERT encoding the system logs, a text feature vector Log can be obtained. feature = [0.01, 0.88, ..., 0.44], with a dimension of 768, these feature values ​​can be used to represent key semantic information in system logs.

[0052] Another possible approach is to use Temporal Convolutional Networks (TCNs) or Long Short-Term Memory Networks (LSTMs) to capture trends and temporal dependencies in the data, specifically for temperature, power consumption, and fan speed. This process extracts features reflecting changes in system operating status from temperature, power consumption, and fan speed based on long-term dependencies in the time-series data. For example, after processing the time-series data using a TCN or LSTM, a temporal feature vector TS can be obtained. feature = [0.25, 0.35, ..., 0.12], with a dimension of 128. These feature values ​​can be used to represent the trend changes of various time series indicators at different time points.

[0053] Another possible implementation is to use Mel-frequency cepstral coefficients (MFCC) to characterize fan speed variations and abnormal noise frequency bands. MFCC can capture frequency and energy distribution information in the sound signal, allowing for the identification of normal fan operation sounds and abnormal noises. For example, after processing the fan audio signal with MFCC, a sound feature vector AD can be obtained. feature = [0.34, 0.43, ..., 0.23], with a dimension of 256. These feature values ​​can be used to represent the energy distribution of fan sound at different frequency bands.

[0054] S103 fuses the modal features of all modes in the multimodal data to predict the fault information of the node server.

[0055] The fault information includes: all fault types and the fault probability corresponding to each fault type.

[0056] One possible implementation involves temporal alignment and feature fusion of all modal features to construct a feature vector. The trained fault prediction model is then used to classify and predict faults using this feature vector, determining all fault types of the node server and the corresponding fault probability for each type.

[0057] Specifically, after extracting modal features from all modalities in the multimodal data, the modal features of all modalities are first aligned in the time dimension to ensure consistency across all modalities. After time alignment, a Transformer or a multimodal fusion network (such as MM-Transformer) is used to integrate the modal features of all modalities to construct a fused input tensor, i.e., a feature vector.

[0058] The trained fault prediction model is then used to classify and predict faults in the fused feature vectors, determining all fault types of the node server and the fault probability corresponding to each fault type.

[0059] For example, a multilayer perceptron (MLP) or a Transformer decoder is used to predict the fused vector, outputting the fault type (such as "Fan Failure", "GPU Overheat", "Power Anomaly", "Normal") and its corresponding fault probability. For example, the output results could be: [Fan Failure: 0.72, GPU Overheat: 0.11, Power Anomaly: 0.03, Normal: 0.14].

[0060] Furthermore, in order to improve the accuracy and sensitivity of fault prediction, in the process of feature fusion of modal features of all modalities to construct feature vectors, this application embodiment uses a modal attention mechanism and a modal feature volatility estimation mechanism to dynamically guide the allocation of attention weights, so that the attention allocation is not only based on modal features, but also on abnormal fluctuations in modal data, making fault prediction particularly sensitive to sudden events in GPU servers, such as sudden abnormal noises or abnormal power consumption fluctuations.

[0061] One possible implementation involves determining the fluctuation intensity of any target mode within the current acquisition window, given all modalities. The fluctuation intensity represents the degree of change in the modal features corresponding to the target mode within the current acquisition window. The attention weights of the target mode are then adjusted using the fluctuation intensity to obtain the corresponding fluctuation weights. Finally, a feature vector is determined based on the modal features of all modalities and their corresponding fluctuation weights.

[0062] Specifically, for any target mode among all modes, the fluctuation intensity, i.e., the fluctuation coefficient, of the target mode within the current acquisition window is first determined. The fluctuation intensity represents the degree of change of the modal characteristics corresponding to the target mode within the current acquisition window. It can be determined using the following formula.

[0063]

[0064] in, Let the standard deviation of this mode be within the sliding window. The mean, Let be a small constant, and i be the target mode.

[0065] Then, the attention weights of the target mode are adjusted using the fluctuation coefficient to obtain the fluctuation weights corresponding to the target mode. .

[0066]

[0067] Among them, W i and W j Let h be the weight matrix. i h is the eigenvector. j Let be the sum of the eigenvectors, where i is the index of the eigenvector and j is the index of the summation of the i eigenvectors.

[0068] This application embodiment uses a modal attention mechanism and a modal feature volatility estimation mechanism to dynamically guide the allocation of attention weights, so that the attention allocation is not only based on modal content, but also takes into account modal abnormal fluctuations. This makes the model more sensitive to situations such as sudden abnormal noises in the GPU server and abnormal power consumption fluctuations, and can detect potential faults in the server earlier.

[0069] Furthermore, in order to improve the accuracy of the fault information predicted by the fault prediction model, this application also uses a pre-constructed fault knowledge graph to determine the confidence level of the fault information, thereby determining the true accuracy of the fault information prediction.

[0070] The fault knowledge graph can be constructed based on modal features, fault types, equipment types, and feedback information from historical fault prediction processes. These modal features, fault types, equipment types, and feedback information are organized into multiple nodes and edges, thus forming a directed fault knowledge graph.

[0071] This node can include modal feature nodes, fault type nodes, device type nodes, and feedback information nodes. Modal feature nodes correspond to multimodal data. Fault type nodes correspond to fault types. Device type nodes represent the device models of devices in the node server. Devices include GPUs and fans. Feedback information nodes represent feedback from operations and maintenance personnel regarding fault information.

[0072] For example, the fault type node could be a fan noise or memory overheating. The modal characteristic node could be a high-frequency noise peak or thermal map asymmetry. The device type node could be the GPU model or fan specifications. The feedback information node could be confirmation or false alarm.

[0073] Edges are used to represent the relationships between connected nodes.

[0074] For example, the edge type can be "cause_of", "has_feature", "validated_by", "similar_to", etc.

[0075] After the graph modeling is completed, modal features, fault types, equipment types and feedback information from the historical fault prediction process are constructed and their corresponding representations are filled into the graph to build a complete fault knowledge graph.

[0076] For example, such as Figure 3 As shown, Figure 3 This is a specific example diagram of a fault knowledge graph provided in an embodiment of this application. In this fault knowledge graph, yellow nodes are modal feature nodes, pink nodes are fault type nodes, green nodes are device type nodes, and purple nodes are feedback information nodes. The device type determined during historical fault prediction is GPU model Mthead S4000, and fan model Fan S23. The extracted modal features are a fan peak frequency of 3kHz and a GPU throttling issue; the fault type is fan malfunction; and the feedback information is "feedback is true".

[0077] Specifically, when using a fault knowledge graph to determine the confidence level of fault information, a fault knowledge graph pre-built on the server platform can be obtained. This fault knowledge graph is then used to determine the confidence level of the fault information. The confidence level is used to verify the authenticity of the fault information.

[0078] Specifically, based on the predicted fault information and the equipment type and modal characteristics determined during the fault information prediction process, a corresponding directed graph is determined in the fault knowledge graph using the reasoning capability of the fault knowledge graph. If the feedback information indicated by the directed graph is true, the confidence level of this fault information prediction is increased.

[0079] For example, if the collected fan peak value is 3.2kHz and the corresponding modal features are extracted as fan anomalies, the corresponding feedback information can be determined as true based on the fault knowledge graph, thereby improving the confidence of the fault information prediction.

[0080] It should be noted that the fault knowledge graph can also be used as a training sample generator. Soft-label data can be constructed through the nodes in the fault knowledge graph and the correspondence between the nodes. Using the soft-label data for knowledge distillation and reinforcement training can improve the reliability of fault information predicted by the fault prediction model.

[0081] S104 maps the fault information onto the thermal image to obtain the fault risk map of the node server.

[0082] Among them, the fault risk map is used to overlay and display fault information on the thermal image.

[0083] One possible implementation involves determining the corresponding hotspot activation map based on the thermal image. The hotspot activation map represents the image region in the thermal image that is related to fault information. By mapping the fault information to the thermal image using the hotspot activation map, a fault risk map for the node server can be determined.

[0084] For example, Grad-CAM is performed on a thermal image. First, the thermal image is forward-propagated through a CNN until the last convolutional layer, and the output of the feature map of the last convolutional layer is recorded. Forward propagation continues until the network output layer, and the score for the target class is calculated. Backpropagation is then performed on the target class score to calculate the gradient of the feature map of the last convolutional layer. This gradient represents the contribution of each location on the feature map to the target class score. Global average pooling is performed on the gradient of each channel of the last convolutional layer to obtain the weights for each channel. The feature map of each channel is multiplied by its corresponding weight, and the results of all channels are summed to obtain a single thermal image. The thermal image is normalized and mapped onto the thermal image. By adjusting the transparency factor (β) and applying a pseudo-color transformation function (colormap), fault information is superimposed onto the thermal image in a color-coded form to obtain a fault risk map. The fault risk map can be generated using the following formula.

[0085]

[0086] Among them, Irisk (x, y) represents the pixel value of the fault risk map at coordinates (x, y), I original (x, y) represents the pixel value of the thermal image at coordinates (x, y), β is the transparency adjustment factor, colormap is the pseudo-color conversion function, and G... cam (x,y) represents the pixel value of the active heatmap of the thermal channel at coordinates (x,y).

[0087] This process generates a fault risk map by overlaying fault information onto a thermal image, for example, such as... Figure 4 The diagram shown is a specific example of a fault risk map provided in an embodiment of this application. Maintenance personnel can use this fault risk map to quickly identify and locate faulty areas, such as modules G1 and G5.

[0088] Furthermore, the system receives feedback from operations and maintenance personnel regarding fault information. This feedback is then bound to multimodal data to determine the labeled samples for each node server. Using the device types of multiple node servers on the server platform and their corresponding labeled samples, the node weight for each node server is determined. Based on the node weight and the model parameters of the fault prediction model for each node server, the global parameters of the server platform's global fault prediction model are adjusted. This global fault prediction model serves as the indicator for adjusting the fault prediction model for each node server.

[0089] Specifically, after completing fault information prediction, detailed prediction information can be automatically provided to operations and maintenance personnel. For example, the prediction information may include all fault types predicted during the fault information prediction process, the probability of each fault type, and the contribution of key modes. Key modes are those with fluctuation weights exceeding a threshold; for example, the sound mode accounts for 78%. Operations and maintenance personnel need to confirm these prediction results; feedback options may include "confirm fault," "false alarm," and "ignore."

[0090] Then, the system records the feedback information and the corresponding modal feature bindings to form labeled samples. Using the device types of multiple node servers on the server platform and the corresponding labeled samples, the node weight for each node server is determined. This node weight can be determined using the following formula.

[0091]

[0092] Where, ω k Let λk be the weight of the k-th server, λ1, λ2, and λ3 be the weight coefficients, and n be the weight of the server. k Var is the local sample size of the k-th server, where n is the total sample size of all servers. kGPUTypeFactor is the variance of the prediction loss for the k-th server. k This represents the sensitivity factor of the device type of the k-th server to the target fault (different models of GPUs have different sensitivity factors to faults; for example, the L40 model GPU is more sensitive than the T4 model GPU).

[0093] Finally, based on the node weight corresponding to each node server and the model parameters of the fault prediction model corresponding to each node server, the global parameters of the global fault prediction model of the server platform are adjusted. These global parameters can be determined by the following formula.

[0094]

[0095] Where θ is a global parameter, θ k These are the model parameters for the k-th server.

[0096] Furthermore, during the adjustment of the global parameters of the global fault prediction model on the server platform, when a server is detected to have triggered a preset number of alarms (e.g., more than three fan failures), the model parameters for that server can be uploaded to the server platform in advance. The server platform then adjusts the global parameters based on these model parameters. In this process, the global model can quickly adapt to fault information from sudden failures, improving the real-time performance and accuracy of fault prediction.

[0097] Furthermore, by utilizing fault information from node servers, modal features corresponding to each modality in multimodal data, device types of node servers, and feedback from maintenance personnel regarding fault information, the fault knowledge graph of the server platform is updated.

[0098] Specifically, by utilizing fault information from node servers, modal features corresponding to each modality in multimodal data, device types of node servers, and feedback information from maintenance personnel regarding fault information, the types of nodes and edges between nodes in the fault knowledge graph are updated.

[0099] If any modality data is found to be missing in the multimodal data, the missing modality data is determined by reasoning based on the relationships in a pre-built fault knowledge graph. The missing modality data is then filled into the multimodal data accordingly.

[0100] Specifically, during data acquisition, the system can detect any missing modal data, such as thermal images due to camera malfunction or data transmission issues. When modal data loss is detected, inference can be performed using relationships within a fault knowledge graph. For example, when a thermal image is missing, modal data from other modes, such as fan speed, temperature, and power consumption, along with known relationships with the thermal image, can be used to infer the missing image. The incomplete multimodal data is then input into the fault prediction model to fill in the inference blind spots in modal loss scenarios.

[0101] The above primarily describes the solutions provided in this application from the perspective of the device's working principle. It is understood that, in order to achieve the above functions, the multimodal server fault prediction device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the algorithm steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0102] This application embodiment can divide the multimodal server fault prediction device into functional modules according to the above method example. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module.

[0103] It should be noted that the module division in this embodiment is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. When dividing functional modules according to their respective functions, Figure 5 A schematic diagram illustrating a possible composition of the multimodal server fault prediction device described above and in the embodiments is shown. Figure 5 As shown, the multimodal server fault prediction device 500 may include: a data acquisition module 501, an extraction module 502, a prediction module 503, and a mapping module 504.

[0104] The acquisition module 501 is used to support the execution of the multimodal server fault prediction device 500. Figure 1 S201 in the schematic multimodal server fault prediction method.

[0105] Extraction module 502 is used to support the execution of multimodal server fault prediction device 500. Figure 1 S202 in the schematic multimodal server fault prediction method.

[0106] Prediction module 503 is used to support the execution of multimodal server fault prediction device 500. Figure 1 S203 in the schematic multimodal server fault prediction method.

[0107] Mapping module 504 is used to support the execution of multimodal server fault prediction device 500. Figure 1 S204 in the schematic multimodal server fault prediction method.

[0108] One possible implementation involves fault information including all fault types and the corresponding fault probability for each fault type. Specifically, the multimodal server fault prediction device performs temporal alignment and feature fusion of all modal features to construct a feature vector. The trained fault prediction model then uses this feature vector to classify and predict faults, determining all fault types of the node server and the corresponding fault probability for each fault type.

[0109] One possible implementation involves a multimodal server fault prediction device that specifically acquires a pre-built fault knowledge graph from the server platform. The fault knowledge graph comprises multiple nodes. These nodes include modal feature nodes, fault type nodes, device type nodes, and feedback information nodes. Modal feature nodes correspond to multimodal data. Fault type nodes correspond to fault types. Device type nodes represent the device models of devices within the node servers. Devices include GPUs and fans. Feedback information nodes represent feedback from maintenance personnel regarding fault information. The fault knowledge graph is used to determine the confidence level of the fault information. This confidence level is then used to verify the authenticity of the fault information.

[0110] One possible implementation involves a multimodal server fault prediction device that, for any target modality among all modalities, determines the fluctuation intensity of the target modality within the current acquisition window. The fluctuation intensity represents the degree of change of the modal features corresponding to the target modality within the current acquisition window. The attention weights of the target modality are adjusted using the fluctuation intensity to obtain the fluctuation weights corresponding to the target modality. Based on the modal features of all modalities and their corresponding fluctuation weights, a feature vector is determined.

[0111] One possible implementation involves a multimodal server fault prediction device that determines a thermal channel activation hotspot map corresponding to a thermal image. The thermal channel activation hotspot map represents the image region in the thermal image related to fault information. By mapping the fault information to the thermal image using the thermal channel activation hotspot map, a fault risk map for the node server is determined.

[0112] One possible implementation involves a multimodal server fault prediction device that receives feedback from maintenance personnel regarding fault information. This feedback is then bound to multimodal data to determine labeled samples for each node server. Utilizing the device types of multiple node servers on the server platform and their corresponding labeled samples, the node weight for each node server is determined. Based on the node weights and model parameters of the fault prediction model for each node server, the global parameters of the server platform's global fault prediction model are adjusted. This global fault prediction model serves as the indicator for adjusting the fault prediction model for each node server.

[0113] One possible implementation is that the multimodal server fault prediction device is used to update the fault knowledge graph of the server platform by utilizing the fault information of the node server, the modal features corresponding to each modality in the multimodal data, the device type of the node server, and the feedback information of the operation and maintenance personnel on the fault information.

[0114] One possible implementation involves a multimodal server fault prediction device that, upon detecting a missing modality in the multimodal data, uses relationships from a pre-built fault knowledge graph to infer the missing modality. The missing modality is then filled into the multimodal data accordingly.

[0115] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0116] The multimodal server fault prediction device 500 provided in this application embodiment is used to perform the above-mentioned... Figure 1 The multimodal server fault prediction method shown can achieve the same effect as the multimodal server fault prediction method described above.

[0117] This application also provides a multimodal server fault prediction device, which can execute the multimodal server fault prediction method and related steps described in the above method embodiments.

[0118] This application also provides a computer-readable storage medium storing instructions thereon, which, when executed, perform the multimodal server fault prediction method and related steps described in the above method embodiments.

[0119] This application also provides a computer program product that, when run on a computer, causes the computer to execute the multimodal server fault prediction method and related steps described in the above method embodiments.

[0120] In some embodiments, the methods shown in this application can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of art.

[0121] This application also provides a multimodal server fault prediction system 600, such as... Figure 6 As shown, the multimodal server fault prediction system 600 includes at least one processor 601 and at least one interface circuit 602.

[0122] As an example, when the multimodal-based server fault prediction system 600 includes a processor and an interface circuit, the processor can be... Figure 6 The processor 601 shown in the solid box (or the processor 601 shown in the dashed box) can be an interface circuit. Figure 6 The interface circuit 602 is shown in the solid box (or the dashed box). When the multimodal server fault prediction system 600 includes two processors and two interface circuits, the two processors include... Figure 6 The processor 601 shown in the solid box and the processor 601 shown in the dashed box, these two interface circuits include Figure 6 Interface circuit 602 is shown in both solid and dashed boxes. No limitations are imposed on this.

[0123] Processor 601 and interface circuit 602 can be interconnected via a line. For example, interface circuit 602 can be used to receive signals. Alternatively, interface circuit 602 can be used to send signals to other devices (e.g., processor 601). For instance, interface circuit 602 can read computer instructions stored in memory and send those instructions to processor 601. Processor 601 executes the instructions and, in conjunction with input / output devices, implements the various steps in the above embodiments, such as implementing... Figures 1-4 The steps performed in any of the method embodiments shown herein. Of course, this multimodal server fault prediction system may also include other discrete components, and the embodiments of this application do not specifically limit this.

[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0126] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0128] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to it, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A server fault prediction method based on multimodality, characterized in that, Applied to a server platform, the server platform including at least one node server, the method includes: Collect multimodal data from the node server; the multimodal data includes at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio. Modal features corresponding to each mode are extracted from the multimodal data; The modal features of all modes in the multimodal data are fused to predict the fault information of the node server; The heatmap is forward-propagated through a CNN until the last convolutional layer, and the output of the feature map of the last convolutional layer is recorded. Forward propagation continues until the network output layer, and the score of the target class is calculated. Backpropagation is performed on the target class score to calculate the gradient of the feature map of the last convolutional layer. The gradient of the feature map of this convolutional layer is used to represent the contribution of each position on the feature map to the target class score. Global average pooling is performed on the gradient of each channel of the last convolutional layer to obtain the weight of each channel. The feature map of each channel is multiplied by its corresponding weight, and then the results of all channels are summed to obtain a single heatmap. The heatmap is normalized and mapped onto the thermal image. By adjusting the transparency adjustment factor and applying the pseudo-color conversion function, the fault information is superimposed onto the thermal image in the form of color encoding to determine the fault risk map of the node server. The fault risk map is used to overlay and display the fault information on the thermal image.

2. The method according to claim 1, characterized in that, The fault information includes: all fault types and the fault probability corresponding to each fault type; the step of fusing the modal features of all modes in the multimodal data to predict the fault information of the node server includes: Modal features of all modalities are time-aligned and feature fused to construct feature vectors; The trained fault prediction model is used to classify and predict faults in the feature vector, thereby determining all fault types of the node server and the fault probability corresponding to each fault type.

3. The method according to claim 2, characterized in that, The method further includes: The system acquires a pre-built fault knowledge graph from the server platform. The fault knowledge graph includes multiple nodes, each comprising a modal feature node, a fault type node, a device type node, and a feedback information node. The modal feature nodes correspond to the multimodal data. The fault type nodes correspond to the fault types. The device type nodes represent the device models of the devices in the node servers. The devices include GPUs and fans. The feedback information nodes represent feedback from maintenance personnel regarding fault information. The confidence level of the fault information is determined using the fault knowledge graph; the confidence level is used to verify the authenticity of the fault information.

4. The method according to claim 2, characterized in that, The process of temporal alignment and feature fusion of all modal features to construct a feature vector includes: For any target mode among all modes, determine the fluctuation intensity of the target mode within the current acquisition window; the fluctuation intensity is used to represent the degree of change of the modal feature corresponding to the target mode within the current acquisition window; The attention weight of the target mode is adjusted using the fluctuation intensity to obtain the fluctuation weight corresponding to the target mode; The feature vector is determined based on the modal characteristics of all modes and the fluctuation weights corresponding to the modes.

5. The method according to claim 1, characterized in that, The method further includes: Receive feedback information from maintenance personnel regarding the fault information; The feedback information is bound to the multimodal data to determine the labeled samples of the node server; By utilizing the device types of multiple node servers under the server platform and the corresponding labeled samples, the node weight corresponding to each node server is determined; Based on the node weight corresponding to each node server and the model parameters of the fault prediction model corresponding to each node server, the global parameters of the global fault prediction model of the server platform are adjusted; the global fault prediction model is used to indicate the adjustment of the fault prediction model for each node server.

6. The method according to claim 5, characterized in that, After receiving feedback from maintenance personnel regarding the fault information, the method further includes: The fault knowledge graph of the server platform is updated by utilizing the fault information of the node server, the modal features corresponding to each modality in the multimodal data, the device type of the node server, and the feedback information of the operation and maintenance personnel on the fault information.

7. The method according to claim 1, characterized in that, The method further includes: If any modality data is found to be missing in the multimodal data, the missing modality data is determined by reasoning based on the relationships in the pre-built fault knowledge graph. The missing model data will be filled into the multimodal data.

8. A server fault prediction device based on multimodality, characterized in that, Applied to a server platform, the server platform including at least one node server, the device includes: The acquisition module is used to acquire multimodal data from the node server; the multimodal data includes at least two of the following: thermal images, system logs, GPU utilization, power consumption, temperature, fan speed, and fan audio. The extraction module is used to extract modal features corresponding to each modality from the multimodal data; The prediction module is used to fuse the modal features of all modes in the multimodal data to predict the fault information of the node server. The mapping module is used to propagate the heatmap forward through a CNN until the last convolutional layer, and record the output of the feature map of the last convolutional layer; continue forward propagation until the network output layer, and calculate the score of the target category; perform backpropagation on the target category score, and calculate the gradient of the feature map of the last convolutional layer; wherein, the gradient of the feature map of this convolutional layer is used to represent the contribution of each position on the feature map to the target category score; perform global average pooling on the gradient of each channel of the last convolutional layer to obtain the weight of each channel; multiply the feature map of each channel by its corresponding weight, and then add the results of all channels to obtain a single heatmap; normalize the heatmap, and map the normalized heatmap onto the heatmap; by adjusting the transparency adjustment factor and applying the pseudo-color conversion function, the fault information is superimposed on the heatmap in the form of color encoding to determine the fault risk map of the node server; the fault risk map is used to superimpose and display the fault information on the heatmap.

9. A server fault prediction device based on multimodal characteristics, characterized in that, The multimodal server fault prediction device includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, and the processor executing the machine-executable instructions to implement the multimodal server fault prediction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Danger detection method and device

    CN114581377A

  • Server fault early warning method, device and equipment and storage medium

    CN116361132A

  • Multi-source heterogeneous data collaborative fault prediction method and system for 10KV substation equipment

    CN120031549A

  • Power equipment fault prediction method based on multi-modal data and related equipment

    CN120197059A