Fault diagnosis methods, devices, media and program products

By obtaining the timing data of computer components, extracting timing characteristics and using multiple models for comprehensive diagnosis, the problem of GPU fault diagnosis is solved, and efficient and accurate fault diagnosis is achieved to ensure the stability and user experience of the server.

CN119847809BActive Publication Date: 2025-08-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510336312.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-08-08
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

In the prior art, GPU fault diagnosis takes a long time and has low accuracy, which affects server stability and user experience.

Method used

By obtaining the timing data of computer components, extracting timing characteristics, and using multiple models for troubleshooting, the diagnostic results are finally determined through fusion processing, including comprehensive diagnosis of time domain features, frequency domain features and event sequence features.

Benefits of technology

It realizes efficient and accurate GPU fault diagnosis, ensures the stability and user experience of the server, and can identify short- and long-term faults, and is suitable for GPUs of different manufacturers and models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847809B_ABST
    Figure CN119847809B_ABST
Patent Text Reader

Abstract

The present application discloses a fault diagnosis method, device, medium and program product, which relates to the field of computer technology, including obtaining time series data of computer components; time series data refers to a series of ordered data points generated and / or collected by computer components within a set time interval; extracting at least one feature reflecting the fault state of the computer component from the time series data to obtain a time series feature; performing fault diagnosis based on the time series feature through multiple pre-built models to obtain multiple output results; the output results include diagnostic results and diagnostic values corresponding to the diagnostic results, and the diagnostic values are quantitative representations of the diagnostic results generated by the model; and fusing the multiple output results to determine the final diagnostic result. By extracting time series features representing different fault states and using multiple models to make comprehensive diagnostic decisions, the technical problems of long diagnostic time and low accuracy are solved, and the technical effect of efficient and accurate fault diagnosis is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a fault diagnosis method, device, medium, and program product. Background Art

[0002] With the widespread application of artificial intelligence in various fields, the demand for servers has increased significantly. Servers consist of multiple components, such as central processing units (CPUs), graphics processing units (GPUs), and hard drives, and the demand for each of these components has also increased dramatically. However, taking GPUs as an example, manufacturing defects may occur during the GPU manufacturing process, or the electronic components may age due to long-term use, resulting in poor heat dissipation, unstable power supplies, driver incompatibilities, firmware vulnerabilities, resource scheduling errors, and prolonged high-load operation. These problems can all lead to GPU failures.

[0003] GPU failures can cause image anomalies, system crashes, computational errors, and transient downtime. Failure to effectively identify and diagnose a faulty GPU can lead to severe server failures, significantly impacting both enterprise production and user experience. Therefore, there is an urgent need for efficient and accurate fault diagnosis methods for computer components to ensure server stability and reliability, and guarantee a positive user experience. Summary of the Invention

[0004] The present application provides a fault diagnosis method, device, medium and program product to at least solve the problem of long fault diagnosis time and low accuracy in related technologies.

[0005] The present application provides a method for diagnosing a fault of a computer component, comprising:

[0006] Obtaining time series data of a computer component; wherein time series data refers to a series of ordered data points generated and / or collected by the computer component within a set time interval;

[0007] Extracting at least one feature reflecting a fault state of a computer component from the time series data to obtain a time series feature;

[0008] Fault diagnosis is performed based on time series features using multiple pre-built models to obtain multiple output results. The output results include the diagnosis result and the corresponding diagnostic value. The diagnostic value is a quantitative representation of the diagnosis result generated by the model.

[0009] Multiple output results are fused to determine the final diagnosis result.

[0010] The present application also provides a computer component fault diagnosis device, comprising:

[0011] An acquisition module, configured to acquire time series data of a computer component; wherein time series data refers to a series of ordered data points generated and / or collected by the computer component within a set time interval;

[0012] A time series feature extraction module is used to extract at least one feature reflecting the fault state of the computer component from the time series data to obtain a time series feature;

[0013] A fault diagnosis module is used to perform fault diagnosis based on time series features using multiple pre-built models to obtain multiple output results. The output results include diagnostic results and corresponding diagnostic values, which are quantitative representations of the diagnostic results generated by the model.

[0014] The fusion output module is used to fuse multiple output results to determine the final diagnosis result.

[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned computer component fault diagnosis methods when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned computer component fault diagnosis methods are implemented.

[0017] The present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned computer component fault diagnosis methods when executed by a processor.

[0018] Through this application, timing features that can represent different fault states of components are extracted. Comprehensive diagnostic decisions are made based on timing features through multiple models to obtain accurate fault diagnosis results. This can solve the technical problems of long diagnosis time and low accuracy, achieve the technical effect of efficient and accurate fault diagnosis, and further ensure the stability and reliability of the server, guaranteeing a good user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1A flowchart of a fault diagnosis method provided in an embodiment of the present application;

[0021] Figure 2 A technical framework diagram of a fault diagnosis method provided in an embodiment of the present application;

[0022] Figure 3 A technical framework diagram of a time domain feature extraction method provided in an embodiment of the present application;

[0023] Figure 4 A technical framework diagram of a frequency domain feature extraction method provided in an embodiment of the present application;

[0024] Figure 5 A flowchart of a fault diagnosis method provided in an embodiment of the present application;

[0025] Figure 6 A schematic diagram of the structure of a residual module provided in an embodiment of the present application;

[0026] Figure 7 A schematic diagram of the structure of a temporal convolutional network provided in an embodiment of the present application;

[0027] Figure 8 A schematic diagram of the structure of a long short-term memory neural network provided in an embodiment of the present application;

[0028] Figure 9 A schematic diagram of the structure of a fault diagnosis device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0030] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the computer component fault diagnosis method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0033] Currently, GPUs may experience the following problems during use due to various factors, including hardware design, software drivers, thermal management, power supply stability, and the complexity of computing loads:

[0034] 1) Since GPUs are based on advanced processes (such as 5nm and 4nm), high-density transistor integration is susceptible to microscopic defects. Electromigration may occur after long-term use, resulting in circuit disconnection or short circuit.

[0035] 2) Aging of metal-oxide-semiconductor field-effect transistors (MOSFETs) or capacitors in multi-phase power supply designs can cause excessive voltage ripple, leading to unstable power supply to the GPU core or video memory, resulting in calculation errors or momentary system downtime.

[0036] 3) The thermal design power (TDP) of high-performance GPUs can reach over 400W, and the heat flux density exceeds the limits of traditional cooling solutions. Hot spots (such as SM clusters) are prone to local overheating, triggering thermal throttling (TDP) or even hardware fuses.

[0037] 4) Resource competition in parallel computing tasks (such as register bank conflicts and shared memory overflows) may cause undefined behavior. Insufficient exception handling mechanisms at the driver layer can lead to GPU context loss (GPU hang).

[0038] 5) Very long instruction streams (such as Monte Carlo simulations) may trigger the GPU watchdog timer (WatchdogTimer) to misjudge it as a hardware hang.

[0039] 6) Improper PCIe (Peripheral Component Interconnect Express) lane splitting in multi-GPU systems (e.g., misconfiguring x8 / x8 mode as x16 / x0) can lead to bandwidth contention and Transaction Layer Packet (TLP) transmission errors.

[0040] 7) GPUs that support ECC use Hamming codes to correct single-bit errors and record multi-bit error rates, but the accumulation of uncorrected errors may indicate problems such as memory module failure.

[0041] In summary, with the successive upgrades of GPUs, new models of GPUs may have other unknown faults, and the diagnostic coverage of complex GPU faults is not high. In order to avoid abnormal server operation due to GPU failures, accurate and efficient diagnosis of GPU failures will be an issue that needs to be solved urgently.

[0042] An embodiment of the present application provides a computer component fault diagnosis method, and the method is described in detail in conjunction with the execution flow of the computer component fault diagnosis method. Figure 1 A schematic diagram of a fault diagnosis method provided in an embodiment of the present application, specifically including the following steps: Figure 1 The following steps are shown:

[0043] For example, see Figure 2 , Figure 2 This is a technical framework diagram of a fault diagnosis method provided in an embodiment of the present application. The fault diagnosis method is executed by a fault diagnosis system. Figure 2 This can also be understood as a framework diagram of a fault diagnosis system. The fault diagnosis system includes a log collection module, a time series feature extraction module, a multi-model decision engine, and a fusion decision module. The log collection module collects log data from computer components and their associated components. Computer components can be GPUs and / or hard drives. The following embodiments describe GPU fault diagnosis in detail. In the hard drive fault diagnosis scenario, time series features can be established based on the hard drive's Smart parameters, temperature, voltage, and other information. A multi-model decision algorithm and a multi-model decision engine are then used to comprehensively diagnose whether the hard drive is faulty. The time series feature extraction module extracts features from the collected log data to obtain time series features. Time series features include time domain features, frequency domain features, and event sequence features. The detailed feature extraction process is described in the following embodiments. The multi-model decision engine generates multiple diagnostic results based on the same time series features using multiple models. These models include a short-term feature model, a long-term dependency model, and a knowledge rule engine (also known as a knowledge rule inference model). The fusion decision module fuses the diagnostic results from these multiple models and outputs the final diagnostic result. This is explained in detail in the following embodiments.

[0044] Troubleshooting methods include:

[0045] S101. Obtain timing data of computer components.

[0046] Here, time series data refers to a series of ordered data points generated and / or collected by computer components within a set time interval.

[0047] It can be understood that time series data can be log data, which contains information such as GPU temperature, voltage, power consumption, memory load rate, ECC (Error Correction Code) error count, PCIe (Peripheral Component Interconnect Express) transmission error rate, and PCIe bandwidth. Time series data can be understood as all log data generated within a certain time period. The set time interval refers to the time unit for data collection, which can be seconds, minutes, hours, days, etc. The collected data points are arranged in the order of their collection time. Each data point includes a timestamp and the corresponding measurement value, such as the temperature collected at a specific moment.

[0048] It is understandable that the sampling frequency can also be set according to the GPU load. For example, when the GPU is under low load, the sampling frequency is 1HZ. In this case, due to the low GPU activity, the lower sampling frequency is sufficient to capture the main performance characteristics without losing important information. At the same time, it also helps to reduce the occupation of system resources, such as reducing CPU usage and reducing memory consumption. When the GPU is under high load, the sampling frequency is 10HZ. This situation means that there are a lot of computing tasks in progress, its utilization, temperature changes, and possible bottlenecks, so the sampling frequency can be increased.

[0049] As you can understand, time series data is collected using a sliding time window within a set time interval, with a series of data points collected in each window. For example, if the set time interval is 10 seconds, the first window is 0-5 seconds, the second window is 1-6 seconds, the third window is 2-7 seconds, and so on, completing the collection of time series data.

[0050] S102: Extract at least one feature reflecting the failure state of the computer component from the time series data to obtain a time series feature.

[0051] It is understandable that, based on the above S101, the information in the log data that can reflect the GPU fault status is continuous, periodic, and discrete. In order to make full use of the log data to more comprehensively diagnose GPU faults, the useful information in the log data needs to be converted into time series features. Among them, time series features include time domain features, frequency domain features, and event sequence features. Extracting multiple features can more comprehensively and accurately characterize the fault information. Specifically, the time domain feature is to calculate the arithmetic mean, variance, kurtosis, and adjacent window differences of all data points in a sliding time window of a fixed or dynamic length. The frequency domain feature is to extract the energy ratio of the main frequency component through FFT to identify periodic anomalies, such as the regular temperature rise caused by the vibration of the cooling fan. The event sequence feature is to encode discrete events (such as ECC error bursts) into a timestamp sequence. It is understandable that the extraction order of the three features is not limited. They can be extracted simultaneously or sequentially.

[0052] Among them, the time series features include time domain features.

[0053] Optionally, extracting at least one feature reflecting the fault state of a computer component from the time series data to obtain a time series feature can be specifically achieved through the following steps:

[0054] Calculate the mean and variance of the time series data; calculate the kurtosis based on the time series data, mean and variance; where the kurtosis reflects the data distribution of the time series data; obtain the mean and variance of the previous data; calculate the mean difference and variance ratio based on the mean and variance of the time series data and the mean and variance of the previous data; where the time domain features include the mean, variance, kurtosis, mean difference and variance ratio of the time series data.

[0055] It is understandable that the mean and variance of the data points within each sliding time window (hereinafter referred to as the window) are calculated. The mean reflects the average level of the data points within the window. For example, the continuous increase in the mean of the GPU core temperature within the window may indicate a decrease in heat dissipation efficiency or load anomaly. The mean calculation formula is shown in Formula (1). The variance measures the degree of dispersion of the data points within the window from the mean and can capture the volatility of the data points. For example, a sudden increase in voltage variance may be caused by an unstable power module or a momentary short circuit. The contrast calculation formula is shown in Formula (2). Kurtosis can describe the steepness of the distribution of data points within the window, reflect the probability of extreme values, and identify extreme events. For example, an abnormal kurtosis of the memory error rate may indicate a hardware failure (such as damage to the memory particles). The kurtosis calculation formula is shown in Formula (3). The mean and variance of the previous data refer to the mean and variance corresponding to the previous window, or the mean and variance of the time series data collected within the previous set time period. The following embodiment uses the mean and variance corresponding to the previous window and the mean and variance corresponding to the current window in the same time series data as an example to explain in detail. The mean and variance of the current window and the previous window can be calculated by formula (1) and formula (2). The mean difference and variance ratio calculated based on the mean and variance of adjacent windows (current window and previous window) can be understood as the adjacent window difference (hereinafter referred to as difference). The difference can quantify the difference in statistical features between adjacent windows and is used to detect state mutations or gradual trends and capture state migration. The difference is mainly reflected in two aspects: mean difference and variance ratio. The mean difference represents the absolute difference or relative change rate of the adjacent window means. For example, if the voltage mean difference of adjacent windows exceeds a threshold (such as 10%), it may trigger a power module fault warning. The variance ratio represents the ratio of the variances of adjacent windows and detects fluctuation changes.

[0056] Formula (1)

[0057] Where, is the mean, N is the window size, is the i-th data point in the window.

[0058] Formula (2)

[0059] Where, is the variance.

[0060] Formula (3)

[0061] Where, is the kurtosis, where Indicates that the distribution of data points is steeper than the normal distribution (peaked and fat-tailed), and there may be sudden anomalies. Indicates that the data points are distributed relatively flatly and the fluctuations are gentle.

[0062] , Formula (4)

[0063] Where, is the mean difference, VR is the variance ratio, and are the mean and variance of the current time window, and are the mean and variance of the previous window.

[0064] For example, see Figure 3 , Figure 3 This is a technical framework diagram of a time domain feature extraction method provided in an embodiment of the present application. After completing the collection of log data, the time domain features of the log data are extracted. The time domain features include the mean, variance, kurtosis and adjacent window difference calculated for each window.

[0065] The time series features include frequency domain features, the set time interval includes at least one time window, and the time series data includes multiple data obtained in at least one time window.

[0066] Optionally, extracting at least one feature reflecting the fault state of a computer component from the time series data to obtain a time series feature can be specifically achieved through the following steps:

[0067] Frequency domain conversion is performed on multiple data obtained from the time window to obtain a complex spectrum corresponding to the time window; the main frequency component of the complex spectrum is extracted to obtain the main frequency energy ratio corresponding to the time window; the mean and standard deviation of the main frequency energy ratio corresponding to at least one time window are counted, and the abnormal threshold range is calculated based on the mean, standard deviation and preset threshold; the target main frequency energy ratio within the abnormal threshold range is determined as the frequency domain feature.

[0068] It is understandable that both time domain features and frequency domain features are extracted for the data points within each window. The steps for extracting frequency domain features for a certain window are as follows: perform data preprocessing on the log data. Data preprocessing includes signal acquisition, de-operation processing, and segmented windowing processing. If the log data is obtained, then obtain time series data from the log data, such as temperature, voltage, load rate, etc., while ensuring the sampling frequency to obtain the sampled signal. Subsequently, use a sliding average filter or wavelet denoising to eliminate high-frequency noise in the sampled signal to avoid interference with frequency domain analysis. Split the continuous signal into fixed lengths and apply a Hamming window to reduce spectrum leakage. After completing the preprocessing of the log data, perform a fast Fourier transform (FFT) on the preprocessed data. That is, perform FFT on the data points within the window and perform frequency domain conversion to obtain a complex spectrum. The frequency domain conversion is shown in formula (5). Subsequently, extract the main frequency component of the complex spectrum after FFT conversion to obtain the main frequency energy ratio corresponding to the window. The main frequency component extraction is shown in formula (6). Subsequently, the mean and standard deviation of the main frequency energy ratios corresponding to all time windows are calculated, that is, the mean and variance of all main frequency energy ratios are calculated. The abnormal threshold range is calculated based on the mean, standard deviation, and preset threshold. The abnormal threshold range is specifically shown in Formula (7). Subsequently, the target main frequency energy ratio within the abnormal threshold range is determined as a frequency domain feature. That is, if the main frequency energy ratio value calculated in real time exceeds the abnormal threshold range, an alarm is triggered. For example, the periodic freezing of the GPU cooling fan can be detected or periodic abnormal phenomena caused by mechanical wear, circuit aging, etc. can be identified.

[0069] Formula (5)

[0070] Where, is the complex spectrum, is the sampling signal after data preprocessing.

[0071] , Formula (6)

[0072] Where, is a single frequency component, Recorded as amplitude , is the total energy, is the main frequency energy ratio, and M represents the window.

[0073] Formula (7)

[0074] Where T is the abnormal threshold range, and are the mean and variance of the main frequency energy ratio, 3 is the preset threshold, and the preset threshold can be determined according to user needs.

[0075] For example, see Figure 4 , Figure 4 This is a technical framework diagram of a frequency domain feature extraction method provided in an embodiment of the present application. After completing log data collection, the frequency domain features of the log data are extracted. The specific implementation steps are as follows: data preprocessing is performed on the log data, which includes at least one processing method such as signal sampling, denoising, and segmented windowing; the preprocessed sampled signal is subjected to fast Fourier transform, main frequency component extraction, and abnormality determination logic processing. The specific implementation steps are referred to in the above embodiment and are not repeated here.

[0076] Among them, time series data includes discrete data and continuous data reflecting discrete events, and time series features include event sequence features.

[0077] Optionally, extracting at least one feature reflecting the fault state of a computer component from the time series data to obtain a time series feature can be specifically achieved through the following steps:

[0078] Determine the event type of discrete events; where discrete events refer to non-continuous events with timestamps that occur at a certain time point within a set time interval; encode discrete events into timestamp sequences according to the event type to obtain discrete features; concatenate discrete features and continuous features in the feature dimension to obtain event sequence features; where continuous features are extracted from continuous data.

[0079] Discrete data is data that can only take on specific values, typically countable values. For example, Hypertext Transfer Protocol (HTTP) status codes (e.g., 200, 404, 500) in log data. Continuous data is a data type that can take on arbitrary values and typically involves metrics or counts, such as response time and CPU utilization. Event sequence features encode discrete events (e.g., ECC error bursts) as a sequence of timestamps and combine them with continuous data to construct multimodal (multi-model) inputs. Discrete events are non-continuous events that occur suddenly at a specific point in time and have a clear timestamp. Examples include ECC errors, PCIe retransmission requests, and driver crash logs. ECC errors refer to sudden errors in the graphics memory error correction code (e.g., single-event upsets). PCIe retransmission requests refer to retransmission events triggered by data transmission errors. Driver crash logs record the time points when the GPU driver abnormally exits. The event type of a discrete event is determined based on its nature and purpose, or can be customized based on user needs. Subsequently, different encoding methods are used according to the event type to encode discrete events into timestamp sequences and obtain event sequence features. Finally, the encoded discrete event sequence (event sequence features) is fused with continuous features as a unified input. Continuous features are extracted from continuous data (temperature, voltage, etc.) collected by continuous sensors. Event sequence features can also be understood as discrete features. In the process of fusing discrete features and continuous features, time alignment and feature splicing are also required. Among them, time alignment refers to the division of continuous data and discrete data in the same window, for example, the mean and variance of temperature and voltage sampled in real time in the same window (continuous features) and the statistical ECC error count (discrete features). Feature splicing refers to splicing continuous features with discrete features in the feature dimension. For example, continuous features include temperature mean = 75 , variance = 2.1, power consumption = 150W (dimension = 3), discrete features include ECC error count = 4, time difference mean = 10s (dimension = 2), continuous features and discrete features are concatenated into an input vector of dimension 5, that is, 75, 2.1, 150, 4, 10.

[0080] Optionally, based on the event type, discrete events are encoded into a timestamp sequence to obtain discrete features. This can be achieved through the following steps:

[0081] If the event type is a single event, binary coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; wherein binary coding refers to marking whether the event occurs at each sampling time point; or, if the event type is a high-frequency event, count coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; wherein, a high-frequency event refers to an event that occurs more than a first set threshold number of times in a first time period, and count coding refers to counting the number of occurrences of the event in each time window; or, if the event type is an intermittent event, time difference coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; wherein, an intermittent event refers to an event that occurs discontinuously and has periodicity, and time difference coding refers to recording the time interval between adjacent events; or, if the event type is a complex event, time embedding coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; wherein, a complex event refers to an event that occurs with a period greater than a second set threshold and relies on multiple data for trend analysis, and time embedding coding refers to mapping the timestamp of the event into a continuous vector.

[0082] Discrete events can be classified into single events, high-frequency events, intermittent events, and complex events. A single event refers to a single type of event, such as a driver crash. A high-frequency event refers to an event that occurs more than a first threshold in a short period of time. These events are not sensitive to time intervals and have a high frequency of occurrence, such as multiple ECC errors within a short period of time. An intermittent event refers to an event that occurs discontinuously, with a periodic or discontinuous pattern, or a sudden change in pattern, such as an intermittent failure that changes from sparse to dense, such as a periodic cooling fan freeze. A complex event refers to an event with a periodicity greater than a second threshold and requires trend analysis based on multiple data sets. These events require long-term analysis, such as a periodic increase in error rate due to aging. After determining the event type, coarse-grained detection of single events uses binary coding, which marks whether the event occurred at each sampling time point. For high-frequency events, count coding is used, which counts the number of event occurrences within a sliding window. For intermittent events, time difference coding is used, which records the time intervals between adjacent events to reflect the rhythm of the event occurrence. Time embedding coding is used for complex events. Time embedding coding refers to mapping timestamps into continuous vectors to capture periodicity / trend characteristics.

[0083] S103. Perform fault diagnosis based on time series features using multiple pre-built models to obtain multiple output results.

[0084] The output results include the diagnosis results and the diagnosis values corresponding to the diagnosis results. The diagnosis values are quantitative representations of the diagnosis results generated by the model.

[0085] As can be understood, based on the above S103, multiple pre-built models (hereinafter referred to as "multi-models") are obtained. Different models are used to diagnose whether a fault exists in the time series features according to different types of time series features. Each model will produce an output result, and the input of the multi-model is all time series features. Each of the multiple output results of the multi-model includes a diagnosis result and a corresponding diagnostic value. The diagnostic value is a quantitative representation of the diagnosis result generated by the model, such as a probability value and confidence level.

[0086] Among them, the various models include short-term feature models, long-term dependency models and knowledge rule reasoning models.

[0087] Understandably, a model that detects sudden anomalies in a relatively short period of time is called a short-term feature model, such as a sudden anomaly caused by a voltage drop. A model that captures performance degradation trends across hours is called a long-term dependency model, such as the temperature drift caused by silicone grease aging. A model that requires the injection of prior rules to improve interpretability is called a knowledge rule engine (knowledge rule reasoning model). For example, "a video memory error rate > 1e-5 accompanied by PCIe retransmission" is judged as a video memory failure.

[0088] S104: Fusing multiple output results to determine a final diagnosis result.

[0089] It is understandable that, based on the above S103, multiple diagnosis results output by multiple models such as the short-term feature model, the long-term dependency model, and the knowledge rule reasoning model are further integrated to output the final diagnosis result.

[0090] Optionally, multiple output results are fused to determine the final diagnosis result, which can be achieved through the following steps:

[0091] Determine whether the multiple output results include only one diagnostic result; if the multiple output results include only one first diagnostic result, determine the first diagnostic result as the final diagnostic result; or, if the multiple output results include multiple diagnostic results, compare and analyze the second diagnostic result in the third output result, the third diagnostic result in the first output result, and / or the fourth diagnostic result in the second output result to determine the final diagnostic result.

[0092] It is understandable that it is necessary to judge whether the multiple diagnostic results are consistent, that is, whether there is only one diagnostic result among the multiple diagnostic results. For example, the multiple diagnostic results all show that the GPU temperature is too high. In this case, the first diagnostic result is directly used as the final diagnostic result, wherein the first diagnostic result is the only diagnostic result that exists among the multiple diagnostic results. If the multiple output results include at least two diagnostic results, the diagnostic results output by each model are compared and analyzed to determine the final diagnostic result, wherein the third output result output by the knowledge rule reasoning model includes the second diagnostic result. The first output result output by the short-term feature model includes the third diagnostic result. The second output result output by the long-term dependency model includes the fourth diagnostic result for comparison and analysis, and there are at least two diagnostic results among the second diagnostic result, the third diagnostic result and the fourth diagnostic result.

[0093] Optionally, the second diagnostic result in the third output result, the third diagnostic result in the first output result, and / or the fourth diagnostic result in the second output result are compared and analyzed to determine a final diagnostic result, which can be specifically achieved by the following steps:

[0094] When the second diagnostic result is determined based on a rule case, the second diagnostic result is determined as the final diagnostic result; when the second diagnostic result is not determined based on a rule case and the second diagnostic result is different from the third diagnostic result and the fourth diagnostic result, the target confidence in the third output result is weighted averaged with the first probability value in the first output result and the second probability value in the second output result, and the second diagnostic result, the third diagnostic result and the fourth diagnostic result are determined as the final diagnostic results; when the second diagnostic result is not determined based on a rule case and the second diagnostic result is the same as the third diagnostic result or the fourth diagnostic result, the target confidence is compared with the first probability value or the second probability value, and the diagnostic result corresponding to the maximum value is determined as the final diagnostic result.

[0095] It is understandable that in the process of diagnostic result fusion, if there is a clear rule case in the knowledge rule reasoning model, that is, the second diagnostic result is determined based on a real case, then the second diagnostic result is directly output, and the diagnostic results of other models are no longer output; if the second diagnostic result is inconsistent with the third diagnostic result and / or the fourth diagnostic result, the target confidence and the first probability value in the first output result and / or the second probability value in the second output result are weighted averaged, and then the diagnostic results are output separately, that is, for the same fault diagnosis result, each model must perform a weighted average of the confidence and probability values when outputting the diagnostic value. The final diagnostic result includes the second diagnostic result, the third diagnostic result, and / or the fourth diagnostic result, as well as the diagnostic value after weighted average of various diagnostic results. For example, the second diagnostic result is diagnosis A with a confidence level of a, the third diagnostic result is diagnosis B with a probability value of b, and the third diagnostic result is C with a probability value of c. In this case, the three diagnostic results are all different. For the diagnostic result of diagnosis A, the probability values of the other two models are obtained, and the confidence level a and the probability values of the other two models are weighted averaged to obtain the diagnostic value A corresponding to diagnosis A. Diagnosis A and diagnostic value A are output, and so on to calculate the diagnostic values of other diagnostic results. If the second diagnostic result is consistent with the diagnostic result in the model, the diagnostic result with the higher confidence level (probability value) is output. For example, the second diagnostic result is diagnosis A1 with a confidence level of a, the third diagnostic result is diagnosis B with a probability value of b, and the third diagnostic result is A2 with a probability value of c. In this case, if the confidence level a is greater than the probability value c, diagnosis A1 is determined as the final diagnostic result. If the confidence level a is less than the probability value c, diagnosis A2 is determined as the final diagnostic result. Other possible fusion decision-making methods are not described in detail.

[0096] The fault diagnosis method provided in this application converts useful information in log data into time-series features, such as time-domain features, frequency-domain features, and event sequence features. This method is capable of diagnosing not only short-term faults on the millisecond level, but also long-term faults spanning months. Secondly, a sliding time window is used in time-domain feature extraction to calculate the arithmetic mean, variance, kurtosis, and adjacent window differences of all data points within the sliding window, enabling more accurate fault diagnosis using continuous data. In frequency-domain feature extraction, the energy ratio of the main frequency component is extracted using FFT to identify periodic anomalies. In event sequence feature extraction, discrete event coding and multimodal input fusion strategies are used to clarify discontinuous events. By extracting multiple time-series features, the method is highly versatile and can diagnose faults for GPUs of different manufacturers and models. Furthermore, when fusion decisions are made on multiple output results, a learning model and a knowledge-based rule inference model are used to collaboratively determine the final diagnostic result, achieving high-precision, interpretable fault diagnosis and component-level localization, demonstrating strong practicality.

[0097] Based on the above embodiments, Figure 5 A flow chart of a fault diagnosis method provided in an embodiment of the present application is provided. Optionally, fault diagnosis is performed based on time series features using multiple pre-built models to obtain multiple output results, including: Figure 5 The following steps are shown:

[0098] S501: Using the time series feature as the input of a short-term feature model to diagnose a sudden abnormality through the short-term feature model to obtain a first output result.

[0099] Among them, the short-term feature model adopts the temporal convolutional network.

[0100] Understandably, the short-term feature model uses a lightweight temporal convolutional network (TCN) to detect sudden anomalies. This is because the TCN has the advantages of millisecond-level response requirements (e.g., voltage sags must be reported within 10ms), lightweight (model parameter count <1M, inference memory usage <50MB), and accurate capture of short-term mutations (e.g., temperature spikes, sudden increases in ECC error rates). It also has the characteristic of suppressing noise interference. The input of the short-term feature model is the time series feature, and the output is the diagnostic value and the corresponding diagnostic result. It is used to diagnose sudden abnormal faults reflected by the time series feature. The TCN is shown in Formula (8).

[0101] Formula (8)

[0102] Where, It is the diagnostic value output by the temporal convolutional network, that is, the probability value. is the filter, is the time series feature, d is The expansion factor at .

[0103] Optionally, the time series features are used as input to the short-term feature model to diagnose sudden abnormalities using the short-term feature model to obtain a first output result. This can be achieved by the following steps:

[0104] Sudden anomalies are diagnosed by adjusting the value of a dilation factor in a temporal convolutional network to obtain a first output result; wherein the dilation factor is used to enable the temporal convolutional network to capture pattern changes in temporal features at different time scales.

[0105] Understandably, in the process of diagnosing sudden abnormal faults through short-term feature models, it is also possible to diagnose whether the GPU is faulty by adjusting the value of the dilation factor or by detecting long-term time series features. The dilation factor is a parameter in the dilated convolution that determines the size of the interval between elements in the convolution kernel. By adjusting the dilation factor, the receptive field of the convolution operation can be controlled, that is, the time span of the input data that the network can "see". In other words, the dilation factor is mainly used to adjust the time period of the time series features, such as detecting sudden abnormalities from the first 1 second to the first 10 minutes in the time series features. For example, in the server monitoring log, if the CPU usage suddenly soars to 90% at a certain moment, while it usually remains at around 20%, this phenomenon may be regarded as a sudden abnormality.

[0106] Among them, the temporal convolutional network includes a residual module, which includes multiple convolutional layers, convolution kernel weights, activation functions, and connection methods between network layers.

[0107] Optionally, before diagnosing the sudden abnormality using the short-term feature model and obtaining the first output result, the method further includes:

[0108] Set the relationship between the expansion factor and the number of convolutional layers; the association means that the expansion factor increases with the number of network layers; decompose the weight into two independent parameters, direction and amplitude, to constrain the weight distribution; set the gated activation function; set the connection mode to skip connection; skip connection refers to a network structure that adds the input to the output.

[0109] Understandably, in TCNs, deep networks are prone to vanishing or exploding gradients. To address this, the residual module needs to be redesigned. Specifically, the relationship between the expansion factor and the number of convolutional layers is set. For example, the expansion factor is set to increase exponentially with the number of network layers, d×2, and the convolution kernel size is set to K=3. This setting expands the causal convolutional layer. By decomposing the weights into two independent parameters, direction and amplitude, the weight distribution is directly constrained, thereby accelerating model convergence and avoiding small-batch noise. A gated activation function is set to enhance nonlinear expression capabilities while suppressing irrelevant features. This means that the features to be retained, amplified, or suppressed can be customized. Skip connections are set to address the problem of small gradients. A skip connection is a network structure that adds input to output.

[0110] For example, Figure 6 This is a schematic diagram of the structure of a residual module provided in an embodiment of the present application. The residual module involves configurations such as multiple convolutional layers, weights, activation functions, and the connection between network layers. Specifically, the residual module performs settings such as convolutional layer expansion, weight normalization, gated activation, and skip connections to eliminate the problem of vanishing or exploding gradients. For specific configuration instructions, please refer to the above embodiment and will not be repeated here.

[0111] For example, Figure 7 This is a structural diagram of a temporal convolutional network provided in an embodiment of the present application. In the temporal convolutional network, temporal features pass through the input layer and enter the 3-layer residual module ( Figure 7 The hidden layer in the convolutional layer is the residual module), and finally reaches the output layer. In the first residual module, d = 2, K = 3 are set. In the second residual module, d = 4, K = 3 are set. In the third residual module, d = 8, K = 3 are set. The convolutional layer is expanded by increasing the dilation factor. The residual module can be added according to user needs. The setting method of the residual module is not described in detail here. The final fault diagnosis result is determined in the output layer. Figure 7 It can be inferred that after processing by the time series convolutional network, irrelevant feature information in the time series features can be filtered out layer by layer, and the feature information that can reflect GPU related information can be amplified, and finally the diagnosis results and the probability of failure can be output.

[0112] S502: Using the time series feature as an input of a long-term dependency model to diagnose abnormal performance trends through the long-term dependency model to obtain a second output result.

[0113] Among them, the long-term dependency model adopts long short-term memory neural network.

[0114] Understandably, the long-term dependency model uses long short-term memory (LSTM) neural networks because LSTM networks can capture time series features that exhibit performance degradation trends over hours. For example, the aging of silicone grease reduces its thermal conductivity, causing the GPU core temperature to slowly increase over time; the aging of capacitors causes the power supply module's capacitance to decrease, leading to increased voltage fluctuations; and the wear of fan bearings reduces heat dissipation efficiency, requiring a 5%-10% increase in fan speed under the same load. LSTM uses logic control within gate units to determine whether to update or discard data. This overcomes the drawbacks of excessive weight influence and the susceptibility to vanishing and exploding gradients, enabling better and faster network convergence and effectively improving prediction accuracy. The input of the long-term dependency model is also time series features, and the output is a diagnostic value and the corresponding diagnostic result, which is also a probability value.

[0115] It is understandable that the order of reasoning of the three models is not limited.

[0116] For example, Figure 8This is a schematic diagram of the structure of a long short-term memory neural network provided in an embodiment of the present application. LSTM has three gates: a forget gate, an input gate, and an output gate, which determine whether information is remembered or forgotten at each moment. The input gate determines how much new information is added at each moment, the forget gate controls whether information is forgotten at each moment, and the output gate determines whether any information is output at each moment. The calculation formula for the forget gate is shown in Formula (9).

[0117] Formula (9)

[0118] Where, is the hidden state at the previous moment, is the current input, that is, the time series feature, and are the weight and bias of the forget gate.

[0119] The calculation formula of the input gate is shown in formula (10).

[0120] , Formula (10)

[0121] Where, is a candidate memory, and are the weights and biases of the input gate, and are the weights and biases of the candidate memories.

[0122] The calculation formula for memory update is shown in formula (11).

[0123] Formula (11)

[0124] Where, is the candidate memory of the previous moment, It is the candidate memory after updating at the current moment.

[0125] The calculation formula of the output gate is shown in formula (12).

[0126] , Formula (12)

[0127] Where, is the final hidden state.

[0128] Optionally, in order to more accurately predict the fault trend over a long period of time in the time series features and output reasonable fault handling suggestions in the fault diagnosis results, it is necessary to add potential fitting loss to the predicted diagnosis results to correct the predicted trend so that the given diagnosis results are more realistic and reliable. The calculation formula for the trend term is shown in Formula (13).

[0129] Formula (13)

[0130] Where, is a parameter that controls the smoothness of the trend.

[0131] Optionally, the time series features are used as input to the long-term dependency model to diagnose abnormal performance trends through the long-term dependency model and obtain a second output result. This can be achieved through the following steps:

[0132] Taking the time series features as the input of the long-term dependency model, the hidden states at different moments calculated by the output gate in the long short-term memory neural network are obtained; wherein the output gate is used to determine whether there is information output at each moment; based on the hidden state at the current moment, the hidden state at the previous moment and the set parameters, the trend of the time series features is calculated to obtain the predicted state at the current moment; wherein the set parameters refer to the parameters that control the smoothness of the trend; the loss is calculated based on the hidden state at the current moment and the predicted state at the current moment; the loss is compared with the set range to determine the diagnosis result to obtain a second output result.

[0133] It is understandable that the time series features are used as the input of the long-term dependency model, and the hidden states at different times are obtained through the input gate, forget gate and output gate. The hidden state is the output of the output gate, which is the output of formula (12). Then, the hidden state at the current moment, the hidden state at the previous moment, the time series features and the set parameters are obtained, the trend of the time series features is calculated, and the predicted state at the current moment is obtained. The predicted state is also the predicted trend output by formula (13) , the setting parameter refers to the formula (13) Subsequently, the loss is calculated based on the predicted state and the hidden state at the current moment, as shown in Formula (14). It is determined whether the loss at the current moment is within the set range, which can be understood as the acceptable range, and the diagnosis result is determined to obtain the second output result, which is the output result of the long-term dependency model.

[0134] Formula (14)

[0135] Where, For loss, 720 means that there are 720 hours in a month, and the specific time can be determined according to user needs.

[0136] For example, assuming that the aging of silicone grease causes the GPU temperature to rise by 0.3 degrees per month, The value of is 14400, the actual GPU temperature trend (hidden state) is [75.0, 75.1, 75.2, ..., 77.4], the predicted GPU temperature trend (predicted state) is [75.0, 75.05, 75.15, ..., 77.3], and the calculated loss is 0.02. If the loss is within an acceptable range, the diagnosis result can be obtained that "the GPU temperature rises by 0.3 degrees per month due to aging of silicone grease."

[0137] S503: Using the time series feature as input to the knowledge rule reasoning model to diagnose anomalies using a plurality of predefined rules in the knowledge rule reasoning model to obtain a third output result;

[0138] The multiple output results include a first output result, a second output result, and a third output result.

[0139] Understandably, the knowledge rule inference model contains multiple predefined rules, primarily derived from expert experience, historical failure analysis, and vendor documentation. Expert experience refers to failure modes summarized by hardware engineers (e.g., the correlation between memory error rates and PCIe retransmissions). Historical failure analysis involves mining frequently co-occurring event combinations from failure logs (e.g., high temperature and insufficient fan speed), which are based on real-world examples. Vendor documentation refers to the troubleshooting guides provided by GPU manufacturers (e.g., hardware issues corresponding to specific error codes). The knowledge rule inference model uses time series features as input and tertiary output as output, primarily through rule-based fault diagnosis.

[0140] As you can understand, the knowledge rule reasoning model supports logical judgments between multiple rules: AND, OR, NOT, and so on. For example, when diagnosing multiple pairs of timing features, if Rule 1 ("Video memory error rate > 1e-5") and Rule 2 ("PCIe retransmission rate > 0.1") are triggered simultaneously, a diagnosis of a video memory hardware fault can be generated. In the knowledge rule reasoning model, each rule also includes information such as rule change history, rule source, effective time, verification status, priority, and confidence level. The specific content of each rule is not limited.

[0141] Optionally, the time series features are used as input to a knowledge rule reasoning model to diagnose anomalies using multiple rules predefined in the knowledge rule reasoning model to obtain a third output result. This can be achieved through the following steps:

[0142] Match the temporal features with multiple predefined rules in the knowledge rule reasoning model to determine at least one target rule; obtain the rule relationship between the multiple rules; wherein the rule relationship reflects the degree of correlation between different rules; determine the calculation strategy of at least one target rule based on the rule relationship, and calculate the target confidence based on the calculation strategy; wherein the third output result includes the confidence and the diagnostic result corresponding to the target confidence.

[0143] As can be understood, the extracted time series features include information such as the mean, variance, kurtosis, and adjacent window differences within the window. These time series features are then matched against multiple rules to identify at least one target rule. To avoid the pitfalls of overestimating rule redundancy (if rules are strongly correlated, confidence is repeatedly calculated after integration, resulting in an inflated confidence) and underestimating complementary rules (if rules describe the same fault from different dimensions, the independence assumption can underestimate the true confidence) when multiple matched target rules point to the same conclusion, a classification fusion strategy based on rule relationships is adopted in the knowledge rule reasoning model to address these two issues. Specifically, the rule relationships between multiple predefined rules are obtained. Rule relationships reflect the degree of correlation between different rules. For example, the first and second rules are independent, the first and third rules are completely correlated, and the first and fourth rules are partially correlated. Subsequently, based on the rule relationships, the degree of correlation between the matched target rules is determined. Confidence is then calculated based on the corresponding calculation strategy, with different calculation strategies corresponding to different degrees of correlation.

[0144] The at least one target rule includes a first rule and a second rule. A calculation strategy for the at least one target rule is determined based on the rule relationship, and the target confidence is calculated based on the calculation strategy. This can be specifically achieved through the following steps:

[0145] If the rule relationship between the first rule and the second rule is independent of each other, the target confidence is calculated based on the first confidence of the first rule and the second confidence of the second rule; or, if the rule relationship is strongly correlated, the correlation coefficient of the first rule and the second rule is determined; and the target confidence is calculated based on the correlation coefficient, the first confidence and the second confidence, wherein strong correlation refers to a relationship in which the correlation coefficient between the rules is greater than a third set threshold; or, if the rule relationship is a complementary relationship, the complementary weight between the first rule and the second rule is determined; and the target confidence is calculated based on the complementary weight, the correlation coefficient, the first confidence and the second confidence.

[0146] It is understandable that the confidence calculation is described in detail by taking the first rule and the second rule in multiple target rules as an example. Specifically, if the rule relationship between the first rule and the second rule is independent of each other, that is, the two rules are completely unrelated, then the target confidence is calculated based on the confidence of each rule, where the confidence of the independent rule is shown in formula (15). If the rule relationship between the first rule and the second rule is strongly correlated, then the correlation coefficient of each rule is determined based on the rule relationship, mainly by adding the correlation coefficient in the rule. ,in, , if the first rule and the second rule are independent of each other, then ; If the first rule and the second rule are completely correlated, then ; If the first rule and the second rule have a certain correlation, then , where strong correlation means that the correlation coefficient between the rules is greater than the third set threshold. The third set threshold can be set according to user needs. For example, the third set threshold is 0.5. If the correlation coefficient between the first rule and the second rule is If the correlation coefficient is greater than 0.5, it means that the relationship between the first rule and the second rule is strongly correlated. If the correlation coefficient is greater than 0.5, it also indicates that the relationship between the first rule and the second rule is strongly correlated. If it is less than 0.5, it means that the rule relationship between the first rule and the second rule is not strongly correlated and can be defined as weakly correlated. After determining the correlation coefficient, the target confidence is calculated based on the correlation coefficient, the first confidence level, and the second confidence level. The calculation formula is shown in formula (16). If the rule relationship between the first rule and the second rule is a complementary relationship, that is, the two rules are complementary, in this case, the complementary weight between the first rule and the second rule is determined based on the rule relationship. The complementary weight is used to reduce the effective weight of the first rule when it is complementary to other rules, so as to avoid repeated calculation of the confidence level. After determining the correlation coefficient and the complementary weight, the target confidence is calculated based on the complementary weight, the correlation coefficient, the first confidence level, and the second confidence level. The specific calculation formula is shown in formula (17).

[0147] Formula (15)

[0148] Where, is the target confidence, is the confidence of the rule, where is the first confidence level, is the second confidence level.

[0149] Formula (16)

[0150] Where, is the correlation coefficient between rule i and rule j, where rule i is the first rule and rule j is the second rule.

[0151] Formula (17)

[0152] Where, is the complementary weight of the first rule.

[0153] It is understandable that in the knowledge rule reasoning model, the confidence of the final diagnosis result will be output by selecting an appropriate calculation strategy based on the association relationship of the diagnosis rules.

[0154] The fault diagnosis method provided by the embodiment of the present disclosure introduces weight normalization, gated activation, and skip connection in the residual module of the short-term feature model to solve the problem of gradient vanishing or exploding easily due to deep networks in TCN. In order to be able to more accurately predict the trend of time series features across long periods of time and output reasonable fault handling opinions in the fault diagnosis results, potential fitting loss is added to the prediction results output by the long-term dependent model to correct the prediction trend. Confidence, priority, correlation coefficient, and weight are added to the diagnostic rules in the knowledge rule reasoning model. If the diagnostic results of multiple diagnostic rules point to the same conclusion, the confidence of the diagnostic results is calculated using a classification fusion strategy based on rule relationships.

[0155] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0156] The embodiment of the present application also provides a computer component fault diagnosis device, Figure 9 This is a schematic diagram of the structure of a fault diagnosis device provided in an embodiment of the present application. The fault diagnosis device includes an acquisition module 901, a time series feature extraction module 902, a fault diagnosis module 903, and a fusion output module 904, wherein:

[0157] An acquisition module 901 is configured to acquire time series data of a computer component; wherein time series data refers to a series of ordered data points generated and / or collected by the computer component within a set time interval;

[0158] A time series feature extraction module 902 is configured to extract at least one feature reflecting a computer component failure state from the time series data to obtain a time series feature;

[0159] Fault diagnosis module 903, configured to perform fault diagnosis based on time series features using multiple pre-built models, and obtain multiple output results; wherein the output results include a diagnosis result and a corresponding diagnosis value of the diagnosis result, which is a quantitative representation of the diagnosis result generated by the model;

[0160] The fusion output module 904 is used to fuse multiple output results to determine the final diagnosis result.

[0161] Among them, the time series features include time domain features.

[0162] Optionally, the time series feature extraction module 902 is used to:

[0163] Calculate the mean and variance of time series data;

[0164] Calculate the kurtosis based on the time series data, mean, and variance; the kurtosis reflects the data distribution of the time series data;

[0165] Obtaining the mean and variance of previous data; wherein the previous data refers to the time series data of the computer component obtained before a set time interval;

[0166] Calculate the mean difference and variance ratio based on the mean and variance of the time series data and the mean and variance of the previous data;

[0167] Among them, the time domain features include the mean, variance, kurtosis, mean difference and variance ratio of time series data.

[0168] The time series features include frequency domain features, the set time interval includes at least one time window, and the time series data includes multiple data obtained in each time window.

[0169] Optionally, the time series feature extraction module 902 is used to:

[0170] Perform frequency domain conversion on the multiple data obtained in each time window to obtain the complex spectrum corresponding to each time window;

[0171] Extract the main frequency component of the complex spectrum and obtain the main frequency energy ratio corresponding to each time window;

[0172] Count the mean and standard deviation of the main frequency energy ratio corresponding to all time windows, and calculate the abnormal threshold range based on the mean, standard deviation and preset threshold;

[0173] The target main frequency energy ratio within the abnormal threshold range is determined as the frequency domain feature.

[0174] Among them, time series data includes discrete data and continuous data reflecting discrete events, and time series features include event sequence features.

[0175] Optionally, the time series feature extraction module 902 is used to:

[0176] Determine the event type of a discrete event; a discrete event is a non-continuous event with a timestamp that occurs at a certain point in time within a set time interval;

[0177] Encode discrete events into time stamp sequences according to event types to obtain discrete features;

[0178] The discrete features and continuous features are spliced in the feature dimension to obtain event sequence features; among them, the continuous features are extracted from continuous data.

[0179] Optionally, the time series feature extraction module 902 is used to:

[0180] If the event type is a single event, binary coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; binary coding means marking whether the event occurs at each sampling time point; or,

[0181] If the event type is a high-frequency event, count coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; wherein a high-frequency event refers to an event that occurs more than a first set threshold number of times within a first time period, and count coding refers to counting the number of occurrences of the event within each time window; or,

[0182] If the event type is an intermittent event, time difference coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; intermittent events refer to events that occur discontinuously and have periodicity, and time difference coding refers to recording the time interval between adjacent events; or,

[0183] If the event type is a complex event, time embedding coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; among them, a complex event refers to an event whose occurrence period is greater than the second set threshold and relies on multiple data for trend analysis, and time embedding coding refers to mapping the timestamp of the event into a continuous vector.

[0184] Among them, the various models include short-term feature models, long-term dependency models and knowledge rule reasoning models.

[0185] Optionally, the fault diagnosis module 903 is used to:

[0186] Using the time series features as input of the short-term feature model to diagnose sudden abnormalities through the short-term feature model to obtain a first output result;

[0187] Using the time series features as input to the long-term dependency model to diagnose abnormal performance trends through the long-term dependency model and obtain a second output result;

[0188] Using the time series features as input to the knowledge rule reasoning model to diagnose anomalies through a plurality of rules predefined in the knowledge rule reasoning model to obtain a third output result;

[0189] The multiple output results include a first output result, a second output result, and a third output result.

[0190] Among them, the short-term feature model adopts the temporal convolutional network.

[0191] Optionally, the fault diagnosis module 903 is used to:

[0192] Sudden anomalies are diagnosed by adjusting the value of a dilation factor in a temporal convolutional network to obtain a first output result; wherein the dilation factor is used to enable the temporal convolutional network to capture pattern changes in temporal features at different time scales.

[0193] Among them, the temporal convolutional network includes a residual module, which includes multiple convolutional layers, convolution kernel weights, activation functions, and connection methods between network layers.

[0194] Optionally, the fault diagnosis method 900 is further used to:

[0195] Set the relationship between the expansion factor and the number of convolutional layers; the relationship means that the expansion factor increases with the number of layers;

[0196] Decompose the weight into two independent parameters, direction and magnitude, to constrain the weight distribution;

[0197] Set the gate activation function;

[0198] Set the connection mode to skip connection; a skip connection is a network structure that adds input to output.

[0199] Among them, the long-term dependency model adopts long short-term memory neural network.

[0200] Optionally, the fault diagnosis module 903 is used to:

[0201] The time series features are used as the input of the long-term dependency model to obtain the hidden states at different moments calculated by the output gate of the long short-term memory neural network. The output gate is used to determine whether there is information output at each moment.

[0202] Based on the hidden state at the current moment, the hidden state at the previous moment, and the set parameters, the trend of the time series features is calculated to obtain the predicted state at the current moment; where the set parameters refer to the parameters that control the smoothness of the trend;

[0203] Calculate the loss based on the hidden state at the current moment and the predicted state at the current moment;

[0204] The loss is compared with a set range, and a diagnosis result is determined to obtain a second output result.

[0205] Optionally, the fault diagnosis module 903 is used to:

[0206] Match the temporal features with multiple predefined rules in the knowledge rule reasoning model to determine at least one target rule;

[0207] Obtaining the rule relationships between multiple rules; wherein the rule relationships reflect the degree of correlation between different rules;

[0208] A calculation strategy for at least one target rule is determined according to the rule relationship, and the target confidence is calculated according to the calculation strategy; wherein the third output result includes the confidence and the diagnosis result corresponding to the target confidence.

[0209] The at least one target rule includes a first rule and a second rule.

[0210] Optionally, the fault diagnosis module 903 is used to:

[0211] If the rule relationship between the first rule and the second rule is independent of each other, the target confidence is calculated according to the first confidence of the first rule and the second confidence of the second rule; or,

[0212] If the rule relationship is strongly correlated, determine the correlation coefficient between the first rule and the second rule; and calculate the target confidence level based on the correlation coefficient, the first confidence level, and the second confidence level, wherein a strong correlation refers to a relationship in which the correlation coefficient between the rules is greater than a third set threshold; or,

[0213] If the rule relationship is a complementary relationship, a complementary weight between the first rule and the second rule is determined; and a target confidence is calculated based on the complementary weight, the correlation coefficient, the first confidence and the second confidence.

[0214] Optionally, the fusion output module 904 is used to:

[0215] Determining whether the multiple output results include only one diagnosis result;

[0216] If the multiple output results include only one first diagnosis result, the first diagnosis result is determined as the final diagnosis result; or,

[0217] If the multiple output results include multiple diagnosis results, the second diagnosis result in the third output result, the third diagnosis result in the first output result, and / or the fourth diagnosis result in the second output result are compared and analyzed to determine a final diagnosis result.

[0218] Optionally, the fusion output module 904 is used to:

[0219] In the case where the second diagnosis result is determined based on a rule case, the second diagnosis result is determined as the final diagnosis result; or,

[0220] In the case where the second diagnostic result is not determined based on a rule case and the second diagnostic result is different from the third diagnostic result and / or the fourth diagnostic result, a weighted average of the target confidence in the third output result and the first probability value in the first output result and / or the second probability value in the second output result is performed, and the second diagnostic result, the third diagnostic result and the fourth diagnostic result are determined as the final diagnostic result; or,

[0221] When the second diagnostic result is not determined based on a rule case and the second diagnostic result is the same as the third diagnostic result and the fourth diagnostic result, the target confidence, the first probability value and the second probability value are compared, and the diagnostic result corresponding to the maximum value is determined as the final diagnostic result.

[0222] For the description of the features in the embodiment corresponding to the computer component fault diagnosis device, reference can be made to the relevant description of the embodiment corresponding to the computer component fault diagnosis method, which will not be repeated here.

[0223] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned computer component fault diagnosis method embodiments.

[0224] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned computer component fault diagnosis method embodiments when running.

[0225] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0226] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned computer component fault diagnosis method embodiments are implemented.

[0227] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned computer component fault diagnosis method embodiments.

[0228] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0229] The above is a detailed introduction to the fault diagnosis of a computer component provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. It should be noted that, for those skilled in the art, without departing from the principles of this application, various improvements and modifications may be made to this application, and such improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A method for diagnosing a fault of a computer component, characterized in that: include: Acquiring time series data of a computer component; wherein the time series data refers to a series of ordered data points generated and / or collected by the computer component within a set time interval; Extracting multiple features that comprehensively reflect the fault status of the computer component from the time series data to obtain time series features; wherein the time series features include time domain features, frequency domain features, and event sequence features; Fault diagnosis is performed based on the time series features using multiple pre-built models to obtain multiple output results; wherein the output results include diagnostic results and diagnostic values corresponding to the diagnostic results, the diagnostic values being quantitative representations of the diagnostic results generated by the models, the inputs of the multiple models being the time series features, the multiple models including a short-term feature model for diagnosing sudden anomalies, a long-term dependency model for diagnosing performance anomaly trends, and a knowledge rule reasoning model for diagnosing anomalies of multiple predefined rules; the diagnostic value of the short-term feature model is a probability value, the diagnostic value of the long-term dependency model is a probability value, and the diagnostic value of the knowledge rule reasoning model is a confidence level; Fusing the multiple output results to determine a final diagnosis result; The short-term feature model adopts a temporal convolutional network, which includes a residual module. The residual module includes multiple convolutional layers, convolution kernel weights, activation functions, and connection methods between network layers. Before diagnosing a sudden abnormality using the short-term characteristic model and obtaining an output result, the method further includes: Setting an association between the expansion factor in the short-term feature model and the number of convolutional layers; wherein the association means that the expansion factor increases with the number of network layers; decomposing the weight into two independent parameters, direction and amplitude, to constrain the weight distribution; setting a gated activation function; setting the connection mode to a skip connection; wherein the skip connection refers to a network structure in which the input is added to the output; The fusing of the multiple output results to determine a final diagnosis result includes: Determine whether the multiple output results include only one diagnostic result; if the multiple output results include only one diagnostic result, determine the one diagnostic result as the final diagnostic result; if the multiple output results include multiple diagnostic results, compare and analyze the diagnostic result in the third output result output by the knowledge rule reasoning model, the diagnostic result in the first output result output by the short-term feature model, and the diagnostic result in the second output result output by the long-term dependency model to determine the final diagnostic result, including: In a case where the diagnosis result in the third output result is determined based on a rule case, determining the diagnosis result in the third output result as the final diagnosis result; If the diagnosis result in the third output result is not determined based on the rule case and the diagnosis results in the multiple output results are all different, performing a weighted average of the confidence value in the third output result, the probability value in the first output result, and the probability value in the second output result, and determining the diagnosis result in the third output result, the diagnosis result in the first output result, and the diagnosis result in the second output result as the final diagnosis result; When the diagnosis result in the third output result is not determined based on the rule case, and the diagnosis result in the third output result is the same as the diagnosis result in the first output result or the diagnosis result in the second output result, the confidence level is compared with the probability value in the first output result or the probability value in the second output result, and the diagnosis result corresponding to the maximum value is determined as the final diagnosis result.

2. The method according to claim 1, characterized in that The step of extracting at least one feature reflecting the fault state of the computer component from the time series data to obtain a time series feature includes: Calculating the mean and variance of the time series data; Calculating kurtosis based on the time series data, the mean, and the variance; wherein the kurtosis reflects the data distribution of the time series data; Obtaining a mean and a variance of previous data; wherein the previous data refers to time series data of the computer component obtained before the set time interval; Calculating a mean difference and a variance ratio based on the mean and variance of the time series data and the mean and variance of the previous data; Among them, the time domain features include the mean, variance, kurtosis, mean difference and variance ratio of the time series data.

3. The method according to claim 1, characterized in that The set time interval includes at least one time window, the time series data includes a plurality of data acquired in the time window, and extracting at least one feature reflecting the fault state of the computer component from the time series data to obtain the time series feature includes: Performing frequency domain conversion on the multiple data acquired in the time window to obtain a complex spectrum corresponding to the time window; Extracting a main frequency component from the complex spectrum to obtain a main frequency energy ratio corresponding to the time window; Counting the mean and standard deviation of the main frequency energy ratio corresponding to the at least one time window, and calculating the abnormal threshold range according to the mean, the standard deviation and a preset threshold; The target main frequency energy ratio within the abnormal threshold range is determined as the frequency domain feature.

4. The method according to claim 1, wherein The time series data includes discrete data and continuous data reflecting discrete events, and extracting at least one feature reflecting the fault state of the computer component from the time series data to obtain a time series feature includes: Determining an event type of the discrete event; wherein the discrete event refers to a non-continuous event with a timestamp that occurs at a certain time point within the set time interval; Encoding the discrete event into a timestamp sequence according to the event type to obtain discrete features; The discrete features and the continuous features are spliced in the feature dimension to obtain event sequence features; wherein the continuous features are extracted from the continuous data.

5. The method according to claim 4, characterized in that The step of encoding the discrete event into a timestamp sequence according to the event type to obtain discrete features includes: If the event type is a single event, the discrete event is encoded into a timestamp sequence using binary coding to obtain discrete features; wherein the binary coding means marking whether the event occurs at each sampling time point; or, If the event type is a high-frequency event, the discrete event is encoded into a timestamp sequence using count coding to obtain discrete features; wherein the high-frequency event refers to an event that occurs more than a first set threshold number of times within a first time period, and the count coding refers to the number of occurrences of the event within a statistical time window; or, If the event type is an intermittent event, the discrete event is encoded into a timestamp sequence using time difference coding to obtain discrete features; wherein the intermittent event refers to an event that occurs discontinuously and has periodicity, and the time difference coding refers to recording the time interval between adjacent events; or, If the event type is a complex event, time embedding coding is used to encode the discrete event into a timestamp sequence to obtain discrete features; wherein the complex event refers to an event whose occurrence period is greater than a second set threshold and relies on multiple data for trend analysis, and the time embedding coding refers to mapping the timestamp of the event into a continuous vector.

6. The method according to claim 1, characterized in that The fault diagnosis is performed based on the time series characteristics using the pre-built multiple models to obtain multiple output results, including: Using the time series feature as an input of the short-term feature model to diagnose a sudden abnormality through the short-term feature model to obtain a first output result; Using the time series features as input to the long-term dependency model, so as to diagnose abnormal performance trends through the long-term dependency model and obtain a second output result; Using the time series feature as an input of the knowledge rule reasoning model to diagnose anomalies according to a plurality of rules predefined in the knowledge rule reasoning model, thereby obtaining a third output result; The multiple output results include the first output result, the second output result and the third output result.

7. The method according to claim 6, characterized in that The step of using the time series feature as the input of the short-term feature model to diagnose a sudden abnormality through the short-term feature model to obtain a first output result includes: A sudden abnormality is diagnosed by adjusting the value of the dilation factor in the temporal convolutional network to obtain a first output result; wherein the dilation factor is used to enable the temporal convolutional network to capture pattern changes in the temporal features at different time scales.

8. The method according to claim 6, characterized in that The long-term dependency model adopts a long short-term memory neural network, and the time series features are used as inputs of the long-term dependency model to diagnose abnormal performance trends through the long-term dependency model to obtain a second output result, including: Using the time series features as input to the long-term dependency model, obtaining hidden states at different moments calculated by an output gate in the long short-term memory neural network; wherein the output gate is used to determine whether information is output at each moment; Based on the hidden state at the current moment, the hidden state at the previous moment, and a set parameter, the trend of the time series feature is calculated to obtain the predicted state at the current moment; wherein the set parameter refers to a parameter that controls the smoothness of the trend; Calculating a loss based on the hidden state at the current moment and the predicted state at the current moment; The loss is compared with a set range to determine a diagnosis result to obtain a second output result.

9. The method according to claim 6, characterized in that The time series feature is used as the input of the knowledge rule reasoning model to diagnose the anomaly through a plurality of rules predefined in the knowledge rule reasoning model to obtain a third output result, including: Matching the temporal features with a plurality of predefined rules in the knowledge rule reasoning model to determine at least one target rule; Obtaining rule relationships between the multiple rules; wherein the rule relationships reflect the degree of correlation between different rules; A calculation strategy for the at least one target rule is determined according to the rule relationship, and a target confidence is calculated according to the calculation strategy; wherein the third output result includes the confidence and a diagnosis result corresponding to the target confidence.

10. The method according to claim 9, characterized in that The at least one target rule includes a first rule and a second rule, and determining a calculation strategy for the at least one target rule according to the rule relationship, and calculating the target confidence according to the calculation strategy, includes: If the rule relationship between the first rule and the second rule is independent of each other, the target confidence is calculated according to the first confidence of the first rule and the second confidence of the second rule; or, If the rule relationship is a strong correlation, determining the correlation coefficient between the first rule and the second rule; and calculating a target confidence level based on the correlation coefficient, the first confidence level, and the second confidence level, wherein the strong correlation refers to a relationship in which the correlation coefficient between the rules is greater than a third set threshold; or, If the rule relationship is a complementary relationship, a complementary weight between the first rule and the second rule is determined; and a target confidence is calculated based on the complementary weight, the correlation coefficient, the first confidence, and the second confidence.

11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the computer component fault diagnosis method according to any one of claims 1 to 10 when executing the computer program.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault diagnosis method for a computer component according to any one of claims 1 to 10.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the fault diagnosis method for a computer component according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Fault diagnosis method of electric drive axle and training method of neural network model

    CN119513813A