Hardware fault positioning method and device, equipment and storage medium
By obtaining server hardware operation data, using transfer learning and LSTM model combined with Euclidean distance and cosine similarity to accurately locate fault hardware, the inefficiency problem in traditional methods is solved, and fast and accurate hardware fault location is achieved, reducing business interruptions and economic losses.
Patent Information
- Application Number
- CN202510830600.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Traditional server hardware fault location methods rely on experience and manual troubleshooting, are inefficient and easy to misjudgment, and are difficult to quickly and accurately locate faulty hardware, resulting in business interruptions and economic losses.
By acquiring hardware operation data, preprocessing it, using transfer learning and long and short-term memory network model (LSTM) to initially locate the faulty hardware, and combines Euclidean distance and cosine similarity to accurately locate the faulty hardware, and dynamically adjust the weights to determine the target fault location.
It realizes rapid and accurate positioning of server hardware failures, reduces business interruptions and economic losses, and improves the efficiency and accuracy of fault location.
Smart Images

Figure CN120336989A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fault prediction, and particularly to a method, device, equipment and storage medium for hardware fault location. Background Art
[0002] In today's digital age, servers, as the core devices for data storage, processing and transmission, are widely used in various fields such as the Internet, finance, healthcare, and scientific research. With the continuous improvement of server performance and the increasing complexity of the architecture, its hardware components are more precise, which significantly increases the probability of hardware failures. Once a server hardware failure occurs, it will not only cause business interruption and serious economic losses, but also may affect the enterprise's reputation and user experience. Traditional server hardware fault location methods mainly rely on the experience of operation and maintenance personnel and manual troubleshooting, with low efficiency and prone to misjudgment. Although some existing monitoring systems can monitor some hardware parameters, they lack in-depth analysis and effective utilization of historical data, making it difficult to quickly and accurately locate the faulty hardware, and even less able to provide detailed repair guidance for operation and maintenance personnel.
[0003] Therefore, in view of the shortcomings of the existing technical solutions, the present invention provides a method for hardware fault location. Summary of the Invention
[0004] This application provides a method, device, equipment and storage medium for hardware fault location, so as to at least solve the problems of business interruption and economic losses caused by traditional methods relying on experience and manual troubleshooting.
[0005] This application provides a method for hardware fault location. The method includes: obtaining the hardware operation data corresponding to multiple hardware collected by a controller; preprocessing the multiple hardware operation data according to the importance degree of the multiple hardware operation data; determining the faulty hardware based on the preprocessed hardware operation data according to transfer learning and long short-term memory network model; obtaining the real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extracting the data features of the real-time fault data and the data features of the historical fault data, and determining the real-time vector and the historical vector; calculating the Euclidean distance and cosine similarity based on the dynamic weights according to the real-time vector and the historical vector of the multiple candidate fault locations, and determining the target fault location.
[0006] The present application also provides a hardware fault location device, which includes: a first processing module for obtaining hardware operation data corresponding to multiple hardware components collected by a controller; a second processing module for preprocessing the multiple hardware operation data according to the importance level of the multiple hardware operation data; a third processing module for determining a faulty hardware component based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; a fourth processing module for obtaining real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware component within a preset time, extracting data features of the real-time fault data and data features of the historical fault data, and determining a real-time vector and a historical vector; a fifth processing module for calculating Euclidean distance and cosine similarity based on the real-time vector and the historical vector of the multiple candidate fault locations and determining a target fault location based on dynamic weights.
[0007] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the following steps when executing the computer program: obtaining hardware operation data corresponding to multiple hardware components collected by a controller; preprocessing the multiple hardware operation data according to the importance level of the multiple hardware operation data; determining a faulty hardware component based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; obtaining real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware component within a preset time, extracting data features of the real-time fault data and data features of the historical fault data, and determining a real-time vector and a historical vector; calculating Euclidean distance and cosine similarity based on the real-time vector and the historical vector of the multiple candidate fault locations and determining a target fault location based on dynamic weights.
[0008] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the following steps: obtaining hardware operation data corresponding to multiple hardware components collected by a controller; preprocessing the multiple hardware operation data according to the importance level of the multiple hardware operation data; determining a faulty hardware component based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; obtaining real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware component within a preset time, extracting data features of the real-time fault data and data features of the historical fault data, and determining a real-time vector and a historical vector; calculating Euclidean distance and cosine similarity based on the real-time vector and the historical vector of the multiple candidate fault locations and determining a target fault location based on dynamic weights.
[0009] The present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps: obtaining hardware operation data corresponding to multiple hardware components collected by a controller; preprocessing the multiple hardware operation data according to the importance levels of the multiple hardware operation data; determining a faulty hardware component based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; obtaining real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware component within a preset time, extracting data features of the real-time fault data and data features of the historical fault data, and determining a real-time vector and a historical vector; calculating an Euclidean distance and a cosine similarity based on the real-time vector and the historical vector of the multiple candidate fault locations based on dynamic weights, and determining a target fault location.
[0010] Through the present application, since the hardware operation data corresponding to multiple hardware components collected by the controller is obtained; the multiple hardware operation data is preprocessed according to the importance levels of the multiple hardware operation data; the faulty hardware component is determined based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; the real-time fault data and the historical fault data of the multiple candidate fault locations of the faulty hardware component within a preset time are obtained, the data features of the real-time fault data and the data features of the historical fault data are extracted, and the real-time vector and the historical vector are determined; the Euclidean distance and the cosine similarity are calculated based on the real-time vector and the historical vector of the multiple candidate fault locations based on dynamic weights, and the target fault location is determined, therefore, the faulty hardware component is initially located through the LSTM model, and then the target fault location is accurately located by combining the Euclidean distance and the cosine similarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a flowchart showing the process of a hardware fault location method provided by an embodiment of the present application; Figure 2 It is a flowchart showing the process of determining a quantile interval of a hardware fault location method provided by an embodiment of the present application; Figure 3 It is a flowchart showing the process of machine learning algorithm modeling and optimization of a hardware fault location method provided by an embodiment of the present application; Figure 4 It is a flowchart showing the process of hardware fault location of a hardware fault location method provided by an embodiment of the present application; Figure 5Block diagram of a hardware fault location device provided by an embodiment of the present application; Figure 6 Internal structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0015] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing steps, and do not particularly refer to the meaning of order or sequence, nor are they used to limit the present application. They are only used to conveniently describe the method of the present application and cannot be understood as indicating the order of steps. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0016] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0017] An embodiment of the present application provides a hardware fault location method. The method will be described in detail in combination with the execution flow of the hardware fault location method.
[0018] S101: Obtain the hardware operation data corresponding to multiple hardware collected by the controller.
[0019] Here, the controller can be an OPEN BMC (Open Baseboard Management Controller) or a BMC (Baseboard Management Controller).
[0020] Here, the hardware is the hardware in the server.
[0021] Among them, the hardware operation data corresponding to one piece of hardware can include one piece of data or multiple pieces of data.
[0022] Among them, the hardware operation data is obtained by collecting sensors. For example, OpenBMC has a dedicated process for collecting sensor data: history-record, and this process will record various sensor data into a file.
[0023] Here, the hardware operation data can be the temperature of the CPU (Central Processing Unit), the temperature of the PSU (Power Supply Unit), and the usage rate data of various types of hardware; among them, the CPU temperature can include CPU temperature, CPU0_Temp, CPU1_Temp, CPU_Margin_Temp, CPU_DIMM_Temp, and CPU_VR_Temp, etc.; the PSU temperature can include PSU0_Inlet_Temp and PSU1_Inlet_Temp, etc.; the usage rate data of various types of hardware can include CPU usage rate, memory usage rate, and IO (Input / Output) usage rate, etc.
[0024] In one embodiment, when collecting sensor data, different attributes can be configured for each sensor, and different sampling times can be configured. For example, a shorter sampling time can be configured for the sensors that are key concerns, and a longer sampling time can be configured for the sensors that are not key concerns.
[0025] S102: Preprocess multiple pieces of hardware operation data according to the importance levels of the multiple pieces of hardware operation data.
[0026] Here, the preprocessing can include normalization, missing value processing, duplicate value processing, outlier detection and processing, standardization, etc.
[0027] Among them, the importance level of the hardware operation data can be determined according to the fault correlation, influence range, fluctuation controllability, safety threshold clarity, monitoring priority, etc. of the hardware operation data. Exemplarily, assume that the hardware operation data is the CPU temperature. High temperature is highly correlated with system downtime, and the fault correlation is strong; it directly affects system stability, and the influence range is large; there is a relatively narrow reasonable range under normal operation, and the fluctuation is controllable; the manufacturer provides TJMAX (maximum junction temperature), and the safety threshold is clear; most systems have a temperature alarm mechanism, and the monitoring priority is high. To sum up, the CPU temperature is important hardware operation data.
[0028] Specifically, the importance degree of the hardware operation data can be obtained through a scoring model, the analytic hierarchy process, the fuzzy comprehensive evaluation method, machine learning methods, etc.
[0029] Specifically, these hardware operation data can be divided according to a time window with a fixed duration to form a continuous sequence. For example, for the hardware operation data of CPU temperature, the temperature data within every 5 minutes is taken as a sequence segment. In this way, a multi-dimensional time series data set containing multiple hardware operation data is constructed.
[0030] S103: Based on the preprocessed hardware operation data, determine the faulty hardware according to transfer learning and the long short-term memory network model.
[0031] Here, transfer learning is a machine learning method aimed at applying the knowledge learned in one field or task to a different but related field or task.
[0032] Among them, the LSTM model is improved and optimized through transfer learning, making the LSTM model a model for classifying and diagnosing server hardware faults.
[0033] Here, the LSTM (Long Short-Term Memory) model is a special recurrent neural network aimed at solving the problems of gradient disappearance and gradient explosion in long sequence training.
[0034] Specifically, the network architecture of the LSTM model includes: an input layer, an LSTM layer, and a fully connected layer. The input layer is responsible for receiving preprocessed multi-dimensional time series data, and the data dimension n depends on the types of the collected hardware operation data. For example, if n parameters such as CPU temperature, memory usage rate, and hard disk read / write speed are collected at the same time, the dimension of the time series data received by the input layer is n, which can ensure that the model can make full use of multi-source hardware operation data for fault diagnosis; the LSTM layer is configured with two layers of LSTM cells, each layer contains 128 neurons, and the activation function is selected as Tanh. The unique gating mechanism of LSTM (including input gate, forget gate, and output gate) enables it to effectively capture long-term dependencies in the time series. The input gate determines the retention degree of the current input information, the forget gate controls the discarding of old information in the cell state, and the output gate determines the final output content. When processing the data of CPU temperature changing over time, LSTM can remember the temperature trend in the past period of time. Even if there are short-term fluctuations in the middle, it can accurately judge whether the temperature change is abnormal, thus providing strong support for fault diagnosis; the output layer of the fully connected layer uses the Softmax function, whose role is to map the feature vector output by the LSTM layer to each fault category label. For example, the fault categories can be set as "CPU fault", "power module fault", "memory fault", etc. The Softmax function will calculate the probability of each fault category, and the category with the highest probability is the fault type predicted by the model, realizing the classification diagnosis of server hardware faults.
[0035] S104: Obtain the real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extract the data features of the real-time fault data and the data features of the historical fault data, and determine the real-time vector and the historical vector.
[0036] Here, the candidate fault location refers to the specific module in the faulty hardware where a fault may exist. For example, if the faulty hardware is a CPU, the candidate fault locations may be the CPU voltage regulator module, cache, control unit, input / output interface, etc.
[0037] Among them, different fault locations correspond to different fault data.
[0038] Here, the fault data can include various data, such as CPU temperature, CPU usage rate, etc.
[0039] Among them, one candidate fault location corresponds to one real-time vector and one historical vector. The real-time vector and the historical vector are n-dimensional vectors, including the features of multiple fault data.
[0040] S105: Based on the real-time vectors and historical vectors of multiple candidate fault locations, calculate the Euclidean distance and cosine similarity based on dynamic weights, and determine the target fault location.
[0041] Here, the dynamic weight can dynamically adjust the weights of different fault data under different circumstances.
[0042] Here, the Euclidean distance is a common way to measure the absolute distance between two vectors in a multi-dimensional space.
[0043] Here, the cosine similarity measures the similarity between vectors by measuring the cosine value of the angle between two vectors.
[0044] Among them, the smaller the Euclidean distance between two vectors, the more similar the two vectors are; the larger the cosine similarity between two vectors, the more similar the two vectors are.
[0045] Among them, the value range of the cosine similarity is between [-1, 1]. The closer the value of the cosine similarity is to 1, the more similar vectors X and Y are; the closer the value of the cosine similarity is to -1, the more opposite their directions are.
[0046] Specifically, after determining the faulty hardware, from the preset fault association database, obtain multiple fault locations corresponding to the faulty hardware and the fault data corresponding to each fault location, collect the real-time fault data corresponding to each fault location, extract the features of the real-time fault data to obtain a real-time vector, obtain the historical fault data, extract the features of the historical fault data to obtain a historical vector, compare the similarity between the real-time vector and the historical vector of each fault location, and use the fault location with the greatest similarity between the real-time vector and the historical vector as the target fault location.
[0047] In one embodiment, after determining the target fault location, obtain the solution corresponding to the fault location from the fault association database, and perform fault repair according to the instructions provided by the fault association database. Exemplarily, these repair guidance steps are carefully organized and optimized, covering the whole process from preparing repair tools, disassembling the faulty hardware, installing new hardware to debugging and testing. For example, for a CPU fault, the required screwdriver model, the correct order of disassembling the CPU radiator, precautions when installing a new CPU, and the specific methods and indicators for debugging and testing will be clearly listed.
[0048] It should be noted that this application initially locates the faulty hardware through the LSTM model, and then combines the Euclidean distance and cosine similarity to accurately locate the fault location.
[0049] In some specific embodiments, according to the importance of multiple hardware operation data, preprocess the multiple hardware operation data, including: Determine the quantile intervals corresponding to the multiple hardware operation data according to the importance of the multiple hardware operation data; Preprocess the multiple hardware operation data according to the quantile intervals.
[0050] Here, a quantile refers to dividing a data set into equal parts. Common quantiles include percentiles, deciles, and quartiles.
[0051] Among them, for important hardware operation data (such as CPU temperature, PSU voltage, etc.), a narrow quantile interval (such as 5% - 95%) is selected. The narrow quantile interval can eliminate abnormal jump values of sensors and retain the detailed differences in the real business interval. If the 95th percentile of the CPU temperature historical data is 75°C and the 5th percentile is 25°C, then this feature is normalized to the [0, 1] interval corresponding to [25°C, 75°C], and extreme values outside this range are regarded as noise and filtered.
[0052] Among them, for unimportant hardware operation data (such as fan speed, etc.), a wide quantile interval (such as 1% - 99%) is selected. Unimportant hardware operation data allows a certain range of abnormal fluctuations, and the wide quantile interval can avoid over-filtering normal extreme values (such as the instantaneous high speed when the fan runs at full speed).
[0053] Here, preprocessing can include normalization, mapping each hardware operation data to the [0, 1] interval, so that different parameters are on the same scale, which helps the model converge faster and improve the training effect. Normalization can be achieved through formula (1): (1) Among them, x is the original data, and are respectively the minimum and maximum values of this hardware operation data in the data set.
[0054] Specifically, the data set is sorted according to the data size, the hardware operation data is segmented through the quantile interval, the hardware operation data within the quantile interval is obtained, and the hardware operation data within the quantile interval is normalized.
[0055] In one embodiment, after determining the quantile interval, it is possible to judge whether the quantile interval meets the requirements by calculating the F1 value and false alarm rate of the quantile interval on the validation set.
[0056] In this way, while removing outliers, the fluctuation details of the data under normal operating conditions can be retained.
[0057] In some specific embodiments, according to the importance levels of multiple hardware operation data, determining the quantile intervals corresponding to the multiple hardware operation data further includes: Obtaining the life cycle and real-time load of the hardware corresponding to the multiple hardware operation data; Based on the life cycles and real-time loads of the multiple hardware, adaptively adjusting the quantile intervals of the multiple hardware operation data.
[0058] Here, the life cycle can include a new machine stage and an aging stage, and the real-time load can include a high load and a low load.
[0059] Among them, when the hardware corresponding to the hardware operation data is in the new machine stage, a conservative quantile interval (such as 10%-90%) is adopted to avoid misjudging normal fluctuations during the running-in period as anomalies; when the hardware corresponding to the hardware operation data is in the aging stage, an aggressive quantile interval (such as 5%-95%) is selected to capture subtle changes in early performance degradation.
[0060] Among them, when the hardware corresponding to the hardware operation data is in a high-load scenario (such as big data computing), the quantile interval is automatically widened to avoid normal peaks being normalized and compressed; when the hardware corresponding to the hardware operation data is in a low-load scenario, the quantile interval is narrowed to sensitively capture small anomalies and improve the sensitivity of anomaly detection.
[0061] In one embodiment, factors such as time period, environmental variables, and task type can also be considered to adjust the quantile interval.
[0062] In this way, dynamically adjusting the quantile interval considering the actual situation can enhance the robustness and adaptability of normalization.
[0063] In some specific embodiments, based on the preprocessed hardware operation data, according to transfer learning and long short-term memory network models, the faulty hardware is determined, including: Obtain general hardware operation data and preprocess the general hardware operation data; Through the preprocessed general hardware operation data, pre-train a preset long short-term memory network model until the preset long short-term memory network model converges to obtain a first model; Fine-tune the first model according to the hardware operation data to obtain a target model; Input the processed hardware operation data into the target model and output the faulty hardware.
[0064] Here, the general hardware operation data can include the hardware operation data of servers of multiple brands and multiple models.
[0065] Here, fine-tuning is a part of transfer learning. Fine-tuning means training a model through a general dataset, and then fine-tuning the model according to the hardware characteristics and fault modes of a specific model of server, so that the model can better adapt to the current server. Specifically, freeze the underlying parameters of the LSTM. These parameters have learned the general hardware operation mode in the pre-training stage and do not need to be trained again. Only for the fully connected layer.
[0066] Specifically, first run the data on general-purpose hardware to pre-train the underlying parameters of the LSTM, obtaining a general LSTM model. Determine the corresponding server based on the hardware operation data, and use approximately 500 sets of the server's hardware operation data to fine-tune the LSTM model to obtain a fine-tuned model.
[0067] Specifically, preprocess the general-purpose hardware operation data. According to the importance of multiple general-purpose hardware operation data, determine the quantile interval corresponding to the general-purpose hardware operation data, obtain the life cycle and real-time load of the hardware corresponding to multiple general-purpose hardware operation data, adaptively adjust the quantile interval of multiple general-purpose hardware operation data based on the life cycle and real-time load of multiple hardware, and preprocess the general-purpose hardware operation data according to the quantile interval.
[0068] In one embodiment, for the case where the general-purpose hardware operation data samples are few, data augmentation techniques can be used to expand the general-purpose hardware operation data set. For time series data, new samples can be generated by means such as translation, scaling, and adding noise; for numerical data, variants of the generative adversarial network (GAN), such as the conditional generative adversarial network (CGAN), are used to generate more representative general-purpose hardware operation data samples with the hardware failure type as the condition, enabling the model to learn richer failure features during the training process and improving the diagnostic ability for rare failures.
[0069] In one embodiment, during the actual operation of the target model, a reinforcement learning algorithm can be used for dynamic optimization. Using the fault location accuracy, repair time, etc. as reward metrics, the target model continuously adjusts its own parameters and decision-making strategies according to the results of each fault diagnosis and repair. For example, when the target model makes misjudgments multiple times in a certain type of fault diagnosis, the reinforcement learning mechanism will give negative rewards, prompting the target model to re-learn the features of this type of fault, optimize the diagnostic rules, and gradually improve the accuracy and efficiency of diagnosis.
[0070] In this way, the LSTM model can quickly adapt to the characteristics of a specific model of server, complete model adaptation with less data, and improve the accuracy and pertinence of diagnosis.
[0071] In some specific embodiments, pre-train a preset long short-term memory network model with the preprocessed general-purpose hardware operation data, including: Input the preprocessed hardware operation data into the preset long short-term memory network model to obtain a loss function; Judge whether the preset long short-term memory network model converges according to the loss function; If not converged, update the parameters of the preset long short-term memory network model through an optimization algorithm.
[0072] Here, the loss function is used to measure the difference between the model's predicted value and the true value.
[0073] Specifically, the loss function can be the cross - entropy loss, and the cross - entropy loss is calculated by formula (2): (2) where C is the number of fault categories, y i is the true label (0 or 1), p i is the probability that the model predicts the i - th type of fault.
[0074] Here, the optimization algorithm calculates the gradient of the loss function with respect to the model parameters and updates the parameters in the opposite direction of the gradient, so that the loss value gradually decreases, making the model's prediction closer and closer to the true value.
[0075] Specifically, the optimization algorithm can be the Adam algorithm. The Adam algorithm is an adaptive learning rate method for optimizing the weights of a neural network. It adjusts the learning rate of each parameter in the model by calculating the first - moment estimate and second - moment estimate of the gradient. Exemplarily, the learning rate can be set to 0.001. The Adam algorithm can adaptively adjust the learning rate, accelerate the model convergence speed, and improve the training efficiency.
[0076] In one embodiment, in the fine - tuning stage, the model can also be trained through the loss function and the optimization algorithm.
[0077] In one embodiment, to prevent the model from overfitting, the Dropout technique is introduced, and the dropout probability is set to 0.3. Dropout randomly discards some neurons during the training process, avoiding the model's over - reliance on certain specific neurons and enhancing the model's generalization ability. At the same time, L2 regularization is adopted, and the coefficient is set to 0.01. L2 regularization constrains the weights, making the model smoother and reducing the overfitting risk caused by overly large weights.
[0078] In this way, the accuracy and generalization ability of the model can be improved.
[0079] In some specific embodiments, based on the real - time vectors and historical vectors of multiple candidate fault locations, and based on dynamic weights, the Euclidean distance and cosine similarity are calculated to determine the target fault location, including: Determine the weights of multiple fault data according to the real - time vector and the historical vector, and calculate the Euclidean distance between the real - time vector and the historical vector; Determine the weights of multiple time windows according to the real - time vector and the historical vector, and calculate the cosine similarity between the real - time vector and the historical vector based on the dynamic weights; Calculate the similarity scores of multiple candidate fault locations according to the Euclidean distance and the cosine similarity; In response to the existence of a similarity score greater than or equal to a preset score, the candidate fault location with the highest similarity score is taken as the target fault location.
[0080] Here, the fault data corresponding to each candidate fault location is different, and each candidate fault location corresponds to a vector.
[0081] Among them, a historical vector is selected as a comparison object. The more similar the real-time vector is to the historical vector, the higher the possibility of a similar fault occurring at the current fault location.
[0082] Specifically, through the real-time vector, the historical vector, and the dynamic weight, the Euclidean distance and cosine similarity between the real-time vector and the historical vector of each candidate fault location are calculated. Combining the Euclidean distance and cosine similarity to obtain a similarity score, comparing the similarity scores of each candidate fault location, and determining whether there is a similarity score greater than or equal to the preset score. When there is a similarity score greater than or equal to the preset score, the candidate fault location with the highest similarity score is selected as the target fault location.
[0083] In one embodiment, it is possible to determine to detect sudden faults or predict progressive faults according to requirements. When detecting sudden faults, the change rate of the fault data is detected, and according to the change rate of the fault data combined with other factors, the weights of multiple pieces of fault data are determined. The time window is divided into small time periods, divided into millisecond-level windows, to capture instantaneous signal jumps, and the weights of the millisecond-level windows are determined. When predicting progressive faults, the change rate and moving average of the fault data are detected, and according to the change rate and moving average of the fault data combined with other factors, the weights of multiple pieces of fault data are determined. The time window is divided into long time periods, divided into hour-level windows, and the weights of the hour-level windows are determined. In this way, sudden faults can be detected or progressive faults can be predicted, improving the fault response speed and enhancing the prediction ability.
[0084] In this way, by combining multiple measurement criteria, the accuracy of judgment can be improved, and the false alarm rate caused by the fluctuation of a single index can be reduced.
[0085] In some specific embodiments, according to the Euclidean distance and cosine similarity, calculating the similarity scores of multiple candidate fault locations includes: Calculating the fusion coefficient of the Euclidean distance and cosine similarity through the fault data training set; Based on the fusion coefficient, Euclidean distance, and cosine similarity, a similarity score is obtained through weighted summation.
[0086] Specifically, the similarity score can be calculated by formula (3): (3) Among them, β is the fusion coefficient, and its value range is [0, 1], is the cosine similarity between the real-time vector and the historical vector, d DWED is the Euclidean distance between the real-time vector and the historical vector.
[0087] In one embodiment, the Euclidean distance and the cosine similarity can also be fused through an adaptive gating mechanism.
[0088] In this way, the absolute numerical difference and the relative change trend of the features are considered simultaneously, improving the comprehensiveness of fault matching; when there is data noise or partial sensor failure, the hybrid metric can be compensated by another metric, reducing the risk of misjudgment.
[0089] In some specific embodiments, according to the real-time vector and the historical vector, the weights of multiple fault data are determined, and the Euclidean distance between the real-time vector and the historical vector is calculated, including: Determine the initial weights of multiple fault data according to the importance levels of multiple fault data; Determine the adjusted weights of multiple fault data according to the fault thresholds of multiple fault data and / or the correlation coefficients between multiple fault data; Calculate the Euclidean distance between the real-time vector and the historical vector according to the adjusted weights of multiple fault data.
[0090] Specifically, the Euclidean distance can be calculated by formula (4): (4) where, w i (0 < w i ≤ 1)is the weight of each fault data, x i and x j are the i-th dimensional feature values of the real-time fault data and the historical fault data, respectively.
[0091] In this way, by dynamically adjusting the weights of fault data, it can be adjusted in real time with the accumulation of new fault data to adapt to changes in the hardware architecture.
[0092] In some specific embodiments, according to the fault thresholds of multiple fault data, the adjusted weights of multiple fault data are determined, including: Obtain the fault thresholds of multiple fault data; Determine the sensitivity factors corresponding to multiple fault data according to multiple fault data and the fault thresholds; Calculate the adjusted weights of multiple fault data based on the sensitivity factors.
[0093] Here, different fault data correspond to different fault thresholds. For example, the fault threshold for CPU temperature is 90°C.
[0094] Specifically, the sensitivity factor can be calculated by formula (5): (5) where x i is the fault data, T i is the fault threshold, and λ is the sensitivity coefficient.
[0095] Among them, different hardware corresponds to different sensitivity coefficients, and the sensitivity coefficients of different hardware are fixed.
[0096] Specifically, the initial weight of each fault data is multiplied by the sensitivity factor to obtain the corrected weight of each fault data.
[0097] In this way, through the sensitivity factor, the weight contains both the fault critical distance information, and the influence of abnormal features is amplified in advance.
[0098] In some specific embodiments, according to the correlation coefficients between multiple fault data, the corrected weights of multiple fault data are determined, including: Calculating the correlation coefficients between multiple fault data based on the historical fault data of multiple fault data; Constructing an association matrix according to the correlation coefficients; Calculating the corrected weights of multiple fault data based on the association matrix.
[0099] Here, the correlation coefficient represents the association strength between two fault data. For example, the association strength between CPU temperature and CPU usage rate.
[0100] Specifically, the mutual information or correlation coefficient between parameters is calculated through historical BMC data to generate an association matrix C, where C i,j represents the parameter x i and x j the association strength (0 ≤ C i,j ≤ 1).
[0101] Specifically, the corrected weight can be calculated by formula (6): (6) where W i and W j are the fault data xi The initial weight of x j , is the weight of the fault data after association constraint . C i,j Indicates the fault data x i and x jA correlation strength.
[0102] In this way, introducing the feature correlation matrix to correct the weight can avoid the weight of a single feature deviating too much from the importance of its associated features.
[0103] In some specific embodiments, according to the real-time vector and the historical vector, the weights of multiple time windows are determined, and based on the dynamic weights, the cosine similarity between the real-time vector and the historical vector is calculated, including: According to the preset window size, the real-time vector and the historical vector are divided into multiple time windows in chronological order; Based on the attention mechanism, the weights of multiple time windows are calculated; According to the weights of multiple time windows, the cosine similarity between the real-time vector and the historical vector is calculated.
[0104] Specifically, the preset window size can be 5 minutes.
[0105] Specifically, the real-time data and the historical data are divided into T windows in chronological order, and the feature vectors in each window are represented as and , where n represents the dimension of the vector, that is, the number of elements in the vector, is the vector X t each element of, respectively represents the value corresponding to different features in the time window t; is the vector Y t each element of, respectively represents the value corresponding to different features in the time window t.
[0106] Specifically, the Transformer attention mechanism is used to calculate the weights of each time window , reflecting the importance degree of this window. Example: The window weight within 30 minutes before the fault occurs is higher than that in earlier periods.
[0107] Specifically, the cosine similarity can be calculated by the following formula (7): (7) where T represents the total number of time windows, is the weight of each time window, ≥ 0 and ; X t is a real - time vector, Y t is a historical vector.
[0108] In this way, by assigning different weights to the features of different time periods, the early warning signals before the fault occurrence can be accurately matched with the peak data at the time of fault occurrence.
[0109] In some specific embodiments, calculating the similarity scores of multiple candidate fault locations according to the Euclidean distance and cosine similarity further includes: In response to the non - existence of a similarity score greater than or equal to a preset score, obtaining the hardware operation data of multiple hardware devices; Cross - validating the faulty hardware according to the hardware operation data of multiple hardware devices; Determining the target fault location according to the cross - validation result.
[0110] Specifically, when locating a CPU fault, if there is no similarity score greater than or equal to the preset score, at this time, comprehensively analyze other relevant data, such as the output voltage and current data of the PSU, the read - write error rate of the memory, etc. If an abnormal fluctuation in the PSU output voltage is detected, further verify the inference that a fault in the CPU voltage regulator module may cause unstable power supply.
[0111] In this way, through the combination of the LSTM model and multi - dimensional data cross - validation, the accuracy and efficiency of fault location are greatly improved.
[0112] In one embodiment, Figure 2 is a flow diagram of determining the quantile interval in the embodiment of the present application. As Figure 2 shown, the process of determining the quantile interval in the present application includes: collecting hardware operation data, determining whether the hardware operation data is important data. When the hardware operation data is important data, select a narrow quantile interval. When the hardware operation data is non - important data, select a wide quantile interval; dynamically adjust the quantile interval in combination with the life cycle and real - time load, and generate the final quantile interval through five - fold cross - validation of the quantile interval.
[0113] In one embodiment, Figure 3 is a flow diagram of machine learning algorithm modeling and optimization in the embodiment of the present application. As Figure 3 shown, the machine learning algorithm modeling and optimization process in the present application includes: BMC collects data, splits the data to obtain a multi - dimensional time - series data set, performs data normalization, outputs the probability of the fault type through the LSTM model, and optimizes the strategy through transfer learning to quickly use a specific model server.
[0114] In one embodiment, Figure 4 is a schematic flowchart of hardware fault location in the embodiments of the present application. As Figure 4 shown, the process of hardware fault location in the present application includes: server fault triggering, the BMC collects hardware operation data in real time and transmits it to the fault diagnosis module. The fault diagnosis module preliminarily diagnoses the fault type and probability with the help of the LSTM model, compares the real-time data with the historical data in the fault correlation database, calculates indexes such as Euclidean distance and cosine similarity, and determines whether the similarity reaches the preset threshold. If the similarity reaches the preset threshold, the faulty hardware is determined. If the similarity does not reach the preset threshold, multi-dimensional data such as the PSU output voltage and the memory read / write error rate are comprehensively analyzed for cross-verification to finally determine the faulty hardware.
[0115] It should be understood that although Figures 1-4 the steps in the flowchart of Figures 1-4 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0117] The embodiments of the present application also provide a hardware fault location device. The device includes: a first processing module 501 for obtaining the hardware operation data corresponding to multiple hardware collected by the controller; a second processing module 502 for preprocessing the multiple hardware operation data according to the importance of the multiple hardware operation data; a third processing module 503 for determining the faulty hardware based on the preprocessed hardware operation data according to transfer learning and the long short-term memory network model; a fourth processing module 504 for obtaining the real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extracting the data features of the real-time fault data and the data features of the historical fault data, and determining the real-time vector and the historical vector; a fifth processing module 505 for calculating the Euclidean distance and cosine similarity based on the dynamic weights according to the real-time vector and the historical vector of multiple candidate fault locations to determine the target fault location.
[0118] As a preferred implementation manner, in the embodiment of the present application, the second processing module is specifically configured to: determine quantile intervals corresponding to multiple hardware operation data according to the importance degrees of the multiple hardware operation data; and preprocess the multiple hardware operation data according to the quantile intervals.
[0119] As a preferred implementation manner, in the embodiment of the present application, the second processing module is further specifically configured to: obtain the life cycle and real-time load of the hardware corresponding to the multiple hardware operation data; and adaptively adjust the quantile intervals of the multiple hardware operation data based on the life cycles and real-time loads of the multiple hardware.
[0120] As a preferred implementation manner, in the embodiment of the present application, the third processing module is specifically configured to: obtain general hardware operation data, and preprocess the general hardware operation data; pre-train a preset long short-term memory network model through the preprocessed general hardware operation data until the preset long short-term memory network model converges to obtain a first model; fine-tune the first model according to the hardware operation data to obtain a target model; and input the processed hardware operation data into the target model to output a faulty hardware.
[0121] As a preferred implementation manner, in the embodiment of the present application, the third processing module is further specifically configured to: input the preprocessed hardware operation data into a preset long short-term memory network model to obtain a loss function; determine whether the preset long short-term memory network model converges according to the loss function; and if not, update the parameters of the preset long short-term memory network model through an optimization algorithm.
[0122] As a preferred implementation manner, in the embodiment of the present application, the fifth processing module is specifically configured to: determine the weights of multiple fault data according to a real-time vector and a historical vector, and calculate the Euclidean distance between the real-time vector and the historical vector; determine the weights of multiple time windows according to the real-time vector and the historical vector, and calculate the cosine similarity between the real-time vector and the historical vector based on dynamic weights; calculate the similarity scores of multiple candidate fault positions according to the Euclidean distance and the cosine similarity; and in response to the existence of a similarity score greater than or equal to a preset score, use the candidate fault position with the highest similarity score as the target fault position.
[0123] As a preferred implementation manner, in the embodiment of the present application, the fifth processing module is further specifically configured to: calculate a fusion coefficient of the Euclidean distance and the cosine similarity through a fault data training set; and obtain the similarity score through weighted summation based on the fusion coefficient, the Euclidean distance, and the cosine similarity.
[0124] As a preferred embodiment, in the embodiment of the present application, the fifth processing module is further specifically configured to: determine the initial weights of multiple fault data according to the importance levels of the multiple fault data; determine the adjusted weights of the multiple fault data according to the fault thresholds of the multiple fault data and / or the correlation coefficients between the multiple fault data; and calculate the Euclidean distance between the real-time vector and the historical vector according to the adjusted weights of the multiple fault data.
[0125] As a preferred embodiment, in the embodiment of the present application, the fifth processing module is further specifically configured to: obtain the fault thresholds of multiple fault data; determine the sensitivity factors corresponding to the multiple fault data according to the multiple fault data and the fault thresholds; and calculate the adjusted weights of the multiple fault data based on the sensitivity factors.
[0126] As a preferred embodiment, in the embodiment of the present application, the fifth processing module is further specifically configured to: calculate the correlation coefficients between multiple fault data according to the historical fault data of the multiple fault data; construct an association matrix according to the correlation coefficients; and calculate the adjusted weights of the multiple fault data based on the association matrix.
[0127] As a preferred embodiment, in the embodiment of the present application, the fifth processing module is further specifically configured to: divide the real-time vector and the historical vector into multiple time windows in chronological order according to a preset window size; calculate the weights of the multiple time windows based on an attention mechanism; and calculate the cosine similarity between the real-time vector and the historical vector according to the weights of the multiple time windows.
[0128] As a preferred embodiment, in the embodiment of the present application, the fifth processing module is further specifically configured to: in response to the non-existence of a similarity score greater than or equal to a preset score, obtain the hardware operation data of multiple hardware components; perform cross-verification on the faulty hardware according to the hardware operation data of the multiple hardware components; and determine the target fault location according to the cross-verification result.
[0129] For the descriptions of the features in the corresponding embodiment of the hardware fault location device, reference may be made to the relevant descriptions in the corresponding embodiment of the hardware fault location method, which will not be elaborated herein one by one.
[0130] The embodiment of the present application further provides an electronic device, which may be a terminal, and its internal structure diagram may be as Figure 6As shown in the figure. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a hardware fault location method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0131] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0132] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: S1: Obtain the hardware operation data corresponding to multiple hardware collected by the controller; S2: Preprocess the multiple hardware operation data according to the importance of the multiple hardware operation data; S3: Based on the preprocessed hardware operation data, determine the faulty hardware according to transfer learning and the long short-term memory network model; S4: Obtain the real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extract the data features of the real-time fault data and the data features of the historical fault data, and determine the real-time vector and the historical vector; S5: According to the real-time vector and the historical vector of multiple candidate fault locations, calculate the Euclidean distance and cosine similarity based on the dynamic weight, and determine the target fault location.
[0133] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Determine the quantile interval corresponding to the multiple hardware operation data according to the importance of the multiple hardware operation data; Preprocess the multiple hardware operation data according to the quantile interval.
[0134] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Obtain the life cycle and real-time load of the hardware corresponding to the multiple hardware operation data; Adaptively adjust the quantile interval of the multiple hardware operation data based on the life cycle and real-time load of the multiple hardware.
[0135] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining general hardware operation data, and preprocessing the general hardware operation data; pre-training a preset long short-term memory network model with the preprocessed general hardware operation data until the preset long short-term memory network model converges, to obtain a first model; fine-tuning the first model according to the hardware operation data, to obtain a target model; inputting the processed hardware operation data into the target model, and outputting a faulty hardware.
[0136] In one embodiment, when the processor executes the computer program, the following steps are further implemented: inputting the preprocessed hardware operation data into a preset long short-term memory network model, to obtain a loss function; judging whether the preset long short-term memory network model converges according to the loss function; if not converging, updating the parameters of the preset long short-term memory network model through an optimization algorithm.
[0137] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining the weights of multiple fault data according to a real-time vector and a historical vector, and calculating the Euclidean distance between the real-time vector and the historical vector; determining the weights of multiple time windows according to the real-time vector and the historical vector, and calculating the cosine similarity between the real-time vector and the historical vector based on the dynamic weights; calculating the similarity scores of multiple candidate fault positions according to the Euclidean distance and the cosine similarity; in response to there being a similarity score greater than or equal to a preset score, taking the candidate fault position with the highest similarity score as the target fault position.
[0138] In one embodiment, when the processor executes the computer program, the following steps are further implemented: calculating a fusion coefficient of the Euclidean distance and the cosine similarity through a fault data training set; obtaining the similarity score through weighted summation based on the fusion coefficient, the Euclidean distance, and the cosine similarity.
[0139] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining the initial weights of multiple fault data according to the importance levels of the multiple fault data; determining the corrected weights of the multiple fault data according to the fault thresholds of the multiple fault data and / or the correlation coefficients between the multiple fault data; calculating the Euclidean distance between the real-time vector and the historical vector according to the corrected weights of the multiple fault data.
[0140] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining the fault thresholds of multiple fault data; determining the sensitive factors corresponding to the multiple fault data according to the multiple fault data and the fault thresholds; calculating the corrected weights of the multiple fault data based on the sensitive factors.
[0141] In one embodiment, when the processor executes the computer program, the following steps are further implemented: calculating the correlation coefficient between multiple fault data according to the historical fault data of the multiple fault data; constructing an association matrix according to the correlation coefficient; calculating the corrected weights of the multiple fault data based on the association matrix.
[0142] In one embodiment, when the processor executes the computer program, the following steps are further implemented: dividing the real-time vector and the historical vector into multiple time windows in chronological order according to the preset window size; calculating the weights of the multiple time windows based on the attention mechanism; calculating the cosine similarity between the real-time vector and the historical vector according to the weights of the multiple time windows.
[0143] In one embodiment, when the processor executes the computer program, the following steps are further implemented: in response to the non-existence of a similarity score greater than or equal to the preset score, obtaining the hardware operation data of multiple hardware; cross-verifying the faulty hardware according to the hardware operation data of the multiple hardware; determining the target fault location according to the cross-verification result.
[0144] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: S1: obtaining the hardware operation data corresponding to multiple hardware collected by a controller; S2: preprocessing the multiple hardware operation data according to the importance degree of the multiple hardware operation data; S3: determining the faulty hardware based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; S4: obtaining the real-time fault data and the historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extracting the data features of the real-time fault data and the data features of the historical fault data, and determining the real-time vector and the historical vector; S5: calculating the Euclidean distance and the cosine similarity based on the dynamic weights according to the real-time vector and the historical vector of the multiple candidate fault locations, and determining the target fault location.
[0145] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determining the quantile intervals corresponding to the multiple hardware operation data according to the importance degree of the multiple hardware operation data; preprocessing the multiple hardware operation data according to the quantile intervals.
[0146] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining the life cycle and the real-time load of the hardware corresponding to the multiple hardware operation data; adaptively adjusting the quantile intervals of the multiple hardware operation data based on the life cycles and the real-time loads of the multiple hardware.
[0147] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining general hardware operation data, and preprocessing the general hardware operation data; pre-training a preset long short-term memory network model with the preprocessed general hardware operation data until the preset long short-term memory network model converges to obtain a first model; fine-tuning the first model according to the hardware operation data to obtain a target model; inputting the processed hardware operation data into the target model to output a faulty hardware.
[0148] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: inputting the preprocessed hardware operation data into a preset long short-term memory network model to obtain a loss function; judging whether the preset long short-term memory network model converges according to the loss function; if not converged, updating the parameters of the preset long short-term memory network model through an optimization algorithm.
[0149] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determining the weights of multiple fault data according to the real-time vector and the historical vector, and calculating the Euclidean distance between the real-time vector and the historical vector; determining the weights of multiple time windows according to the real-time vector and the historical vector, and calculating the cosine similarity between the real-time vector and the historical vector based on the dynamic weights; calculating the similarity scores of multiple candidate fault positions according to the Euclidean distance and the cosine similarity; in response to the existence of a similarity score greater than or equal to a preset score, taking the candidate fault position with the highest similarity score as the target fault position.
[0150] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: calculating a fusion coefficient of the Euclidean distance and the cosine similarity through a fault data training set; obtaining a similarity score through weighted summation based on the fusion coefficient, the Euclidean distance, and the cosine similarity.
[0151] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determining the initial weights of multiple fault data according to the importance levels of the multiple fault data; determining the corrected weights of the multiple fault data according to the fault thresholds of the multiple fault data and / or the correlation coefficients between the multiple fault data; calculating the Euclidean distance between the real-time vector and the historical vector according to the corrected weights of the multiple fault data.
[0152] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining the fault thresholds of multiple fault data; determining the sensitive factors corresponding to the multiple fault data according to the multiple fault data and the fault thresholds; calculating the corrected weights of the multiple fault data based on the sensitive factors.
[0153] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: calculating a correlation coefficient between multiple fault data according to historical fault data of the multiple fault data; constructing an association matrix according to the correlation coefficient; and calculating corrected weights of the multiple fault data based on the association matrix.
[0154] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: dividing a real-time vector and a historical vector into multiple time windows in chronological order according to a preset window size; calculating weights of the multiple time windows based on an attention mechanism; and calculating a cosine similarity between the real-time vector and the historical vector according to the weights of the multiple time windows.
[0155] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining hardware operation data of multiple hardware devices in response to the non-existence of a similarity score greater than or equal to a preset score; performing cross-verification on a faulty hardware device according to the hardware operation data of the multiple hardware devices; and determining a target fault location according to the cross-verification result.
[0156] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0158] The above embodiments only illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A hardware fault location method, characterized in that, The method includes: Obtaining hardware operation data corresponding to multiple hardware collected by a controller; Preprocessing the multiple hardware operation data according to the importance levels of the multiple hardware operation data; Based on the preprocessed hardware operation data, determining faulty hardware according to transfer learning and a long short-term memory network model; Obtaining real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extracting data features of the real-time fault data and data features of the historical fault data, and determining a real-time vector and a historical vector; Based on the real-time vector and the historical vector of the multiple candidate fault locations, calculating an Euclidean distance and a cosine similarity based on dynamic weights, and determining a target fault location.
2. The hardware fault location method according to claim 1, wherein The preprocessing the multiple hardware operation data according to the importance levels of the multiple hardware operation data includes: Determining quantile intervals corresponding to the multiple hardware operation data according to the importance levels of the multiple hardware operation data; Preprocessing the multiple hardware operation data according to the quantile intervals.
3. The hardware fault location method according to claim 2, wherein The determining the quantile intervals corresponding to the multiple hardware operation data according to the importance levels of the multiple hardware operation data further includes: Obtaining the life cycles and real-time loads of the hardware corresponding to the multiple hardware operation data; Based on the life cycles and real-time loads of the multiple hardware, adaptively adjusting the quantile intervals of the multiple hardware operation data.
4. The hardware fault location method according to claim 1, wherein The determining faulty hardware according to transfer learning and a long short-term memory network model based on the preprocessed hardware operation data includes: Obtaining general hardware operation data and preprocessing the general hardware operation data; Pre-training a preset long short-term memory network model with the preprocessed general hardware operation data until the preset long short-term memory network model converges to obtain a first model; Fine-tuning the first model according to the hardware operation data to obtain a target model; Inputting the processed hardware operation data into the target model and outputting the faulty hardware.
5. The hardware fault location method according to claim 4, wherein The pre-training the preset long short-term memory network model with the preprocessed general hardware operation data includes: Inputting the preprocessed hardware operation data into the preset long short-term memory network model to obtain a loss function; Judging whether the preset long short-term memory network model converges according to the loss function; If not converged, updating parameters of the preset long short-term memory network model through an optimization algorithm.
6. The hardware fault location method according to claim 1, characterized in that, The calculating an Euclidean distance and a cosine similarity based on dynamic weights, and determining a target fault location based on the real-time vector and the historical vector of the multiple candidate fault locations includes: Determining weights of multiple fault data according to the real-time vector and the historical vector, and calculating the Euclidean distance between the real-time vector and the historical vector; Determining weights of multiple time windows according to the real-time vector and the historical vector, and calculating the cosine similarity between the real-time vector and the historical vector based on dynamic weights; Calculating similarity scores of the multiple candidate fault locations according to the Euclidean distance and the cosine similarity; In response to the existence of a similarity score greater than or equal to a preset score, the candidate fault location with the highest similarity score is used as the target fault location.
7. The hardware fault location method according to claim 6, characterized in that The calculating the similarity scores of the multiple candidate fault locations according to the Euclidean distance and the cosine similarity includes: Calculating a fusion coefficient of the Euclidean distance and the cosine similarity through a fault data training set; Based on the fusion coefficient, the Euclidean distance, and the cosine similarity, obtaining the similarity score through weighted summation.
8. The hardware fault location method according to claim 6, wherein The determining the weights of the multiple fault data according to the real-time vector and the historical vector and calculating the Euclidean distance between the real-time vector and the historical vector includes: Determining initial weights of the multiple fault data according to the importance levels of the multiple fault data; Determining the corrected weights of the multiple fault data according to the fault thresholds of the multiple fault data and / or the correlation coefficients between the multiple fault data; Calculating the Euclidean distance between the real-time vector and the historical vector according to the corrected weights of the multiple fault data.
9. The hardware fault location method according to claim 8, characterized in that, The determining the corrected weights of the multiple fault data according to the fault thresholds of the multiple fault data includes: Obtaining the fault thresholds of the multiple fault data; Determining sensitive factors corresponding to the multiple fault data according to the multiple fault data and the fault thresholds; Calculating the corrected weights of the multiple fault data based on the sensitive factors.
10. The hardware fault location method according to claim 8, wherein The determining the corrected weights of the multiple fault data according to the correlation coefficients between the multiple fault data includes: Calculating the correlation coefficients between the multiple fault data according to the historical fault data of the multiple fault data; Constructing an association matrix according to the correlation coefficients; Calculating the corrected weights of the multiple fault data based on the association matrix.
11. The hardware fault location method according to claim 6, characterized in that, The determining the weights of multiple time windows according to the real-time vector and the historical vector and calculating the cosine similarity between the real-time vector and the historical vector based on dynamic weights includes: Dividing the real-time vector and the historical vector into multiple time windows in chronological order according to a preset window size; Calculating the weights of the multiple time windows based on an attention mechanism; Calculating the cosine similarity between the real-time vector and the historical vector according to the weights of the multiple time windows.
12. The hardware fault location method according to claim 6, wherein The calculating the similarity scores of the multiple candidate fault locations according to the Euclidean distance and the cosine similarity further includes: In response to the non-existence of a similarity score greater than or equal to a preset score, obtaining the hardware operation data of multiple hardware devices; Performing cross-validation on the faulty hardware according to the hardware operation data of the multiple hardware devices; Determining the target fault location according to the cross-validation result.
13. A hardware fault location device, characterized in that, The device includes: A first processing module, configured to obtain the hardware operation data corresponding to multiple hardware devices collected by a controller; A second processing module, configured to preprocess the multiple hardware operation data according to the importance levels of the multiple hardware operation data; A third processing module, configured to determine a faulty hardware device based on the preprocessed hardware operation data according to transfer learning and a long short-term memory network model; A fourth processing module, configured to obtain real-time fault data and historical fault data of multiple candidate fault locations of the faulty hardware within a preset time, extract data features of the real-time fault data and data features of the historical fault data, and determine a real-time vector and a historical vector; A fifth processing module, configured to calculate an Euclidean distance and a cosine similarity based on dynamic weights according to the real-time vector and the historical vector of the multiple candidate fault locations, and determine a target fault location.
14. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the hardware fault location method according to any one of claims 1 to 12 when executing the computer program.
15. A computer-readable storage medium, characterized in that, A computer program is stored in a computer-readable storage medium, wherein the computer program implements the steps of the hardware fault location method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Equipment fault monitoring method and device, computer equipment and storage medium
CN113570473A
Intelligent auxiliary diagnosis and maintenance method and system based on multi-path recall
CN119357787A
Communication network operation and maintenance fault positioning and tracking method and system
CN119420639A
Cited By
Terminal equipment management method and device based on large model and electronic equipment
CN120743614A
Large model-based terminal device management method and apparatus, and electronic device
CN120743614B
Data processing system, device and method for end-side system
CN122527090A