Fault handling method, electronic device, and storage medium

By building a closed-loop learning mechanism of detection-dispatch-feedback-optimization in the data center, and by leveraging the collaborative work of cloud servers and edge servers, the problems of low efficiency and high false alarm rate in traditional operation and maintenance methods have been solved. This has enabled efficient and accurate fault detection and model updates, thereby improving operation and maintenance efficiency.

CN120498979BActive Publication Date: 2026-01-27LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510983937.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2026-01-27
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Traditional data center operations and maintenance rely on manual inspections and static rule-based alarm systems, which are inefficient and have high false alarm and false negative rates. The single training data results in insufficient model generalization ability, making it difficult to accurately detect faults.

Method used

A closed-loop learning mechanism of detection-dispatch-feedback-optimization is constructed. Fault log data reported by edge servers is received by cloud servers, training samples with real labels are constructed, local fault detection models are trained using training samples, and lightweight fault detection models are obtained through knowledge distillation and deployed on edge servers for fault detection.

Benefits of technology

It improves the accuracy of data center fault detection, achieves real-time and precise fault detection, reduces computing resource consumption, and enhances data security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498979B_ABST
    Figure CN120498979B_ABST
Patent Text Reader

Abstract

The application discloses a fault processing method, an electronic device and a storage medium, and relates to the technical field of computers. In order to effectively improve the accuracy of fault detection, the application constructs a training sample with a real label. The training sample is used to train a local fault detection large model. A light fault detection model is obtained by knowledge distillation on the trained local fault detection large model. After the cloud server sends the fault detection model to the edge server, the edge server can use the fault detection model to detect faults in newly collected log data, thereby more accurately detecting faults.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to fault handling methods, electronic devices, and storage media. Background Technology

[0002] Traditional data center operations and maintenance mainly rely on manual inspections and static rule alarm systems. However, manual inspections are inefficient and difficult to detect momentary faults in a timely manner, while static rule alarm systems have a high false alarm rate and a high false alarm rate.

[0003] Currently, some operation and maintenance solutions combine edge computing and federated learning, but the lack of single training data still results in insufficient coverage of diverse fault scenarios, leading to insufficient generalization ability of the model and difficulty in accurately detecting faults.

[0004] Improving the accuracy of fault detection in data centers is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] This application provides a fault handling method, electronic equipment, and storage medium. This application can form a closed-loop learning mechanism of detection-dispatch-feedback-optimization to continuously improve the accuracy of fault detection.

[0006] A fault handling method, applied to a cloud server, includes:

[0007] Receive fault log data reported by the edge server; the fault log data includes log data and fault detection results, wherein the fault detection results are the results obtained by the edge server using the latest fault detection model to detect the log data;

[0008] After sending a fault work order that matches the fault log data to the operation and maintenance terminal, the system receives the actual fault data fed back by the operation and maintenance terminal.

[0009] Using the real fault data, the log data, and the fault detection results, training samples with real labels are constructed;

[0010] Under the condition that the training conditions are met, the local fault detection model is trained using the training samples, and knowledge distillation is performed on the trained local fault detection model to obtain a new fault detection model.

[0011] The new fault detection model is sent to the edge server so that the edge server can use the latest fault detection model to perform fault detection on the newly collected log data.

[0012] A fault handling method, applied to an edge server, includes:

[0013] Collect log data;

[0014] The log data is analyzed using the latest local fault detection model to obtain fault detection results;

[0015] By concatenating the log data and the fault detection results, fault log data is obtained;

[0016] The fault log data is reported to the cloud server so that the cloud server can perform the steps of the fault handling method described above.

[0017] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described fault handling methods when executing the computer program.

[0018] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described fault handling methods.

[0019] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described fault handling methods.

[0020] The cloud server in this application can receive fault log data reported by the edge server. This fault log data includes the log data collected by the edge server implementing edge computing, as well as the fault detection results obtained by the edge server using a local fault detection model. The cloud server can send fault work orders matching the fault log data to the maintenance terminal and receive real fault data fed back by the maintenance terminal. To effectively improve the accuracy of fault detection, this application combines real fault data, log data, and fault detection results to construct training samples with real labels. That is, these training samples can reflect real fault conditions. Under the condition of meeting the training requirements, the local fault detection model is trained using the training samples. To enable the edge server to support the detection model, a lightweight fault detection model can be obtained by knowledge distillation of the trained local fault detection model. After the cloud server sends the fault detection model to the edge server, the edge server can use the fault detection model to perform fault detection on newly collected log data.

[0021] In other words, this application constructs training samples with realistic labels by combining feedback data corresponding to fault work orders. These realistic labels can reflect the actual fault situation. Based on these training samples, a large local fault detection model is trained on a cloud server. Through knowledge distillation, a lightweight fault detection model that can be deployed on edge servers is obtained, enabling edge servers to perform fault detection and thus detect faults more accurately.

[0022] This application has a technical effect that can form a closed-loop learning mechanism of detection-dispatch-feedback-optimization, continuously improving the accuracy of fault detection. Attached Figure Description

[0023] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart of a fault handling method provided in an embodiment of this application;

[0025] Figure 2 A schematic diagram of the structure of an intelligent operation and maintenance and work order coordination system based on federated learning and edge computing provided for an embodiment of this application;

[0026] Figure 3 A flowchart of another fault handling method provided in this application embodiment;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0028] Figure 5 This is a schematic diagram of the specific structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0030] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] Please refer to Figure 1 , Figure 1 This is a flowchart of a fault handling method according to an embodiment of this application. The method is applied to a cloud server, which can be... Figure 2 The server corresponding to the cloud layer in the system shown.

[0033] For details, please refer to Figure 2 , Figure 2 This application provides an intelligent operation and maintenance and work order coordination system based on federated learning and edge computing. The system is mainly divided into cloud and edge components. The cloud component has cloud servers, and the edge component has edge servers. Edge servers are distributed in customer data centers across various locations and can collect hardware and system logs from computing devices, network devices, storage devices, power supply equipment, and cooling equipment in real time. A fault detection model can run on the edge servers, which can be used to detect faults in the collected log data, thereby obtaining fault detection results. Based on the fault detection results, data that needs to be reported is selected. The reported data includes fault detection results with a confidence level greater than a dynamic threshold and the corresponding log data; in this paper, the data reported in this case is simply referred to as fault log data.

[0034] After receiving the fault log data, the cloud server can perform fault detection. If an anomaly is detected, a fault work order is generated and dispatched to the operations and maintenance terminal. This allows the operations and maintenance personnel associated with the terminal to process the fault work order and report the actual fault data.

[0035] The cloud server constructs training samples with real labels based on real fault data, log data, and fault detection results. It then uses these training samples to train a large local fault detection model. After training, the model is compressed into a lightweight fault detection model and distributed to edge servers. This allows for updating of the fault detection model on the edge servers, thereby improving the accuracy of fault detection.

[0036] Specifically, the fault handling method provided in this application includes the following steps.

[0037] S101, Receive fault log data reported by the edge server.

[0038] The fault log data includes log data and fault detection results. The fault detection results are obtained by the edge server using the latest fault detection model to detect the log data.

[0039] Edge servers can collect logs from devices within their data center. These logs include, but are not limited to, hardware logs and system logs. The devices whose logs are collected include, but are not limited to, computing devices, network devices, storage devices, power supply equipment, and cooling equipment.

[0040] When collecting logs, edge servers can set differentiated collection cycles based on device type and log characteristics. For example, BMC (Baseboard Management Controller) logs and accelerator card logs can be collected from servers every 5 minutes, RAID (Redundant Array of Independent Disks) logs every hour, and SMART (Self-Monitoring, Analysis and Reporting Technology) logs daily; network devices can collect hardware and system logs every 10 minutes via SNMP (Simple Network Management Protocol).

[0041] After collecting log data, the edge server can anonymize the log data, removing potentially sensitive user information or other irrelevant data that is not related to fault prediction. Then, the edge server uses its latest local fault detection model to perform fault detection on the log data, thereby obtaining the fault detection results.

[0042] To reduce the pressure on the cloud, after obtaining the fault detection results, the data can be filtered according to the confidence level. Only the fault detection results with a confidence level greater than the dynamic threshold and their corresponding log data are selected and reported to the cloud server.

[0043] When reporting fault log data, edge servers can encrypt and upload this data. During the upload process, the data can also be compressed and uploaded in chunks. Combined with a hash algorithm, the integrity of the chunks can be verified and the order of the chunks can be determined.

[0044] After receiving the fault log data, the cloud server can re-detect the fault in the fault log data. If an anomaly is detected, it indicates that there is a fault problem that needs to be handled, and a corresponding fault work order can be generated at this time.

[0045] In one specific embodiment of this application, receiving fault log data reported by an edge server includes: obtaining compressed blocks corresponding to the fault log data using an encrypted channel with the edge server; assembling the compressed blocks using their corresponding IDs; and then decompressing them to obtain the fault log data. In other words, when reporting fault log data, the edge server can protect the data through encryption or reduce data transmission volume using compression technology. Specifically, Protocol Buffers can be used to define a log upload service interface, the protoc tool can be used to generate gRPC (Remote Procedure Call) client and server code, and OpenSSL can be used to generate root certificates, cloud server certificates, and edge server certificates. The cloud server certificate is deployed in the cloud, the edge server certificate is deployed at the edge, and the root certificate is pre-installed at both the edge and the cloud for two-way authentication. Both the server and edge servers enable TLS 1.3 (Transport Layer Security 1.3), using cipher suites such as TLS_AES_128_GCM_SHA256. Gzip (a data compression algorithm) is used to compress logs from Module 2 with a confidence level greater than a threshold. The compressed logs are then chunked, with each chunk generating a unique chunk_id, recording the chunk's order within the complete log. The edge server records the chunk upload status to the database, adds chunks to be uploaded to a queue, and processes them asynchronously using an independent thread pool to avoid blocking the main thread. Before each transmission, the database is queried, skipping successfully uploaded chunks. If a transmission is interrupted, the failed chunk is recorded, and the failed chunk is retried after a delay. If three consecutive uploads fail, transmission is paused and an alarm is recorded. The cloud server receives the chunks and verifies the hash value (comparing the chunk_id with the calculated value). If verification fails, an error code is returned requesting retransmission. The complete log file is reassembled according to the sequence field order. After reassembly, decompression (Gzip) and cloud inference are triggered.

[0046] S102. After sending a fault work order that matches the fault log data to the operation and maintenance terminal, receive the actual fault data fed back by the operation and maintenance terminal.

[0047] In this embodiment, the maintenance terminal can specifically be a mobile device bound to the maintenance personnel. Fault tickets are sent to the maintenance terminal so that the maintenance personnel can process them, i.e., identify and troubleshoot the fault. During and after processing the fault ticket, the maintenance personnel can report the actual fault situation to the cloud server. In this embodiment, the data / information corresponding to the actual fault situation reported by the maintenance terminal is referred to as real fault data. Real fault data may include whether the fault is real (i.e., whether it is a false alarm), the fault type, the faulty device, and the repair measures.

[0048] In one specific embodiment of this application, sending a fault work order matching the fault log data to the operation and maintenance terminal includes: performing anomaly detection on the fault log data using the latest local fault detection big data model; generating a fault work order corresponding to the anomaly if an anomaly is detected; obtaining the real-time status information and skill information of the operation and maintenance personnel; determining the appropriate target operation and maintenance personnel by combining the real-time status information and skill information; and sending the fault work order to the operation and maintenance terminal that has a binding relationship with the target operation and maintenance personnel.

[0049] Specifically, after receiving logs reported by various data centers (i.e., fault log data reported by edge servers), the cloud server uses a local fault detection model to infer and detect anomalies. When an anomaly is detected, a work order generation process is triggered. For example, based on the fault type and device type, a description template is matched from a preset template library. For example, the template might be: Fault Description: {Data Center Location} {Device Model} Device {Device Location} detected {Fault Type}, Suggested Actions: {Suggested Actions}, Associated Skill Tags: {Skill Tags}. Then, the historical maintenance records of the device are queried from the database, and a work order template is output, such as: {"Device Identifier": "Device 1", "Fault Type": "RAID Card Anomaly", "Fault Description": "Dell device in a certain data center detected a RAID card anomaly", "Suggested Actions": "Please check the physical connection status of the RAID card. If there is no improvement, replace the RAID card", "Skill Tags": ["Storage Device Repair", "RAID Configuration"], "Additional Information": { "Historical Maintenance Records": ["Record 1"]}}.

[0050] After a fault work order is generated, the database is checked to see if a similar work order with an unprocessed status exists, based on the fault type and device ID of the newly generated work order. If it already exists, it will not be reassigned. If it does not exist, the work order will be dynamically assigned based on the real-time status information of the maintenance personnel (such as location and load) and technical information (such as the types of skills mastered in the remarks). Specifically, this is done using the formula... Dynamically assign work orders. Among these, the score is: 1 / Distance: the reciprocal of the geographical distance between the operator's current location and the data center where the faulty device is located; the closer the distance, the larger the value and the higher the score; Skill_Match: the degree of match between the skill tags required for the work order and the operator's skill tags; 1 / Workload: the reciprocal of the operator's current workload; the fewer the number of work orders, the larger the value and the higher the score.

[0051] θ, η, and ξ are weighting coefficients, with default values ​​of 0.5, 0.3, and 0.2, respectively. For example, for an emergency work order: θ=0.7, η=0.2, ξ=0.1; for a complex fault: θ=0.3, η=0.5, ξ=0.2.

[0052] The Haversine formula is used to calculate the straight-line distance between the GPS coordinates of the operations and maintenance personnel and the GPS coordinates of the data center. This is then used to calculate the required skill tags for the work order (e.g., ["Storage Device Repair", "RAID Configuration"]) and the operation and maintenance personnel's skill tags (e.g., ["Storage Device Repair", "Network Equipment Repair"]). Jaccard similarity is then calculated. The reciprocal of the number of work orders currently assigned to the operations and maintenance personnel (e.g., 3) is taken as the load score. The personnel with the highest score are selected to assign work orders, and notifications are pushed through the mobile app (i.e., the operations and maintenance terminal). The work order status is updated to the database in real time (e.g., assigned, processing). When the number of work orders exceeds the system's processing capacity, they are temporarily stored in a buffer queue and assigned according to priority.

[0053] After processing a work order, maintenance personnel report the actual fault type, repair measures, and whether there is a false alarm (i.e., real fault data) via mobile device.

[0054] S103. Construct training samples with real labels using real fault data, log data, and fault detection results.

[0055] After obtaining real fault data, training samples that can reflect real fault conditions can be constructed using real fault data, log data, and fault detection results.

[0056] In other words, after the edge server collects logs, it records the input features of the logs and the model output (logits or soft labels), encapsulating {input feature hash value, fault type logits, confidence level, timestamp}. After processing a fault ticket, maintenance personnel report the actual fault type, remedial measures, and whether there was a false alarm to the cloud via a mobile device. The cloud associates the real fault data with the original logits to generate training samples with real labels.

[0057] Specifically, the cloud receives logits and associated feature hash values ​​uploaded by each edge server, clusters logits by fault type and device type, and builds a global knowledge base to train a large local fault detection model.

[0058] In one specific embodiment of this application, the method further includes: obtaining false alarm samples from the training samples; adjusting the dynamic threshold using the false alarm samples; and sending the dynamic threshold to the edge server so that the edge server only reports fault log data corresponding to fault detection results with a confidence level greater than the dynamic threshold. False alarm samples can then be used to adjust the dynamic threshold corresponding to the fault detection model of the edge server.

[0059] S104. Under the condition that the training conditions are met, the local fault detection model is trained using training samples, and knowledge distillation is performed on the trained local fault detection model to obtain a new fault detection model.

[0060] In this application, training conditions can be set, such as periodic periods, or specific conditions that trigger training, such as at least one of the conditions that the false positive rate is lower than the false positive threshold and the false negative rate is higher than the false negative threshold.

[0061] Under the condition that the training requirements are met, the local fault detection model can be trained using training samples, enabling it to achieve a more accurate fault detection rate. To allow the local fault detection model to run on an edge server, in this embodiment, knowledge distillation can be performed on the local fault detection model to obtain a lightweight new fault detection model.

[0062] In one specific embodiment of this application, training a large local fault detection model using training samples includes: obtaining the divergence of the training samples; if the divergence exceeds a divergence threshold, adding a distribution alignment term to the loss function to enhance the loss function; and training the large local fault detection model using the training samples and the enhanced loss function.

[0063] Regarding the loss function, cross-entropy loss (based on hard labels from work order feedback): ,in, Cross-entropy loss measures the difference between the cloud model's prediction and the actual fault label. The smaller the value, the more accurate the model prediction.

[0064] : The one-hot encoding of the true label, representing the true label of the sample belonging to the i-th type of fault.

[0065] Model prediction probability: The probability value predicted by the cloud model for a sample belonging to the i-th type of fault, ranging from [0,1] and satisfying the following conditions: .

[0066] For logarithmic probability, for predicted probability Take the natural logarithm, when When it is close to the real label, The losses have been reduced.

[0067] KL divergence loss (soft label based on edge model logits): ;in, Divergence loss measures the distribution predicted by the cloud-based model. soft label distribution with edge model The differences.

[0068] : Sum of all fault types, C = total number of fault types.

[0069] : The soft label probability output by the edge model, representing the prediction probability of the edge model for the i-th type of fault.

[0070] Cloud-based model prediction probability; cloud-based prediction probability of the same sample.

[0071] : A measure of probability distribution difference, penalizing situations where the cloud model distribution deviates from the edge model distribution.

[0072] Total loss: (ρ is dynamically adjusted with each training epoch, initially ρ=0.3, later ρ=0.7), where Ltotal: the total mixed loss value, the final optimization target for cloud model training. Lhard: hard label loss, equivalent to cross-entropy loss L. CE Lsoft: Soft label loss, equivalent to KL divergence loss L. KL ρ: Hard label loss weights; 1-ρ: Soft label loss weights.

[0073] Clustered logits can be used as soft labels to guide BERT-Large (Bidirectional Encoder Representations from Transformers, a natural language processing model, corresponding to a large local fault detection model) in learning the knowledge of edge models. Supervised training is performed using real fault labels from work order feedback (such as repair results reported by maintenance personnel). FedAvg aggregation is performed periodically (e.g., weekly) on the first few layers of parameters of each edge model to update the cloud-shared layer parameters, which are then distributed to the edge. After processing a work order, maintenance personnel report the actual fault type, repair measures, and whether there are false alarms via mobile devices. The cloud associates the feedback data with the original logits to generate training samples with real labels. For false alarm samples, the classification threshold of the edge model is adjusted; for missed alarm samples, the loss weight of that category is increased during cloud training. Global model retraining is triggered periodically (e.g., weekly) to calculate the Jensen-Shannon divergence (JSD). , ;in, Jensen-Shannon divergence of high / low confidence sample distributions measures the degree of difference between the probability distributions of the two classes of samples. The larger the value, the more significant the difference in distribution. This is the probability distribution of fault categories for high-confidence samples. When the edge model infers local logs, it represents the probability distribution of fault categories for samples whose confidence scores exceed the dynamic threshold. : Fault category probability distribution for low-confidence samples; probability distribution of fault categories for samples with a margin model confidence score less than the dynamic threshold. M: Intermediate distribution between high / low confidence distributions. and The arithmetic mean distribution.

[0074] The KL divergence of the distribution relative to M measures the degree of difference between the high-confidence distribution and the intermediate distribution.

[0075] The KL divergence of the distribution relative to M measures the degree of difference between the low-confidence distribution and the intermediate distribution.

[0076] The weighting coefficients are averaged over the KL divergence of the two terms to ensure the symmetry of the JSD.

[0077] Distribution alignment is triggered when JSD exceeds a threshold (e.g., 0.3).

[0078] Loss function enhancement, that is, in the original Based on this, add a distribution alignment term: ;in: This represents the distribution alignment loss value; These are the weighting coefficients for the MMD loss; The weighting coefficients of the loss; Maximum mean difference measures the distribution distance between high / low confidence samples in the feature space; Attention cosine similarity measures the consistency of the model's attention focus on high / low confidence samples; MMD (maximum mean difference) measures the difference in feature space distribution between high / low confidence samples, and is calculated as follows:

[0079] Where, Xhigh: high-confidence sample set, samples with marginal model confidence > threshold. Xlow: low-confidence sample set, samples with marginal model confidence ≤ threshold. Nhigh, Nlow: number of samples, Nhigh = m, Nlow = n. Kernel function mapping maps input features to the reproducing kernel Hilbert space (RKHS). Mean embedding of high-confidence samples in RKHS. ∶ Low-confidence samples in the mean embedding of RKHS. Hilbert space norm: calculates the distance between two mean embeddings.

[0080] CosSim (cosine similarity) calculates the attention similarity of BERT-Large samples to high / low confidence levels. ;in, Attention feature vectors of the samples, extracted from the cloud model; Vector dot product measures directional similarity; Norm, a normalization factor, constrains the output range. Weighting coefficients, for example, λ=0.2, μ=0.1.

[0081] Final loss function: .

[0082] Regarding the knowledge distillation of the large local fault detection model to obtain a lightweight fault detection model, the specific implementation includes: using the updated BERT-Large as the teacher model, and re-distilling to generate a lightweight edge model. For example, progressive distillation can be adopted: first round: distill to BERT-Medium (12 layers) to retain more teacher knowledge; second round: further compress to BERT-Tiny (6 layers) to adapt to edge servers.

[0083] S105. Send the new fault detection model to the edge server so that the edge server can use the latest fault detection model to perform fault detection on the newly collected log data.

[0084] In one specific embodiment of this application, the teacher model parameters in the cloud are compared with the current model parameters at the edge, the changed parameters are extracted, the difference data is compressed using Gzip, and then sent to the edge server using a Protocol Buffers-defined transmission protocol. In other words, when sending the fault detection model to the edge server, it is not necessary to send the complete fault detection model; only the latest fault detection model currently used by the edge server and the newly obtained fault detection model from the cloud (i.e., the teacher model parameters in the cloud) are needed to determine the changed parameters. These changed parameters are then compressed and sent to the edge server to reduce data transmission volume, reduce system resource consumption, and ensure effective business operation. The edge server adjusts its local fault detection model based on these changed parameters to update the fault detection model, enabling the edge server to use the latest fault detection model to perform fault detection on newly collected log data.

[0085] The cloud server in this application can receive fault log data reported by the edge server. This fault log data includes the log data collected by the edge server implementing edge computing, as well as the fault detection results obtained by the edge server using a local fault detection model. The cloud server can send fault work orders matching the fault log data to the maintenance terminal and receive real fault data fed back by the maintenance terminal. To effectively improve the accuracy of fault detection, this application combines real fault data, log data, and fault detection results to construct training samples with real labels. That is, these training samples can reflect real fault conditions. Under the condition of meeting the training requirements, the local fault detection model is trained using the training samples. To enable the edge server to support the detection model, a lightweight fault detection model can be obtained by knowledge distillation of the trained local fault detection model. After the cloud server sends the fault detection model to the edge server, the edge server can use the fault detection model to perform fault detection on newly collected log data.

[0086] In other words, this application constructs training samples with realistic labels by combining feedback data corresponding to fault work orders. These realistic labels can reflect the actual fault situation. Based on these training samples, a large local fault detection model is trained on a cloud server. Through knowledge distillation, a lightweight fault detection model that can be deployed on edge servers is obtained, enabling edge servers to perform fault detection and thus detect faults more accurately.

[0087] Please refer to Figure 3 , Figure 3 This is a flowchart of another fault handling method in an embodiment of this application, which can be applied to, for example... Figure 2 The method for an edge server within the edge layer of the system shown includes the following steps.

[0088] S201. Collect log data.

[0089] Log data refers to the logs corresponding to several devices monitored by the edge server, including but not limited to hardware logs and system logs.

[0090] In this embodiment, different collection frequencies can be used for different types of logs.

[0091] In one specific embodiment of this application, collecting log data includes: collecting hardware logs and system logs of the target device within the data center where the edge server is located; searching for content to be hidden in the hardware logs and system logs using a business-sensitive word library and a matching algorithm; replacing the content to be hidden in the hardware logs and system logs; and formatting the hardware logs and system logs to obtain log data. The content to be hidden can be data that requires privacy protection and is unrelated to fault prediction.

[0092] Edge servers are deployed in customer data centers across various locations to collect hardware and system logs from computing, network, storage, power supply, and cooling equipment in real time. Differentiated collection cycles are set based on device type and log characteristics (e.g., servers collect BMC and accelerator card logs every 5 minutes, RAID logs every hour, and SMART logs daily; network devices collect hardware and system logs every 10 minutes via SNMP). Logs used for fault prediction generally do not involve customer application or business data. To ensure customer data security, sensitive content is filtered from collected logs using a multi-pattern matching algorithm (Aho-Corasick) based on a pre-defined business-sensitive keyword library (e.g., customer names, internal database table names, business paths).

[0093] Specifically, keywords are broken down by character and a trie is constructed. A breadth-first search (BFS) is used to traverse the trie and calculate a failure pointer for each node. The log content is then traversed for matching. For the matched sensitive content (i.e., a specific type of content to be hidden), the start and end positions of the sensitive words in the text are obtained. The characters at the start-1 and end+1 positions are checked to see if they are boundary characters (such as spaces, punctuation marks, or special symbols). If they meet the boundary conditions, they are marked as valid matches. The valid sensitive content is then replaced, and the filtered log is formatted.

[0094] Log formatting can be structured as follows:

[0095] {"device_id": "Server_01",

[0096] "log_type": "BMC SEL",

[0097] "timestamp": "2023-10-01 14:30:00",

[0098] "log_detail":"3ef9 | Processor #0xff | IERR | Asserted "

[0099] }

[0100] The meaning of this log structure is as follows:

[0101] {"Device ID": "Server_01",

[0102] Log type: "BMC System Event Log",

[0103] "Timestamp": "2023-10-01 14:30:00",

[0104] "Log details": "3ef9 | Processor #0xff | IERR (Internal Error) | Asserted"}.

[0105] In practical applications, log collection and processing can be separated into independent threads and executed in batches through an asynchronous thread pool to improve efficiency.

[0106] S202. Use the latest local fault detection model to detect log data and obtain fault detection results.

[0107] A fault detection model is deployed on the edge server. This fault detection model can be trained on the cloud server and then the pre-trained BERT-Large model can be distilled into a lightweight model through knowledge and deployed on the edge server. When device log information is collected, the model can infer and detect anomalies from the log information.

[0108] S203. Combine log data and fault detection results to obtain fault log data.

[0109] Log data is concatenated with fault detection results to obtain fault log data. This fault log data includes not only log data but also fault detection results.

[0110] In one specific embodiment of this application, concatenating log data and fault detection results to obtain fault log data includes: if the confidence level of the fault detection result is greater than a dynamic threshold, then concatenating log data and fault detection results to obtain fault log data; wherein, the process of determining the dynamic threshold includes: obtaining local false negative rate, false positive rate, and alarm statistics, and adjusting the old threshold using the false negative rate, false positive rate, and alarm statistics to obtain a new threshold; or, in the case of receiving a dynamic threshold sent by a cloud server, determining the received dynamic threshold as the new threshold.

[0111] When the confidence level output by the fault detection model exceeds the dynamic threshold, the log data and inference results are uploaded to the cloud for collaborative processing; otherwise, an alarm notification is pushed to the data center contact person through the message queue, and the feature hash is uploaded to the cloud.

[0112] Regarding the generation of dynamic thresholds: Dynamic thresholds are generated by considering the device type, fault level (fatal / critical / warning / general), and criticality (core / critical / important / general). The initial value is calculated by adding a criticality correction (core devices -10%, general devices +5%) and a fault level correction (fatal -10%, general +5%) to a baseline value (70%). The calculation results are then used to create a mapping table. Every 24 hours, dynamic thresholds are adjusted based on the false alarm rate (FP Rate) and false negative rate (FN Rate) according to the following adjustment formula: Where FP is the number of false alarms and FP Rate is the false alarm rate; FN is the number of false alarms and FN Rate is the false alarm rate. The target FP Rate and target FN Rate are pre-set ideal target values ​​for false alarm rate and false alarm rate, used to adjust the dynamic thresholds to achieve a balance between false alarms and false alarms in the system.

[0113] FP and FN are used to calculate the actual FP Rate and FN Rate. These actual values ​​are then compared with the target values ​​to generate a threshold adjustment signal.

[0114] Adaptive optimization is performed based on dynamic thresholds, where A: the total number of logs that trigger alarms daily (i.e., the number of logs with confidence > threshold), B: the total number of logs that do not trigger alarms daily (i.e., the number of logs with confidence ≤ threshold), γ: the proportion of actual faults among those that do not trigger alarms (default γ = 5%), and adjustment coefficients α and β control the sensitivity of false alarms and false negatives respectively (default α = 0.3, β = 0.5). A threshold range is set (e.g., 50%~80%) to prevent over-adjustment.

[0115] S204. Report the fault log data to the cloud server so that the cloud server can perform the following actions: Figure 1 The steps of the described fault handling method.

[0116] After collecting logs, the edge servers record the input features and model output (logits or soft labels), encapsulate the content into {input feature hash value, fault type logits, confidence score, timestamp}, and upload it. The logits are compressed using sparse coding (e.g., retaining the top-3 probability values) to reduce transmission volume, and then uploaded to the cloud via a gRPC+TLS encrypted channel. The cloud receives the logits and associated feature hash values ​​uploaded by each edge server, clusters the logits by fault type and device type, and builds a global knowledge base.

[0117] The cloud server can pre-train BERT-Large using historical global fault log data (such as public datasets or early anonymized data) as the initial teacher model. The training objective is fault classification (such as fault type identification). Knowledge distillation is performed on BERT-Large to generate a lightweight edge server model. The first four layers (word embeddings + the first three Transformer layers) of the cloud model and the edge model are forced to be consistent, while subsequent layers are independent. The edge server uses locally collected logs to fine-tune the lightweight model. The training objective is to minimize the loss function of local fault detection (such as cross-entropy loss) and update the model parameters. The training frequency is adjusted according to the device log collection cycle; for example, high-frequency logs trigger fine-tuning hourly, while low-frequency logs trigger fine-tuning daily.

[0118] Fine-tuning refers to performing small-scale training on the fault detection model sent by the cloud server for the fault detection task, thereby making minor adjustments to the parameters of the fault detection model and ultimately obtaining a fault detection model that is adapted to the fault detection task and data.

[0119] It should be noted that the technical content and effects related to the fault handling method provided in this embodiment are different from those of the previous embodiment. Figure 1 The corresponding troubleshooting methods can be used interchangeably, and will not be elaborated on here.

[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0121] Corresponding to the above method embodiments, this application also provides an electronic device. The electronic device described below and the fault handling method described above can be referred to each other.

[0122] See Figure 4 As shown, the electronic device includes:

[0123] Memory 332 is used to store computer programs;

[0124] The processor 322 is used to implement the steps of the fault handling method in the above method embodiment when executing a computer program.

[0125] For details, please refer to Figure 5 , Figure 5This is a schematic diagram of the specific structure of an electronic device provided in this embodiment. The electronic device can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and a memory 332. The memory 332 stores one or more computer programs 342 or data 344. The memory 332 can be temporary or permanent storage. The program stored in the memory 332 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the data processing device. Furthermore, the processor 322 may be configured to communicate with the memory 332 and execute the series of instruction operations stored in the memory 332 on the electronic device 301.

[0126] Electronic device 301 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.

[0127] The steps in the fault handling method described above can be implemented by the structure of the electronic device.

[0128] Specifically, this electronic device can be a server, and when deployed at the edge layer, the server can implement, for example... Figure 1 The steps of the troubleshooting method shown; when the server is deployed in the cloud layer, the server can implement as follows Figure 3 The steps of the fault handling method shown are as follows.

[0129] It simultaneously possesses edge servers deployed at the edge layer and cloud servers deployed at the cloud layer, and implements corresponding fault handling methods for each, which is... Figure 2 The system shown.

[0130] In other words, Figure 2 In the system shown, implementations such as... can be carried out in the edge server. Figure 3 The steps of the fault handling method shown are implemented simultaneously on the cloud server as follows. Figure 1The steps of the fault handling method are shown. This improves the real-time performance and accuracy of fault detection. By deploying a lightweight fault detection model on edge nodes and combining it with a dynamic threshold optimization algorithm, local real-time anomaly detection is achieved, significantly reducing the latency issues of traditional centralized cloud processing. Simultaneously, knowledge distillation technology is used to compress the large cloud model into a lightweight model suitable for edge devices, reducing computational resource consumption while maintaining detection accuracy, effectively addressing the need for rapid response to transient faults. Enhanced data security and privacy protection: Sensitive content in logs is accurately identified and anonymized using multi-pattern matching algorithms (such as Aho-Corasick) combined with boundary character verification, avoiding misjudgments or omissions in traditional regular expression matching. Furthermore, chunked transmission, asynchronous retries, and optimized encrypted communication are employed to significantly reduce the computational and communication overhead of edge devices while ensuring data transmission security, ensuring efficient system operation in resource-constrained environments. Cross-data center collaborative learning and model generalization optimization are achieved. Through a federated learning framework, the cloud aggregates feature hash values ​​and model outputs (logits) uploaded by each edge node. A hybrid loss function (cross-entropy + KL divergence + distribution alignment term) is then used to dynamically optimize the global model, significantly improving its generalization ability. Simultaneously, sparse coding compression technology reduces data transmission volume, enabling efficient knowledge sharing and collaborative evolution while protecting the data privacy of edge nodes. The work order allocation strategy is optimized: based on a dynamic scoring algorithm, considering the real-time location, skill matching, and workload of maintenance personnel, accurate work order allocation is achieved, significantly shortening the average fault repair time and improving customer satisfaction. A closed-loop learning mechanism is constructed for continuous system optimization: feedback data from work order processing results (such as false alarm / missed alarm markers) dynamically adjusts the edge model thresholds and adds loss weights during cloud training, forming a closed-loop learning mechanism of detection-dispatch-feedback-optimization. This mechanism continuously improves model accuracy, prevents the recurrence of similar faults, and achieves system self-improvement and long-term stability.

[0131] Corresponding to the above method embodiments, this application also provides a readable storage medium. The readable storage medium described below corresponds to the fault handling method described above. This application also provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the fault handling method embodiments described above when run.

[0132] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0133] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described fault handling method embodiments.

[0134] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described fault handling method embodiments.

[0135] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0136] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A fault handling method, characterized in that, Applications to cloud servers include: Receive fault log data reported by the edge server; the fault log data includes log data and fault detection results, wherein the fault detection results are the results obtained by the edge server using the latest fault detection model to detect the log data; After sending a fault work order matching the fault log data to the operation and maintenance terminal, the system receives the actual fault data fed back by the operation and maintenance terminal; wherein, the actual fault data includes the actual fault type, repair measures, and whether it is a false alarm; Using the real fault data, the log data, and the fault detection results, training samples with real labels are constructed; Under the condition that the training conditions are met, the local fault detection model is trained using the training samples, and knowledge distillation is performed on the trained local fault detection model to obtain a new fault detection model; the training samples include missed samples, which are used to improve the loss weight of the fault type to which the missed sample belongs when training in the cloud. The new fault detection model is sent to the edge server so that the edge server can use the latest fault detection model to perform fault detection on the newly collected log data; This includes sending a fault work order matching the fault log data to the operation and maintenance terminal, including: The latest local fault detection model is used to detect anomalies in the fault log data. When an anomaly is detected, a fault work order corresponding to the anomaly is generated; based on the fault type and equipment type, a description template is matched from a preset template library, and the historical maintenance records of the equipment are queried from the database to output a work order template; wherein, the template includes: fault description, suggested measures, and associated skill tags; Obtain real-time status and skill information of operation and maintenance personnel; real-time status information includes location and load, and skill information refers to the types of skills they have mastered; By combining the real-time status information and the skill information, suitable target maintenance personnel can be determined; The fault work order is sent to the maintenance terminal that is bound to the target maintenance personnel; The training of the local fault detection model using the training samples includes: Jensen-Shannon divergence (JSD): , ; in, Jensen-Shannon divergence of high / low confidence sample distributions measures the degree of difference between the probability distributions of the two classes of samples; For high-confidence samples, the probability distribution of fault categories is given. When the edge model infers local logs, the probability distribution of fault categories for samples whose confidence scores exceed the dynamic threshold is given. : Fault category probability distribution for low-confidence samples; probability distribution of fault categories for samples with a margin model confidence score less than the dynamic threshold; M: intermediate distribution between high / low confidence distributions. and The arithmetic mean distribution; The KL divergence of the distribution relative to M measures the degree of difference between the high-confidence distribution and the intermediate distribution; The KL divergence of the distribution relative to M measures the degree of difference between the low-confidence distribution and the intermediate distribution; The weighting coefficients are averaged over the KL divergence of the two terms to ensure the symmetry of the JSD. When the JSD exceeds the threshold, distribution alignment is triggered; Loss function enhancement, in Based on this, add a distribution alignment term: Where: Ltotal is the total mixed loss, Lhard is the cross-entropy loss, and Lsoft is the KL divergence loss; This represents the distribution alignment loss value; These are the weighting coefficients for the MMD loss; The weighting coefficients of the loss; Maximum mean difference measures the distribution distance between high / low confidence samples in the feature space; Attention cosine similarity measures the consistency of the model's attention focus on high / low confidence samples; MMD maximum mean difference measures the difference in feature space distribution between high / low confidence samples, calculated as follows: Where, Xhigh: high-confidence sample set, samples with margin model confidence > threshold; Xlow: low-confidence sample set, samples with margin model confidence ≤ threshold; Nhigh, Nlow: number of samples, Nhigh=m, Nlow=n; Kernel function mapping maps input features to the reproducing kernel Hilbert space (RKHS). Mean embedding of high-confidence samples in RKHS; : Mean embedding of low-confidence samples in RKHS; Hilbert space norm: used to calculate the distance between two mean embeddings; CosSim (cosine similarity) calculates the attention similarity of BERT-Large samples to high / low confidence levels. ;in, Attention feature vectors of the samples, extracted from the cloud model; Vector dot product measures directional similarity; Norm, normalization factor, constrains the output range; weighting coefficients λ=0.2, μ=0.1; Final loss function: .

2. The method according to claim 1, characterized in that, The local fault detection model is trained using the training samples, including: Obtain the divergence of the training samples; If the divergence exceeds the divergence threshold, a distribution alignment term is added to the loss function to enhance the loss function; The local fault detection model is trained using the training samples and the enhanced loss function.

3. The method according to claim 1, characterized in that, Receive fault log data reported by the edge server, including: Using the encrypted channel with the edge server, compressed blocks corresponding to the fault log data are obtained; Using the IDs corresponding to the compressed blocks, the compressed blocks are assembled and then decompressed to obtain the fault log data.

4. The method according to any one of claims 1 to 3, characterized in that, Also includes: False positive samples are obtained from the training samples; Adjust the dynamic threshold using the false alarm samples; The dynamic threshold is sent to the edge server so that the edge server only reports fault log data corresponding to fault detection results with a confidence level greater than the dynamic threshold.

5. A fault handling method, characterized in that, Applications to edge servers include: Collect log data; The log data is analyzed using the latest local fault detection model to obtain fault detection results; By concatenating the log data and the fault detection results, fault log data is obtained; The fault log data is reported to the cloud server so that the cloud server can perform the steps of the fault handling method as described in any one of claims 1 to 4.

6. The method according to claim 5, characterized in that, Collect log data, including: Collect hardware logs and system logs of the target devices within the data center where the edge server is located; By combining a business-sensitive word database and a matching algorithm, the content to be hidden is searched in the hardware logs and the system logs; Replace the content to be hidden in the hardware log and the system log; The hardware log and the system log are formatted to obtain the log data.

7. The method according to claim 5 or 6, characterized in that, By concatenating the log data and the fault detection results, fault log data is obtained, including: If the confidence level of the fault detection result is greater than the dynamic threshold, then the log data and the fault detection result are concatenated to obtain the fault log data; The process of determining the dynamic threshold includes: Obtain the local false negative rate, false positive rate, and alarm statistics, and adjust the old threshold using the false negative rate, false positive rate, and alarm statistics to obtain a new threshold; Alternatively, if a dynamic threshold is received from the cloud server, the received dynamic threshold may be determined as the new threshold.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the fault handling method as described in any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the fault handling method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Transformer fault diagnosis method based on knowledge distillation and collaborative incremental learning

    CN119202705A

  • Fault diagnosis data labeling method and device

    CN119202708A