Electronic computer fault detection method and system based on cloud computing

By building a virtual environment with diverse fault scenarios, combining cloud computing and virtualization technology, using XGBoost predictive analysis and fault classification model, the problem of inaccurate fault detection of electronic computers in the existing technology is solved, and fast and accurate fault detection and positioning is achieved to ensure the stable operation of the system.

CN120492192AActive Publication Date: 2025-08-15SHENZHEN JIE ENTROPY TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510568071.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In the prior art, the accuracy of electronic computer fault detection is not high, especially for complex hardware failures or intermittent failures, it is difficult to accurately detect and locate.

Method used

By building a virtual environment containing diverse failure scenarios, the virtualization technology and containerization technology of the cloud computing platform are used to simulate hardware and software failure scenarios, combined with XGBoost prediction and analysis of the expected operating data of computer nodes, calculate the comprehensive deviation value, and use a pre-trained fault classification model to match and locate fault types.

Benefits of technology

Improve the accuracy and efficiency of fault detection, ensure the system to quickly detect and locate faults, reduce downtime, and ensure the stable operation of critical infrastructure and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492192A_ABST
    Figure CN120492192A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic computer fault detection method and system based on cloud computing, and the method comprises the steps: constructing a virtual environment containing diversified fault scenes, carrying out the prediction, obtaining the expected operation data corresponding to a performance index in the virtual environment, obtaining the real-time operation data of the system, and obtaining the real-time operation data of the system; calculating a comprehensive deviation value of the expected operation data and the real-time operation data in combination with the scene weighting factor and the time attenuation coefficient, calculating a priority score according to the comprehensive deviation value in combination with the influence level of the fault scene on the service continuity, and generating a trigger task queue sorted according to the priority score; simulating a fault scene corresponding to the fault detection task with the highest priority score through the virtual environment, and obtaining simulation operation data; and performing fault matching on the simulated operation data by adopting a pre-trained fault classification model, and outputting a fault type and fault positioning information. According to the invention, the fault of the electronic computer can be accurately detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a computer fault detection method and system based on cloud computing. Background Art

[0002] Computer system reliability research is a core area of modern information technology, directly related to the stable operation of critical infrastructure and data security. As computer systems become increasingly complex, fault detection and diagnosis become key to ensuring high system availability.

[0003] In one existing technology, fault detection is mainly performed using power-on self-test and hardware diagnostic software. Error and warning information in the system log is viewed through the system's built-in event viewer or log analysis tool. The computer's network connection status is checked using a network diagnostic tool. Functional tests are performed on the operating system and application programs to check for software faults.

[0004] However, some diagnostic software in the prior art needs to be run in a specific environment, and may not be able to accurately detect or locate some complex hardware faults or intermittent faults, resulting in low accuracy in fault detection of electronic computers. Summary of the Invention

[0005] The present invention provides a computer fault detection method and system based on cloud computing to solve the problem of low accuracy of computer fault detection in the prior art.

[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a computer fault detection method based on cloud computing, comprising:

[0007] Use virtualization technology to build a virtual environment that includes various fault scenarios, obtain system operation logs and hardware performance indicators, and predict the expected operation data corresponding to the performance indicators in the virtual environment;

[0008] Acquire real-time operating data of the system, and calculate based on the scenario weighting factor, the time decay coefficient, and the expected operating data to obtain a comprehensive deviation value;

[0009] Calculating a priority score based on the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the level of impact of the fault scenario on business continuity, and generating a trigger task queue sorted by priority score based on the priority score; the trigger task queue is a collection of fault detection tasks including trigger conditions, priority scores, and target fault scenarios;

[0010] Extracting the highest priority fault detection task from the trigger task queue, and executing the corresponding fault scenario simulation in the virtual environment to obtain simulated operation data of the fault scenario;

[0011] A pre-trained fault classification model is used to perform fault pattern matching on the simulated operation data, and fault type and fault location information are output.

[0012] Preferably, the process of constructing a virtual environment containing diverse fault scenarios through virtualization technology, obtaining system operation logs and hardware performance indicators, and predicting expected operation data corresponding to the performance indicators in the virtual environment includes:

[0013] Based on the virtualized resources of the cloud computing platform, containerization technology is used to deploy simulation nodes, configure multi-level hardware structures and heterogeneous operating systems;

[0014] By injecting hardware and software failure scenarios and coupling them through tools and fault injection frameworks, a virtual environment containing diverse failure scenarios is obtained;

[0015] Obtain the operation logs of each node in the virtual environment through log collection tools, and use monitoring tools to collect hardware performance indicators in real time;

[0016] Extract key state features using a feature extraction algorithm based on the node operation log and the hardware performance indicators to obtain a state feature set;

[0017] According to the state feature set, the structured features are processed by XGBoost to predict the expected operating state of each node in the virtual environment and generate expected operating data.

[0018] Preferably, the real-time operation data of the acquisition system is calculated in combination with the scenario weighting factor, the time decay coefficient and the expected operation data to obtain a comprehensive deviation value, including:

[0019] Based on the deployed containerized nodes, log collection tools are used to extract operation logs containing timestamps to obtain real-time operation data of the system.

[0020] Based on the impact of the fault scenario on the business, the initial weight is assigned and the weight is dynamically updated based on the real-time business load or external events to obtain the scenario weighting factor;

[0021] Obtaining the time interval from the occurrence of the event to the current calculation, and calculating the time decay coefficient based on the time interval and a preset decay rate;

[0022] A comprehensive deviation value is calculated based on the real-time operation data, the expected operation data, the combined scenario weighting factor and the time attenuation coefficient.

[0023] Preferably, the calculation formula of the time decay coefficient is:

[0024] β=e -λt

[0025] Where β is the time attenuation coefficient, λ is the attenuation rate, and the unit is S -1 , t is the time interval from the occurrence of the event to the current calculation, in S;

[0026] The calculation formula of the comprehensive deviation value is:

[0027]

[0028] Among them, Δ is the comprehensive deviation value; w is the scene weighting factor; β is the time attenuation coefficient; Now is the real-time operation value; Pre is the expected operation value; and both Now and Pre are values between 0 and 1.

[0029] Preferably, the calculating of the priority score according to the magnitude by which the comprehensive deviation value exceeds the preset deviation threshold and the impact level of the fault scenario on business continuity, and generating a trigger task queue sorted by the priority score according to the priority score, includes:

[0030] Calculating the deviation amplitude based on the comprehensive deviation value and the preset deviation threshold;

[0031] Classify the impact of different failure scenarios on business continuity and obtain weight scores corresponding to different levels;

[0032] Calculate the priority scores of different fault scenarios based on the deviation magnitude and the weight score;

[0033] A fault detection task including a trigger condition, a priority, and a target fault scenario is generated, and all the fault detection tasks are sorted according to the priority score to obtain a trigger task queue.

[0034] Preferably, extracting the highest priority fault detection task from the trigger task queue, executing the corresponding fault scenario simulation in a virtual environment, and obtaining simulated operation data of the fault scenario includes:

[0035] Obtain the fault detection task with the highest priority score from the trigger task queue to obtain the fault scenario to be simulated;

[0036] According to the fault scenario to be simulated, corresponding operating parameters are configured and loaded in the virtual environment to obtain a virtual environment of the fault scenario to be simulated;

[0037] Simulating the fault scenario to be simulated in the virtual environment, and using a monitoring tool to collect operation logs and hardware performance indicators of the containerized node in real time to obtain simulated operation data;

[0038] After the simulation is completed, the virtual environment is reset to allow for the next fault simulation.

[0039] Preferably, the use of a pre-trained fault classification model to perform fault pattern matching on the simulated operation data and outputting fault type and fault location information includes:

[0040] Extracting the operation logs of the containerized nodes in the simulation operation data using a log collection tool based on the simulation operation data, and performing structured processing using a log parsing tool to obtain structured simulation operation data;

[0041] Using a pre-trained fault classification model to match the structured simulation operation data to obtain a fault pattern matching result;

[0042] According to the fault pattern matching results, a rule engine is used to classify the matched fault patterns, and the corresponding fault types are output through a preset fault type mapping table;

[0043] The node identifier of the containerized node in the simulation operation data is extracted to obtain fault location information including the node name and IP address.

[0044] Preferably, the training process of the pre-trained fault classification model includes:

[0045] Obtain historical system fault data, diverse fault data coupled with the virtual environment, and open source fault data as pre-training data for the model;

[0046] Labeling the corresponding fault types and location information for different fault data in the pre-training data;

[0047] According to the pre-training data, combined with the fault type and the location information, a dual-channel deep learning architecture is used to train the model;

[0048] Calculate the precision, recall and F1 score indicators as evaluation results;

[0049] A regularization method is adopted to adjust the model according to the evaluation results to obtain a trained fault classification model.

[0050] Preferably, the accuracy, recall and F1 score are calculated using the following formula:

[0051] The calculation formula of the accuracy is:

[0052] Accuracy=(TP+TN) / (TP+TN+FP+FN)

[0053] The calculation formula of the recall rate Recall is:

[0054] Recall = TP / (TP+FN)

[0055] The calculation formula of the F1 score F1-Score is:

[0056]

[0057] Among them, TP is a true positive example, TN is a true negative example, FP is a false positive example, FN is a false negative example, F1-Score is the harmonic mean of precision and recall, and precision = TP / (TP+FP).

[0058] In a second aspect, the present invention provides an electronic computer fault detection system based on cloud computing, comprising:

[0059] The virtual environment construction module is used to build a virtual environment that includes various fault scenarios through virtualization technology, obtain system operation logs and hardware performance indicators, and predict the expected operation data corresponding to the performance indicators in the virtual environment;

[0060] An operation deviation calculation module is used to obtain real-time operation data of the system, and calculate the comprehensive deviation value by combining the scenario weighting factor, the time decay coefficient and the expected operation data;

[0061] a trigger task generation module, configured to calculate a priority score based on the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the level of impact of the fault scenario on business continuity, and generate a trigger task queue sorted by priority score based on the priority score; the trigger task queue is a collection of fault detection tasks including trigger conditions, priority scores, and target fault scenarios;

[0062] A fault scenario simulation module is used to extract the highest priority fault detection task from the trigger task queue, and execute the corresponding fault scenario simulation in a virtual environment to obtain simulated operation data of the fault scenario;

[0063] The matching result output module is used to use a pre-trained fault classification model to perform fault pattern matching on the simulated operation data and output fault type and fault location information.

[0064] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements any one of the above-mentioned cloud computing-based electronic computer fault detection methods.

[0065] In a fourth aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned cloud computing-based electronic computer fault detection methods.

[0066] Compared with the prior art, the present invention has the following beneficial effects:

[0067] 1. The present invention integrates cloud computing and virtualization technologies to build a virtual environment that includes multiple fault scenarios. It combines containerized node simulation with XGBoost to predict and analyze the expected operating value of the indicator data of each node of the computer, and then calculates the comprehensive deviation value of the indicator data of each node in combination with the real-time operating value. The nodes corresponding to the indicator data whose comprehensive deviation value exceeds the preset deviation threshold are marked as potential fault nodes, and a fault detection task is generated and simulated in the virtual environment. According to the simulated operating data and combined with the fault classification model, the fault type can be matched more accurately and the fault location can be located, thereby improving the accuracy of fault detection.

[0068] 2. The virtual environment of the present invention includes various fault scenarios, which can fully simulate the fault conditions that may occur during system operation. It monitors the operating data of each deployed containerized node and can quickly and accurately locate the fault node after matching the fault mode using a pre-trained and continuously optimized fault classification model, thereby improving the efficiency of fault detection.

[0069] 3. The present invention ensures that the system can be quickly detected after a fault occurs and then promptly and effectively processed by quickly and accurately detecting the fault type and fault location information of the electronic computer, and can be restored to normal operation as soon as possible, thereby reducing system downtime and data loss risks caused by faults, ensuring the stable operation and data security of key infrastructure, and improving the overall reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 The present invention is a flowchart of a method for detecting electronic computer faults based on cloud computing.

[0071] Figure 2 The figure is a module diagram of a cloud computing-based electronic computer fault detection system of the present invention. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0073] Reference Figure 1 The first embodiment of the present invention provides a flowchart of a method for detecting computer faults based on cloud computing, comprising the following steps:

[0074] S11, using virtualization technology to build a virtual environment that includes various fault scenarios, obtain system operation logs and hardware performance indicators, and predict the expected operation data corresponding to the performance indicators in the virtual environment;

[0075] S12, obtaining real-time operating data of the system, and performing calculations based on the scenario weighting factor, the time decay coefficient, and the expected operating data to obtain a comprehensive deviation value;

[0076] S13, calculating a priority score based on the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the impact level of the fault scenario on business continuity, and generating a trigger task queue sorted by priority score based on the priority score; the trigger task queue is a collection of fault detection tasks including trigger conditions, priority scores, and target fault scenarios;

[0077] S14, extracting the highest priority fault detection task from the trigger task queue, and executing a corresponding fault scenario simulation in a virtual environment to obtain simulated operation data of the fault scenario;

[0078] S15, using a pre-trained fault classification model to perform fault pattern matching on the simulated operation data, and outputting fault type and fault location information.

[0079] In step S11, a virtual environment containing various fault scenarios is constructed using virtualization technology, system operation logs and hardware performance indicators are obtained, and expected operation data corresponding to the performance indicators in the virtual environment are predicted, including:

[0080] Based on the virtualized resources of the cloud computing platform, containerization technology is used to deploy simulation nodes, configure multi-level hardware structures and heterogeneous operating systems;

[0081] By injecting hardware and software failure scenarios and coupling them through tools and fault injection frameworks, a virtual environment containing diverse failure scenarios is obtained;

[0082] Obtain the operation logs of each node in the virtual environment through log collection tools, and use monitoring tools to collect hardware performance indicators in real time;

[0083] Extract key state features using a feature extraction algorithm based on the node operation log and the hardware performance indicators to obtain a state feature set;

[0084] According to the state feature set, the structured features are processed by XGBoost to predict the expected operating state of each node in the virtual environment and generate expected operating data.

[0085] By building a virtual environment, it is possible to perform fault simulations when operation deviations occur in subsequent nodes and obtain simulated operation data so that the fault classification model can match the fault type without occupying computer resources at the time of the fault.

[0086] The cloud computing platform can be based on mainstream cloud services such as AWS EKS, Alibaba Cloud ACK, Azure Kubernetes Service, or private cloud platforms such as OpenStack; the virtualized resources can be virtual machines, networks, storage volumes, etc. The application of the containerization technology can be categorized as container orchestration, using Kubernetes K8s to manage containerized simulation nodes, defining resources such as Pods, Deployments, Services, or lightweight containers through declarative YAML files, and deploying container images such as Ubuntu, CentOS, and Alpine for simulation nodes using the Docker or containerd runtime.

[0087] The multi-layer hardware limits container resources, such as CPU or memory, through Kubernetes' Resource Quotas and Limit Ranges to simulate different hardware specifications; the heterogeneous operating system can be nested in the container using QEMU or Firecracker micro-virtual machines to simulate specific hardware architectures, such as the coexistence of ARM nodes and x86 nodes.

[0088] The tools and fault injection framework are used to inject hardware and software fault scenarios and couple them. For software faults, Chaos Mesh or Litmus can be used to inject Pod-level faults in the K8s cluster, such as process termination and network delay. For hardware faults, CPU or memory can be limited through cgroups, or disk IO errors can be simulated using the fault-injection kernel module. Fault coupling can be achieved by writing scripts such as Python or Ansible to coordinate multiple fault injection tools to achieve "cascading failures". For example, while triggering network packet loss, the CPU quota of a container can be limited.

[0089] Obtaining the operation logs of each node in the virtual environment through a log collection tool and using a monitoring tool to collect hardware performance indicators in real time can be achieved through the following steps: using log aggregation tools such as Fluentd and Logstash to collect operating system logs, application logs, such as Nginx error logs, and middleware logs, such as MySQL slow query logs; deploying a monitoring agent such as Prometheus Node Exporter to collect CPU usage, memory usage, disk I / O, network throughput and other indicators in real time.

[0090] The key state features are extracted by the feature extraction algorithm to obtain the state feature set, which can be obtained by the following steps:

[0091] Clean the collected node operation logs to remove meaningless characters, blank lines, and duplicate log records; for example, use regular expressions to match and filter out garbled characters and irrelevant symbols in the logs, and convert the logs into a standard text format to facilitate subsequent processing; fill in missing values and smooth the hardware performance indicator data; for example, for occasionally missing CPU usage data points, the average value of the adjacent data points can be used to fill in; for fluctuating memory usage data, use the moving average method for smoothing to reduce the impact of data noise.

[0092] For log data, text feature extraction algorithms can be used; for example, using the Bag of Words model to convert log text into word frequency vectors, by counting the various keywords in the log, such as "error",

[0093] The frequency of occurrence of "warning", "start", "stop", etc. forms a high-dimensional word frequency vector space, where each dimension corresponds to the number of occurrences of a keyword.

[0094] Statistical feature extraction methods can be used for hardware performance indicator data. For example, statistical features such as the mean, variance, maximum, and minimum values of CPU usage can be calculated. For memory usage, features such as the average and growth rate of the usage can be extracted. For network performance indicators, features such as bandwidth utilization, packet loss rate, and latency can be extracted.

[0095] Features extracted from log data and hardware performance metrics are fused to form a comprehensive state feature set. For example, the word frequency vectors of log data are combined with the statistical features of hardware performance metrics to form a multidimensional feature vector that includes both textual and numerical features. Dimensionality reduction techniques such as principal component analysis (PCA) can be used to reduce the dimensionality of the fused features to remove redundant features and reduce computational complexity.

[0096] The generated expected operating data involves building multiple trees using XGBoost to learn input features, predicting the state feature set of each node in the current virtual environment, and obtaining the expected operating state of each node. Based on the expected operating state, the corresponding expected operating data is generated. For example, for expected operating data in a normal state, system operation logs and hardware performance indicator data that conform to the historical data distribution during normal operation can be generated based on the data distribution. For expected operating data in a fault state, corresponding data patterns under historical fault scenarios can be referenced for generation.

[0097] In step S12, the real-time operating data of the system is obtained, and the scenario weighting factor, the time decay coefficient, and the expected operating data are combined to calculate and obtain a comprehensive deviation value, including:

[0098] Based on the deployed containerized nodes, log collection tools are used to extract operation logs containing timestamps to obtain real-time operation data of the system.

[0099] Based on the impact of the fault scenario on the business, the initial weight is assigned and the weight is dynamically updated based on the real-time business load or external events to obtain the scenario weighting factor;

[0100] Obtaining the time interval from the occurrence of the fault to the current calculation, and calculating the time decay coefficient based on the time interval and a preset decay rate;

[0101] A comprehensive deviation value is calculated based on the real-time operation data, the expected operation data, the scenario weighting factor, and the time attenuation coefficient.

[0102] By calculating the comprehensive deviation value of the indicator data of each node, a fault detection task is generated for the node that exceeds the preset deviation threshold. After simulation in a virtual environment, simulated operation data is obtained. After matching the fault type according to the fault classification model, the node where the fault occurs can be quickly obtained and the fault location information can be obtained, thereby improving the accuracy of fault detection.

[0103] Among them, the extraction of operation logs containing timestamps based on the deployed containerized nodes in combination with the log collection tool to obtain the real-time operation data of the system can be achieved through the following steps: using Filebeat as the log collection tool, configuring the Filebeat configuration file inside the container, and specifying the log file path and log output format to be monitored; the log collection tool reads the log files generated by each application or system service in real time during the operation of the container; the log collection tool transmits the collected operation logs with timestamps to a centralized log processing system or data storage platform, performs preliminary formatting and organization on the logs to ensure the consistency and integrity of the data; and obtains the system operation data in real time from the centralized log processing system or data storage platform.

[0104] For example, each record in the access log of the web server contains the timestamp of the client request, with a format similar to "[2024-10-01 10:00:00]Received request from client IP 192.168.1.100"; the application log may also record log entries with timestamps such as "[2024-10-01 10:05:30]Application started"; Filebeat transmits the collected logs to the Logstash data processing pipeline tool through the network, parses, filters and converts the logs in Logstash, extracts key information, and stores it uniformly in the Elasticsearch distributed search engine and analysis engine; by querying the latest log data in Elasticsearch, you can obtain performance indicators such as real-time CPU usage, memory usage, network request response time, etc. of each containerized node, as well as status information such as whether the application has any exceptions or errors.

[0105] The method of allocating initial weights based on the impact of the fault scenario on the business and dynamically updating the weights based on real-time business load or external events to obtain the scenario weighting factors is as follows: determining the weight definition, and the business impact level is based on the impact of the fault scenario on the business, such as a weight of 0.9 for core service downtime and a weight of 0.3 for edge service delay to allocate initial weights; dynamically adjusting the rules based on real-time business load, such as increasing the weight during peak hours; or external events, such as increasing the weight of related scenarios when a security vulnerability breaks out, to dynamically update the weights.

[0106] It's important to note that the decay rate λ needs to be adjusted based on the specific fault. For high-frequency anomalies, such as multiple disk I / O errors within a short period of time, the decay rate is increased to quickly reduce the impact of older anomalies. For low-frequency critical events, such as database master node downtime, the decay rate is slowed to prolong the impact.

[0107] For example, the data of the last 5 minutes is retained, and the weight of data outside the time window is reset to zero. If a node experiences CPU overload three times in a row within 10 seconds, λ is adjusted from 0.1 to 0.3 to accelerate the decay of old data.

[0108] The time interval from the occurrence of the acquisition fault to the current calculation can be obtained by a system timer.

[0109] It is worth noting that the calculation formula of the time attenuation coefficient is:

[0110] β=e -λt

[0111] Where β is the time attenuation coefficient, λ is the attenuation rate, and the unit is S-1 , t is the time interval from the occurrence of the event to the current calculation, in S;

[0112] The calculation formula of the comprehensive deviation value is:

[0113]

[0114] Among them, Δ is the comprehensive deviation value; w is the scenario weighting factor; β is the time attenuation coefficient; Now is the real-time operating value; Pre is the expected operating value; and both Now and Pre are values between (0,1) and expressed as percentages.

[0115] For example, if the real-time CPU utilization is 98% and the expected value is 70%, the weight is 0.9, the decay coefficient is 0.8, and the CPU deviation is calculated to be 0.288. If the real-time disk IOPS is 50 and the expected value is 200, the weight is 0.5, the decay coefficient is 0.6, and the disk deviation is calculated to be 0.225, for a total deviation of 0.513. Based on the comprehensive deviation results, the weights and decay coefficients can be adjusted inversely. For example, if a certain type of failure occurs frequently, its weight can be increased. Reinforcement learning, such as Q-Learning, can be used to automatically optimize the adjustment strategy.

[0116] In step S13, a priority score is calculated based on the magnitude by which the comprehensive deviation value exceeds the preset deviation threshold and the impact level of the fault scenario on business continuity, and a trigger task queue sorted by priority score is generated based on the priority score; the trigger task queue is a collection of fault detection tasks including trigger conditions, priority scores, and target fault scenarios, including:

[0117] Calculating the deviation amplitude based on the comprehensive deviation value and the preset deviation threshold;

[0118] Classify the impact of different failure scenarios on business continuity and obtain weight scores corresponding to different levels;

[0119] Calculate the priority scores of different fault scenarios based on the deviation magnitude and the weight score;

[0120] A fault detection task including a trigger condition, a priority score, and a fault scenario is generated, and all the fault detection tasks are sorted according to the priority score to obtain a trigger task queue.

[0121] By calculating the priority score, we can ensure that system resources are concentrated on handling the most urgent and most impactful faults, and accordingly generate corresponding fault detection tasks containing trigger conditions, priority scores and target fault scenarios, so as to obtain simulated operation data for fault matching.

[0122] The magnitude by which the comprehensive deviation value exceeds the preset deviation threshold can be calculated by dividing the difference between the comprehensive deviation value and the preset deviation threshold by the preset deviation threshold and converting it into a percentage. For example, if the preset deviation threshold is set to 0.5 and the actual deviation value is 0.8, the calculated deviation magnitude is 60%. Nodes corresponding to indicator data with comprehensive deviation values exceeding the preset deviation threshold are marked as potential fault nodes, and a fault detection task is generated and entered into the trigger task queue.

[0123] It should be noted that the impact of failure scenarios on business continuity can be divided into the following levels: Impact Level 1 is the unavailability of core services, such as payment and database, with a corresponding weight score of 5; Impact Level 2 is the degradation of secondary services, such as logging services, with a corresponding weight score of 3; Impact Level 3 is the delay of non-critical services, such as monitoring alarms, with a corresponding weight score of 1.

[0124] It is worth noting that the priority score PS can be calculated using the following formula:

[0125] PS=α×DS+β×BIL

[0126] Among them, PS is the priority score, α and β are weight coefficients. The default α is 0.6 and β is 0.4, which can be dynamically adjusted according to the business scenario. DS is the deviation amplitude, and BIL is the weight score corresponding to the business continuity impact level.

[0127] The priority sorting can use a maximum heap data structure, sorting from high to low by priority score. When a new task node is inserted at the end of the heap, the heap structure is adjusted upward through node value comparison and exchange operations, so that the top node of the heap always maintains the fault detection task corresponding to the maximum priority score. The triggered task queue is a collection of fault detection tasks including trigger conditions, priority scores and target fault scenarios.

[0128] In step S14, the fault detection task with the highest priority is extracted from the trigger task queue, and the corresponding fault scenario simulation is performed in the virtual environment to obtain the simulated operation data of the fault scenario, including:

[0129] Obtain the fault detection task with the highest priority score from the trigger task queue to obtain the fault scenario to be simulated;

[0130] According to the fault scenario to be simulated, corresponding operating parameters are configured and loaded in the virtual environment to obtain a virtual environment of the fault scenario to be simulated;

[0131] Simulating the fault scenario to be simulated in the virtual environment, and using a monitoring tool to collect operation logs and hardware performance indicators of the containerized node in real time to obtain simulated operation data;

[0132] After the simulation is completed, the virtual environment is reset to allow for the next fault simulation.

[0133] By generating corresponding fault detection tasks and performing fault simulation for each node that exceeds the preset deviation threshold according to the priority score and obtaining simulated operation data, the fault classification model can be used for matching and the type of fault can be accurately detected.

[0134] The fault detection task with the highest priority score can be directly extracted to trigger the task at the top of the task queue to obtain the fault scenario to be simulated.

[0135] According to the fault scenario to be simulated, the corresponding operating parameters are configured and loaded in the virtual environment, and the virtual environment for the fault scenario to be simulated needs to generate a corresponding simulation configuration according to the fault scenario determined in the fault detection task, and configure and load the corresponding operating parameters in the virtual environment to simulate the system state under the fault scenario. For example, for a database server disk I / O failure, the simulation configuration may include reducing the I / O performance of the virtual disk, setting a higher read and write delay, etc. The virtualization service provided by the virtualization tool to the cloud platform will create or adjust the virtual environment based on the generated simulation configuration. This involves setting the hardware parameters of the virtual machine, such as disk I / O rate, network bandwidth, etc. and software configuration, such as operating system settings, application parameters, etc., to ensure that the virtual environment can accurately simulate the expected fault scenario.

[0136] The monitoring tool is used to collect the operation logs and hardware performance indicators of the containerized nodes in real time to obtain simulated operation data. The deployed Filebeat log collection tool and hardware performance monitoring tools such as Prometheus Node Exporter and Grafana Agent can be used to collect the operation logs of the containerized nodes and hardware performance indicators such as CPU usage, memory usage, disk I / O, and network traffic to obtain simulated operation data.

[0137] The virtual environment may be reset by calling a snapshot management module of the virtualization environment, such as VMware's Snapshots API, to roll back the system state to the base image before the simulation and securely erase the remaining temporary data, such as coredump files and cache logs, for the next simulation.

[0138] In step S15, the pre-trained fault classification model is used to perform fault pattern matching on the simulated operation data, and the fault type and fault location information are output, including:

[0139] Extracting the operation logs of the containerized nodes in the simulation operation data using a log collection tool based on the simulation operation data, and performing structured processing using a log parsing tool to obtain structured simulation operation data;

[0140] Using a pre-trained fault classification model to match the structured simulation operation data to obtain a fault pattern matching result;

[0141] According to the fault pattern matching results, a rule engine is used to classify the matched fault patterns, and the corresponding fault types are output through a preset fault type mapping table;

[0142] The node identifier of the containerized node in the simulation operation data is extracted to obtain fault location information including the node name and IP address.

[0143] By accurately matching the simulated operation data with the fault classification model, the accurate fault type is obtained, and the fault location information is accurately extracted based on the source node of the simulated operation data, thereby improving the accuracy of fault detection.

[0144] Among them, according to the simulated operation data, a log collection tool is used to extract the operation log of the containerized node, and a log parsing tool is used to perform structured processing. The structured simulated operation data can be obtained by the following steps: using the Filebeat log collection tool to extract the operation log that has been collected and saved as simulated operation data, and then using log parsing tools such as Logstash and Fluentd to extract key information from the log through regular expressions, grok parsing patterns, etc., and convert it into a structured format, such as JSON, CSV, etc.

[0145] For example, configure the grok plugin in Logstash and define parsing rules to match fields in the log, such as the timestamp, log level, and message content. For the log entry "[2024-10-01 10:00:00] ERROR: Paymentprocessing failed for order ID 12345 due to connection timeout", grok parsing can convert it into structured data in JSON format:

[0146]

[0147] The method of matching the structured simulated operation data with the pre-trained fault classification model to obtain the fault pattern matching result is to extract features from the structured simulated operation data. The model identifies key features in the data based on the existing fault data, such as error codes and warning information in the log, abnormal values in hardware performance indicators, etc.; and compares and matches the extracted features with corresponding features in the real-time operation data to obtain the fault pattern matching result.

[0148] According to the fault pattern matching result, a rule engine is used to classify the matched fault pattern, and the corresponding fault type is output through a preset fault type mapping table. By selecting and configuring a suitable rule engine, such as Drools, CLIPS, etc., rules for classifying faults based on the characteristics and business logic of the fault pattern can be defined in the rule engine; the fault pattern matching result is input into the rule engine, and the rule engine evaluates the input fault pattern matching result according to predefined rules to check whether the characteristics in the matching result meet the conditions specified in the rule. For example, a rule may be defined as "If the fault pattern includes a database connection timeout and the affected component is the order service, then the fault is classified as an order service database connection failure."

[0149] The extraction of the node identifier of the containerized node in the simulation operation data to obtain the fault location information including the node name and IP address requires extracting the unique identifier, node name and IP address of the containerized node to form complete fault location information. For example, the output result of the fault location information combined with the fault type can be expressed as: an order service database connection failure occurred at node n1 with IP address 192.168.1.101.

[0150] It should be noted that the fault type mapping table maps the classification results in the rule engine to specific fault type descriptions. It can be a simple dictionary or database table, where the key is the classification identifier of the rule engine and the value is the corresponding fault type name.

[0151] For example, the mapping table may contain the following mapping relationships: classification identifier 1 corresponds to network connection failure; classification identifier 2 corresponds to hardware performance failure; classification identifier 3 corresponds to application failure; when the rule engine outputs the classification result, the result is used as the key to query the fault type mapping table to obtain the corresponding fault type name.

[0152] It is worth noting that the pre-trained fault classification model can be obtained by training through the following steps: obtaining historical system fault data, diversified fault data coupled with a virtual environment, and open source fault data as pre-training data for the model; labeling different fault data in the pre-training data with corresponding fault types and location information; based on the pre-training data, combined with the fault type and the location information, using a dual-channel deep learning architecture to train the model; calculating the accuracy, recall rate, and F1 score indicators as evaluation results; adopting a regularization method to adjust the model according to the evaluation results to obtain a trained fault classification model.

[0153] Among them, the acquisition of the pre-training data can be achieved by extracting information from the system historical fault database and various open source fault information, and inputting them into the virtual environment for coupling, so as to obtain diversified fault data including system historical fault data, open source fault data and virtual environment coupled as model pre-training data, and labeling different fault data with corresponding fault types and location information according to the acquired fault information.

[0154] It should be noted that the dual-channel deep learning architecture is divided into two channels. The first channel mainly processes the structured features of the pre-training data, using the fully connected layer Dense Layer as the basis to perform weight calculation and nonlinear transformation on the structured features; the second channel mainly processes text data such as system logs, and can adopt a one-dimensional convolutional neural network architecture.

[0155] It is worth noting that the evaluation indicators of the model can be divided into accuracy, recall, and F1-Score. The corresponding calculation formulas are as follows:

[0156] The calculation formula for accuracy is:

[0157] Accuracy=(TP+TN) / (TP+TN+FP+FN)

[0158] The calculation formula for recall is:

[0159] Recall = TP / (TP+FN)

[0160] The calculation formula of F1 score F1-Score is:

[0161]

[0162] Among them, TP is a true positive example, TN is a true negative example, FP is a false positive example, FN is a false negative example, F1-Score is the harmonic mean of precision and recall, and Precision = TP / (TP+FP).

[0163] The regularization method is adopted to prevent overfitting, and L1 or L2 regularization methods can be used. L1 regularization adds an absolute value weight term to the loss function, making some weights of the model become 0, which plays a role in feature selection; L2 regularization adds a square weight term to limit the size of the weight, preventing the model from being too complex.

[0164] In summary, the present invention integrates cloud computing and virtualization technology to construct a virtual environment containing multiple fault scenarios, combines containerized node simulation with XGBoost to predict and analyze the expected operating value of the indicator data of each node of the computer, and then calculates the comprehensive deviation value of the indicator data of each node in combination with the real-time operating value. For the nodes corresponding to the indicator data whose comprehensive deviation value exceeds the preset deviation threshold, a fault detection task is generated and simulated in the virtual environment. According to the simulated operating data, the fault type is accurately matched through the fault classification model, thereby improving the accuracy of electronic computer fault detection.

[0165] Reference Figure 2 The second embodiment of the present invention provides a computer fault detection system based on cloud computing, comprising:

[0166] The virtual environment construction module is used to build a virtual environment that includes various fault scenarios through virtualization technology, obtain system operation logs and hardware performance indicators, and predict the expected operation data corresponding to the performance indicators in the virtual environment;

[0167] An operation deviation calculation module is used to obtain real-time operation data of the system, and calculate the comprehensive deviation value by combining the scenario weighting factor, the time decay coefficient and the expected operation data;

[0168] a trigger task generation module, configured to calculate a priority score based on the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the level of impact of the fault scenario on business continuity, and generate a trigger task queue sorted by priority score based on the priority score; the trigger task queue includes a set of fault detection tasks for trigger conditions, priority scores, and target fault scenarios;

[0169] A fault scenario simulation module is used to extract the highest priority fault detection task from the trigger task queue, and execute the corresponding fault scenario simulation in a virtual environment to obtain simulated operation data of the fault scenario;

[0170] The matching result output module is used to use a pre-trained fault classification model to perform fault pattern matching on the simulated operation data and output fault type and fault location information.

[0171] It should be noted that the cloud computing-based electronic computer fault detection system provided in an embodiment of the present invention is used to execute all the process steps of the cloud computing-based electronic computer fault detection method in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.

[0172] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as an XGBoost prediction analysis program. When the processor executes the computer program, the steps in the above-mentioned embodiments of the electronic computer fault detection method based on cloud computing are implemented, for example Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as running the deviation calculation module.

[0173] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0174] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.

[0175] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire electronic device using various interfaces and lines.

[0176] The memory can be used to store the computer programs and / or modules, and the processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0177] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of each of the above-mentioned method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0178] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0179] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A computer fault detection method based on cloud computing, characterized in that: include: Use virtualization technology to build a virtual environment that includes various fault scenarios, obtain system operation logs and hardware performance indicators, and predict the expected operation data corresponding to the performance indicators in the virtual environment; Acquire real-time operating data of the system, and calculate based on the scenario weighting factor, the time decay coefficient, and the expected operating data to obtain a comprehensive deviation value; Calculating a priority score based on the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the level of impact of the fault scenario on business continuity, and generating a trigger task queue sorted by priority score based on the priority score; the trigger task queue is a collection of fault detection tasks including trigger conditions, priority scores, and target fault scenarios; Extracting the highest priority fault detection task from the trigger task queue, and executing the corresponding fault scenario simulation in the virtual environment to obtain simulated operation data of the fault scenario; A pre-trained fault classification model is used to perform fault pattern matching on the simulated operation data, and fault type and fault location information are output.

2. The electronic computer fault detection method based on cloud computing according to claim 1, characterized in that: The virtual environment containing various fault scenarios is constructed by virtualization technology, system operation logs and hardware performance indicators are obtained, and the expected operation data corresponding to the performance indicators in the virtual environment are predicted, including: Based on the virtualized resources of the cloud computing platform, containerization technology is used to deploy simulation nodes, configure multi-level hardware structures and heterogeneous operating systems; By injecting hardware and software failure scenarios and coupling them through tools and fault injection frameworks, a virtual environment containing diverse failure scenarios is obtained; Obtain the operation logs of each node in the virtual environment through log collection tools, and use monitoring tools to collect hardware performance indicators in real time; Extract key state features using a feature extraction algorithm based on the node operation log and the hardware performance indicators to obtain a state feature set; According to the state feature set, the structured features are processed by XGBoost to predict the expected operating state of each node in the virtual environment and generate expected operating data.

3. The electronic computer fault detection method based on cloud computing according to claim 1, characterized in that: The real-time operation data of the acquisition system is calculated in combination with the scenario weighting factor, the time decay coefficient and the expected operation data to obtain a comprehensive deviation value, including: Based on the deployed containerized nodes, log collection tools are used to extract operation logs containing timestamps to obtain real-time operation data of the system. Based on the impact of the fault scenario on the business, the initial weight is assigned and the weight is dynamically updated based on the real-time business load or external events to obtain the scenario weighting factor; Obtaining the time interval from the occurrence of the fault to the current calculation, and calculating the time decay coefficient based on the time interval and a preset decay rate; A comprehensive deviation value is calculated based on the real-time operation data, the expected operation data, the scenario weighting factor, and the time attenuation coefficient.

4. The electronic computer fault detection method based on cloud computing according to claim 3, characterized in that: The calculation formula of the time decay coefficient is: β=e -λt Where β is the time attenuation coefficient, λ is the attenuation rate, and the unit is S -1 , t is the time interval from the occurrence of the event to the current calculation, in S; The calculation formula of the comprehensive deviation value is: Among them, Δ is the comprehensive deviation value; w is the scene weighting factor; β is the time attenuation coefficient; Now is the real-time operation value; Pre is the expected operation value; and both Now and Pre are values between 0 and 1.

5. The electronic computer fault detection method based on cloud computing according to claim 1, characterized in that: The calculating of a priority score according to the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the impact level of the fault scenario on business continuity, and generating a trigger task queue sorted by priority score according to the priority score, includes: Calculating the deviation amplitude based on the comprehensive deviation value and the preset deviation threshold; Classify the impact of different failure scenarios on business continuity and obtain weight scores corresponding to different levels; Calculate the priority scores of different fault scenarios based on the deviation magnitude and the weight score; A fault detection task including a trigger condition, a priority score, and a fault scenario is generated, and all the fault detection tasks are sorted according to the priority score to obtain a trigger task queue.

6. The electronic computer fault detection method based on cloud computing according to claim 1, characterized in that: The extracting the highest priority fault detection task from the trigger task queue, and executing the corresponding fault scenario simulation in the virtual environment to obtain the simulated operation data of the fault scenario includes: Obtain the fault detection task with the highest priority score from the trigger task queue to obtain the fault scenario to be simulated; According to the fault scenario to be simulated, corresponding operating parameters are configured and loaded in the virtual environment to obtain a virtual environment of the fault scenario to be simulated; Simulating the fault scenario to be simulated in the virtual environment, and using a monitoring tool to collect operation logs and hardware performance indicators of the containerized node in real time to obtain simulated operation data; After the simulation is completed, the virtual environment is reset to allow for the next fault simulation.

7. The electronic computer fault detection method based on cloud computing according to claim 1, characterized in that: The method of using a pre-trained fault classification model to perform fault pattern matching on the simulated operation data and outputting fault type and fault location information includes: Extracting the operation logs of the containerized nodes in the simulation operation data using a log collection tool based on the simulation operation data, and performing structured processing using a log parsing tool to obtain structured simulation operation data; Using a pre-trained fault classification model to match the structured simulation operation data to obtain a fault pattern matching result; According to the fault pattern matching results, a rule engine is used to classify the matched fault patterns, and the corresponding fault types are output through a preset fault type mapping table; The node identifier of the containerized node in the simulation operation data is extracted to obtain fault location information including the node name and IP address.

8. The electronic computer fault detection method based on cloud computing according to claim 7, characterized in that: The training process of the pre-trained fault classification model includes: Obtain historical system fault data, diverse fault data coupled with the virtual environment, and open source fault data as pre-training data for the model; Labeling the corresponding fault types and location information for different fault data in the pre-training data; According to the pre-training data, combined with the fault type and the location information, a dual-channel deep learning architecture is used to train the model; Calculate the precision, recall and F1 score indicators as evaluation results; A regularization method is adopted to adjust the model according to the evaluation results to obtain a trained fault classification model.

9. The electronic computer fault detection method based on cloud computing according to claim 8, characterized in that: The accuracy, recall, and F1 score are calculated using the following formulas: The calculation formula of the accuracy is: Accuracy=(TP+TN) / (TP+TN+FP+TN) The calculation formula of the recall rate Recall is: Recall = TP / (TP+FN) The calculation formula of the F1 score F1-Score is: Among them, TP is a true positive example, TN is a true negative example, FP is a false positive example, FN is a false negative example, F1-Score is the harmonic mean of precision and recall, and Precision = TP / (TP+FP).

10. An electronic computer fault detection system based on cloud computing, characterized in that: include: The virtual environment construction module is used to build a virtual environment that includes various fault scenarios through virtualization technology, obtain system operation logs and hardware performance indicators, and predict the expected operation data corresponding to the performance indicators in the virtual environment; An operation deviation calculation module is used to obtain real-time operation data of the system, and calculate the comprehensive deviation value by combining the scenario weighting factor, the time decay coefficient and the expected operation data; A trigger task generation module is used to calculate a priority score based on the magnitude by which the comprehensive deviation value exceeds a preset deviation threshold and the impact level of the fault scenario on business continuity, and generate a trigger task queue sorted by priority score based on the priority score; The trigger task queue is a collection of fault detection tasks including trigger conditions, priority scores and target fault scenarios; A fault scenario simulation module is used to extract the highest priority fault detection task from the trigger task queue, and execute the corresponding fault scenario simulation in a virtual environment to obtain simulated operation data of the fault scenario; The matching result output module is used to use a pre-trained fault classification model to perform fault pattern matching on the simulated operation data and output fault type and fault location information.

Citation Information

Patent Citations

  • Abnormality detection model training method and device and fault scene positioning method and device

    CN116304909A

  • Fault analysis system for big data cloud computing

    CN116594801A

  • System kernel reliability evaluation method and system based on fault sequence generation

    CN119088715A

  • Fault processing method and device of cloud computing platform, electronic equipment and storage medium

    CN119557134A

  • KR20240163217A