Server fault diagnosis method and device

By collecting the actual operation data of the server and historical fault data, and using the fault path generation model to dynamically generate diagnostic paths, the problem of process solidification in server fault diagnosis is solved, efficient and accurate fault location and diagnosis is achieved, and the risk of production interruption is reduced.

CN120354178BActive Publication Date: 2025-08-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510846497.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-29
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The existing server fault diagnosis method has a single process, and cannot be dynamically adjusted according to the real-time operating status and fault conditions, resulting in low diagnostic efficiency and poor accuracy. Once the fault cannot be solved quickly, it is easy to cause production interruptions, causing economic losses to the enterprise.

Method used

By collecting the actual operating data of the server and historical fault data, the fault path generation model is used to generate diagnostic paths dynamically, including detecting server faults, generating diagnostic rules, labels and types, and realizing dynamic adjustment and accurate diagnosis.

Benefits of technology

Dynamically generate accurate and efficient diagnostic paths, shorten diagnosis time, reduce energy consumption and maintenance costs, improve fault positioning accuracy, and reduce the risk of production interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354178B_ABST
    Figure CN120354178B_ABST
Patent Text Reader

Abstract

The present invention discloses a server fault diagnosis method and device, which relate to the technical field of electronic digital data processing. The method comprises collecting actual operation data of a target server, obtaining configuration information and historical fault data, and detecting whether a fault occurs in the target server; when a fault occurs, inputting the configuration information, actual operation data and historical fault data into a pre-built fault path generation model to output a corresponding diagnostic path; and determining diagnostic rules, diagnostic labels and fault types when diagnosing the target server according to the diagnostic path to generate a current fault diagnosis result. The method solves the technical problems of a rigid and single process, inability to dynamically adjust, lack of flexibility and pertinence, low diagnostic efficiency and poor diagnostic accuracy, thereby achieving the technical effects of shortening diagnostic time, reducing server energy consumption and maintenance costs, improving fault location accuracy, reducing fault resolution time and realizing automated fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a method and device for diagnosing server faults. Background Art

[0002] In related technologies, server faults can be diagnosed according to a specific diagnostic process. For example, the server is initially configured according to factory specifications; the server's basic hardware is then tested; the server's functionality is tested; and finally, the server's core hardware is stress-tested, thereby fully ensuring the server's stable and efficient operation. Furthermore, during the diagnosis process, manual intervention can be used to ensure that the server product meets factory specifications.

[0003] However, in related technologies, the process is rigid and single, and cannot be dynamically adjusted according to the real-time operating status and fault conditions of the server. It lacks flexibility and pertinence. Especially when the server operation is abnormal, the diagnosis direction and depth cannot be changed in time, resulting in low diagnostic efficiency, reduced diagnostic accuracy, and serious waste of resources. In addition, once the fault cannot be resolved quickly, it is very easy to cause production interruption, causing huge economic losses to the enterprise, and it is in urgent need of improvement. Summary of the Invention

[0004] The present invention provides a server fault diagnosis method and device to at least solve the problems in the related art, such as the rigid and single process, the inability to dynamically adjust according to the real-time operating status and fault conditions of the server, the lack of flexibility and pertinence, the low diagnostic efficiency and poor diagnostic accuracy. In addition, if the fault cannot be quickly resolved, it is very likely to cause production interruption, resulting in huge economic losses to the enterprise.

[0005] The present invention provides a server fault diagnosis method, comprising the following steps: collecting actual operation data of a target server, and obtaining configuration information and at least one of historical fault data of the target server, so as to use the actual operation data and at least one of the configuration information and the historical fault data to detect whether a fault occurs on the target server; when a fault is detected on the target server, inputting the configuration information, the actual operation data and the historical fault data into a pre-built fault path generation model, so as to output a diagnostic path when a fault occurs on the target server; determining a diagnostic rule, a diagnostic label and a fault type when diagnosing the target server according to the diagnostic path, so as to generate a current fault diagnosis result of the target server based on the diagnostic rule, the diagnostic label and the fault type.

[0006] The present invention also provides a server fault diagnosis device, comprising: a detection module, used to collect actual operation data of the target server, and obtain the configuration information and at least one of the historical fault data of the target server, so as to use the actual operation data and at least one of the configuration information and the historical fault data to detect whether the target server has a fault; an output module, used to input the configuration information, the actual operation data and the historical fault data into a pre-built fault path generation model when a fault is detected in the target server, so as to output a diagnostic path when a fault occurs in the target server; a diagnostic module, used to determine the diagnostic rules, diagnostic labels and fault types when diagnosing the target server according to the diagnostic path, so as to generate the current fault diagnosis result of the target server based on the diagnostic rules, the diagnostic labels and the fault type.

[0007] The present invention also provides a server, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server fault diagnosis methods when executing the computer program.

[0008] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server fault diagnosis methods are implemented.

[0009] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned server fault diagnosis methods when executed by a processor.

[0010] Through the present invention, the actual operation data, configuration information and historical fault data of the target server can be collected, and then at least one of the actual operation data, configuration information and historical fault data can be used to detect whether the target server has a fault. When a fault is detected on the target server, the configuration information, actual operation data and historical fault data are input into a pre-built fault path generation model, and then the diagnostic path when the target server has a fault is output, thereby determining the diagnostic rules, diagnostic labels and fault types when diagnosing the target server, and then generating the current fault diagnosis results of the target server. Therefore, it can solve the problems of rigid and single processes, inability to dynamically adjust according to the real-time operation status and fault conditions of the server, lack of flexibility and pertinence, low diagnostic efficiency and poor diagnostic accuracy. In addition, once the fault cannot be quickly resolved, it is very easy to cause production interruption, causing huge economic losses to the enterprise. The technical problem is achieved by dynamically generating accurate and efficient diagnostic paths, shortening diagnostic time, reducing server energy consumption and maintenance costs, improving fault location accuracy, reducing fault resolution time, and realizing the technical effect of automated fault diagnosis for different types of servers. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 A flowchart of a server fault diagnosis method provided according to an embodiment of the present invention;

[0013] Figure 2 A flowchart of constructing a fault path generation model according to an embodiment of the present invention;

[0014] Figure 3 A schematic diagram of a server product configuration provided according to an embodiment of the present invention;

[0015] Figure 4 A schematic diagram of a server product production failure repair according to an embodiment of the present invention;

[0016] Figure 5 A flowchart of the working principle of a server fault diagnosis method provided by one embodiment of the present invention;

[0017] Figure 6 The present invention is a block diagram of a server fault diagnosis device according to an embodiment of the present invention.

[0018] Reference numerals:

[0019] Among them, 10 is a server fault diagnosis device; 100 is a detection module, 200 is an output module, and 300 is a diagnosis module. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0022] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0023] An embodiment of the present invention provides a server fault diagnosis method, and the method is described in detail in conjunction with the execution flow of the server fault diagnosis method.

[0024] Specifically, Figure 1 The present invention provides a flowchart of a method for diagnosing server faults according to an embodiment of the present invention.

[0025] like Figure 1 As shown, the fault diagnosis method of the server includes the following steps:

[0026] In step S101, actual operation data of the target server is collected, and at least one of the configuration information and historical fault data of the target server is obtained, so as to detect whether a fault occurs in the target server using the actual operation data and at least one of the configuration information and historical fault data.

[0027] It should be understood that in embodiments of the present invention, configuration information may include, but is not limited to, target server model information, hardware information (such as firmware version, hardware model, network mode, RAID (Redundant Array of Independent Disks) status, etc., which is not specifically limited in the present invention), CPU (Central Processing Unit) data (such as CPU threads and frequency, etc., which are not specifically limited in the present invention), hard disk data (such as hard disk capacity and transmission rate, etc., which are not specifically limited in the present invention), memory data (such as memory capacity and transmission rate, etc., which are not specifically limited in the present invention), etc., which are not specifically limited in the present invention; actual operation data may include, but is not limited to, target server log files, ambient temperature, operating voltage, performance indicators, etc., which are not specifically limited in the present invention; historical fault data may include, but is not limited to, maintenance data (such as repair measures, etc., which are not specifically limited in the present invention) and fault data (such as fault time, error code, fault type, occurrence frequency, etc., which are not specifically limited in the present invention) within a certain period of time, which is not specifically limited in the present invention. The certain period of time can be set by those skilled in the art based on actual conditions and is not specifically limited in the present invention.

[0028] Furthermore, embodiments of the present invention can collect real-time operational data from target servers through a combination of IPMI (Intelligent Platform Management Interface) and BMC (Baseboard Management Controller). The sampling frequency can be increased from once every 10 minutes to once every 1 minute. The specific setting can be determined by those skilled in the art based on practical circumstances and is not specifically limited by the present invention. Furthermore, embodiments of the present invention can record hardware configuration information at intervals (e.g., every 10 minutes) and transmit this hardware configuration information to a database to create a hardware status snapshot.

[0029] In some embodiments, embodiments of the present invention may collect configuration information, actual operation data, and historical fault data of a target server, and then detect whether a fault occurs in the target server based on the configuration information, actual operation data, and historical fault data.

[0030] In some embodiments, the embodiments of the present invention can collect actual operation data of the target server and obtain historical fault data of the target server, and then use the actual operation data and historical fault data to detect whether a fault occurs in the target server.

[0031] Exemplarily, an embodiment of the present invention can use a combination of IPMI and BMC to collect the actual operation data of the target server in real time, with a sampling frequency of 1 time / 1 minute, and record hardware configuration information every 10 minutes, transmit it to the database, and take a hardware status snapshot, and then combine it with historical fault data to detect whether the target server has a fault.

[0032] The embodiments of the present invention deeply integrate multi-source data such as server configuration information, actual operation data and historical fault data, breaking through the limitations of traditional single reliance on preset strategies. Instead of adopting a "one-size-fits-all" fixed diagnostic process, a diagnostic path is tailored based on the unique data combination of each server.

[0033] Optionally, in one embodiment of the present invention, before inputting the configuration information, actual operation data and historical fault data into a pre-built fault path generation model, it also includes: based on the configuration information, actual operation data and historical fault data, respectively calculating the hardware matching degree between the configuration information and the corresponding diagnostic rules, the operation matching degree between the actual operation data and the corresponding diagnostic rules, and the fault matching degree between the historical fault data and the corresponding diagnostic rules; based on the hardware matching degree, the operation matching degree and the fault matching degree, respectively determining the weight coefficients corresponding to the hardware matching degree, the operation matching degree and the fault matching degree; based on the weight coefficient, the hardware matching degree, the operation matching degree and the fault matching degree, determining the matching degree of the corresponding diagnostic rule; and constructing a fault path generation model based on different diagnostic rules and the matching degrees of different diagnostic rules.

[0034] It is understood that the embodiment of the present invention can uniformly manage different diagnostic rules and store them in a rule base in json format. Each diagnostic rule is targeted at different machine model information, hardware configuration information and application scenarios. Contains multiple diagnostic rules ,in, ,and is an integer. Each diagnostic rule includes trigger conditions, actions, priority weights, etc., which are not specifically limited in the present invention.

[0035] In some embodiments, embodiments of the present invention may calculate the hardware matching degree between configuration information and corresponding diagnostic rules according to different diagnostic rules.

[0036] In some embodiments, embodiments of the present invention may calculate the operation matching degree between actual operation data and corresponding diagnosis rules according to different diagnosis rules.

[0037] In some embodiments, embodiments of the present invention may calculate the fault matching degree between historical fault data and corresponding diagnostic rules according to different diagnostic rules.

[0038] Furthermore, the embodiment of the present invention can calculate the matching degree (0-1) between different diagnostic rules and configuration information, actual operation data, and historical fault data based on the hardware matching degree, operation matching degree, and fault matching degree. The calculation formula can be, but is not limited to, the following:

[0039] ,

[0040] in, , , is the weight coefficient (wherein, in the embodiment of the present invention, , , Set to , , , the present invention does not make specific restrictions); The hardware matching degree between the hardware configuration required by the diagnostic rule and the current hardware configuration (e.g., if an NVMe (Non-Volatile Memory Express) hard drive exists, the matching degree is 1; otherwise, it is 0, which is not specifically limited in the present invention); The matching degree between the actual operating data and the diagnostic rule triggering condition (e.g., if the temperature exceeds the threshold, the matching degree is 1, which is not specifically limited in the present invention); is the normalized value of the correlation between historical fault data and diagnostic rules (e.g., if the frequency of a fault associated with a diagnostic rule is 30%, the matching degree is 0.3, which is not specifically limited in the present invention); The current configuration information of the server; The actual operation data; This is historical fault data.

[0041] Furthermore, the embodiment of the present invention can construct a fault path generation model based on different diagnostic rules and the matching degrees of different diagnostic rules. The expression thereof can be, but is not limited to,:

[0042] ,

[0043] in, Indicates the The priority weight of the rule, For rule index, the diagnosis rule with higher weight occupies a more important position in path generation and affects the priority sorting of the diagnosis process; For the candidate paths, the objective function is calculated by The weighted matching degree of the path with the highest matching degree is selected as the final diagnosis path to achieve dynamic optimization. It is the index identifier of the candidate path, used to distinguish different diagnostic path solutions.

[0044] For example, the present invention can be combined with Figure 2 As shown in the figure, a fault path generation model is constructed, the main contents of which are:

[0045] Step S201: Obtain the configuration information, actual operation data and historical fault data of the server.

[0046] Step S202: Acquire different diagnostic rules.

[0047] Step S203: Calculate the matching degree between different diagnosis rules and configuration information, actual operation data and historical fault data.

[0048] Step S204: sorting different diagnostic rules based on different matching degrees.

[0049] In the embodiment of the present invention, based on a certain threshold, diagnostic rules with a low matching degree can be filtered out, and the matching degree can be sorted in descending order. The certain threshold can be set by those skilled in the art according to actual conditions and is not specifically limited by the present invention.

[0050] Step S205: constructing an initial fault path.

[0051] The embodiment of the present invention can construct an initial fault path through greedy selection.

[0052] Step S206: Real-time monitoring and dynamic adjustment.

[0053] The embodiment of the present invention can perform real-time monitoring and dynamic adjustment based on different diagnostic rules and corresponding initial fault paths.

[0054] Step S207: Update the matching degree.

[0055] The embodiment of the present invention can recalculate the matching degree of the current diagnosis rule according to the situation where the current diagnosis rule fails to execute or the condition changes.

[0056] Step S208: Adjust the initial fault path to obtain the fault path.

[0057] The embodiment of the present invention can adjust the initial fault path by inserting new diagnostic rules or removing invalid diagnostic rules, thereby obtaining the fault path.

[0058] The embodiment of the present invention improves the diagnostic accuracy through a triple data verification mechanism of hardware matching, operation matching and fault matching, dynamically adjusts weights to ensure diagnostic precision, pre-generates fault paths, and meets the real-time requirements of industrial control scenarios.

[0059] Optionally, in one embodiment of the present invention, before inputting the configuration information, actual operation data and historical fault data into the pre-built fault path generation model, it also includes: generating trigger conditions for triggering the diagnostic rules in the pre-built fault path generation model based on the configuration information, actual operation data and historical fault data; judging whether the diagnostic rules exist based on the trigger conditions; if the diagnostic rules exist, filtering the diagnostic rules according to preset conditions to obtain diagnostic rules applicable to when a fault occurs in the target server; if the diagnostic rules do not exist, establishing corresponding diagnostic rules based on the preset conditions, configuration information, actual operation data and historical fault data to obtain the established diagnostic rules.

[0060] It is understandable that the embodiments of the present invention can establish corresponding diagnostic rules through the collected configuration information, actual operation data and historical fault data, and timely update or delete the corresponding diagnostic rules according to different trigger conditions to ensure the accuracy and effectiveness of the diagnostic rules.

[0061] In some embodiments, embodiments of the present invention can generate trigger conditions for triggering diagnostic rules in a pre-built fault path generation model based on configuration information, actual operation data, and historical fault data, and then determine whether the diagnostic rules exist based on the trigger conditions; if they exist, the diagnostic rules are filtered according to certain conditions to obtain diagnostic rules applicable to when a target server fails; if they do not exist, corresponding diagnostic rules are established based on the certain conditions, configuration information, actual operation data, and historical fault data to obtain the established diagnostic rules. The certain conditions can be set by those skilled in the art based on actual circumstances and are not specifically limited by the present invention.

[0062] In the embodiment of the present invention, the triggering conditions may include, but are not limited to, conditional judgments, processing actions, and dependency conditions between diagnostic processes, etc., and the present invention does not impose specific limitations.

[0063] Exemplarily, the embodiment of the present invention can formulate different diagnostic rules according to different server failure situations. During the diagnostic process, the corresponding diagnostic rules are adjusted based on configuration information, actual operation data and historical failure data. For example, if a trigger condition is triggered in the embodiment of the present invention, the diagnostic path can be adjusted according to the processing action of the diagnostic rule, and the corresponding diagnostic tasks can be increased or decreased. For example, in the NVMe hard drive performance test, the embodiment of the present invention detects that the hard drive read and write rate does not meet the baseline standard, stops the test and adds a hard drive comprehensive inspection task.

[0064] Among them, in an embodiment of the present invention, the diagnostic rules corresponding to the NVMe hard disk performance test can be: first obtain the diagnostic rules for different NVMe hard disk operations, then introduce the purpose of each diagnostic rule, and then set the triggering conditions for triggering different diagnostic rules, such as the executed operations and related historical failure conditions, etc. The present invention does not make specific restrictions. Finally, based on the configuration information, actual operation data and historical failure data, determine the corresponding processing actions, priorities and dependencies.

[0065] Furthermore, when adjusting diagnostic rules, embodiments of the present invention can suspend or adjust diagnostic rules that depend on abnormal components based on the dependencies and priorities between different diagnostic rules. For example, if a CPU failure is detected during a CPU stress test, after replacing the CPU, the CPU status check, CPU performance check, and stress test can be performed sequentially based on the "xx" parameter in the diagnostic rule. This is not specifically limited by the present invention.

[0066] By generating trigger conditions, the embodiments of the present invention can locate relevant diagnostic rules more quickly, avoid wasting time on irrelevant rules, and improve diagnostic efficiency. When diagnostic rules do not exist, new rules can be established based on preset conditions and data to enhance system adaptability. By screening diagnostic rules applicable to the current fault, computing resources can be optimized and diagnostic accuracy can be improved.

[0067] Optionally, in one embodiment of the present invention, before inputting the configuration information, actual operation data and historical fault data into a pre-built fault path generation model, it also includes: processing the configuration information, actual operation data and historical fault data separately to obtain configuration information, actual operation data and historical fault data that meet preset data format conditions.

[0068] It is understandable that the embodiment of the present invention can use a combination of a distributed file system and a distributed database to store configuration information, actual operation data and historical fault data. For example, the embodiment of the present invention can store structured data such as hardware information, fault time, error code, etc. (the present invention does not impose specific restrictions) in a Hive table and use the SQL-like query language in the Hive table for efficient query. The storage diagram is shown in FIG. Figure 3 As shown; for unstructured data, such as fault type, maintenance measures, log files, etc., the present invention does not make specific restrictions, and is stored in the document database MongoDB and stored in json format. The storage diagram is as shown Figure 4 shown.

[0069] In some embodiments, embodiments of the present invention can process configuration information, actual operation data, and historical fault data to obtain configuration information, actual operation data, and historical fault data that meet certain data format conditions. The certain data format conditions can be set by those skilled in the art based on actual circumstances and are not specifically limited by the present invention.

[0070] For example, an embodiment of the present invention can use a big data collection tool to collect server log files, maintenance measures, fault types and fault times in real time, and clean, transform and integrate the collected data through an ETL (Extract Transform Load) process to make it meet a certain data format, and then uniformly store it in a Hive table to facilitate subsequent query and analysis.

[0071] In addition, it should be noted that when analyzing log files, fault types and fault times, the embodiments of the present invention can use big data analysis tools to replace inefficient manual intervention, and use the xx algorithm to count the fault frequencies of different hardware information in different diagnostic rules. The possibility of fault occurrence is predicted based on hardware information, actual operation data, etc., and the analysis results are fed back to the diagnostic rules, thereby optimizing the fault path and diagnostic rules and improving the accuracy and efficiency of diagnosis.

[0072] The embodiments of the present invention improve data quality through data processing, ensure the accuracy of subsequent processing, enhance model reliability and compatibility, and save resources.

[0073] In step S102 , when a target server failure is detected, configuration information, actual operation data, and historical failure data are input into a pre-built failure path generation model to output a diagnostic path when the target server failure occurs.

[0074] It can be understood that in an embodiment of the present invention, the pre-built fault path generation model may include multiple diagnostic rules, each diagnostic rule is targeted at the model information, hardware information and actual operation data of different servers. In addition, each diagnostic rule has clear trigger conditions, condition judgments, processing actions and dependency conditions between diagnostic processes, etc. The present invention does not impose specific restrictions.

[0075] As a possible implementation method, an embodiment of the present invention can input configuration information, actual operation data and historical failure data into a pre-built failure path generation model when a failure occurs in the target server, and then output a diagnostic path when the failure occurs in the target server.

[0076] Illustratively, an embodiment of the present invention can determine the diagnostic rule A of the target server in fault state 1 based on configuration information A, actual operation data A and historical fault data A when fault 1 occurs on the target server, and then input the configuration information A, actual operation data A, historical fault data A and diagnostic rule A into a pre-built fault path generation model, and then output the diagnostic path when fault 1 occurs on the target server.

[0077] The embodiment of the present invention can determine the diagnostic rule B of the target server in the second fault state based on the configuration information B, actual operation data B and historical fault data B when the target server has the second fault, and then input the configuration information B, actual operation data B, historical fault data B and diagnostic rule B into a pre-built fault path generation model, and then output the diagnostic path when the target server has the second fault.

[0078] Optionally, in one embodiment of the present invention, configuration information, actual operation data and historical fault data are input into a pre-built fault path generation model to output a diagnostic path when a fault occurs in the target server, including: determining whether the diagnostic rules meet the preset fault diagnosis conditions; if the diagnostic rules meet the preset fault diagnosis conditions, determining the diagnostic path based on the diagnostic rules.

[0079] As one possible implementation, embodiments of the present invention can determine whether a diagnostic rule meets certain fault diagnosis conditions when inputting configuration information, actual operating data, and historical fault data into a pre-built fault path generation model. If so, a diagnostic path is determined based on the diagnostic rule. The fault diagnosis conditions can be configured by those skilled in the art based on actual circumstances and are not specifically limited by the present invention.

[0080] Illustratively, the embodiment of the present invention can determine whether the diagnosis rule meets certain fault diagnosis conditions according to the type of the server, and determine the diagnosis path if the conditions are met.

[0081] Furthermore, in the case of a newly launched server, embodiments of the present invention can select corresponding diagnostic rules based on the server's hardware information, thereby determining a diagnostic path that meets certain fault diagnosis conditions. For example, if the server is equipped with an NVMe hard drive, by selecting diagnostic rules related to the NVMe hard drive, such as firmware upgrade checks, property checks, and performance tests, and if the diagnostic rules meet certain fault diagnosis conditions, a diagnostic path is determined based on the corresponding diagnostic rules.

[0082] In the embodiment of the present invention, when the server is a mass-produced server, diagnostic rules that meet certain fault diagnosis conditions can be screened based on the server's model information, hardware information, actual operation data, and historical fault data, thereby determining a diagnostic path.

[0083] In addition, it should be noted that when determining the diagnostic path, the embodiment of the present invention follows the priority of server registration, firmware upgrade and inspection, power redundancy test, hardware performance test and inspection, hardware stress test, human-computer interaction and server factory settings.

[0084] The embodiments of the present invention ensure that only relevant and effective rules are applied through certain fault diagnosis conditions, thereby reducing misdiagnosis, improving diagnostic accuracy, avoiding unnecessary rule checks, optimizing resource utilization, shortening diagnosis time, adjusting rule application according to different fault conditions, improving flexibility, supporting complex fault handling, and generating a more comprehensive diagnostic path.

[0085] Optionally, in one embodiment of the present invention, configuration information, actual operation data and historical fault data are input into a pre-built fault path generation model to output a diagnostic path when a fault occurs in the target server, including: determining a diagnostic label when a fault occurs in the target server based on the configuration information, actual operation data and historical fault data; identifying the fault type of the target server based on the diagnostic label; and calculating the diagnostic path using the pre-built fault path generation model based on the diagnostic label and the fault type.

[0086] It is understood that embodiments of the present invention can add different diagnostic tags to a target server experiencing a fault based on different configuration information, actual operating data, and historical fault data, thereby quickly identifying the fault issue. For example, embodiments of the present invention can add hardware diagnostic tags, such as CPU, memory, hard disk, GPU (Graphics Processing Unit), power supply, etc., without specific limitations in the present invention; software diagnostic tags, such as BIOS (Basic Input / Output System) version, firmware version, driver version, OS (Operating System) type, etc., without specific limitations in the present invention; environmental diagnostic tags, such as temperature, humidity, voltage fluctuation, load status, etc., without specific limitations in the present invention; and fault severity diagnostic tags, such as minor alarm, severe fault, system crash, etc., without specific limitations in the present invention.

[0087] Furthermore, the embodiment of the present invention can redefine the diagnostic path and uniformly classify the mixed diagnostic rules into corresponding diagnostic tags. Each diagnostic tag has a single and clear function, and the internal codes are closely related, which reduces the dependencies between different diagnostic tags and reduces the coupling degree.

[0088] During the actual execution process, the embodiment of the present invention can determine the diagnostic label when a fault occurs in the target server based on configuration information, actual operation data and historical fault data, and then identify the fault type of the target server, and then calculate the diagnostic path based on the diagnostic label and fault type using a pre-built fault path generation model.

[0089] Illustratively, an embodiment of the present invention may divide the diagnostic rules into multiple hardware diagnostic tags based on the hardware information of the server and the nature of the diagnostic rules, wherein the hardware diagnostic tags may include, but are not limited to, CPU diagnostic tags, memory diagnostic tags, and hard disk diagnostic tags, etc., and the present invention does not impose specific restrictions.

[0090] The CPU diagnostic tag may be responsible for CPU firmware upgrades, parameter checks, performance tests, and stress tests, etc., which are not specifically limited in the present invention.

[0091] The memory diagnostic tag can check the capacity, in-place status, read and write speed, and CE fault screening of the memory, etc., and the present invention does not impose specific limitations.

[0092] The hard drive diagnostic tag can perform hard drive attribute checks, performance tests, stress tests, firmware upgrades, and fault diagnosis on different types of hard drives (such as NVMe hard drives, etc., which are not specifically limited in this invention). This invention does not impose specific restrictions.

[0093] Furthermore, the embodiment of the present invention takes the hard disk diagnostic tag as an example, and its diagnostic process can be: executing the upgrade operation of the NVMe hard disk according to the firmware version, upgrade diagnosis, upgrade result, etc. of the NVMe hard disk; executing the attribute check diagnostic operation of the NVMe hard disk by reading the attribute information of the NVMe hard disk, such as capacity, health status, etc.; executing the stress test diagnostic operation of the NVMe hard disk by setting stress test parameters, running test tools, etc., and then accurately instantiating and calling the corresponding diagnostic rules, and meeting the personalized diagnostic needs of servers with different hardware configurations through the combination of diagnostic tags.

[0094] It should also be noted that, to facilitate the management and access of diagnostic tags, this embodiment of the present invention establishes a diagnostic tag management system responsible for the registration, loading, and querying of diagnostic tags. Furthermore, in this embodiment of the present invention, the diagnostic tag management system utilizes a configuration file that defines the name, class name, and path of the diagnostic tag. This is not specifically limited in the present invention, and unified processing improves efficiency.

[0095] The embodiment of the present invention enables unified processing of data from different sources through diagnostic tags, thereby narrowing the scope of diagnosis, reducing unnecessary examinations, improving efficiency, and achieving more accurate path calculation.

[0096] Optionally, in one embodiment of the present invention, before determining the diagnostic rules, diagnostic labels and fault types when diagnosing the target server based on the diagnostic path, it also includes: judging whether the matching degree of the diagnostic path is less than a preset matching degree value; if the matching degree of the diagnostic path is less than the preset matching degree value, regenerating the diagnostic path using a pre-built fault path generation model until the matching degree of the regenerated diagnostic path is greater than or equal to the preset matching degree value.

[0097] During actual implementation, before determining the diagnostic rules, diagnostic labels, and fault types for the target server based on the diagnostic path, embodiments of the present invention may also determine whether the matching degree of the diagnostic path is less than a certain matching degree value. If so, the diagnostic path is regenerated using a pre-built fault path generation model until the matching degree of the regenerated diagnostic path is greater than or equal to the certain matching degree value. The certain matching degree value can be set by those skilled in the art based on actual circumstances and is not specifically limited by the present invention.

[0098] Exemplarily, in an embodiment of the present invention, when an NVMe hard drive fails, , In the case of , diagnostic rules The corresponding diagnostic path 1 matching degree is , follow diagnostic path 1 to perform firmware upgrade, property check, performance test, and stress test in sequence. During the performance test, it is detected that the read and write rate of the NVMe hard disk is lower than the threshold (for example, , the present invention does not make specific restrictions), if a certain matching value is not met, then update , re-associate diagnostic rules , and recalculate the matching degree of the diagnostic path, adjust the diagnostic path to diagnostic path 2, and perform firmware upgrade, attribute check, performance test (failure), firmware version check, and link status check in sequence according to diagnostic path 2. Among them, in diagnostic path 2, you can remove the stress test (because the performance does not meet the standard, there is no need to continue).

[0099] The embodiments of the present invention ensure that only paths with high matching degrees are adopted through a dynamic adjustment mechanism, thereby reducing misdiagnosis and missed diagnosis, improving diagnostic accuracy and reliability, and automatically regenerating paths to reduce manual intervention and speed up fault diagnosis. When faced with complex or new faults, more accurate paths can be generated through iteration, thereby enhancing the adaptability of the system and improving user satisfaction.

[0100] In step S103 , the diagnosis rules, diagnosis labels and fault types when diagnosing the target server are determined according to the diagnosis path, so as to generate a current fault diagnosis result of the target server based on the diagnosis rules, diagnosis labels and fault types.

[0101] From the above analysis, it can be seen that the embodiment of the present invention can determine the corresponding diagnostic rules, diagnostic labels and fault types when diagnosing the target server according to different diagnostic paths, and then generate the current fault diagnosis results of the target server based on the diagnostic rules, diagnostic labels and fault types.

[0102] Optionally, in one embodiment of the present invention, the method further includes: determining a fault level when a fault occurs in the target server based on the fault diagnosis result; and generating corresponding fault handling measures based on different fault levels.

[0103] It can be understood that the embodiments of the present invention can divide the fault level into a first-level fault level, such as core business interruption; a second-level fault level, such as critical business degradation; a third-level fault level, such as business performance degradation; a fourth-level fault level, such as no business impact, etc. The present invention does not impose specific limitations.

[0104] Furthermore, embodiments of the present invention can perform standby machine switching, resource isolation, etc. in the case of a first-level fault level; perform node migration, etc. in the case of a second-level fault level; perform component replacement, etc. in the case of a third-level fault level; and perform monitoring threshold adjustment, etc. in the case of a fourth-level fault level. The present invention does not impose specific limitations.

[0105] As a possible implementation method, the embodiment of the present invention can determine the fault level based on the fault diagnosis result and then generate corresponding fault handling measures.

[0106] The embodiment of the present invention adopts different processing measures for faults of different levels, optimizes resource allocation, avoids resource waste, reduces manual intervention, and improves response speed. Each fault can be handled according to a standard process, improving the consistency and reliability of processing and optimizing resource allocation.

[0107] The working principle of the server fault diagnosis method proposed in the embodiment of the present invention is introduced below with reference to a specific embodiment.

[0108] Figure 5 The present invention is a flowchart of the working principle of a server fault diagnosis method provided according to an embodiment of the present invention.

[0109] Step S501: Collect the configuration information, actual operation data and historical fault data of the server.

[0110] Step S502: Determine whether the server type is a new online server.

[0111] In this embodiment of the present invention, if the server type is a newly launched server, step S503 is executed; if the server type is a mass-produced server, step S504 is executed.

[0112] Step S503: Determine a diagnostic path according to the hardware information of the server.

[0113] Step S504: Determine a diagnostic path based on the server's model information, hardware information, actual operation data, and historical fault data.

[0114] Step S505: Determine whether the server fails.

[0115] In this embodiment of the present invention, if the server does not fail, step S506 is executed; otherwise, step S507 is executed.

[0116] Step S506: determining the diagnosis rules, diagnosis labels and fault types when diagnosing the target server according to the diagnosis path, so as to generate the current fault diagnosis result of the target server based on the diagnosis rules, diagnosis labels and fault types.

[0117] Step S507: using the pre-built fault path generation model to adjust the diagnosis path.

[0118] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0119] The server fault diagnosis method proposed in an embodiment of the present invention can collect the actual operation data, configuration information and historical fault data of the target server, and then use the actual operation data, configuration information and at least one of the historical fault data to detect whether the target server has a fault. When a fault is detected in the target server, the configuration information, actual operation data and historical fault data are input into a pre-built fault path generation model, and then the diagnostic path when the target server has a fault is output, thereby determining the diagnostic rules, diagnostic labels and fault types when diagnosing the target server, and then generating the current fault diagnosis result of the target server. Therefore, it can solve the technical problems of rigid and single processes, inability to dynamically adjust according to the real-time operation status and fault conditions of the server, lack of flexibility and pertinence, low diagnostic efficiency and poor diagnostic accuracy. In addition, if the fault cannot be quickly resolved, it is very likely to cause production interruption and bring huge economic losses to the enterprise. The method achieves the technical effect of dynamically generating accurate and efficient diagnostic paths, shortening diagnosis time, reducing server energy consumption and maintenance costs, improving fault location accuracy, reducing fault resolution time, and achieving automated fault diagnosis for different types of servers.

[0120] An embodiment of the present invention also provides a server fault diagnosis device.

[0121] Figure 6The present invention is a block diagram of a server fault diagnosis device according to an embodiment of the present invention.

[0122] like Figure 6 As shown, the server fault diagnosis device 10 includes: a detection module 100 , an output module 200 and a diagnosis module 300 .

[0123] Among them, the detection module 100 is used to collect the actual operation data of the target server and obtain at least one of the configuration information and historical fault data of the target server, so as to use the actual operation data and at least one of the configuration information and historical fault data to detect whether a fault occurs in the target server.

[0124] The output module 200 is used to input configuration information, actual operation data and historical fault data into a pre-built fault path generation model when a fault is detected on the target server, so as to output a diagnostic path when the fault occurs on the target server.

[0125] The diagnosis module 300 is used to determine the diagnosis rules, diagnosis labels and fault types when diagnosing the target server according to the diagnosis path, so as to generate the current fault diagnosis result of the target server based on the diagnosis rules, diagnosis labels and fault types.

[0126] Optionally, in one embodiment of the present invention, it further includes: a calculation module, a first determination module, a second determination module and a construction module.

[0127] Among them, the calculation module is used to calculate the hardware matching degree between the configuration information and the corresponding diagnostic rules, the operation matching degree between the actual operation data and the corresponding diagnostic rules, and the fault matching degree between the historical fault data and the corresponding diagnostic rules based on the configuration information, actual operation data and historical fault data before inputting the configuration information, actual operation data and historical fault data into the pre-built fault path generation model.

[0128] The first determination module is used to determine weight coefficients corresponding to the hardware matching degree, the operation matching degree and the fault matching degree respectively based on the hardware matching degree, the operation matching degree and the fault matching degree.

[0129] The second determination module is used to determine the matching degree of the corresponding diagnosis rule based on the weight coefficient, the hardware matching degree, the operation matching degree and the fault matching degree.

[0130] The building module is used to build a fault path generation model based on different diagnosis rules and the matching degrees of different diagnosis rules.

[0131] Optionally, in one embodiment of the present invention, it further includes: a first generating module, a first judging module, a second generating module and a third generating module.

[0132] Among them, the first generation module is used to generate trigger conditions for triggering diagnostic rules in the pre-built fault path generation model based on the configuration information, actual operation data and historical fault data before inputting the configuration information, actual operation data and historical fault data into the pre-built fault path generation model.

[0133] The first judgment module is used to judge whether the diagnosis rule exists based on the trigger condition.

[0134] The second generating module is used to screen the diagnosis rules according to preset conditions when the diagnosis rules exist, so as to obtain the diagnosis rules applicable to the target server when a fault occurs.

[0135] The third generating module is used to establish corresponding diagnostic rules based on preset conditions, configuration information, actual operation data and historical fault data when the diagnostic rules do not exist, so as to obtain the established diagnostic rules.

[0136] Optionally, in one embodiment of the present invention, the output module 200 includes: a judgment unit and a first determination unit.

[0137] The judgment unit is used to judge whether the diagnosis rule meets the preset fault diagnosis conditions.

[0138] The first determining unit is configured to determine a diagnosis path based on the diagnosis rule when the diagnosis rule meets a preset fault diagnosis condition.

[0139] Optionally, in one embodiment of the present invention, it further includes: a second judgment module and a fourth generation module.

[0140] The second judgment module is configured to judge whether the matching degree of the diagnostic path is less than a preset matching degree value before determining the diagnostic rule, diagnostic label and fault type when diagnosing the target server according to the diagnostic path.

[0141] The fourth generation module is used to regenerate the diagnostic path using a pre-built fault path generation model when the matching degree of the diagnostic path is less than a preset matching degree value, until the matching degree of the regenerated diagnostic path is greater than or equal to the preset matching degree value.

[0142] Optionally, in one embodiment of the present invention, the output module 200 includes: a second determination unit, an identification unit, and a generation unit.

[0143] The second determining unit is configured to determine a diagnostic tag when a fault occurs on the target server based on at least one of configuration information, actual operation data, and historical fault data.

[0144] The identification unit is used to identify the fault type of the target server based on the diagnosis tag.

[0145] The generation unit is used to calculate the diagnosis path based on the diagnosis label and the fault type by using a pre-built fault path generation model.

[0146] Optionally, in one embodiment of the present invention, it further includes: a processing module.

[0147] Among them, the processing module is used to process the configuration information, actual operation data and historical fault data respectively before inputting them into the pre-built fault path generation model to obtain the configuration information, actual operation data and historical fault data that meet the preset data format conditions.

[0148] Optionally, in one embodiment of the present invention, it further includes: a third determining module and a fifth generating module.

[0149] The third determination module is used to determine the fault level of the target server when a fault occurs based on the fault diagnosis result.

[0150] The fifth generation module is used to generate corresponding fault handling measures based on different fault levels.

[0151] For the description of the features in the embodiment corresponding to the server fault diagnosis device, please refer to the relevant description of the embodiment corresponding to the server fault diagnosis method, which will not be repeated here.

[0152] The server fault diagnosis device proposed in accordance with an embodiment of the present invention can collect actual operation data, configuration information, and historical fault data of a target server, and then use the actual operation data, configuration information, and at least one of the historical fault data to detect whether a fault occurs on the target server. When a fault is detected on the target server, the configuration information, actual operation data, and historical fault data are input into a pre-built fault path generation model, and then a diagnostic path is output when the target server fails, thereby determining the diagnostic rules, diagnostic labels, and fault types when diagnosing the target server, and then generating the current fault diagnosis result of the target server. Therefore, it can solve the technical problems of rigid and single processes, inability to dynamically adjust according to the real-time operation status and fault conditions of the server, lack of flexibility and pertinence, low diagnostic efficiency, and poor diagnostic accuracy. In addition, if a fault cannot be quickly resolved, it is very likely to cause production interruption and bring huge economic losses to the enterprise. The system can achieve the technical effect of dynamically generating accurate and efficient diagnostic paths, shortening diagnostic time, reducing server energy consumption and maintenance costs, improving fault location accuracy, reducing fault resolution time, and achieving automated fault diagnosis for different types of servers.

[0153] An embodiment of the present invention further provides a server, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above server fault diagnosis method embodiments.

[0154] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned server fault diagnosis method embodiments when running.

[0155] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0156] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server fault diagnosis method embodiments are implemented.

[0157] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps of any of the above-mentioned server fault diagnosis method embodiments.

[0158] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0159] The above is a detailed introduction to the server fault diagnosis method provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A server fault diagnosis method, characterized in that: The following steps are involved: Collecting actual operation data of a target server, and obtaining configuration information of the target server and historical fault data corresponding to the configuration information, so as to detect whether a fault occurs on the target server using the actual operation data, the configuration information, and the historical fault data; In the case where a fault is detected on the target server, the configuration information, the actual operation data, and the historical fault data are input into a pre-built fault path generation model, so as to determine, by using the pre-built fault path generation model, actual hardware matching degrees between different diagnostic rules and the configuration information, actual operation matching degrees between the different diagnostic rules and the actual operation data, and actual fault matching degrees between the different diagnostic rules and the historical fault data, and output a diagnostic path when the fault occurs on the target server based on the actual hardware matching degrees, the actual operation matching degrees, and the actual fault matching degrees; determining, according to the diagnostic path, a diagnostic rule, a diagnostic label, and a fault type when diagnosing the target server, so as to generate a current fault diagnosis result of the target server based on the diagnostic rule, the diagnostic label, and the fault type, wherein the diagnostic label includes at least one of a hardware diagnostic label, a software diagnostic label, an environmental diagnostic label, and a fault severity diagnostic label; Before inputting the configuration information, the actual operation data and the historical fault data into the pre-built fault path generation model, the method further includes: generating, based on the configuration information, the actual operation data, and the historical fault data, a triggering condition for triggering a diagnostic rule in the pre-built fault path generation model; Based on the trigger condition, determining whether the diagnostic rule exists; If the diagnosis rule exists, screening the diagnosis rule according to the preset conditions to obtain a diagnosis rule applicable to when a fault occurs on the target server; If the diagnosis rule does not exist, a corresponding diagnosis rule is established based on the preset condition, the configuration information, the actual operation data and the historical fault data to obtain an established diagnosis rule.

2. The server fault diagnosis method according to claim 1, characterized in that: Before inputting the configuration information, the actual operation data, and the historical fault data into the pre-built fault path generation model, the method further includes: Based on the configuration information, the actual operation data, and the historical fault data, respectively calculating a hardware matching degree between the configuration information and a corresponding diagnostic rule, an operation matching degree between the actual operation data and the corresponding diagnostic rule, and a fault matching degree between the historical fault data and the corresponding diagnostic rule; Based on the hardware matching degree, the operation matching degree, and the fault matching degree, respectively determine weight coefficients corresponding to the hardware matching degree, the operation matching degree, and the fault matching degree; Determining a matching degree of the corresponding diagnostic rule based on the weight coefficient, the hardware matching degree, the operation matching degree, and the fault matching degree; A fault path generation model is constructed based on different diagnosis rules and the matching degrees of the different diagnosis rules.

3. The server fault diagnosis method according to claim 2, characterized in that: Inputting the configuration information, the actual operation data, and the historical fault data into a pre-built fault path generation model to output a diagnostic path when a fault occurs on the target server includes: Determining whether the diagnostic rule meets a preset fault diagnosis condition; If the diagnosis rule satisfies the preset fault diagnosis condition, the diagnosis path is determined based on the diagnosis rule.

4. The server fault diagnosis method according to claim 1, characterized in that: Before determining the diagnostic rule, diagnostic label, and fault type for diagnosing the target server according to the diagnostic path, the method further includes: Determining whether the matching degree of the diagnostic path is less than a preset matching degree value; If the matching degree of the diagnostic path is less than the preset matching degree value, the diagnostic path is regenerated using the pre-built fault path generation model until the matching degree of the regenerated diagnostic path is greater than or equal to the preset matching degree value.

5. The server fault diagnosis method according to claim 1, characterized in that: Inputting the configuration information, the actual operation data, and the historical fault data into a pre-built fault path generation model to output a diagnostic path when a fault occurs on the target server includes: Determining a diagnostic tag when a fault occurs on the target server based on the configuration information, the actual operation data, and the historical fault data; Based on the diagnostic tag, identifying a fault type of the target server; The diagnostic path is calculated based on the diagnostic label and the fault type using the pre-built fault path generation model.

6. The server fault diagnosis method according to claim 1, characterized in that: Before inputting the configuration information, the actual operation data, and the historical fault data into the pre-built fault path generation model, the method further includes: The configuration information, the actual operation data and the historical fault data are processed respectively to obtain the configuration information, the actual operation data and the historical fault data that meet preset data format conditions.

7. The server fault diagnosis method according to claim 1, characterized in that: Also includes: Determining a fault level when a fault occurs on the target server based on the fault diagnosis result; Generate corresponding fault handling measures based on different fault levels.

8. A server fault diagnosis device, characterized in that: include: a detection module, configured to collect actual operation data of a target server, and obtain configuration information of the target server and historical fault data corresponding to the configuration information, so as to detect whether a fault occurs on the target server using the actual operation data, the configuration information, and the historical fault data; an output module for, when a fault is detected on the target server, inputting the configuration information, the actual operation data, and the historical fault data into a pre-built fault path generation model, so as to determine, by using the pre-built fault path generation model, actual hardware matching degrees between different diagnostic rules and the configuration information, actual operation matching degrees between the different diagnostic rules and the actual operation data, and actual fault matching degrees between the different diagnostic rules and the historical fault data, and outputting, based on the actual hardware matching degrees, the actual operation matching degrees, and the actual fault matching degrees, a diagnostic path for the target server when the fault occurs; a diagnosis module, configured to determine, according to the diagnosis path, a diagnosis rule, a diagnosis tag, and a fault type for diagnosing the target server, and generate a current fault diagnosis result for the target server based on the diagnosis rule, the diagnosis tag, and the fault type, wherein the diagnosis tag includes at least one of a hardware diagnosis tag, a software diagnosis tag, an environment diagnosis tag, and a fault severity diagnosis tag; Among them, also include: a first generating module, configured to generate, before inputting the configuration information, the actual operating data, and the historical fault data into the pre-built fault path generation model, a triggering condition for triggering a diagnostic rule in the pre-built fault path generation model based on the configuration information, the actual operating data, and the historical fault data; A first judgment module, configured to judge whether the diagnosis rule exists based on the trigger condition; A second generating module is used to filter the diagnostic rules according to preset conditions when the diagnostic rules exist, so as to obtain diagnostic rules applicable to when a fault occurs on the target server; The third generating module is used to establish a corresponding diagnostic rule based on the preset conditions, the configuration information, the actual operation data and the historical fault data when the diagnostic rule does not exist, so as to obtain the established diagnostic rule.

9. The server fault diagnosis device according to claim 8, characterized in that: Also includes: a calculation module, configured to calculate, before inputting the configuration information, the actual operation data, and the historical fault data into the pre-built fault path generation model, a hardware matching degree between the configuration information and a corresponding diagnostic rule, an operation matching degree between the actual operation data and the corresponding diagnostic rule, and a fault matching degree between the historical fault data and the corresponding diagnostic rule based on at least one of the configuration information, the actual operation data, and the historical fault data; A first determining module is configured to determine weight coefficients corresponding to the hardware matching degree, the operation matching degree, and the fault matching degree, respectively, based on the hardware matching degree, the operation matching degree, and the fault matching degree; a second determining module, configured to determine a matching degree of the corresponding diagnostic rule based on the weight coefficient, the hardware matching degree, the operation matching degree, and the fault matching degree; The construction module is used to construct a fault path generation model based on different diagnosis rules and the matching degrees of the different diagnosis rules.

10. The server fault diagnosis device according to claim 9, characterized in that: The output module includes: A judgment unit, configured to judge whether the diagnosis rule satisfies a preset fault diagnosis condition; The first determining unit is configured to determine the diagnostic path based on the diagnostic rule when the diagnostic rule satisfies the preset fault diagnosis condition.

11. The server fault diagnosis device according to claim 8, characterized in that: Also includes: A second judgment module is configured to judge whether a matching degree of the diagnostic path is less than a preset matching degree value before determining a diagnostic rule, a diagnostic label, and a fault type for diagnosing the target server according to the diagnostic path; The fourth generation module is used to regenerate the diagnostic path using the pre-built fault path generation model when the matching degree of the diagnostic path is less than the preset matching degree value, until the matching degree of the regenerated diagnostic path is greater than or equal to the preset matching degree value.

12. A server, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the server fault diagnosis method according to any one of claims 1 to 7.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the server fault diagnosis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Server fault diagnosis method and device, equipment and medium

    CN117093405A

  • Server fault diagnosis method and device, storage medium and electronic equipment

    CN117667479A