Fault diagnosis method and device for server
By collecting and analyzing the actual operating data and historical fault data of the server, dynamically generates diagnostic paths, solving the problem of process solidification in server fault diagnosis, achieving efficient and accurate fault diagnosis, and reducing the risk and cost of production interruptions.
Patent Information
- Application Number
- CN202510846497.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the prior art, the server fault diagnosis process is solidified in a single manner and cannot be dynamically adjusted according to the real-time operating status and fault conditions, resulting in low diagnostic efficiency and poor accuracy. Once the fault cannot be solved quickly, it is easy to cause production interruptions, causing economic losses to the enterprise.
By collecting the actual operating data of the server and historical fault data, using the fault path generation model to dynamically generate diagnostic paths, determine diagnostic rules, labels and fault types, and achieve dynamic and accurate fault diagnosis.
It improves the flexibility and pertinence of fault diagnosis, shortens diagnosis time, reduces server energy consumption and maintenance costs, improves fault positioning accuracy, and reduces the risk of production interruptions.
Smart Images

Figure CN120354178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular, to a method and device for fault diagnosis of a server. Background Art
[0002] In the related art, the faults of a server can be diagnosed according to a certain diagnosis process. For example, first, the server is preliminarily set according to the factory specifications; then, the basic hardware of the server is detected; then, the server is subjected to a function test; and finally, a stress test of the core hardware of the server is performed, so as to comprehensively ensure the stable and efficient operation of the server. In addition, during diagnosis, the server product can also be made to meet the factory specifications through manual intervention.
[0003] However, in the related art, the process is fixed and single, and cannot be dynamically adjusted according to the real-time operation state and fault conditions of the server, lacking flexibility and pertinence. Especially when the server runs abnormally, the diagnosis direction and depth cannot be changed in time, resulting in low diagnosis efficiency, reduced diagnosis accuracy, and serious waste of resources. In addition, once the fault cannot be quickly resolved, it is extremely easy to cause production interruption, bringing huge economic losses to the enterprise, and there is an urgent need for improvement. Summary of the Invention
[0004] The present invention provides a method and device for fault diagnosis of a server, so as to at least solve the problems in the related art that the process is fixed and single, cannot be dynamically adjusted according to the real-time operation state and fault conditions of the server, lacks flexibility and pertinence, has low diagnosis efficiency, poor diagnosis accuracy, and in addition, once the fault cannot be quickly resolved, it is extremely easy to cause production interruption, bringing huge economic losses to the enterprise.
[0005] The present invention provides a method for fault diagnosis of a server, including the following steps: collecting the actual operation data of a target server, and obtaining at least one of the configuration information and historical fault data of the target server, so as to detect whether the target server has a fault by using the actual operation data and at least one of the configuration information and the historical fault data; in the case of detecting that the target server has a fault, inputting the configuration information, the actual operation data, and the historical fault data into a pre-constructed fault path generation model to output a diagnosis path when the target server has a fault; determining a diagnosis rule, a diagnosis label, and a fault type for diagnosing the target server according to the diagnosis path, so as to generate a current fault diagnosis result of the target server based on the diagnosis rule, the diagnosis label, and the fault type.
[0006] The present invention also provides a fault diagnosis device for a server, including: a detection module, configured to collect the actual operation data of a target server, and obtain at least one of the configuration information and historical fault data of the target server, so as to detect whether the target server has a fault by using the actual operation data and at least one of the configuration information and the historical fault data; an output module, configured to, when it is detected that the target server has a fault, input the configuration information, the actual operation data and the historical fault data into a pre-constructed fault path generation model, so as to output a diagnosis path when the target server has a fault; a diagnosis module, configured to determine a diagnosis rule, a diagnosis label and a fault type for diagnosing the target server according to the diagnosis path, so as to generate a current fault diagnosis result of the target server based on the diagnosis rule, the diagnosis label and the fault type.
[0007] The present invention also provides a server, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above-mentioned server fault diagnosis methods when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned server fault diagnosis methods are implemented.
[0009] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above-mentioned server fault diagnosis methods are implemented.
[0010] Through the present invention, the actual operation data, configuration information and historical fault data of the target server can be collected, and then at least one of the actual operation data, configuration information and historical fault data is used to detect whether the target server has a fault. When it is detected that the target server has a fault, the configuration information, actual operation data and historical fault data are input into a pre-constructed fault path generation model, and then the diagnosis path when the target server has a fault is output, so as to determine the diagnosis rule, diagnosis label and fault type for diagnosing the target server, and then generate the current fault diagnosis result of the target server. Therefore, the technical problems of fixed and single processes, inability to dynamically adjust according to the real-time operation state and fault conditions of the server, lack of flexibility and pertinence, low diagnosis efficiency, poor diagnosis accuracy, and in addition, once the fault cannot be quickly solved, it is extremely easy to cause production interruption and bring huge economic losses to the enterprise can be solved, and the technical effects of dynamically generating accurate and efficient diagnosis paths, shortening the diagnosis time, reducing the energy consumption and maintenance cost of the server, improving the fault location accuracy rate, reducing the fault solution time, and realizing automatic fault diagnosis for different types of servers can be achieved. Description of the Drawings
[0011] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It is a flowchart of a server fault diagnosis method provided according to an embodiment of the present invention; Figure 2 It is a flowchart of constructing a fault path generation model provided according to an embodiment of the present invention; Figure 3 It is a schematic diagram of a certain server product configuration provided according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the production fault repair of a certain server product provided according to an embodiment of the present invention; Figure 5 It is a flowchart of the working principle of a server fault diagnosis method provided according to an embodiment of the present invention; Figure 6 It is a block diagram of a server fault diagnosis device provided according to an embodiment of the present invention.
[0013] Reference numerals: Among them, 10 - server fault diagnosis device; 100 - detection module, 200 - output module, 300 - diagnosis module. Detailed implementation manners
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0015] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0016] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0017] An embodiment of the present invention provides a method for diagnosing faults of a server. The method will be described in detail in combination with the execution process of the method for diagnosing faults of the server.
[0018] Specifically, Figure 1 FIG. is a flowchart of a method for diagnosing faults of a server according to an embodiment of the present invention.
[0019] As Figure 1 shown, the method for diagnosing faults of the server includes the following steps: In step S101, actual operation data of the target server is collected, and at least one of the configuration information and historical fault data of the target server is obtained, so as to detect whether the target server has a fault by using the actual operation data and at least one of the configuration information and historical fault data.
[0020] It can be understood that in the embodiment of the present invention, the configuration information may include, but is not limited to, the model information of the target server, hardware information (such as firmware version, hardware model, network mode, RAID (Redundant Array of Independent Disks) status, etc., which are not specifically limited in the present invention), CPU (Central Processing Unit) data (such as the threads and frequencies of the CPU, etc., which are not specifically limited in the present invention), hard disk data (such as the capacity and transfer rate of the hard disk, etc., which are not specifically limited in the present invention), memory data (such as the capacity and transfer rate of the memory, etc., which are not specifically limited in the present invention), etc., which are not specifically limited in the present invention; the actual operation data may include, but is not limited to, the log file, environmental temperature, working voltage, performance indicators, etc. of the target server, which are not specifically limited in the present invention; the historical fault data may include, but is not limited to, maintenance data (such as maintenance measures, etc., which are not specifically limited in the present invention), fault data (such as fault time, error code, fault type, occurrence frequency, etc., which are not specifically limited in the present invention), etc., within a certain period of time, which are not specifically limited in the present invention. Among them, the certain period of time can be set by those skilled in the art according to the actual situation, which is not specifically limited in the present invention.
[0021] Furthermore, embodiments of the present invention can collect the actual operation data of the target server in real time through the combination of IPMI (Intelligent Platform Management Interface) and BMC (Baseboard Management Controller). The sampling frequency can be increased from 1 time per 10 minutes to 1 time per 1 minute, which can be specifically set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations. In addition, embodiments of the present invention can record the hardware configuration information every once in a while, such as 10 minutes, which is not specifically limited by the present invention, and transmit the hardware configuration information to the database for a hardware status snapshot.
[0022] In some embodiments, embodiments of the present invention can collect the configuration information, actual operation data, and historical fault data of the target server, and then detect whether the target server has a fault based on the configuration information, actual operation data, and historical fault data.
[0023] In some embodiments, embodiments of the present invention can collect the actual operation data of the target server and obtain the historical fault data of the target server, and then use the actual operation data and the historical fault data to detect whether the target server has a fault.
[0024] Exemplarily, embodiments of the present invention can collect the actual operation data of the target server in real time by combining IPMI and BMC, with a sampling frequency of 1 time per 1 minute, record the hardware configuration information every 10 minutes, transmit it to the database, and perform a hardware status snapshot, and then combine the historical fault data to detect whether the target server has a fault.
[0025] Embodiments of the present invention deeply integrate multi-source data such as the configuration information, actual operation data, and historical fault data of the server, break through the limitations of traditional single reliance on preset strategies, no longer adopt a "one-size-fits-all" fixed diagnosis process, but customize the diagnosis path according to the unique data combination of each server.
[0026] Optionally, in an embodiment of the present invention, before inputting the configuration information, actual operation data, and historical fault data into a pre-constructed fault path generation model, it further includes: calculating the hardware matching degree between the configuration information and the corresponding diagnostic rule, the operation matching degree between the actual operation data and the corresponding diagnostic rule, and the fault matching degree between the historical fault data and the corresponding diagnostic rule based on the configuration information, actual operation data, and historical fault data respectively; determining the weight coefficients corresponding to the hardware matching degree, operation matching degree, and fault matching degree based on the hardware matching degree, operation matching degree, and fault matching degree respectively; determining the matching degree of the corresponding diagnostic rule based on the weight coefficients, hardware matching degree, operation matching degree, and fault matching degree; constructing a fault path generation model based on different diagnostic rules and the matching degrees of different diagnostic rules.
[0027] It can be understood that the embodiments of the present invention can uniformly manage different diagnostic rules and store them in a rule library in json format. Each diagnostic rule is for different model information, hardware configuration information, and application scenarios. It can be understood that the rule library contains multiple diagnostic rules , where , and is an integer. Each diagnostic rule includes a trigger condition, an action, a priority weight, etc., which are not specifically limited in the present invention.
[0028] In some embodiments, the embodiments of the present invention can calculate the hardware matching degree between the configuration information and the corresponding diagnostic rule according to different diagnostic rules.
[0029] In some embodiments, the embodiments of the present invention can calculate the operation matching degree between the actual operation data and the corresponding diagnostic rule according to different diagnostic rules.
[0030] In some embodiments, the embodiments of the present invention can calculate the fault matching degree between the historical fault data and the corresponding diagnostic rule according to different diagnostic rules.
[0031] Further, the embodiments of the present invention can calculate the matching degrees (0-1) of different diagnostic rules with the configuration information, actual operation data, and historical fault data according to the hardware matching degree, operation matching degree, and fault matching degree. The calculation formula can be but is not limited to: , where , , are weight coefficients (where the embodiments of the present invention can set , , to , , respectively, which are not specifically limited in the present invention); The hardware matching degree between the hardware configuration required by the diagnostic rule and the current hardware configuration (for example, if there is an NVMe (Non-Volatile Memory Express) hard disk, the matching degree is 1; otherwise, it is 0. The present invention does not make specific limitations); The matching degree between the actual operation data and the diagnostic rule trigger condition (for example, if the temperature exceeds the threshold, the matching degree is 1. The present invention does not make specific limitations); The normalized value of the correlation degree between the historical fault data and the diagnostic rule (for example, if the occurrence frequency of a fault associated with a certain diagnostic rule is 30%, the matching degree is 0.3. The present invention does not make specific limitations); The current configuration information of the server; The actual operation data; The historical fault data.
[0032] Furthermore, embodiments of the present invention can construct a fault path generation model based on different diagnostic rules and the matching degrees of different diagnostic rules. Its expression can be but is not limited to: , Among them, represents the priority weight of the th rule. is the rule index. The diagnostic rule with a higher weight plays a more important role in path generation and affects the priority sorting of the diagnostic process; is the th candidate path. The objective function calculates the weighted matching degree of each and selects the path with the highest matching degree as the final diagnostic path to achieve dynamic optimization. is the index identifier of the candidate path, used to distinguish different diagnostic path schemes.
[0033] Exemplarily, embodiments of the present invention can construct a fault path generation model in combination with Figure 2 as shown. Its main content is: Step S201: Obtain the configuration information, actual operation data, and historical fault data of the server.
[0034] Step S202: Obtain different diagnostic rules.
[0035] Step S203: Calculate the matching degrees of different diagnostic rules with the configuration information, actual operation data, and historical fault data.
[0036] Step S204: Sort different diagnostic rules based on different matching degrees.
[0037] Among them, the embodiments of the present invention can screen out diagnostic rules with low matching degrees based on a certain threshold and arrange them in descending order according to the magnitudes of the matching degrees. The certain threshold can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0038] Step S205: Construct an initial fault path.
[0039] Among them, the embodiments of the present invention can construct an initial fault path through greedy selection.
[0040] Step S206: Real-time monitoring and dynamic adjustment.
[0041] Among them, the embodiments of the present invention can perform real-time monitoring and dynamic adjustment based on different diagnostic rules and corresponding initial fault paths.
[0042] Step S207: Update the matching degree.
[0043] Among them, the embodiments of the present invention can recalculate the matching degree of the current diagnostic rule when the current diagnostic rule fails to execute or the conditions change.
[0044] Step S208: Adjust the initial fault path to obtain a fault path.
[0045] Among them, the embodiments of the present invention can adjust the initial fault path in the case of inserting a new diagnostic rule or removing an invalid diagnostic rule, and then obtain a fault path.
[0046] The embodiments of the present invention adopt a triple data verification mechanism of hardware matching degree, operation matching degree, and fault matching degree to improve the diagnostic accuracy rate, dynamically adjust the weights, ensure the diagnostic precision, pre-generate the fault path, and meet the real-time requirements of industrial control scenarios.
[0047] Optionally, in an embodiment of the present invention, before inputting the configuration information, actual operation data, and historical fault data into a pre-constructed fault path generation model, it further includes: generating a trigger condition for triggering the diagnostic rules in the pre-constructed fault path generation model based on the configuration information, actual operation data, and historical fault data; judging whether the diagnostic rules exist based on the trigger condition; if the diagnostic rules exist, screening the diagnostic rules according to preset conditions to obtain diagnostic rules applicable to the occurrence of a fault in the target server; if the diagnostic rules do not exist, establishing corresponding diagnostic rules based on the preset conditions, configuration information, actual operation data, and historical fault data to obtain the established diagnostic rules.
[0048] It is understandable that the embodiments of the present invention can establish corresponding diagnostic rules based on the collected configuration information, actual operation data, and historical fault data, and update or delete the corresponding diagnostic rules in a timely manner according to different triggering conditions to ensure the accuracy and effectiveness of the diagnostic rules.
[0049] In some embodiments, the embodiments of the present invention can generate triggering conditions for diagnosing rules in a pre-constructed fault path generation model based on configuration information, actual operation data, and historical fault data, and then determine whether the diagnostic rules exist based on the triggering conditions; if they exist, the diagnostic rules are screened according to certain conditions to obtain diagnostic rules applicable to the occurrence of faults in the target server; if they do not exist, corresponding diagnostic rules are established based on certain conditions, configuration information, actual operation data, and historical fault data, and then the established diagnostic rules are obtained. Among them, the certain conditions can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0050] Among them, in the embodiments of the present invention, the triggering conditions can include, but are not limited to, condition judgments, processing actions, and dependency conditions between diagnostic processes, etc., and the present invention does not make specific limitations.
[0051] Exemplarily, the embodiments of the present invention can formulate different diagnostic rules according to different server fault situations, and adjust the corresponding diagnostic rules based on configuration information, actual operation data, and historical fault data during the diagnosis process. For example, if a certain triggering condition is triggered in the embodiments of the present invention, the diagnostic path can be adjusted according to the processing action of the diagnostic rule, and the corresponding diagnostic tasks can be increased or decreased. For example, in the NVMe hard disk performance test, when it is detected that the hard disk read and write rate does not meet the baseline standard, the test is stopped and a comprehensive hard disk inspection task is added.
[0052] Among them, in the embodiments of the present invention, the diagnostic rules corresponding to the NVMe hard disk performance test can be: first, obtain the diagnostic rules for different NVMe hard disk operations, then introduce the uses of each diagnostic rule, then set the triggering conditions for different diagnostic rules, such as the executed operations and relevant historical fault situations, etc., and the present invention does not make specific limitations. Finally, based on the configuration information, actual operation data, and historical fault data, determine the corresponding processing actions, priorities, and dependencies.
[0053] In addition, when adjusting the diagnostic rules, the embodiments of the present invention can pause or adjust the diagnostic rules that depend on abnormal components according to the dependency relationships and priorities between different diagnostic rules. For example, during the CPU stress test, when a CPU fault is detected and the CPU is replaced, according to the "xx" parameter in the diagnostic rule, the CPU status check, CPU performance check, and stress test can be performed in sequence, and the present invention does not make specific limitations.
[0054] In the embodiments of the present invention, by generating trigger conditions, relevant diagnostic rules can be located more quickly, avoiding wasting time on irrelevant rules and improving diagnostic efficiency. When the diagnostic rules do not exist, new rules can be established based on preset conditions and data, enhancing the system adaptability. By screening diagnostic rules applicable to the current fault, computing resources are optimized and diagnostic accuracy is improved.
[0055] Optionally, in an embodiment of the present invention, before inputting the configuration information, actual operation data, and historical fault data into a pre-constructed fault path generation model, it further includes: respectively processing the configuration information, actual operation data, and historical fault data to obtain configuration information, actual operation data, and historical fault data that meet the preset data format conditions.
[0056] It can be understood that the embodiments of the present invention can store the configuration information, actual operation data, and historical fault data in a combined manner of a distributed file system and a distributed database. For example, the embodiments of the present invention can store structured data, such as hardware information, fault time, error codes, etc., which are not specifically limited in the present invention, into a Hive table, and use the SQL-like query language in the Hive table for efficient query. Its storage schematic diagram is as Figure 3 shown; for unstructured data, such as fault types, maintenance measures, log files, etc., which are not specifically limited in the present invention, store them in a document-oriented database MongoDB and store them in json format. Its storage schematic diagram is as Figure 4 shown.
[0057] In some embodiments, the embodiments of the present invention can process the configuration information, actual operation data, and historical fault data to obtain configuration information, actual operation data, and historical fault data that meet certain data format conditions. Among them, the certain data format conditions can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0058] Exemplarily, the embodiments of the present invention can use big data collection tools to collect the log files, maintenance measures, fault types, and fault times of the server in real time, and perform cleaning, transformation, and integration on the collected data through the ETL (Extract Transform Load) process to make it meet a certain data format, and then uniformly store it in a Hive table for subsequent query and analysis.
[0059] In addition, it should be noted that in the embodiments of the present invention, when analyzing log files, fault types, fault times, etc., big data analysis tools can be used to replace inefficient manual intervention. By using the xx algorithm, the fault frequencies of different hardware information in different diagnostic rules are statistically analyzed, the possibility of a fault occurring is predicted based on the hardware information, actual operation data, etc., and the analysis results are fed back to the diagnostic rules, thereby optimizing the fault path and diagnostic rules and improving the accuracy and efficiency of diagnosis.
[0060] The embodiments of the present invention improve data quality through data processing, ensure the accuracy of subsequent processing, enhance the reliability and compatibility of the model, and save resources.
[0061] In step S102, when it is detected that the target server fails, the configuration information, actual operation data, and historical fault data are input into a pre-constructed fault path generation model to output the diagnostic path when the target server fails.
[0062] It can be understood that in the embodiments of the present invention, the pre-constructed fault path generation model may include multiple diagnostic rules. Each diagnostic rule is directed to the model information, hardware information, and actual operation data of different servers. In addition, each diagnostic rule has clear trigger conditions, condition judgments, processing actions, and dependency conditions between diagnostic processes, etc., which are not specifically limited in the present invention.
[0063] As a possible implementation manner, in the embodiments of the present invention, when the target server fails, the configuration information, actual operation data, and historical fault data can be input into a pre-constructed fault path generation model, and then the diagnostic path when the target server fails is output.
[0064] Exemplarily, in the embodiments of the present invention, when the target server has fault one, the diagnostic rule A for the target server in the fault one state can be determined according to configuration information A, actual operation data A, and historical fault data A. Then, the configuration information A, actual operation data A, historical fault data A, and diagnostic rule A are input into a pre-constructed fault path generation model, and then the diagnostic path when the target server has fault one is output.
[0065] In the embodiments of the present invention, when the target server has fault two, the diagnostic rule B for the target server in the fault two state can be determined according to configuration information B, actual operation data B, and historical fault data B. Then, the configuration information B, actual operation data B, historical fault data B, and diagnostic rule B are input into a pre-constructed fault path generation model, and then the diagnostic path when the target server has fault two is output.
[0066] Optionally, in an embodiment of the present invention, configuration information, actual operation data, and historical fault data are input into a pre-constructed fault path generation model to output a diagnostic path when the target server fails, including: determining whether the diagnostic rule meets the preset fault diagnosis condition; if the diagnostic rule meets the preset fault diagnosis condition, determining the diagnostic path based on the diagnostic rule.
[0067] As a possible implementation manner, when inputting configuration information, actual operation data, and historical fault data into the pre-constructed fault path generation model in the embodiment of the present invention, it can be determined whether the diagnostic rule meets a certain fault diagnosis condition, and when it is met, the diagnostic path is determined based on the diagnostic rule. Among them, the certain fault diagnosis condition can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0068] Exemplarily, in the embodiment of the present invention, it can be determined whether the diagnostic rule meets a certain fault diagnosis condition according to the type of the server, and when it is met, the diagnostic path is determined.
[0069] Furthermore, in the case where the server is a newly launched server in the embodiment of the present invention, the corresponding diagnostic rule can be selected according to the hardware information of the server, and then the diagnostic path that meets a certain fault diagnosis condition is determined. For example, if the server is equipped with an NVMe hard disk, by selecting the diagnostic rules related to the NVMe hard disk, such as firmware upgrade check, attribute check, performance test, etc., and when the diagnostic rule meets a certain fault diagnosis condition, the diagnostic path is determined based on the corresponding diagnostic rule.
[0070] In the case where the server is a mass-produced server in the embodiment of the present invention, the diagnostic rules that meet a certain fault diagnosis condition can be screened according to the model information, hardware information, actual operation data, and historical fault data of the server, and then the diagnostic path is determined.
[0071] In addition, it should be noted that when determining the diagnostic path in the embodiment of the present invention, the priorities of server registration, firmware upgrade and check, power redundancy test, hardware performance test and check, hardware stress test, human-computer interaction, and server factory settings are followed.
[0072] The embodiment of the present invention ensures that only relevant and effective rules are applied through a certain fault diagnosis condition, reduces misdiagnosis, improves diagnostic accuracy, avoids unnecessary rule checks, optimizes resource utilization, shortens the diagnostic time, adjusts rule application according to different fault situations, improves flexibility, supports complex fault handling, and generates a more comprehensive diagnostic path.
[0073] Optionally, in an embodiment of the present invention, configuration information, actual operation data, and historical fault data are input into a pre-constructed fault path generation model to output a diagnostic path when the target server fails, including: determining a diagnostic label when the target server fails based on the configuration information, actual operation data, and historical fault data; identifying the fault type of the target server based on the diagnostic label; and calculating the diagnostic path using the pre-constructed fault path generation model based on the diagnostic label and the fault type.
[0074] It can be understood that the embodiments of the present invention can add different diagnostic labels to the target server that fails according to different configuration information, actual operation data, and historical fault data, thereby quickly identifying fault problems. For example, the embodiments of the present invention can add hardware diagnostic labels, such as CPU, memory, hard disk, GPU (Graphics Processing Unit), power supply, etc., which are not specifically limited in the present invention; software diagnostic labels, such as BIOS (Basic Input / Output System) version, firmware version, driver version, OS (Operating System) type, etc., which are not specifically limited in the present invention; environmental diagnostic labels, such as temperature, humidity, voltage fluctuation, load status, etc., which are not specifically limited in the present invention; fault severity diagnostic labels, such as minor alarm, serious fault, system crash, etc., which are not specifically limited in the present invention.
[0075] Furthermore, the embodiments of the present invention can redefine the diagnostic path, uniformly classify the hybrid diagnostic rules into the corresponding diagnostic labels, each diagnostic label has a single and clear function, the internal code is closely related, reduce the dependency relationship between different diagnostic labels, and reduce the coupling degree.
[0076] During the actual execution process, the embodiments of the present invention can determine the diagnostic label when the target server fails based on the configuration information, actual operation data, and historical fault data, and then identify the fault type of the target server, so as to calculate the diagnostic path using the pre-constructed fault path generation model based on the diagnostic label and the fault type.
[0077] Exemplarily, the embodiments of the present invention can divide the diagnostic rules into multiple hardware diagnostic labels according to the hardware information of the server and the nature of the diagnostic rules. Among them, the hardware diagnostic labels can include, but are not limited to, CPU diagnostic labels, memory diagnostic labels, and hard disk diagnostic labels, etc., which are not specifically limited in the present invention.
[0078] Among them, the CPU diagnostic label can be responsible for firmware upgrade, parameter check, performance test, stress test, etc. of the CPU, which are not specifically limited in the present invention.
[0079] The memory diagnostic label can check the memory capacity, in-position status, read / write speed, CE fault screening, etc., and the present invention does not make specific limitations.
[0080] The hard disk diagnostic label can perform hard disk attribute checks, performance tests, stress tests, firmware upgrades, and fault diagnoses, etc. for different types of hard disks (such as NVMe hard disks, etc., and the present invention does not make specific limitations), and the present invention does not make specific limitations.
[0081] Furthermore, in the embodiment of the present invention, taking the hard disk diagnostic label as an example, its diagnostic process can be: perform the upgrade operation of the NVMe hard disk according to the firmware version, upgrade diagnosis, upgrade result, etc. of the NVMe hard disk; perform the attribute check and diagnosis operation of the NVMe hard disk by reading the attribute information of the NVMe hard disk, such as capacity, health status, etc.; perform the stress test and diagnosis operation of the NVMe hard disk by setting stress test parameters and running test tools, etc., and then accurately instantiate and call the corresponding diagnostic rules, and meet the personalized diagnostic needs of servers with different hardware configurations through the combination of diagnostic labels.
[0082] In addition, it should be noted that in the embodiment of the present invention, in order to facilitate the management and invocation of diagnostic labels, a diagnostic label management system is established, which is responsible for the registration, loading, and query of diagnostic labels. In addition, in the embodiment of the present invention, the diagnostic label management system adopts the form of a configuration file, which can define the name, class name, path, etc. of the diagnostic label, and the present invention does not make specific limitations, and unified processing is performed to improve efficiency.
[0083] In the embodiment of the present invention, through diagnostic labels, data from different sources can be uniformly processed, thereby narrowing the diagnostic scope, reducing unnecessary inspections, improving efficiency, and realizing more accurate path calculation.
[0084] Optionally, in an embodiment of the present invention, before the diagnostic rules, diagnostic labels, and fault types when determining the diagnostic target server according to the diagnostic path, it further includes: determining whether the matching degree of the diagnostic path is less than a preset matching degree value; if the matching degree of the diagnostic path is less than the preset matching degree value, then use the pre-constructed fault path generation model to re-generate the diagnostic path until the matching degree of the re-generated diagnostic path is greater than or equal to the preset matching degree value.
[0085] During the actual execution process, before the diagnostic rules, diagnostic labels, and fault types when determining the diagnostic target server according to the diagnostic path in the embodiment of the present invention, it can also be determined whether the matching degree of the diagnostic path is less than a certain matching degree value, and when it is less, use the pre-constructed fault path generation model to re-generate the diagnostic path until the matching degree of the re-generated diagnostic path is greater than or equal to a certain matching degree value. Among them, the certain matching degree value can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0086] Exemplarily, in the embodiment of the present invention, when the NVMe hard disk fails and , in this case, the weight is set, and the matching degree of the diagnostic path 1 corresponding to the diagnostic rule is . Then, firmware upgrade, attribute check, performance test, and stress test are sequentially performed according to diagnostic path 1. When performing the performance test, it is detected that the read / write rate of the NVMe hard disk is lower than the threshold (for example, , which is not specifically limited in the present invention). If the certain matching degree value is not satisfied, then is updated, the diagnostic rule is re-associated, and the matching degree of the diagnostic path is recalculated. The diagnostic path is adjusted to diagnostic path 2, and firmware upgrade, attribute check, performance test (failed), firmware version check, and link status check are sequentially performed according to diagnostic path 2. Among them, in diagnostic path 2, the stress test can be removed (because there is no need to continue due to unqualified performance).
[0087] The embodiment of the present invention ensures that only the path with high matching degree is adopted through the dynamic adjustment mechanism, reduces misdiagnosis and missed diagnosis, improves the diagnostic accuracy and reliability. The automatic regeneration of the path reduces manual intervention, speeds up the fault diagnosis speed. When facing complex or new faults, more accurate paths can be generated iteratively, enhancing the adaptability of the system and improving user satisfaction.
[0088] In step S103, the diagnostic rule, diagnostic label, and fault type when determining the diagnostic target server according to the diagnostic path are used to generate the current fault diagnosis result of the target server based on the diagnostic rule, diagnostic label, and fault type.
[0089] Through the above analysis, it can be seen that the embodiment of the present invention can determine the diagnostic rule, diagnostic label, and fault type corresponding to the diagnostic target server according to different diagnostic paths, and then generate the current fault diagnosis result of the target server based on the diagnostic rule, diagnostic label, and fault type.
[0090] Optionally, in an embodiment of the present invention, it further includes: determining the fault level when the target server fails based on the fault diagnosis result; generating corresponding fault handling measures based on different fault levels.
[0091] It can be understood that the embodiment of the present invention can divide the fault level into a first-level fault level, such as core service interruption; a second-level fault level, such as key service degradation; a third-level fault level, such as service performance decline; a fourth-level fault level, such as no impact on the service, etc., which is not specifically limited in the present invention.
[0092] Further, embodiments of the present invention can perform standby machine switching, resource isolation, etc. in the case of a primary fault level; perform node migration, etc. in the case of a secondary fault level; perform component replacement, etc. in the case of a tertiary fault level; perform monitoring threshold adjustment, etc. in the case of a quaternary fault level. The present invention does not make specific limitations.
[0093] As a possible implementation manner, embodiments of the present invention can determine a fault level based on a fault diagnosis result, and further generate corresponding fault handling measures.
[0094] Embodiments of the present invention adopt different handling measures for different levels of faults, optimize resource allocation, avoid resource waste, reduce manual intervention, improve response speed, and each fault can be processed according to a standard process, improving the consistency and reliability of processing and optimizing resource allocation.
[0095] The working principle of the fault diagnosis method for a server proposed by embodiments of the present invention is introduced below in combination with a specific embodiment.
[0096] Figure 5 FIG. is a flowchart of the working principle of the fault diagnosis method for a server according to an embodiment of the present invention.
[0097] Step S501: Collect the configuration information, actual operation data, and historical fault data of the server.
[0098] Step S502: Determine whether the type of the server is a newly launched server.
[0099] Wherein, in embodiments of the present invention, if the type of the server is a newly launched server, step S503 is executed; if the type of the server is a mass-produced server, step S504 is executed.
[0100] Step S503: Determine a diagnosis path according to the hardware information of the server.
[0101] Step S504: Determine a diagnosis path according to the model information, hardware information, actual operation data, and historical fault data of the server.
[0102] Step S505: Determine whether the server has a fault.
[0103] Wherein, in embodiments of the present invention, when the server has no fault, step S506 is executed; otherwise, step S507 is executed.
[0104] Step S506: Determine the diagnosis rules, diagnosis tags, and fault types when determining the diagnosis target server according to the diagnosis path, so as to generate the current fault diagnosis result of the target server based on the diagnosis rules, diagnosis tags, and fault types.
[0105] Step S507: Adjust the diagnostic path by using the pre-constructed fault path generation model.
[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0107] According to the server fault diagnosis method proposed by the embodiments of the present invention, the actual operation data, configuration information, and historical fault data of the target server can be collected, and then at least one of the actual operation data, configuration information, and historical fault data is used to detect whether the target server has a fault. When it is detected that the target server has a fault, the configuration information, actual operation data, and historical fault data are input into the pre-constructed fault path generation model, and then the diagnostic path when the target server has a fault is output, so as to determine the diagnostic rules, diagnostic labels, and fault types when diagnosing the target server, and then generate the current fault diagnosis result of the target server. Therefore, it can solve the technical problems of single and fixed processes, inability to dynamically adjust according to the real-time operation status and fault conditions of the server, lack of flexibility and pertinence, low diagnostic efficiency, poor diagnostic accuracy. In addition, once the fault cannot be quickly solved, it is extremely easy to cause production interruption, bringing huge economic losses to the enterprise, and achieve the technical effects of dynamically generating accurate and efficient diagnostic paths, shortening the diagnostic time, reducing the energy consumption and maintenance costs of the server, improving the accuracy of fault location, reducing the fault resolution time, and realizing automatic fault diagnosis for different types of servers.
[0108] The embodiments of the present invention also provide a server fault diagnosis device.
[0109] Figure 6 It is a block diagram of the server fault diagnosis device provided according to the embodiments of the present invention.
[0110] As Figure 6 shown, the server fault diagnosis device 10 includes: a detection module 100, an output module 200, and a diagnosis module 300.
[0111] Among them, the detection module 100 is used to collect the actual operation data of the target server and obtain at least one of the configuration information and historical fault data of the target server, so as to detect whether the target server has a fault by using at least one of the actual operation data, configuration information, and historical fault data.
[0112] The output module 200 is used to input the configuration information, actual operation data, and historical fault data into the pre-constructed fault path generation model when it is detected that the target server has a fault, so as to output the diagnostic path when the target server has a fault.
[0113] The diagnosis module 300 is configured to determine the diagnosis rules, diagnosis tags, and fault types when determining the diagnosis target server according to the diagnosis path, so as to generate the current fault diagnosis result of the target server based on the diagnosis rules, diagnosis tags, and fault types.
[0114] Optionally, in an embodiment of the present invention, it further includes: a calculation module, a first determination module, a second determination module, and a construction module.
[0115] Among them, the calculation module is configured to calculate the hardware matching degree between the configuration information and the corresponding diagnosis rule, the operation matching degree between the actual operation data and the corresponding diagnosis rule, and the fault matching degree between the historical fault data and the corresponding diagnosis rule respectively based on the configuration information, actual operation data, and historical fault data before inputting the configuration information, actual operation data, and historical fault data into the pre-constructed fault path generation model.
[0116] The first determination module is configured to determine the weight coefficients corresponding to the hardware matching degree, operation matching degree, and fault matching degree respectively based on the hardware matching degree, operation matching degree, and fault matching degree.
[0117] The second determination module is configured to determine the matching degree of the corresponding diagnosis rule based on the weight coefficient, hardware matching degree, operation matching degree, and fault matching degree.
[0118] The construction module is configured to construct a fault path generation model based on different diagnosis rules and the matching degrees of different diagnosis rules.
[0119] Optionally, in an embodiment of the present invention, it further includes: a first generation module, a first judgment module, a second generation module, and a third generation module.
[0120] Among them, the first generation module is configured to generate a trigger condition for triggering the diagnosis rule in the pre-constructed fault path generation model based on the configuration information, actual operation data, and historical fault data before inputting the configuration information, actual operation data, and historical fault data into the pre-constructed fault path generation model.
[0121] The first judgment module is configured to judge whether the diagnosis rule exists based on the trigger condition.
[0122] The second generation module is configured to screen the diagnosis rule according to a preset condition when the diagnosis rule exists, so as to obtain a diagnosis rule applicable to the occurrence of a fault in the target server.
[0123] The third generation module is configured to establish a corresponding diagnosis rule based on the preset condition, configuration information, actual operation data, and historical fault data when the diagnosis rule does not exist, so as to obtain the established diagnosis rule.
[0124] Optionally, in an embodiment of the present invention, the output module 200 includes: a judgment unit and a first determination unit.
[0125] Among them, the judgment unit is used to judge whether the diagnostic rule meets the preset fault diagnosis condition.
[0126] The first determination unit is used to determine the diagnostic path based on the diagnostic rule when the diagnostic rule meets the preset fault diagnosis condition.
[0127] Optionally, in an embodiment of the present invention, it further includes: a second judgment module and a fourth generation module.
[0128] Among them, the second judgment module is used to judge whether the matching degree of the diagnostic path is less than the preset matching degree value before determining the diagnostic rule, diagnostic label and fault type of the diagnostic target server according to the diagnostic path.
[0129] The fourth generation module is used to regenerate the diagnostic path by using the pre-constructed fault path generation model when the matching degree of the diagnostic path is less than the preset matching degree value until the matching degree of the regenerated diagnostic path is greater than or equal to the preset matching degree value.
[0130] Optionally, in an embodiment of the present invention, the output module 200 includes: a second determination unit, an identification unit and a generation unit.
[0131] Among them, the second determination unit is used to determine the diagnostic label when the target server fails based on at least one of the configuration information, actual operation data and historical fault data.
[0132] The identification unit is used to identify the fault type of the target server based on the diagnostic label.
[0133] The generation unit is used to calculate the diagnostic path by using the pre-constructed fault path generation model based on the diagnostic label and the fault type.
[0134] Optionally, in an embodiment of the present invention, it further includes: a processing module.
[0135] Among them, the processing module is used to process the configuration information, actual operation data and historical fault data respectively before inputting them into the pre-constructed fault path generation model to obtain the configuration information, actual operation data and historical fault data that meet the preset data format conditions.
[0136] Optionally, in an embodiment of the present invention, it further includes: a third determination module and a fifth generation module.
[0137] The third determination module is configured to determine the fault level when the target server fails based on the fault diagnosis result.
[0138] The fifth generation module is configured to generate corresponding fault handling measures based on different fault levels.
[0139] For the description of the features in the corresponding embodiments of the fault diagnosis device of the server, reference can be made to the relevant descriptions in the corresponding embodiments of the fault diagnosis method of the server, which will not be elaborated here one by one.
[0140] According to the fault diagnosis device of the server proposed by the embodiments of the present invention, the actual operation data, configuration information, and historical fault data of the target server can be collected, and then at least one of the actual operation data, configuration information, and historical fault data is used to detect whether the target server fails. In the case where it is detected that the target server fails, the configuration information, actual operation data, and historical fault data are input into a pre-constructed fault path generation model, and then the diagnosis path when the target server fails is output, so as to determine the diagnosis rules, diagnosis tags, and fault types when diagnosing the target server, and then generate the current fault diagnosis result of the target server. Therefore, it can solve the technical problems of single and fixed process, inability to dynamically adjust according to the real-time operation state and fault conditions of the server, lack of flexibility and pertinence, low diagnosis efficiency, poor diagnosis accuracy. In addition, once the fault cannot be quickly solved, it is extremely easy to cause production interruption, bringing huge economic losses to the enterprise, and achieve the technical effects of dynamically generating accurate and efficient diagnosis paths, shortening the diagnosis time, reducing the energy consumption and maintenance cost of the server, improving the accuracy rate of fault location, reducing the fault resolution time, and realizing automatic fault diagnosis for different types of servers.
[0141] An embodiment of the present invention further provides a server, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the embodiments of the above-mentioned fault diagnosis method of the server.
[0142] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the embodiments of the above-mentioned fault diagnosis method of the server when running.
[0143] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (abbreviated as ROM), random access memory (abbreviated as RAM), mobile hard disk, magnetic disk, or optical disc and other various media that can store computer programs.
[0144] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server fault diagnosis method.
[0145] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server fault diagnosis method.
[0146] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0147] The above has introduced in detail a server fault diagnosis method provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A fault diagnosis method for a server, characterized in that, Including the following steps: Collect the actual operation data of the target server, and obtain at least one of the configuration information and historical fault data of the target server, so as to detect whether the target server fails by using the actual operation data and at least one of the configuration information and the historical fault data; In the case of detecting that the target server fails, input the configuration information, the actual operation data and the historical fault data into a pre-constructed fault path generation model, so as to output the diagnostic path when the target server fails; Determine the diagnostic rules, diagnostic labels and fault types for diagnosing the target server according to the diagnostic path, so as to generate the current fault diagnosis result of the target server based on the diagnostic rules, the diagnostic labels and the fault types.
2. The fault diagnosis method of the server according to claim 1, wherein Before inputting the configuration information, the actual operation data and the historical fault data into the pre-constructed fault path generation model, it further includes: Based on the configuration information, the actual operation data and the historical fault data, calculate the hardware matching degree between the configuration information and the corresponding diagnostic rules, the operation matching degree between the actual operation data and the corresponding diagnostic rules, and the fault matching degree between the historical fault data and the corresponding diagnostic rules respectively; Based on the hardware matching degree, the operation matching degree and the fault matching degree, determine the weight coefficients corresponding to the hardware matching degree, the operation matching degree and the fault matching degree respectively; Based on the weight coefficients, the hardware matching degree, the operation matching degree and the fault matching degree, determine the matching degree of the corresponding diagnostic rules; Based on different diagnostic rules and the matching degrees of the different diagnostic rules, construct a fault path generation model.
3. The fault diagnosis method of the server according to claim 2, wherein Before inputting the configuration information, the actual operation data and the historical fault data into the pre-constructed fault path generation model, it further includes: Based on the configuration information, the actual operation data and the historical fault data, generate a trigger condition for triggering the diagnostic rules in the pre-constructed fault path generation model; Based on the trigger condition, judge whether the diagnostic rules exist; If the diagnostic rules exist, screen the diagnostic rules according to preset conditions to obtain the diagnostic rules applicable to the failure of the target server; If the diagnostic rules do not exist, establish corresponding diagnostic rules based on the preset conditions, the configuration information, the actual operation data and the historical fault data to obtain the established diagnostic rules.
4. The fault diagnosis method of the server according to claim 2, wherein The step of inputting the configuration information, the actual operation data and the historical fault data into a pre-constructed fault path generation model to output the diagnostic path when the target server fails includes: Judge whether the diagnostic rules meet the preset fault diagnosis conditions; If the diagnostic rules meet the preset fault diagnosis conditions, determine the diagnostic path based on the diagnostic rules.
5. The fault diagnosis method of the server according to claim 1, characterized in that, Before determining the diagnostic rules, diagnostic labels and fault types for diagnosing the target server according to the diagnostic path, it further includes: Determine whether the matching degree of the diagnosis path is less than a preset matching degree value; If the matching degree of the diagnosis path is less than the preset matching degree value, regenerate the diagnosis path using the pre-constructed fault path generation model until the matching degree of the regenerated diagnosis path is greater than or equal to the preset matching degree value.
6. The fault diagnosis method of the server according to claim 1, characterized in that, The inputting the configuration information, the actual operation data, and the historical fault data into a pre-constructed fault path generation model to output a diagnosis path when the target server fails includes: Based on the configuration information, the actual operation data, and the historical fault data, determine the diagnosis label when the target server fails; Based on the diagnosis label, identify the fault type of the target server; Based on the diagnosis label and the fault type, calculate the diagnosis path using the pre-constructed fault path generation model.
7. The fault diagnosis method of the server according to claim 1, characterized in that Before inputting the configuration information, the actual operation data, and the historical fault data into the pre-constructed fault path generation model, it further includes: Process the configuration information, the actual operation data, and the historical fault data respectively to obtain the configuration information, the actual operation data, and the historical fault data that meet the preset data format conditions.
8. The fault diagnosis method of the server according to claim 1, wherein It further includes: Based on the fault diagnosis result, determine the fault level when the target server fails; Generate corresponding fault handling measures based on different fault levels.
9. A fault diagnosis device for a server, characterized in that, It includes: A detection module, configured to collect the actual operation data of the target server and obtain at least one of the configuration information and the historical fault data of the target server, so as to detect whether the target server fails by using the actual operation data and at least one of the configuration information and the historical fault data; An output module, configured to, when it is detected that the target server fails, input the configuration information, the actual operation data, and the historical fault data into a pre-constructed fault path generation model to output a diagnosis path when the target server fails; A diagnosis module, configured to determine the diagnosis rule, the diagnosis label, and the fault type when diagnosing the target server according to the diagnosis path, so as to generate the current fault diagnosis result of the target server based on the diagnosis rule, the diagnosis label, and the fault type.
10. The fault diagnosis device of the server according to claim 9, characterized in that, It further includes: A calculation module, configured to, before inputting the configuration information, the actual operation data, and the historical fault data into the pre-constructed fault path generation model, calculate the hardware matching degree between the configuration information and the corresponding diagnosis rule, the operation matching degree between the actual operation data and the corresponding diagnosis rule, and the fault matching degree between the historical fault data and the corresponding diagnosis rule respectively based on at least one of the configuration information, the actual operation data, and the historical fault data; A first determination module, configured to determine the weight coefficients corresponding to the hardware matching degree, the operation matching degree, and the fault matching degree respectively based on the hardware matching degree, the operation matching degree, and the fault matching degree; A second determination module, configured to determine the matching degree of the corresponding diagnosis rule based on the weight coefficient, the hardware matching degree, the operation matching degree, and the fault matching degree; A construction module, configured to construct a fault path generation model based on different diagnosis rules and the matching degrees of the different diagnosis rules.
11. The fault diagnosis device of the server according to claim 10, characterized in that, It further includes: A first generation module, configured to generate a triggering condition for triggering a diagnosis rule in the pre-constructed fault path generation model based on the configuration information, the actual operation data, and the historical fault data before inputting the configuration information, the actual operation data, and the historical fault data into the pre-constructed fault path generation model; A first judgment module, configured to judge whether the diagnosis rule exists based on the triggering condition; A second generation module, configured to screen the diagnosis rule according to a preset condition when the diagnosis rule exists, so as to obtain a diagnosis rule applicable to the occurrence of a fault in the target server; A third generation module, configured to establish a corresponding diagnosis rule based on the preset condition, the configuration information, the actual operation data, and the historical fault data when the diagnosis rule does not exist, so as to obtain the established diagnosis rule.
12. The fault diagnosis device of the server according to claim 10, wherein The output module includes: A judgment unit, configured to judge whether the diagnosis rule meets a preset fault diagnosis condition; A first determination unit, configured to determine the diagnosis path based on the diagnosis rule when the diagnosis rule meets the preset fault diagnosis condition.
13. The fault diagnosis device of the server according to claim 9, wherein, It further includes: A second judgment module, configured to judge whether the matching degree of the diagnosis path is less than a preset matching degree value before determining the diagnosis rule, diagnosis label, and fault type for diagnosing the target server according to the diagnosis path; A fourth generation module, configured to regenerate the diagnosis path by using the pre-constructed fault path generation model when the matching degree of the diagnosis path is less than the preset matching degree value until the matching degree of the regenerated diagnosis path is greater than or equal to the preset matching degree value.
14. A server, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the server fault diagnosis method according to any one of claims 1-8.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement the server fault diagnosis method according to any one of claims 1-8.
Citation Information
Patent Citations
Construction method of deep neural network model and fault diagnosis method and system
CN111342997A
Server fault diagnosis method and device and related equipment
CN112988537A
Server fault diagnosis method and device, equipment and medium
CN117093405A
Server fault diagnosis method and device, storage medium and electronic equipment
CN117667479A
Hardware fault analysis system and method
WO2016188175A1
Cited By
Server hardware detection method and electronic equipment
CN120670240A
Server hardware testing methods and electronic equipment
CN120670240B
Fault repairing method of server single board and electronic equipment
CN120723571A
Server hardware link diagnosis method and system
CN121501556A
Server hardware link diagnostic method and system
CN121501556B