Fault diagnosis rule maintenance method, system, terminal and storage medium

By generating and sorting diagnostic rule discrimination values ​​based on log data and common faults, the problem of difficulty in updating fault diagnosis rules in the existing technology is solved, the accuracy and speed of fault diagnosis are improved, and the efficiency of server operation and maintenance is achieved.

CN115185725BActive Publication Date: 2025-06-06INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210784120.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-06-06
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

In the prior art, it is difficult to update the fault diagnosis rules, resulting in inaccurate fault diagnosis.

Method used

Log discrimination values ​​and fault discrimination values ​​are generated based on the degree of matching between diagnostic rules and log data and common faults, the variance of each diagnostic rule is calculated and the weight is assigned. Finally, the diagnostic rule base is sorted according to the combined discrimination values ​​and updated.

Benefits of technology

It improves the accuracy and speed of fault diagnosis, saves manpower and time, and ensures efficient server operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115185725B_ABST
    Figure CN115185725B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of server diagnosis technology, and specifically provides a fault diagnosis rule maintenance method, system, terminal and storage medium, including: generating a log discrimination value of each diagnostic rule based on the matching degree of each diagnostic rule and log data; generating a fault discrimination value of each diagnostic rule based on the commonness of each diagnostic rule and the corresponding fault; calculating the log discrimination value variance and the fault discrimination value variance of each diagnostic rule, and taking the inverse of the log discrimination value variance as the log discrimination weight, and taking the inverse of the fault discrimination value variance as the fault discrimination weight; calculating the combined discrimination value of the diagnostic rule based on the log discrimination value, log discrimination weight, fault discrimination value and fault discrimination weight of the diagnostic rule; sorting all diagnostic rules from high to low according to the combined discrimination value, and putting all diagnostic rules before the specified ranking into the diagnostic rule library. The present invention greatly improves the diagnosis rate, saving manpower and time for server operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of server operation and maintenance, and specifically relates to a fault diagnosis rule maintenance method, system, terminal and storage medium. Background Art

[0002] A large amount of server log data of various types is used in the server operation and maintenance and fault diagnosis process. The server fault information is output through the server diagnostic tool based on the existing diagnostic rules. The diagnostic rules are accumulated by the long-term work of server R&D personnel, testers and operation and maintenance personnel. The rule base has not been updated for a long time, and there are some cases where some rules have few application scenarios or the output faults are not common. Summary of the invention

[0003] In view of the problem in the prior art that it is difficult to update fault diagnosis rules resulting in inaccurate fault diagnosis, the present invention provides a fault diagnosis rule maintenance method, system, terminal and storage medium to solve the above technical problem.

[0004] In a first aspect, the present invention provides a fault diagnosis rule maintenance method, comprising:

[0005] Generate a log discrimination value for each diagnostic rule based on the matching degree between each diagnostic rule and the log data;

[0006] Generate a fault discrimination value for each diagnostic rule based on the commonness of each diagnostic rule and the corresponding fault;

[0007] Calculate the log discriminant value variance and fault discriminant value variance of each diagnostic rule, and use the inverse of the log discriminant value variance as the log discriminant weight, and use the inverse of the fault discriminant value variance as the fault discriminant weight;

[0008] Calculate the combined discrimination value of the diagnosis rule based on the log discrimination value, the log discrimination weight, the fault discrimination value and the fault discrimination weight of the diagnosis rule;

[0009] All diagnostic rules are sorted from high to low according to the combined discrimination value, and all diagnostic rules before the specified ranking are put into the diagnostic rule base.

[0010] Furthermore, based on the matching degree between each diagnostic rule and the log data, a log discrimination value of each diagnostic rule is generated, including:

[0011] Parse the log data into multiple minimal log files, and save the multiple minimal log files as a log data file package;

[0012] Extract the file content of the smallest log file in the log data file package, extract the key information in the diagnosis rule, match the file content and the key information, and count the number of successful matches as the initial log discrimination value of the diagnosis rule for the log data;

[0013] Define log data indicators and analyze log data indicator data, including the difficulty of collection method, server module coverage and collection time;

[0014] A nonparametric model is established by taking the three indicator data as independent variables and the initial log discriminant value of the diagnostic rule as the dependent variable. Three fitting values ​​are obtained based on the nonparametric model, and the mean of the three fitting values ​​is output as the log discriminant value of the diagnostic rule.

[0015] Furthermore, log data indicators are defined and analyzed, including the difficulty of collection method, server module coverage and collection time, including:

[0016] The indicator data analysis rules for the difficulty of the collection method include the difficulty of the system built-in command collection method as 1, the difficulty of the IPMI tool collection method as 2, the difficulty of the RESTful interface collection method as 3, the difficulty of the system third-party tool collection method as 4, and the difficulty of other collection methods as 5;

[0017] The indicator data analysis rules for server module coverage include setting equal initial coverage values ​​according to the total number of server modules involved in the fault diagnosis, and reducing the coverage by one each time a module is detected to be not covered by log data;

[0018] The analysis rule of the indicator data of collection time consumption is to use the collection time of log data as the indicator data of collection time consumption.

[0019] Furthermore, a non-parametric model is established by taking the three indicator data as independent variables and the initial log discriminant value of the diagnostic rule as the dependent variable, three fitting values ​​are obtained based on the non-parametric model, and the mean of the three fitting values ​​is output as the log discriminant value of the diagnostic rule, including:

[0020] Construct a nonparametric model:

[0021]

[0022] Among them, X k represents the kth indicator data, the value of k is 1, 2 or 3, X 1 Indicates the difficulty value of the collection method, X 2 Indicates the server module coverage value, X 3 Indicates the collection time value; X kj represents the kth indicator data of the jth log data; X ki Represents the kth indicator data of the i-th log data; R 1_ L i represents the initial log discrimination value of the first diagnostic rule for the i-th log data;

[0023] The kernel function selects the second-order Gaussian kernel fitting kernel function: Calculate the three fitting values ​​corresponding to the three indicator data;

[0024] The mean of the three fitted values ​​is calculated to obtain the log discriminant value of the first diagnostic rule for the i-th log data.

[0025] Furthermore, the fault discrimination value of each diagnostic rule is generated based on the commonness of each diagnostic rule and the corresponding fault, including:

[0026] Collect fault information output within a specified period, wherein the fault information includes fault level, handling suggestions and fault module;

[0027] The keywords in the extracted processing suggestions are matched with the source log content in the diagnostic rules. The matching degree base is 0, and 1 is added for each successful match. Finally, the initial fault discrimination value of the diagnostic rule for the fault information is obtained.

[0028] Assigning a value to the fault level, and generating a fault level matching degree between the diagnosis rule and the fault information according to the fault level contained in the fault information and the fault level corresponding to the diagnosis rule;

[0029] Generate a fault module corresponding value of the fault information according to the modules involved in the fault information, the modules involved in the diagnosis rule and the correlation between the modules specified by the diagnosis platform;

[0030] A non-parametric model is established with the fault level matching degree and the fault module corresponding value as independent variables and the initial fault discrimination value of the fault information as the dependent variable. Two fault parameter fitting values ​​are obtained based on the non-parametric model, and the mean of the two fault parameter fitting values ​​is output as the fault discrimination value of the diagnostic rule.

[0031] Furthermore, a value is assigned to the fault level, and a fault level matching degree between the diagnosis rule and the fault information is generated according to the fault level included in the fault information and the fault level corresponding to the diagnosis rule, including:

[0032] Set the scores corresponding to different fault levels;

[0033] Assign values ​​to the c fault levels of the fault information to obtain the fault diversity set {d 1 …d c};

[0034] Assign D to the fault level of the diagnostic rule;

[0035] The matching degree between the diagnostic rule and the fault level of the fault information is:

[0036]

[0037] Where d represents the maximum fault level score.

[0038] Furthermore, the log discriminant value variance and the fault discriminant value variance of each diagnostic rule are calculated, and the inverse of the log discriminant value variance is used as the log discriminant weight, and the inverse of the fault discriminant value variance is used as the fault discriminant weight, including:

[0039] Based on the log discriminant values ​​of the diagnosis rule for the multiple log data, the log discriminant value variance of the multiple log discriminant values ​​of the diagnosis rule is calculated, and the inverse of the log discriminant value variance is used as the log discriminant weight;

[0040] Based on the fault discrimination values ​​of the diagnosis rule for the plurality of fault information, the fault discrimination value variance of the plurality of fault discrimination values ​​of the diagnosis rule is calculated, and the inverse of the fault discrimination value variance is used as the fault discrimination weight.

[0041] In a second aspect, the present invention provides a fault diagnosis rule maintenance system, comprising:

[0042] A log discrimination unit, used to generate a log discrimination value of each diagnostic rule based on the matching degree between each diagnostic rule and the log data;

[0043] A fault discrimination unit, used to generate a fault discrimination value for each diagnostic rule based on each diagnostic rule and the commonness of the corresponding fault;

[0044] A weight calculation unit, used to calculate the log discrimination value variance and the fault discrimination value variance of each diagnostic rule, and use the inverse of the log discrimination value variance as the log discrimination weight, and use the inverse of the fault discrimination value variance as the fault discrimination weight;

[0045] A comprehensive discrimination unit, used for calculating a combined discrimination value of a diagnostic rule based on a log discrimination value, a log discrimination weight, a fault discrimination value and a fault discrimination weight of the diagnostic rule;

[0046] The rule screening unit is used to sort all the diagnosis rules from high to low according to the combined discrimination value, and put all the diagnosis rules before the specified ranking into the diagnosis rule library.

[0047] In a third aspect, a terminal is provided, including:

[0048] processor, memory, wherein:

[0049] The memory is used to store computer programs.

[0050] The processor is used to call and run the computer program from the memory, so that the terminal executes the above-mentioned terminal method.

[0051] In a fourth aspect, a computer storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the computer-readable storage medium is run on a computer, the computer executes the methods described in the above aspects.

[0052] The beneficial effects of the present invention are:

[0053] The fault diagnosis rule maintenance method, system, terminal and storage medium provided by the present invention establish a log-rule system and a fault rule system after processing input log data and output fault information, and select diagnostic rules with higher application value from both input and output directions, whose logs used for diagnosis are easier to collect and fault scenarios are more common, thereby providing an effective rule discrimination value for subsequent fault diagnosis. The present invention greatly improves the diagnosis rate and saves manpower and time for server operation and maintenance.

[0054] In addition, the invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0056] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0057] Figure 2 is a schematic block diagram of a system according to an embodiment of the present invention.

[0058] Figure 3 A schematic diagram of the structure of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0059] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0060] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject may be a fault diagnosis rule maintenance system.

[0061] like Figure 1 As shown, the method includes:

[0062] Step 110, generating a log discrimination value of each diagnostic rule based on the matching degree of each diagnostic rule and the log data; Step 120, generating a fault discrimination value of each diagnostic rule based on the commonness of each diagnostic rule and the corresponding fault;

[0063] Step 130, calculating the log discrimination value variance and the fault discrimination value variance of each diagnostic rule, and taking the inverse of the log discrimination value variance as the log discrimination weight, and taking the inverse of the fault discrimination value variance as the fault discrimination weight;

[0064] Step 140, calculating a combined discrimination value of the diagnosis rule based on the log discrimination value, the log discrimination weight, the fault discrimination value and the fault discrimination weight of the diagnosis rule;

[0065] Step 150: sort all the diagnosis rules from high to low according to the combined discrimination value, and put all the diagnosis rules before the specified ranking into the diagnosis rule library.

[0066] To facilitate understanding of the present invention, the fault diagnosis rule maintenance method provided by the present invention is further described below based on the principle of the fault diagnosis rule maintenance method of the present invention and in combination with the process of maintaining the fault diagnosis rules in the embodiment.

[0067] Specifically, the fault diagnosis rule maintenance method includes:

[0068] S1. Generate a log discrimination value for each diagnostic rule based on the matching degree between each diagnostic rule and log data.

[0069] L represents log data, and a total of a log data input recently are collected {L 1 ,L 2 ,…,L a}, log data generally exists in tar or zip format, iteratively decompress the log data until each log data is organized into a single log file as the smallest unit.

[0070] The log discrimination value refers to the degree of application of the input log data to the rules, which is measured by the matching degree between the rules and the log data. Based on the log parsing module, the file content of the log file is extracted, and the key information of the source log content indicator in the diagnostic rule is extracted to match the file content of the log. The matching degree base is 0, and each successful match is increased by 1 to obtain the initial log discrimination value R_L

[0071]

[0072] Where R i _L=[R i _L 1… R i _L a ] represents all the initial log discriminant values ​​of the i-th diagnostic rule, R i _L j Represents the initial log discrimination value of the j-th log data corresponding to the i-th diagnostic rule.

[0073] Define indicators for log data: difficulty of collection method, server module coverage, and collection time. Extract the above indicator data for each log data in a single rule.

[0074] (1) Collection method difficulty X 1 : The main collection methods include system built-in commands, system third-party commands, ipmi tools, restful interfaces, etc. Among them, the collection of system built-in commands is the most convenient and fast, followed by ipmi commands to collect out-of-band logs with a faster response, the restful interface method requires login and has a slow response, the system third-party commands require pre-installed dependencies, and the versions of dependencies of different operating systems vary greatly, which is the most complicated. In summary, the difficulty of the collection method is:

[0075] Table 1- Log collection method weight table

[0076]

[0077] The difficulty level of each log file collection method in a log data is added together to obtain the collection difficulty level of the log data.

[0078] (2) Server module coverage X 2 :The server modules involved in the fault diagnosis include CPU, memory, fan, power supply, PCIE, etc., a total of C modules, then the initial value of coverage is C. Every time it is detected that the log data does not cover a module, the coverage is reduced by one.

[0079] (3) Collection time X 3 : Current log data collection time.

[0080] For example, according to the first rule, after extracting a log data, we get:

[0081]

[0082] The three indicator data are used as independent variables, the log discriminant value of the rule is used as the dependent variable, and a non-parametric model is established. The log discriminant value is fitted according to the Nadaraya-Watson estimation:

[0083]

[0084] Among them, X krepresents the kth indicator data, the value of k is 1, 2 or 3, X 1 Indicates the difficulty value of the collection method, X 2 Indicates the server module coverage value, X 3 Indicates the collection time value; X kj represents the kth indicator data of the jth log data; X ki Represents the kth indicator data of the i-th log data; R 1_ L i It represents the initial log discriminant value of the first diagnostic rule for the i-th log data. The kernel function selects the second-order Gaussian kernel fitting kernel function: Calculate three fitted values ​​corresponding to the three indicator data.

[0085] For example, the log discriminant value of the first indicator and the first diagnostic rule:

[0086]

[0087] The kernel function selects the second-order Gaussian kernel fitting kernel function: Then the mean of the three fitted values ​​is calculated to obtain the log discriminant value of the first diagnostic rule, and the first one is expressed as

[0088]

[0089] S2. Generate a fault discrimination value for each diagnostic rule based on each diagnostic rule and the commonness of the corresponding fault.

[0090] E represents fault information, and a total of b fault information output recently is collected {E 1 ,E 2 ,…,E b}, the fault information includes indicators such as fault level, processing suggestions, and fault module. The keywords in the processing suggestions are extracted and matched with the source log content in the diagnosis rules. The matching degree base is 0, and each successful match is increased by 1. Initial fault discrimination value R_E

[0091]

[0092] Where R i _E=[R i _E 1 … R i _E a ] represents all initial fault discrimination values ​​of the i-th rule, R i _E j Represents the initial fault judgment value of the j-th fault information corresponding to the ith diagnosis rule.

[0093] Assign a value to the fault level, and generate the fault level matching degree Z between the diagnosis rule and the fault information according to the fault level contained in the fault information and the fault level corresponding to the diagnosis rule 1 , methods include:

[0094] (1) Set the scores corresponding to different fault levels. There are d fault levels in total, and the level values ​​are assigned from 1 to d from low to high.

[0095] (2) Assign values ​​to the c fault levels of the fault information to obtain the fault diversity set {d 1 …d c}.

[0096] (3) Assign D to the fault level of the diagnostic rule.

[0097] (4) The matching degree between the diagnostic rule and the fault level of the fault information is:

[0098]

[0099] Where d represents the maximum fault level score.

[0100] Generate the fault module corresponding value Z of the fault information according to the modules involved in the fault information, the modules involved in the diagnosis rules and the correlation between the modules specified by the diagnosis platform 2 , methods include:

[0101] Fault module corresponding value Z 2 : refers to the corresponding value between the fault module in the fault information and the server module in the diagnosis rule. The initial value is 0. Assuming that the diagnosis information contains the module {CPU, memory, PCIE}, and the module in the fault information is {CPU, fan}, then the modules in the fault information are traversed. If the CPU is a completely corresponding module, the corresponding value is increased by one. The fan has no direct corresponding module. According to the correlation degree of each module provided by the diagnosis platform (between 0 and 1), plus the correlation degree between the fan and the memory and the correlation degree between the fan and the PCIE, the corresponding value of the fault module is obtained. If there is a fault module that is of particular concern, the corresponding value of the module and the diagnosis rule module is calculated, and added to the previous corresponding value as the final corresponding value of the fault module.

[0102] A non-parametric model is established with the fault level matching degree and the fault module corresponding value as independent variables and the initial fault discrimination value of the fault information as the dependent variable. Two fault parameter fitting values ​​are obtained based on the non-parametric model, and the mean of the two fault parameter fitting values ​​is output as the fault discrimination value of the diagnostic rule.

[0103] The nonparametric model is:

[0104]

[0105] Among them, Zk represents the kth fault indicator data, Z 1 Indicates the fault level matching degree, Z 2 Indicates the corresponding value of the fault module; Z kj represents the kth fault indicator data of the jth fault information; Z ki represents the kth fault indicator data of the i-th fault information; R n _E i It represents the initial fault discrimination value of the nth diagnostic rule for the ith fault information. The kernel function selects the second-order Gaussian kernel fitting kernel function:

[0106] The nonparametric model can be used to obtain the nth diagnostic rule. and by For example:

[0107]

[0108] Then the fault judgment value of the nth diagnostic rule for the ith fault information is:

[0109]

[0110] Since there are b pieces of fault information, there are b fault discrimination values ​​of the diagnosis rule.

[0111] S3. Calculate the log discrimination value variance and the fault discrimination value variance of each diagnosis rule, and use the inverse of the log discrimination value variance as the log discrimination weight, and use the inverse of the fault discrimination value variance as the fault discrimination weight.

[0112] In step S1 , a log discrimination values ​​of each diagnostic rule are calculated, and in step S2 , b fault discrimination values ​​of each diagnostic rule are calculated.

[0113] Calculate the variance of the log discriminant value of the diagnostic rule (log discriminant value variance), taking the first diagnostic rule as an example:

[0114]

[0115] Calculate the variance of the fault discrimination value of the diagnostic rule (fault discrimination variance), taking the first diagnostic rule as an example:

[0116]

[0117] Then the log discrimination weight is:

[0118]

[0119] The fault discrimination weight is:

[0120]

[0121] S4. Calculate the combined discrimination value of the diagnosis rule based on the log discrimination value, log discrimination weight, fault discrimination value and fault discrimination weight of the diagnosis rule.

[0122] In step S3, the log discrimination weight and fault discrimination weight of the first diagnostic rule are known, and the combined discrimination value of the first diagnostic rule is:

[0123]

[0124] The calculation method of the combined discriminant value of other diagnostic rules is the same as that of the first diagnostic rule, and only the weight variable, log discriminant value variable and fault discriminant value variable are replaced accordingly.

[0125] S5. Sort all the diagnosis rules from high to low according to the combined discrimination value, and put all the diagnosis rules before the specified ranking into the diagnosis rule library.

[0126] All diagnostic rules are sorted from high to low according to the combined discrimination value, and all diagnostic rules before the specified ranking are put into the diagnostic rule base.

[0127] like Figure 2 As shown, the system 200 includes:

[0128] The log identification unit 210 is used to generate a log identification value of each diagnostic rule based on the matching degree between each diagnostic rule and the log data;

[0129] A fault determination unit 220, for generating a fault determination value for each diagnostic rule based on each diagnostic rule and the commonness of the corresponding fault;

[0130] The weight calculation unit 230 is used to calculate the log discrimination value variance and the fault discrimination value variance of each diagnostic rule, and use the inverse of the log discrimination value variance as the log discrimination weight, and use the inverse of the fault discrimination value variance as the fault discrimination weight;

[0131] A comprehensive discrimination unit 240, for calculating a combined discrimination value of a diagnostic rule based on a log discrimination value, a log discrimination weight, a fault discrimination value and a fault discrimination weight of the diagnostic rule;

[0132] The rule screening unit 250 is used to sort all the diagnosis rules from high to low according to the combined discrimination value, and put all the diagnosis rules before the specified ranking into the diagnosis rule library.

[0133] Figure 3 The present invention provides a schematic diagram of the structure of a terminal 300 provided in an embodiment of the present invention. The terminal 300 can be used to execute the fault diagnosis rule maintenance method provided in an embodiment of the present invention.

[0134] The terminal 300 may include: a processor 310, a memory 320 and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention, and it may be a bus structure or a star structure, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0135] The memory 320 can be used to store the execution instructions of the processor 310, and the memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can perform some or all of the steps in the following method embodiments.

[0136] The processor 310 is the control center of the storage terminal, and uses various interfaces and lines to connect various parts of the entire electronic terminal. It runs or executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic terminal and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of a plurality of packaged ICs with the same or different functions. For example, the processor 310 can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0137] The communication unit 330 is used to establish a communication channel so that the storage terminal can communicate with other terminals, receive user data sent by other terminals or send user data to other terminals.

[0138] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0139] Therefore, the present invention establishes a log-rule system and a fault rule system after processing the input log data and the output fault information, and screens the diagnostic logs used for diagnosis from both the input and output directions, which are easier to collect and have more common fault scenarios and higher application value diagnostic rules, thereby providing an effective rule discrimination value for subsequent fault diagnosis. The present invention greatly improves the diagnostic rate, saves manpower and time for server operation and maintenance, and the technical effects that can be achieved by this embodiment can be found in the description above, which will not be repeated here.

[0140] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention, in essence or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes, including several instructions for enabling a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention.

[0141] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the terminal embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

[0142] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.

[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0145] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of these shall be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A fault diagnosis rule maintenance method, It is characterized in that include: Generate a log discrimination value for each diagnostic rule based on the matching degree between each diagnostic rule and the log data; Generate a fault discrimination value for each diagnostic rule based on the commonness of each diagnostic rule and the corresponding fault; Calculate the log discriminant value variance and fault discriminant value variance of each diagnostic rule, and use the inverse of the log discriminant value variance as the log discriminant weight, and use the inverse of the fault discriminant value variance as the fault discriminant weight; Calculate the combined discrimination value of the diagnosis rule based on the log discrimination value, the log discrimination weight, the fault discrimination value and the fault discrimination weight of the diagnosis rule; Sort all diagnostic rules from high to low according to the combined discrimination value, and put all diagnostic rules before the specified ranking into the diagnostic rule library; The log discrimination value of each diagnostic rule is generated based on the matching degree between each diagnostic rule and the log data, including: Parse the log data into multiple minimal log files, and save the multiple minimal log files as a log data file package; Extract the file content of the smallest log file in the log data file package, extract the key information in the diagnosis rule, match the file content and the key information, and count the number of successful matches as the initial log discrimination value of the diagnosis rule for the log data; Define log data indicators and analyze log data indicator data, including the difficulty of collection method, server module coverage and collection time; A nonparametric model is established by taking the three indicator data as independent variables and the initial log discriminant value of the diagnostic rule as the dependent variable, three fitting values ​​are obtained based on the nonparametric model, and the mean of the three fitting values ​​is output as the log discriminant value of the diagnostic rule; The fault discrimination value of each diagnostic rule is generated based on the commonness of each diagnostic rule and the corresponding fault, including: Collect fault information output within a specified period, wherein the fault information includes fault level, handling suggestions and fault module; The keywords in the extracted processing suggestions are matched with the source log content in the diagnostic rules. The matching degree base is 0, and 1 is added for each successful match. Finally, the initial fault discrimination value of the diagnostic rule for the fault information is obtained. Assigning a value to the fault level, and generating a fault level matching degree between the diagnosis rule and the fault information according to the fault level contained in the fault information and the fault level corresponding to the diagnosis rule; Generate a fault module corresponding value of the fault information according to the modules involved in the fault information, the modules involved in the diagnosis rule and the correlation between the modules specified by the diagnosis platform; A non-parametric model is established with the fault level matching degree and the fault module corresponding value as independent variables and the initial fault discrimination value of the fault information as the dependent variable. Two fault parameter fitting values ​​are obtained based on the non-parametric model, and the mean of the two fault parameter fitting values ​​is output as the fault discrimination value of the diagnostic rule.

2. The method according to claim 1, It is characterized in that Define log data indicators and analyze log data indicators, including the difficulty of collection method, server module coverage and collection time, including: The indicator data analysis rules for the difficulty of the collection method include the difficulty of the system built-in command collection method as 1, the difficulty of the IPMI tool collection method as 2, the difficulty of the RESTful interface collection method as 3, the difficulty of the system third-party tool collection method as 4, and the difficulty of other collection methods as 5; The indicator data analysis rules for server module coverage include setting equal initial coverage values ​​according to the total number of server modules involved in the fault diagnosis, and reducing the coverage by one each time a module is detected to be not covered by log data; The analysis rule of the indicator data of collection time consumption is to use the collection time of log data as the indicator data of collection time consumption.

3. The method according to claim 1, It is characterized in that The three indicator data are respectively used as independent variables, and the initial log discriminant value of the diagnostic rule is used as the dependent variable to establish a non-parametric model, three fitting values ​​are obtained based on the non-parametric model, and the mean of the three fitting values ​​is output as the log discriminant value of the diagnostic rule, including: Construct a nonparametric model: Among them, X k represents the kth indicator data, the value of k is any one of 1, 2 or 3, X 1 Indicates the difficulty value of the collection method, X 2 Indicates the server module coverage value, X 3 Indicates the collection time value; X kj represents the kth indicator data of the jth log data; X ki Represents the kth indicator data of the i-th log data; R n_ L i It represents the initial log discrimination value of the nth diagnostic rule for the i-th log data; The kernel function selects the second-order Gaussian kernel fitting kernel function: , calculate the three fitting values ​​corresponding to the three indicator data; The mean of the three fitted values ​​is calculated to obtain the log discriminant value of the first diagnostic rule for the i-th log data.

4. The method according to claim 1, It is characterized in that Assign a value to the fault level, and generate a fault level matching degree between the diagnosis rule and the fault information according to the fault level contained in the fault information and the fault level corresponding to the diagnosis rule, including: Set the scores corresponding to different fault levels; Assign values ​​to the c fault levels of the fault information to obtain the fault diversity set of the fault information ; Assign D to the fault level of the diagnostic rule; Then the fault level matching degree between the diagnostic rule and the fault information is: Where d represents the maximum fault level score.

5. The method according to claim 1, It is characterized in that The log discriminant value variance and fault discriminant value variance of each diagnostic rule are calculated, and the inverse of the log discriminant value variance is used as the log discriminant weight, and the inverse of the fault discriminant value variance is used as the fault discriminant weight, including: Based on the log discriminant values ​​of the diagnosis rule for the multiple log data, the log discriminant value variance of the multiple log discriminant values ​​of the diagnosis rule is calculated, and the inverse of the log discriminant value variance is used as the log discriminant weight; Based on the fault discrimination values ​​of the diagnosis rule for the plurality of fault information, the fault discrimination value variance of the plurality of fault discrimination values ​​of the diagnosis rule is calculated, and the inverse of the fault discrimination value variance is used as the fault discrimination weight.

6. A fault diagnosis rule maintenance system, It is characterized in that include: A log discrimination unit, used to generate a log discrimination value of each diagnostic rule based on the matching degree between each diagnostic rule and the log data; A fault discrimination unit, used to generate a fault discrimination value for each diagnostic rule based on each diagnostic rule and the commonness of the corresponding fault; A weight calculation unit, used to calculate the log discrimination value variance and the fault discrimination value variance of each diagnostic rule, and use the inverse of the log discrimination value variance as the log discrimination weight, and use the inverse of the fault discrimination value variance as the fault discrimination weight; A comprehensive discrimination unit, used for calculating a combined discrimination value of a diagnosis rule based on a log discrimination value, a log discrimination weight, a fault discrimination value and a fault discrimination weight of the diagnosis rule; A rule screening unit is used to sort all the diagnosis rules from high to low according to the combined discrimination value, and put all the diagnosis rules before the specified ranking into the diagnosis rule library; The log discrimination value of each diagnostic rule is generated based on the matching degree between each diagnostic rule and the log data, including: Parse the log data into multiple minimal log files, and save the multiple minimal log files as a log data file package; Extract the file content of the smallest log file in the log data file package, extract the key information in the diagnosis rule, match the file content and the key information, and count the number of successful matches as the initial log discrimination value of the diagnosis rule for the log data; Define log data indicators and analyze log data indicator data, including the difficulty of collection method, server module coverage and collection time; A nonparametric model is established by taking the three indicator data as independent variables and the initial log discriminant value of the diagnostic rule as the dependent variable, three fitting values ​​are obtained based on the nonparametric model, and the mean of the three fitting values ​​is output as the log discriminant value of the diagnostic rule; The fault discrimination value of each diagnostic rule is generated based on the commonness of each diagnostic rule and the corresponding fault, including: Collect fault information output within a specified period, wherein the fault information includes fault level, handling suggestions and fault module; The keywords in the extracted processing suggestions are matched with the source log content in the diagnostic rules. The matching degree base is 0, and 1 is added for each successful match. Finally, the initial fault discrimination value of the diagnostic rule for the fault information is obtained. Assigning a value to the fault level, and generating a fault level matching degree between the diagnosis rule and the fault information according to the fault level contained in the fault information and the fault level corresponding to the diagnosis rule; Generate a fault module corresponding value of the fault information according to the modules involved in the fault information, the modules involved in the diagnosis rule and the correlation between the modules specified by the diagnosis platform; A non-parametric model is established with the fault level matching degree and the fault module corresponding value as independent variables and the initial fault discrimination value of the fault information as the dependent variable. Two fault parameter fitting values ​​are obtained based on the non-parametric model, and the mean of the two fault parameter fitting values ​​is output as the fault discrimination value of the diagnostic rule.

7. A terminal, It is characterized in that include: processor; A memory for storing execution instructions of the processor; The processor is configured to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Fault diagnosis method and system based on server log

    CN111737035A

  • Server fault diagnosis rule screening method based on GRA

    CN113791924A