A Method, System, Device, and Storage Medium for Batch Analysis of Server Failures
By performing feature extraction and grouping of server fault logs, we determine the fault representative matrix and its type, solving the problem of low fault analysis efficiency in server batch failures, and achieving fast and efficient fault handling.
Patent Information
- Application Number
- CN202210764720.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-01
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-01
AI Technical Summary
When the server fails in batches, it is difficult for the existing technology to conduct fault analysis efficiently, resulting in low fault processing efficiency and long time-consuming, and the failure to deal with the faulty server in time.
By extracting the fault logs of N fault servers, a fault feature matrix is obtained and divided into K packets. According to the fault feature matrix with a high degree of similarity, it is placed into the same packet, and the fault representative matrix and its fault type of each packet are determined.
It realizes efficient fault analysis of server batch failures, quickly determines the fault types of each faulty server, and improves the efficiency of fault handling.
Smart Images

Figure CN115129501B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault analysis, and particularly to a method, system, device and storage medium for batch analysis of server faults. Background Art
[0002] As a computer with relatively high hardware configuration and relatively fixed application scenarios, a server has higher stability compared to an ordinary personal computer and can operate reliably without downtime for a long time. However, hardware has a lifespan loss, software also has vulnerabilities and BUGs, and during the operation of the server, it is in a high-pressure load state for most of the time to release its computing performance. In addition, the operation of the hardware generates heat, and the accumulated heat causes the temperature to rise, and too high a temperature is also likely to damage the hardware. Therefore, even though the server has good stability, there is still a certain risk of failure during long-term operation. In actual application scenarios, server downtime caused by failures often causes huge economic losses. In server operation and maintenance, fault analysis is a very important link.
[0003] Currently, the logic for server fault analysis is to first obtain the fault logs of the faulty servers, then extract the corresponding fault characteristics from them, and then calculate the fault factors to compare with the known fault types, that is, to match the corresponding fault types. Although this method of making comparisons one by one is feasible, in actual applications, if there are a large number of server faults, the efficiency of fault analysis is relatively low and it takes a very long time, which is not conducive to timely handling of the faults of the faulty servers. For example, in the case of extremely high temperature weather, the temperature in the computer room may be too high, which may cause large-scale server downtime. Another example is that the sudden power outage in the server computer room causes a large number of servers to be overloaded with voltage, resulting in hardware failures.
[0004] In summary, for the batch faults of servers, how to efficiently perform fault analysis and quickly determine the respective fault types of each faulty server is a technical problem that those skilled in the art urgently need to solve at present. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system, device and storage medium for batch analysis of server faults, so as to efficiently perform fault analysis for the batch faults of servers and quickly determine the respective fault types of each faulty server.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] A method for batch analysis of server faults includes:
[0008] Feature extraction is respectively performed on the fault logs of each of the N faulty servers to obtain N fault feature matrices of each of the N faulty servers; N is a positive integer not less than 2, and K is a positive integer and K ≤ N;
[0009] The N fault feature matrices are divided into K groups; wherein, when grouping, the fault feature matrices with high similarity are placed in the same group as the grouping principle;
[0010] For any one of the K groups, based on each of the fault feature matrices partitioned into the group, a fault representative matrix for reflecting the fault type of the group is determined;
[0011] For any one of the K groups, the fault type of the fault representative matrix of the group is determined, and the fault type is used as the fault type of the faulty servers corresponding to each of the fault feature matrices in the group.
[0012] Preferably, the dividing the N fault feature matrices into K groups includes:
[0013] From the N fault feature matrices, any K fault feature matrices are arbitrarily selected as K initial mass points;
[0014] For any one of the N - K fault feature matrices that are not mass points currently, the fault feature matrix is divided into the group where the mass point closest to itself is located to obtain a total of K groups;
[0015] For any one of the current K groups, the mean value of each of the fault feature matrices in the group is determined as the current verification mass point of the group;
[0016] For any one of the current K groups, when the distance between the current mass point of the group and the current verification mass point of the group is greater than the first threshold, the current verification mass point of the group is used as the new mass point after the update of the group, and the operation of dividing any one of the N - K fault feature matrices that are not mass points currently into the group where the mass point closest to itself is located to obtain a total of K groups is returned to be executed for the update of the grouping until when the distance between the current mass point of any group and the current verification mass point of the group is not greater than the first threshold, the current K groups are used as the final groups.
[0017] Preferably, when dividing the fault feature matrix into the group where the mass point closest to itself is located, the Euclidean distance calculation formula is used to determine the distance between the fault feature matrix and each mass point.
[0018] Preferably, after selecting the K initial mass points, it further includes:
[0019] Timing is performed by a timer, and when the timing duration reaches the first duration and the K final groups have not been determined yet, the value of K is adjusted and the operation of arbitrarily selecting K fault feature matrices from the N fault feature matrices as K initial mass points is re-executed.
[0020] Preferably, for any one of the K groups, based on each fault feature matrix divided into the group, determining a fault representative matrix for reflecting the fault type of the group includes:
[0021] For any one of the K groups, the mean value of each fault feature matrix divided into the group is used as the determined fault representative matrix for reflecting the fault type of the group.
[0022] Preferably, the feature extraction of the fault logs of each of the N fault servers respectively includes:
[0023] The feature extraction of the fault logs of each of the N fault servers is respectively performed by means of outlier capture of the fault logs and / or by means of abnormal information frequency statistics of the fault logs.
[0024] Preferably, for any one of the K groups, determining the fault type of the fault representative matrix of the group includes:
[0025] Construct a fault sample set with known fault types and divide it into a training set and a validation set;
[0026] Under the condition that the set value Q of the number of neighbors is different, based on the KNN algorithm, use the training set to judge the fault types of the samples in the validation set, and determine the value of the number of neighbors Q that makes the fault type judgment accuracy rate the highest; Q is a positive integer;
[0027] Based on the determined value Q of the number of neighbors, based on the KNN algorithm, use the constructed fault sample set with known fault types to judge the fault types of each fault representative matrix, and obtain the fault types of each fault representative matrix.
[0028] A batch analysis system for server faults includes:
[0029] A fault feature matrix extraction module for respectively performing feature extraction on the fault logs of each of the N fault servers to obtain N fault feature matrices of each of the N fault servers; N is a positive integer not less than 2, and K is a positive integer and K ≤ N;
[0030] A grouping module, configured to divide N fault feature matrices into K groups; wherein, when grouping, the fault feature matrices with high similarity are placed in the same group as the grouping principle;
[0031] A fault representative matrix determination module, configured to, for any one of the K groups, determine a fault representative matrix for reflecting the fault type of the group based on each of the fault feature matrices divided into the group;
[0032] A grouped fault type determination module, configured to, for any one of the K groups, determine the fault type of the fault representative matrix of the group, and use the fault type as the fault type of the fault servers corresponding to each of the fault feature matrices in the group.
[0033] A batch analysis device for server faults, comprising:
[0034] A memory, configured to store a computer program;
[0035] A processor, configured to execute the computer program to implement the steps of the batch analysis method for server faults as described in any one of the above.
[0036] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the batch analysis method for server faults as described in any one of the above are implemented.
[0037] Applying the technical solution provided by the embodiments of the present invention, considering that when there are batch failures in the server, some of the failed servers have commonalities, that is, the failure types of some servers may be similar. Therefore, in this application, the failure logs of each of the N failed servers are first subjected to feature extraction to obtain the failure feature matrices of each of the N failed servers. After that, according to the grouping principle that the failure feature matrices with high similarity are placed in the same group, the N failure feature matrices are divided into K groups. That is to say, for any one of the K groups, the same failure type can be used to represent each of the failed servers corresponding to this group. Therefore, for any one of the K groups, based on the respective failure feature matrices divided into this group, this application can determine a failure representative matrix for reflecting the failure type of this group. After subsequently determining the failure types of the respective failure representative matrices of each group, the failure type of the failure representative matrix can be used as the failure type of each of the failed servers corresponding to the respective failure feature matrices in the corresponding group. It can be seen that in the solution of this application, it is not necessary to individually determine the failure types of each failure feature matrix, but only to determine the failure types of the K failure representative matrices, and then these can be used to represent the failure types of the K groups. Therefore, for the batch failures of the server, the solution of this application realizes efficient failure analysis, that is, it can quickly determine the failure types of each of the failed servers. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0039] Figure 1 It is a flowchart of the implementation of a method for batch analysis of server failures in the present invention;
[0040] Figure 2 It is a schematic structural diagram of a system for batch analysis of server failures in the present invention;
[0041] Figure 3 It is a schematic structural diagram of a device for batch analysis of server failures in the present invention. Detailed Embodiments
[0042] The core of the present invention is to provide a method for batch analysis of server failures, which realizes efficient failure analysis for batch failures of the server, that is, it can quickly determine the failure types of each of the failed servers.
[0043] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0044] Please refer to Figure 1 , Figure 1 which is the implementation flowchart of a method for batch analysis of server failures in the present invention. The method for batch analysis of server failures may include the following steps:
[0045] Step S101: Extract features from the respective fault logs of N faulty servers to obtain the respective fault feature matrices of the N faulty servers. N is a positive integer not less than 2, and K is a positive integer and K ≤ N.
[0046] Specifically, the fault log is text data. When performing subsequent fault analysis, that is, when performing operations such as grouping and determining the fault type of the fault representative matrix in the future, the program algorithm is not convenient for processing unstructured data. Therefore, it is necessary to perform data modeling on the existing text data according to certain rules for subsequent data analysis and processing. In the solution of the present application, the specific data modeling method adopted is to extract features from the respective fault logs of N faulty servers, so as to obtain the respective fault feature matrices of the N faulty servers.
[0047] It can be seen that for the fault log of each faulty server, the fault feature matrix of the faulty server can be obtained through feature extraction.
[0048] When performing feature extraction, there are various specific methods, as long as the extracted fault feature matrix can effectively reflect the content of the fault log.
[0049] For example, in a specific embodiment of the present invention, the step of extracting features from the respective fault logs of N faulty servers described in step S101 may specifically include:
[0050] Extract features from the respective fault logs of N faulty servers by means of capturing outliers in the fault log and / or by means of statistically counting the frequency of abnormal information in the fault log.
[0051] Outlier capture is a relatively convenient feature extraction method. For example, for the fans of a server, although the fan speed will automatically adjust with the current temperature of the server, when the load is normal, the fan speed mostly remains at the level of 30% - 70%, roughly around 5000 RPM (Revolutions Per Minute). If the fan is in a full-load or extremely low-load situation for a long time, it indicates that there may be a malfunction in the fan control module or temperature sensor. Therefore, for the abnormal value of the fan speed in the fault log, fault feature extraction can be performed. And considering the convenience of subsequent calculations, normalization operations can usually be carried out when performing fault feature extraction, that is, converting the fault features into the interval of [0, 1].
[0052] Abnormal information frequency statistics can perform feature extraction on the fault information presented in non-numerical form in the fault log. For example, for server network card failures, the abnormal information "unable to obtain network information" may appear in the fault log and may appear multiple times in the collected fault logs. Therefore, the frequency of its occurrence can be statistically analyzed, and then according to the rule that the higher the occurrence frequency, the higher the probability of failure, this feature can be transposed into the interval of [0, 1] to achieve feature extraction.
[0053] It can be seen that in this implementation method, through the method of outlier capture of the fault log and / or through the method of abnormal information frequency statistics of the fault log, feature extraction of the fault log can be effectively carried out. And in practical applications, a more commonly used solution is to simultaneously use these two methods for feature extraction to ensure the accuracy and comprehensiveness of feature extraction.
[0054] For the fault log of each faulty server, the fault feature matrix of the faulty server can be obtained through feature extraction. The fault feature matrix can be represented as a one-dimensional matrix, that is, the fault feature matrix x can be represented as the matrix [x 1 , x 2 , x 3 ... x n , where n needs to be set in advance, representing the number of columns of the fault feature matrix, that is, the number of types of fault feature items.
[0055] It can be understood that for the fault logs of different servers, the structures of the extracted fault feature matrices are the same, that is, they can all be represented in the form of the above matrix [x 1 , x 2 , x 3 ... x n . And because the specific contents of the fault logs of different servers are different, the values of each element in different fault feature matrices can be different.
[0056] For example, in a specific scenario, only a fan failure of server 1 is reflected in the fault log, and there are no other faults. Then, in the fault feature matrix corresponding to server 1, for example, x 1 is the value after fault feature extraction, while the remaining x 2 to x n can all be taken as 0.
[0057] Step S102: Divide the N fault feature matrices into K groups; among them, when grouping, the fault feature matrices with high similarity are placed in the same group as the grouping principle.
[0058] In the solution of this application, only the fault types of the K fault representative matrices need to be determined subsequently to represent the fault types of the K groups respectively. Therefore, it is required that the determined fault representative matrices can effectively reflect the fault types of the corresponding groups, and for each group, the fault feature matrices in this group should have similarity, which means that when grouping, the fault servers with the same fault type need to be grouped together. Therefore, in the solution of this application, when grouping, the fault feature matrices with high similarity are placed in the same group as the grouping principle.
[0059] The specific grouping method can be set according to the actual situation, as long as it can, as described above, the set grouping method can place the fault feature matrices with high similarity in the same group.
[0060] For example, in a specific embodiment of the present invention, step S102 may specifically include the following steps:
[0061] Step 1: Arbitrarily select K fault feature matrices from the N fault feature matrices as K initial mass points;
[0062] Step 2: For any one of the N - K fault feature matrices that are not mass points currently, divide the fault feature matrix into the group where the mass point closest to itself is located to obtain a total of K groups;
[0063] Step 3: For any one of the current K groups, determine the mean value of the fault feature matrices in the group as the current verification mass point of the group;
[0064] Step 4: For any one of the current K groups, when the distance between the current mass point of the group and the current verification mass point of the group is greater than the first threshold, use the current verification mass point of the group as the new mass point after group update, and return to execute the operation of step 2 to update the group until when the distance between the current mass point of any group and the current verification mass point of the group is not greater than the first threshold, the current K groups are used as the final groups.
[0065] In the solution of this application, the value of K can be preset and adjusted as needed. In practical applications, each time a batch analysis of server failures is performed, the staff can, based on practical experience, such as a preliminary review of the failure logs, estimate a more appropriate value of K for this batch analysis of server failures.
[0066] For example, if there are 100 faulty servers to be analyzed for failures this time, 100 fault feature matrices can be determined. For example, if K is set to 7, then 7 fault feature matrices can be arbitrarily selected, such as randomly selecting 7 fault feature matrices. These K fault feature matrices can be called K initial mass points.
[0067] Since 7 out of 100 fault feature matrices are selected as mass points, there are 93 remaining fault feature matrices. The operation in step two performed later is to divide each of these 93 fault feature matrices into the corresponding mass point. That is to say, through the operation in step two, including 7 mass points, 100 fault feature matrices can be divided into 7 groups.
[0068] For any one of the 93 fault feature matrices, it is determined which group where the mass point is located to divide this fault feature matrix according to the distance between this fault feature matrix and the 7 mass points. This is because the closer the distance between the fault feature matrix and the mass point, the higher the similarity between the two, that is, the more likely the servers corresponding to the two have the same type of failure.
[0069] When step three is executed for the first time, the mass points of each of the 7 groups are the 7 initial mass points randomly selected in step one. For any one of the current K groups, this application will determine the mean value of the fault feature matrices of this group as the current verification mass point of this group, and then determine whether the distance between the current mass point of this group and the current verification mass point of this group is greater than the first threshold. If it is greater than the first threshold, it means that the current mass point of this group and the current verification mass point of this group have a large gap, and this group is not very reasonable, that is, the mass point is not selected very reasonably, and the operation in step two can be returned for execution.
[0070] And when returning to execute the operation in step two, for the group that fails the verification, the mass point of this group will be updated. The specific update method is to use the current verification mass point of this group as the new mass point after the update of this group. For example, in the above example, when step three is executed for the first time, 5 out of 7 groups pass the verification, and for 2 groups, the distance between the mass point and the corresponding verification mass point is greater than the first threshold. Then the verification mass points of these 2 groups can be used as the new mass points after the update of these 2 groups respectively, and the operation in step two can be returned for execution.
[0071] It should be noted that when returning to perform the operation of step two, even if 5 out of the 7 groups in the above example pass the verification, the grouping still needs to be re-divided. That is, when re-performing step two, the grouping determined in the previous execution of step two can be discarded at this time. At this time, for the current N-K non-particle fault feature matrices with respect to the current 7 mass points, grouping needs to be performed again according to the distance, and thus a total of K groups can be obtained again. Then, the operations of step three and step four are triggered again. After looping one or more times like this, when the grouping is completed at a certain time, for each group, the distance between the current mass point of the group and the current verification mass point of the group is not greater than the first threshold, and then the current K groups can be used as the final grouping.
[0072] In the above embodiment, considering that the distance between different fault feature matrices can reflect the similarity degree, a grouping method is designed to meet the requirement of "grouping the fault feature matrices with high similarity degree into the same group as the grouping principle", and it is also relatively convenient to execute, and the finally determined grouping is also relatively appropriate. That is, the similarity degrees of the fault feature matrices in the same group are relatively high, which is beneficial to improving the accuracy of batch analysis of server faults in the solution of this application.
[0073] In addition, it should be noted that in the above embodiment, the mean value of each fault feature matrix to be grouped needs to be determined as the current verification mass point of the grouping. For example, if there are a total of 3 fault feature matrices including the mass point in a group, which are respectively represented as [x 11 ,x 12 ,x 13 , [x 21 ,x 22 ,x 23 and [x 31 ,x 32 ,x 33 , then the verification mass point can be represented as [(x 11 +x 21 +x 31 ) / 3, (x 12 +x 22 +x 32 ) / 3, (x 13 +x 23 +x 33 ) / 3].
[0074] During the execution of the above embodiment, it is necessary to determine the distance between the fault feature matrix and the mass point, that is, to determine the distance between the fault feature matrices. As described above, considering that the fault feature matrix x can be represented as a one-dimensional matrix [x 1 ,x 2 ,x 3...x n , so the distance between different fault feature matrices can be determined by the Euclidean distance. That is, in a specific embodiment of the present invention, when dividing the fault feature matrix into the group where the mass point with the closest distance to itself is located, the Euclidean distance calculation formula can be used to determine the distance between the fault feature matrix and each mass point.
[0075] For example, a certain fault feature matrix is represented as [x 11 , x 12 , x 13 , and a certain mass point is represented as [x 21 , x 22 , x 23 , then the Euclidean distance between the two
[0076] In a specific embodiment of the present invention, after selecting K initial mass points in the above step one, it may further include:
[0077] Timing is performed by a timer, and when the timing duration reaches the first duration and the K final groups have not been determined yet, the value of K is adjusted and the operation of arbitrarily selecting K fault feature matrices from N fault feature matrices as K initial mass points is re-executed.
[0078] As described above, the mass points and groups can be updated through several loops, and finally the K final groups can be determined. However, in some cases, after several loops, the condition for exiting the loop, that is, "the distance between any current mass point of any group and the current verification mass point of the group is not greater than the first threshold", may not be satisfied. This may be caused by an unreasonable setting of K. For example, during a batch failure of servers, there are 30 types of faults for these faulty servers, and the set K = 5, which is much smaller than the actual number of fault types 30. As a result, after several loops, a suitable grouping still cannot be obtained.
[0079] In response to this, in this embodiment, timing is performed by a timer. When the timing duration reaches the first duration, the value of K is adjusted and the operation of the above step one is re-executed. There are various ways to adjust the value of K. For example, a common solution is to increase the value of K according to a set step size.
[0080] Step S103: For any one of the K groups, based on the fault feature matrices divided into the group, determine a fault representative matrix for reflecting the fault type of the group.
[0081] After determining the K groups, the fault representative matrices of these K groups can be determined. The fault representative matrix is a matrix used to reflect the fault types of the groups. There are various specific selection methods. For example, a more appropriate and convenient method is to take the average value to determine. That is, in a specific embodiment of the present invention, step S103 may specifically include:
[0082] For any one of the K groups, the average value of the respective fault feature matrices divided into the group is used as the determined fault representative matrix for reflecting the fault type of the group.
[0083] In this embodiment, for any one of the K groups, the average value of the respective fault feature matrices divided into the group is used as the fault representative matrix of the group. Determining the average value of the respective fault feature matrices divided into the same group is the same as the calculation method for determining the verification mass points of the group in the above text, which is very simple in calculation, and the selected fault representative matrix is also very reasonable.
[0084] Step S104: For any one of the K groups, determine the fault type of the fault representative matrix of the group, and use the fault type as the fault type of the respective fault servers corresponding to the respective fault feature matrices in the group.
[0085] Since the fault representative matrix can reflect the fault type of the group, for the K groups, in this application, it is not necessary to determine the fault types of the respective fault feature matrices in the group one by one. Instead, only the fault types of the K fault representative matrices need to be determined, which can represent the fault types of the K groups. That is, the fault types of the determined K fault representative matrices are used as the fault types of the respective fault servers corresponding to the respective fault feature matrices in the corresponding groups.
[0086] For example, if there are 10 fault feature matrices in a group and the determined fault type of the fault representative matrix of the group is a fan fault, it can be considered that the respective fault servers corresponding to these 10 fault feature matrices are all fan faults.
[0087] When determining the fault type of the fault representative matrix of the group, there are various specific methods. For example, by calculating the fault factor and comparing it with the known fault types to match the corresponding fault type. However, in a specific embodiment of the present invention, considering that the accuracy of such a method is not high enough, when performing "for any one of the K groups, determine the fault type of the fault representative matrix of the group" in step S104, it may specifically include the following steps:
[0088] Construct a fault sample set of known fault types and divide it into a training set and a validation set;
[0089] When the set value Q of the number of neighbors is different, based on the KNN (K Nearest Neighbors) algorithm, the training set and the samples in the validation set are used to determine the fault type, and the value of the number of neighbors Q that maximizes the accuracy of the fault type judgment is determined.
[0090] Based on the determined value Q of the number of neighbors, based on the KNN algorithm, using the constructed fault sample set with known fault types, the fault types of each fault representative matrix are judged, and the fault types of each fault representative matrix are obtained.
[0091] It should be noted that here, in order to distinguish the grouping number K in the above text, the value of the number of neighbors to be set in the KNN algorithm is represented by Q, and Q is a positive integer.
[0092] Specifically, a fault sample set with known fault types can be constructed and divided into a training set and a validation set. For example, in a specific scenario, it is divided into a training set and a validation set according to a ratio of 6:4, so that cross-validation can be carried out subsequently.
[0093] For example, when Q = 3, according to the KNN algorithm, the samples in the validation set are selected in turn. After each selection of 1 sample in the validation set, the 3 samples in the training set that are closest to this sample can be determined, and the fault type with the most occurrences among these 3 samples is counted. If the counted fault type is the same as the fault type of this sample in the validation set, then this sample in the validation set can be regarded as accurately judged, otherwise it is regarded as inaccurate. It can be understood that after the samples in the validation set are selected in turn, the accuracy of the fault type judgment when Q = 3 can be obtained.
[0094] Similarly, for other values of Q, the accuracy of the fault type judgment can also be obtained. For example, if the value range of Q is from 1 to 20, a curve with Q on the horizontal axis and the accuracy of the fault type judgment on the vertical axis can be obtained. The Q value when this curve reaches the maximum is used as the value of the required number of neighbors Q.
[0095] After the value of the number of neighbors Q is determined, similarly, based on the KNN algorithm, using the constructed fault sample set with known fault types, the fault types of each fault representative matrix are judged, and the fault types of each fault representative matrix are obtained.
[0096] The principle of KNN is that when predicting a new value x, it is judged which category x belongs to according to the categories of the Q points closest to it.
[0097] Therefore, for example, if the determined number value Q of neighbors is 9, then for any one fault representative matrix, 9 fault samples that are closest to this fault representative matrix need to be selected from the constructed fault sample set of known fault types, and the fault type that appears most frequently among these 9 samples is used as the fault type of the determined fault representative matrix.
[0098] Applying the technical solution provided by the embodiments of the present invention, considering that when there are batch faults in the server, some of the faulty servers have commonalities, that is, the fault types of some servers may be similar. Therefore, in this application, the fault logs of N faulty servers are first respectively subjected to feature extraction to obtain the fault feature matrices of the N faulty servers. After that, according to the grouping principle that the fault feature matrices with high similarity are placed in the same group, the N fault feature matrices are divided into K groups. That is to say, for any one of the K groups, the same fault type can be used to represent each faulty server corresponding to this group. Therefore, for any one of the K groups, this application can determine a fault representative matrix for reflecting the fault type of this group based on the respective fault feature matrices divided into this group. After subsequently determining the fault types of the respective fault representative matrices of each group, the fault type of the fault representative matrix can be used as the fault type of the respective faulty servers corresponding to each fault feature matrix in the corresponding group. It can be seen that in the solution of this application, it is not necessary to individually determine the fault types of each fault feature matrix, but only to determine the fault types of K fault representative matrices, and then these can be used to represent the fault types of the K groups. Therefore, for the batch faults of the server, the solution of this application realizes efficient fault analysis, that is, it can quickly determine the fault types of each faulty server.
[0099] Corresponding to the above method embodiments, the embodiments of the present invention also provide a batch analysis system for server faults, which can be mutually corresponding and referenced with the above text.
[0100] See Figure 2 As shown in the structure schematic diagram of a batch analysis system for server faults in the present invention, it includes:
[0101] A fault feature matrix extraction module 201, configured to respectively perform feature extraction on the fault logs of N faulty servers to obtain the fault feature matrices of the N faulty servers; N is a positive integer not less than 2, and K is a positive integer and K ≤ N;
[0102] A grouping module 202, configured to divide the N fault feature matrices into K groups; wherein, when grouping, according to the grouping principle that the fault feature matrices with high similarity are placed in the same group;
[0103] The fault representative matrix determination module 203 is configured to, for any one of the K groups, determine a fault representative matrix for reflecting the fault type of the group based on each fault feature matrix partitioned into the group;
[0104] The group fault type determination module 204 is configured to, for any one of the K groups, determine the fault type of the fault representative matrix of the group, and use the fault type as the fault type of the fault server corresponding to each fault feature matrix in the group.
[0105] In a specific embodiment of the present invention, the grouping module 202 is specifically configured to:
[0106] Arbitrarily select K fault feature matrices from the N fault feature matrices as K initial mass points;
[0107] For any one of the N - K fault feature matrices that are not mass points currently, partition the fault feature matrix into the group where the mass point closest to itself is located, obtaining a total of K groups;
[0108] For any one of the current K groups, determine the mean value of each fault feature matrix in the group as the current verification mass point of the group;
[0109] For any one of the current K groups, when the distance between the current mass point of the group and the current verification mass point of the group is greater than the first threshold, use the current verification mass point of the group as the new mass point after the group is updated, and return to execute the operation of partitioning any one of the N - K fault feature matrices that are not mass points currently into the group where the mass point closest to itself is located to obtain a total of K groups for updating the group until when the distance between the current mass point of any group and the current verification mass point of the group is not greater than the first threshold, use the current K groups as the final groups.
[0110] In a specific embodiment of the present invention, when partitioning the fault feature matrix into the group where the mass point closest to itself is located, the Euclidean distance calculation formula is used to determine the distance between the fault feature matrix and each mass point.
[0111] In a specific embodiment of the present invention, it further includes:
[0112] The K value update module is configured to, after selecting K initial mass points, time through a timer, and when the timing duration reaches the first duration and the K final groups have not been determined yet, adjust the value of K and re - execute the operation of arbitrarily selecting K fault feature matrices from the N fault feature matrices as K initial mass points.
[0113] In a specific embodiment of the present invention, the fault representative matrix determination module 203 is specifically configured to:
[0114] For any one of the K groups, take the mean of the respective fault feature matrices divided into the group as the determined fault representative matrix for reflecting the fault type of the group.
[0115] In a specific embodiment of the present invention, the fault feature matrix extraction module 201 is specifically configured to:
[0116] Extract features from the respective fault logs of the N fault servers by means of outlier capture of the fault logs and / or by means of abnormal information frequency statistics of the fault logs.
[0117] In a specific embodiment of the present invention, the grouped fault type determination module 204 is specifically configured to:
[0118] Construct a fault sample set of known fault types and divide it into a training set and a validation set;
[0119] Under the condition that the set value Q of the number of neighbors is different, based on the KNN algorithm, use the training set to judge the fault types of the samples in the validation set, and determine the value of the number of neighbors Q that makes the fault type judgment accuracy rate the highest; Q is a positive integer;
[0120] Based on the determined value Q of the number of neighbors, based on the KNN algorithm, use the constructed fault sample set of known fault types to judge the fault types of the respective fault representative matrices, and obtain the respective fault types of the respective fault representative matrices.
[0121] Corresponding to the above method and system embodiments, the embodiments of the present invention also provide a batch analysis device for server faults and a computer-readable storage medium, which can be correspondingly referred to above.
[0122] Reference can be made to Figure 3 , and the batch analysis device for server faults may include:
[0123] A memory 301 for storing a computer program;
[0124] A processor 302 for executing the computer program to implement the steps of the batch analysis method for server faults in any of the above embodiments.
[0125] A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps of the method for batch analysis of server failures in any of the above embodiments. The computer-readable storage medium mentioned here includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well-known in the technical field.
[0126] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0127] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0128] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the technical solution and its core idea of the present invention. It should be pointed out that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A method for batch analysis of server failures, characterized in that, it includes: respectively extracting features from the fault logs of N faulty servers to obtain N fault feature matrices of the respective faulty servers; N N is a positive integer not less than 2, and K is a positive integer and K ≤ N; dividing the N fault feature matrices into K groups; wherein, when grouping, the fault feature matrices with high similarity are placed in the same group as the grouping principle; for any one of the K groups, based on the respective fault feature matrices divided into the group, determining a fault representative matrix for reflecting the fault type of the group; for any one of the K groups, determining the fault type of the fault representative matrix of the group, and taking the fault type as the fault type of the respective faulty servers corresponding to the respective fault feature matrices in the group; the dividing the N fault feature matrices into K groups includes: arbitrarily selecting K fault feature matrices from the N fault feature matrices as K initial mass points; for any one of the current N - K fault feature matrices that are not mass points, dividing the fault feature matrix into the group where the mass point closest to itself is located to obtain a total of K groups; for any one of the current K groups, determining the mean value of the respective fault feature matrices of the group as the current verification mass point of the group; for any one of the current K groups, when the distance between the current mass point of the group and the current verification mass point of the group is greater than the first threshold, taking the current verification mass point of the group as the new mass point after the update of the group, and returning to execute the operation of dividing any one of the current N - K fault feature matrices that are not mass points into the group where the mass point closest to itself is located to obtain a total of K groups for group update until when the distance between the current mass point of any group and the current verification mass point of the group is not greater than the first threshold, taking the current K groups as the final groups; after selecting K initial mass points, it further includes: timing through a timer, and when the timing duration reaches the first duration and the K final groups have not been determined yet, adjusting the value of K and re - executing the operation of arbitrarily selecting K fault feature matrices from the N fault feature matrices as K initial mass points.
2. The method for batch analysis of server failures according to claim 1, characterized in that, when dividing the fault feature matrix into the group where the mass point closest to itself is located, the Euclidean distance calculation formula is used to determine the distance between the fault feature matrix and each mass point.
3. The method for batch analysis of server failures according to claim 1, characterized in that, the determining, for any one of the K groups, a fault representative matrix for reflecting the fault type of the group based on the respective fault feature matrices divided into the group includes: For any one of the K groups, the mean value of each fault feature matrix divided into the group is used as the determined fault representative matrix for reflecting the fault type of the group.
4. The method for batch analysis of server faults according to claim 1, wherein, the step of respectively extracting features from the fault logs of N faulty servers includes: extracting features from the fault logs of N faulty servers respectively by means of outlier capture of the fault logs and / or by means of abnormal information frequency statistics of the fault logs.
5. The method for batch analysis of server faults according to any one of claims 1 to 4, wherein, the step of determining the fault type of the fault representative matrix of the group for any one of the K groups includes: constructing a fault sample set with known fault types and dividing it into a training set and a validation set; under the condition that the set neighbor number value Q is different, based on the KNN algorithm, using the training set to judge the fault type of the samples in the validation set, and determining the value of the neighbor number value Q that makes the fault type judgment accuracy rate the highest; Q is a positive integer; based on the determined neighbor number value Q, based on the KNN algorithm, using the constructed fault sample set with known fault types, judging the fault type of each fault representative matrix, and obtaining the fault type of each fault representative matrix; The grouping module is specifically used for: arbitrarily selecting K fault feature matrices from the N fault feature matrices as K initial mass points; for any one of the N-K fault feature matrices that are not mass points currently, dividing the fault feature matrix into the group where the mass point closest to itself is located, and obtaining a total of K groups; for any one of the current K groups, determining the mean value of each fault feature matrix of the group as the current verification mass point of the group; for any one of the current K groups, when the distance between the current mass point of the group and the current verification mass point of the group is greater than the first threshold, using the current verification mass point of the group as the new mass point after the group is updated, and returning to execute the operation of dividing any one of the N-K fault feature matrices that are not mass points currently into the group where the mass point closest to itself is located to obtain a total of K groups for group update until when the distance between the current mass point of any group and the current verification mass point of the group is not greater than the first threshold, taking the current K groups as the final groups; further includes: a K value update module, configured to start timing by a timer after selecting K initial mass points, and when the timing duration reaches the first duration and the K final groups have not been determined yet, adjust the value of K and re-execute the operation of arbitrarily selecting K fault feature matrices from the N fault feature matrices as K initial mass points.
6. A system for batch analysis of server faults, wherein, comprises: A fault feature matrix extraction module, configured to extract features from the fault logs of each of the N faulty servers respectively, so as to obtain the fault feature matrices of each of the N faulty servers; N N is a positive integer not less than 2, and K is a positive integer and K ≤ N; A grouping module, configured to divide the N fault feature matrices into K groups; wherein, when grouping, the fault feature matrices with high similarity are placed in the same group as the grouping principle; A fault representative matrix determination module, configured to, for any one of the K groups, determine a fault representative matrix for reflecting the fault type of the group based on each of the fault feature matrices divided into the group; A grouped fault type determination module, configured to, for any one of the K groups, determine the fault type of the fault representative matrix of the group, and use the fault type as the fault type of each of the faulty servers corresponding to the fault feature matrices in the group.
7. A batch analysis device for server faults Characterized in that It includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the steps of the batch analysis method for server faults according to any one of claims 1 to 5.
8. A computer-readable storage medium Characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the batch analysis method for server faults according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Maglev train, and levitation system fault detection method and system of maglev train
CN111460392A
Server fault analysis method and device
CN113835918A