Downtime fault analysis method and device, electronic equipment and medium

By obtaining the status information of the downtime server of the cloud computing platform, performing feature extraction and rule matching, and automatically determining the root cause of the downtime, solving the problems of low analysis efficiency and insufficient accuracy in the existing technology, and achieving efficient and accurate downtime failure analysis.

CN120386650APending Publication Date: 2025-07-29HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410119261.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the analysis of downtime failures of cloud computing platforms relies on manual experience, resulting in low analysis efficiency and difficult to ensure accuracy.

Method used

By obtaining the status information of the downtime server, performing feature extraction, establishing the correspondence between feature rules and the root cause of failure, and automatically determining the root cause of downtime.

Benefits of technology

Improves the efficiency and accuracy of downtime failure analysis and reduces the dependence of manual analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386650A_ABST
    Figure CN120386650A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a downtime fault analysis method and device, electronic equipment and a storage medium. The downtime fault analysis method comprises the following steps: acquiring downtime state information of a downtime server; the downtime state information is used for describing the running state information of the downtime server when the downtime fault occurs; performing feature extraction on the downtime state information to obtain downtime features of the downtime server; obtaining a feature rule and a corresponding relation between the feature rule and the fault root cause; and determining a target feature rule matched with the downtime feature, and determining a target fault root cause corresponding to the target feature rule as a downtime root cause of the downtime server based on the corresponding relationship. According to the embodiment of the invention, the analysis efficiency and accuracy of downtime fault analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, and a storage medium for analyzing a downtime fault. Background Art

[0002] When a server experiences a downtime fault, it will cause relatively serious impacts. For example, due to its advantages such as high security, scalability, and rapid deployment, cloud computing technology is favored by more and more users. A cloud computing platform includes a server cluster composed of a large number of physical servers. Virtual machines can be deployed in the above physical servers to provide corresponding computing services to different users. When a downtime fault occurs in the above physical server or virtual machine in the cloud computing platform, it will cause a poor user experience, thereby affecting user stickiness. When a downtime fault occurs, root cause analysis needs to be carried out to repair the problem as soon as possible.

[0003] In related technologies, the analysis of downtime faults usually relies on experienced experts to conduct manual analysis, which takes a long time, has a high cost, and it is difficult to guarantee the accuracy of the analysis.

[0004] Therefore, there is an urgent need for an efficient and accurate downtime fault analysis solution. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a method for analyzing a downtime fault, including:

[0006] Obtain the downtime status information of the downtime server; the downtime status information is used to describe the operating status information of the downtime server when the downtime fault occurs;

[0007] Extract features from the downtime status information to obtain the downtime features of the downtime server;

[0008] Obtain feature rules and the corresponding relationship between the feature rules and the root causes of faults;

[0009] Determine the target feature rule matched by the downtime features, and based on the corresponding relationship, determine the target root cause of the fault corresponding to the target feature rule as the root cause of the downtime of the downtime server.

[0010] According to the second aspect of the embodiments of the present application, another method for analyzing a downtime fault is provided, including:

[0011] Receive a downtime fault analysis request; the downtime fault analysis request includes a memory image file obtained through a host in a cloud computing platform; the memory image file is generated when a target virtual machine deployed in the host experiences a downtime fault;

[0012] Parse the process memory image file to obtain the crash status information; extract features from the crash status information to obtain crash features; determine the target feature rule matched by the crash features from the preset feature rules, and determine the root cause of the crash of the target virtual machine based on the corresponding relationship between the feature rules and the root causes of the faults;

[0013] Return the analysis result, where the analysis result includes the root cause of the crash of the target virtual machine.

[0014] According to the third aspect of the embodiments of the present application, another method for analyzing a crash fault is provided, including:

[0015] Receive a crash fault analysis request sent by a host in a cloud computing platform; the crash fault analysis request includes the memory image file generated by the host during the historical stage when the crash fault occurs;

[0016] Parse the process memory image file to obtain the crash status information; extract features from the crash status information to obtain crash features; determine the target feature rule matched by the crash features from the preset feature rules, and determine the historical root cause of the crash of the host based on the corresponding relationship between the feature rules and the root causes of the faults;

[0017] Return the analysis result to the host, where the analysis result includes the historical root cause of the crash.

[0018] According to the fourth aspect of the embodiments of the present application, a device for analyzing a crash fault is provided, including:

[0019] An information acquisition module, configured to acquire the crash status information of a crashed server; the crash status information is used to describe the running status information of the crashed server when the crash fault occurs;

[0020] A feature extraction module, configured to extract features from the crash status information to obtain the crash features of the crashed server;

[0021] A correspondence acquisition module, configured to acquire feature rules and the corresponding relationship between the feature rules and the root causes of the faults;

[0022] A root cause determination module, configured to determine the target feature rule matched by the crash features, and based on the corresponding relationship, determine the target root cause of the fault corresponding to the target feature rule as the root cause of the crash of the crashed server.

[0023] According to a fifth aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method described in any one of the first aspect to the third aspect.

[0024] According to a sixth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the method described in any one of the first aspect to the third aspect.

[0025] For the downtime fault analysis solution provided by the embodiments of the present application, after obtaining the downtime status information of the downtime server, feature extraction is performed on the downtime status information to obtain the downtime features of the downtime server. In addition, a pre-established feature rule and the corresponding relationship between the feature rule and the root cause of the fault are also obtained. Then, the target feature rule matched by the downtime features is determined, and further, according to the above corresponding relationship, the root cause of the downtime of the downtime server is obtained. In the embodiments of the present application, a rule-based matching method is adopted to automatically classify the downtime faults of the downtime server, so as to determine the root cause of the downtime. Compared with the manual analysis method, the analysis efficiency and accuracy of the downtime fault analysis are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a flowchart of the steps of a downtime fault analysis method according to Embodiment 1 of the present application;

[0028] Figure 2 For Figure 1 The corresponding analysis process schematic diagram of the illustrated embodiment;

[0029] Figure 3 It is a comparison schematic diagram of different classification scopes of action;

[0030] Figure 4 It is a flowchart of the classification of the root cause of the fault based on the stack matching rule;

[0031] Figure 5 It is a flowchart of the similar classification method;

[0032] Figure 6Schematic diagram of the process for classifying the root cause of a fault based on feature matching rules;

[0033] Figure 7 Schematic diagram of the process for classifying the root cause of a fault based on the combination of stack matching rules and feature matching rules;

[0034] Figure 8 Schematic diagram of the process for obtaining downtime status information;

[0035] Figure 9 Schematic diagram of the process for analyzing and obtaining a downtime status file in the Linux operating system;

[0036] Figure 10 Schematic diagram of the process for analyzing and obtaining a downtime status file in the Windows operating system;

[0037] Figure 11 Flowchart of the steps of a method for analyzing downtime faults according to Embodiment 2 of the present application;

[0038] Figure 12 Flowchart of the steps of a method for analyzing downtime faults according to Embodiment 3 of the present application;

[0039] Figure 13 Block diagram of the structure of a device for analyzing downtime faults according to Embodiment 4 of the present application;

[0040] Figure 14 Block diagram of the structure of a device for analyzing downtime faults according to Embodiment 5 of the present application;

[0041] Figure 15 Block diagram of the structure of a device for analyzing downtime faults according to Embodiment 6 of the present application;

[0042] Figure 16 Schematic diagram of the structure of an electronic device according to Embodiment 7 of the present application. Detailed implementation

[0043] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.

[0044] Embodiment 1

[0045] Refer to Figure 1 , Figure 1The flowchart of the steps of a downtime failure analysis method according to Embodiment 1 of the present application. Specifically, the downtime failure analysis method provided in this embodiment includes the following steps:

[0046] Step 102, obtain the downtime status information of the downtime server; the downtime status information is used to describe the running status information of the downtime server when the downtime failure occurs.

[0047] Specifically, the downtime server in the present application can be any server to be analyzed for downtime failure. For example: it can be a conventional server or a server in a cloud computing platform. Further, in terms of the cloud computing platform, the downtime server in this step can be a physical host or a virtual machine - cloud server instance deployed in the physical host.

[0048] The present application embodiment does not limit the specific method for obtaining the downtime status information of the downtime server. For example: usually, when a downtime failure occurs, a file representing the process execution status (such as a dump file) can be automatically generated inside the downtime server. Therefore, the above file can be obtained, and by performing pre - processing operations on the file (such as: content parsing, redundant information removal, format conversion, etc.), the downtime status information can be obtained; in addition, in some relatively special scenarios, the above file may not be generated. In this case, the kernel log information generated in the downtime server when the failure occurs can be obtained, and based on the above kernel log information, the downtime status information in this step can be obtained.

[0049] Step 104, perform feature extraction on the downtime status information to obtain the downtime features of the downtime server.

[0050] Specifically, the downtime status information usually contains a lot of content. Some content is related to the downtime failure, and some content is not related to the downtime failure. In this step, the process of performing feature extraction on the downtime status information can be a process of information screening or information refinement from the downtime status information, and the obtained downtime features can be the information in the downtime status information that has an association relationship with the downtime failure.

[0051] The downtime features in the embodiments of this application can be set according to the actual situation or expert experience. Further, in order to improve the accuracy of the fault analysis results, the downtime status information can be feature-extracted from multiple different dimensions, and then a variety of different downtime features can be obtained. For example: not only can the downtime status information be feature-extracted from the dimension of stack information to obtain stack information, but also, feature extraction can be performed from the perspective of the installed modules (such as third-party software, etc.) to obtain the information of the installed modules on the downtime server, and / or, key module information that may have an important impact on the downtime (or, has an associated relationship with the downtime fault). Of course, feature extraction can also be performed from the dimension of function calls to obtain the function call feature information of the downtime server, and so on. In the embodiments of this application, the specific manner of feature extraction, as well as the specific content of the obtained downtime features, are not limited and can be custom-set according to the actual situation.

[0052] Step 106, obtain the feature rules and the corresponding relationship between the feature rules and the root cause of the fault.

[0053] Specifically, the feature rules obtained in this step refer to the rules satisfied between features or combinations of features. Before performing the downtime fault analysis, one or more feature rules can be pre-generated, and for each feature rule, its corresponding root cause of the fault is correspondingly generated.

[0054] In practical applications, the definition of feature rules can be based on various factors such as expert experience, known laws, historical data, etc. And during the application process, the defined feature rules can also be continuously optimized and updated to improve the accuracy and reliability of the downtime root cause positioning.

[0055] The corresponding relationship between the feature rules and the root cause of the fault obtained in the embodiments of this application means that when the features of the downtime server (such as the above-mentioned downtime features) satisfy a certain feature rule, the downtime cause of the downtime server can be determined as the root cause of the fault corresponding to the above-mentioned feature rule. For example: Suppose the downtime features of the downtime server extracted through Step 104 include Feature A and Feature B, and the corresponding relationship between a certain feature rule and the root cause of the fault can indicate that when the downtime features satisfy the feature rule that the value of Feature A is a and the value of Feature B is b, the downtime cause of the downtime server can be determined as the root cause of the fault 1.

[0056] Step 108, determine the target feature rule matched by the downtime features, and based on the corresponding relationship, determine the target root cause of the fault corresponding to the target feature rule as the downtime root cause of the downtime server.

[0057] Specifically, in the embodiments of the present application, there is no limitation on the specific matching method adopted when matching the downtime characteristics with the feature rules obtained in step 106, and it can be custom-set according to the actual situation. For example, to improve the accuracy of the matching result, a complete matching method can be used to determine the target feature rule, that is: when the downtime characteristics are exactly the same as the target feature rule, it is determined that they match each other; another example is that to improve the flexibility of the matching process, a similarity matching method can also be used to determine the target feature rule, that is: when the similarity between the downtime characteristics and the target feature rule is relatively high, it is determined that they match each other, and so on.

[0058] As can be seen from the above steps 106 and 108, in the embodiments of the present application, a downtime fault classification algorithm based on rules (feature rules) is used to classify the downtime status information to determine the root cause category to which the root cause of the downtime server belongs. See Figure 2 , Figure 2 For Figure 1 the schematic diagram of the analysis process corresponding to the illustrated embodiment. The following combines Figure 2 to illustrate the process of the downtime fault analysis method of the above embodiments of the present application:

[0059] Specifically, the fault analysis process is as follows: Obtain the downtime status information of the downtime server, and this downtime status information can be used as the input condition of the downtime fault classification algorithm in the embodiments of the present application; perform downtime fault classification based on the downtime status information. Specifically: extract the features of the downtime status information to obtain the downtime characteristics of the downtime server; based on the matching between the features and the feature rules, perform downtime fault classification, that is: determine the target feature rule matched by the downtime characteristics, and based on the corresponding relationship between the feature rule and the root cause of the fault, determine the target root cause of the fault corresponding to the target feature rule as the downtime root cause of the downtime server; output the determined root cause.

[0060] The downtime fault analysis solution provided by the embodiments of the present application, after obtaining the downtime status information of the downtime server, extracts the features of the downtime status information to obtain the downtime characteristics of the downtime server; in addition, it also obtains the pre-established feature rules and the corresponding relationship between the feature rules and the root cause of the fault; then, determines the target feature rule matched by the downtime characteristics, and further obtains the downtime root cause of the downtime server according to the above corresponding relationship. The embodiments of the present application adopt a regularized matching method to automatically classify the downtime faults of the downtime server, thereby determining the downtime root cause. Compared with the manual analysis method, the analysis efficiency and accuracy of the downtime fault analysis are improved.

[0061] The downtime fault analysis method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, PCs, etc.

[0062] Optionally, in some embodiments, the crash features at least include: installed module information of the crashed server, critical module information, and critical stack information;

[0063] Feature extraction is performed on the crash status information to obtain the crash features of the crashed server, including:

[0064] By keyword search, determine the location of the installed module information of the crashed server and the location of the critical stack information in the crash status information;

[0065] Obtain the installed module information of the crashed server from the crash status information according to the location of the installed module information;

[0066] Obtain the critical stack information from the crash status information according to the location of the critical stack information.

[0067] Specifically, the above critical module information may be information of modules (such as third-party software, etc.) that have an associated relationship with the crash fault, that is, the modules that can cause the server to crash.

[0068] The critical stack information may be information describing the running position of the system code and the function call path when the crash fault occurs. Specifically, during the running process of the system code, stack information can be generated to record the current code running position and the function call path information. Those skilled in the art can understand that different code segments may be run and different functions may be called at different times. Therefore, over time, multiple stack information may be generated. In order to accurately analyze the crash fault, in the embodiments of the present application, the critical stack information corresponding to the moment when the crash fault occurs can be extracted from the crash status information for subsequent crash fault analysis.

[0069] The critical stack information can be complete stack information or partial data in the complete stack, for example: the last 3 layers of stack data, the last 5 layers of stack data, etc. The specific content of the critical stack information in the embodiments of the present application is not limited.

[0070] Furthermore, for different operating systems, their crash features may also be different. When performing feature extraction to obtain the crash features, a certain number of features can be selected from multiple features based on information such as the operating system type, experience, historical rules, etc. as the crash features. For example: Table 1 below shows the crash features corresponding to different operating systems selected according to experience:

[0071] Table 1

[0072]

[0073] The following subsystems are used to give an illustrative example of the process of feature extraction to obtain crash features:

[0074] For the Linux operating system, the process of feature extraction can be as follows: 1. Search for PanicInfo information, that is, the key information of the panic line, from the crash status information (which can be a text file); 2. When there are multiple stacks in the crash status information, the position of the key stack that has a greater impact on the crash failure can be located through the PanicInfo information and the CallTrace keyword; 3. Preprocess the key stack, including removing irrelevant information such as useless function lines and function lines with?, and parsing out the complete stack, the last 5 layers of stack data, the last 3 layers of stack data, etc. from the key stack; 4. Locate and parse the function information of the last call from multiple lines before and after the key stack through the "Rip" keyword; 5. Parse out the key module information from the key stack or the Rip line; 6. Locate and obtain the installed module list information through the "ModulesLinkedin" keyword.

[0075] For the Windows operating system, the process of feature extraction can be as follows: 1. When there are multiple stacks in the crash status information, locate the position of the key stack through the "KeBugCheckEx" keyword; 2. Preprocess the key stack, including removing irrelevant information such as useless function lines, and parsing out the complete stack, the last 3 layers, etc. from the key stack; 3. Parse out the key module Module, Rsp, bugCheckCode, etc. information from the key stack; 4. Locate and obtain the InstallModules installed module list information through the "lm" keyword.

[0076] Optionally, in some embodiments, the crash failure analysis method may further include: obtaining the system features of the crashed server; the system features characterize the characteristics of the operating system of the crashed server;

[0077] Determine the target feature rules matched by the crash features, including:

[0078] Determine the target feature rules satisfied by the crash features and the operating system features.

[0079] Specifically, in the embodiments of the present application, the specific content of the system features is not limited. For example, it may include at least one of the following: system type (such as Windows, Linux, etc.), release version (such as Windows 10, Ubuntu 20.04, etc.), system name, and kernel version, etc.

[0080] As described above, the downtime characteristics of different operating systems may vary. Therefore, before performing feature extraction operations, system features can be obtained first, and then corresponding downtime features can be adaptively extracted according to different system features.

[0081] In addition, in the process of classifying downtime status information using the downtime fault classification algorithm based on feature rules in step 108 above to determine the root cause category to which the root cause of the downtime server belongs, adding system features helps to improve the accuracy of the classification result, that is, it helps to improve the accuracy of the fault analysis result. When system features are obtained, the feature rules can refer to the rules satisfied between system features, downtime features, between system features and downtime features, between various system features, or between various downtime features; correspondingly, when performing rule matching, the target feature rule that matches both the downtime feature and the system feature can be determined, and the target root cause corresponding to the target feature rule can be determined as the downtime root cause of the downtime server.

[0082] For step 106 above, two different methods can be used to obtain the feature rules and the corresponding relationship between the feature rules and the root cause of the fault, that is, the root cause of the fault can be classified based on two different rules.

[0083] Among them, the first method is: the root cause classification method based on stack matching rules. Specifically: Based on the existing manually labeled data, a mapping table is established. When the downtime feature of a downtime status information is input, it can be searched in the mapping table to match the feature and determine the corresponding downtime root cause. This method relies on experience and existing manually labeled data to construct a mapping table to associate the feature combination with the downtime root cause.

[0084] The second method is: the root cause classification of faults based on feature matching rules. Specifically: Based on experience and existing knowledge, some feature rules are defined manually; when the combination of downtime features and / or system features conforms to these rules, the downtime fault is attributed to a certain downtime root cause. This method does not rely on manual labeling, but is classified based on human judgment and matching of features.

[0085] Next, the above-mentioned first root cause classification method of faults based on stack matching rules will be described in detail:

[0086] Optionally, in some embodiments, the corresponding relationship between the feature rules and the root cause of the fault is presented in the form of a mapping table, the feature rules are the keys in the mapping table, and the root cause of the fault is the value in the mapping table;

[0087] Determining the target feature rule that the downtime feature matches, and based on the corresponding relationship, determining the target root cause corresponding to the target feature rule as the downtime root cause of the downtime server includes:

[0088] Compare the downtime characteristics of the downtime server with the keys in the mapping table to determine the target key that matches the downtime characteristics of the downtime server;

[0089] Determine the value corresponding to the target key in the mapping table as the root cause of the downtime of the downtime server.

[0090] Specifically, in the above process, after obtaining the downtime characteristics, it is possible to search in the mapping table, that is, match the downtime characteristics with each key included in the mapping table, so as to determine the value corresponding to the matching target key as the root cause of the downtime of the downtime server.

[0091] Refer to Table 2 below. Table 2 is an exemplary mapping table provided by an embodiment of the present application.

[0092] Table 2

[0093]

[0094] As described above, the fault root cause classification method based on the stack matching rule depends on experience and existing labeled data to construct the mapping table. Hereinafter, the construction process of the mapping table, that is, the process of adding new key-value pairs to the mapping table, will be explained:

[0095] Optionally, in some embodiments, the feature rule includes features and the corresponding feature values; the process of adding new key-value pairs to the mapping table may include:

[0096] Obtain the downtime status information to be processed;

[0097] Extract features from the downtime status information to be processed to obtain multiple downtime characteristics included in the downtime status information to be processed;

[0098] Select at least some features from the multiple downtime characteristics included in the downtime status information to be processed, so as to determine at least some features and the corresponding feature values as the keys in the new key-value pair;

[0099] Determine the root cause of the downtime corresponding to the downtime status information to be processed, and use the determined root cause of the downtime as the value in the new key-value pair.

[0100] Specifically, the above process involves the determination of the following three types of information: classification basis, classification scope of action, and root cause. Among them, the classification basis is used to indicate which downtime characteristics are used for classifying the downtime status information to be processed; the classification scope of action mainly characterizes the further limitation of the system characteristics; the root cause is the determined fault root cause corresponding to the downtime status information to be processed.

[0101] The specific processing process is as follows: After extracting features from the downtime status information to be processed and obtaining multiple downtime features, at least some of the downtime features can be marked as classification criteria; at least some of the system features can be determined as the classification scope (this step is optional); in addition, a corresponding fault root cause can be set for the downtime status information to be processed. At this point, the processing process ends, and accordingly, a new key-value pair can be formed in the mapping table, wherein the above-set classification basis and classification scope jointly determine the key in the above-added key-value pair, and the set root cause determines the value in the above-added key-value pair.

[0102] The following is a detailed explanation of the classification scope: Setting the classification scope means adding system features that restrict the system in addition to the downtime features to be matched with the feature rules. Specifically, when setting the classification scope, you can set a specific kernel version (such as 3.10 kernel + Centos7 series), a specific distribution version (such as Centos, Ubuntu, etc.), a specific OS type (such as RHel, Debian, Suse, etc.), and of course, a specific system (Linux), etc. See Figure 3 , Figure 3 A comparative diagram of the scope of different classifications; Figure 3 It can be seen that the above five setting methods gradually relax the restrictions on the system, that is, when performing feature rule matching, the matching range gradually expands.

[0103] Furthermore, when the classification scope is fixed, selecting different downtime characteristics as the classification basis may also result in rule matching conditions of varying degrees of strictness. See Table 3 below, which shows the correspondence between classification basis and matching condition strictness.

[0104] Table 3

[0105]

[0106] Table 3 shows that the stringency of the matching conditions becomes increasingly relaxed as we move from top to bottom. This is because if two stacks have identical FullCallTrace values, they must also have identical data in the last three layers of the stack or the last five layers of the stack. However, the reverse is not true: the probability that two stacks have identical data in the last three layers of the stack is higher than the probability that their full stacks are identical. Therefore, requiring a full stack match is more stringent than requiring a match in the last three layers of the stack.

[0107] See also Figure 4 , Figure 4 This is a flowchart of fault root cause classification based on stack matching rules; Figure 4The process of classifying the above root causes of failures is described as follows:

[0108] After extracting the downtime characteristics from the current downtime status information and obtaining the system characteristics, based on the above characteristics, the current downtime status information can be compared with the downtime status information that has been labeled (processed) in the database to check whether certain feature combinations match the keys in the key-value pairs obtained based on the labeled downtime status information in the mapping table. If there is a match, the classification is successful, and the target value corresponding to the matched target key is determined as the root cause of the failure corresponding to the current downtime status information; otherwise, if there is no matching target key, the classification fails, indicating that the root cause of the failure corresponding to the current downtime status information has not been determined. Additionally, for the obtained current downtime status information, it can also be used as unlabeled downtime status information, and by means of manual marking, a key-value pair corresponding to the unlabeled downtime status information is added to the mapping table. After the key-value pair is constructed, the status of the above unlabeled downtime status information changes from the unclassified status (unlabeled status) to the classified status (labeled status).

[0109] Furthermore, regarding the step of comparing the downtime characteristics of the downtime server with the keys in the mapping table to determine the target key that matches the downtime characteristics of the downtime server, it can be achieved through the following two methods:

[0110] First, the "pattern classification method". The core of this method lies in that when two different downtime status information have the same downtime characteristics and system characteristics, it is considered that these two downtime status information have the same root cause of downtime. That is to say, in this method, the downtime characteristics and system characteristics of the downtime server are compared with the keys in the mapping table. When a certain key is exactly the same as the values of the downtime characteristics and system characteristics, then this key is determined as the target key.

[0111] Second, the "similarity classification method". The specific process is as follows: perform vector quantization encoding operations on each key in the mapping table to obtain the key vectors corresponding to each key in the mapping table; perform vector quantization encoding operations on the downtime characteristics and / or system characteristics of the downtime server to obtain the feature vectors; based on the similarity between the feature vectors and the key vectors corresponding to each key in the mapping table, determine the target key that matches the downtime characteristics.

[0112] In comparison, the first method above is based on exact matching. The advantage is that it can accurately match known data and provide highly accurate matching results; while the second method above uses a similarity detection algorithm to evaluate the similarity degree between different downtime status information by calculating the similarity between them. When the similarity exceeds a certain threshold, it can be considered that two downtime status information are very similar and may have the same root cause. Through this method, the algorithm can be generalized to improve the accuracy and adaptability of the matching algorithm.

[0113] See Figure 5 , Figure 5 which is a schematic flowchart of the similarity classification method. The following will explain the process of the above "similarity classification method" in combination with Figure 5 as follows:

[0114] Extract each key from the mapping table; perform encoding operations on each key to obtain key vectors, and store the obtained key vectors in the key vector library; obtain the total number of downtime status information to be classified; determine whether the number of classified downtime status information is less than the above total number; if the number of classified downtime status information is less than the above total number, obtain a piece of downtime status information to be classified, and search for similar key vectors in the above vector library; if there is a target key vector with a similarity greater than the preset threshold, modify the status of the downtime status information to be classified to the classified downtime status information, and determine the value corresponding to the above target key vector as the root cause of the failure corresponding to the classified downtime status information; if there is no target key vector with a similarity greater than the preset threshold, return to the step of determining whether the number of classified downtime status information is less than the above total number; if the number of classified downtime status information is equal to the above total number, end the process.

[0115] The following will first explain in detail the above second method for classifying the root cause of failure based on feature matching rules:

[0116] As described above, this method for classifying the root cause of failure does not rely on manual tagging, but defines some feature rules based on human experience and existing knowledge. When the combination of downtime features and / or system features meets these rules, the downtime failure is attributed to a certain root cause of downtime. In this case, the feature rules can be understood as the preset conditions satisfied by the features or feature combinations. Correspondingly, the process of determining the target feature rule matched by the downtime features and, based on the corresponding relationship, determining the target root cause of failure corresponding to the target feature rule as the root cause of downtime of the downtime server may include:

[0117] Determine the target preset conditions satisfied by the downtime features and / or system features, and based on the corresponding relationship between the preset conditions and the root cause of failure, determine the root cause of failure corresponding to the target preset conditions as the root cause of downtime of the downtime server.

[0118] See Figure 6 , Figure 6 which is a schematic flowchart of the classification of the root cause of failure based on feature matching rules. The following will explain and illustrate the process of classifying the root cause of failure based on feature matching rules in combination with Figure 6 as follows:

[0119] Obtain the to-be-classified (current) downtime status information, extract features from the to-be-classified downtime status information to obtain downtime features, and obtain system features; check whether the above features meet certain preset conditions; if so, the classification is successful, and determine the fault root cause corresponding to the above preset conditions as the downtime root cause corresponding to the to-be-classified downtime status information; if not, the classification fails, and the downtime root cause corresponding to the to-be-classified downtime status information is not analyzed.

[0120] In addition, when classifying the fault root cause, the above fault root cause classification method based on the stack matching rule and the above fault root cause classification based on the feature matching rule can also be combined to improve the accuracy of the classification result. See Figure 7 , Figure 7 is a schematic diagram of the fault root cause classification process for the combination of the stack matching rule and the feature matching rule. See Figure 7 , and the specific process may include: obtaining the downtime status information of the downtime server; extracting features from the above downtime status information to obtain downtime features, and obtaining system features; based on the above features, execute the fault root cause classification scheme based on the feature matching rule, if the classification is successful, end the process; if the classification fails, then based on the above features, execute the fault root cause classification scheme based on the stack matching rule again; if the classification is successful, end the process.

[0121] In the embodiments of the present application, the execution order of the above two schemes for the fault root cause classification based on the feature matching rule and the fault root cause classification based on the stack matching rule is not limited. Figure 7 Taking the example of first executing the fault root cause classification process of the feature matching rule and then executing the fault root cause classification process of the stack matching rule in

[0122] Optionally, in some embodiments, the process of obtaining the downtime status information of the downtime server in step 102 above may include:

[0123] Obtain the process memory image file; the process memory image file is a file representing the execution state of the process generated in the downtime server when the downtime fault occurs;

[0124] Execute a preset debugger instruction on the process memory image file to obtain a downtime status file, and the downtime status file contains downtime status information.

[0125] Specifically, usually when a downtime failure occurs, a process memory image file, such as a dump file, can be automatically generated inside the downtime server. Therefore, the above file can be obtained, and by executing a preset debugger instruction on the file, a downtime status file containing the downtime status information can be obtained. In the embodiments of the present application, the specific file format of the downtime status file is not limited. For the convenience of subsequent fault classification and analysis, its format can be set to a text format.

[0126] In addition, in some relatively special scenarios, the above file may not be generated in the downtime server. At this time, the kernel log information generated in the downtime server when the fault occurs can be obtained, and based on the above kernel log information (such as serial port log, dmsg log), the downtime status information in this step can be obtained.

[0127] See Figure 8 , Figure 8 is a schematic diagram of the acquisition process of the downtime status information. Specifically: for the case where there is a process memory image file, a preset debugger instruction can be executed on the above process memory image file to obtain a downtime status file containing the downtime status information; for the case where there is no process memory image file, the kernel log information generated in the downtime server when the fault occurs can be queried, and then a downtime status file containing the downtime status information can be obtained.

[0128] Optionally, in some embodiments, if the operating system type of the downtime server is Linux, executing a preset debugger instruction on the process memory image file to obtain a downtime status file includes:

[0129] Determine a file analysis environment corresponding to the version of the operating system; different file analysis environments correspond to different operating system versions, and different file analysis environments are encapsulated in different containers;

[0130] In the file analysis environment corresponding to the version of the operating system, execute a preset debugger instruction corresponding to the version of the operating system to obtain a downtime status file.

[0131] Specifically, for the Linux operating system, the crash state file is obtained by executing a series of fixed crash commands on the process memory image file. Due to the diversity of the Linux operating system, it is somewhat difficult to adapt to the correct analysis environment (including Crash and DebugInfo corresponding to the operating system version). To solve this problem, in the embodiments of the present application, a containerized analysis solution is adopted to separately construct a file analysis environment corresponding to the operating system version, and isolation between different file analysis environments is achieved. Therefore, it can effectively ensure that the analysis environment completely matches the operating system version, thereby improving the acquisition efficiency and accuracy of the crash state file.

[0132] See Figure 9 , Figure 9 FIG. is a schematic flowchart of obtaining a crash state file in the Linux operating system. Specifically: after obtaining the process memory image file, a containerized file analysis environment corresponding to the version of the operating system can be adapted according to the specific version information of the operating system; and then, preset debugger commands are executed in the above analysis environment to obtain the crash state file.

[0133] Optionally, in some embodiments, if the operating system type of the crashed server is Windows, executing preset debugger commands on the process memory image file to obtain a crash state file includes:

[0134] Obtain a customized script file; the customized script file is used to execute multiple debugger commands on the process memory image file, where during the execution of the debugger commands, the type and parameters of the (i + 1)-th debugger command are determined by the execution result of the i-th debugger command, and i is a natural number;

[0135] Run the customized script file to obtain the crash state file.

[0136] Specifically, for the Windows operating system, the crash state file is obtained by executing a series of windbg commands on the process memory image file. Different from Linux, preparing the analysis environment for Windows is relatively simple, but the number of debugging commands is large and not fixed, and needs to be adjusted according to specific situations. Therefore, in the embodiments of the present application, by pre-generating a customized script file, the process of customized file analysis according to specific situations is realized, so as to ensure that the crash state file can be obtained efficiently and accurately.

[0137] See Figure 10 , Figure 10It is a schematic flow chart of analyzing the crash status file in the Windows operating system. Specifically: an analysis environment can be prepared in advance. After obtaining the process memory image file, a customized script file can be run in the above analysis environment to obtain the crash status file.

[0138] Embodiment 2

[0139] Refer to Figure 11 , Figure 11 It is a step flow chart of a crash fault analysis method according to Embodiment 2 of the present application. Figure 11 The application scenario of the crash fault analysis method shown can be: in a cloud computing platform, a virtual machine deployed on a certain host machine crashes. At this time, the Figure 11 shown method can be used for crash fault analysis. Specifically, the crash fault analysis method provided in this embodiment includes the following steps:

[0140] Step 1102, receive a crash fault analysis request; the crash fault analysis request contains a memory image file obtained through the host machine in the cloud computing platform; the memory image file is generated when the target virtual machine deployed on the host machine crashes.

[0141] Step 1104, parse the process memory image file to obtain crash status information; extract features from the crash status information to obtain crash features; determine the target feature rule matched by the crash features from the preset feature rules, and based on the corresponding relationship between the feature rules and the root causes of the faults, determine the root cause of the crash of the target virtual machine.

[0142] Step 1106, return the analysis result, and the analysis result contains the root cause of the crash of the target virtual machine.

[0143] The crash fault analysis solution provided in the embodiment of the present application, after obtaining the crash status information of the target virtual machine, extracts features from the crash status information to obtain the crash features of the target virtual machine; in addition, the pre-established feature rules and the corresponding relationship between the feature rules and the root causes of the faults are also obtained; then, the target feature rule matched by the crash features is determined, and further the root cause of the crash of the target virtual machine is obtained according to the above corresponding relationship. In the embodiment of the present application, a rule-based matching method is adopted to automatically classify the crash faults of the target virtual machine, so as to determine the root cause of the crash. Compared with the manual analysis method, the analysis efficiency and accuracy of the crash fault analysis are improved.

[0144] Embodiment 3

[0145] Refer to Figure 12 , Figure 12 It is a step flow chart of a crash fault analysis method according to Embodiment 3 of the present application.Figure 12 The application scenario of the presented downtime fault analysis method can be as follows: In a cloud computing platform, a certain host machine experiences a downtime fault. After that, the host machine restarts. At this time, the method shown in Figure 12 can be used for downtime fault analysis. Specifically, the downtime fault analysis method provided in this embodiment includes the following steps:

[0146] Step 1202: Receive a downtime fault analysis request sent by a host machine in the cloud computing platform; the downtime fault analysis request includes the memory image file generated by the host machine during the historical stage when the downtime fault occurred.

[0147] Step 1204: Analyze the process memory image file to obtain downtime status information; extract features from the downtime status information to obtain downtime features; determine the target feature rule that the downtime features match from the preset feature rules, and based on the corresponding relationship between the feature rules and the root cause of the fault, determine the historical root cause of the downtime of the host machine.

[0148] Step 1206: Return the analysis result to the host machine, and the analysis result includes the historical root cause of the downtime.

[0149] In the downtime fault analysis solution provided by the embodiments of this application, after obtaining the downtime status information of the host machine, features are extracted from the downtime status information to obtain the downtime features of the host machine; in addition, the pre-established feature rules and the corresponding relationship between the feature rules and the root cause of the fault are also obtained; then, the target feature rule that the downtime features match is determined, and further, the root cause of the downtime of the host machine is obtained according to the above corresponding relationship. In the embodiments of this application, a rule-based matching method is adopted to automatically classify the downtime faults of the host machine, thereby determining the root cause of the downtime. Compared with the manual analysis method, the analysis efficiency and accuracy of the downtime fault analysis are improved.

[0150] Embodiment 4

[0151] Figure 13 It is a structural block diagram of a downtime fault analysis device according to Embodiment 4 of this application. The downtime fault analysis device provided by the embodiments of this application includes:

[0152] An information acquisition module 1302, configured to acquire the downtime status information of the downtime server; the downtime status information is used to describe the running status information of the downtime server when the downtime fault occurs;

[0153] A feature extraction module 1304, configured to extract features from the downtime status information to obtain the downtime features of the downtime server;

[0154] A corresponding relationship acquisition module 1306, configured to acquire feature rules and the corresponding relationship between the feature rules and the root cause of the fault;

[0155] A root cause determination module 1308, configured to determine a target feature rule that matches the crash feature, and based on the corresponding relationship, determine the target fault root cause corresponding to the target feature rule as the crash root cause of the crashed server.

[0156] Optionally, in some embodiments, the corresponding relationship between the feature rule and the fault root cause is presented in the form of a mapping table, where the feature rule is the key in the mapping table and the fault root cause is the value in the mapping table;

[0157] The root cause determination module 1308 is specifically configured to: compare the crash feature of the crashed server with the keys in the mapping table to determine a target key that matches the crash feature of the crashed server;

[0158] Determine the value corresponding to the target key in the mapping table as the crash root cause of the crashed server.

[0159] Optionally, in some embodiments, the feature rule includes a feature and the corresponding feature value; the crash fault analysis device further includes:

[0160] A new key-value pair adding module, configured to obtain multiple crash features to be processed; select at least some features from the multiple crash features included in the crash status information to be processed, so as to determine at least some features and the corresponding feature values as the keys in the new key-value pair; determine the crash root cause corresponding to the crash status information to be processed, and determine the determined crash root cause as the value in the new key-value pair.

[0161] Optionally, in some embodiments, the crash feature at least includes: the installed module information of the crashed server, the key module information, and the key stack information;

[0162] The feature extraction module 1304 is specifically configured to: locate the location of the installed module information of the crashed server and the location of the key stack information in the crash status information through keyword search; obtain the installed module information of the crashed server from the crash status information according to the location of the installed module information; obtain the key stack information from the crash status information according to the location of the key stack information.

[0163] Optionally, in some embodiments, the information acquisition module 1302 is specifically configured to:

[0164] Obtain a process memory image file; the process memory image file is a file generated in the crashed server during the occurrence of the crash fault, which represents the execution status of the process;

[0165] Execute a preset debugger instruction on the process memory image file to obtain a crash status file, and the crash status file contains crash status information.

[0166] Optionally, in some embodiments, if the operating system type of the crashed server is Linux, when the information acquisition module 1302 executes the step of executing a preset debugger instruction on the process memory image file to obtain a crash status file, it specifically is used for:

[0167] Determine a file analysis environment corresponding to the version of the operating system; different file analysis environments correspond to different operating system versions, and different file analysis environments are encapsulated in different containers;

[0168] In the file analysis environment corresponding to the version of the operating system, execute a preset debugger instruction corresponding to the version of the operating system to obtain a crash status file.

[0169] Optionally, in some embodiments, if the operating system type of the crashed server is Windows, when the information acquisition module 1302 executes the step of executing a preset debugger instruction on the process memory image file to obtain a crash status file, it specifically is used for:

[0170] Obtain a customized script file; the customized script file is used to execute multiple debugger instructions on the process memory image file, wherein during the execution of the debugger instructions, the type and parameters of the (i + 1)-th debugger instruction are determined by the execution result of the i-th debugger instruction, and i is a natural number;

[0171] Run the customized script file to obtain a crash status file.

[0172] Optionally, in some embodiments, the crash fault analysis device further includes:

[0173] A system feature acquisition module, configured to acquire system features of the crashed server; the system features characterize the characteristics of the operating system of the crashed server;

[0174] Correspondingly, when the root cause determination module 1308 executes the step of determining the target feature rule matched by the crash feature, it specifically is used for: determining the target feature rule satisfied by the crash feature and the operating system feature.

[0175] The crash fault analysis device of this embodiment is used to implement the corresponding crash fault analysis method in the foregoing Embodiment 1, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the crash fault analysis device of this embodiment can refer to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0176] Embodiment 5

[0177] Figure 14The structural block diagram of a downtime fault analysis device according to Embodiment 5 of the present application. The downtime fault analysis device provided by the embodiment of the present application includes:

[0178] A first request receiving module 1402, configured to receive a downtime fault analysis request; the downtime fault analysis request includes a memory image file obtained through a host in a cloud computing platform; the memory image file is generated when a target virtual machine deployed in the host has a downtime fault;

[0179] A first analysis module 1404, configured to parse the process memory image file to obtain downtime status information; extract features from the downtime status information to obtain downtime features; determine a target feature rule matched by the downtime features from a preset feature rule, and based on the correspondence between the feature rule and the root cause of the fault, determine the root cause of the downtime of the target virtual machine;

[0180] A first result returning module 1406, configured to return an analysis result, and the analysis result includes the root cause of the downtime of the target virtual machine.

[0181] The downtime fault analysis device in this embodiment is used to implement the corresponding downtime fault analysis method in Embodiment 2 above, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the downtime fault analysis device in this embodiment can refer to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0182] Embodiment 6

[0183] Figure 15 The structural block diagram of a downtime fault analysis device according to Embodiment 6 of the present application. The downtime fault analysis device provided by the embodiment of the present application includes:

[0184] A second request receiving module 1502, configured to receive a downtime fault analysis request sent by a host in a cloud computing platform; the downtime fault analysis request includes a memory image file generated by the host in the historical stage of the downtime fault;

[0185] A second parsing module 1504, configured to parse the process memory image file to obtain downtime status information; extract features from the downtime status information to obtain downtime features; determine a target feature rule matched by the downtime features from a preset feature rule, and based on the correspondence between the feature rule and the root cause of the fault, determine the historical root cause of the downtime of the host;

[0186] A second result returning module 1506, configured to return an analysis result to the host, and the analysis result includes the historical root cause of the downtime.

[0187] The downtime fault analysis device of this embodiment is used to implement the corresponding downtime fault analysis method in the third embodiment above, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the downtime fault analysis device of this embodiment can refer to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0188] Embodiment VII

[0189] Referring to Figure 16 , a schematic structural diagram of an electronic device according to Embodiment VII of the present application is shown. The specific implementation of the electronic device in the specific embodiments of the present application is not limited.

[0190] As Figure 16 shown, the electronic device may include: a processor 1602, a communication interface 1604, a memory 1606, and a communication bus 1608.

[0191] Among them:

[0192] The processor 1602, the communication interface 1604, and the memory 1606 communicate with each other through the communication bus 1608.

[0193] The communication interface 1604 is used to communicate with other electronic devices.

[0194] The processor 1602 is used to execute the program 1610, and specifically can execute the relevant steps in the foregoing downtime fault analysis method embodiment.

[0195] Specifically, the program 1610 may include program code, and the program code includes computer operation instructions.

[0196] The processor 1602 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0197] The memory 1606 is used to store the program 1610. The memory 1606 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0198] The program 1610 may include multiple computer instructions. Specifically, the program 1610 may cause the processor 1602 to perform the operations corresponding to the downtime fault analysis method described in any one of the foregoing method embodiments through the multiple computer instructions.

[0199] For the specific implementation of each step in the program 1610, reference may be made to the corresponding descriptions in the corresponding steps and units in the foregoing method embodiments, and they have corresponding beneficial effects, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.

[0200] The embodiment of the present application also provides a computer storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method described in any one of the foregoing method embodiments. The computer storage medium includes but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, magneto-optical disk, etc.

[0201] The embodiment of the present application also provides a computer program product, including computer instructions, which direct a computing device to perform the operations corresponding to any one of the downtime fault analysis methods in the foregoing multiple method embodiments.

[0202] In addition, it should be noted that the information related to users (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data for training the model, data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the users or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0203] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of the components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0204] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored on such a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)) for software processing. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a Random Access Memory (RAM), a Read-Only Memory (ROM), a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0205] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0206] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A method for analyzing downtime faults, comprising: Obtaining the downtime status information of the downed server; The downtime status information is used to describe the running status information of the downed server when the downtime fault occurs; Performing feature extraction on the downtime status information to obtain the downtime features of the downed server; Obtaining feature rules and the corresponding relationship between the feature rules and the root causes of faults; Determining the target feature rule that matches the downtime feature, and based on the corresponding relationship, determining the target root cause of the fault corresponding to the target feature rule as the root cause of the downtime of the downed server.

2. The method according to claim 1, wherein The corresponding relationship between the feature rules and the root causes of faults is presented in the form of a mapping table. The feature rules are the keys in the mapping table, and the root causes of faults are the values in the mapping table; The determining the target feature rule that matches the downtime feature, and based on the corresponding relationship, determining the target root cause of the fault corresponding to the target feature rule as the root cause of the downtime of the downed server includes: Comparing the downtime features of the downed server with the keys in the mapping table to determine the target key that matches the downtime features of the downed server; Determining the value corresponding to the target key in the mapping table as the root cause of the downtime of the downed server.

3. The method according to claim 2, wherein The feature rules include features and the corresponding feature values; The process of adding a new key-value pair to the mapping table includes: Obtaining the downtime status information to be processed; Performing feature extraction on the downtime status information to be processed to obtain multiple downtime features included in the downtime status information to be processed; Selecting at least some features from the multiple downtime features included in the downtime status information to be processed, so as to determine the at least some features and the corresponding feature values as the keys in the new key-value pair; Determining the root cause of the downtime corresponding to the downtime status information to be processed, and using the determined root cause of the downtime as the value in the new key-value pair.

4. The method according to any one of claims 1-3, wherein The downtime features at least include: installed module information of the downed server, key module information, and key stack information; The performing feature extraction on the downtime status information to obtain the downtime features of the downed server includes: Determining the location of the installed module information of the downed server and the location of the key stack information in the downtime status information through keyword search; Obtaining the installed module information of the downed server from the downtime status information according to the location of the installed module information; Obtaining the key stack information from the downtime status information according to the location of the key stack information.

5. The method according to any one of claims 1 to 3, wherein The obtaining the downtime status information of the downed server includes: Obtaining a process memory image file; the process memory image file is a file generated in the downed server when the downtime fault occurs, which represents the execution status of the process; Executing a preset debugger instruction on the process memory image file to obtain a downtime status file, and the downtime status file contains downtime status information.

6. The method according to claim 5, wherein If the operating system type of the downed server is Linux, the executing a preset debugger instruction on the process memory image file to obtain a downtime status file includes: Determine a file analysis environment corresponding to the version of the operating system; different file analysis environments correspond to different operating system versions, and different file analysis environments are encapsulated in different containers; In the file analysis environment corresponding to the version of the operating system, execute a preset debugger instruction corresponding to the version of the operating system to obtain a crash status file.

7. The method according to claim 5, wherein If the operating system type of the crashed server is Windows, the executing a preset debugger instruction on the process memory image file to obtain a crash status file includes: Obtain a customized script file; the customized script file is used to execute multiple debugger instructions on the process memory image file, where during the execution of the debugger instruction, the type and parameters of the (i + 1)-th debugger instruction are determined by the execution result of the i-th debugger instruction, and i is a natural number; Run the customized script file to obtain a crash status file.

8. The method according to claim 1, wherein The method further includes: obtaining system characteristics of the crashed server; the system characteristics characterize the characteristics of the operating system of the crashed server; The determining the target feature rule matched by the crash characteristics includes: Determine the target feature rule satisfied by the crash characteristics and the operating system characteristics.

9. A method for analyzing a crash fault, including: Receive a crash fault analysis request; The crash fault analysis request includes a memory image file obtained through a host in a cloud computing platform; The memory image file is generated when a target virtual machine deployed in the host has a crash fault; Parse the process memory image file to obtain crash status information; perform feature extraction on the crash status information to obtain crash characteristics; Determine a target feature rule matched by the crash characteristics from preset feature rules, and determine the root cause of the crash of the target virtual machine based on the correspondence between the feature rule and the root cause of the fault; Return an analysis result, where the analysis result includes the root cause of the crash of the target virtual machine.

10. A method for analyzing a crash fault, including: Receive a crash fault analysis request sent by a host in a cloud computing platform; The crash fault analysis request includes a memory image file generated by the host during a historical stage of a crash fault; Parse the process memory image file to obtain crash status information; perform feature extraction on the crash status information to obtain crash characteristics; Determine a target feature rule matched by the crash characteristics from preset feature rules, and determine the historical root cause of the crash of the host based on the correspondence between the feature rule and the root cause of the fault; Return the analysis result to the host, where the analysis result includes the historical root cause of the crash.

11. A crash fault analysis device, including: An information acquisition module, configured to acquire crash status information of a crashed server; The crash status information is used to describe the running status information of the crashed server when a crash fault occurs; A feature extraction module, configured to perform feature extraction on the crash status information to obtain crash characteristics of the crashed server; A correspondence acquisition module, configured to acquire feature rules and the correspondence between the feature rules and the root cause of the fault; A root cause determination module, configured to determine a target feature rule that matches the downtime feature, and based on the corresponding relationship, determine the target fault root cause corresponding to the target feature rule as the downtime root cause of the downtime server.

12. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method according to any one of claims 1-10.

13. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to any one of claims 1-10.

14. A computer program product, including computer instructions, where the computer instructions direct a computing device to perform operations corresponding to the method according to any one of claims 1-10.