Kernel downtime diagnosis method and device, electronic equipment and storage medium

By obtaining kernel dump files and loading the downtime detection rule library, automatically analyzing the causes of Linux kernel downtime, solving the problem of inefficient rapid diagnosis of kernel downtime problems in the existing technology, and achieving efficient kernel downtime repair.

CN120276889APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410032898.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, it is difficult to quickly confirm the cause of kernel downtime caused by Linux kernel errors, and the threshold for manpower analysis is high, resulting in low processing efficiency. Operation personnel need to re-check after replacement.

Method used

By obtaining the kernel dump file, loading the downtime detection engine, obtaining the downtime detection rule library from the yum source, generating the kernel downtime diagnosis results based on the memory status information matching, and sending the results to the background server for repair processing, or generating a diagnostic work order to hand it over to the kernel development node to update the downtime detection rule library of the local and yum sources.

Benefits of technology

It realizes automatic analysis and diagnosis of kernel downtime problems, improves diagnostic efficiency, reduces labor costs, and ensures stable operation of the operating system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276889A_ABST
    Figure CN120276889A_ABST
Patent Text Reader

Abstract

The invention discloses a kernel downtime diagnosis method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a kernel dump file which stores memory state information of an operating system of a target machine during kernel downtime; loading a downtime detection engine, and obtaining a downtime detection rule base from the yum source based on the downtime detection engine; the downtime detection rule base comprises a plurality of downtime detection rules for indicating different preset kernel downtime reasons, and each downtime detection rule comprises abnormal memory state information corresponding to the corresponding preset kernel downtime reason and operating system repair logic; and based on the matching condition between the memory state information in the kernel dump file and the abnormal memory state information in each downtime detection rule, generating a kernel downtime diagnosis result and sending the kernel downtime diagnosis result to the background server, so that the background server repairs the operating system of the target machine based on the kernel diagnosis result. According to the method, the diagnosis efficiency of the kernel downtime problem is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method, apparatus, electronic device, and storage medium for diagnosing kernel crashes. Background Art

[0003] In the related art, for the problem of kernel crashes caused by Linux kernel errors, the cause of the kernel error (bug) is mainly determined through manual analysis. Not only can the cause of the bug not be confirmed within a short time, but also due to the high analysis threshold, there are high technical requirements for operation personnel. After a kernel crash problem that has been investigated and resolved occurs again, if the operation personnel are replaced, it may lead to the need for re-investigation, greatly reducing the processing efficiency of kernel crash problems. Summary of the Invention

[0004] To solve the problems of the prior art, embodiments of this application provide a method, apparatus, electronic device, and storage medium for diagnosing kernel crashes. The technical solutions are as follows:

[0005] On the one hand, a method for diagnosing kernel crashes is provided. The method includes:

[0006] Obtain a kernel dump file, where the kernel dump file stores memory state information of the operating system of the target machine when the kernel crashes;

[0007] Load a crash detection engine, and obtain a crash detection rule library from the yum source based on the crash detection engine; the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash causes, and each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash cause and an operating system repair logic;

[0008] Generate a kernel crash diagnosis result based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule;

[0009] Send the kernel crash diagnosis result to a background server so that the background server performs repair processing on the operating system of the target machine based on the kernel diagnosis result.

[0010] On the other hand, a method for diagnosing kernel crashes is provided. The method includes:

[0011] Obtain the kernel crash diagnosis result sent by the target machine; the kernel crash diagnosis result is obtained by the target machine by acquiring a kernel dump file, loading a crash detection engine, and obtaining a crash detection rule library from the yum source based on the crash detection engine. The crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons. Each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash reason and an operating system repair logic, and is generated based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule;

[0012] In the case where the kernel crash diagnosis result indicates that no preset kernel crash reason is matched, generate a kernel crash diagnosis work order based on the memory state information in the kernel dump file carried by the kernel crash diagnosis result, and send the kernel crash diagnosis work order to the kernel development node;

[0013] Obtain the new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, and add the new crash detection rule to the local crash detection rule library to obtain an updated local crash detection rule library;

[0014] Send the updated local crash detection rule library to the yum source to update the crash detection rule library of the yum source.

[0015] On the other hand, a kernel crash diagnosis device is provided. The device includes:

[0016] A kernel dump file acquisition module for acquiring a kernel dump file, where the kernel dump file stores the memory state information of the operating system of the target machine at the time of kernel crash;

[0017] A crash rule library acquisition module for loading a crash detection engine and obtaining a crash detection rule library from the yum source based on the crash detection engine; the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons, and each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash reason and an operating system repair logic;

[0018] A kernel crash diagnosis module for generating a kernel crash diagnosis result based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule;

[0019] A diagnosis result sending module for sending the kernel crash diagnosis result to a background server so that the background server performs a repair process on the operating system of the target machine based on the kernel diagnosis result.

[0020] In an exemplary embodiment, the kernel crash diagnosis module includes:

[0021] A rule library traversal module for traversing the crash detection rules in the obtained crash detection rule library;

[0022] A first acquisition module for, for the currently traversed crash detection rule, based on the memory state type of the abnormal memory state information corresponding to the current crash detection rule, acquiring target memory state information corresponding to the memory state type from the kernel dump file;

[0023] A kernel crash cause determination module for, when the target memory state information matches the abnormal memory state information, determining the preset kernel crash cause corresponding to the current crash detection rule as the target kernel crash cause that causes the operating system to crash, and ending the traversal of the crash detection rule library;

[0024] A first diagnosis result generation module for generating a first kernel crash diagnosis result based on the current crash detection rule.

[0025] In an exemplary embodiment, the background server stores the crash detection rule library; the first diagnosis result generation module includes:

[0026] An identification information acquisition module for acquiring the target identification information of the current crash detection rule;

[0027] A first diagnosis result generation sub-module for generating a first kernel crash diagnosis result based on the target identification information; when the background server performs a repair process on the operating system of the target machine based on the kernel diagnosis result, it is used to determine a matching target kernel crash detection rule from the local crash detection rule library based on the target identification information, and acquire the operating system repair logic in the target kernel crash detection rule, and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0028] In an exemplary embodiment, the kernel crash diagnosis module further includes:

[0029] A loop execution module for, when the target memory state information does not match the abnormal memory state information and there are un-traversed crash detection rules in the crash detection rule library, returning to execute the step of traversing the crash detection rules in the crash detection rule library until there are no un-traversed crash detection rules in the crash detection rule library;

[0030] A second diagnostic result generation module, configured to generate a second kernel crash diagnostic result based on the kernel dump file, where the second kernel crash diagnostic result indicates that no preset kernel crash cause is matched; when the background server repairs the operating system of the target machine based on the kernel diagnostic result, it is configured to generate a kernel crash diagnostic work order based on the memory status information in the kernel dump file, send the kernel crash diagnostic work order to the kernel development node, obtain a new crash detection rule returned by the kernel development node for the kernel crash diagnostic work order, and send the operating system repair logic in the new crash detection rule to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0031] On the other hand, a kernel crash diagnosis device is provided, and the device includes:

[0032] A diagnostic result acquisition module, configured to acquire a kernel crash diagnostic result sent by the target machine; the kernel crash diagnostic result is obtained by the target machine by acquiring a kernel dump file, loading a crash detection engine, obtaining a crash detection rule library from the yum source based on the crash detection engine, the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash causes, each crash detection rule includes abnormal memory status information corresponding to the corresponding preset kernel crash cause and operating system repair logic, and is generated based on the matching situation between the memory status information in the kernel dump file and the abnormal memory status information in each crash detection rule;

[0033] A crash diagnostic work order generation module, configured to generate a kernel crash diagnostic work order based on the memory status information in the kernel dump file carried by the kernel crash diagnostic result and send the kernel crash diagnostic work order to the kernel development node when the kernel crash diagnostic result indicates that no preset kernel crash cause is matched;

[0034] A local crash rule library update module, configured to obtain a new crash detection rule returned by the kernel development node for the kernel crash diagnostic work order, and add the new crash detection rule to the local crash detection rule library to obtain an updated local crash detection rule library;

[0035] A yum source crash rule library update module, configured to send the updated local crash detection rule library to the yum source to update the crash detection rule library of the yum source.

[0036] In an exemplary embodiment, the yum source crash rule library update module includes:

[0037] An RPM package generation module, configured to generate an RPM software package based on the full updated local kernel crash detection rule library in response to an RPM package generation instruction for updating the local kernel crash detection rule library;

[0038] An RPM package push module, configured to send the RPM software package to the Yum source.

[0039] In an exemplary embodiment, the apparatus further includes:

[0040] A target kernel crash detection rule determination module, configured to determine a matching target kernel crash detection rule from the local kernel crash detection rule library based on the target identification information carried in the kernel crash diagnosis result when the kernel crash diagnosis result indicates a match with the preset kernel crash cause;

[0041] An operating system repair module, configured to obtain the operating system repair logic in the target kernel crash detection rule and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0042] In an exemplary embodiment, the apparatus further includes:

[0043] A notification message generation module, configured to generate a kernel crash notification message corresponding to the target machine based on the kernel crash diagnosis result;

[0044] A notification message sending module, configured to send the kernel crash notification message to an operation and maintenance node.

[0045] On the other hand, an electronic device is provided, including a processor and a memory, where at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the kernel crash diagnosis method in any of the above aspects.

[0046] On the other hand, a computer-readable storage medium is provided, where at least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the kernel crash diagnosis method as in any of the above aspects.

[0047] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the kernel crash diagnosis method in any of the above aspects.

[0048] In an embodiment of the present application, by obtaining a kernel dump file which stores the memory state information of the operating system of the target machine when the kernel crashes, and then loading a crash detection engine, a crash detection rule library is obtained from the yum source based on the crash detection engine. The crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons. Each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash reason and an operating system repair logic. Based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule, a kernel crash diagnosis result is generated, and the kernel crash diagnosis result is sent to a background server, so that the background server can perform a repair process on the operating system of the target machine based on the kernel crash diagnosis result, thereby realizing timely and automatic analysis and diagnosis of the kernel crash problem, and greatly improving the diagnosis efficiency of the kernel crash problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0051] Figure 2 It is a schematic flowchart of a kernel crash diagnosis method provided by an embodiment of the present application;

[0052] Figure 3 It is a schematic flowchart of another kernel crash diagnosis method provided by an embodiment of the present application;

[0053] Figure 4 It is a schematic flowchart of another kernel crash diagnosis method provided by an embodiment of the present application;

[0054] Figure 5 It is an architecture example for implementing the kernel crash diagnosis method provided by an embodiment of the present application;

[0055] Figure 6 It is a structural block diagram of a kernel crash diagnosis device provided by an embodiment of the present application;

[0056] Figure 7 It is a structural block diagram of another kernel crash diagnosis device provided by an embodiment of the present application;

[0057] Figure 8 It is a hardware structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0059] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0060] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0061] It can be understood that in the specific implementation manners of the present application, when it comes to data related to user information, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0062] The following will explain the terms involved in the embodiments of the present application.

[0063] Kernel: In computer technology, a kernel is a computer program used to manage the data input / output (IO) requests sent by software, translate these requests into instructions for data processing, and hand them over to the central processing unit (CPU) and computer components for processing. It is the most basic part of an operating system.

[0064] Downtime: It refers to the situation where the operating system cannot recover from a serious system error, or there are serious problems at the system hardware level, resulting in the system being unresponsive for a long time and having to restart the computer. Downtime will cause service interruption.

[0065] Kernel downtime: It refers to the downtime caused by errors in kernel operation.

[0066] Vmcore: In the Linux system, it refers to the kernel dump file generated when the system crashes or encounters serious errors. This file contains information such as the memory image, register status, and stack at the time of system crash.

[0067] Crash: A powerful command-line tool that can be used to analyze vmcore files. After installing the crash tool, you can use the crash command to load the vmcore file and execute various commands to obtain information related to system crashes, such as process status, memory mapping, stack, etc.

[0068] RPM: Red Hat Package Manager, a software package manager used to install, uninstall, and update software packages in operating systems based on Red Hat Linux. It includes a software package file format, a set of tools for managing software packages, and some metadata included in the software packages. RPM software packages usually end with the.rpm file extension, and these files contain binary files, library files, configuration files, documentation, etc. for installing or upgrading software packages.

[0069] Yum source: A software repository containing RPM software packages, used for automatically installing and upgrading RPM software packages, automatically finding and handling dependencies between RPM packages, and installing all dependent software packages at once without the need for cumbersome downloads and installations one by one.

[0070] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a machine cluster 110, a background server 120, and a yum source 130. Among them, each machine (such as 111, 112, 113) in the machine cluster 110, the background server 120, and the yum source can be connected and communicate through wired or wireless networks.

[0071] Among them, the machines (such as 111, 112, 113) in the machine cluster 110 can be servers with a Linux operating system deployed in a computer room, can also be servers with a Linux operating system deployed in a multi-cloud environment, and can also be various types of hosts such as physical machines and virtual machines.

[0072] In the embodiments of the present application, an agent program is installed on each machine (such as 111, 112, 113), and the machine (such as 111, 112, 113) communicates with the background server 120 through the agent program.

[0073] The latest downtime detection rule library is stored in the yum source 130, and the downtime detection rule library is stored in the yum source in the form of an rpm software package. The downtime detection rule library includes multiple downtime detection rules indicating different preset kernel downtime causes. Each downtime detection rule includes abnormal memory status information corresponding to the corresponding preset kernel downtime cause and operating system repair logic. In practical applications, each downtime detection rule can be formed by a C language script, that is, the downtime detection rule library can be composed of several C language scripts.

[0074] The background server 120 can receive the kernel downtime diagnosis result sent by the agent program on the corresponding machine, and perform repair processing on the operating system of the machine based on the kernel downtime diagnosis result. In addition, the background server 120 also stores a local downtime detection rule library locally, and when the local downtime detection rule library changes, it pushes the changed full-scale local downtime detection rule library to the yum source 130 in the form of an rpm software package to update the downtime detection rule library on the yum source 130, so that the latest downtime detection rule library is always stored on the yum source.

[0075] It can be understood that Figure 1 The illustrated implementation environment is only an example, and in practical applications, it may also include other electronic devices that assist in implementing the technical solutions of the embodiments of the present application.

[0076] It should be noted that the server involved in the embodiments of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0077] In an exemplary embodiment, the machine cluster 110, the server 120, and the yum source 130 can all be node devices in a blockchain system, capable of sharing the information obtained and generated with other node devices in the blockchain system, so as to realize information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain, which is composed of multiple blocks, and adjacent blocks before and after have an association relationship, so that when the data in any block is tampered with, it can be detected by the next block, thereby avoiding the data in the blockchain from being tampered with and ensuring the security and reliability of the data in the blockchain.

[0078] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0079] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to realize the calculation, storage, processing, and sharing of data. It is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application. It can form a resource pool and be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture-based websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system backup support, which can only be achieved through cloud computing.

[0080] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for as needed. As a provider of basic cloud computing capabilities, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices. According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be deployed directly on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0081] See also Figure 2 , which is a flow chart of a kernel downtime diagnosis method provided by an embodiment of the present application, which can be applied to Figure 1 In the system architecture shown. It should be noted that this specification provides method operation steps as described in the embodiments or flow charts, but more or fewer operation steps may be included based on routine or non-creative work. The order of steps listed in the embodiments is only one way of executing the steps among many orders, and does not represent the only order of execution. When the actual system or product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or drawings (for example, in a parallel processor or multi-threaded processing environment). Specifically, Figure 2 As shown, the method may include:

[0082] S201, the agent program of the target machine obtains a kernel dump file, where the kernel dump file stores memory status information of the operating system of the target machine when the kernel crashes.

[0083] The target machine may be any machine in the Linux machine cluster where a kernel crash occurs.

[0084] Specifically, after the operating system of the target machine crashes, the kdump service will be automatically triggered to generate a Vmcore file, i.e., a kernel dump file. The Vmcore file contains memory state information such as the memory image, register status, and stack at the time of the operating system crash, which can be used to analyze the cause of the operating system crash.

[0085] After the target machine restarts, the agent program on the target machine will scan the Vmcore file, i.e., the kernel dump file, generated by the target machine in the case of a kernel crash through the crash tool, and then obtain the kernel dump file.

[0086] Among them, the kernel dump file can include memory state information of different memory state types, such as stack, register status, dmesg logs, etc.

[0087] S203, the agent program of the target machine loads the crash detection engine, and obtains the crash detection rule library from the yum source based on the crash detection engine.

[0088] Among them, the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons. Each crash detection rule includes the abnormal memory state information corresponding to the corresponding preset kernel crash reason and the operating system repair logic. It can be understood that the crash detection rule can also include other information, such as crash type, crash description information, etc.

[0089] In specific implementation, the crash detection rule can be a C script, that is, the crash detection rule library is composed of multiple C scripts. The crash detection rule library in the yum source can be formed by rpm software packages, so that when obtaining the crash detection rule library from the yum source, the latest crash detection rule library can be obtained from the yum source by automatically upgrading the rpm software package.

[0090] Among them, the crash detection engine can be an eppic engine.

[0091] S205, the agent program of the target machine generates a kernel crash diagnosis result based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule.

[0092] Among them, the kernel crash diagnosis result indicates whether a preset kernel crash reason is matched.

[0093] Specifically, if there is memory state information in the kernel dump file that matches any abnormal memory state information, a first kernel crash diagnosis result indicating that a preset kernel crash cause has been matched can be generated, and the preset kernel crash cause that has been matched is the preset kernel crash cause corresponding to the crash detection rule to which the matched abnormal memory state information belongs; conversely, if there is no memory state information in the kernel dump file that matches any abnormal memory state information, a second kernel crash diagnosis result indicating that no preset kernel crash cause has been matched can be generated.

[0094] S207, the proxy program of the target machine sends the kernel crash diagnosis result to the background server.

[0095] Specifically, in the case of generating the first kernel crash diagnosis result, the first kernel crash diagnosis result is sent to the background server; in the case of generating the second kernel crash diagnosis result, the second kernel crash diagnosis result is sent to the background server, so that the background server performs repair processing on the operating system of the target machine based on the kernel diagnosis result.

[0096] In some exemplary embodiments, as Figure 3 shown, the foregoing step S205 may include, when implemented:

[0097] S301, traverse the crash detection rules in the obtained crash detection rule library.

[0098] S303, for the currently traversed crash detection rule, obtain target memory state information corresponding to the memory state type of the abnormal memory state information corresponding to the currently traversed crash detection rule from the kernel dump file.

[0099] S305, match the target memory state information with the above abnormal memory state information.

[0100] Specifically, if the target memory state information is the same as the above abnormal memory state information, it indicates that the two match; conversely, if the target memory state information is different from the above abnormal memory state information, it indicates that the two do not match.

[0101] In the embodiments of the present application, when the target memory state information matches the above abnormal memory state information, the following steps S307 to S309 may be executed; when the target memory state information does not match the above abnormal memory state information, the following step S311 may be executed.

[0102] S307, determine the preset kernel crash cause corresponding to the currently traversed crash detection rule as the target kernel crash cause that causes the operating system to crash, and end the traversal of the crash detection rule library.

[0103] For example, the preset kernel crash cause corresponding to a crash detection rule in the crash detection rule library is the sysrq crash problem, and the abnormal memory state information it contains is: the stack contains the string feature "sysrq_handle_crash". Then, when traversing to this crash detection rule, it can be determined that the memory state type corresponding to its abnormal memory state information is the stack. Then, the stack information in the kernel dump file can be detected by function call to check whether the string feature "sysrq_handle_crash" is contained in the stack information. If the string feature is contained, it can be determined as the sysrq crash problem.

[0104] S309, generate a first kernel crash diagnosis result based on the current crash detection rule.

[0105] S311, determine whether there is an un-traversed crash detection rule in the crash detection rule library.

[0106] Specifically, if the judgment result is that there is an un-traversed crash detection rule in the crash detection rule library, then return to execute the foregoing steps S301 to S311; conversely, if the judgment result is that there is no un-traversed crash detection rule in the crash detection rule library, then the following step S313 can be executed.

[0107] S313, generate a second kernel crash diagnosis result based on the kernel dump file, and the second kernel crash diagnosis result indicates that no preset kernel crash cause is matched.

[0108] Specifically, the second kernel crash diagnosis result carries the memory state information in the kernel dump file.

[0109] The above implementation method traverses the crash detection rules in the crash detection rule library, and then determines whether the preset kernel crash cause is matched based on the matching situation between the abnormal memory state information in the traversed crash detection rule and the memory state information of the corresponding memory state type in the kernel dump file, improving the efficiency of kernel crash cause detection.

[0110] Correspondingly, the background server obtains the kernel crash diagnosis result sent by the target machine.

[0111] Specifically, if the first kernel crash diagnosis result is sent by the proxy program of the target machine, the kernel crash diagnosis result obtained by the background server is this first kernel crash diagnosis result; if the second kernel crash diagnosis result is sent by the proxy program of the target machine, the kernel crash diagnosis result obtained by the background server is this second kernel crash diagnosis result.

[0112] In some examples, after the background server obtains the kernel crash diagnosis result sent by the target machine, it can send the kernel crash diagnosis result to the kernel development node, and the kernel development node can repair the operation information of the target machine based on the kernel crash diagnosis result of the target machine.

[0113] The above technical solution of the embodiment of the present application realizes the automatic analysis and diagnosis of the kernel crash problem, improves the efficiency of the kernel crash diagnosis, and further improves the processing efficiency of the kernel crash file.

[0114] In some exemplary embodiments, continue to refer to Figure 3 , if the background server receives the second kernel crash diagnosis result indicating that no preset kernel crash reason is matched, it can perform the following steps S315 to S319:

[0115] S315, generate a kernel crash diagnosis work order based on the memory status information in the kernel dump file carried by the kernel crash diagnosis result, and send the kernel crash diagnosis work order to the kernel development node.

[0116] Specifically, if the background server receives the second crash diagnosis result that no preset kernel crash reason is matched, it can generate a kernel crash diagnosis work order based on the memory status information in the kernel dump file, and send the kernel crash diagnosis work order to the kernel development node. Thus, the kernel personnel of the kernel development node can analyze based on the memory status information in the kernel dump file to determine the kernel crash reason that causes the operating system to crash and the operating system repair logic for solving the kernel crash reason, and then can repair the operating system of the target machine based on the operating system repair logic.

[0117] In the embodiment of the present application, after the kernel development node determines the kernel crash reason that causes the operating system to crash and the operating system repair logic for solving the kernel crash reason, it can create a new crash detection rule based on this and send the new crash detection rule to the background server.

[0118] S317, the background server obtains the new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, and adds the new crash detection rule to the local crash detection rule library to obtain an updated local crash detection rule library.

[0119] In the embodiments of the present application, the background server locally stores a full set of downtime detection rule libraries, hereinafter referred to as the local downtime detection rule library. After the background server obtains the new downtime detection rules returned by the kernel development node for the kernel downtime diagnosis work order of the target machine, it adds the new downtime detection rules to the local downtime detection rule library, thereby updating the local downtime detection rule library. The updated local downtime detection rule library is hereinafter referred to as the updated local downtime rule library.

[0120] S319. The background server sends the above-mentioned updated local downtime detection rule library to the yum source to update the downtime detection rule library of the yum source.

[0121] Specifically, after the background server updates the local downtime detection rule library, it immediately sends the updated local downtime detection rule library to the yum source, thereby updating the downtime detection rule library on the yum source to the latest downtime detection rule library.

[0122] Since the latest downtime detection rule library includes the aforementioned new downtime detection rules, then the next time there is a machine operating system crash in the machine cluster caused by the preset kernel downtime reason corresponding to the new downtime detection rules, it can be quickly and automatically diagnosed based on the aforementioned steps S201 to S207, thereby improving the efficiency of kernel downtime diagnosis.

[0123] In some exemplary embodiments, in order to further improve the efficiency of kernel downtime diagnosis, step S319 may include when implemented:

[0124] Responding to the rpm package generation instruction for the updated local downtime detection rule library, generating an rpm software package based on the full set of the updated local downtime detection rule library;

[0125] Sending the rpm software package to the yum source.

[0126] Specifically, when it is necessary to send the updated local downtime detection rule library to the yum source, an rpm package generation instruction can be initiated for the updated local downtime detection rule library. Thus, the background server responds to the rpm package generation instruction, automatically generates an rpm software package from the full set of the updated local downtime detection rule library, and pushes the rpm software package to the yum source, thereby realizing the automatic update of the downtime detection rule library in the yum source by automatically upgrading the rpm software package.

[0127] It can be understood that when the background server detects other types of changes in the local downtime detection rule library, such as modifying the downtime detection rules, deleting one or more downtime detection rules, etc., the background server can send the updated local downtime detection rule library to the yum source based on the foregoing step S319 to update the downtime detection rule library of the yum source.

[0128] In some exemplary embodiments, to improve the efficiency of handling kernel downtime problems based on kernel downtime diagnosis results, such as Figure 4 As shown, in the foregoing step S309, when the agent program of the target machine generates the first kernel downtime diagnosis result based on the current downtime detection rule, it may include:

[0129] S401, obtain the target identification information of the current downtime detection rule.

[0130] Specifically, a corresponding identification information (ID) can be assigned to each downtime detection rule in the downtime detection rule library. This identification information can uniquely identify a downtime detection rule, and then the identification information of the matched current downtime detection rule can be obtained, hereinafter referred to as the target identification information.

[0131] S403, generate the first kernel downtime diagnosis result based on the target identification information.

[0132] Specifically, the first kernel downtime diagnosis result carries the target identification information. In practical applications, the first kernel downtime diagnosis result may also include the memory status information in the kernel dump file for subsequent review of the downtime detection rule corresponding to the target identification information to ensure the accuracy of the kernel downtime diagnosis result. Exemplarily, the format of the first kernel downtime diagnosis result may be: {ID: specific downtime rule ID, one ID corresponds to one downtime rule; panicmsg: output information of kernel panic; calltrace: function call stack; system: system information; subsystem: register and stack information}.

[0133] In the above embodiments, generating the first kernel downtime diagnosis result sent to the background server based on the identification information of the matched downtime detection rule can reduce the amount of data transmitted, improve the data transmission efficiency, and thus improve the diagnosis efficiency of kernel downtime.

[0134] Based on this, in some exemplary embodiments, continue to refer to Figure 4 , when the background server receives the above first kernel downtime diagnosis result indicating that a preset kernel downtime cause is matched, it may execute the following steps S405 to S407:

[0135] S405. Based on the target identification information carried in the first kernel crash diagnosis result, determine a matching target crash detection rule from the local crash detection rule library.

[0136] Specifically, the background server can parse the first kernel crash diagnosis result to obtain the target identification information, and then search for the crash detection rule corresponding to the target identification information from the local crash detection rule library, which is hereinafter referred to as the target crash detection rule.

[0137] S407. Obtain the operating system repair logic in the target crash detection rule, and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0138] Specifically, after the background server finds the target crash detection rule, it can obtain the corresponding operating system repair logic from it, and then send the operating system repair logic to the proxy program on the target machine, so that the proxy program on the target machine repairs the operating system of the target machine based on the operating system repair logic, thereby realizing the automatic repair of the operating system of the target machine and improving the processing efficiency of the kernel crash problem.

[0139] In some exemplary embodiments, in order to improve the flexibility of handling kernel crash problems, after the background server obtains the kernel crash diagnosis result sent by the target machine, the method may further include:

[0140] Generate a kernel crash notification message corresponding to the target machine based on the kernel crash diagnosis result;

[0141] Send the kernel crash notification message to the operation and maintenance node.

[0142] Specifically, when the kernel crash diagnosis result is the first kernel crash diagnosis result, the kernel crash notification message may carry the target preset kernel crash reason, so that the operation and maintenance node can know the reason for the crash of the operating system of the target machine when receiving the kernel crash notification message; when the kernel crash diagnosis result is the second kernel crash diagnosis result, the kernel crash notification message may prompt an unknown kernel crash reason, so that the operation and maintenance node can interact with the kernel development node in time when receiving the kernel crash notification message, which is conducive to improving the processing efficiency of the kernel crash problem.

[0143] In the above technical solution of the embodiment of the present application, when any machine in the machine cluster has a kernel crash, the kernel crash problem can be automatically analyzed and diagnosed. It is possible to realize the analysis and diagnosis of kernel crash faults such as MCE (Machine Check Exception), UAF (Use After Free), Double Free, Dead Lock, and memory leak, without manual analysis and troubleshooting, reducing labor costs and time costs, improving the efficiency of kernel crash diagnosis, facilitating the rapid elimination of kernel crash problems generated during the operation of the operating system, and ensuring the long-term stable operation of the operating system.

[0144] To facilitate the understanding of the technical solution of the embodiment of the present application, the following Figure 5 will describe the kernel crash diagnosis method of the embodiment of the present application.

[0145] As Figure 5 shown, after the machine system crashes, kdump starts to run, saves the operating system memory parameters and stack into a vmcore file, and then the machine restarts; after the machine restarts, it will pull up the crash tool through the agent to obtain the vmcore file, and then load the analysis engine. The analysis engine first obtains the latest crash detection rule library (i.e., multiple C language scripts) from the yum source by upgrading the rpm software package and loads it into memory, and then analyzes the symbol table of the vmcore file, matches it according to the crash detection rule library, and obtains the kernel crash diagnosis result; after that, the proxy program reports the kernel crash diagnosis result and some memory status information at the time of the crash to the background server in the cloud through the data channel.

[0146] After receiving the reported data, the background server in the cloud can store it in the database. In addition, for known kernel crash problems (i.e., hitting the crash detection rules in the crash detection rule library), the background server in the cloud can directly generate corresponding repair solutions for the operation and maintenance personnel to perform repair and upgrade. For unknown kernel crash problems (i.e., not hitting the crash detection rules in the crash detection rule library), a kernel crash diagnosis work order is generated, notifying the kernel development to analyze and solve the problem. After solving the problem, a new crash detection rule is generated and the local crash detection rule library is updated and published to the yum source to ensure that subsequent machines can load the latest crash detection rule library.

[0147] It should be noted that the background server in the cloud maintains a set of crash detection rule libraries. Kernel personnel can add, delete, and modify the crash detection rule libraries to change the content of the crash detection rule libraries. After the change, through the function of generating an rpm software package with one key, the full-scale local crash detection rule library can be automatically generated into an rpm software package and pushed to the yum source.

[0148] In addition, all kernel crash events can be notified to the corresponding operation and maintenance personnel through the mail system, which facilitates the operation and maintenance personnel to perform repairs in a timely manner. Among them, the corresponding operating system repair logic in the new crash detection rules will also be notified to the corresponding operation and maintenance personnel through the mail system, so that the operation and maintenance personnel can repair the operating system of the machine based on this operating system repair logic.

[0149] Corresponding to the kernel crash diagnosis methods provided in the above several embodiments, the embodiments of the present application also provide a kernel crash diagnosis device. Since the kernel crash diagnosis device provided in the embodiments of the present application corresponds to the kernel crash diagnosis methods provided in the above several embodiments, the implementation manners of the foregoing kernel crash diagnosis methods are also applicable to the kernel crash diagnosis device provided in this embodiment and will not be described in detail in this embodiment.

[0150] Please refer to Figure 6 , which shows a schematic structural diagram of a kernel crash diagnosis device provided in an embodiment of the present application. This device has the function of implementing the kernel crash diagnosis method on the Linux machine side in the above method embodiment. This function can be implemented by hardware or by hardware executing corresponding software. As Figure 6 shown, the kernel crash diagnosis device 600 may include:

[0151] A kernel dump file acquisition module 610, configured to acquire a kernel dump file, where the kernel dump file stores memory state information of the operating system of the target machine when the kernel crashes;

[0152] A crash rule library acquisition module 620, configured to load a crash detection engine and acquire a crash detection rule library from the yum source based on the crash detection engine; the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons, and each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash reason and an operating system repair logic;

[0153] A kernel crash diagnosis module 630, configured to generate a kernel crash diagnosis result based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule;

[0154] A diagnosis result sending module 640, configured to send the kernel crash diagnosis result to a background server, so that the background server repairs the operating system of the target machine based on the kernel diagnosis result.

[0155] In an exemplary implementation manner, the kernel crash diagnosis module 630 includes:

[0156] A rule library traversal module for traversing the crash detection rules in the obtained crash detection rule library;

[0157] A first acquisition module for, for the currently traversed crash detection rule, based on the memory state type of the abnormal memory state information corresponding to the current crash detection rule, acquiring target memory state information corresponding to the memory state type from the kernel dump file;

[0158] A kernel crash cause determination module for, when the target memory state information matches the abnormal memory state information, determining the preset kernel crash cause corresponding to the current crash detection rule as the target kernel crash cause that causes the operating system to crash, and ending the traversal of the crash detection rule library;

[0159] A first diagnosis result generation module for generating a first kernel crash diagnosis result based on the current crash detection rule.

[0160] In an exemplary embodiment, the background server stores the crash detection rule library; the first diagnosis result generation module includes:

[0161] An identification information acquisition module for acquiring the target identification information of the current crash detection rule;

[0162] A first diagnosis result generation sub-module for generating a first kernel crash diagnosis result based on the target identification information; when the background server performs a repair process on the operating system of the target machine based on the kernel diagnosis result, it is used to determine a matching target kernel crash detection rule from the local crash detection rule library based on the target identification information, and acquire the operating system repair logic in the target kernel crash detection rule, and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0163] In an exemplary embodiment, the kernel crash diagnosis module 630 further includes:

[0164] A loop execution module for, when the target memory state information does not match the abnormal memory state information and there are un-traversed crash detection rules in the crash detection rule library, returning to execute the step of traversing the crash detection rules in the crash detection rule library until there are no un-traversed crash detection rules in the crash detection rule library;

[0165] A second diagnosis result generation module, configured to generate a second kernel crash diagnosis result based on the kernel dump file, where the second kernel crash diagnosis result indicates that no preset kernel crash cause is matched; when the background server repairs the operating system of the target machine based on the kernel diagnosis result, it is configured to generate a kernel crash diagnosis work order based on the memory status information in the kernel dump file, send the kernel crash diagnosis work order to the kernel development node, obtain a new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, and send the operating system repair logic in the new crash detection rule to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0166] Please refer to Figure 7 , which shows a schematic structural diagram of another kernel crash diagnosis device provided by an embodiment of the present application. The device has the function of implementing the kernel crash diagnosis method on the background server side in the above method embodiment. This function can be implemented by hardware or by hardware executing corresponding software. As Figure 7 shown, the kernel crash diagnosis device 700 may include:

[0167] A diagnosis result acquisition module 710, configured to acquire a kernel crash diagnosis result sent by the target machine; the kernel crash diagnosis result is obtained by the target machine by acquiring a kernel dump file, loading a crash detection engine, and obtaining a crash detection rule library from a yum source based on the crash detection engine. The crash detection rule library includes multiple crash detection rules indicating different preset kernel crash causes. Each crash detection rule includes abnormal memory status information corresponding to the corresponding preset kernel crash cause and an operating system repair logic, and is generated based on the matching situation between the memory status information in the kernel dump file and the abnormal memory status information in each crash detection rule;

[0168] A crash diagnosis work order generation module 720, configured to generate a kernel crash diagnosis work order based on the memory status information in the kernel dump file carried by the kernel crash diagnosis result and send the kernel crash diagnosis work order to the kernel development node when the kernel crash diagnosis result indicates that no preset kernel crash cause is matched;

[0169] A local crash rule library update module 730, configured to acquire a new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, add the new crash detection rule to the local crash detection rule library to obtain an updated local crash detection rule library;

[0170] A yum source crash rule library update module 740, configured to send the updated local crash detection rule library to the yum source to update the crash detection rule library of the yum source.

[0171] In an exemplary embodiment, the yum source downtime rule library update module 740 includes:

[0172] An rpm package generation module, configured to generate an rpm software package based on the full updated local downtime detection rule library in response to an rpm package generation instruction for updating the local downtime detection rule library.

[0173] An rpm package push module, configured to send the rpm software package to the yum source.

[0174] In an exemplary embodiment, the apparatus 700 further includes:

[0175] A target downtime detection rule determination module, configured to determine a matching target downtime detection rule from the local downtime detection rule library based on the target identification information carried in the kernel downtime diagnosis result when the kernel downtime diagnosis result indicates a match with the preset kernel downtime cause.

[0176] An operating system repair module, configured to obtain the operating system repair logic in the target downtime detection rule and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

[0177] In an exemplary embodiment, the apparatus 700 further includes:

[0178] A notification message generation module, configured to generate a kernel downtime notification message corresponding to the target machine based on the kernel downtime diagnosis result.

[0179] A notification message sending module, configured to send the kernel downtime notification message to an operation and maintenance node.

[0180] It should be noted that, when implementing its functions, the apparatus provided in the above embodiments is only illustrated by dividing the above functional modules. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process thereof can be found in the method embodiments and will not be elaborated here.

[0181] An embodiment of the present application provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement any one of the kernel downtime diagnosis methods provided in the above method embodiments.

[0182] The memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0183] The method embodiments provided in the embodiments of the present application can be executed on a computer terminal, a server, or a similar computing device, that is, the above-mentioned electronic device can include a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 8 is a hardware structure block diagram of an electronic device that runs a kernel crash diagnosis method provided in the embodiments of the present application. As Figure 8 shown, the server 800 can vary greatly due to configuration or performance differences and can include one or more central processing units (CPUs) 810 (the processor 810 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 830 for storing data, and one or more storage media 820 for storing application programs 823 or data 822 (such as one or more mass storage devices). Among them, the memory 830 and the storage media 820 can be short-term storage or persistent storage. The programs stored in the storage media 820 can include one or more modules, and each module can include a series of instruction operations on the server. Further, the central processor 810 can be configured to communicate with the storage media 820 and execute a series of instruction operations in the storage media 820 on the server 800. The server 800 can also include one or more power supplies 860, one or more wired or wireless network interfaces 850, one or more input / output interfaces 840, and / or one or more operating systems 821, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0184] The input / output interface 840 can be used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by the communication provider of the server 800. In one example, the input / output interface 840 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the input / output interface 840 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0185] Those of ordinary skill in the art can understand that Figure 8 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the server 800 may further include more or fewer components than those shown Figure 8 shown, or have a different configuration from that Figure 8 shown.

[0186] Embodiments of the present application also provide a computer-readable storage medium, which can be arranged in an electronic device to store at least one instruction or at least one segment of a program related to implementing a method for repairing kernel crash. The at least one instruction or the at least one segment of the program is loaded and executed by the processor to implement any one of the kernel crash diagnosis methods provided by the above method embodiments.

[0187] Embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes any one of the kernel crash diagnosis methods provided by the above method embodiments.

[0188] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0189] It should be noted that: The above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the specific embodiments of the present specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0190] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments.

[0191] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0192] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A method for diagnosing kernel crashes, characterized in that, The method includes: Obtaining a kernel dump file, which stores the memory state information of the operating system of the target machine when the kernel crashes; Loading a crash detection engine, and obtaining a crash detection rule library from the yum source based on the crash detection engine; the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons, and each crash detection rule includes the abnormal memory state information corresponding to the corresponding preset kernel crash reason and the operating system repair logic; Generating a kernel crash diagnosis result based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule; Sending the kernel crash diagnosis result to a background server, so that the background server performs repair processing on the operating system of the target machine based on the kernel diagnosis result.

2. The method according to claim 1, wherein The generating a kernel crash diagnosis result based on the matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule includes: Traversing the crash detection rules in the obtained crash detection rule library; For the currently traversed crash detection rule, obtaining target memory state information corresponding to the memory state type from the kernel dump file based on the memory state type of the abnormal memory state information corresponding to the current crash detection rule; If the target memory state information matches the abnormal memory state information, determining the preset kernel crash reason corresponding to the current crash detection rule as the target kernel crash reason causing the operating system to crash, and ending the traversal of the crash detection rule library; Generating a first kernel crash diagnosis result based on the current crash detection rule.

3. The method according to claim 2, wherein The background server stores the crash detection rule library; the generating a first kernel crash diagnosis result based on the current crash detection rule includes: Obtaining the target identification information of the current crash detection rule; Generating a first kernel crash diagnosis result based on the target identification information; when the background server performs repair processing on the operating system of the target machine based on the kernel diagnosis result, it is used to determine a matching target kernel crash detection rule from the local crash detection rule library based on the target identification information, and obtain the operating system repair logic in the target kernel crash detection rule, and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

4. The method according to claim 2, characterized in that, The method further includes: If the target memory state information does not match the abnormal memory state information, and there are un-traversed crash detection rules in the crash detection rule library, then return to execute the step of traversing the crash detection rules in the crash detection rule library until there are no un-traversed crash detection rules in the crash detection rule library; Generate a second kernel crash diagnosis result based on the kernel dump file, where the second kernel crash diagnosis result indicates that no preset kernel crash cause is matched; when the background server repairs the operating system of the target machine based on the kernel diagnosis result, it is used to generate a kernel crash diagnosis work order based on the memory status information in the kernel dump file, send the kernel crash diagnosis work order to the kernel development node, obtain the new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, and send the operating system repair logic in the new crash detection rule to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

5. A kernel crash diagnosis method, characterized in that, The method includes: Obtain the kernel crash diagnosis result sent by the target machine; the kernel crash diagnosis result is obtained by the target machine by acquiring the kernel dump file, loading the crash detection engine, obtaining the crash detection rule library from the yum source based on the crash detection engine, where the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash causes, each crash detection rule includes the abnormal memory status information corresponding to the corresponding preset kernel crash cause and the operating system repair logic, and is generated based on the matching situation between the memory status information in the kernel dump file and the abnormal memory status information in each crash detection rule; In the case where the kernel crash diagnosis result indicates that no preset kernel crash cause is matched, generate a kernel crash diagnosis work order based on the memory status information in the kernel dump file carried by the kernel crash diagnosis result, and send the kernel crash diagnosis work order to the kernel development node; Obtain the new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, and add the new crash detection rule to the local crash detection rule library to obtain an updated local crash detection rule library; Send the updated local crash detection rule library to the yum source to update the crash detection rule library of the yum source.

6. The method according to claim 5, wherein The sending the updated local crash detection rule library to the yum source includes: In response to the rpm package generation instruction for the updated local crash detection rule library, generate an rpm software package based on the full amount of the updated local crash detection rule library; Send the rpm software package to the yum source.

7. The method according to claim 5, characterized in that The method further includes: In the case where the kernel crash diagnosis result indicates that a preset kernel crash cause is matched, determine the matching target crash detection rule from the local crash detection rule library based on the target identification information carried by the kernel crash diagnosis result; Obtain the operating system repair logic in the target crash detection rule, and send the operating system repair logic to the target machine so that the target machine repairs the operating system based on the operating system repair logic.

8. The method according to claim 5, wherein After obtaining the kernel crash diagnosis result sent by the target machine, the method further includes: Generate a kernel crash notification message corresponding to the target machine based on the kernel crash diagnosis result; Send the kernel crash notification message to the operation and maintenance node.

9. A kernel crash diagnosis device, characterized in that, The device includes: A kernel dump file acquisition module, configured to acquire a kernel dump file, where the kernel dump file stores memory state information of the operating system of the target machine when a kernel crash occurs; A crash rule library acquisition module, configured to load a crash detection engine and acquire a crash detection rule library from the yum source based on the crash detection engine; the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons, and each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash reason and an operating system repair logic; A kernel crash diagnosis module, configured to generate a kernel crash diagnosis result based on a matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule; A diagnosis result sending module, configured to send the kernel crash diagnosis result to a background server, so that the background server performs a repair process on the operating system of the target machine based on the kernel diagnosis result.

10. A kernel crash diagnosis device, characterized in that The device includes: A diagnosis result acquisition module, configured to acquire a kernel crash diagnosis result sent by the target machine; the kernel crash diagnosis result is obtained by the target machine by acquiring a kernel dump file, loading a crash detection engine, acquiring a crash detection rule library from the yum source based on the crash detection engine, where the crash detection rule library includes multiple crash detection rules indicating different preset kernel crash reasons, and each crash detection rule includes abnormal memory state information corresponding to the corresponding preset kernel crash reason and an operating system repair logic, and generating based on a matching situation between the memory state information in the kernel dump file and the abnormal memory state information in each crash detection rule; A crash diagnosis work order generation module, configured to generate a kernel crash diagnosis work order based on the memory state information in the kernel dump file carried in the kernel crash diagnosis result and send the kernel crash diagnosis work order to a kernel development node when the kernel crash diagnosis result indicates that no preset kernel crash reason is matched; A local crash rule library update module, configured to acquire a new crash detection rule returned by the kernel development node for the kernel crash diagnosis work order, and add the new crash detection rule to a local crash detection rule library to obtain an updated local crash detection rule library; A yum source crash rule library update module, configured to send the updated local crash detection rule library to the yum source to update the crash detection rule library of the yum source.

11. An electronic device, characterized in that, It includes a processor and a memory, where at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the kernel crash diagnosis method according to any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the kernel crash diagnosis method according to any one of claims 1 to 8.

13. A computer program, characterized in that, When the computer program is executed by a processor, it implements the kernel crash diagnosis method described in any one of claims 1 to 8.

Citation Information

Cited By

  • Operating system state root cause inference method, device, equipment, medium and product

    CN122450731A