Kernel downtime repairing method and device, electronic equipment and storage medium

By obtaining and executing the downtime analysis results of the target machine and automatically creating and executing repair tasks, the problem that Linux kernel downtime cannot be repaired in large quantities is solved, and efficient kernel downtime repair and system stability are achieved.

CN120276887APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410024612.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the existing technology, the downtime problem caused by Linux kernel errors cannot be automatically repaired in large quantities, resulting in low repair efficiency and the kernel downtime problem of operating system cannot be eliminated in time, affecting the stable operation of the system.

Method used

By obtaining the downtime analysis results of the target machine, determining the automatic repair permissions, and creating repair tasks based on the operating system repair logic, automatically performing repair processing, and obtaining repair log information to achieve large-scale machine kernel repair without human intervention.

Benefits of technology

Improve the efficiency of kernel downtime repair, ensure the long-term and stable operation of the operating system, and eliminate the kernel downtime problem in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276887A_ABST
    Figure CN120276887A_ABST
Patent Text Reader

Abstract

The invention discloses a kernel downtime repairing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a downtime analysis result reported by an agent program on a target machine, the downtime analysis result indicating an operating system repairing logic corresponding to a kernel downtime event; under the condition that the automatic repair permission information of the target machine indicates that automatic repair is authorized, creating a target operating system repair task based on the operating system repair logic and writing the target operating system repair task into a repair task queue; when the target operating system repair task is read from the repair task queue, calling a task execution service of the agent program to execute the target operating system repair task, so as to repair the operating system of the target machine based on the operating system repair logic; and in response to a received repair processing result indicating whether repair succeeds or not, calling a log service of the agent program to obtain repair log information generated in a repair processing process. According to the method and the device, the kernel downtime repairing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for repairing kernel crashes. Background Art

[0002] With the development of computer technology, the machine clusters of the Linux operating system are becoming increasingly large. The number of machines in the cluster has currently reached millions, and the number of crash events reaches thousands per day, among which a large number of crashes are caused by Linux kernel bugs.

[0003] In the related art, for the crash problems caused by Linux kernel errors, when repairing the crash problems, manual repair is mainly relied on, and the kernel repair of machines cannot be carried out in large quantities, resulting in low repair efficiency. Furthermore, the kernel crash problems generated during the operation of the Linux operating system cannot be eliminated in time, and the long-term stable operation of the operating system cannot be ensured. Summary of the Invention

[0004] To solve the problems of the prior art, embodiments of this application provide a method, apparatus, electronic device, and storage medium for repairing kernel crashes. The technical solutions are as follows:

[0005] On the one hand, a method for repairing kernel crashes is provided. The method includes:

[0006] Obtain the crash analysis result reported by the proxy program on the target machine for the kernel crash event of the target machine; the crash analysis result indicates the operating system repair logic corresponding to the kernel crash event;

[0007] Determine the automatic repair permission information corresponding to the target machine. When the automatic repair permission information indicates authorized automatic repair, create a target operating system repair task corresponding to the target machine based on the operating system repair logic;

[0008] Write the target operating system repair task into the repair task queue;

[0009] When the target operating system repair task is read from the repair task queue, call the task execution service of the proxy program to execute the target operating system repair task, so as to repair the operating system of the target machine based on the operating system repair logic;

[0010] In response to receiving the repair processing result returned by the proxy program, call the log service of the proxy program to obtain the repair log information generated during the repair processing; the repair processing result indicates whether the repair is successful.

[0011] On the other hand, a method for repairing kernel crashes is provided, and the method includes:

[0012] Obtain a call request from a repair server to the task execution service of the proxy program, where the call request carries a target operating system repair task; the target operating system repair task is created by the repair server based on the operating system repair logic when determining that the automatic repair permission information of the target machine indicates authorized automatic repair; the operating system repair logic is obtained based on the crash analysis result reported by the proxy program on the target machine for the kernel crash event of the target machine;

[0013] Execute the target operating system repair task through the task execution service to repair the operating system of the target machine based on the operating system repair logic;

[0014] When the execution of the target operating system repair task is completed, return a repair processing result to the repair server; the repair processing result indicates whether the repair is successful;

[0015] In response to a call request from the repair server to the log service of the proxy program, send the repair log information generated during the repair processing to the repair server through the log service.

[0016] On the other hand, a device for repairing kernel crashes is provided, and the device includes:

[0017] A crash analysis result acquisition module, configured to obtain a crash analysis result reported by a proxy program on a target machine for a kernel crash event of the target machine; the crash analysis result indicates an operating system repair logic corresponding to the kernel crash event;

[0018] A first repair task creation module, configured to determine the automatic repair permission information corresponding to the target machine, and create a target operating system repair task corresponding to the target machine based on the operating system repair logic when the automatic repair permission information indicates authorized automatic repair;

[0019] A repair task writing module, configured to write the target operating system repair task into a repair task queue;

[0020] A repair task distribution module, configured to, when reading the target operating system repair task from the repair task queue, call the task execution service of the proxy program to execute the target operating system repair task to repair the operating system of the target machine based on the operating system repair logic;

[0021] A repair log acquisition module, configured to, in response to receiving the repair processing result returned by the agent program, call the log service of the agent program to obtain the repair log information generated during the repair processing; the repair processing result indicates whether the repair is successful.

[0022] In an exemplary embodiment, the apparatus further includes:

[0023] A repair failure type determination module, configured to, when the repair processing result indicates that the repair fails, determine the failure type of the repair failure based on the repair log information

[0024] A retry task determination module, configured to generate a retry task for the target operating system repair task when the failure type is a preset failure type;

[0025] A retry module, configured to execute the retry task to write the target operating system repair task into the repair task queue again.

[0026] In an exemplary embodiment, the repair task writing module includes:

[0027] A task execution mode determination module, configured to determine the task execution mode corresponding to the target operating system repair task;

[0028] A timing time acquisition module, configured to, when the task execution mode is scheduled execution, acquire the timing time;

[0029] A first repair task writing sub-module, configured to write the target operating system repair task into the repair task queue when the current time matches the timing time.

[0030] In an exemplary embodiment, the repair task writing module further includes:

[0031] A second repair task writing sub-module, configured to, when the task execution mode is one-time execution, write the target operating system repair task into the repair task queue at the current time.

[0032] In an exemplary embodiment, the apparatus further includes:

[0033] A repair interface display module, configured to, when the automatic repair permission information indicates unauthorized automatic repair, display a repair interface through an associated terminal; the repair interface includes an automatic repair control;

[0034] A second repair task creation module, configured to, in response to an automatic repair request triggered based on the automatic repair control, create a target operating system repair task for the target machine based on the operating system repair logic.

[0035] In an exemplary embodiment, the downtime analysis result acquisition module includes:

[0036] A downtime analysis message acquisition module, configured to acquire a target downtime analysis message from a downtime message queue, where the target downtime analysis message is obtained by serializing a downtime analysis result corresponding to a kernel downtime event of the target machine by an agent program on the target machine;

[0037] A deserialization processing module, configured to perform deserialization processing on the target downtime analysis message to obtain a downtime analysis result corresponding to the kernel downtime event of the target machine.

[0038] On the other hand, a kernel downtime repair device is provided, and the device includes:

[0039] An execution service call request acquisition module, configured to acquire a call request from a repair server for a task execution service of an agent program, where the call request carries a target operating system repair task; the target operating system repair task is created by the repair server based on an operating system repair logic when determining that the automatic repair permission information of the target machine indicates authorized automatic repair; the operating system repair logic is obtained based on a downtime analysis result reported by the agent program on the target machine for the kernel downtime event of the target machine;

[0040] A repair task execution module, configured to execute the target operating system repair task through the task execution service to perform repair processing on the operating system of the target machine based on the operating system repair logic;

[0041] A repair processing result return module, configured to return a repair processing result to the repair server when the target operating system repair task is completed; the repair processing result indicates whether the repair is successful;

[0042] A repair log sending module, configured to, in response to a call request from the repair server for a log service of the agent program, send repair log information generated during the repair processing to the repair server through the log service.

[0043] In an exemplary embodiment, the repair task execution module includes:

[0044] A queue writing module, configured to write the target operating system repair task into a to-be-executed task queue corresponding to a task execution thread pool;

[0045] A thread execution module, configured to execute the target operating system repair task in the to-be-executed task queue based on the task execution thread pool, store the corresponding repair processing result in a memory queue, and write the repair log information generated during the repair processing into a task log file;

[0046] A repair processing result reporting module, configured to report the repair processing result in the memory queue to the repair server based on a callback thread.

[0047] In an exemplary embodiment, the repair log sending module is specifically configured to: read the task log file through the log service to obtain the repair log information, and send the read repair log information to the repair server.

[0048] On the other hand, an electronic device is provided, including a processor and a memory, where at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the kernel crash repair method in any of the above aspects.

[0049] On the other hand, a computer-readable storage medium is provided, where at least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the kernel crash repair method as described in any of the above aspects.

[0050] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the kernel crash repair method in any of the above aspects.

[0051] In the embodiment of the present application, the agent program on the target machine reports the crash analysis result for the kernel crash event on it, which indicates the operating system repair logic corresponding to the kernel crash event. The repair server obtains the crash analysis result and determines the automatic repair permission information corresponding to the target machine. When the automatic repair permission information indicates authorized automatic repair, the repair server creates a target operating system repair task for the target machine based on the operating system repair logic indicated by the crash analysis result, and then writes the target operating system repair task into the repair task queue. When the target operating system repair task is read from the repair task queue, the task execution service of the above agent program is called to execute the target operating system repair task, so as to repair the operating system of the target machine based on the operating system repair logic. In response to receiving the repair processing result returned by the agent program, the log service of the agent program is called to obtain the repair log information during the repair processing. Thus, automatic repair is issued according to the operating system repair logic, and the entire repair process does not require manual intervention. It can achieve the repair of a large number of machine kernels, greatly improving the repair efficiency. Furthermore, it can promptly eliminate the kernel crash problem generated during the operation of the operating system, ensuring the long-term stable operation of the operating system. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0054] Figure 2 is a schematic flowchart of a method for repairing a kernel crash provided by an embodiment of the present application;

[0055] Figure 3 is a schematic flowchart of another method for repairing a kernel crash provided by an embodiment of the present application;

[0056] Figure 4 is a schematic flowchart of another method for repairing a kernel crash provided by an embodiment of the present application;

[0057] Figure 5 is a schematic flowchart of another method for repairing a kernel crash provided by an embodiment of the present application;

[0058] Figure 6 is an architecture example for implementing a method for repairing a kernel crash provided by an embodiment of the present application;

[0059] Figure 7 It is a structural block diagram of a kernel crash repair device provided by an embodiment of the present application;

[0060] Figure 8 It is a structural block diagram of another kernel crash repair device provided by an embodiment of the present application;

[0061] Figure 9 It is a hardware structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0063] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used in appropriate cases can be interchanged so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0064] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0065] It can be understood that in the specific implementation manners of the present application, when it comes to data related to user information, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.

[0066] The terms involved in the embodiments of the present application are explained below.

[0067] Kernel: In computer technology, a kernel is a computer program used to manage data input / output (IO) requests sent by software, translate these requests into data processing instructions, and hand them over to the Central Processing Unit (CPU) and computer components for processing. It is the most basic part of an operating system.

[0068] System crash: It refers to the situation where the operating system cannot recover from a serious system error, or there are serious problems at the system hardware level, resulting in the system being unresponsive for a long time and having to restart the computer. A system crash will cause service interruption.

[0069] Kernel crash: It refers to a system crash caused by an error in kernel operation.

[0070] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by the embodiments of the present application. The implementation environment includes a machine cluster 110 and a repair server 120. Among them, each machine (such as 111, 112, 113) in the machine cluster 110 can be connected and communicate with the repair server 120 through a wired or wireless network.

[0071] Among them, the machines (such as 111, 112, 113) in the machine cluster 110 can be servers with a Linux operating system deployed in a computer room, or servers with a Linux operating system deployed in a multi-cloud environment, or various types of hosts such as physical machines and virtual machines.

[0072] In the embodiments of the present application, an agent program is installed on each machine (such as 111, 112, 113), and the machine (such as 111, 112, 113) communicates with the background repair server 120 through this agent program.

[0073] The repair server 120 can automatically repair the operating system of the corresponding machine based on the crash analysis results reported by the agent program for kernel crash events.

[0074] In some specific application scenarios, the repair server 120 may include a downtime analysis server and a repair task scheduling server. Among them, the downtime analysis server can obtain the downtime analysis results reported by the proxy programs on each machine in the case of kernel downtime, and the downtime analysis results can indicate the operating system repair logic for the corresponding kernel downtime event. In some examples, the operating system repair logic may be carried in the downtime analysis results, and thus the downtime analysis server can directly obtain the operating system repair logic corresponding to the kernel downtime event on the corresponding machine from the downtime analysis results; in other examples, the downtime analysis results may include downtime description information (for example, it may include the reason for downtime, the type of downtime, etc.), and the downtime analysis server can maintain the correspondence between the downtime description information and the operating system repair logic. Thus, the downtime analysis server can find the matching operating system repair logic from this correspondence based on the downtime description information in the downtime analysis results, and determine the matching operating system repair logic as the operating system repair logic for the kernel downtime event on the corresponding machine.

[0075] After determining the operating system repair logic corresponding to the kernel downtime event on the corresponding machine, the downtime analysis server can generate a downtime work order corresponding to the corresponding machine based on the operating system repair logic, and send the downtime work order to the repair task scheduling server.

[0076] After receiving the downtime work order corresponding to the corresponding machine, the repair task scheduling server can obtain the corresponding operating system repair logic based on the downtime work order, and create an operating system repair task corresponding to the corresponding machine based on the operating system repair logic. The operating system repair task instructs to perform the repair processing of the operating system based on the operating system repair logic. Thus, the repair task invocation server can send the operating system repair task to the proxy program on the corresponding machine, and the proxy program can repair the operating system of the corresponding machine by executing the operating system repair task and return the repair processing result to the repair task scheduling server. After receiving the repair processing result, the repair task scheduling server can obtain the repair log information generated during the corresponding repair processing from the proxy program for storage.

[0077] It can be understood that Figure 1 The illustrated implementation environment is only an example. In actual applications, there may also be other electronic devices that assist in implementing the technical solutions of the embodiments of the present application. For example, there may also be an interaction server, and the repair server interacts with each machine in the machine cluster through the interaction server.

[0078] It should be noted that the server involved in the embodiments of the present application may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0079] In an exemplary embodiment, both the machine cluster 110 and the repair server 120 can be node devices in a blockchain system, capable of sharing the information obtained and generated with other node devices in the blockchain system, so as to achieve information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain, which consists of multiple blocks, and adjacent blocks have an associated relationship, so that when the data in any block is tampered with, it can be detected by the next block, thereby avoiding the data in the blockchain from being tampered with and ensuring the security and reliability of the data in the blockchain.

[0080] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0081] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. It is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system backup support, which can only be achieved through cloud computing.

[0082] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely scalable and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices. Logically divided, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. The SaaS layer can also be directly deployed on the IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, mass SMS senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0083] Please refer to Figure 2 , which shows a schematic flowchart of a method for repairing kernel crashes provided by an embodiment of the present application. This method can be applied to Figure 1 the repair server in. It should be noted that this specification provides method operation steps as described in the embodiments or flowcharts, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or product is executed, it can be executed in the order of the method shown in the embodiments or the drawings, or executed in parallel (for example, in an environment with parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method may include:

[0084] S201, obtain the proxy program on the target machine and the crash analysis result reported for the kernel crash event of the target machine.

[0085] Among them, the crash analysis result indicates the operating system repair logic corresponding to the kernel crash event. The target machine can be any machine in the Linux machine cluster that experiences a kernel crash.

[0086] In some examples, the crash analysis result may carry corresponding operating system repair logic, which can be determined by the agent program on the target machine when detecting a kernel crash event. The repair server can directly obtain the operating system repair logic from the crash analysis result.

[0087] In other examples, the crash analysis result includes crash description information, such as the reason for the crash, the type of the crash, etc. This crash description information can be determined by the agent program on the target machine when detecting a kernel crash event. The repair server can maintain the correspondence between the crash description information and the operating system repair logic, and then determine the operating system repair logic that matches the crash description information in the crash analysis result from this correspondence, and determine the operating system repair logic corresponding to the kernel crash event of the target machine as the matching operating system repair logic.

[0088] In some possible implementation manners, when implementing the above step S201, it may include:

[0089] Obtain a target crash analysis message from the crash message queue. The target crash analysis message is obtained by serializing the crash analysis result corresponding to the kernel crash event of the target machine by the agent program on the target machine;

[0090] Deserialize the target crash analysis message to obtain the crash analysis result corresponding to the kernel crash event of the target machine.

[0091] Specifically, after determining the crash analysis result of the kernel crash event, the agent program on the target machine can serialize the crash analysis result in a preset serialization manner to obtain a target crash analysis message, and then send the target crash analysis message to the crash message queue. The repair server can read the target crash analysis message from the crash message queue, and then obtain the crash analysis result by deserializing the target crash analysis message.

[0092] Among them, the preset serialization manner can be the Protobuf manner. Protobuf uses a structured schema definition language to define the structure of the data to be serialized. Correspondingly, the Protobuf manner can be used for deserialization processing.

[0093] Among them, the crash message queue can be a distributed publish-subscribe message system, such as a kafka message queue.

[0094] Through the crash message queue and the serialization and deserialization processing of the crash analysis result, the above implementation manner can improve the information transmission efficiency between the agent program on the target machine and the repair server, and thus improve the repair efficiency of the kernel crash.

[0095] In some examples, when the agent program on the target machine reports the downtime analysis result, it can further consider the size of the data volume of the downtime analysis result. When the data volume is small, such as less than a set data volume threshold, it can be transmitted through the TCP (Transmission Control Protocol) data stream direct transmission mode. When the data volume is large, such as greater than the set data volume threshold, it can be transmitted through the BT (BitTorrent, a peer-to-peer transmission protocol) protocol. Among them, the data volume threshold can be set based on actual experience. For example, the data volume threshold can be 10KB. Then, when the data volume of the downtime analysis result is less than 10KB, it can be transmitted through the TCP data stream direct transmission mode. When the data volume of the downtime analysis result exceeds 10KB, it can be transmitted through the BT protocol.

[0096] S203. Determine the automatic repair permission information corresponding to the target machine. When the automatic repair permission information indicates authorized automatic repair, create a target operating system repair task corresponding to the target machine based on the above operating system repair logic.

[0097] Specifically, the repair server can maintain an automatic repair permission mapping relationship to represent whether the machines in the machine cluster are authorized for automatic repair. For example, the automatic repair permission information of each machine in the machine cluster can be stored. The automatic repair permission information can include authorized automatic repair or unauthorized automatic repair. For example, the value "0" indicates unauthorized automatic repair, and the value "1" indicates authorized automatic repair. Then, when determining the automatic repair permission information corresponding to the target machine, the automatic repair permission information corresponding to the target machine can be found from the automatic repair permission mapping relationship. For example, if the found automatic repair permission information is "1", it indicates that the target machine is authorized for automatic repair. On the contrary, if the found automatic repair permission information is "0", it indicates that the target machine is not authorized for automatic repair.

[0098] In some examples, in order to improve the flexibility of kernel downtime repair, the authorization valid time can also be stored. For example, if the value found from the automatic repair permission mapping relationship is "1" and the current time is within the authorization valid time, it can be determined that the target machine is authorized for automatic repair at the current time. On the contrary, if the value found from the automatic repair permission mapping relationship is "1", but the current time is not within the authorization valid time, it can be determined that the target machine is not authorized for automatic repair at the current time.

[0099] When the automatic repair permission information corresponding to the target machine of the repair server indicates authorized automatic repair, the repair server can determine the operating system repair logic based on the obtained downtime analysis result, and then create a target operating system repair task for the corresponding target machine based on the operating system repair logic. The target operating system repair task is a task indicating to repair the operating system of the target machine based on the operating system repair logic. Among them, the target operating system repair task can carry the operating system repair logic.

[0100] S205, write the target operating system repair task into the repair task queue.

[0101] Among them, the repair task queue is used to store the operating system repair tasks to be dispatched.

[0102] In some possible implementation manners, in order to improve the flexibility of kernel downtime repair, when implementing the above step S205, it may include:

[0103] Determine the task execution method corresponding to the target operating system repair task;

[0104] When the task execution method is scheduled execution, obtain the scheduled time;

[0105] When the current time matches the scheduled time, write the target operating system repair task into the repair task queue.

[0106] Among them, the scheduled time can be preset based on actual needs. For example, it can be set to 18:00 every day. It can be understood that the scheduled time can be a specific time point or a time range indicating a period of time. In practical applications, a scheduled time can be uniformly set for all operating system repair tasks to facilitate the management of kernel downtime repair. Of course, the scheduled time can also be set differently for one or more operating system repair tasks according to actual needs.

[0107] Specifically, after the repair server creates the target operating system repair task, it can determine the task execution method corresponding to the target operating system repair task. If the task execution method is scheduled execution, the corresponding scheduled time can be obtained, and then the current time is matched with the scheduled time. If the current time matches the scheduled time, for example, the scheduled time is 18:00 every day and the current time is 18:00 on xx year / xx month / xx day, it means that the current time matches the scheduled time. At this time, the target operating system repair task can be written into the repair task queue; otherwise, if the current time does not match the scheduled time, it enters the waiting-for-queueing stage until the current time matches the scheduled time, and then the target operating system repair task is written into the repair task queue.

[0108] In some other examples, if the task execution mode is one-time execution, the target operating system repair task can be written into the repair task queue at the current time.

[0109] Specifically, if the task execution mode corresponding to the target operating system repair task is one-time execution, after the repair server creates the target operating system repair task, it can write the target operating system repair task into the repair task queue.

[0110] S207, when the target operating system repair task is read from the repair task queue, call the task execution service of the agent program on the target machine to execute the target operating system repair task, so as to repair the operating system of the target machine based on the operating system repair logic.

[0111] Specifically, the repair server reads the operating system repair task to be dispatched from the repair task queue for dispatching. When the currently read task is the target operating system repair task, the repair server calls the task execution service of the agent program on the target machine to execute the target operating system repair task, so as to dispatch the target operating system repair task to the target machine for execution. In a specific implementation, the repair server can send a repair task execution request to the task execution service of the agent program on the target machine. The repair task execution request carries the target operating system repair task. The task execution service parses the repair task execution request to obtain the target operating system repair task, and then executes the target operating system repair task, so as to repair the operating system of the target machine through the operating system repair logic therein, and generate a repair processing result when the target operating system repair task is executed. The repair processing result can indicate whether the repair is successful. For example, it can be represented by the value "1" for successful repair and the value "0" for failed repair. It can be understood that the repair processing result can be returned to the repair server.

[0112] S209, in response to receiving the repair processing result returned by the agent program on the target machine, call the log service of the agent program to obtain the repair log information generated during the repair processing.

[0113] Among them, the repair processing result indicates whether the repair is successful.

[0114] Specifically, when the repair server receives the repair processing result returned by the agent program on the target machine, it can send a repair log acquisition request to the log service of the agent program. Then, the log service responds to the repair log acquisition request and returns the repair log information generated during the repair processing in the foregoing step S207 to the repair server. Correspondingly, the repair server obtains and stores the repair log information generated during the repair processing of the operating system of the target machine.

[0115] In some exemplary embodiments, to further improve the efficiency of kernel crash repair, in the case where the repair processing result corresponding to the target machine indicates a repair failure, such as Figure 3 shown, after step S207, the method may further include:

[0116] S301, based on the repair log information, determine the failure type of the repair failure.

[0117] Specifically, when the repair processing result of the target machine indicates a repair failure, the repair server may determine the failure type of the repair failure based on the corresponding repair log information, and this failure type may characterize the reason for the repair failure.

[0118] S303, if the failure type is a preset failure type, generate a retry task for the repair task of the target operating system.

[0119] Among them, the preset failure type can be set based on actual experience. For example, the preset failure type can be network interruption, interruption of task execution service on the machine, etc.

[0120] Among them, the retry task instructs to send the repair task of the target operating system to the agent program of the target machine again.

[0121] S305, execute the retry task to write the repair task of the target operating system into the repair task queue again.

[0122] Specifically, after generating a retry task for the repair task of the target operating system, the repair server may execute the retry task to write the repair task of the target operating system into the repair task queue again, and when reading the repair task of the target operating system from the repair task queue, it may call the task execution service of the agent program on the target machine again to execute the repair task of the target operating system, so as to perform a re - repair process on the operating system of the target machine based on the operating system repair logic, thereby enabling automatic retry after a repair failure and improving the repair efficiency of the kernel crash.

[0123] In some exemplary embodiments, to further improve the flexibility of kernel crash repair, such as Figure 4 shown, the method may further include:

[0124] S401, when the automatic repair permission information of the target machine indicates that automatic repair is not authorized, display a repair interface through an associated terminal, and the repair interface includes an automatic repair control.

[0125] S403, in response to an automatic repair instruction triggered based on the automatic repair control, create a repair task of the target operating system corresponding to the target machine based on the operating system repair logic.

[0126] Specifically, when the automatic repair permission information of the target machine indicates unauthorized automatic repair, the repair server can display a repair interface through the associated terminal. An automatic repair control is displayed on the repair interface. For example, the automatic repair control is "One-click Repair". If the user triggers (such as clicking or long-pressing) the automatic repair control, an automatic repair instruction is issued. After receiving the automatic repair instruction, the associated terminal sends an automatic repair request triggered by the automatic repair control to the repair server. In response to the automatic repair request, the repair server creates a target operating system repair task for the target machine based on the aforementioned operating system repair logic.

[0127] As can be seen from the above technical solutions of the embodiments of the present application, in the embodiments of the present application, the repair server obtains the downtime analysis result, determines the automatic repair permission information corresponding to the target machine, and when the automatic repair permission information indicates authorized automatic repair, creates a target operating system repair task for the target machine based on the operating system repair logic indicated by the downtime analysis result. Then, the target operating system repair task is written into the repair task queue. When the target operating system repair task is read from the repair task queue, the task execution service of the aforementioned proxy program is called to execute the target operating system repair task, so as to repair the operating system of the target machine based on the operating system repair logic, and in response to receiving the repair processing result returned by the proxy program, the log service of the proxy program is called to obtain the repair log information during the repair processing. Thus, automatic repair is issued according to the operating system repair logic, and the entire repair process does not require manual intervention, enabling the repair of a large number of machine kernels, greatly improving the repair efficiency, and then timely eliminating the kernel downtime problem generated during the operation of the operating system, ensuring the long-term stable operation of the operating system.

[0128] Please refer to Figure 5 , which shows a schematic flowchart of another method for repairing kernel downtime provided by the embodiments of the present application. This method takes the target machine as the execution subject and specifically includes:

[0129] S501, obtain a call request from the repair server to the task execution service of the proxy program on the target machine.

[0130] Among them, the call request carries a target operating system repair task, which is created by the repair server based on the operating system repair logic when determining that the automatic repair permission information of the target machine indicates authorized automatic repair. The operating system repair logic is obtained based on the downtime analysis result reported by the proxy program on the target machine for the kernel downtime event of the target machine.

[0131] S503. Execute the target operating system repair task through the task execution service to repair the operating system of the target machine based on the operating system repair logic.

[0132] Exemplarily, executing the target operating system repair task through the task execution service can be to write the target operating system repair task into the to-be-executed task queue corresponding to the task execution thread pool, then execute the target operating system repair task in the to-be-executed task queue based on this task execution thread pool, store the corresponding repair result in the memory queue, write the repair log information generated during the repair process into the task log file, and then report the repair result in the memory queue to the repair server based on the callback thread, thereby improving the execution efficiency of the target operating system repair task and further improving the repair efficiency of the kernel crash.

[0133] S505. When the target operating system repair task is completed, return the repair result to the repair server, and the repair result indicates whether the repair is successful.

[0134] Specifically, if the operating system of the target machine is repaired successfully, a repair result indicating success is generated; conversely, if the operating system of the target machine is repaired failed, a repair result indicating failure is generated.

[0135] S507. In response to the repair server's call request for the log service of the agent program, send the repair log information generated during the repair process to the repair server through this log service.

[0136] Specifically, the repair log information can be obtained by reading the task log file through the log service and sending the read repair log information to the repair server.

[0137] The above implementation manner obtains the repair server's call request for the task execution service of the agent program, executes the target operating system repair task issued based on this call request through this task execution service to repair the operating system of the target machine based on the operating system repair logic in the target operating system repair task, and returns the repair result indicating whether the repair is successful, and in response to the repair server's call request for the log service of the agent program, sends the repair log information generated during the repair process to the repair server through this log service, thereby realizing the automatic repair of the memory crash. The entire repair process requires no manual intervention, greatly improving the repair efficiency, and further can eliminate the kernel crash problem generated during the operation of the operating system in a timely manner, ensuring the long-term stable operation of the operating system.

[0138] To facilitate the understanding of the technical solutions of the embodiments of the present application, the following Figure 6 describes the kernel crash repair method of the embodiments of the present application.

[0139] As Figure 6 shown, an agent program is installed on each Linux machine, and the Linux machine interacts with the cloud server through this agent program. Specifically, the interaction can be carried out through the Restful API method.

[0140] The interaction server serves as the underlying interaction channel and mainly provides three types of service capabilities: file service, naming service, and data service. Among them, for small files, the file service is transmitted through the TCP data stream direct transmission mode, and for large files, it is transmitted through the BT protocol. Files smaller than 10KB can be regarded as small files, and files larger than 10KB can be regarded as large files; the data service mainly receives the data reported by the agent through the kafka message queue and transmits the data in the protobuf manner; the command service is used to implement the repair task scheduling server to issue tasks to the agent program through the Restful API method.

[0141] Each time the Linux machine restarts, the Agent program scans whether there is a kernel dump file generated due to downtime on this machine. If so, it starts the crash tool for analysis, generates a downtime analysis result, and reports it to the downtime analysis server.

[0142] The downtime analysis server manages and maintains the correspondence between the downtime description information and the operating system repair logic. After reading the reported information, the downtime analysis server determines the corresponding operating system repair logic by comparing this correspondence, and then generates a corresponding downtime work order based on this. In specific implementation, the above correspondence maintained by the downtime analysis server can be downtime rules. The main information of each downtime rule includes: downtime type, affected Linux kernel version, operating system repair logic, downtime description information, downtime cause, etc. The ID of the matched downtime rule can be carried in the downtime analysis result. Then, the downtime analysis server can determine the target downtime rule from the multiple downtime rules it maintains based on this ID, and then obtain the operating system repair logic from this target downtime rule. For example, the downtime analysis result can include the following fields: {ID: the specific downtime rule ID, one ID corresponds to one downtime rule; panicmsg: the output information of the kernel panic; calltrace: the function call stack; system: system information; subsystem: register and stack information}.

[0143] Each downtime work order will create a corresponding operating system repair task through the repair task scheduling server, and the repair task scheduling server will issue subsequent tasks to repair the downtime problem on the corresponding Linux machine.

[0144] Corresponding to the kernel crash repair methods provided in the above several embodiments, an embodiment of the present application also provides a kernel crash repair device. Since the kernel crash repair device provided in the embodiment of the present application corresponds to the kernel crash repair methods provided in the above several embodiments, the implementation manners of the foregoing kernel crash repair methods are also applicable to the kernel crash repair device provided in this embodiment and will not be described in detail in this embodiment.

[0145] Please refer to Figure 7 , which shows a schematic structural diagram of a kernel crash repair device provided in an embodiment of the present application. The device has the function of implementing the kernel crash repair method on the repair server side in the above method embodiment. The function can be implemented by hardware or by hardware executing corresponding software. As Figure 7 shown, the kernel crash repair device 700 may include:

[0146] A crash analysis result acquisition module 710, configured to acquire a proxy program on a target machine and obtain a crash analysis result reported for a kernel crash event of the target machine; the crash analysis result indicates an operating system repair logic corresponding to the kernel crash event;

[0147] A first repair task creation module 720, configured to determine automatic repair permission information corresponding to the target machine, and create a target operating system repair task corresponding to the target machine based on the operating system repair logic when the automatic repair permission information indicates authorization for automatic repair;

[0148] A repair task writing module 730, configured to write the target operating system repair task into a repair task queue;

[0149] A repair task distribution module 740, configured to, when reading the target operating system repair task from the repair task queue, call a task execution service of the proxy program to execute the target operating system repair task, so as to perform a repair process on the operating system of the target machine based on the operating system repair logic;

[0150] A repair log acquisition module 750, configured to, in response to receiving a repair processing result returned by the proxy program, call a log service of the proxy program to obtain repair log information generated during the repair processing; the repair processing result indicates whether the repair is successful.

[0151] In an exemplary implementation manner, the device further includes:

[0152] A repair failure type determination module, configured to determine a failure type of the repair failure based on the repair log information when the repair processing result indicates that the repair fails

[0153] A retry task determination module, configured to generate a retry task for the target operating system repair task when the failure type is a preset failure type;

[0154] A retry module, configured to execute the retry task to write the target operating system repair task into the repair task queue again.

[0155] In an exemplary embodiment, the repair task writing module 730 includes:

[0156] A task execution mode determination module, configured to determine the task execution mode corresponding to the target operating system repair task;

[0157] A timing time acquisition module, configured to acquire the timing time when the task execution mode is scheduled execution;

[0158] A first repair task writing sub-module, configured to write the target operating system repair task into the repair task queue when the current time matches the timing time.

[0159] In an exemplary embodiment, the repair task writing module 730 further includes:

[0160] A second repair task writing sub-module, configured to write the target operating system repair task into the repair task queue at the current time when the task execution mode is one-time execution.

[0161] In an exemplary embodiment, the device further includes:

[0162] A repair interface display module, configured to display a repair interface through an associated terminal when the automatic repair permission information indicates unauthorized automatic repair; the repair interface includes an automatic repair control;

[0163] A second repair task creation module, configured to create a target operating system repair task for the target machine based on the operating system repair logic in response to an automatic repair request triggered based on the automatic repair control.

[0164] In an exemplary embodiment, the crash analysis result acquisition module 710 includes:

[0165] A crash analysis message acquisition module, configured to acquire a target crash analysis message from the crash message queue, where the target crash analysis message is obtained by serializing the crash analysis result corresponding to the kernel crash event of the target machine by an agent program on the target machine;

[0166] A deserialization processing module, configured to perform deserialization processing on the target crash analysis message to obtain the crash analysis result corresponding to the kernel crash event of the target machine.

[0167] Please refer to Figure 8 , which shows a schematic structural diagram of another kernel crash repair device provided by an embodiment of the present application. This device has the function of implementing the kernel crash repair method on the target machine side in the above method embodiment. This function can be implemented by hardware or by hardware executing corresponding software. As Figure 8 shown, the kernel crash repair device 800 may include:

[0168] A service call request acquisition module 810, configured to acquire a call request from a repair server for the task execution service of the proxy program. This call request carries a target operating system repair task. The target operating system repair task is created by the repair server based on the operating system repair logic when determining that the automatic repair permission information of the target machine indicates authorized automatic repair. The operating system repair logic is obtained based on the crash analysis result reported for the kernel crash event of the target machine by the proxy program on the target machine;

[0169] A repair task execution module 820, configured to execute the target operating system repair task through the task execution service, so as to perform repair processing on the operating system of the target machine based on the operating system repair logic;

[0170] A repair processing result return module 830, configured to return a repair processing result to the repair server when the target operating system repair task is completed. The repair processing result indicates whether the repair is successful;

[0171] A repair log sending module 840, configured to send the repair log information generated during the repair processing to the repair server through the log service in response to a call request from the repair server for the log service of the proxy program.

[0172] In an exemplary embodiment, the repair task execution module 820 includes:

[0173] A queue writing module, configured to write the target operating system repair task into the to-be-executed task queue corresponding to the task execution thread pool;

[0174] A thread execution module, configured to execute the target operating system repair task in the to-be-executed task queue based on the task execution thread pool, store the corresponding repair processing result in a memory queue, and write the repair log information generated during the repair processing into a task log file;

[0175] A repair processing result reporting module, configured to report the repair processing result in the memory queue to the repair server based on a callback thread.

[0176] In an exemplary embodiment, the repair log sending module 840 is specifically configured to: read the task log file through the log service to obtain the repair log information, and send the read repair log information to the repair server.

[0177] It should be noted that for the device provided in the above embodiment, when implementing its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process can be found in the method embodiment, which will not be elaborated here.

[0178] An embodiment of the present application provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement any one of the kernel crash repair methods provided in the above method embodiments.

[0179] The memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0180] The method embodiments provided in the embodiments of the present application can be executed on a computer terminal, a server or a similar computing device, that is, the above electronic device can include a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 9 is a hardware structure block diagram of an electronic device for running a kernel crash repair method provided in an embodiment of the present application, as Figure 9As shown, the server 900 can vary significantly due to different configurations or performances, and may include one or more central processing units (CPUs) 910 (the processor 910 may include, but is not limited to, processing devices such as a microprocessor MCU or a field programmable gate array FPGA), a memory 930 for storing data, and one or more storage media 920 (such as one or more mass storage devices) for storing application programs 923 or data 922. Among them, the memory 930 and the storage media 920 can be transient storage or persistent storage. The program stored in the storage media 920 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processor 910 can be configured to communicate with the storage media 920 and execute a series of instruction operations in the storage media 920 on the server 900. The server 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0181] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the server 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 940 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0182] Those of ordinary skill in the art can understand that Figure 9 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the server 900 may also include more or fewer components than Figure 9 shown therein, or have a different configuration from Figure 9 that shown.

[0183] An embodiment of the present application also provides a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one segment of program related to implementing a method for repairing kernel crash. The at least one instruction or the at least one segment of program is loaded and executed by the processor to implement any one of the methods for repairing kernel crash provided by the above method embodiments.

[0184] Embodiments of the present application further provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes any one of the kernel crash repair methods provided in the above method embodiments.

[0185] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0186] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0187] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0188] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0189] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for repairing kernel crashes, characterized in that The method includes: Obtaining the crash analysis result reported by the proxy program on the target machine for the kernel crash event of the target machine; the crash analysis result indicates the operating system repair logic corresponding to the kernel crash event. Determining the automatic repair permission information corresponding to the target machine, and when the automatic repair permission information indicates authorization for automatic repair, creating a target operating system repair task for the target machine based on the operating system repair logic. Writing the target operating system repair task into the repair task queue. When reading the target operating system repair task from the repair task queue, calling the task execution service of the proxy program to execute the target operating system repair task, so as to repair the operating system of the target machine based on the operating system repair logic. In response to receiving the repair processing result returned by the proxy program, calling the log service of the proxy program to obtain the repair log information generated during the repair processing; the repair processing result indicates whether the repair is successful.

2. The method according to claim 1, characterized in that When the repair processing result indicates that the repair fails, the method further includes: Determining the failure type of the repair failure based on the repair log information. If the failure type is a preset failure type, generating a retry task for the target operating system repair task. Executing the retry task to write the target operating system repair task into the repair task queue again.

3. The method according to claim 1, wherein The writing the target operating system repair task into the repair task queue includes: Determining the task execution mode corresponding to the target operating system repair task. When the task execution mode is scheduled execution, obtaining the scheduled time. When the current time matches the scheduled time, writing the target operating system repair task into the repair task queue.

4. The method according to claim 3, wherein The method further includes: When the task execution mode is one-time execution, writing the target operating system repair task into the repair task queue at the current time.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: When the automatic repair permission information indicates that automatic repair is not authorized, displaying a repair interface through an associated terminal; the repair interface includes an automatic repair control. In response to an automatic repair request triggered based on the automatic repair control, creating a target operating system repair task for the target machine based on the operating system repair logic.

6. The method according to claim 1, characterized in that The obtaining the crash analysis result reported by the proxy program on the target machine for the kernel crash event of the target machine includes: Obtaining a target crash analysis message from the crash message queue, where the target crash analysis message is obtained by serializing the crash analysis result corresponding to the kernel crash event of the target machine by the proxy program on the target machine. Performing deserialization processing on the target crash analysis message to obtain the crash analysis result corresponding to the kernel crash event of the target machine.

7. A method for repairing kernel crash, characterized in that, The method includes: Obtain a call request from the repair server for the task execution service of the proxy program, where the call request carries a target operating system repair task; the target operating system repair task is created by the repair server based on the operating system repair logic when determining that the automatic repair permission information of the target machine indicates authorized automatic repair; the operating system repair logic is obtained based on the crash analysis result reported by the proxy program on the target machine for the kernel crash event of the target machine; Execute the target operating system repair task through the task execution service to perform repair processing on the operating system of the target machine based on the operating system repair logic; When the target operating system repair task is completed, return a repair processing result to the repair server; the repair processing result indicates whether the repair is successful; In response to the call request from the repair server for the log service of the proxy program, send the repair log information generated during the repair processing to the repair server through the log service.

8. The method according to claim 7, characterized in that, The executing the target operating system repair task through the task execution service includes: Write the target operating system repair task into the to-be-executed task queue corresponding to the task execution thread pool; Execute the target operating system repair task in the to-be-executed task queue based on the task execution thread pool, store the corresponding repair processing result in the memory queue, and write the repair log information generated during the repair processing into the task log file; Report the repair processing result in the memory queue to the repair server based on the callback thread.

9. The method according to claim 8, characterized in that The sending the repair log information generated during the repair processing to the repair server through the log service includes: Read the task log file through the log service to obtain the repair log information, and send the read repair log information to the repair server.

10. A repair device for kernel crash, characterized in that, The device includes: A crash analysis result acquisition module, configured to obtain the crash analysis result reported by the proxy program on the target machine for the kernel crash event of the target machine; the crash analysis result indicates the operating system repair logic corresponding to the kernel crash event; A repair task generation module, configured to determine the automatic repair permission information corresponding to the target machine, and create a target operating system repair task corresponding to the target machine based on the operating system repair logic when the automatic repair permission information indicates authorized automatic repair; A repair task writing module, configured to write the target operating system repair task into the repair task queue; A repair task distribution module, configured to, when the target operating system repair task is read from the repair task queue, call the task execution service of the proxy program to execute the target operating system repair task to perform repair processing on the operating system of the target machine based on the operating system repair logic; A repair log acquisition module, configured to, in response to receiving a repair processing result returned by the agent program, call the log service of the agent program to obtain repair log information generated during the repair processing; the repair processing result indicates whether the repair is successful.

11. A repair device for kernel crash, characterized in that, The device includes: An execution service call request acquisition module, configured to obtain a call request from a repair server for the task execution service of the agent program, the call request carrying a target operating system repair task; the target operating system repair task is created by the repair server based on an operating system repair logic when determining that the automatic repair permission information of the target machine indicates authorized automatic repair; the operating system repair logic is obtained based on the crash analysis result reported by the agent program on the target machine for a kernel crash event of the target machine; A repair task execution module, configured to execute the target operating system repair task through the task execution service to repair the operating system of the target machine based on the operating system repair logic; A repair processing result return module, configured to, when the execution of the target operating system repair task is completed, return a repair processing result to the repair server; the repair processing result indicates whether the repair is successful; A repair log sending module, configured to, in response to a call request from the repair server for the log service of the agent program, send the repair log information generated during the repair processing to the repair server through the log service.

12. An electronic device, characterized in that, It includes a processor and a memory, and at least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the method for repairing a kernel crash according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that, At least one instruction or at least one segment of program is stored in the computer-readable storage medium, and the at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the method for repairing a kernel crash according to any one of claims 1 to 9.

14. A computer program, characterized in that, The computer program, when executed by a processor, implements the method for repairing a kernel crash according to any one of claims 1 to 9.