An eBPF-based operating system cluster monitoring method and an electronic device

By using eBPF-based monitoring methods in the operating system cluster, the monitoring program is generated and injected, and the problems of poor monitoring universality and low efficiency in the prior art are solved, and efficient and accurate operating system cluster monitoring and fault information collection are achieved.

CN115827370BActive Publication Date: 2025-06-24INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211372386.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-06-24
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

The existing operating system cluster monitoring methods are poorly versatile and inefficient, making it difficult to effectively collect and analyze operating system failure information in the cluster.

Method used

The eBPF-based operating system cluster monitoring method is adopted, and the monitoring program is generated for the operating system by displaying the monitoring page on the user side, and the eBPF mechanism is used to inject the program into the core of the business server to monitor and collect operating system data in real time.

Benefits of technology

It improves the accuracy and efficiency of operating system cluster monitoring, and can deploy different data acquisition points and analysis sources according to actual needs, accurately determine the source of the problem, thereby saving operation and maintenance time and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115827370B_ABST
    Figure CN115827370B_ABST
Patent Text Reader

Abstract

The present application discloses an eBPF-based operating system cluster monitoring method and an electronic device, belonging to the technical field of device monitoring. Among them, the method includes displaying a monitoring page to the user side; generating a first monitoring program for the operating system according to the first input of the user on the monitoring page; injecting the first monitoring program into the kernels of each business server based on the eBPF mechanism for each business server to monitor its own operating system according to the first monitoring program; receiving the monitoring data returned by each business server according to the first monitoring program. Through the monitoring page, the present application can deploy different data collection points and analysis sources according to actual needs, accurately determine the source of problems, improve the ability of operating system cluster monitoring and collecting fault information, thereby saving the operation and maintenance time and operation and maintenance costs of the operating system cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of device monitoring, and particularly relates to a method for monitoring an operating system cluster based on eBPF and an electronic device. Background Art

[0002] With the continuous development of information technology, the operation and maintenance of operating system clusters have become increasingly important.

[0003] Currently, when an operating system in the cluster encounters an anomaly, it is necessary to log in to the background of each operating system to add kernel modules and recompile kernel code to collect the information required for troubleshooting. The above methods are easily restricted by customer resources, with poor generality and low efficiency. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method for monitoring an operating system cluster based on eBPF and an electronic device, which can solve the problems of poor generality and low efficiency in the existing methods for monitoring operating system clusters.

[0005] To solve the above technical problems, this application is implemented as follows:

[0006] In a first aspect, the embodiments of this application provide a method for monitoring an operating system cluster based on eBPF, which is applied to a management server. The management server establishes a server cluster with multiple business servers. The method includes:

[0007] Display a monitoring page to the user side;

[0008] Generate a first monitoring program for the operating system according to the first input of the user on the monitoring page;

[0009] Based on the eBPF mechanism, inject the first monitoring program into the kernels of each business server for each business server to monitor its own operating system according to the first monitoring program;

[0010] Receive the monitoring data returned by each business server according to the first monitoring program.

[0011] Optionally, in the method for monitoring an operating system cluster based on eBPF, generating a first monitoring program for the operating system according to the first input of the user on the monitoring page includes:

[0012] Define a return data structure, user-space monitoring logic, and kernel-space monitoring logic according to the first input;

[0013] Generate the first monitoring program according to the return data structure, user-space monitoring logic, and kernel-space monitoring logic.

[0014] Optionally, before generating the first monitoring program according to the user's first input on the monitoring page, the method further includes:

[0015] On the monitoring page, display the supported kernel modules, the function list corresponding to the kernel modules, and the supported mounting methods;

[0016] Receive the user's second input on the monitoring page, and determine the target kernel module, the target function, and the target mounting method;

[0017] Generate a program template in the format of the BCC compilation chain according to the target kernel module, the target function, and the target mounting method;

[0018] The first input includes a first sub-input and a second sub-input. Generating the first monitoring program according to the user's first input on the monitoring page includes:

[0019] Receive the user's first sub-input for the program template;

[0020] In response to the first sub-input, define the return data structure, the user-space monitoring logic, and the kernel-space monitoring logic;

[0021] Receive the user's second sub-input for the monitoring page;

[0022] In response to the second sub-input, generate the first monitoring program based on the defined return data structure, the user-space monitoring logic, and the kernel-space monitoring logic.

[0023] Optionally, before injecting the first monitoring program into the kernels of each business server based on the eBPF mechanism, the method further includes:

[0024] Receive a third input for the first monitoring program;

[0025] In response to the third input, start verifying the first monitoring program;

[0026] In the case where the first monitoring program starts successfully, execute the step of injecting the first monitoring program into the kernels of each business server based on the eBPF mechanism.

[0027] Optionally, after injecting the first monitoring program into the kernels of each business server, the method further includes:

[0028] Based on the eBPF mechanism, send a start command to each business server to control the start of the first monitoring program.

[0029] Optionally, after displaying the monitoring page, the method further includes:

[0030] Receive a fourth input for the second monitoring program in the monitoring page;

[0031] In response to the fourth input, based on the eBPF mechanism, send an uninstall command to each business server to control the uninstallation of the second monitoring program.

[0032] Optionally, the method further includes:

[0033] Store the monitoring data in a preset format and provide a data query interface.

[0034] Optionally, in the eBPF-based operating system cluster monitoring method, the monitoring data includes program startup error information;

[0035] After receiving the monitoring data returned by each business server, the method further includes:

[0036] Display the program startup error information, which is returned by the business server to the management server by logging when the first monitoring program fails to start.

[0037] Optionally, in the eBPF-based operating system cluster monitoring method, the monitoring data includes node exception information;

[0038] After receiving the monitoring data returned by each business server, the method further includes:

[0039] Display the node exception information, which is returned by the business server to the management server when an exception occurs.

[0040] Optionally, in the eBPF-based operating system cluster monitoring method, the monitoring data includes heartbeat information;

[0041] After receiving the monitoring data returned by each business server, the method further includes:

[0042] Obtain the heartbeat information sent by each business server;

[0043] If the heartbeat information of a first business server not included in the list of live servers is received, add the first business server to the list of live servers;

[0044] If the heartbeat information of a second business server is not received within a preset duration, delete the second business server from the list of live servers.

[0045] In a second aspect, an embodiment of the present application provides an eBPF-based operating system cluster monitoring device, which is applied to a management server, and the management server establishes a server cluster with multiple business servers. The device includes:

[0046] A page display module, configured to display a monitoring page to a user terminal;

[0047] A program construction module, configured to generate a first monitoring program for an operating system according to a first input of a user on the monitoring page;

[0048] A program injection module, configured to inject the first monitoring program into the kernels of respective business servers based on the eBPF mechanism, so that each of the business servers monitors its own operating system according to the first monitoring program;

[0049] A receiving module, configured to receive monitoring data returned by each of the business servers according to the first monitoring program.

[0050] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0051] In a fourth aspect, an embodiment of the present application provides a non-volatile readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0052] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0053] In the embodiment of the present application, first, a monitoring page is displayed to a user terminal; then, according to a first input of the user on the monitoring page, a first monitoring program for an operating system is generated; then, based on the eBPF mechanism, the first monitoring program is injected into the kernels of respective business servers, so that each of the business servers monitors its own operating system according to the first monitoring program; and monitoring data returned by each of the business servers according to the first monitoring program is received. In the above monitoring method, a user can use the kernel eBPF mechanism to write a monitoring program through a page for dynamic instrumentation, collect and analyze monitoring data for multiple modules of an operating system in a cluster. Therefore, different data collection points and analysis sources can be deployed according to actual needs through the monitoring page, the problem occurrence source can be accurately judged, the ability of the operating system cluster to monitor and collect fault information is improved, and thus the operation and maintenance time and operation and maintenance cost of the operating system cluster are saved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a flowchart of a method for monitoring an operating system cluster based on eBPF provided by an embodiment of the present application;

[0055] Figure 2 is the system structure diagram for implementing the cluster monitoring method provided in the embodiments of the present application;

[0056] Figure 3 is the schematic structural diagram of the operating system cluster monitoring device based on eBPF provided in the embodiments of the present application;

[0057] Figure 4 is the schematic structural diagram of the electronic device provided in the embodiments of the present application. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0059] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0060] Next, the eBPF-based operating system cluster monitoring method provided in the embodiments of the present application will be described in detail in conjunction with the accompanying drawings through specific embodiments and their application scenarios.

[0061] Please refer to Figure 1 , which shows the step flowchart of an eBPF-based operating system cluster monitoring method provided in the embodiments of the present application. Among them, it is applied to a management server, and the management server establishes a server cluster with multiple business servers. The method may include steps S100 to S400.

[0062] Step 100: Display a monitoring page to the user side.

[0063] In the above step 100, the monitoring page is the management interface of the management server, which can be used by users to edit the monitoring program, control the monitoring program, and view the monitoring data. That is, users can log in to the management server cluster in a web manner to collect operating system running data. Among them, the above monitoring page can be a Web page.

[0064] In practical applications, users can access the management server through a browser and display the above monitoring page.

[0065] Step 200: Generate a first monitoring program for the operating system according to the user's first input on the above monitoring page.

[0066] In the above step 200, the first monitoring program is a program for monitoring the operating system of the business server; the first input is an operation for generating a monitoring program, specifically, a monitoring program editing operation performed by the user on the monitoring page; thus, a monitoring program for the operating system, that is, the above first monitoring program, can be generated according to the above first input.

[0067] Step 300: Based on the eBPF mechanism, inject the above first monitoring program into the kernels of each business server so that each of the above business servers can monitor its own operating system according to the above first monitoring program.

[0068] In the above step 300, that is, by using the extended Berkeley Packet Filter (eBPF) mechanism for kernel extension of the business server, the monitoring program written by the user through the monitoring page is converted into bytecode, and then dynamic instrumentation is performed on the kernels of each business server in the cluster. The operating systems of each business server can be monitored by using this monitoring program to obtain monitoring data, and then the monitoring data is fed back to the management server to achieve the collection and analysis of monitoring data for multiple modules of the operating system in the cluster.

[0069] Step 400: Receive the monitoring data returned by each of the above business servers according to the above first monitoring program.

[0070] In the above step 400, because the above first monitoring program not only defines the modules to be monitored but also defines the feedback logic of the corresponding monitoring data, after the first monitoring program is started, the business server runs the first monitoring program to collect monitoring data and feeds it back to the management server, so that users can monitor and analyze the operating system running conditions of each business server on the management server side.

[0071] The operating system cluster monitoring method provided by the embodiment of this application first displays a monitoring page to the user side; then generates a first monitoring program for the operating system according to the first input of the user on this monitoring page; and then injects the first monitoring program into the kernels of each business server based on the eBPF mechanism for each business server to monitor its own operating system according to this first monitoring program; and receives the monitoring data returned by each business server according to the first monitoring program. In the above monitoring method, the user can use the kernel eBPF mechanism to write a monitoring program through the page for dynamic instrumentation, collect and analyze monitoring data for multiple modules of the operating system in the cluster, so that different data collection points and analysis sources can be deployed according to actual needs, accurately determine the source of the problem, and improve the ability to monitor the operating system cluster and collect fault information, thereby saving the operation and maintenance time and cost of the operating system cluster.

[0072] Optionally, in an implementation manner, the above step 200 includes steps 201 to 202:

[0073] Step 201, define a return data structure, user-space monitoring logic, and kernel-space monitoring logic according to the above first input.

[0074] In the above step 201, the user-space monitoring logic is used to specify the analysis logic after receiving data in the user space; the kernel-space monitoring logic is used to specify the monitoring objects and monitoring methods in the kernel space; the return data structure is used to define the data format, data type, etc. of the data that needs to be returned to the management server; among them, because the first input is the monitoring program editing operation performed by the user on the monitoring page to generate the monitoring program, and the monitoring program is jointly composed of the return data structure, user-space monitoring logic, and kernel-space monitoring logic, the user needs to define the return data structure, user-space monitoring logic, and kernel-space monitoring logic through the above first input to form the monitoring program.

[0075] Step 202, generate the above first monitoring program according to the above return data structure, user-space monitoring logic, and kernel-space monitoring logic.

[0076] In the above step 202, that is, according to the return data structure, user-space monitoring logic, and kernel-space monitoring logic customarily edited by the user on the monitoring page, a monitoring file is formed, and then converted into computer-readable bytecode to form the above first monitoring program.

[0077] In the above implementation manner, the user customizes the return data structure, user-space monitoring logic, and kernel-space monitoring logic on the monitoring page according to the monitoring needs to form a monitoring program for each operating system in the cluster, so as to perform operating system cluster monitoring and problem data collection, which provides convenience for analyzing the operating status of the operating systems in the cluster.

[0078] Optionally, in one implementation, before step 200 in the method for monitoring an operating system cluster based on eBPF provided by the embodiments of the present application, steps 101 to 103 are further included; the first input includes a first sub-input and a second sub-input, and step 200 includes steps 211 to 214:

[0079] Step 101: On the monitoring page, display the supported kernel modules, the function list corresponding to the kernel modules, and the supported mounting methods.

[0080] In step 101, that is, on the monitoring page, display the supported kernel modules, the corresponding function list, and the supported mounting points to facilitate the user to customize the page and write the monitoring processing logic online. Optionally, the function parameters corresponding to each function in the function list can also be displayed on the monitoring page to facilitate the user to refer to and select when writing the monitoring processing logic.

[0081] In practical applications, the above kernel modules include System Call Interface, File System, Virtual Memory, TCP / UDP protocol stack, IP layer in the protocol stack, etc.

[0082] Step 102: Receive the second input of the user on the monitoring page, and determine the target kernel module, the target function, and the target mounting method.

[0083] In step 102, the second input is an operation for the user to select the kernel module, function, and mounting method to be monitored. Therefore, when the above second input is received, the selected kernel module is determined as the target kernel module, the selected function is used as the target function, and the selected mounting method is determined as the above target mounting method.

[0084] Step 103: Generate a program template in the BCC compilation chain format according to the above target kernel module, target function, and target mounting method.

[0085] In step 103, BCC (BPF Compiler Collection) is a toolset for tracking kernels and operating programs based on eBPF. Tools such as ebpf / kprobe / uprobe are integrated in its software package, which can be encapsulated using Python and some common functions are integrated. Since BCC is an open-source Linux dynamic tracking tool without third-party module dependencies, and this tool inherits the functions of the virtual machine in BPF, it can track programs efficiently and safely.

[0086] In step 103, according to the compilation chain format required by BCC, the kernel modules, functions, and mounting methods selected by the user are compiled into a program template.

[0087] Step 211: Receive the first sub-input from the user for the above program template.

[0088] In this step 211, the first sub-input is that the user performs a custom editing operation on the program template on the monitoring page.

[0089] Step 212: In response to the above first sub-input, define the return data structure, user-space monitoring logic, and kernel-space monitoring logic.

[0090] In this step 212, based on the user's custom editing operation on the program template, determine the return data structure, user-space monitoring logic, and kernel-space monitoring logic in the monitoring program.

[0091] Step 213: Receive the second sub-input from the user for the above monitoring page.

[0092] In this step 213, the second sub-input is the input for determining the generation of the monitoring program. In practical applications, the above second sub-input is a click or touch operation on the first control on the monitoring page, and the first control is a control used to indicate saving the currently edited monitoring logic as the monitoring program.

[0093] Step 214: In response to the above second sub-input, generate the above first monitoring program based on the defined return data structure, user-space monitoring logic, and kernel-space monitoring logic.

[0094] In this step 214, because the second sub-input is the input for determining the generation of the monitoring program, when the above second sub-input is received, it means that the user needs to form the edited monitoring logic code into a monitoring program. Therefore, based on the return data structure, user-space monitoring logic, and kernel-space monitoring logic defined by the user on the monitoring page, the corresponding monitoring program, that is, the above first monitoring program, is generated.

[0095] In the above implementation, the supported kernel modules, the corresponding function list, and the supported mount points are displayed on the monitoring page for the user to select. Then, a program template can be generated according to the user's selection in accordance with the compilation chain format required by BCC, which is convenient for the user to customize and write the monitoring processing logic online.

[0096] Optionally, after the first monitoring program is injected into the business server kernel, the kernel-space program in the first monitoring program is verified. Specifically, verifications such as no illegal memory access, no crashing of the kernel, and no infinite loops are performed. Only after the verification is successful, the first monitoring program is loaded and run in the kernel, otherwise, relevant logs are recorded and displayed to the user through the page.

[0097] In the embodiment of the present application, when the first monitoring program is injected into the kernel, BCC is responsible for loading and starting by using the verification mechanism of eBPF itself. If the startup fails, the error message is read and a startup error prompt is given.

[0098] Optionally, in an implementation manner, the eBPF-based operating system cluster monitoring method provided by the embodiment of the present application further includes steps 104 to 106 before step 200:

[0099] Step 104: Receive a third input for the above first monitoring program.

[0100] In step 104, the third input is an input for starting the second monitoring program; in practical applications, the above third input is a click or touch operation on the second control on the monitoring page, and the second control is a control for indicating the start of the second monitoring program.

[0101] Step 105: In response to the above third input, start verifying the above first monitoring program.

[0102] In step 105, when the third input is received, it indicates that the user hopes to start the first monitoring program. Therefore, the first monitoring program is first verified for startup on the management server side to avoid injecting a monitoring program that cannot be started into the kernel. In practical applications, specifically, the user-mode program in the first monitoring program is verified for startup because the user-mode program is a Python program, and Python will perform verification. If there are syntax errors, etc., the startup will fail.

[0103] Step 106: In the case where the above first monitoring program starts successfully, execute the step of injecting the above first monitoring program into the kernels of each business server based on the eBPF mechanism.

[0104] In step 106, that is, only after the first monitoring program starts successfully on the management server side, the first monitoring program will be injected and loaded into the kernels of each business server to run, and then the operating system is monitored.

[0105] In the above implementation manner, by verifying the startup of the monitoring program on the management server side before starting the monitoring program and injecting the monitoring program into the kernel of the business server only after successful startup, it can effectively avoid injecting a monitoring program that cannot be started into the kernel.

[0106] Optionally, in a specific implementation manner, the eBPF-based operating system cluster monitoring method provided by the embodiment of the present application further includes step 107 after step 400:

[0107] Step 107: Based on the eBPF mechanism, send the startup command to each business server to control the startup of the above first monitoring program.

[0108] In this step 107, when the second input is received, it indicates that the user hopes to start the first monitoring program. Therefore, based on the eBPF mechanism, the startup command is sent to each business server to control each business server to start the first monitoring program, so as to monitor the operating system of the business server according to the monitoring logic of the first monitoring program.

[0109] In the embodiment of the present application, the monitoring rule configuration can be stored as a monitoring program, and all the stored monitoring programs are displayed on the monitoring page, so that the user can select to start again when needed, and after starting, the business server is re-injected based on the eBPF mechanism to monitor its operating system.

[0110] Optionally, in a specific implementation manner, after step 100 in the method for monitoring an operating system cluster based on eBPF provided by the embodiment of the present application, steps 111 to 112 are further included.

[0111] In the embodiment of the present application, all the monitoring programs injected into the kernel of the business server can be displayed on the monitoring page for the user to select for uninstallation. In practical applications, the number of monitoring programs injected and running in the kernel of the business server can be one or more.

[0112] Step 111: Receive a fourth input for the second monitoring program on the above monitoring page.

[0113] In this step 111, the second monitoring program is a monitoring program injected into the kernel of the business server, and the second monitoring program can be the above first monitoring program or other monitoring programs; the fourth input is an input for uninstalling the second monitoring program; in practical applications, the above fourth input is a click or touch operation on the third control on the monitoring page, and the third control is a control for indicating the uninstallation of the second monitoring program.

[0114] Step 112: In response to the above fourth input, based on the eBPF mechanism, send an uninstallation command to each business server to control the uninstallation of the above second monitoring program.

[0115] In this step 112, when the fourth input is received, it indicates that the user hopes to uninstall the second monitoring program. Therefore, based on the eBPF mechanism, the uninstallation command is sent to each business server to control each business server to uninstall the second monitoring program.

[0116] In the above implementation manner, by performing the uninstallation operation of the monitoring program on the monitoring page, the corresponding monitoring program can be uninstalled on each server in the cluster.

[0117] Optionally, in one embodiment, the method for monitoring an operating system cluster based on eBPF provided by the embodiments of the present application further includes step 500:

[0118] Step 500: Store the above monitoring data in a preset format and provide a data query interface.

[0119] In this embodiment, the monitoring data sent back by the business server is stored in a database in a unified format, and a query interface is provided to facilitate users to query and analyze the operating conditions of the operating systems of each business server.

[0120] In this embodiment, after receiving the monitoring data, the data is converted and saved to the database. Specifically, when implementing, the lightweight in-memory database SQLite3 can be used for saving. Its in-memory database has high read and write performance, and the database file can also be encrypted conveniently to ensure the confidentiality and reliability of the data.

[0121] In the embodiments of the present application, an alarm module and a message module are constructed at the management server and are respectively displayed on the monitoring page. Among them, the alarm module displays all the operating system cluster data collected by the monitoring, and provides search methods such as keyword query and time sorting; the message module is used to push the data monitored by each monitoring program in the last hour to the monitoring page for real-time alarm, which is convenient for users to process in time.

[0122] In the embodiments of the present application, the user logs in to the management server cluster through a web page, and can online write a monitoring and problem collection program in a page form according to the monitoring needs, and use the eBPF / BCC compilation chain mechanism to perform dynamic instrumentation on each business server to collect the operating data of each operating system. Among them, the collected data is the monitoring data, and the content it contains is defined by the monitoring program.

[0123] Optionally, in one embodiment, in the method for monitoring an operating system cluster based on eBPF provided by the embodiments of the present application, the above monitoring data includes program startup error information; after the above step 400, the method further includes step 600:

[0124] Step 600: Display the above program startup error information, where the above program startup error information is sent back to the above management server by the above business server through logging when the above first monitoring program fails to start.

[0125] In step 600, after the user defines the data structure for backhaul, the user-mode program, and the kernel-mode monitoring logic respectively according to the program template on the monitoring page, the generated monitoring program is deployed to other business servers through the file channel; then when the monitoring program needs to be started, a start instruction is transmitted to the business server through the command channel. After receiving the instruction on the business server, the corresponding monitoring program is started through the eBPF mechanism. If the start fails, the attempt to start continues. If the monitoring program still fails to start after the number of attempt times exceeds the default maximum number, the start error information is logged and sent back to the management server.

[0126] Optionally, in an implementation, in the method for monitoring an operating system cluster based on eBPF provided by the embodiments of the present application, the above monitoring data includes node exception information; after the above step 400, the above method further includes step 700:

[0127] Step 700, display the above node exception information, where the above node exception information is sent back to the above management server by the above business server in the event of an exception.

[0128] In this step, when the user customizes and edits the monitoring program on the monitoring page, the user can define the monitoring of each business node. Then, after injecting the monitoring program into the kernel, each business node is continuously monitored. If a node has an exception, data is collected according to the monitoring logic and transmitted to the management server through the eBPF data channel, and then displayed to the user through the visualization device.

[0129] In practical applications, since data is generated when the business node triggers the function execution, it can be determined whether there is an exception in the business node based on this data.

[0130] Optionally, in an implementation, in the method for monitoring an operating system cluster based on eBPF provided by the embodiments of the present application, the above monitoring data includes heartbeat information; after the above step 400, the above method further includes steps 801 to 803:

[0131] Step 801, obtain the heartbeat information sent by each business server.

[0132] In this step 801, the above heartbeat information is specifically a heartbeat signal, which is a data packet sent by the business server in the cluster to the management server every predetermined time for the management server to determine whether the communication link with the current business server has been disconnected. Specifically, the above preset actually needs to be defined according to the business scenario, for example, it is 5s.

[0133] Step 802, if the heartbeat information of the first business server not included in the list of surviving servers is received, add the above first business server to the list of surviving servers.

[0134] In step 802, when the heartbeat information of a business server that does not exist in the list of surviving servers is received, it indicates that a communication link has been established between the business server and the management server, and thus it can be monitored. Therefore, it is added to the list of surviving servers.

[0135] Step 803: If the heartbeat information of the second business server is not received within a preset time period, the above-mentioned second business server is deleted from the list of surviving servers.

[0136] In this step 803, the second business server is any business server in the list of surviving servers; when the heartbeat information of the second business server is not received within the preset time period, it indicates that the communication link between the business server and the management server has been disconnected and it cannot be monitored. Therefore, it is deleted from the list of surviving servers.

[0137] In the above implementation manner, the heartbeat information sent by the business server is obtained from the received monitoring data to maintain the list of surviving servers managed in the current cluster.

[0138] Please refer to Figure 2 , which shows the system structure diagram for implementing the cluster monitoring method provided in the embodiments of the present application.

[0139] As Figure 2 shown, the system includes a rule configuration and function customization device, a transmission device, an eBPF controller, a monitoring data management device, and a visualization device. Among them, the rule configuration and function customization device, the monitoring data management device, and the visualization device are only configured on the management server, while the transmission device and the eBPF controller are configured at both the management server and the business server;

[0140] Among them, the rule configuration and function customization device is used to implement the configuration selection and browsing of the monitoring module through a Web page, and to configure the monitoring module in the operating system kernel state; display the supported kernel modules, the corresponding function lists, the supported mounting methods, and the corresponding function parameters through the Web page; when a specific module and function are selected, call the eBPF controller interface to obtain the processing templates of the user-mode and kernel-mode programs generated by the eBPF controller and display them on the Web page, and the user can directly customize and edit the monitoring logic; after the user defines the backhaul data structure, the user-mode program, and the kernel-mode monitoring logic editing according to the program template respectively, the generated file is deployed to other business servers through the file channel of the transmission device.

[0141] Among them, the transmission device is the transmission channel of the entire cluster monitoring system, responsible for transmitting files, commands, and data between clusters. It includes a file channel, a command channel, and a data channel. The file channel is used to transmit user-defined monitoring programs, the command channel is used to transmit start commands to start monitoring and send heartbeat information of the business server to the management server; the data channel transmits the collected data returned by the business server.

[0142] Among them, the eBPF controller is the control center of the entire cluster monitoring system. Its main function is to configure according to the received rules and customize the instructions of the functional device, start or uninstall the corresponding monitoring program in each server in the cluster, and use the existing BCC compilation chain mechanism to inject the monitoring program into the kernel.

[0143] Specifically, when the management server receives the command to start the module to be monitored, it locally starts the monitoring program configured by the user. If the user-mode program fails to start or the kernel-mode injection fails, relevant logs are recorded and the error information is displayed to the user;

[0144] At the same time, the eBPF controller on the management server side receives the heartbeat information sent by the transmission device in real time, maintains the list of alive servers managed in the current cluster. If the business server does not send heartbeat information within the specified time, it is deleted from the current server list. If the server recovers, it is added to the list of alive servers again.

[0145] In addition, if the monitoring program starts successfully on the management server, the command is sent to each business server in the list of alive servers through the command channel and the corresponding module is started.

[0146] Among them, the monitoring data management device is the storage center of the entire cluster monitoring system, used to store the monitoring rules and monitoring data configured by the user, and provide a unified data storage and query interface. After receiving the collected data information from the transmission device, it converts the data information and saves it to the database.

[0147] Among them, the main function of the visualization device is to read and display data from the monitoring data management device.

[0148] Specifically, the visualization device includes an alarm center, a message center, and a service backend; the alarm center is used to display all operating system data in the monitored cluster and provide search methods such as keyword query and time sorting; the message center is used to perform real-time alarms on the data monitored in the last hour; the service backend is used to obtain the monitored data from the database of the monitoring data management device.

[0149] It should be noted that for the eBPF-based operating system cluster monitoring method provided in the embodiments of the present application, the execution entity can be a server, or a control module in the server for executing the eBPF-based operating system cluster monitoring method. In the embodiments of the present application, taking the server executing the eBPF-based operating system cluster monitoring method as an example, the eBPF-based operating system cluster monitoring device provided in the embodiments of the present application is described.

[0150] Please refer to Figure 3 , which shows a schematic structural diagram of the eBPF-based operating system cluster monitoring device provided in the embodiments of the present application. It is applied to a management server, and the management server establishes a server cluster with multiple business servers. As Figure 3 shown, the device 30 includes:

[0151] A page display module 31, configured to display a monitoring page to the client;

[0152] A program construction module 32, configured to generate a first monitoring program for the operating system according to a first input of the user on the above monitoring page;

[0153] A program injection module 33, configured to inject the above first monitoring program into the kernels of each business server based on the eBPF mechanism, so that each of the above business servers can monitor its own operating system according to the above first monitoring program;

[0154] A first receiving module 34, configured to receive the monitoring data returned by each of the above business servers according to the above first monitoring program.

[0155] Optionally, in the eBPF-based operating system cluster monitoring device, the program construction module 33 includes:

[0156] A first definition unit, configured to define a return data structure, a user-space monitoring logic, and a kernel-space monitoring logic according to the above first input;

[0157] A first generation unit, configured to generate the above first monitoring program according to the above return data structure, user-space monitoring logic, and kernel-space monitoring logic.

[0158] Optionally, in one implementation, the eBPF-based operating system cluster monitoring device further includes:

[0159] A first display module, configured to display the supported kernel modules, the function list corresponding to the above kernel modules, and the supported mounting methods on the above monitoring page before generating a first monitoring program for the operating system according to a first input of the user on the above monitoring page;

[0160] The first determination module is configured to receive a second input from the user on the above monitoring page, and determine a target kernel module, a target function, and a target mounting method;

[0161] The template generation module is configured to generate a program template in the BCC compilation chain format according to the above target kernel module, target function, and target mounting method;

[0162] The above first input includes a first sub-input and a second sub-input; the program construction module 32 includes:

[0163] The first receiving unit is configured to receive a first sub-input from the user for the above program template;

[0164] The first definition unit is configured to define a return data structure, a user-space monitoring logic, and a kernel-space monitoring logic in response to the above first sub-input;

[0165] The second node unit is configured to receive a second sub-input from the user for the above monitoring page;

[0166] The second generation unit is configured to generate the above first monitoring program based on the defined return data structure, user-space monitoring logic, and kernel-space monitoring logic in response to the above second sub-input.

[0167] Optionally, in one implementation, the eBPF-based operating system cluster monitoring device further includes:

[0168] The second receiving module is configured to receive a third input for the above first monitoring program before injecting the above first monitoring program into the kernels of each business server based on the eBPF mechanism;

[0169] The verification module is configured to start verifying the above first monitoring program in response to the above third input;

[0170] The execution module is configured to execute the step of injecting the above first monitoring program into the kernels of each business server based on the eBPF mechanism when the above first monitoring program is successfully started.

[0171] Optionally, in one implementation, the eBPF-based operating system cluster monitoring device further includes:

[0172] The first sending module is configured to send a start command to each business server based on the eBPF mechanism after injecting the above first monitoring program into the kernels of each business server, so as to control the start of the above first monitoring program.

[0173] Optionally, in one implementation, the eBPF-based operating system cluster monitoring device further includes:

[0174] A third receiving module, configured to receive a fourth input to the second monitoring program in the monitoring page after presenting the monitoring page to the user side;

[0175] A second sending module, configured to send an uninstallation command to each business server based on the eBPF mechanism in response to the fourth input, so as to control the uninstallation of the second monitoring program.

[0176] Optionally, in an implementation manner, the eBPF-based operating system cluster monitoring device further includes:

[0177] A storage module, configured to store the monitoring data in a preset format and provide a data query interface.

[0178] Optionally, in an implementation manner, the monitoring data includes program startup error information;

[0179] The eBPF-based operating system cluster monitoring device further includes:

[0180] A second display module, configured to present the program startup error information after receiving the monitoring data returned by each business server, where the program startup error information is returned by the business server to the management server by logging in the case of failure to start the first monitoring program.

[0181] Optionally, in an implementation manner, the monitoring data includes node exception information;

[0182] The eBPF-based operating system cluster monitoring device further includes:

[0183] A third display module, configured to present the node exception information after receiving the monitoring data returned by each business server, where the node exception information is returned by the business server to the management server in the case of an exception.

[0184] Optionally, in an implementation manner, the monitoring data includes heartbeat information;

[0185] The eBPF-based operating system cluster monitoring device further includes:

[0186] An obtaining module, configured to obtain the heartbeat information sent by each business server after receiving the monitoring data returned by each business server;

[0187] An adding module, configured to add the first business server to the list of alive servers if the heartbeat information of the first business server that is not included in the list of alive servers is received;

[0188] A deleting module, configured to delete the second business server from the list of alive servers if the heartbeat information of the second business server is not received within a preset duration.

[0189] The operating system cluster monitoring device based on eBPF provided by the embodiment of the present application first has the page display module 31 display a monitoring page to the user side; then the program construction module 32 generates a first monitoring program for the operating system according to the first input of the user on the monitoring page; and then the program injection module 33 injects the first monitoring program into the kernels of each business server based on the eBPF mechanism for each business server to monitor its own operating system according to the first monitoring program; the first receiving module 34 receives the monitoring data returned by each business server according to the first monitoring program. In the above monitoring method, the user can utilize the kernel eBPF mechanism to write a monitoring program through the page for dynamic instrumentation, collect and analyze monitoring data for multiple modules of the operating system in the cluster, and thus can deploy different data collection points and analysis sources according to actual needs, accurately determine the source of problems, improve the ability to monitor the operating system cluster and collect fault information, thereby saving the operation and maintenance time and cost of the operating system cluster.

[0190] The operating system cluster monitoring device based on eBPF in the embodiment of the present application is a device with an operating system. The operating system can be a linux operating system or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0191] The operating system cluster monitoring device based on eBPF in the embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., which are not specifically limited in the embodiment of the present application.

[0192] The operating system cluster monitoring device based on eBPF provided by the embodiment of the present application can implement Figures 1 to 2 each process implemented by the method embodiment. To avoid repetition, it will not be elaborated here.

[0193] Optionally, as Figure 4As shown in the figure, an embodiment of the present application further provides an electronic device 40, including a processor 41, a memory 42, and a program or instruction stored on the memory 42 and executable on the processor 41. When the program or instruction is executed by the processor 41, it implements each process of the above-mentioned embodiment of the eBPF-based operating system cluster monitoring method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0194] An embodiment of the present application further provides a non-volatile readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the eBPF-based operating system cluster monitoring method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0195] Wherein, the processor is the processor in the server described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0196] Another embodiment of the present application provides a chip. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned embodiment of the eBPF-based operating system cluster monitoring method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0197] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0198] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0200] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A method for monitoring an operating system cluster based on eBPF, characterized in that, Applied to a management server, the management server establishes a server cluster with multiple business servers, and the method includes: Display a monitoring page to the client; Generate a first monitoring program for the operating system according to the first input of the user on the monitoring page; Based on the eBPF mechanism, inject the first monitoring program into the kernels of each business server, so that each business server monitors its own operating system according to the first monitoring program; Receive the monitoring data returned by each business server according to the first monitoring program; Before generating a first monitoring program for the operating system according to the first input of the user on the monitoring page, the method further includes: On the monitoring page, display the supported kernel modules, the function list corresponding to the kernel modules, and the supported mounting methods; Receive the second input of the user on the monitoring page, and determine the target kernel module, the target function, and the target mounting method; Generate a program template in the BCC compilation chain format according to the target kernel module, the target function, and the target mounting method; The first input includes a first sub-input and a second sub-input. Generating a first monitoring program for the operating system according to the first input of the user on the monitoring page includes: Receive the first sub-input of the user for the program template; In response to the first sub-input, define a return data structure, user-space monitoring logic, and kernel-space monitoring logic; Receive the second sub-input of the user on the monitoring page; In response to the second sub-input, generate the first monitoring program based on the defined return data structure, user-space monitoring logic, and kernel-space monitoring logic.

2. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein Before injecting the first monitoring program into the kernels of each business server based on the eBPF mechanism, the method further includes: Receive a third input for the first monitoring program; In response to the third input, start verifying the first monitoring program; In the case where the first monitoring program starts successfully, execute the step of injecting the first monitoring program into the kernels of each business server based on the eBPF mechanism.

3. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein After injecting the first monitoring program into the kernels of each business server, the method further includes: Based on the eBPF mechanism, send a start command to each business server to control the start of the first monitoring program.

4. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein After displaying the monitoring page to the client, the method further includes: Receive a fourth input for a second monitoring program in the monitoring page; In response to the fourth input, based on the eBPF mechanism, send an uninstall command to each business server to control the uninstallation of the second monitoring program.

5. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein The method further includes: Store the monitoring data in a preset format and provide a data query interface.

6. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein The monitoring data includes program startup error information; After receiving the monitoring data returned by each business server, the method further includes: Display the program startup error information, which is returned by the business server to the management server by recording a log in the case where the first monitoring program fails to start.

7. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein The monitoring data includes node exception information; After receiving the monitoring data returned by each business server, the method further includes: The node abnormality information is displayed, and the node abnormality information is transmitted back to the management server by the business server when an abnormality occurs.

8. The method for monitoring an operating system cluster based on eBPF according to claim 1, wherein The monitoring data includes heartbeat information; After receiving the monitoring data returned by each service server, the method further includes: Get the heartbeat information sent by each business server; If the heartbeat information of the first service server not included in the surviving server list is received, the first service server is added to the surviving server list; If the heartbeat information of the second service server is not received within the preset time period, the second service server is deleted from the surviving server list.

9. An operating system cluster monitoring device based on eBPF, characterized in that, Applied to a management server, the management server and multiple business servers establish a server cluster, the device includes: Page display module, used to display the monitoring page to the user end; A program building module, used for generating a first monitoring program for the operating system according to a first input of a user on the monitoring page; A program injection module, used for injecting the first monitoring program into the kernel of each service server based on the eBPF mechanism, so that each service server can monitor its own operating system according to the first monitoring program; A first receiving module, used for receiving monitoring data returned by each of the business servers according to the first monitoring program; The device further comprises: a first display module, for displaying supported kernel modules, a function list corresponding to the kernel modules, and supported mounting methods on the monitoring page before generating a first monitoring program for the operating system according to a first input of a user on the monitoring page; A first determination module is used to receive a second input from a user on the monitoring page, and determine a target kernel module, a target function, and a target mounting method; A template generation module, used to generate a program template in accordance with the BCC compilation chain format according to the target kernel module, target function and target mounting mode; The first input includes a first sub-input and a second sub-input; the program construction module includes: A first receiving unit, configured to receive a first sub-input of the program template by a user; A first definition unit, configured to define a return data structure, a user-mode monitoring logic, and a kernel-mode monitoring logic in response to the first sub-input; A second node unit, used for receiving a second sub-input of the user on the monitoring page; The second generating unit is used to generate the first monitoring program in response to the second sub-input based on the defined return data structure, user-mode monitoring logic and kernel-mode monitoring logic.

10. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the eBPF-based operating system cluster monitoring method as described in any one of claims 1 to 8 are implemented.

11. A non-volatile readable storage medium, characterized in that, The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the eBPF-based operating system cluster monitoring method as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • eBPF-based micro-service system performance detection method, device and system

    CN112256542A

  • Operating system fault monitoring device and method

    CN113778788A