EBPF-based parallel file system high availability method, system and equipment
By using eBPF technology in parallel file systems to monitor and repair the read-only mode of the local file system, and switching new object storage targets when repair fails, the shortcomings of fault handling and read-only mode processing in distributed storage environments are solved, and high availability and business continuity are achieved.
Patent Information
- Application Number
- CN202510091247.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
AI Technical Summary
Existing parallel file systems are difficult to effectively distinguish temporary and permanent failures in distributed storage environments, and lack flexible processing mechanisms for the read-only mode of local file systems, resulting in high availability and business continuity being affected.
Using an eBPF-based method, data is striped in parallel to the storage volume through the parallel file client, the read-only mode status of the local file system is monitored, and it is repaired through a high-availability proxy. When repair fails, select the new object storage target to remount the failed storage volume.
It realizes more efficient fault detection, repair and resource switching in a distributed storage environment, ensuring high availability and business continuity of parallel file systems.
Smart Images

Figure CN119938385A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage technology, and in particular to a high-availability method, system and device for a parallel file system based on eBPF. Background Art
[0002] With the rapid development of information technology, the demand for data storage and processing is increasing, especially in large-scale distributed computing environments, parallel file systems have been widely used. Parallel file systems significantly improve data read and write performance and system throughput by distributing file data on multiple storage nodes and allowing clients to access these nodes in parallel. However, existing parallel file systems still have some problems in high availability, especially in the face of failures in distributed storage environments, the availability and continuity of the system are difficult to be effectively guaranteed.
[0003] A typical parallel file system architecture includes the client, metadata server (MDS), object storage target (OST), and back-end distributed block storage. The client distributes IO requests in parallel to the local file system (LocalFS) of each node through the network, and the local file system writes the data to the back-end distributed block storage. While this architecture improves performance, it also introduces complex failure modes and high availability challenges.
[0004] Existing high availability solutions are usually based on node-level fault detection and resource takeover mechanisms. For example, when a storage node (such as Host1) fails, the high availability agent (HA Agent) on other nodes (such as Host2 and Host3) will detect the failure and take over the storage volume on the failed node to continue providing file services to clients. However, this solution has obvious shortcomings when dealing with the read-only mode of the local file system.
[0005] When a local file system encounters a read-write error (such as EIO, input / output error) when processing an IO request, the file system is usually marked as read-only mode to protect the integrity of the file system. This read-only mode will cause the parallel file system to be unable to continue to provide normal write operation services, thereby affecting business continuity. However, in a distributed block storage environment, an IO read-write failure may be caused by temporary problems such as network failure, host failure, or disk failure, and these problems do not necessarily mean that the entire storage volume is unavailable. After the distributed storage cluster returns to normal through its own repair mechanism, the storage volume can still continue to provide services. However, at this time, since the local file system has entered read-only mode, the existing high availability mechanism cannot effectively restore it to a writable state, resulting in service interruption. The shortcomings of this existing high availability solution are mainly due to the fact that the existing technology fails to effectively distinguish between temporary and permanent failures in a distributed storage environment, and also lacks a flexible processing mechanism for the read-only mode of the local file system.
[0006] Therefore, how to achieve more efficient fault detection, repair, and resource switching in a distributed storage environment to ensure the high availability of parallel file systems is a technical problem that needs to be solved urgently. Summary of the invention
[0007] In view of this, in order to overcome the deficiencies of the prior art, the present application aims to provide a high-availability method, system and device for a parallel file system based on eBPF.
[0008] According to a first aspect of the present application, a parallel file system high availability method based on eBPF is provided, the method comprising: The parallel file client stores data in parallel stripes to the storage volume; When the distributed storage layer fails, the local file system corresponding to the storage volume is adjusted to read-only mode; Monitor the read-only mode status of the local file system and repair the local file system in read-only mode through the high-availability agent; When the local file system fails to be repaired, a new object storage target is selected through the high-availability proxy to remount the failed storage volume.
[0009] Optionally, in the high availability method of the parallel file system based on eBPF of the present application, the parallel file client stores the data in parallel stripes in the storage volume, including: The parallel file client sends a file read and write request to the metadata server and obtains file layout information from the metadata server; The parallel file client sends data write requests in parallel to the local file systems of multiple object storage targets according to the acquired file layout information; The local file system of the object storage target stores the corresponding data in stripes in the storage volume according to the received data write request.
[0010] Optionally, in the eBPF-based parallel file system high availability method of the present application, the parallel file client stores the data in parallel stripes in the storage volume, and also includes: the storage volume sends read and write IO requests to the storage nodes of the back-end distributed storage layer according to the stored data.
[0011] Optionally, in the eBPF-based parallel file system high availability method of the present application, when a distributed storage layer fails, the local file system corresponding to the storage volume is adjusted to a read-only mode state, including: when a distributed storage layer fails, an error code corresponding to the failure is sent to the storage volume, and when the local file system detects the error code in the storage volume, it enters a read-only mode state.
[0012] Optionally, in the high availability method of the parallel file system based on eBPF of the present application, monitoring the read-only mode state of the local file system is triggered, and repairing the local file system in the read-only mode state through the high availability agent includes: Set a Hook monitoring point in the local file system through the eBPF program, and monitor the trigger of the local file system entering the read-only mode in real time through the set Hook monitoring point; When the local file system is monitored to enter the read-only mode, the eBPF program reports the read-only event in real time in user mode to the high-availability agent deployed in the object storage target; After receiving the read-only event, the high-availability agent repairs the local file system.
[0013] Optionally, in the eBPF-based parallel file system high availability method of the present application, the local file system is repaired, including: performing a local file system check and attempting to remount the local file system in read-write mode.
[0014] Optionally, in the eBPF-based parallel file system high availability method of the present application, when the local file system repair fails, a new object storage target is selected by a high availability agent to remount the failed storage volume, including: When the local file system fails to be repaired, the high-availability agent in the failed object storage target selects another object storage target from the storage cluster to take over and mount the storage volume in the failed object storage target, and starts the relevant storage services to verify the running status of the local file system; The parallel file client reconnects to the new object storage target and continues the file read and write operations.
[0015] Optionally, in the eBPF-based parallel file system high availability method of the present application, the high availability agent in the faulty object storage target selects other object storage targets from the storage cluster to take over and mount the storage volumes in the faulty object storage target, including: sending a switching request to the high availability agents in other object storage targets through the high availability agent in the faulty object storage target, and the high availability agent that receives the switching request takes over and mounts the faulty storage volume in the object storage target where it is located.
[0016] According to a second aspect of the present application, a parallel file system high availability system based on eBPF is provided, the system comprising: A parallel storage module is used to call a parallel file client to store data in parallel stripes in a storage volume; A read-only mode state triggering module is used to adjust the local file system corresponding to the storage volume to a read-only mode state when a failure occurs in the distributed storage layer; The monitoring and repair module is used to monitor the read-only mode status of the local file system and repair the local file system in the read-only mode through the high-availability agent; The switch mount module is used to select a new object storage target through the high-availability agent to remount the failed storage volume when the local file system repair fails.
[0017] According to a third aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect of the present application when executing the program.
[0018] The parallel file system high availability method and system of the present application uses eBPF to intercept the relevant logic of the local file system processing input / output errors at runtime. Once the read-only mode is triggered, the high-availability agent promptly triggers the switching of resource mounting by monitoring the read-only event to prevent the read-only mode from causing service unavailability. In the scenario of parallel file docking storage volumes, this embodiment can intercept the "read-only" error handling logic of the local file system through eBPF, perform local repairs in time, and trigger high-availability switching if the repair fails, thereby ensuring high service availability. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1This is an example diagram of the architecture of a parallel file system high availability system based on eBPF according to an embodiment of the present application; Figure 2 A flowchart of a method for high availability of a parallel file system based on eBPF according to an embodiment of the present application; Figure 3 An example architecture for executing a parallel file system high availability method based on eBPF according to an embodiment of the present application; Figure 4 A schematic diagram of the structure of the device provided in this application. DETAILED DESCRIPTION
[0021] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0022] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.
[0023] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0024] Figure 1 FIG. 1 is an example diagram of an architecture of a parallel file system high availability system based on eBPF according to an embodiment of the present application, such as Figure 1 As shown, the system of this embodiment includes: A parallel storage module is used to call a parallel file client to store data in parallel stripes in a storage volume; A read-only mode state triggering module is used to adjust the local file system corresponding to the storage volume to a read-only mode state when a failure occurs in the distributed storage layer; The monitoring and repair module is used to monitor the read-only mode status of the local file system and repair the local file system in the read-only mode through the high-availability agent; The switch mount module is used to select a new object storage target through the high-availability agent to remount the failed storage volume when the local file system repair fails.
[0025] Based on the above system, an embodiment of the present application provides a parallel file system high availability method based on eBPF, which is illustrated by the following embodiments.
[0026] Figure 2 The present invention is a flowchart of a method for high availability of a parallel file system based on eBPF according to an embodiment of the present application. Figure 3 The following is an example architecture for executing a high-availability method for a parallel file system based on eBPF according to an embodiment of the present application. Figure 2 and Figure 3 As shown, the high availability method of the parallel file system based on eBPF in this embodiment includes the following steps: Step S101: The parallel file client stores the data in parallel stripes in a storage volume.
[0027] As an optional example, in this embodiment, the parallel file client sends a file read and write request to the metadata server and obtains the file layout information from the metadata server; the parallel file client sends the data write request in parallel to the local file systems of multiple object storage targets based on the obtained file layout information; the local file system of the object storage target stores the corresponding data in stripes in the storage volume based on the received data write request. The storage volume sends a read and write IO request to the storage node of the back-end distributed storage layer based on the stored data. In actual applications, the storage volume sends a read and write IO request to the back-end distributed storage node through the NVMe oF protocol. NVMe oF (NVMe over Fabrics) is a network protocol that allows access to NVMe storage devices over the network, improving data transmission speed and efficiency.
[0028] In this embodiment, the metadata server is used to manage the metadata of the file system, including information such as file name, directory structure, and permissions. The object storage target is a physical storage device that stores file data, and each object storage target stores a portion of the data of a file. The local file system is a file system running on the object storage target, which is used to manage and operate the mounted storage volume and map the IO requests of the parallel file client to a specific physical storage device. A storage volume is a storage unit composed of multiple physical storage devices, which provides a unified storage space for the file system to use.
[0029] In this embodiment, striped storage refers to dividing file data into multiple stripes and distributing and storing them on multiple storage devices, thereby achieving the goal of improving data reading and writing performance.
[0030] Step S102: When a failure occurs in the distributed storage layer, the local file system corresponding to the storage volume is adjusted to a read-only mode.
[0031] As an optional example, in this embodiment, when a fault occurs in the distributed storage layer, an error code corresponding to the fault is sent to the storage volume, and when the local file system detects the error code in the storage volume, it enters the read-only mode state. In this embodiment, the read-only mode state indicates that the local file system enters the protection mode, allowing only read operations and not write operations, to prevent data corruption.
[0032] Step S103: monitoring the read-only mode state of the local file system to be triggered, and repairing the local file system in the read-only mode state through a high-availability agent.
[0033] As an optional example, in this embodiment, a Hook monitoring point is set in the local file system through the eBPF program, and the trigger of the local file system entering the read-only mode state is monitored in real time through the set Hook monitoring point; when the local file system is monitored to enter the read-only mode state, the eBPF program reports the read-only event in real time to the high availability agent (HA Agent, High Availability Agent) deployed in the object storage target in the user state; after receiving the read-only event, the high availability agent repairs the local file system, for example, performs a local file system check and attempts to remount the local file system in read-write mode. The file system check can check and repair errors in the local file system to ensure the consistency and integrity of the local file system.
[0034] In this embodiment, the eBPF program is a kernel technology that allows user-defined programs to run in the operating system kernel and monitor and debug kernel behavior in real time. The high-availability agent is an agent program deployed in the object storage target, which is used to monitor the status of the local file system, perform repair and switching operations, and ensure high availability of the system.
[0035] Step S104: When the local file system fails to be repaired, a new object storage target is selected through a high-availability proxy to remount the failed storage volume.
[0036] As an optional example, in this embodiment, when the local file system fails to be repaired, the high-availability agent in the failed object storage target selects other object storage targets from the storage cluster to take over and mount the storage volume in the failed object storage target, and starts the relevant storage services to verify the operation status of the local file system; the parallel file client reconnects to the new object storage target to continue the file read and write operations. For example, in this embodiment, the high-availability agent in the failed object storage target sends a switching request to the high-availability agents in other object storage targets, and the high-availability agent that receives the switching request takes over the failed storage volume and mounts it in the object storage target where it is located.
[0037] The high-availability method and system of the parallel file system based on eBPF in this embodiment uses eBPF to intercept the logic related to the local file system processing input / output errors at runtime. Once the read-only mode is triggered, the high-availability agent promptly triggers the switching of resource mounting by monitoring the read-only event to prevent the read-only mode from causing service unavailability. In the scenario of parallel file docking storage volumes, this embodiment can intercept the "read-only" error handling logic of the local file system through eBPF, perform local repairs in time, and trigger high-availability switching if the repair fails, thereby ensuring high availability of the business.
[0038] like Figure 4 As shown, the present application also provides a device, including a processor 210, a communication interface 220, a memory 230 for storing a computer program executable by the processor, and a communication bus 240. The processor 210, the communication interface 220, and the memory 230 communicate with each other through the communication bus 240. The processor 210 implements the above-mentioned high-availability method of parallel file system based on eBPF by running the executable computer program.
[0039] Among them, the computer program in the memory 230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0040] The system embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected based on actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0041] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on such an understanding, the above technical solution can essentially or in other words be embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.
[0042] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A high availability method for a parallel file system based on eBPF, characterized in that: The method comprises: The parallel file client stores data in parallel stripes to the storage volume; When the distributed storage layer fails, the local file system corresponding to the storage volume is adjusted to read-only mode; Monitor the read-only mode status of the local file system and repair the local file system in read-only mode through the high-availability agent; When the local file system fails to be repaired, a new object storage target is selected through the high-availability proxy to remount the failed storage volume.
2. The high availability method for parallel file system based on eBPF according to claim 1, characterized in that: The parallel file client stripes data in parallel to the storage volume, including: The parallel file client sends a file read and write request to the metadata server and obtains file layout information from the metadata server; The parallel file client sends data write requests in parallel to the local file systems of multiple object storage targets according to the acquired file layout information; The local file system of the object storage target stores the corresponding data in stripes in the storage volume according to the received data write request.
3. The high availability method for parallel file system based on eBPF according to claim 1, characterized in that: The parallel file client stores the data in parallel stripes in the storage volume, and also includes: the storage volume sends a read and write IO request to the storage node of the back-end distributed storage layer according to the stored data.
4. The high availability method for parallel file system based on eBPF according to claim 1, characterized in that: When a fault occurs in the distributed storage layer, the local file system corresponding to the storage volume is adjusted to a read-only mode, including: when a fault occurs in the distributed storage layer, an error code corresponding to the fault is sent to the storage volume, and when the local file system detects the error code in the storage volume, it enters a read-only mode.
5. The high availability method for parallel file system based on eBPF according to claim 1, characterized in that: Monitor the read-only mode status of the local file system and repair the local file system in read-only mode through the high-availability agent, including: Set a Hook monitoring point in the local file system through the eBPF program, and monitor the trigger of the local file system entering the read-only mode in real time through the set Hook monitoring point; When the local file system is monitored to enter the read-only mode, the eBPF program reports the read-only event in real time in user mode to the high-availability agent deployed in the object storage target; After receiving the read-only event, the high-availability agent repairs the local file system.
6. The high availability method for parallel file system based on eBPF according to claim 5, characterized in that: Repair the local file system, including checking the local file system and attempting to remount the local file system in read-write mode.
7. The high availability method for parallel file system based on eBPF according to claim 1, characterized in that: When the local file system fails to be repaired, a new object storage target is selected through the high availability proxy to remount the failed storage volume, including: When the local file system fails to be repaired, the high-availability agent in the failed object storage target selects another object storage target from the storage cluster to take over and mount the storage volume in the failed object storage target, and starts the relevant storage services to verify the running status of the local file system; The parallel file client reconnects to the new object storage target and continues the file read and write operations.
8. The high availability method for parallel file system based on eBPF according to claim 7, characterized in that: The high-availability agent in the failed object storage target selects other object storage targets from the storage cluster to take over and mount the storage volume in the failed object storage target, including: sending a switching request to high-availability agents in other object storage targets through the high-availability agent in the failed object storage target, and the high-availability agent that receives the switching request takes over and mounts the failed storage volume in its own object storage target.
9. A parallel file system high availability system based on eBPF, characterized in that: The system comprises: A parallel storage module is used to call a parallel file client to store data in parallel stripes in a storage volume; A read-only mode state triggering module is used to adjust the local file system corresponding to the storage volume to a read-only mode state when a failure occurs in the distributed storage layer; The monitoring and repair module is used to monitor the read-only mode status of the local file system and repair the local file system in the read-only mode through the high-availability agent; The switch mount module is used to select a new object storage target through the high-availability agent to remount the failed storage volume when the local file system repair fails.
10. A computer device, characterized in that: The computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 8 when executing the program.