NFS service high availability method, system and device based on state detection and medium
By deploying CephFS and keepalived services on the NFS-Ganesha server side, and directly detecting the NFS server state in combination with the local loopback address mount policy, the low reliability problem caused by abnormal working status of the server in the NFS high-availability system is solved, and fast failover and high availability are achieved.
Patent Information
- Application Number
- CN202510157897.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-23
AI Technical Summary
The problem of low NFS service reliability caused by abnormal working status of NFS server in NFS high-availability systems.
By deploying CephFS as a storage backend on the NFS-Ganesha server, the keepalived service is used for health monitoring and failover, and the local loopback address mount policy is used to directly detect the working status of the NFS server, achieving rapid failover.
It improves the high availability and reliability of NFS services, ensuring that NFS services can quickly switch to alternate nodes when the server is working abnormally, and avoids unavailability of services.
Smart Images

Figure CN120034553A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a method, system, device and medium for high availability of NFS services based on state detection. Background Art
[0002] Network File System (NFS) is a widely used network file sharing protocol that works based on a client-server model, making remote file access as simple as local operations. With the development of cloud computing and big data technologies, the requirements for NFS have been upgraded, including but not limited to dynamic scalability, cross-platform compatibility, and seamless integration of emerging storage technologies. To this end, NFS-Ganesha was developed. It is a high-performance user-mode NFS server. Different from the traditional kernel-mode implementation, it has a modular plug-in architecture and can flexibly connect to a variety of storage backends. NFS-Ganesha has become an important choice for network file sharing in modern data centers and cloud environments due to its ease of use, flexibility, and support for enterprise-level features (such as high availability and security).
[0003] There are some problems with using NFS directly. The most important one is that the NFS server may cause the NFS service to be unavailable when a single point failure occurs, that is, NFS itself does not have high availability capabilities. In order to ensure the reliability of the network file system, the existing technology usually uses an NFS high availability solution based on VIP (Virtual IP Address). The NFS high availability solution based on VIP requires configuring two or more NFS servers to form a cluster and share storage resources; setting VIP as a unified entrance for client access, hiding the physical address of the actual server, and facilitating failover; using software such as Keepalived or Heartbeat to continuously monitor the status of each NFS server. Once the main server is abnormal, the fault response is immediately triggered, automatically switching to the backup node and configuring VIP to provide services to the outside world. Keepalived can determine that the main server of NFS is abnormal, but it cannot directly detect the working status of the NFS server. If the NFS server responds too slowly or stops responding, that is, it cannot process requests normally, it is also necessary to switch between the main and backup nodes to ensure the availability and reliability of the network file system. Therefore, the focus of the NFS high availability solution becomes detecting the working status of the NFS server. Check whether the NFS server is working normally. Generally, the client tries to write a file in the shared directory. If successful, the server is running normally. If failed, the server is considered to be abnormal. However, this method cannot directly distinguish whether it is a problem with the client, server, or network. Additional steps are required to locate the specific cause of the fault.
[0004] When using NFS service, a single point failure of NFS server or failure of NFS server to process requests normally may cause service unavailability. Therefore, how to avoid the defect of low reliability of NFS service caused by abnormal working status of NFS server in NFS high availability system and ensure stable operation of NFS service is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical task of the present invention is to provide a method, system, device and medium for high availability of NFS service based on status detection to solve the problem of low reliability of NFS service caused by abnormal working status of NFS server in NFS high availability system.
[0006] The technical task of the present invention is achieved in the following manner: a method for high availability of NFS service based on state detection, the method is as follows:
[0007] Use NFS-Ganesha as the user-mode NFS server and integrate CephFS as the backend storage system. The master and backup nodes share the same data. The NFS high-availability solution does not require a backend data synchronization mechanism. At the same time, deploy NFS-Ganesha services on all server nodes in the cluster and export CephFS to NFS through FSAL.
[0008] Deploy the keepalived service on all server nodes in the cluster for health monitoring and failover. All nodes share the same virtual IP address.
[0009] The local loopback address mounting strategy is adopted. By mounting the shared directory exported by the server node, the IP address of the NFS service is specified as the local loopback address to detect the server's ability to process requests.
[0010] Use monitoring scripts to perform status checks regularly, including using the ganesha_stats tool to obtain runtime status information of the NFS-Ganesha service and analyzing the exported NFS protocol usage and operation count changes;
[0011] When an abnormal working status of the server is detected, the keepalived failover mechanism is triggered by adjusting the node priority, automatically switching the VIP to the standby node, and gracefully restarting the NFS-Ganesha process on the original primary node.
[0012] As a preferred method, the NFS-Ganesha service is installed by compiling the source code, adding -DUSE_FSAL_CEPH=ON and -DUSE_RADOS_RECOV=ON in the compilation options, and setting the FSAL type to Ceph in the export (EXPORT) configuration to export CephFS to NFS through FSAL;
[0013] At the same time, add -DUSE_ADMIN_TOOLS=ON in the compilation options to use the state management tool provided by NFS-Ganesha.
[0014] Preferably, the NFS-Ganesha services deployed on the active and standby nodes use the same configuration file, and all exported configurations are written to objects in the Ceph RADOS pool.
[0015] Preferably, the client accesses the NFS server through VIP. If a single point failure occurs in the primary server being accessed, the keepalived service automatically switches to the backup node and configures VIP to provide external services. The client accesses the new server and obtains the last connection status to restore the connection. If the connection fails to be restored, the connection is re-established.
[0016] Preferably, the keepalived service deployed on all server nodes in the cluster uses the same configuration file, all host states are set to BACKUP, the priority is the same, the preemptive mode is selected, and the same virtual IP address is specified;
[0017] The keepalived service provides an application layer health check method to check whether the server is running normally according to the user's settings. At the same time, it defines a monitoring script to detect the working status of the NFS server and determine whether the NFS-Ganesha service is available. The monitoring script is executed once a minute, the timeout period is 30s, and the monitoring script weight is a negative value.
[0018] Preferably, the NFS-Ganesha service modifies the exported configuration to allow access rights to the client with the IP address 127.0.0.1.
[0019] Preferably, the NFS server working status detection is as follows:
[0020] Get the status of all exports of the NFS-Ganesha service: Use the ganesha_stats json export command to get the status statistics of all exports of the NFS-Ganesha service. The status statistics of all exports of the NFS-Ganesha service include the unique identifier of each export, the shared directory, and the usage status of the NFSv3, NFSv4.0, NFSv4.1, and NFSv4.2 protocols. The usage status of each exported NFS protocol represents whether there is a client mounting the corresponding exported shared directory using the corresponding protocol, which is represented by the values 0 and 1. 0 means that no client mounts the corresponding version of the NFS protocol, and 1 means that a client mounts the corresponding version of the NFS protocol. The number of clients must be at least 1.
[0021] Traverse all exports, exclude exports with identifier 0 and exports with NFS protocol usage status accumulated to 0, i.e., exports with no client mounting, to reduce the time and workload of status detection; the standby node executes the monitoring script and exits normally, because no client establishes a connection with the standby node, so the NFS protocol usage status accumulated to 0;
[0022] Perform a local mount test based on whether the monitoring operation count changes to determine whether the server is in an abnormal working state. Specifically:
[0023] Get the client operation counts of all versions of the currently exported NFS protocol: Use the ganesha_stats jsontotal id command to get the operation counts of different versions of the currently exported NFS protocol. The operation count is the total number of request operations of all clients that mount the corresponding export with the corresponding version of the NFS protocol processed by the NFS-Ganesha service;
[0024] Determine whether the operation counts of all versions of the currently exported NFS protocol have changed within the set detection time:
[0025] If the client operation counts of all versions do not change within the detection time, network problems, configuration problems, permission problems, and statistical tool problems are ruled out. This means that all currently exported clients have no active requests or the NFS server is in an abnormal working state and cannot process requests normally.
[0026] The shared directory exported by mounting the local loopback address has all been mounted for a specified number of times. The server node is used to perform a local mount test on the currently exported shared directory to directly detect the working status of the NFS server. The mount operation specifies that the IP address of the NFS service is the local loopback address 127.0.0.1, and sets the mount operation timeout.
[0027] If all mounts fail within the set number of tests, the NFS server is considered to be in an abnormal working state and cannot process requests normally, and the monitoring script execution is terminated with a non-zero exit status code;
[0028] If the keepalived service detects that the monitoring script exits abnormally, it will lower the priority value of the local machine. The BACKUP node has a higher priority, so it will preempt the MASTER node, VIP drift, and the MASTER node will change to BACKUP state. This preemption is very fast to ensure service continuity and availability; and configure the notification script for the MASTER node to change to BACKUP state, and the notification script will execute an elegant restart of the NFS-Ganesha process; when the original MASTER node executes the monitoring script and exits normally, that is, after it returns to normal, it will restore the original priority, but the configuration is BACKUP state and the priority is the same as the current MASTER node, so it will not be preempted again.
[0029] A high-availability NFS service system based on state detection, the system is used to implement the above-mentioned high-availability NFS service method based on state detection; the system comprises:
[0030] Deployment module 1 is used to deploy NFS-Ganesha services on all server nodes, and the server nodes share the CephFS storage system;
[0031] Deployment module 2 is used to deploy the keepalived service on all server nodes, configured with the same virtual IP address;
[0032] Status detection module, used to periodically execute status monitoring scripts, directly detect the working status of the NFS server by analyzing the output data of ganesha_stats and the local loopback address mount test;
[0033] The failover module is used to integrate into the keepalived service. When an abnormality is detected on the primary server, the service is automatically switched to the standby node and the necessary service restart is performed.
[0034] An electronic device comprising: a memory and at least one processor;
[0035] Wherein, the memory stores computer-executable instructions;
[0036] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the above-mentioned NFS service high availability method based on status detection.
[0037] A computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the above-mentioned NFS service high-availability method based on status detection is implemented.
[0038] The method, system, device and medium for high availability of NFS service based on status detection of the present invention have the following advantages:
[0039] (1) The present invention uses NFS-Ganesha as a high-performance, user-mode NFS server, integrates CephFS as the storage backend, and implements efficient fault switching by directly detecting the working status of the NFS server;
[0040] (ii) The present invention deploys NFS-Ganesha and keepalived services on all server nodes, and uses the same configuration to ensure service consistency;
[0041] (III) The present invention proposes a novel local loopback address mounting test strategy, which directly mounts the shared directory exported by its own NFS on the server node, eliminates the interference of network factors with the help of local loopback communication, and accurately determines the working status of the NFS server; if an abnormality is detected, the keepalived failover mechanism is triggered by changing the node priority, and the VIP is quickly switched to the backup node, and the NFS-Ganesha process is gracefully restarted on the original primary node to restore the service;
[0042] (iv) The present invention ensures high availability and reliability of NFS services through refined status detection and instant master-slave switching strategy;
[0043] (V) The present invention solves the problem of low reliability of NFS service caused by abnormal working state of NFS server in NFS high availability system. To ensure the reliability of network file system, the present invention directly detects the working state of the server on the NFS server. Once it is determined that the working state of the server is abnormal, it automatically switches to the standby node to provide external services, thereby ensuring stable operation of NFS service.
[0044] (VI) The present invention can timely discover abnormal working status of the NFS server, quickly complete the master-slave switch, improve the availability and reliability of the NFS service, effectively solve the problem of service unavailability caused by abnormal status of the NFS service, simplify the complexity of status detection, strengthen the system monitoring and fault recovery capabilities, and ensure the continuous and stable operation of the NFS service;
[0045] (VII) The status detection module of the present invention uses the JSON output of the ganesha_stats tool to perform in-depth data analysis to provide a more accurate basis for judging the working status of the server;
[0046] (VIII) The present invention also includes a notification script, which is automatically executed to achieve a graceful restart of the NFS-Ganesha service when the master node is downgraded to a backup node due to unavailability;
[0047] (IX) The present invention aims to provide an efficient and reliable NFS service high availability solution, which significantly improves the availability and reliability of the network file system through direct status detection and fast fault switching mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The present invention is further described below in conjunction with the accompanying drawings.
[0049] Attached Figure 1 The schematic diagram of the NFS service high availability method based on state detection;
[0050] Attached Figure 2 This is a flowchart of the NFS server working status detection process. DETAILED DESCRIPTION
[0051] The method, system, device and medium for high availability of NFS service based on status detection of the present invention are described in detail below with reference to the drawings and specific embodiments of the specification.
[0052] Embodiment 1:
[0053] This embodiment provides a method for high availability of NFS service based on status detection, and the method is specifically as follows:
[0054] S1. Use NFS-Ganesha as the user-mode NFS server and integrate CephFS as the backend storage system. The master and backup nodes share the same data. The NFS high-availability solution does not require a backend data synchronization mechanism. At the same time, deploy NFS-Ganesha services on all server nodes in the cluster and export CephFS to NFS through FSAL. CephFS (Ceph FileSystem) is a component of the distributed storage system Ceph. It is built on Ceph's RADOS object storage and implements a POSIX-compatible file system interface. NFS-Ganesha is a high-performance, scalable user-mode NFS server that can operate CephFS through the Ceph library libcephfs and export CephFS to NFS through the File System Abstraction Layer (FSAL).
[0055] S2. Deploy the keepalived service on all server nodes in the cluster for health monitoring and failover. All nodes share the same virtual IP address.
[0056] S3, adopt the local loopback address mounting strategy, by mounting the shared directory exported by itself on the server node, specifying the IP address of the NFS service as the local loopback address, to detect the server's ability to process requests;
[0057] S4. Use monitoring scripts to perform status checks regularly, including using the ganesha_stats tool to obtain runtime status information of the NFS-Ganesha service and analyzing the exported NFS protocol usage and operation count changes;
[0058] S5. When an abnormal working status of the server is detected, the keepalived failover mechanism is triggered by adjusting the node priority, automatically switching the VIP to the standby node, and gracefully restarting the NFS-Ganesha process on the original primary node.
[0059] The NFS-Ganesha service in step S1 of this embodiment is installed by compiling the source code, adding -DUSE_FSAL_CEPH=ON and -DUSE_RADOS_RECOV=ON in the compilation options, and setting the FSAL type to Ceph in the export (EXPORT) configuration to export CephFS to NFS through FSAL;
[0060] At the same time, add -DUSE_ADMIN_TOOLS=ON in the compilation options to use the state management tool provided by NFS-Ganesha.
[0061] This embodiment uses the distributed file system CephFS as the storage backend, and the primary and backup nodes share the same data, so the NFS high availability solution does not require a backend data synchronization mechanism.
[0062] The NFS-Ganesha service deployed by the active and standby nodes in step S2 of this embodiment uses the same configuration file, and writes all exported configurations into objects in the RADOS pool of Ceph respectively.
[0063] NFSv4 is a stateful protocol. The client and server maintain the file system status. The high availability of stateful services must maintain state consistency. When the server restarts, all state information and lock information will be lost. The NFSv4 protocol provides a recovery mechanism to re-establish the state, which allows the client to obtain the last connection state from the new server to restore the connection. Of course, if the time required to restore the connection exceeds the grace period of the server, the client can only re-establish the connection, as shown in the attached Figure 1 shown.
[0064] Keepalived is the mainstream lightweight high-availability software, which can perform health checks on servers or services and monitor their status. If a failure is detected in the master node (MASTER), it will automatically remove the node from the cluster and trigger a failover through the VRRP (Virtual Routing Redundancy Protocol) protocol, using a healthy backup node (BACKUP) to take over the work, and then rejoin the failed node to the cluster after it recovers.
[0065] The client in this embodiment accesses the NFS server through VIP. If a single point failure occurs in the primary server being accessed, the keepalived service automatically switches to the backup node and configures VIP to provide external services. The client accesses the new server and obtains the last connection status to restore the connection. If it fails to restore, the connection is re-established.
[0066] In this embodiment, the keepalived service deployed on all server nodes in the cluster uses the same configuration file, all host states are set to BACKUP, the priority is the same, the preemptive mode is set, and the same virtual IP address is specified;
[0067] Since there is no MASTER node at the beginning and the BACKUP nodes have the same priority, VRRP makes each BACKUP node participate in the election, and the winner is the MASTER node. The BACKUP node will not preempt the MASTER node unless its priority is higher. The keepalived service provides a health check method at the application layer, checking whether the server is running normally according to the user's settings; at the same time, it defines a monitoring script to detect the working status of the NFS server and determine whether the NFS-Ganesha service is available; among them, the monitoring script is executed once a minute, the timeout period is 30s, and the monitoring script weight is a negative value.
[0068] In order to directly detect the working status of the NFS server, this embodiment adopts a clever method: mount the shared directory exported by the NFS test on the NFS server node. In order to eliminate the influence of the network on the status detection, the IP address of the NFS service is specified as the local loopback address 127.0.0.1 when the server node performs the mount operation. Because the mount operation is performed on the server node, the client and the server are in the same system, and the loopback interface is used to achieve internal communication. Therefore, the IP address of the client that issues the mount request during this mount process is also 127.0.0.1, and then the NFS-Ganesha service needs to release the access rights of the client with the IP address of 127.0.0.1. This can be achieved by modifying the exported configuration, that is, the NFS-Ganesha service modifies the exported configuration to achieve the release of the access rights of the client with the IP address of 127.0.0.1.
[0069] As attached Figure 2 As shown, the NFS server working status detection in this embodiment is as follows:
[0070] ① Get the status of all exports of the NFS-Ganesha service: Use the ganesha_stats json export command to get the status statistics of all exports of the NFS-Ganesha service. The status statistics of all exports of the NFS-Ganesha service include the unique identifier of each export, the shared directory, and the usage status of the NFSv3, NFSv4.0, NFSv4.1, and NFSv4.2 protocols. The usage status of each exported NFS protocol represents whether there is a client mounting the corresponding exported shared directory with the corresponding protocol, which is represented by the values 0 and 1. 0 means that no client mounts with the corresponding version of the NFS protocol, and 1 means that a client mounts with the corresponding version of the NFS protocol. The number of clients is at least 1.
[0071] ② Traverse all exports, exclude exports with identifier 0 and exports with NFS protocol usage status accumulated to 0, i.e., exports with no client mounting, to reduce the time and workload of status detection; the standby node executes the monitoring script and exits normally, because no client establishes a connection with the standby node, so the NFS protocol usage status accumulated to 0;
[0072] ③ Perform a local mount test based on whether the monitoring operation count changes to determine whether the server is in an abnormal working state, specifically:
[0073] ④ Get the client operation counts of all versions of the currently exported NFS protocol: Use the ganesha_statsjson total id command to get the operation counts of different versions of the currently exported NFS protocol. The operation count is the total number of request operations of all clients that mount the corresponding export with the corresponding version of the NFS protocol processed by the NFS-Ganesha service;
[0074] ⑤ Determine whether the operation counts of all versions of the currently exported NFS protocol have changed within the set detection time:
[0075] If the client operation counts of all versions do not change within the detection time, network problems, configuration problems, permission problems, and statistical tool problems are ruled out. This means that all currently exported clients have no active requests or the NFS server is in an abnormal working state and cannot process requests normally.
[0076] ⑥ The shared directory exported by the local loopback address is mounted, and all mounts are timed out within the specified number of times: the server node is used to locally mount the currently exported shared directory to directly detect the working status of the NFS server. The mount operation specifies that the IP address of the NFS service is the local loopback address 127.0.0.1, and sets the mount operation timeout;
[0077] If all mounts fail within the set number of tests, the NFS server is considered to be in an abnormal working state and cannot process requests normally, and the monitoring script execution is terminated with a non-zero exit status code;
[0078] If the keepalived service detects that the monitoring script exits abnormally, it will lower the priority value of the local machine. The BACKUP node has a higher priority, so it will preempt the MASTER node, VIP drift, and the MASTER node will change to BACKUP state. This preemption is very fast to ensure service continuity and availability; and configure the notification script for the MASTER node to change to BACKUP state, and the notification script will execute an elegant restart of the NFS-Ganesha process; when the original MASTER node executes the monitoring script and exits normally, that is, after it returns to normal, it will restore the original priority, but the configuration is BACKUP state and the priority is the same as the current MASTER node, so it will not be preempted again.
[0079] Embodiment 2:
[0080] This embodiment provides a high-availability NFS service system based on state detection, which is used to implement the high-availability NFS service method based on state detection in Example 1; the system includes:
[0081] Deployment module 1 is used to deploy NFS-Ganesha services on all server nodes, and the server nodes share the CephFS storage system;
[0082] Deployment module 2 is used to deploy the keepalived service on all server nodes, configured with the same virtual IP address;
[0083] Status detection module, used to periodically execute status monitoring scripts, directly detect the working status of the NFS server by analyzing the output data of ganesha_stats and the local loopback address mount test;
[0084] The failover module is used to integrate into the keepalived service. When an abnormality is detected on the primary server, the service is automatically switched to the standby node and the necessary service restart is performed.
[0085] Embodiment 3:
[0086] This embodiment also provides an electronic device, including: a memory and at least one processor;
[0087] Wherein, the memory stores computer-executable instructions;
[0088] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the NFS service high availability method based on status detection as described in any one of the present inventions.
[0089] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.
[0090] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.
[0091] Embodiment 4:
[0092] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor, so that the processor executes the NFS service high availability method based on state detection in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0093] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0094] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code may be downloaded from a server computer via a communication network.
[0095] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0096] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for high availability of NFS service based on state detection, characterized in that: The method is as follows: Use NFS-Ganesha as the user-mode NFS server and integrate CephFS as the backend storage system. The master and slave nodes share the same data. Deploy NFS-Ganesha services on all server nodes in the cluster and export CephFS to NFS through FSAL. Deploy the keepalived service on all server nodes in the cluster for health monitoring and failover. All nodes share the same virtual IP address. The local loopback address mounting strategy is adopted. By mounting the shared directory exported by the server node, the IP address of the NFS service is specified as the local loopback address to detect the server's ability to process requests. Use monitoring scripts to perform status checks regularly, including using the ganesha_stats tool to obtain runtime status information of the NFS-Ganesha service and analyzing the exported NFS protocol usage and operation count changes; When an abnormal working status of the server is detected, the keepalived failover mechanism is triggered by adjusting the node priority, automatically switching the VIP to the standby node, and gracefully restarting the NFS-Ganesha process on the original primary node.
2. The method for high availability of NFS service based on state detection according to claim 1, characterized in that: The NFS-Ganesha service is installed by compiling the source code. Add -DUSE_FSAL_CEPH=ON and -DUSE_RADOS_RECOV=ON in the compilation options, and set the FSAL type to Ceph in the exported configuration to export CephFS to NFS through FSAL. At the same time, add -DUSE_ADMIN_TOOLS=ON in the compilation options to use the state management tool provided by NFS-Ganesha.
3. The NFS service high availability method based on state detection according to claim 1 or 2, characterized in that: The NFS-Ganesha services deployed on the active and standby nodes use the same configuration file and write all exported configurations to objects in the Ceph RADOS pool.
4. The method for high availability of NFS service based on state detection according to claim 3 is characterized in that: The client accesses the NFS server through VIP. If a single point failure occurs in the primary server being accessed, the keepalived service automatically switches to the backup node and configures VIP to provide external services. The client accesses the new server and obtains the last connection status to restore the connection. If it fails to recover, it re-establishes the connection.
5. The method for high availability of NFS service based on state detection according to claim 4 is characterized in that: The keepalived service deployed on all server nodes in the cluster uses the same configuration file, all host states are set to BACKUP, the priority is the same, the preemptive mode is set, and the same virtual IP address is specified; The keepalived service provides an application layer health check method to check whether the server is running normally according to the user's settings. At the same time, it defines a monitoring script to detect the working status of the NFS server and determine whether the NFS-Ganesha service is available. The monitoring script is executed once a minute, the timeout period is 30s, and the monitoring script weight is a negative value.
6. The method for high availability of NFS service based on state detection according to claim 5, characterized in that: The NFS-Ganesha service modifies the exported configuration to allow access permissions for clients with IP address 127.0.0.
1.
7. The method for high availability of NFS service based on state detection according to claim 6, characterized in that: The NFS server working status detection is as follows: Get the status of all exports of the NFS-Ganesha service: Use the ganesha_stats json export command to get the status statistics of all exports of the NFS-Ganesha service. The status statistics of all exports of the NFS-Ganesha service include the unique identifier of each export, the shared directory, and the usage status of the NFSv3, NFSv4.0, NFSv4.1, and NFSv4.2 protocols. The usage status of each exported NFS protocol represents whether there is a client mounting the corresponding exported shared directory using the corresponding protocol, which is represented by the values 0 and 1. 0 means that no client mounts the corresponding version of the NFS protocol, and 1 means that a client mounts the corresponding version of the NFS protocol. The number of clients must be at least 1. Traverse all exports and exclude exports with identifier 0 and exports with NFS protocol usage status accumulated to 0, i.e., no client mounts them; the standby node executes the monitoring script and exits normally; Perform a local mount test based on whether the monitoring operation count changes to determine whether the server is in an abnormal working state. Specifically: Get the client operation counts of all versions of the currently exported NFS protocol: Use the ganesha_stats jsontotal id command to get the operation counts of different versions of the currently exported NFS protocol. The operation count is the total number of request operations of all clients that mount the corresponding export with the corresponding version of the NFS protocol processed by the NFS-Ganesha service; Determine whether the operation counts of all versions of the currently exported NFS protocol have changed within the set detection time: If the client operation counts of all versions do not change within the detection time, network problems, configuration problems, permission problems, and statistical tool problems are ruled out. This means that all currently exported clients have no active requests or the NFS server is in an abnormal working state and cannot process requests normally. The shared directory exported by mounting the local loopback address has all been mounted for a specified number of times. The server node is used to perform a local mount test on the currently exported shared directory to directly detect the working status of the NFS server. The mount operation specifies that the IP address of the NFS service is the local loopback address 127.0.0.1, and sets the mount operation timeout. If all mounts fail within the set number of tests, the NFS server is considered to be in an abnormal working state and cannot process requests normally, and the monitoring script execution is terminated with a non-zero exit status code; If the keepalived service detects that the monitoring script exits abnormally, it will lower the priority value of the local machine. The BACKUP node has a higher priority and will preempt the MASTER node. The VIP will drift and the MASTER node will change to the BACKUP state. A notification script will be configured for the MASTER node to change to the BACKUP state. The notification script will gracefully restart the NFS-Ganesha process. When the original MASTER node exits normally after executing the monitoring script, that is, when it returns to normal, the original priority will be restored.
8. A high availability system for NFS service based on status detection, characterized in that: The system is used to implement the NFS service high availability method based on state detection as described in any one of claims 1 to 7; The system includes: Deployment module 1 is used to deploy NFS-Ganesha services on all server nodes, and the server nodes share the CephFS storage system; Deployment module 2 is used to deploy the keepalived service on all server nodes, configured with the same virtual IP address; Status detection module, used to periodically execute status monitoring scripts, directly detect the working status of the NFS server by analyzing the output data of ganesha_stats and the local loopback address mount test; The failover module is used to integrate into the keepalived service. When an abnormality is detected on the primary server, the service is automatically switched to the standby node and the necessary service restart is performed.
9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the NFS service high availability method based on status detection as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the NFS service high-availability method based on status detection as described in any one of claims 1 to 7 is implemented.