State recovery method and system of cluster type network file system
By taking over the virtual IP address of the target server in the clustered network file system and reading the file status operation request record in the shared recovery module, the problem of inconsistent file status after server failure in the distributed clustered network file system is solved, and high-reliability file status recovery is achieved.
Patent Information
- Application Number
- CN202411904887.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
AI Technical Summary
In a distributed cluster network file system, after a server failure, the file status is inconsistent, which may lead to data conflicts, business interruptions and data loss. The prior art only focuses on the consistency of locks, does not involve the management of other states, and has limitations.
A state recovery method and system for a clustered network file system is provided. By taking over the virtual IP address of the target server, entering a grace period to rebuild the file state, reading the file state operation request record in the shared recovery module, and performing file operation requests in sequence according to these records to restore the file state.
It realizes that even if the server goes down, the file state can be restored, the reliability of file state recovery can be improved, and multiple file states can be restored, and there is no limitation of file state recovery.
Smart Images

Figure CN119988089A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of network file systems, and in particular to a state recovery method and system for a clustered network file system. Background Art
[0002] In the field of NAS (Network Attached Storage) file storage, NFS (Network File System) is an important network file storage protocol under the operating system. Multiple nodes work together to provide distributed cluster network file system services. This system consists of multiple server nodes, each of which can provide network file system services and share the same file system.
[0003] However, in a distributed cluster network file system, when a main server node (such as node A) fails, the system will let the backup server node (such as node B) take over the request of client a in order to ensure high availability. However, when the client accesses the main service node (node A), it may have performed status operations (such as locking) on certain files, and node B does not hold the status data of these files. When client a accesses node B instead, the file status it holds is inconsistent with the status on the server. At this time, if other clients (such as client b) try to lock the same file through another server node (such as node C), they may succeed, resulting in a lock conflict. This may cause the following problems:
[0004] 1) Data inconsistency: Multiple clients modifying the same file at the same time may cause data conflicts or file corruption. Since NFS is a distributed system, operations between clients are not synchronized immediately, which further exacerbates data confusion.
[0005] 2) Business interruption: Lock conflicts usually lead to abnormal file access. The client's file read and write operations may be blocked or interrupted, affecting normal business processes, especially in environments that require a large number of concurrent file accesses.
[0006] 3) Data loss: Due to the unclear locking status, the file system may not be able to properly save the write content of all clients, resulting in data loss.
[0007] Therefore, the publication number CN115934743B "A file lock management method, system, device and computer-readable storage medium" proposes receiving a lock application from the client; when the target file already has a lock, determining the current server holding the lock; after the file lock is transferred, generating and saving the first lock record information, and updating the second lock record information of the current server at the same time, recording the transfer of the lock. By querying the two lock records, the transfer of file locks between different servers can be accurately known, thereby solving the data consistency problem. However, this method saves the lock information in the local data of the node. If the local database is lost for some reason, the file lock status will not be completely restored, and there is a risk. Moreover, the file status not only includes the lock, but also the file open status and lease (such as delegation). However, the prior art only focuses on the consistency of the lock, and does not involve the management of other states, which has limitations. Summary of the invention
[0008] Based on this, in view of the above technical problems, a state recovery method and system for a clustered network file system are provided to solve the problem of low reliability when performing state recovery in the prior art.
[0009] In a first aspect, a state recovery method of a clustered network file system is applied to a current server, and the method includes:
[0010] When the target server fails, after obtaining the virtual IP address of the target server, the service of the target server is taken over and a grace period is entered to rebuild the file status; the client that executes the file operation request through the target server is recorded as the target client; during the grace period, new non-file status reconstruction operation requests sent by the target client are rejected;
[0011] Reading the IP address of the target client associated with the target client stored in the shared recovery module to send a recovery operation request to the target client information;
[0012] Receive a pre-stored file status operation request record sent by a target client, and send the IP address information of the target client to the shared recovery module, so that the shared recovery module updates the IP address information of the client corresponding to the current server; the stored file status operation request record is all file status operation records sent by the target client to the target server;
[0013] Execute all file operation requests in sequence according to the file operation request record, thereby rebuilding the file status;
[0014] The grace period ends and a new file operation request sent by the target client is received.
[0015] In the above solution, optionally, the file operation request includes: file opening, file locking and file lease.
[0016] In the above solution, further optionally, the shared recovery module uses a Ceph cluster, or an HBase cluster, or a GlusterFS cluster as a database.
[0017] In the above solution, optionally, after receiving the pre-stored file status operation request record sent by the target client, the method further includes: storing the file status operation request record.
[0018] In the above solution, further optionally, after receiving the new file operation request sent by the target client, the method further includes storing the new file operation request.
[0019] In the above solution, optionally, before obtaining the virtual IP address of the target server, the following steps are further included:
[0020] If the current server is a newly added node of the cluster network file system, after inserting the record of the current server into the shared recovery module, initialization starts.
[0021] In the above solution, optionally, the shared recovery module also stores a serial number record of the file status operation request.
[0022] In a second aspect, a state recovery system for a clustered network file system is provided, the system comprising:
[0023] The target server virtual IP acquisition module is used to obtain the virtual IP address of the target server after the target server fails, take over the business of the target server, enter a grace period to rebuild the file status; record the client that executes the file operation request through the target server as the target client; and refuse to receive new non-file status reconstruction operation requests sent by the target client during the grace period;
[0024] The recovery operation request sending module is used to read the IP address of the target client associated with the target client stored in the shared recovery module based on the RADOS cluster storage, so as to send a recovery operation request to the target client information;
[0025] The file status operation request record receiving module is used to receive the pre-stored file status operation request record sent by the target client, and send the IP address information of the target client to the shared data recovery module, so that the shared data recovery module updates the IP address information of the client corresponding to the current server; the stored file status operation request record is all the file status operation records sent by the target client to the target server;
[0026] Status reconstruction module: used to execute all file operation requests in sequence according to the file operation request record, so as to rebuild the file status;
[0027] Normal receiving request module: used to end the grace period and receive new file operation requests sent by the target client.
[0028] In a third aspect, a computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the state recovery method of a cluster network file system described in the first aspect are implemented.
[0029] In a fourth aspect, a computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of a state recovery method for a clustered network file system as described in the first aspect.
[0030] This application has at least the following beneficial effects:
[0031] When the present application performs file status operations through the server via the client, the client information processed by the server is stored in the shared data module. After the client sends a file operation request to the server, the file operation request will be recorded in its own storage area. After the current server fails, the current server takes over the virtual IP address of the target server, finds the IP address of the client processed by the target server from the shared data recovery module, and sends a recovery status request to these clients, so that these clients send the stored file operation request records to the target server, so that the current server performs status recovery according to the file operation request records. In this way, even if the server is down, the file status can be restored, improving the reliability of file status recovery. In addition, the present application method can restore multiple file states, and there is no limitation in file status recovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A schematic diagram of a state recovery method for a clustered network file system provided in one embodiment of the present application;
[0033] Figure 2 A diagram showing the connection relationship between the various parts of a network file system provided in one embodiment of the present application;
[0034] Figure 3 A schematic diagram of a grace period provided for an embodiment of the present application;
[0035] Figure 4 A detailed flowchart of a state recovery method for a clustered network file system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0037] In one embodiment, Figure 1 as well as Figure 2 As shown, a state recovery method of a cluster network file system is provided, which is applied to the current server, and the method includes:
[0038] Step S1: When a target server fails, after obtaining the virtual IP address of the target server, take over the business of the target server and enter a grace period to rebuild the file status; the client that executes the file operation request through the target server is recorded as the target client; during the grace period, refuse to receive new non-file status reconstruction operation requests sent by the target client.
[0039] In step S1, in a cluster distributed file system, the server may lose its state after restarting due to a crash. In order to ensure that the server can rebuild the file state, a grace period (graceperiod) needs to be designed. During this period, the system will prohibit the execution of new non-file state recovery requests and only allow requests to restore file state. After the grace period, the server will enter the normal state (normal). In short, after the connection between the client and the server is interrupted, if it times out, it will enter the grace period for recovery, and during the recovery process, the client is prohibited from obtaining a new state, and only the old state is allowed to be recycled. In a cluster network file system, the state changes of nodes are divided into recovery (recovery) and normal (normal) states. The life cycle of each server node can be regarded as a series of transition processes from the recovery period to the normal period. This process is called an "epoch." As Figure 3 As shown in the figure, a grace period includes recovery and normal states, represented by R and N in the figure respectively. In a grace period, there may be multiple recovery (R) phases and one normal (N) phase. If the target server node goes down multiple times, there may be multiple recovery phases (R) in a grace period. During the grace period, when the server restarts and enters the recovery state, it must declare entering a new grace period by increasing the current epoch value and reclaim the client state. The time state changes of the grace period are as follows: Figure 3 shown.
[0040] Step S2: Read the IP address of the target client associated with the target client stored in the shared recovery module to send a recovery operation request to the target client information.
[0041] In step S2, for the server, when a node fails, in order to ensure the persistent preservation of the client status information and effectively save it in the cluster network file system, a shared data module needs to be designed. As a distributed object storage system, the RADOS cluster storage system can better undertake shared storage tasks, save the information of the entire cluster server nodes and the required status recovery information. The shared recovery module includes the GRACE database and the server node information database. The GRACE database records the status of each node and the server node ID. The node status information of the entire cluster can be known through the GRACE database; each service node has its own server node information database, and records the IP address of the client being processed and other related information.
[0042] rados-grace-tool is a command-line tool for operating the GRACE database. It can query or manage the status of each node in the cluster. In addition, before initializing the cluster network file system, a GRACE database must be generated in the RADO S pool through the tool rados-grace-tool, which records the status of all servers in the cluster. With this mechanism, the server node can perceive the storage status of other nodes. When the cluster is expanded, the newly added server nodes need to insert their records into the GRACE database through rado s-grace-tool before they can be initialized and started, otherwise the server nodes cannot start normally.
[0043] Step S3: Receive the pre-stored file status operation request record sent by the target client, and send the IP address information of the target client to the shared recovery module, so that the shared recovery module updates the IP address information of the client corresponding to the current server; the stored file status operation request record is all file status operation records sent by the target client to the target server.
[0044] In step S3, when the client performs a stateful file operation through the server, the client will record the operation. When receiving a recovery notification from the file system server, the client will resend the previously recorded file status operation request to restore the file status of the server. For example, if file f1 is locked, the client will send a request to lock file f1 again. The recovery of the file status depends on the status data stored by the client.
[0045] Step S4: executing all file operation requests in sequence according to the file operation request record, thereby rebuilding the file status;
[0046] Step S5: End the grace period and receive a new file operation request sent by the target client.
[0047] In the above-mentioned state recovery method of a clustered network file system, when the client performs file state operations through the server, the client information processed by the server is stored in the shared data module. After the client sends a file operation request to the server, the file operation request will be recorded in its own storage area. After the current server fails, the current server takes over the virtual IP address of the target server, finds the IP address of the client processed by the target server from the shared data recovery module, and sends a recovery state request to these clients, so that these clients send the stored file operation request records to the target server, so that the current server performs state recovery according to the file operation request records. In this way, even if the server is down, the file state can be restored, improving the reliability of file state recovery. In addition, the present application method can restore a variety of file states, and there is no limitation in file state recovery.
[0048] In one embodiment, the file operation request includes: file opening, file locking, and file lease.
[0049] In one embodiment, the shared recovery module uses a Ceph cluster, an HBase cluster, or a GlusterFS cluster as a database.
[0050] In one embodiment, after receiving the pre-stored file status operation request record sent by the target client, the method further includes: storing the file status operation request record.
[0051] In one embodiment, after receiving the new file operation request sent by the target client, the step further includes storing the new file operation request.
[0052] In this embodiment, when the client sends an operation with a file status, the server should record and correspond to the client information to ensure that the server can notify the corresponding client when the status is restored. When the server performs status recovery, other file modification operations should be rejected. After the server is restarted, it should enter the status recovery phase and perform normal business processing after the status recovery is completed. At the same time, other server nodes in the same cluster also notify other nodes to maintain synchronization when one of the nodes is in the recovery state.
[0053] In one embodiment, before obtaining the virtual IP address of the target server, the method further includes:
[0054] If the current server is a newly added node of the cluster network file system, after inserting the record of the current server into the shared recovery module, initialization starts.
[0055] In one embodiment, the shared data module also stores a serial number record of the file status operation request.
[0056] In the state recovery scheme of the clustered network file system, state recovery is implemented by the client, server, and shared data module. Each module plays a different role and stores its own data, as follows:
[0057] The data recorded by the client and the server are generally similar. For example, when the client opens a file and locks it, the client will record the corresponding operation request, mainly including the two states of opening and locking. The server also stores the two states of opening the file and locking. The main data recorded by both include the following aspects:
[0058] 1. File lock status: Records the file lock information held by the client to ensure that the client can continue to maintain exclusive access to the file after the status is restored, avoiding data conflicts caused by the loss of lock information.
[0059] 2. File open (OPEN) status and close (CLOSE) status: The OPEN operation records the status of the client opening the file, including information such as the file descriptor and access mode, to ensure that the client can directly operate the opened file after recovery without having to re-execute the file open operation. The CLOSE operation records the file closing status and releases the status information accumulated by the OPEN operation. For example, if the file is locked, the CLOSE operation will release the file lock.
[0060] 3. Session status of communication between client and server: Restore the status of NFSv4 session, including information such as the requested sequence ID, to ensure the order and consistency of communication.
[0061] 4. File and Directory Handles: Restore the file and directory handles held by the client to support continued access to related resources.
[0062] 5. Client Lease: When a network partition occurs and the lease time has expired, the server will automatically release all locks held by the client. If the client does not renew the lease within the defined period, all lease states associated with the client will be released by the server. During the file state reconstruction process, the server will restore the client's lease state in order to manage the life cycle of the file state, thereby ensuring the consistency of the state after restoration.
[0063] 6. Delegation State: When the server grants a file delegation to a client, it ensures that the client can exclusively share the file with other clients in certain semantic aspects. During an OPEN operation, the server can provide the client with a read or write delegation for the file. If the client is granted a read delegation, the server will ensure that other clients cannot write to the file during the validity period of the delegation; if the client is granted a write delegation, it can be sure that other clients cannot read or write to the file during the delegation. If another client requests access to the file in a way that conflicts with the granted authorization, the server will notify the initial client and revoke the authorization. This requires a valid callback path between the server and the client; if the callback path does not exist, the delegation cannot be granted. The core of delegation is that it allows clients to perform file operations (such as OPEN, CLOSE, LOCK, LOCKU, READ, WRITE) locally without immediately interacting with the server.
[0064] 7. Layout State: Restore the layout state of NFSv4.1 and pNFS clients to maintain the normal operation of the distributed file system.
[0065] The shared recovery module based on RADOS storage mainly records the following data:
[0066] 1. Client and server node information: including key information such as the client node's IP address and server node ID, which is used to identify and manage client nodes and server nodes.
[0067] 2. Life cycle of server nodes: The life cycle of a server node is represented by the epoch value. The epoch records a series of transitions from the recovery period to the normal period. The module records the latest epoch value of the cluster and the current epoch value of each server node to track its state changes.
[0068] 3. Cluster status flags: When a node joins or is removed from a cluster, its status includes two flags: (1) N (NEED): Indicates whether the node needs to process client status reconstruction operations. If necessary, this flag will be set; (2) E (ENFORCING): Indicates whether the node is enforcing the grace period. After the cluster enters the normal phase, the server will clear these two flags, indicating that the node has returned to normal working status and can continue to process regular I / O requests and client operations. The use of these flags helps to indicate the operating status of the node, thereby achieving unified management and monitoring of the cluster.
[0069] In a distributed NFS cluster network file system, dynamic expansion of the number of nodes is an indispensable function that can effectively improve the scalability and fault tolerance of the system. When the cluster load increases or storage requirements change, dynamic expansion can add new nodes to share the load, improve performance, or expand storage capacity without interrupting services. The following example illustrates the process of cluster node expansion:
[0070] 1. Use the rados-grace-tool tool to insert the ID of the new server node (such as node B) into the GRACE database in the shared recovery module based on RADOS storage to ensure that the node can correctly participate in the state recovery and management of the cluster.
[0071] 2. Start the new server node B and keep it running normally to ensure that it can communicate and synchronize with other nodes in the cluster.
[0072] 3. Server node B is a newly added node. Other nodes in the cluster will automatically detect the addition of the node and trigger the cluster to enter a grace period. During this period, all nodes in the cluster will stop accepting new file operation requests and only allow operations to restore file status to ensure system consistency.
[0073] 4. After the grace period ends, the entire cluster will return to normal state and resume normal file operations and I / O request processing. At this point, the file status of all nodes has been restored, and the system can continue to run and provide complete services.
[0074] In the NFS protocol, state recovery refers to re-establishing the state information between the client and the server to maintain service consistency after a server restart, failure, or network interruption. The steps involved in file state recovery can be summarized as follows: Figure 4 As shown, the server nodes include A and B; the client node includes a; the detailed steps are described as follows:
[0075] 1. Server node B detects that server node A has gone offline due to a fault. At this time, all the status of server node A in memory is lost, and server node A cannot provide normal I / O services.
[0076] 2. Client A communicates with the server node through the virtual IP address. When server node A goes down, the virtual IP will automatically switch, server node B obtains the virtual IP, and takes over the business of server node A. At this time, server node B enters the grace period, the epoch value increases, and switches from normal state to recovery state.
[0077] 3. The entire cluster node enters a grace period, and server node B begins to prohibit the execution of new non-file status reconstruction operation requests sent by client A, and only allows the execution of file status recovery requests. In the recovery phase, the status data of client a held by server node A is synchronized to the database of node B in the cluster.
[0078] 4. The server node reads the RADO S cluster shared recovery module server node information database, obtains the IP and other information of all clients communicating with the failed server node A, and notifies the client node a to send the corresponding file status reconstruction operation request, such as file opening, locking, delegation status, etc.
[0079] 5. Client node a receives the message from server node B and resends the previously recorded file status operation request, thereby re-establishing all file statuses on server node B.
[0080] 6. After receiving the message from client node A, server node B updates and records the corresponding client information. When restoring the file status next time, the server can identify and notify the relevant client, asking it to send a recovery operation request. These requests include operations such as re-locking locked files and reopening opened files to ensure that all states of the file are rebuilt.
[0081] 7. After the server completes file status recovery and re-establishes all status information, it transitions from the recovery phase to the normal phase, and begins to receive normal I / O requests.
[0082] Therefore, this application can realize the recovery and migration of file status for the network file system in the distributed cluster, ensure that the file status is not lost, and maintain the consistency of the file status. This helps to prevent business anomalies in the scenario of stateful file locking. Storing the client information of the server node, etc. in the third-party Ceph cluster RADOS pool can effectively avoid the inability to rebuild the file status due to the loss of client information and ensure the reliability of data.
[0083] In one embodiment, a state recovery system for a clustered network file system includes:
[0084] The target server virtual IP acquisition module is used to obtain the virtual IP address of the target server after the target server fails, take over the business of the target server, enter a grace period to rebuild the file status; record the client that executes the file operation request through the target server as the target client; and refuse to receive new non-file status reconstruction operation requests sent by the target client during the grace period;
[0085] A recovery operation request sending module: used to read the IP address of the target client associated with the target client stored in the shared recovery module, so as to send a recovery operation request to the target client information;
[0086] The file status operation request record receiving module is used to receive the pre-stored file status operation request record sent by the target client, and send the IP address information of the target client to the shared data recovery module, so that the shared data recovery module updates the IP address information of the client corresponding to the current server; the stored file status operation request record is all the file status operation records sent by the target client to the target server;
[0087] Status reconstruction module: used to execute all file operation requests in sequence according to the file operation request record, so as to rebuild the file status;
[0088] Normal receiving request module: used to end the grace period and receive new file operation requests sent by the target client.
[0089] The specific definition of a state recovery system for a clustered network file system can be found in the definition of a state recovery method for a clustered network file system in the above text, which will not be repeated here. Each module in the above-mentioned state recovery system for a clustered network file system can be implemented in whole or in part by software, hardware, and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each of the above modules.
[0090] In one embodiment, a computer device is provided, which may be a server, and includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the state recovery method of the cluster network file system described above is implemented.
[0091] In one embodiment, a computer program product is also provided, including a computer program / instruction. When the computer program / instruction is executed by a processor, it involves all or part of the process in the above-mentioned embodiment method.
[0092] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0093] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0094] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A state recovery method for a cluster network file system, characterized in that: Applied to the current server, the method includes: When the target server fails, after obtaining the virtual IP address of the target server, the service of the target server is taken over and a grace period is entered to rebuild the file status; the client that executes the file operation request through the target server is recorded as the target client; during the grace period, new non-file status reconstruction operation requests sent by the target client are rejected; Reading the IP address of the target client associated with the target client stored in the shared recovery module to send a recovery operation request to the target client information; Receive a pre-stored file status operation request record sent by a target client, and send the IP address information of the target client to the shared recovery module, so that the shared recovery module updates the IP address information of the client corresponding to the current server; the stored file status operation request record is all file status operation records sent by the target client to the target server; Execute all file operation requests in sequence according to the file operation request record, thereby rebuilding the file status; The grace period ends and a new file operation request sent by the target client is received.
2. The state recovery method of a cluster network file system according to claim 1, characterized in that: The file operation request includes: file opening, file locking and file lease.
3. The method according to claim 1, characterized in that The shared recovery module uses a Ceph cluster, an HBase cluster, or a GlusterFS cluster as a database.
4. The state recovery method of a cluster network file system according to claim 1, characterized in that: After receiving the pre-stored file status operation request record sent by the target client, the method further includes: storing the file status operation request record.
5. The state recovery method of a cluster network file system according to claim 4, characterized in that: After receiving the new file operation request sent by the target client, the method further includes storing the new file operation request.
6. The state recovery method of a cluster network file system according to claim 1, characterized in that: Before obtaining the virtual IP address of the target server, it also includes: If the current server is a newly added node of the cluster network file system, after inserting the record of the current server into the shared recovery module, initialization starts.
7. The state recovery method of a cluster network file system according to claim 1, characterized in that: The shared recovery module also stores a sequence number record of the file status operation request.
8. A state recovery system for a clustered network file system, characterized in that: The system comprises: The target server virtual IP acquisition module is used to obtain the virtual IP address of the target server after the target server fails, take over the business of the target server, enter a grace period to rebuild the file status; record the client that executes the file operation request through the target server as the target client; and refuse to receive new non-file status reconstruction operation requests sent by the target client during the grace period; A recovery operation request sending module: used to read the IP address of the target client associated with the target client stored in the shared recovery module, so as to send a recovery operation request to the target client information; The file status operation request record receiving module is used to receive the pre-stored file status operation request record sent by the target client, and send the IP address information of the target client to the shared data recovery module, so that the shared data recovery module updates the IP address information of the client corresponding to the current server; the stored file status operation request record is all the file status operation records sent by the target client to the target server; Status reconstruction module: used to execute all file operation requests in sequence according to the file operation request record, so as to rebuild the file status; Normal receiving request module: used to end the grace period and receive new file operation requests sent by the target client.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
A file lock management method, system, device, and computer-readable storage medium
CN115934743B