File system data processing method and related equipment
By maintaining and processing queues for each directory in the file system, and updating the metadata database and metadata cache in queue order, combined with the parent directory lock mechanism, the problems of metadata update consistency and correctness are solved, and efficient metadata management is achieved.
Patent Information
- Application Number
- CN202311762216.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the file system, how to ensure the consistency and correctness of metadata updates, especially when metadata is stored in two storage spaces: metadata and metadata cache.
By receiving file operation requests, it is added to the processing queue of the corresponding directory, and the metadata cache is updated first according to the order in the queue. Use the parent directory lock mechanism to ensure that file operation requests under the same parent directory are updated in sequence.
It realizes the consistency and correctness of metadata updates, ensures strong consistency between metadata database and metadata cache, and improves the efficiency of metadata updates.
Smart Images

Figure CN120179660A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, in particular to the field of file system technology, and specifically relates to a method for processing file system data, an apparatus for processing file system data, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] A file system is a method for storing and organizing data. File system data can include metadata and file data. Metadata can be used to summarize basic information of file data (such as directory information, attribute information, and storage location information, etc.). Metadata is usually stored in a metadata database. To improve the access efficiency of metadata, metadata can also be stored in a metadata cache. That is to say, metadata is stored in two storage spaces, one is the metadata database, and the other is the metadata cache. When there is a file operation request in the file system, the metadata needs to be updated in response to the file operation request. For metadata stored in two storage spaces, how to ensure the consistency and correctness of metadata updates has become a current research hotspot. Summary of the Invention
[0003] Embodiments of this application provide a method for processing file system data and related devices, which can ensure the consistency and correctness of metadata updates.
[0004] On the one hand, embodiments of this application provide a method for processing file system data. File system data includes metadata and file data. The metadata is stored in a metadata cache and a metadata database; the file data includes multiple files, and the metadata includes description information of each file in the file data; the method for processing the file system data includes:
[0005] Receiving a target file operation request, where the file requested to be operated by the target file operation request is a target file, and the parent directory to which the target file belongs is a first directory;
[0006] Adding the target file operation request to a processing queue corresponding to the first directory; the processing queue corresponding to the first directory includes at least one file operation request. The files requested to be processed by the file operation requests in the processing queue corresponding to the first directory all belong to the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time;
[0007] Updating the metadata database according to the target file operation request;
[0008] Based on the metadata database update result of the target file operation request, updating the metadata cache in the order of the target file operation request in the processing queue corresponding to the first directory.
[0009] Accordingly, an embodiment of the present application provides a processing device for file system data. The file system data includes metadata and file data. The metadata is stored in a metadata cache and a metadata database; the file data includes multiple files, and the metadata includes description information of each file in the file data. The processing device for the file system data includes:
[0010] A communication unit, configured to receive a target file operation request. The file requested to be operated by the target file operation request is a target file, and the parent directory to which the target file belongs is a first directory;
[0011] A processing unit, configured to add the target file operation request to a processing queue corresponding to the first directory; the processing queue corresponding to the first directory includes at least one file operation request. The files requested to be processed by the file operation requests in the processing queue corresponding to the first directory all belong to the first directory as the parent directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time;
[0012] The processing unit is further configured to update the metadata database according to the target file operation request;
[0013] The processing unit is further configured to update the metadata cache based on the metadata database update result of the target file operation request, in the arrangement order of the target file operation request in the processing queue corresponding to the first directory.
[0014] In one implementation, the processing queue corresponding to the first directory is located in a metadata update structure. The metadata update structure at least further includes a processing queue corresponding to a second directory; the first directory and the second directory are different directories;
[0015] The processing queue corresponding to the first directory and the processing queue corresponding to the second directory concurrently execute the operation of updating the metadata; the operation of updating the metadata includes updating the metadata database and updating the metadata cache.
[0016] In one implementation, each file operation request in the processing queue corresponding to the first directory updates the metadata cache in sequence based on the corresponding file operation request; the head of the processing queue corresponding to the first directory has a parent directory lock, and the file operation request that obtains the parent directory lock has the permission to update the metadata cache;
[0017] When the processing unit updates the metadata cache based on the metadata database update result of the target file operation request, in the arrangement order of the target file operation request in the processing queue corresponding to the first directory, it is specifically configured to perform the following steps:
[0018] When the target file operation request is at the head of the queue, add the parent directory lock to the target file operation request;
[0019] If the metadata database update result of the target file operation request indicates that the metadata database is successfully updated according to the target file operation request, then update the metadata cache according to the target file operation request.
[0020] In one implementation, the processing unit is further configured to perform the following steps:
[0021] If the metadata database update result of the target file operation request indicates that the update of the metadata database according to the target file operation request fails, then release the target file operation request from the head of the queue;
[0022] Move the file operation request that is one position behind the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
[0023] In one implementation, the processing unit is further configured to perform the following steps:
[0024] After successfully updating the metadata cache according to the target file operation request, release the target file operation request from the head of the queue;
[0025] Move the file operation request that is one position behind the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
[0026] In one implementation, the operation request type of the target file operation request is the target operation request type. After the target file operation request is structured into the target file operation log corresponding to the target operation request type, it is added to the processing queue corresponding to the first directory; the target file operation log includes the operation data of the target file operation request.
[0027] When the processing unit is used to update the metadata database according to the target file operation request, it is specifically configured to perform the following steps:
[0028] Update the metadata database according to the operation data of the target file operation request in the target file operation log;
[0029] When the processing unit is used to update the metadata cache based on the metadata database update result of the target file operation request in the order of arrangement of the target file operation request in the processing queue corresponding to the first directory, it is specifically configured to perform the following steps:
[0030] Based on the metadata database update result of the target file operation request, in the order of arrangement of the target file operation request in the processing queue corresponding to the first directory, update the metadata cache according to the operation data of the target file operation request in the target file operation log.
[0031] In one implementation, the parent directories to which different files in the file data belong form a directory tree according to the hierarchical relationship, and the metadata processing structure is used to store the processing queues corresponding to the parent directories at different levels in the directory tree; the processing unit, when adding a target file operation request to the processing queue corresponding to the first directory, is specifically used to execute the following steps:
[0032] If there is a processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory;
[0033] If there is no processing queue corresponding to the first directory in the metadata processing structure, after creating the processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory.
[0034] In one implementation, a global lock is added to the metadata processing structure, and obtaining the global lock has the update permission for the metadata processing structure; the processing unit, when creating the processing queue corresponding to the first directory in the metadata processing structure, is specifically used to execute the following steps:
[0035] Obtain the global lock;
[0036] After obtaining the global lock, create the processing queue corresponding to the first directory in the metadata processing structure.
[0037] In one implementation, the processing unit is further used to execute the following steps:
[0038] If it is detected that there is an empty processing queue in the metadata processing structure, obtain the global lock; an empty processing queue refers to a processing queue in which the included file operation requests have all been released;
[0039] After obtaining the global lock, delete the empty processing queue in the metadata processing structure.
[0040] In one implementation, the file operation request is transmitted based on the transmission protocol;
[0041] The communication unit, when receiving the target file operation request, is specifically used to execute the following steps: create a processing thread corresponding to the transmission protocol; call the processing thread to receive the target file operation request;
[0042] The processing unit, when adding the target file operation request to the processing queue corresponding to the first directory, is specifically used to execute the following steps: reuse the processing thread corresponding to the transmission protocol and add the target file operation request to the processing queue corresponding to the first directory;
[0043] A processing unit, when used to update the metadata database according to a target file operation request, is specifically used to perform the following steps: Reuse the processing thread corresponding to the transmission protocol, and update the metadata database according to the target file operation request.
[0044] In one implementation, the processing unit is further used to perform the following steps:
[0045] After updating the metadata cache, reuse the processing thread corresponding to the transmission protocol, and notify the metadata update result corresponding to the target file operation request.
[0046] Correspondingly, an embodiment of the present application provides a computer device, which includes:
[0047] A processor, adapted to implement a computer program;
[0048] A computer-readable storage medium, storing a computer program, the computer program being adapted to be loaded and executed by the processor to perform the above-mentioned method for processing file system data.
[0049] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, storing a computer program, when the computer program is read and executed by the processor of the computer device, the computer device is caused to perform the above-mentioned method for processing file system data.
[0050] Correspondingly, an embodiment of the present application provides a computer program product, which includes a computer program, the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the above-mentioned method for processing file system data.
[0051] In the embodiments of the present application, when a target file operation request is received, the target file operation request can be added to the processing queue corresponding to the first directory. The file for which the operation is requested in the target file operation request is the target file, and the first directory is the parent directory to which the target file belongs. The processing queue corresponding to the first directory includes at least one file operation request. The parent directories of the files to be processed by the file operation requests in the processing queue corresponding to the first directory are all the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time. The metadata database can be updated according to the target file operation request; then, based on the metadata database update result of the target file operation request, the metadata cache can be updated in the arrangement order of the target file operation request in the processing queue corresponding to the first directory. It can be seen that in the embodiments of the present application, when a file operation request is received, the metadata database can be updated first, and then the metadata cache can be updated according to the metadata database update result, which can ensure the consistency of metadata update; moreover, the file operation requests in the processing queue corresponding to the first directory update the metadata cache in the order of the request reception time, which can ensure the correctness of metadata update. That is to say, by adopting the embodiments of the present application, the consistency and correctness of metadata update can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Figure 1 is a schematic diagram of a directory tree structure provided by an embodiment of the present application;
[0054] Figure 2 is a schematic diagram of the architecture of a file system provided by an embodiment of the present application;
[0055] Figure 3 is a schematic diagram of the process of metadata update provided by an embodiment of the present application;
[0056] Figure 4 is a schematic diagram of the process of a batch processing solution provided by an embodiment of the present application;
[0057] Figure 5 is a schematic diagram of the process of a customer request thread reuse solution provided by an embodiment of the present application;
[0058] Figure 6 is a schematic diagram of the process of a partitioning solution provided by an embodiment of the present application;
[0059] Figure 7 It is a schematic flowchart of a processing solution for file system data provided by an embodiment of the present application;
[0060] Figure 8 It is a schematic flowchart of a method for processing file system data provided by an embodiment of the present application;
[0061] Figure 9 It is a schematic flowchart of a process for processing queue control to update metadata cache provided by an embodiment of the present application;
[0062] Figure 10 It is a schematic flowchart of another method for processing file system data provided by an embodiment of the present application;
[0063] Figure 11 It is a schematic structural diagram of a device for processing file system data provided by an embodiment of the present application;
[0064] Figure 12 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0066] In order to more clearly understand the technical solutions provided by the embodiments of the present application, the technical terms involved in the embodiments of the present application are introduced here first:
[0067] (1) File system:
[0068] The file system of a computer device refers to a method of storing or organizing data in the computer device, which makes it easy to access and search for data. The file system uses the abstract logical concepts of files and tree-shaped directories to replace the concept of data blocks used by physical devices such as hard disks and optical discs. When using the file system to save data, there is no need to care about the address of the data block on which the data is actually stored on the hard disk (or optical disc), only need to remember the directory to which the file belongs and the file name. Before writing new data into the file system, there is no need to care about which address data block on the hard disk is not used. The storage space management function on the hard disk (including allocating storage space and releasing storage space) is automatically completed by the file system, and only need to remember which directory the data is written into.
[0069] The file system mentioned in the embodiments of the present application can support POSIX (Portable Operating System Interface) semantics. POSIX is a general term for a series of interrelated standards defined by the IEEE (the Institute of Electrical and Electronics Engineers) to define the API (Application Programming Interface) for software to run on various UNIX operating systems. The embodiments of the present application do not limit the type of the file system. The file system mentioned in the embodiments of the present application can be any file system that supports POSIX semantics. For example, the file system can be a distributed file system that supports POSIX semantics.
[0070] Among them, a distributed file system means that the physical storage resources managed by the file system are not necessarily directly connected to the local node, but are connected to the node (which can be simply understood as a computer device) through a computer network; or it is a complete hierarchical file system formed by combining several different logical disk partitions or volume labels. For example, the distributed file system can include HDFS (Hadoop (a distributed system infrastructure) Distributed File System), CHDFS (CloudHDFS, Cloud HDFS), Alluxio (a distributed file system), JuiceFS (another distributed file system), and CephFS (yet another distributed file system), and so on. CHDFS is a cloud-based distributed file system that provides a standard HDFS access protocol, excellent performance, and a hierarchical namespace; CHDFS mainly solves the storage and analysis of massive data in the big data scenario, and can seamlessly migrate the locally self-built HDFS to the highly available, highly scalable, low-cost, reliable, and secure CHDFS for big data users without changing the existing code.
[0071] (2) File system data:
[0072] The data involved in the file system can be called file system data, and the file system data can include file data and metadata. Among them, file data refers to the data stored or organized with the help of the file system, and the file data can include multiple files. The embodiments of the present application do not limit the type of the file. For example, the file type can include text files, audio files, video files, image files, directory files, device files, pipe files, and socket (socket) files, and so on.
[0073] Generally, metadata refers to "data that describes data", which is defined as data that provides information about certain data unidirectionally or multi - dimensionally; it is used to outline the basic information of the data to simplify the data search process and facilitate its use. In the embodiments of the present application, metadata refers to the basic information that outlines file data, which can be understood as the descriptive information of file data. These descriptive information may include directory information, attribute information (such as the last modification time, user group (owner), and permission data, etc.) of each file in the file data, as well as storage location information (such as disk block location information), etc.
[0074] Metadata can usually be stored in a metadata database. To improve the access efficiency of metadata, metadata can also be stored in a metadata cache. Thus, metadata can be accessed quickly by accessing the metadata cache. The purpose of providing "caching" is to make the access speed of metadata adapt to the processing speed of the CPU (Central Processing Unit). Its principle is based on the "locality behavior of program execution and data access" in memory, that is, within a certain program execution time and space, the accessed code is concentrated in a part. To give full play to the role of the cache, not only rely on "temporarily storing the just - accessed data", but also use instruction prediction and data pre - fetching technologies implemented by hardware to pre - fetch the data to be used from memory into the cache as much as possible. When there is a need to update metadata, it is necessary to ensure the consistency of metadata updates. The consistency of metadata updates means that the updates of metadata in the metadata database and the updates of metadata in the metadata cache need to be kept consistent. In this way, after the metadata is updated, the metadata stored in the metadata database and the metadata stored in the metadata cache can be kept consistent, ensuring the correctness of the metadata.
[0075] (3) Directory tree:
[0076] Metadata can be organized and managed in the form of a directory tree, which refers to a hierarchical or tree-like structure composed of directories at different levels. Specifically, in a computer device, a directory (or folder) is a virtual "container" that holds a digital file system, in which files and other directories are stored; a typical file system may contain thousands of directories, and by storing files in directories, the purpose of organizing the storage of files can be achieved; another directory within a directory is called its subdirectory (or subfolder), that is to say, there is a hierarchical relationship between directories. The outer directory (the outer directory refers to the directory that contains other directories) can be called the parent directory, and the inner directory (the inner directory refers to the directory that is contained by other directories) can be called the subdirectory. For example, if directory 1 contains directory 2, then directory 1 can be called the parent directory to which directory 2 belongs, and directory 2 can be called the subdirectory of directory 1. If directory 1 contains file 1, then directory 1 can be called the parent directory to which file 1 belongs. In this way, these directories at different levels constitute a directory tree with a hierarchical or tree-like structure.
[0077] The directory tree can be structured as a directory tree structure composed of multiple inodes (index nodes). Each inode can represent a directory or a file in the directory tree. That is to say, each inode can be used to identify a directory or a file in the directory tree, and each inode has corresponding attributes indicating whether the inode is used to identify a directory or a file. An inode is a data structure. In the directory tree structure, files are generally represented as leaf nodes in the directory tree, and intermediate path nodes are represented as directories. Figure 1 Shows an exemplary directory tree structure. In Figure 1 The shown directory tree structure may include 3 directories and 2 files. The 3 directories include the root directory " / " and two other directories " / Dir1" and " / Dir1 / Dir2", while the 2 files include the file "Dir1 / File2" and the file " / File1". Each inode can have node metadata, and the node metadata of each inode refers to the description information of the directory or file identified by the inode (for example, the description information can include directory information, attribute information, storage location information, etc.).
[0078] In addition to having node metadata, each inode can also have a parent inode id, an inode id, and an inode name. Among them, the inode id refers to the identifier of the inode, and the file system needs to have the ability to quickly find the inode through the inode id. The parent inode id refers to the inode id of the parent inode of the inode; here, the concepts of the parent inode and the child inode are introduced. Among two inodes with an inclusion relationship in the directory tree structure, the inode included by other inodes can be called the child inode, and the inode that includes other inodes can be called the parent inode. For example, in Figure 1 the directory tree structure shown, the inode " / " is the parent inode of the inode "Dir1", and the inode "Dir1" is the child inode of the inode " / ". The inode name refers to the name of the inode.
[0079] (4) Object Storage:
[0080] Object Storage is a storage architecture for data in computer devices that manages data as objects, different from other storage architectures (for example, the file system manages data as a file hierarchy, while block storage manages data as blocks within sectors and tracks). Each object usually includes the data itself, varying amounts of metadata, and a globally unique identifier. Object Storage can be implemented at multiple levels, including the device level (specifically referring to object storage devices), the system level, and the interface level. In each case, Object Storage attempts to achieve capabilities not available in other storage architectures, such as an interface that can be directly programmed by applications, a namespace that can span multiple physical hardware instances, and data management functions such as data replication and data distribution at the object-level granularity.
[0081] In particular, COS (Cloud Object Storage) is a distributed storage cloud service for storing massive files, with advantages such as high scalability, low cost, reliability, and security. Through diverse methods such as the console, API, SDK (Software Development Kit), and tools, the file system can easily and quickly access COS to upload, download, and manage multi-format files, achieving massive data storage and management.
[0082] Based on the introduction of the above technical terms such as file system, file system data, directory tree, and object storage, the file system can be combined with object storage to store the data in the file system by means of object storage. Specifically, the file data in the file system can be stored by means of object storage. The following introduces the system architecture of the file system and the processing flow of the file system:
[0083] I. System Architecture of the File System:
[0084] As Figure 2 shown, the file system may include a file client 201, an object storage 202, a root server (RS) 203, a proxy server 204, a meta data service device 205, and a meta database 206. Among them:
[0085] File Client 201: The number of file clients 201 can be one or more. The file client can run in a terminal device; the business object can request to operate the files in the file system through the file client 201 to generate a file operation request, thereby triggering metadata update. The file client 201 can provide a POSIX semantic interface. The file client 201 can encapsulate the file system semantic interface for use by the business object. Among them, common methods include HCFS (a method of encapsulating the file system semantic interface) and FUSE (another method of encapsulating the file system semantic interface), etc.
[0086] Object Storage 202: The object storage 202 can be any object storage. The file data stream part of the file client 201 directly accesses the object storage.
[0087] Root Server 203: The root server 203 can provide a management interface externally, which can include management interfaces such as creating and deleting file systems, creating and deleting mount points, and configuring file system attributes. Moreover, the root server 203 can also evenly distribute the FS route (a kind of route) to different meta data services, and can achieve high availability and master-slave disaster tolerance through Raft (a consensus algorithm) election.
[0088] Proxy Server 204: The proxy server 204 can provide a meta data interface externally, which can include meta data interfaces such as creating and deleting file handles and meta data update interfaces. The meta data stream part of the file client 201 directly accesses the proxy server 204, and the proxy server 204 can perform protocol conversion and forward it to the meta data service corresponding to the file system.
[0089] Metadata service device 205: It can provide metadata services, which can be used to manage the metadata of the file system, that is, it can be used to manage the directory tree and the node metadata of each inode in the directory tree. The metadata service device may include a metadata service master node and a metadata service slave node.
[0090] Metadata database 206: The metadata database 206 can be used to store metadata. The metadata database 206 can be a metadata storage cluster (DB Cluster), and can include any type of DB (Data Base), for example, the metadata storage cluster can include databases such as CDB (a kind of database), TDSQL (another kind of database), and CynosDB (yet another kind of database). The metadata database 206 can also be extended to a KV (Key-Value) storage structure.
[0091] It should be noted that the terminal device mentioned in the embodiments of the present application may include any one of the following: smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, smart watches, vehicle-mounted terminals, smart home appliances, and aircraft, etc., but is not limited thereto. The server mentioned in the embodiments of the present application can be a single physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present application do not limit this.
[0092] II. Processing flow of the file system:
[0093] In the creation stage of the file system: The root server 203 can provide a management interface externally, and can allocate the created file system and the configured attributes to the specified metadata service. At the same time, record the routing information of the file system for access.
[0094] During the update phase of the file system: The file client 201 can encapsulate the POSIX interface provided by the proxy server 204 to access the file data and metadata in the file system. When receiving a file operation request from the file client 201, the file data stream part of the file operation requested by the file operation request can directly access the object storage 202; moreover, when updating the file data in the object storage 202, it is not necessary to read the file-related object requested by the operation into the memory for modification and then write it back. Instead, the updated file is written as a new object into the object storage, which can save performance overhead and improve the operation efficiency of the file data stream. When receiving a file operation request from the file client 201, the metadata stream part of the file operation requested by the file operation request can be transmitted to the proxy server 204; the proxy server 204 can provide a metadata update interface externally, and the metadata update interface can route the file operation request to the metadata service recorded by the root server 203; the corresponding metadata service in the metadata service device 205 can update the metadata, specifically, it can update the metadata in the metadata database and the metadata in the metadata cache.
[0095] Based on the above system architecture and processing flow of the file system, it can be known that when a file operation request is received and metadata needs to be updated, to ensure the consistency of metadata update, the metadata update process can include: 1. Abstract the file operation request into a file operation log (XLOG), and the file operation log is an update structure used for updating the source data; 2. Update the metadata database according to the file operation log, and this process can also be called persisting the file operation log into the metadata database; 3. When the metadata database is successfully updated according to the file operation log (i.e., after successful persistence), the metadata cache can be updated according to the file operation log, and this process can also be called Apply memory. Apply memory refers to writing the persisted data into the cache. Different file systems or data have different description methods, and here it generally refers to the operation of updating the cache after data persistence to disk (data persistence to disk means successful update of the metadata database).
[0096] In the above metadata update process, the following two points need to be noted:
[0097] First, file operations can include four types: file creation operation (create), file deletion operation (delete), file modification operation (update), and file renaming operation (rename). Among them: The file creation operation refers to the operation of creating a file. When receiving a file operation request corresponding to the file creation operation, it is necessary to create a file in the file data and add a corresponding inode and the inode's node metadata in the metadata. The file deletion operation refers to the operation of deleting a file. When receiving a file operation request corresponding to the file deletion operation, it is necessary to delete the file in the file data and delete the corresponding inode and the inode's node metadata in the metadata. The file modification operation refers to the operation of modifying a file. When receiving a file operation request corresponding to the file modification operation, it is necessary to modify the file in the file data and modify the inode's node metadata corresponding to the file in the metadata. The file renaming operation refers to the operation of modifying the file name. The file renaming operation can include any one of the following: the operation of modifying the file name without modifying the file's parent directory, the operation of modifying the file's parent directory (for example, in the directory tree structure shown in Figure 1 delete the relevant content of the file "File2" in the directory "Dir1"), and the operation of modifying the file's parent directory and modifying the file name under the modified file's parent directory; when receiving a file operation request corresponding to the file renaming operation, it is necessary to rename the file in the file data and rename the inode's node metadata corresponding to the file in the metadata.
[0098] Second, among the four types of file operations, file creation operation, file deletion operation, and file modification operation belong to unary operations. A unary operation refers to an operation involving the update of corresponding metadata under a single parent directory; while the file renaming operation belongs to a binary operation. A binary operation refers to an operation involving the update of corresponding metadata under two parent directories. The file operation log abstracted from the file operation request can be used to record the operation data of the file operation request. The operation data may include the identifier (inode id) of the index node of the operation and the operation information. Among them, for the operation information, the operation information recorded in the file operation log corresponding to the file creation operation and the file deletion operation is unary operation information; for example, for the file creation operation, the unary operation information is the node metadata of the index node corresponding to the created file; another example is that for the file deletion operation, the unary operation information is the node metadata of the index node corresponding to the deleted file). The operation information recorded in the file operation log corresponding to the file modification operation and the file renaming operation is binary operation information; for example, for the file modification operation, the binary operation information includes the node metadata of the index node corresponding to the file before modification and the node metadata of the index node corresponding to the file after modification; another example is that for the file renaming operation (for example, specifically an operation to modify the file's parent directory), the binary operation information includes the node metadata of the index node corresponding to the source parent directory before renaming and the node metadata of the index node corresponding to the parent directory after renaming).
[0099] Taking the file renaming operation (specifically, the operation of moving the file "FileA" from the parent directory "DirA" to the parent directory "DirB" and changing the file "FileA" to the file "FileB") as an example, the update process of the metadata is as follows Figure 3 shown. The operation information recorded in the file operation log may include the node metadata of the index node corresponding to the file "FileA" and the node metadata of the index node corresponding to the file "FileB". It can be seen that in order to ensure the correctness of the metadata update for a file operation request, it is necessary to update the metadata cache after the metadata database is updated successfully. After the metadata cache is updated successfully, the result is then returned to the client.
[0100] The update efficiency and update correctness of metadata are the concerns in the high-concurrency scenario of file operation requests. The high-concurrency scenario of file operation requests refers to the scenario where a large number of file operation requests are received within a short period of time. Based on this, the embodiments of the present application propose a method for processing file system data. This method for processing file system data can adopt an asynchronous processing method that separates the update of the metadata database and the update of the metadata cache, making full use of the performance of the underlying storage medium to improve the metadata update efficiency. Moreover, this method for processing file system data can, based on the characteristics of the file system, allow file operation requests under different parent directories to perform metadata updates concurrently without affecting each other, thereby improving the metadata update efficiency. In addition, for file operation requests under the same parent directory, this method for processing file system data uses the method of parent directory locks to queue the file operation requests. The file operation requests received earlier under the same parent directory are preferentially responded to for updating the metadata, and the file operation requests received later are responded to for updating the metadata later, which can ensure the correctness of the metadata update. For file operation requests under the same parent directory, after successfully updating the metadata database, it asynchronously notifies the blocked file operation requests to update the metadata cache, which can ensure the strong consistency of the metadata database and the metadata cache while improving the metadata update efficiency and the metadata update correctness.
[0101] The method for processing file system data provided by the embodiments of the present application can be applicable to the storage and organization of big data. Moreover, the storage involved in the embodiments of the present application (such as object storage, metadata cache, and metadata database) can be cloud storage. Among them:
[0102] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth-rate, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has attracted more and more attention. Big data requires special technologies to effectively process a large amount of data that tolerates the elapsed time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.
[0103] Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technologies, and distributed storage file systems and other functions, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside.
[0104] The exploration process of the method for processing file system data provided by the embodiments of the present application is introduced below. Based on this exploration process, the advantages of the method for processing file system data provided by the embodiments of the present application in terms of metadata update can be more clearly understood.
[0105] (1) Batch processing solution:
[0106] Batch processing can also be referred to as batch processing tasks, which refers to the operation of executing a series of programs on a computer device without manual intervention. The idea of the batch processing solution is to merge and flush the file operation requests to disk. The so-called merge and flush to disk means that after packing the file operation requests of a certain amount of data into a request batch (batch table), the metadata database is updated in units of the request batch, and the metadata cache is updated after the metadata database is successfully updated.
[0107] The specific operation of the batch processing solution is as Figure 4 shown: After the file operation requests from the file client are abstracted into file operation logs (XLOGs), they are queued in the log queue (XLOG queue) through the parent directory to which the file operated by the request belongs (i.e., the parent inode id); each log queue corresponds to a thread that asynchronously consumes to form a request batch (batch table). The request batch (batch table) contains a fixed number of file operation logs (XLOGs) of a certain size. The log identifier (XLOG id) of the file operation logs in each request batch (batch table) needs to ensure global unique increment; the request batch (batch table) can be delivered to the persistence queue (commit queue) according to a periodic time interval and a fixed threshold; there is a thread that concurrently flushes and persists the file operation logs (XLOGs) in the request batch (batch table) in the persistence queue (commit queue) to the metadata database (DB). Concurrent flushing and persistence means that the metadata database is updated concurrently for each file operation log in the request batch (batch table). Then, the request batch (batch table) with successful persistence can be delivered to the cache update queue (apply queue), and the metadata cache can be updated asynchronously in order according to the file operation logs, and then the file client is notified of the metadata update result (the metadata update result can include a metadata update result indicating successful metadata update and a metadata update result indicating failed metadata update).
[0108] We combine the batch processing solution with the file system characteristics. After hashing through the parent directory (i.e., parent inode id), file operation requests are divided into the log queue (XLOG queue). File operation logs under different parent inode ids can update metadata (updating metadata includes updating the metadata database and updating the metadata cache) in parallel without affecting each other. The concept of the abstract request batch (batch table) is used to combine file operation logs (XLOG), and then they are processed in units of request batches (batch tables). Finally, we have two steps of updating the metadata database and updating the metadata cache, achieving asynchronous updating of the metadata database and asynchronous updating of the metadata cache through different queues, greatly improving the processing efficiency of file operation requests.
[0109] The batch processing solution mentioned above improves the processing efficiency of file operation requests by means of batch processing and asynchronous updating of the metadata database and metadata cache, but there are still problems in the architecture. For example: 1) Resource consumption problem. Currently, there are three-level queues, and the first-level queue is a multi-level queue hashed according to the parent directory (i.e., parent inode id), and each queue corresponds to a thread for processing. As the number of file systems in the service increases, the resource occupancy of this part will be more; 2) At the same time, the granularity in the request batch (batch table) is the file operation log (XLOG), and the architecture is not clear. Finally, and most importantly, if a file operation log (XLOG) in the request batch (batch table) has a long tail when updating the metadata database (a long tail means that the update result of the metadata database cannot be obtained for a long time), processing in units of the request batch (batch table) will cause the increase of the processing time of other requests; because it is necessary to cache the metadata in order for the file operation logs (XLOG) that have successfully updated the metadata database, if the file operation logs (XLOG) in front of the log queue (XLOG queue) are processed slowly, they will block the processing of the subsequent file operation logs (XLOG). From the above architecture, it can be seen that some structures are not necessary. For example: 1) Allocating a large number of threads to do the same thing as database operations, but the threads for processing client requests are blocked. The client request threads (e.g., threads of RPC (Remote Procedure Call)) can be used to interact with the metadata database (DB), thus proposing a client request thread reuse solution. Among them, the client request thread refers to the thread created for file operation requests from the file client, and the thread is the smallest unit that the operating system can perform operation scheduling.
[0110] (2) Client request thread reuse solution:
[0111] The specific operations of the client request thread reuse solution are asFigure 5 As shown: Each file operation request can be abstracted into a file operation log (XLOG), and a client request thread (e.g., an RPC thread) is assigned to each file operation request. Moreover, the concept of batch processing can be discarded. Each file operation log (XLOG) can be assigned an identifier of the file operation log (XLOGID), and the file operation logs (XLOGs) are queued based on the identifier of the file operation log (XLOG ID). In addition, all processing involved in metadata updates is handled in the client request thread. After the file operation logs (XLOGs) in the log queue (XLOG queue) successfully update the metadata database and successfully update the metadata cache, the client request thread is asynchronously notified to return the metadata update result. In this way, we only retain one queue and only one client request thread to handle all processes, and the processing capacity of the client request thread can be reused as much as possible in a high-concurrency scenario. Additionally, the underlying transaction capabilities can be reused to atomically protect the two operations of writing the file operation log (XLOG) to the log queue (XLOG queue) and updating the metadata database. If the write to the log queue (XLOG queue) is successful, the metadata database is also successfully updated persistently. Therefore, file operation requests can be processed in two steps. The first step is to write to the log queue (XLOG queue) and update the metadata database persistently, and the second step is to update the metadata cache. However, limited by the correctness of the file system semantics, we need to ensure that file operation requests that update the metadata database first are followed by metadata cache updates, which will cause file operation requests that have successfully updated the metadata to be blocked and the metadata cache updates to be performed in sequence. Therefore, if simply queuing according to the identifier of the file operation log (XLOG ID), there will still be a long-tail phenomenon. Therefore, the next direction of our optimization is how to solve the problem of minimizing the increase in latency caused by persistent long tails under the file system semantics, and thus a partitioning scheme is proposed.
[0112] (3) Partitioning scheme:
[0113] The partitioning scheme combines file system characteristics. The file system characteristics refer to that file operation requests under different parent directories can perform source data updates concurrently. The specific operations of the partitioning scheme are as Figure 6As shown below: File operation requests under different parent directories can be concurrent. Therefore, we partition according to the parent directory (i.e., parent inode id), and queue the file operation requests under the same parent directory (i.e., parent inode id). The advantage of this approach is that after the requests under different parent directories (i.e., parent inode id) successfully update the metadata database, they do not need to wait for the file operation requests in other queues to finish updating the metadata cache before returning; they only need to wait for the file request operations that have previously successfully updated the metadata database within the partition queue to complete the update of the metadata database. However, the partitioning scheme may not be applicable to all scenarios. For example, in big data scenarios, most read and write operations are performed under a part of the parent directories (i.e., parent inode id). At the same time, if partitioning is done according to the parent directory (i.e., parent inode id), it depends on the number of partition queues. The fewer the number of partition queues, the more concentrated the data, and the more resources are occupied if there are more. Therefore, to address the situation of intensive operations in big data or other scenarios with the same parent directory (i.e., parent inode id), the data processing scheme for the file system proposed in this application embodiment is presented.
[0114] (4) Data processing scheme for the file system:
[0115] Combined with the characteristics of relatively intensive file updates, whether it is write-then-read or stream-write-then-read, file operations have the characteristic of concentration. The file system more conforms to the scenario where the read frequency is greater than the write frequency. When the system is running stably, it mainly provides read services externally. Here, in order to avoid the long-tail problem brought by the partitioning scheme, we directly refine the dimension to the specific parent directory (i.e., parent inode id), and maintain a dynamic processing queue for the parent directory (i.e., parent inode id) where the leaf nodes (i.e., files) in the directory tree have update operations. At the same time, each parent directory (i.e., parent inode id) is maintained through a parent directory lock to ensure that the file operation logs that have successfully updated the metadata database under the same parent directory (i.e., parent inode id) update the metadata cache in sequence.
[0116] The specific operations of the data processing scheme for the file system are as Figure 7As shown: After abstracting a file operation request (for example, it can be an RPC request) into a file operation (Operator) and corresponding file operation log (XLOG), it can be dynamically inserted into the parent directory tree lock according to the parent directory of its operation (i.e., parent inode id). Each node in the parent directory tree lock maintains a processing queue (parent inode lock chan), and each processing queue has a parent directory lock. After the file operation log (XLOG) with the parent directory lock successfully updates the metadata database, it can perform metadata cache update. After successful insertion, we update the metadata database and the metadata cache. Updating the metadata database can reuse the client request thread to concurrently write to the underlying storage. After the metadata database is successfully updated, the metadata cache is updated. Then, it can notify the processing queue to release the parent directory lock and notify the next file operation log (XLOG) in the processing queue. If the next file operation log (XLOG) has successfully updated the database, it can obtain the parent directory lock to update the metadata cache. After the metadata cache is successfully updated, the metadata update result can be returned. Because of the intensive nature of file system operations, the file operation logs (XLOG) in the processing queue (parent inode lock chan) of the node may all have been processed, and it is not necessary to keep this node for a long time. Therefore, we can add a memory recycling service (server GC) to periodically obtain and clean these nodes.
[0117] The following will introduce in detail the method for processing file system data provided by the embodiments of the present application with reference to the accompanying drawings.
[0118] The embodiments of the present application provide a method for processing file system data. This method for processing file system data mainly introduces the control effect of the parent directory lock in the processing queue on the sequential execution of file operation requests in the processing queue. This method for processing file system data can be executed by a computer device, and the computer device can be, for example, Figure 2 the metadata service device 205 in the file system shown. Please refer to Figure 8 , and this method for processing file system data can include but is not limited to the following steps S801 - step S804:
[0119] S801, receive a target file operation request. The file requested to be operated on by the target file operation request is the target file, and the parent directory to which the target file belongs is the first directory.
[0120] As described above, the file system data may include metadata and file data. The metadata may be stored in a metadata cache and a metadata database. The file data may include multiple files. The metadata may include description information of each file in the file data. The description information may include, but is not limited to, directory information, attribute information (e.g., last modification time, user group, and permission data, etc.), and storage location information (e.g., disk block location information). After receiving a file operation request from a file client, it is necessary to confirm the parent directory to which the file requested by the file operation request belongs. In the embodiments of the present application, it is assumed that the received file operation request is a target file operation request, the file requested by the target file operation request is a target file, and the parent directory to which the target file belongs is a first directory, and the following will be introduced in detail by way of example.
[0121] After receiving the target file operation request, the target file operation request may be structured into a corresponding file operation log. The so-called structuring means converting the file operation request into operation data that meets the format requirements according to the format requirements of the file operation log. The file operation log can be used to record the file operation of the file operation request. Specifically, it can be used to record the operation data of the file operation. The operation data may include the identifier (parent inode id) of the parent index node (i.e., the parent directory) of the requested operation and the operation information. Structuring the file operation request into a corresponding file operation log is because the file operation log is a way of updating metadata. In addition, the file operation log may also have the following uses: when there is a power failure or an abnormal interruption (coredump) during the metadata update process, the metadata after executing the file operation log can be compared with the actual metadata to ensure the consistency and correctness of the metadata update when there is a power failure or an abnormal interruption (coredump) during the metadata update process; and, the file operation log can be persisted to the metadata database and the current file operation log can be recorded. The standby machine can obtain the differential file operation log for playback through the current file operation log and the existing information, and can perform disaster recovery and master-slave strong or weak consistency synchronization through snapshots (Snapshot) and file operation logs.
[0122] Further, the process of structuring the target file operation request into the corresponding file operation log may include: determining the request type of the target file operation request (for example, the request type may be any one of the request types corresponding to file creation operation (create), file deletion operation (delete), file modification operation (update), and file rename operation (rename)). When the request type of the target file operation request is the target operation request type, the target file operation request may be structured into the target file operation log corresponding to the target operation request type, and the target file operation log may be used to record the operation data of the target file operation request, that is to say, the target file operation log may include the operation data of the target file operation request.
[0123] S802, adding the target file operation request to the processing queue corresponding to the first directory.
[0124] After determining that the parent directory to which the target file requested by the target file operation request belongs is the first directory, the target file operation request may be added to the processing queue corresponding to the first directory; here, specifically, it means that after structuring the target file operation request into the target file operation log, the target file operation log is added to the processing queue corresponding to the first directory. The processing queue corresponding to the first directory may include at least one file operation request. The parent directories to which the files requested by the file operation requests in the processing queue corresponding to the first directory belong are all the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time; as described above, the file operation request may be inserted into the processing queue corresponding to the parent directory of the requested operation in the form of a file operation log. Therefore, the processing queue corresponding to the first directory may include at least one file operation log. The parent directories to which the files requested by the file operation logs in the processing queue corresponding to the first directory belong are all the first directory, and the file operation logs in the processing queue corresponding to the first directory are arranged in the order of the corresponding request reception time.
[0125] It should be noted that the processing queue mentioned in the embodiments of the present application may also refer to a processing linked list, specifically a queue implemented in the form of a linked list. Implementing a queue in the form of a linked list has the characteristics of dynamic storage space, the ability to store any type of data, fast insertion and deletion speeds, etc. The parent directories to which different files in the file data belong can form a directory tree according to the hierarchical relationship. There is a dedicated processing structure for storing the processing queues corresponding to different parent directories in the directory tree. This dedicated processing structure can be called a metadata processing structure. A metadata processing structure is a data structure that can be used to store the processing queues corresponding to different levels of parent directories in the directory tree. In the metadata processing structure, the head of the processing queue corresponding to each parent directory has a parent directory lock (xchan lock). A parent directory lock is a control mechanism or control structure that can be used to control the file operation requests (specifically, the file operation logs after structuring the file operation requests) in the corresponding processing queue to update the metadata cache in sequence, ensuring the consistency of metadata updates. Thus, the metadata processing structure can be understood as a tree-shaped lock structure corresponding to the directory tree (corresponding to the parent directory tree lock mentioned above).
[0126] Based on the introduction of the metadata processing structure, before adding a target file operation request to the processing queue corresponding to the first directory, it can be first checked whether there is a processing queue corresponding to the first directory in the metadata processing structure; if there is a processing queue corresponding to the first directory in the metadata processing structure, the target file operation request can be added to the processing queue corresponding to the first directory; if there is no processing queue corresponding to the first directory in the metadata processing structure, after creating a processing queue corresponding to the first directory in the metadata processing structure, the target file operation request can be added to the processing queue corresponding to the first directory.
[0127] Moreover, after receiving a target file operation request, the target file operation request can be structured into a corresponding file operation log, and the target file operation request is added to the processing queue corresponding to the first directory in the form of a file operation log; that is, if there is a processing queue corresponding to the first directory in the metadata processing structure, the target file operation log can be added to the processing queue corresponding to the first directory; if there is no processing queue corresponding to the first directory in the metadata processing structure, after creating a processing queue corresponding to the first directory in the metadata processing structure, the target file operation log can be added to the processing queue corresponding to the first directory.
[0128] S803, update the metadata database according to the target file operation request.
[0129] After receiving a target file operation request, the target file operation request can be structured into a corresponding file operation log, and the metadata database is updated through the file operation log. That is to say, according to the target file operation request, the metadata database is updated, specifically referring to: updating the metadata database according to the operation data of the target file operation request in the target file operation log.
[0130] For example, when the target file operation request is a file creation operation request, an inode corresponding to the target file can be created under the first directory in the metadata database, and the node metadata of the inode is added. When the target file operation request is a file deletion operation request, the inode corresponding to the target file can be deleted under the first directory in the metadata database, and the node metadata of the inode is deleted. When the target file operation request is a file modification operation request, the node metadata of the inode corresponding to the target file can be modified under the first directory in the metadata database. When the target file operation request is a file rename operation request (for example, the file rename operation is specifically an operation to move the target file from the first directory to the target directory), the inode corresponding to the target file and the node metadata of the inode can be deleted under the first directory in the metadata database, and an inode corresponding to the target file and the node metadata of the inode are created under the target directory in the metadata database.
[0131] S804, based on the metadata database update result of the target file operation request, update the metadata cache according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory.
[0132] Similar to the update of the metadata database, the metadata cache can also be updated according to the file operation log. That is to say, the process of updating the metadata cache based on the metadata database update result of the target file operation request and according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory specifically refers to: based on the metadata database update result of the target file operation request, according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory, update the metadata cache according to the operation data of the target file operation request in the target file operation log. It should be noted that the method of updating the metadata cache according to the operation data of the target file operation request in the target file operation log is the same as the method of updating the metadata database according to the operation data of the target file operation request in the target file operation log. For specific details, please refer to the update method of the metadata database and will not be elaborated here.
[0133] Further, each file operation request in the processing queue corresponding to the first directory can, in the order of arrangement, update the metadata cache based on the corresponding file operation request in sequence (in the following content of step S804, if it is mentioned to update the metadata cache according to the file operation request or based on the file operation request, it specifically refers to updating the metadata cache according to the file operation log structured from the file operation request, which is hereby explained). Moreover, the head of the processing queue corresponding to the first directory has a parent directory lock. Among the file operation requests in the processing queue corresponding to the first directory, the file operation request that obtains the parent directory lock can have the permission to update the metadata cache.
[0134] Furthermore, based on the metadata database update result of the target file operation request, updating the metadata cache in the order of arrangement of the target file operation request in the processing queue corresponding to the first directory can include: when the target file operation request is at the head of the queue, adding the parent directory lock to the target file operation request, and the target file operation request can have the permission to update the metadata cache; if the metadata database update result of the target file operation request indicates that the metadata database is successfully updated according to the target file operation request, then the metadata cache can be updated according to the target file operation request.
[0135] If the metadata database update result of the target file operation request indicates that updating the metadata database according to the target file operation request fails, then the target file operation request can be released from the head of the queue; the file operation request that is arranged one position after the target file operation request in the processing queue corresponding to the first directory is moved to the head of the queue. Among them, any of the following situations may cause the failure to update the metadata database according to the target file operation request: the time when updating the metadata database according to the target file operation request ends exceeds the time threshold but the metadata database is still not updated successfully, an error occurs in the metadata database update, or the target file operation log does not exist.
[0136] The following combines Figure 9Exemplarily introduce the overall update process of the metadata database and the metadata cache: File operation requests A, B, and C under the first directory are received. The reception time of file operation request A is earlier than that of file operation request B, and the reception time of file operation request B is earlier than that of file operation request C. File operation request A is structured into file operation log A (XLOGA) and then added to the processing queue corresponding to the first directory, and the metadata database is updated according to file operation log A (XLOGA) through the interface request of the metadata database; File operation request B is structured into file operation log B (XLOGB) and then added to the processing queue corresponding to the first directory, and the metadata database is updated according to file operation log B (XLOGB) through the interface request of the metadata database; File operation request C is structured into file operation log C (XLOGC) and then added to the processing queue corresponding to the first directory, and the metadata database is updated according to file operation log C (XLOGB) through the interface request of the metadata database. Moreover, the file operation log (XLOG) in the processing queue corresponding to the first directory can be called a file operation node (Xlock). That is to say, the processing queue corresponding to the first directory can include the file operation node A (XlockA) corresponding to file operation log A (XLOGA), the file operation node B (XlockB) corresponding to file operation log B (XLOGB), and the file operation node C (XlockC) corresponding to file operation log C (XLOGC) arranged in sequence.
[0137] The head of the processing queue corresponding to the first directory has a parent directory lock (xchan lock), and the file operation node A (XlockA) is located at the head of the processing queue corresponding to the first directory. If the file operation node A (XlockA) successfully updates the metadata database, it can obtain the parent directory lock and update the metadata cache; after the file operation node A (XlockA) successfully updates the metadata cache, it can return the metadata update result regarding file operation request A to the file client. After the file operation node A (XlockA) successfully updates the metadata cache, it can notify the processing queue corresponding to the first directory, triggering the deletion of the file operation node A (XlockA) from the processing queue corresponding to the first directory. After deletion, the file operation node B (XlockB) is located at the head of the processing queue corresponding to the first directory; if the file operation node B (XlockB) successfully updates the metadata database, it can obtain the parent directory lock and update the metadata cache; and so on, until the file operation node C (XlockC) is processed.
[0138] Based on the content of step S804 above, it can be seen that in the embodiment of this application, the file operation requests under the parent directory are controlled by the parent directory lock, and the metadata cache is updated sequentially according to the sorting order. In this control method, the following three points need to be noted (taking the processing queue corresponding to the first directory as an example): First, the parent directory lock is a queue-level lock (xlist lock). The file operation node (Xlock) that obtains the parent directory lock performs metadata cache update. After the metadata cache is successfully updated, it triggers the callback of the processing queue corresponding to the first directory to release the parent directory lock of the next file operation node, and removes the current file operation node. To ensure consistency, a queue-level lock is required. All notifications related to the parent directory lock and operations of the file operation node (Xlock) must be guaranteed by the queue-level lock (xlist lock). Second, the deletion of the file operation node (Xlock) is placed after the metadata cache update is successful, and the request is immediately responded to after the metadata cache is successfully updated. If the callback method is used (the callback method means notifying the processing queue corresponding to the first directory to delete the file operation node (Xlock) that has successfully updated the metadata cache), it cannot be guaranteed whether the file operation node (Xlock) is cleared by the memory when removing it, and there may be a null pointer scenario. Therefore, the file operation node (Xlock) and the callback of the processing queue are directly returned. By deleting the specific file operation node (Xlock) through the callback of the processing queue, the security protection of concurrent operations can be achieved through the directory lock at the xlist level. Third, the lock at the file operation node (Xlock) level cannot be used to ensure that the parent directory lock is notified in order.
[0139] The above steps S801 - S804 introduce the metadata update process after receiving any file operation request. It should be noted that in a high-concurrency scenario, multiple file operation requests can exist simultaneously, and multiple processing queues corresponding to different parent directories can exist simultaneously in the metadata processing structure. In this case, the processing queues corresponding to different parent directories can perform metadata updates concurrently. Specifically, the processing queue corresponding to the first directory is located in the metadata update structure, and the metadata update structure can at least further include a processing queue corresponding to the second directory. The first directory and the second directory are different directories; the processing queue corresponding to the first directory and the processing queue corresponding to the second directory can concurrently execute the operation of updating metadata; the operation of updating metadata can include updating the metadata database and updating the metadata cache; and for any file operation request in a processing structure, the metadata cache is updated after the metadata is successfully updated.
[0140] In the embodiments of the present application, an asynchronous processing method of separating the update of the metadata database and the update of the metadata cache can be adopted to make full use of the performance of the underlying storage medium and improve the metadata update efficiency. Moreover, the method for processing file system data can be based on the characteristics of the file system. File operation requests under different parent directories can perform metadata updates concurrently and do not affect each other, thereby improving the metadata update efficiency. In addition, for file operation requests under the same parent directory, the method for processing file system data uses the method of parent directory lock to queue the file operation requests. The file operation requests received first under the same parent directory are preferentially responded to update the metadata, and the file operation requests received later are responded to update the metadata later, which can ensure the correctness of the metadata update. For file operation requests under the same parent directory, after successfully updating the metadata database, an asynchronous notification is sent to the blocked file operation requests to update the metadata cache, which can ensure the strong consistency of the metadata database and the metadata cache while improving the metadata update efficiency and the metadata update correctness.
[0141] The embodiments of the present application provide a method for processing file system data. This method for processing file system data mainly introduces the control function of the global lock and the reuse of client request threads. This method for processing file system data can be executed by a computer device, and the computer device can be, for example, Figure 2 the metadata service device 205 in the file system shown. Please refer to Figure 10 , this method for processing file system data may include but is not limited to the following steps S1001 - step S1006:
[0142] S1001, receive a target file operation request. The file requested to be operated by the target file operation request is the target file, and the parent directory to which the target file belongs is the first directory.
[0143] In the embodiments of the present application, the execution process of step S1001 is the same as the execution process of step S801 in the above Figure 8 shown embodiment. Specifically, the execution process of step S801 in the above Figure 8 shown embodiment can be referred to and will not be elaborated here.
[0144] S1002, query the processing queue corresponding to the first directory in the metadata processing structure.
[0145] S1003, if there is a processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory.
[0146] S1004, if there is no processing queue corresponding to the first directory in the metadata processing structure, create a processing queue corresponding to the first directory in the data processing structure and then add the target file operation request to the processing queue corresponding to the first directory.
[0147] In steps S1002 - S1004, as described above, file operation requests under the same parent directory are queued. For frequent operations on different parent directories, it will cause frequent swapping in and out of nodes in the metadata processing structure. Here, the node refers to the processing queue corresponding to the parent directory. Here, swapping in means creating the processing queue corresponding to the parent directory in the metadata processing structure, and swapping out means deleting the corresponding processing queue in the metadata processing structure. In order to protect the swapping in and out operations of nodes in the metadata processing structure, a global lock needs to be added. The global lock is similar to the parent directory lock and is a control mechanism or control structure. The metadata processing structure can be added with a global lock. Obtaining the global lock means having the update permission for the metadata processing structure.
[0148] For the scenario of node swapping in, taking the scenario where the processing queue corresponding to the first directory needs to be created and does not exist in the metadata processing structure as an example, if the processing queue corresponding to the first directory does not exist in the metadata processing structure, then the processing queue corresponding to the first directory needs to be created in the metadata processing structure and the global lock can be obtained; after obtaining the global lock and having the permission to create the processing queue, the processing queue corresponding to the first directory can be created in the metadata processing structure.
[0149] For the scenario of node swapping out, for example, when the memory recycling service (server GC) detects that there is an empty processing queue in the metadata processing structure (an empty processing queue refers to a processing queue in which all file operation requests included have been released), the global lock can be obtained. After obtaining the global lock and having the permission to delete the processing queue, the empty processing queue can be deleted in the metadata processing structure.
[0150] It should be noted that after adding the global lock to the metadata processing structure, it may cause performance loss. Moreover, considering the characteristics of big data and file system usage, most operations are on files under the same parent directory. Therefore, it is necessary to make a trade - off according to the actual file processing requirements. Specifically, the number of swapped - in and swapped - out nodes can be counted periodically. The global lock of the metadata processing structure can be released at the start of the period and the number of swapped - in and swapped - out nodes can be counted. If the number of swapped - in and swapped - out nodes counted within the period exceeds the node number threshold, the global lock can be added to the metadata processing structure. If the number of swapped - in and swapped - out nodes counted within the period does not exceed the node number threshold, the metadata processing structure does not need to add the global lock all the time. In this way, the global lock can be added or released according to the file operation requirements, which can better meet the file operation requirements.
[0151] S1005, update the metadata database according to the target file operation request.
[0152] In the embodiments of the present application, the execution process of step S1005 is the same as that of step S803 in the above Figure 8 illustrated embodiment. Specifically, reference may be made to the execution process of step S803 in the above Figure 8 illustrated embodiment, which will not be elaborated herein.
[0153] S1006. Based on the metadata database update result of the target file operation request, update the metadata cache according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory.
[0154] In the embodiments of the present application, the execution process of step S1006 is the same as that of step S804 in the above Figure 8 illustrated embodiment. Specifically, reference may be made to the execution process of step S804 in the above Figure 8 illustrated embodiment, which will not be elaborated herein.
[0155] In steps S1001 - S1006, the customer request thread can be reused to execute each step in steps S1001 - S1006. Specifically, the file operation request can be transmitted based on the transport protocol. The embodiments of the present application do not limit the type of the transport protocol. For example, the transport protocol can be the RPC protocol. A processing thread corresponding to the transport protocol can be created, and the processing thread can be called to receive the target file operation request. The processing thread corresponding to the transport protocol can be reused to add the target file operation request to the processing queue corresponding to the first directory, and the processing thread corresponding to the transport protocol can be reused to update the metadata database according to the target file operation request. In addition, after updating the metadata database based on the target file operation request, the processing thread corresponding to the transport protocol can be reused to notify the metadata update result corresponding to the target file operation request. It can be seen that from the receipt of the target file operation request to the notification of the metadata update result corresponding to the target file operation request, only one thread is involved, and this thread is reused throughout the processing process; in the high-concurrency scenario of requests, through the reuse of threads, the metadata update efficiency in the high-concurrency scenario of requests can be improved.
[0156] In the embodiments of the present application, an asynchronous processing method of separating the update of the metadata database and the update of the metadata cache can be adopted to make full use of the performance of the underlying storage medium and improve the metadata update efficiency. Moreover, the method for processing the file system data can be based on the characteristics of the file system. File operation requests under different parent directories can perform metadata updates concurrently without affecting each other, thereby improving the metadata update efficiency. In addition, for file operation requests under the same parent directory, the method for processing the file system data queues the file operation requests by using a parent directory lock. The file operation request received first under the same parent directory is preferentially responded to update the metadata, and the file operation request received later is responded to update the metadata later, which can ensure the correctness of the metadata update. For file operation requests under the same parent directory, after successfully updating the metadata database, an asynchronous notification is sent to the blocked file operation requests to update the metadata cache, which can ensure the strong consistency of the metadata database and the metadata cache while improving the metadata update efficiency and the correctness of the metadata update.
[0157] The method of the embodiments of the present application is described in detail above. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.
[0158] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a device for processing file system data provided by the embodiments of the present application. The device for processing file system data can be set in the computer device provided by the embodiments of the present application. The computer device can be the Figure 2 metadata service device 205 in the file system shown above. Figure 11 The device for processing file system data shown above can be a computer program running in a computer device. The data processing device can be used to execute Figure 8 or Figure 10 some or all of the steps in the method embodiments shown above. The file system data includes metadata and file data. The metadata is stored in the metadata cache and the metadata database; the file data includes multiple files, and the metadata includes the description information of each file in the file data. Please refer to Figure 11 , the device for processing file system data can include the following units:
[0159] A communication unit 1101, configured to receive a target file operation request. The file requested to be operated by the target file operation request is a target file, and the parent directory to which the target file belongs is a first directory;
[0160] A processing unit 1102 is configured to add a target file operation request to a processing queue corresponding to a first directory; the processing queue corresponding to the first directory includes at least one file operation request, the files requested to be processed by the file operation requests in the processing queue corresponding to the first directory all belong to a parent directory that is the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time;
[0161] The processing unit 1102 is further configured to update a metadata database according to the target file operation request;
[0162] The processing unit 1102 is further configured to update a metadata cache based on the metadata database update result of the target file operation request, in the arrangement order of the target file operation request in the processing queue corresponding to the first directory.
[0163] In one implementation, the processing queue corresponding to the first directory is located in a metadata update structure, and the metadata update structure at least further includes a processing queue corresponding to a second directory; the first directory and the second directory are different directories;
[0164] The processing queue corresponding to the first directory and the processing queue corresponding to the second directory concurrently execute operations to update metadata; the operations to update metadata include updating the metadata database and updating the metadata cache.
[0165] In one implementation, each file operation request in the processing queue corresponding to the first directory updates the metadata cache based on the corresponding file operation request in the arranged order; the head of the processing queue corresponding to the first directory has a parent directory lock, and the file operation request that obtains the parent directory lock has the permission to update the metadata cache;
[0166] When the processing unit 1102 updates the metadata cache based on the metadata database update result of the target file operation request, in the arrangement order of the target file operation request in the processing queue corresponding to the first directory, it is specifically configured to perform the following steps:
[0167] When the target file operation request is at the head of the queue, add the parent directory lock to the target file operation request;
[0168] If the metadata database update result of the target file operation request indicates that the metadata database is successfully updated according to the target file operation request, update the metadata cache according to the target file operation request.
[0169] In one implementation, the processing unit 1102 is further configured to perform the following steps:
[0170] If the metadata database update result of the target file operation request indicates that the update of the metadata database according to the target file operation request fails, release the target file operation request from the head of the queue;
[0171] Move the file operation request that is one position after the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
[0172] In one implementation, the processing unit 1102 is further configured to perform the following steps:
[0173] After successfully updating the metadata cache according to the target file operation request, release the target file operation request from the head of the queue;
[0174] Move the file operation request that is one position after the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
[0175] In one implementation, the operation request type of the target file operation request is the target operation request type. After the target file operation request is structured into the target file operation log corresponding to the target operation request type, it is added to the processing queue corresponding to the first directory; the target file operation log includes the operation data of the target file operation request.
[0176] When the processing unit 1102 is used to update the metadata database according to the target file operation request, it is specifically configured to perform the following steps:
[0177] Update the metadata database according to the operation data of the target file operation request in the target file operation log;
[0178] When the processing unit 1102 is used to update the metadata cache based on the metadata database update result of the target file operation request in the order of arrangement of the target file operation request in the processing queue corresponding to the first directory, it is specifically configured to perform the following steps:
[0179] Based on the metadata database update result of the target file operation request, update the metadata cache according to the operation data of the target file operation request in the target file operation log in the order of arrangement of the target file operation request in the processing queue corresponding to the first directory.
[0180] In one implementation, the parent directories to which different files in the file data belong form a directory tree according to the hierarchical relationship, and the metadata processing structure is used to store the processing queues corresponding to the parent directories at different levels in the directory tree; when the processing unit 1102 is used to add the target file operation request to the processing queue corresponding to the first directory, it is specifically configured to perform the following steps:
[0181] If there is a processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory;
[0182] If there is no processing queue corresponding to the first directory in the metadata processing structure, after creating the processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory.
[0183] In one implementation, a global lock is added to the metadata processing structure. Obtain the global lock to have the update permission for the metadata processing structure. The processing unit 1102 is used to perform the following steps specifically when creating the processing queue corresponding to the first directory in the metadata processing structure:
[0184] Obtain the global lock;
[0185] After obtaining the global lock, create the processing queue corresponding to the first directory in the metadata processing structure.
[0186] In one implementation, the processing unit 1102 is further used to perform the following steps:
[0187] If it is detected that there is an empty processing queue in the metadata processing structure, obtain the global lock; an empty processing queue refers to a processing queue in which the included file operation requests have all been released;
[0188] After obtaining the global lock, delete the empty processing queue in the metadata processing structure.
[0189] In one implementation, the file operation request is transmitted based on the transmission protocol;
[0190] The communication unit 1101 is used to perform the following steps specifically when receiving the target file operation request: create a processing thread corresponding to the transmission protocol; call the processing thread to receive the target file operation request;
[0191] The processing unit 1102 is used to perform the following steps specifically when adding the target file operation request to the processing queue corresponding to the first directory: reuse the processing thread corresponding to the transmission protocol and add the target file operation request to the processing queue corresponding to the first directory;
[0192] The processing unit 1102 is used to perform the following steps specifically when updating the metadata database according to the target file operation request: reuse the processing thread corresponding to the transmission protocol and update the metadata database according to the target file operation request.
[0193] In one implementation, the processing unit 1102 is further used to perform the following steps:
[0194] After updating the metadata cache, reuse the processing thread corresponding to the transmission protocol to notify the metadata update result corresponding to the target file operation request.
[0195] According to another embodiment of the present application, Figure 11Each unit in the processing device for file system data shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the processing device for file system data can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.
[0196] According to another embodiment of this application, it can be achieved by running a computer program capable of executing each step involved in part or all of the methods shown on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM). Figure 8 or Figure 10 shown to construct the processing device for file system data shown in Figure 11 and to implement the method for processing file system data of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above computing device through the computer-readable storage medium, and run therein.
[0197] In the embodiments of this application, when a target file operation request is received, the target file operation request can be added to the processing queue corresponding to the first directory. The file requested to be operated on by the target file operation request is the target file, and the first directory is the parent directory to which the target file belongs; the processing queue corresponding to the first directory includes at least one file operation request. The parent directories of the files requested to be processed by the file operation requests in the processing queue corresponding to the first directory are all the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the time of receiving the requests. The metadata database can be updated according to the target file operation request; then, based on the metadata database update result of the target file operation request, the metadata cache can be updated in the arrangement order of the target file operation request in the processing queue corresponding to the first directory. It can be seen that in the embodiments of this application, when a file operation request is received, the metadata database can be updated first, and then the metadata cache can be updated according to the metadata database update result, which can ensure the consistency of metadata updates; moreover, the file operation requests in the processing queue corresponding to the first directory update the metadata cache in the order of the time of receiving the requests, which can ensure the correctness of metadata updates. That is to say, by adopting the embodiments of this application, the consistency and correctness of metadata updates can be ensured.
[0198] Based on the above method and apparatus embodiments, an embodiment of the present application provides a computer device. Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. Figure 12 The computer device shown at least includes a processor 1201, an input interface 1202, an output interface 1203, and a computer-readable storage medium 1204. Among them, the processor 1201, the input interface 1202, the output interface 1203, and the computer-readable storage medium 1204 can be connected by a bus or other means.
[0199] The computer-readable storage medium 1204 can be stored in the memory of the computer device. The computer-readable storage medium 1204 is used to store a computer program, and the computer program includes computer instructions. The processor 1201 is used to execute the computer program stored in the computer-readable storage medium 1204. The processor 1201 (or CPU (Central Processing Unit, central processing unit)) is the computing core and control core of the computer device, which is suitable for implementing the computer program, specifically suitable for loading and executing the computer program to implement the corresponding method flow or corresponding function.
[0200] An embodiment of the present application also provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device, which is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device, and of course can also include the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the computer device is stored in this storage space. And, a computer program suitable for being loaded and executed by the processor is also stored in this storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0201] The computer device can be the metadata service device 205 in the Figure 2 file system shown above; the file system data includes metadata and file data, and the metadata is stored in the metadata cache and the metadata database; the file data includes multiple files, and the metadata includes the description information of each file in the file data. In a specific implementation, the processor 1201 can be used to load and execute the computer program stored in the computer-readable storage medium 1204 to implement the above-mentioned Figure 8 or Figure 10The corresponding steps in the method for processing file system data shown. In a specific implementation, the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to perform the following steps:
[0202] Receive a target file operation request. The file requested to be operated on by the target file operation request is the target file, and the parent directory to which the target file belongs is the first directory;
[0203] Add the target file operation request to the processing queue corresponding to the first directory; the processing queue corresponding to the first directory includes at least one file operation request. The parent directories of the files requested to be processed by the file operation requests in the processing queue corresponding to the first directory are all the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time;
[0204] Update the metadata database according to the target file operation request;
[0205] Based on the metadata database update result of the target file operation request, update the metadata cache in the order of the arrangement of the target file operation request in the processing queue corresponding to the first directory.
[0206] In one implementation, the processing queue corresponding to the first directory is located in the metadata update structure, and the metadata update structure at least further includes a processing queue corresponding to a second directory; the first directory and the second directory are different directories;
[0207] The processing queue corresponding to the first directory and the processing queue corresponding to the second directory concurrently perform the operation of updating the metadata; the operation of updating the metadata includes updating the metadata database and updating the metadata cache.
[0208] In one implementation, each file operation request in the processing queue corresponding to the first directory updates the metadata cache based on the corresponding file operation request in the arranged order; the head of the processing queue corresponding to the first directory has a parent directory lock, and the file operation request that obtains the parent directory lock has the permission to update the metadata cache;
[0209] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to update the metadata cache based on the metadata database update result of the target file operation request in the order of the arrangement of the target file operation request in the processing queue corresponding to the first directory, it is specifically used to perform the following steps:
[0210] When the target file operation request is at the head of the queue, add the parent directory lock to the target file operation request;
[0211] If the metadata database update result of the target file operation request indicates that the metadata database is successfully updated according to the target file operation request, then update the metadata cache according to the target file operation request.
[0212] In one implementation, the computer program in the computer-readable storage medium 1204 is loaded by the processor 1201 and is also used to execute the following steps:
[0213] If the metadata database update result of the target file operation request indicates that the update of the metadata database fails according to the target file operation request, then release the target file operation request from the head of the queue;
[0214] Move the file operation request that is arranged one position after the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
[0215] In one implementation, the computer program in the computer-readable storage medium 1204 is loaded by the processor 1201 and is also used to execute the following steps:
[0216] After successfully updating the metadata cache according to the target file operation request, release the target file operation request from the head of the queue;
[0217] Move the file operation request that is arranged one position after the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
[0218] In one implementation, the operation request type of the target file operation request is the target operation request type. After the target file operation request is structured into the target file operation log corresponding to the target operation request type, it is added to the processing queue corresponding to the first directory; the target file operation log includes the operation data of the target file operation request;
[0219] When the computer program in the computer-readable storage medium 1204 is loaded by the processor 1201 and executes to update the metadata database according to the target file operation request, it is specifically used to execute the following steps:
[0220] Update the metadata database according to the operation data of the target file operation request in the target file operation log;
[0221] When the computer program in the computer-readable storage medium 1204 is loaded by the processor 1201 and executes to update the metadata cache based on the metadata database update result of the target file operation request and in accordance with the arrangement order of the target file operation request in the processing queue corresponding to the first directory, it is specifically used to execute the following steps:
[0222] Based on the update result of the metadata database for the target file operation request, according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory, update the metadata cache according to the operation data of the target file operation request in the target file operation log.
[0223] In one implementation, the parent directories to which different files in the file data belong form a directory tree in a hierarchical relationship, and the metadata processing structure is used to store the processing queues corresponding to the parent directories at different levels in the directory tree; when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to add the target file operation request to the processing queue corresponding to the first directory, it is specifically used to execute the following steps:
[0224] If there is a processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory;
[0225] If there is no processing queue corresponding to the first directory in the metadata processing structure, after creating the processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory.
[0226] In one implementation, a global lock is added to the metadata processing structure. Obtain the global lock to have the update permission for the metadata processing structure; when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to create the processing queue corresponding to the first directory in the metadata processing structure, it is specifically used to execute the following steps:
[0227] Obtain the global lock;
[0228] After obtaining the global lock, create the processing queue corresponding to the first directory in the metadata processing structure.
[0229] In one implementation, the computer program in the computer-readable storage medium 1204 is also used to be loaded and executed by the processor 1201 to execute the following steps:
[0230] If it is detected that there is an empty processing queue in the metadata processing structure, obtain the global lock; an empty processing queue refers to a processing queue in which the included file operation requests have all been released;
[0231] After obtaining the global lock, delete the empty processing queue in the metadata processing structure.
[0232] In one implementation, the file operation request is transmitted based on a transmission protocol.
[0233] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to receive a target file operation request, it is specifically used to perform the following steps: create a processing thread corresponding to the transmission protocol; call the processing thread to receive the target file operation request;
[0234] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to add the target file operation request to the processing queue corresponding to the first directory, it is specifically used to perform the following steps: reuse the processing thread corresponding to the transmission protocol and add the target file operation request to the processing queue corresponding to the first directory;
[0235] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to update the metadata database according to the target file operation request, it is specifically used to perform the following steps: reuse the processing thread corresponding to the transmission protocol and update the metadata database according to the target file operation request.
[0236] In one implementation, the computer program in the computer-readable storage medium 1204 is loaded and further used by the processor 1201 to perform the following steps:
[0237] After updating the metadata cache, reuse the processing thread corresponding to the transmission protocol to notify the metadata update result corresponding to the target file operation request.
[0238] In the embodiments of the present application, when a target file operation request is received, the target file operation request can be added to the processing queue corresponding to the first directory. The file requested to be operated by the target file operation request is the target file, and the first directory is the parent directory to which the target file belongs; the processing queue corresponding to the first directory includes at least one file operation request. The parent directories of the files requested to be processed by the file operation requests in the processing queue corresponding to the first directory are all the first directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time. The metadata database can be updated according to the target file operation request; then, based on the metadata database update result of the target file operation request, the metadata cache can be updated in the arrangement order of the target file operation request in the processing queue corresponding to the first directory. It can be seen that in the embodiments of the present application, when a file operation request is received, the metadata database can be updated first, and then the metadata cache can be updated according to the metadata database update result, which can ensure the consistency of metadata update; moreover, the file operation requests in the processing queue corresponding to the first directory update the metadata cache in the order of the request reception time, which can ensure the correctness of metadata update. That is to say, by adopting the embodiments of the present application, the consistency and correctness of metadata update can be ensured.
[0239] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for processing file system data.
[0240] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0241] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0242] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0243] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.
Claims
1. A method for processing file system data, characterized in that, The file system data includes metadata and file data. The metadata is stored in a metadata cache and a metadata database. The file data includes multiple files, and the metadata includes description information of each file in the file data. The method includes: Receiving a target file operation request, where the file requested to be operated on by the target file operation request is a target file, and the parent directory to which the target file belongs is a first directory; Adding the target file operation request to a processing queue corresponding to the first directory. The processing queue corresponding to the first directory includes at least one file operation request. The files requested to be processed by the file operation requests in the processing queue corresponding to the first directory all belong to the first directory as their parent directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the time when the requests are received; Updating the metadata database according to the target file operation request; Based on the metadata database update result of the target file operation request, updating the metadata cache in the order of the arrangement of the target file operation request in the processing queue corresponding to the first directory.
2. The method according to claim 1, characterized in that, The processing queue corresponding to the first directory is located in a metadata update structure. The metadata update structure at least further includes a processing queue corresponding to a second directory. The first directory and the second directory are different directories; The processing queue corresponding to the first directory and the processing queue corresponding to the second directory concurrently execute the operation of updating the metadata; The operation of updating the metadata includes updating the metadata database and updating the metadata cache.
3. The method according to claim 1 or 2, characterized in that, Each file operation request in the processing queue corresponding to the first directory updates the metadata cache in sequence based on the corresponding file operation request according to the arrangement order; The head of the processing queue corresponding to the first directory has a parent directory lock. The file operation request that obtains the parent directory lock has the permission to update the metadata cache; The updating the metadata cache based on the metadata database update result of the target file operation request in the order of the arrangement of the target file operation request in the processing queue corresponding to the first directory includes: When the target file operation request is at the head of the queue, adding the parent directory lock to the target file operation request; If the metadata database update result of the target file operation request indicates that the metadata database is successfully updated according to the target file operation request, then updating the metadata cache according to the target file operation request.
4. The method according to claim 3, characterized in that, The method further includes: If the metadata database update result of the target file operation request indicates that updating the metadata database according to the target file operation request fails, then releasing the target file operation request from the head of the queue; Moving the file operation request that is arranged one position after the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
5. The method according to claim 3, characterized in that, The method further includes: After successfully updating the metadata cache according to the target file operation request, releasing the target file operation request from the head of the queue; Move the file operation request that is one position after the target file operation request in the processing queue corresponding to the first directory to the head of the queue.
6. The method according to claim 1 or 2, characterized in that, The operation request type of the target file operation request is the target operation request type. After the target file operation request is structured into the target file operation log corresponding to the target operation request type, it is added to the processing queue corresponding to the first directory. The target file operation log includes the operation data of the target file operation request. The updating of the metadata database according to the target file operation request includes: Updating the metadata database according to the operation data of the target file operation request in the target file operation log. The updating of the metadata cache based on the metadata database update result of the target file operation request, according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory, includes: Based on the metadata database update result of the target file operation request, according to the arrangement order of the target file operation request in the processing queue corresponding to the first directory, update the metadata cache according to the operation data of the target file operation request in the target file operation log.
7. The method according to claim 1 or 2, characterized in that, The parent directories to which different files in the file data belong form a directory tree according to the hierarchical relationship. The metadata processing structure is used to store the processing queues corresponding to the parent directories at different levels in the directory tree. The adding of the target file operation request to the processing queue corresponding to the first directory includes: If there is a processing queue corresponding to the first directory in the metadata processing structure, add the target file operation request to the processing queue corresponding to the first directory. If there is no processing queue corresponding to the first directory in the metadata processing structure, create a processing queue corresponding to the first directory in the metadata processing structure, and then add the target file operation request to the processing queue corresponding to the first directory.
8. The method according to claim 7, characterized in that, A global lock is added to the metadata processing structure. Obtaining the global lock grants the update permission for the metadata processing structure. The creating of the processing queue corresponding to the first directory in the metadata processing structure includes: Obtain the global lock. After obtaining the global lock, create a processing queue corresponding to the first directory in the metadata processing structure.
9. The method according to claim 8, characterized in that, The method further includes: If it is detected that there is an empty processing queue in the metadata processing structure, obtain the global lock. The empty processing queue refers to the processing queue in which all the included file operation requests have been released. After obtaining the global lock, delete the empty processing queue in the metadata processing structure.
10. The method according to claim 1 or 2, characterized in that, The file operation request is transmitted based on the transmission protocol. The receiving of the target file operation request includes: creating a processing thread corresponding to the transmission protocol; calling the processing thread to receive the target file operation request. The adding of the target file operation request to the processing queue corresponding to the first directory includes: reusing the processing thread corresponding to the transmission protocol to add the target file operation request to the processing queue corresponding to the first directory. Updating the metadata database according to the target file operation request includes: reusing the processing thread corresponding to the transport protocol and updating the metadata database according to the target file operation request.
11. The method according to claim 10, characterized in that, The method further includes: After updating the metadata cache, reusing the processing thread corresponding to the transport protocol to notify the metadata update result corresponding to the target file operation request.
12. A processing device for file system data, characterized in that, The file system data includes metadata and file data. The metadata is stored in a metadata cache and a metadata database; the file data includes a plurality of files, and the metadata includes description information of each file in the file data. The apparatus includes: A communication unit, configured to receive a target file operation request, where the file requested to be operated by the target file operation request is a target file, and the parent directory to which the target file belongs is a first directory; A processing unit, configured to add the target file operation request to a processing queue corresponding to the first directory; the processing queue corresponding to the first directory includes at least one file operation request, the files requested to be processed by the file operation requests in the processing queue corresponding to the first directory all belong to the first directory as the parent directory, and the file operation requests in the processing queue corresponding to the first directory are arranged in the order of the request reception time; The processing unit is further configured to update the metadata database according to the target file operation request; The processing unit is further configured to update the metadata cache based on the metadata database update result of the target file operation request in the order of arrangement of the target file operation request in the processing queue corresponding to the first directory.
13. A computer device, characterized in that, The computer device includes: A processor, adapted to implement a computer program; A computer-readable storage medium storing a computer program, the computer program being adapted to be loaded and executed by the processor to perform the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program being adapted to be loaded and executed by a processor to perform the method according to any one of claims 1-11.
15. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-11.