Data processing method and device, equipment, storage medium and computer program product
By adopting a multi-level storage architecture in a distributed file system, and leveraging the advantages of disk engines and memory engines, high-throughput and low-latency large-capacity metadata storage is achieved, solving the contradiction between metadata storage performance and reliability in the existing technology.
Patent Information
- Application Number
- CN202510024969.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to meet the needs of metadata in distributed file systems, namely high throughput, low latency, large capacity, high reliability, and low cost.
It adopts a multi-level storage architecture, including a data layer that stores full metadata and a cache layer that stores hotspot metadata. The data layer adopts a disk engine, and the cache layer adopts a memory engine. Through this architecture, metadata is written to the data layer and/or cache layer, and when reading and operating metadata, it utilizes the high performance of the memory engine and the large capacity and low cost of the disk engine.
It realizes high throughput and low latency metadata storage in distributed file systems, and has the advantages of large capacity, high reliability and low cost, solving the contradiction between metadata storage performance and reliability in the existing technology.
Smart Images

Figure CN120045538A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless communication technologies, and in particular, to a data processing method, apparatus, device, storage medium, and computer program product. Background Art
[0002] Currently, in a distributed file system, its key challenges include larger data sets, more small files, higher throughput, lower latency, and so on. As the index of the file system, metadata is a key link in the performance, scalability, etc. of the distributed file system. As the storage foundation of metadata, the storage engine is an important factor reflecting the characteristics of metadata. From the perspective of the storage medium, the storage engine can be divided into a disk storage engine and a memory storage engine. However, in related technologies, when storing the metadata of files in a distributed file system, the requirements of metadata, namely high throughput, low latency, large capacity, high reliability, and low cost, cannot be met. Summary of the Invention
[0003] In view of this, embodiments of the present application are expected to provide a data processing method, apparatus, device, storage medium, and computer program product.
[0004] The technical solution of the embodiments of the present application is implemented as follows:
[0005] Embodiments of the present application provide a data processing method, which is applied to a storage architecture. The storage architecture includes a data layer for storing all metadata and a cache layer for storing hot metadata. The data layer uses a disk engine, and the cache layer uses a memory engine. The method includes:
[0006] Writing the metadata into the data layer and / or the cache layer;
[0007] Wherein, the metadata includes at least one of the following:
[0008] All metadata; the all metadata are all the metadata corresponding to files and directories in the distributed file system respectively;
[0009] Hot metadata; the hot metadata are the metadata with an access frequency greater than a first preset threshold among all the metadata corresponding to the file and the metadata with an access frequency greater than a second preset threshold among all the metadata corresponding to the directory.
[0010] In addition, according to at least one embodiment of the present application, the storage architecture further includes an access layer;
[0011] Among them, the access layer stores the routing information of the metadata and the semantic information of the file. The routing information of the metadata represents the storage path of the metadata, and the semantic information of the file represents the logic for performing operations on the metadata. The access layer performs read and write interactions of the metadata with the cache layer.
[0012] In addition, according to at least one embodiment of the present application, the data layer includes at least one of the following:
[0013] The first layer; the first layer is used to receive a file metadata request sent by the cache layer. The file metadata request is used to request writing or reading the full amount of metadata. The file metadata request carries a first parameter and permissions. The first parameter represents the attribute information of the file and the directory, and the permissions represent the access permissions of the user. Parse the file metadata request sent by the cache layer to obtain the first parameter and permissions, and verify the first parameter and permissions.
[0014] The second layer; the second layer is used to maintain the routing information of the full amount of metadata stored in the data layer.
[0015] The third layer; the third layer is used to store the semantic information of the file. The semantic information of the file represents the logic for performing operations on the full amount of metadata.
[0016] The fourth layer; the fourth layer is used to define the data structure of the full amount of metadata.
[0017] The fifth layer; the fifth layer is used to encode and decode the full amount of metadata to obtain corresponding key-value pairs.
[0018] The sixth layer; the sixth layer is used to encapsulate the interface for performing basic operations on the seventh layer for the third layer to call.
[0019] The seventh layer; the seventh layer is a disk engine used to store the full amount of metadata.
[0020] In addition, according to at least one embodiment of the present application, the cache layer includes:
[0021] The first layer; the first layer is used to store the routing information of the hot metadata and the semantic information of the file. The routing information of the hot metadata represents the storage path of the hot metadata, and the semantic information of the file represents the logic for performing operations on the hot metadata. And perform read and write interactions of the metadata with the access layer.
[0022] The second layer; the second layer is a memory engine used to store the hot metadata. And perform read and write interactions with the data layer.
[0023] In addition, according to at least one embodiment of the present application, writing the metadata to the data layer and / or the cache layer includes:
[0024] When the metadata is hot metadata, write the hot metadata to the cache layer and the data layer through a first writing method;
[0025] And / or,
[0026] When the metadata is non-hot metadata, write the non-hot metadata to the data layer through a second writing method.
[0027] In addition, according to at least one embodiment of the present application, the method further includes:
[0028] Read the hot metadata from the cache layer through a first reading method. If the hot metadata is not read from the cache layer, synchronously read the hot metadata from the data layer;
[0029] Or,
[0030] Read the hot metadata from the cache layer through a second reading method. If the hot metadata is not read from the cache layer, asynchronously read the hot metadata from the data layer.
[0031] In addition, according to at least one embodiment of the present application, the method further includes:
[0032] Eliminate the hot metadata in the cache layer whose unused duration is greater than a third preset threshold through a first operation;
[0033] Or,
[0034] Eliminate the hot metadata in the cache layer whose access times are less than a fourth preset threshold through a second operation.
[0035] At least one embodiment of the present application provides a data processing device, including:
[0036] A processing module, configured to write metadata to a data layer storing all metadata of the storage architecture and / or a cache layer storing hot metadata; the data layer uses a disk engine, and the cache layer uses a memory engine; wherein, the metadata includes at least one of the following: all metadata; the all metadata is all metadata respectively corresponding to files and directories in a distributed file system; hot metadata; the hot metadata is metadata in all metadata corresponding to the file whose access frequency is greater than a first preset threshold and metadata in all metadata corresponding to the directory whose access frequency is greater than a second preset threshold
[0037] At least one embodiment of the present application provides an electronic device, including a processor and a memory for storing a computer program that can run on the processor.
[0038] Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of the above.
[0039] At least one embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0040] At least one embodiment of the present application provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the method described in any one of the above.
[0041] The data processing method, device, equipment, storage medium and computer program product provided by the embodiments of the present application, the method is applied to a storage architecture, the storage architecture includes a data layer for storing all metadata and a cache layer for storing hot metadata, the data layer uses a disk engine, and the cache layer uses a memory engine; the method includes: writing metadata into the data layer and / or the cache layer; wherein, the metadata includes at least one of the following: all metadata; the all metadata is all the metadata corresponding to files and directories in a distributed file system; hot metadata; the hot metadata is the metadata in all the metadata corresponding to the file whose access frequency is greater than a first preset threshold and the metadata in all the metadata corresponding to the directory whose access frequency is greater than a second preset threshold.
[0042] Adopting the technical solution provided by the embodiments of the present application, a multi-level storage architecture is adopted, which can utilize the characteristics of high throughput and low latency of the memory engine, and at the same time utilize the advantages of large capacity, low cost and high reliability of the disk engine to solve the demand for metadata in high-performance file storage in certain scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic flowchart of the implementation of the data processing method in the embodiments of the present application;
[0044] Figure 2 is a schematic diagram of the storage architecture in the embodiments of the present application;
[0045] Figure 3 is a schematic diagram of the composition structure of the data layer in the embodiments of the present application;
[0046] Figure 4 is a schematic diagram of the reading and writing of metadata in the embodiments of the present application;
[0047] Figure 5 is a schematic diagram of the expansion of the number of replicas in the embodiments of the present applicationFigure 1 ;
[0048] Figure 6 is a schematic illustration of the copy number expansion of the embodiments of the present application Figure 2 ;
[0049] Figure 7 is a schematic illustration of the horizontal splitting expansion of the embodiments of the present application Figure 1 ;
[0050] Figure 8 is a schematic illustration of the horizontal splitting expansion of the embodiments of the present application Figure 2 ;
[0051] Figure 9 is a schematic diagram of the composition structure of the data processing device in the embodiments of the present application;
[0052] Figure 10 is a schematic diagram of the composition structure of the electronic device in the embodiments of the present application. Detailed implementation manners
[0053] Before introducing the technical solutions of the embodiments of the present application, the related technologies will be introduced first.
[0054] In the related technologies, with the rapid development of global artificial intelligence, intelligent technologies are empowering various industries. For example, they have been deeply integrated into fields such as intelligent driving, medical imaging, and financial risk control. However, at the same time, many new challenges have been posed to artificial intelligence. For example, recently, GPT-4 in the field of artificial general intelligence (AGI) has attracted much attention, making large AI models another direction that many countries and organizations are competing to focus on in the field of artificial intelligence. Large AI models can be applied to scenarios such as machine translation, automated question answering, and intelligent customer service, helping enterprises achieve business automation, improve customer response, meet people's growing service needs, and save human resources at the same time.
[0055] Data is the cornerstone of intelligent computing scenarios such as large AI models. More high-quality feature data, a large number of model parameters, deep data structures, etc. all pose more challenges to distributed storage. In the distributed file system, the key challenges include larger data sets, more small files, higher throughput, lower latency, and so on. And as the index of the file system, metadata is a key link in the performance, scalability, etc. of the distributed file system.
[0056] As the storage cornerstone of metadata, the storage engine is an important factor reflecting the characteristics of metadata. From the perspective of storage media, the storage engine can be divided into a disk storage engine and a memory storage engine.
[0057] In the related technologies, the following two solutions are included:
[0058] Solution 1: A distributed metadata solution based on a disk engine. In this solution, disks are mainly used as the storage medium for the metadata of the file system, and memory is only used as an auxiliary medium for a small amount of data buffering and index acceleration. Based on a single-machine disk storage engine, distributed consistency, high availability, etc. are achieved based on Multi Raft or the distributed consensus algorithm (Paxos). Its advantages include: 1) Disks are persistent storage, which can guarantee data persistence. In case of failures such as system crashes, data on a single machine can still be guaranteed not to be lost, and data recovery can be performed through certain means. 2) The capacity of disks is much larger than that of memory. From the perspective of a single machine, more data can be stored. 3) Disks are less costly than memory, which can reduce the total system cost.
[0059] However, the technical defect of Solution 1 is that since the performance of disks is much lower than that of memory, the read and write performance of metadata is poor, mainly manifested in high latency and low throughput. However, in storage scenarios such as intelligent computing with high-performance distributed file systems, one of the core requirements is high throughput and low latency. Therefore, the distributed metadata solution based on a disk engine cannot well meet the requirements of relevant scenarios for high-performance metadata in distributed file systems.
[0060] Solution 2: A distributed metadata solution based on a memory engine. In this solution, metadata is mainly stored in memory, and disks are used as auxiliary storage media such as Append Log. Similarly, based on a single-machine memory storage engine, distributed consistency, high availability, etc. are achieved based on MultiRaft or Paxos. Its advantages are reflected in that the read and write performance of memory is significantly higher than that of disks, mainly manifested in the ability to respond to metadata read and write with lower latency and higher throughput.
[0061] However, the technical defects of Solution 2 are: 1) Since memory is volatile storage, data loss is likely to occur in case of single-machine failures such as power outages, and data reliability is low. 2) The price of memory is high, and the cost per unit of the same storage capacity is much higher than that of disks. 3) In terms of the configurable storage capacity of a single machine, it is also much lower than that of disks.
[0062] In summary, in related technologies, when storing the metadata of files in a distributed file system, the requirements for metadata, namely high throughput, low latency, large capacity, high reliability, and low cost, cannot be met.
[0063] See Figure 1 , Figure 1It is a schematic flowchart of the implementation of the data processing method according to an embodiment of the present application, which is applied to a storage architecture. The storage architecture includes a data layer for storing all metadata and a cache layer for storing hot metadata. The data layer uses a disk engine, and the cache layer uses a memory engine; as Figure 1 shown, the method includes:
[0064] Writing the metadata into the data layer and / or the cache layer;
[0065] Among them, the metadata includes at least one of the following:
[0066] All metadata; the all metadata is all the metadata corresponding to files and directories in a distributed file system respectively;
[0067] Hot metadata; the hot metadata is the metadata whose access frequency is greater than a first preset threshold among all the metadata corresponding to the file and the metadata whose access frequency is greater than a second preset threshold among all the metadata corresponding to the directory.
[0068] It can be understood that the distributed file system is a file system that connects multiple hosts through a network and performs distributed file storage on these hosts. In the distributed file system, the directory (or described as a folder) is used to manage files.
[0069] It can be understood that all the metadata corresponding to the file may refer to all the attribute information of the file, such as file type, permission to perform operations on the file, file size, time when the file is accessed, and so on.
[0070] It can be understood that all the metadata corresponding to the directory may refer to all the attribute information describing the directory data, such as directory name, permission, time, directory structure, and so on.
[0071] It can be understood that the metadata whose access frequency is greater than the first preset threshold among all the metadata corresponding to the file may refer to the attribute information whose access frequency is greater than the first preset threshold among all the attribute information of the file, such as file type, etc.
[0072] It can be understood that the metadata whose access frequency is greater than the second preset threshold among all the metadata corresponding to the directory may refer to the attribute information whose access frequency is greater than the second preset threshold among all the attribute information of the directory data, such as directory name, etc.
[0073] It can be understood that the first preset threshold and the second preset threshold can be set according to the actual situation.
[0074] Here, writing the metadata into the data layer and / or the cache layer includes:
[0075] When the metadata includes the full - volume metadata, write the full - volume metadata to the data layer;
[0076] Alternatively, when the metadata includes the hot - spot metadata, write the hot - spot metadata to the cache layer and the data layer;
[0077] Alternatively, when the metadata includes the full - volume metadata and the hot - spot metadata, write the full - volume metadata to the data layer, and write the hot - spot metadata to the cache layer and the data layer.
[0078] It should be noted that in this application, according to the characteristics of the disk engine, which has large capacity, high scalability, and low cost, using the disk engine to store the full - volume metadata of files and directories can meet the requirements of this metadata, namely large capacity, high reliability, and low cost.
[0079] It should be noted that in this application, according to the characteristics of the memory engine, which has high throughput, low latency and other high - performance characteristics, using the memory engine to store the hot - spot metadata with obvious hot - spot characteristics can meet the requirements of this metadata, namely high throughput and low latency.
[0080] In practical applications, in order to implement the read and write operations on the metadata, the storage architecture may further include an access layer.
[0081] Based on this, in some embodiments, the storage architecture further includes an access layer;
[0082] Among them, the access layer stores the routing information of the metadata and the semantic information of the file. The routing information of the metadata represents the storage path of the metadata, and the semantic information of the file represents the logic for performing operations on the metadata; the access layer performs read and write interactions of the metadata with the cache layer.
[0083] It can be understood that the routing information of the metadata stored by the access layer includes the routing information of the full - volume metadata and the routing information of the hot - spot metadata; among them, the routing information of the full - volume metadata refers to the storage path of the full - volume metadata, and this storage path represents the mapping relationship between the full - volume metadata and the storage node; the routing information of the hot - spot metadata refers to the storage path of the hot - spot metadata, and this storage path represents the mapping relationship between the hot - spot metadata and the storage node.
[0084] It can be understood that the semantic information of the file represents the logic for performing operations on the metadata, and may include the logic for performing read and write and other operations on the full - volume metadata, as well as the logic for performing read and write and other operations on the hot - spot metadata.
[0085] It can be understood that the read and write interaction of the metadata between the access layer and the cache layer may include the read and write interaction between the access layer and the cache layer, and / or the read and write interaction between the access layer and the data layer through the cache layer.
[0086] See Figure 2 , Figure 2 which is a schematic diagram of the storage architecture of an embodiment of the present application. As Figure 2 shown, the present application designs a multi-level storage architecture, which from bottom to top are as follows:
[0087] The bottom layer is the data layer (or described as the full-volume data layer), which uses a disk engine and is used to store the full-volume metadata corresponding to the files and directories in the distributed file system, and has the characteristics of large capacity, high scalability, and low cost.
[0088] It should be noted that the disk engine can be partitioned. For example, it can be divided into multiple partitions, each partition is called a Paritition, one Paritition constitutes a Raft Group, multiple Parititions constitute multiple (Multi) Raft Groups, each partition includes multiple replicas, and the replicas are also called Replicas. Each replica is located on a different node. For example, the replica located on the Leader node is called the Leader Replica, and the replica located on the Follower node is called the Follower Replica.
[0089] The middle layer is the cache layer (or described as the hot data cache layer), which uses a memory engine and is used to store the hot metadata with obvious hotness. The hot metadata is the metadata with an access frequency greater than the first preset threshold in all the metadata corresponding to the file and the metadata with an access frequency greater than the second preset threshold in all the metadata corresponding to the directory, and has high-performance characteristics such as high throughput and low latency.
[0090] It should be noted that the memory engine can be partitioned. For example, it can be divided into multiple partitions, each partition is called a Paritition, one Paritition constitutes a Raft Group, multiple Parititions constitute multiple (Multi) Raft Groups, each partition includes multiple replicas, and the replicas are also called Replicas. Each replica is located on a different node. For example, the replica located on the Leader node is called the Leader Replica, and the replica located on the Follower node is called the Follower Replica.
[0091] The topmost layer is the access layer (or described as the metadata access layer), which is used to maintain the routing information of the metadata (represented by Router), including the routing information of the hot metadata in the cache layer and the routing information of the full metadata in the data layer. Among them, the routing information of the hot metadata refers to the storage path of the hot metadata, and this storage path represents the mapping relationship between the hot metadata and the storage node. The routing information of the full metadata refers to the storage path of the full metadata, and this storage path represents the mapping relationship between the full metadata and the storage node; and it is also used to maintain the semantic information of the file, that is, the semantic logic of the file at the metadata level (represented by File Semantics). It can also directly connect to the cache layer through the SDK to perform read and write interactions of the metadata. That is, the Meta SDK is installed in this access layer, and the Meta SDK refers to a development kit that provides specific functions (such as performing read and write interactions of metadata with the cache layer).
[0092] In actual application, the data layer can adopt a hierarchical logical structure.
[0093] Based on this, in some embodiments, the data layer includes at least one of the following:
[0094] The first layer; the first layer is used to receive the file metadata request sent by the cache layer. The file metadata request is used to request to write or read the full metadata; the file metadata request carries a first parameter and permissions. The first parameter represents the attribute information of the file and the directory, and the permissions represent the access permissions of the user; parse the file metadata request sent by the cache layer to obtain the first parameter and permissions, and verify the first parameter and permissions;
[0095] The second layer; the second layer is used to maintain the routing information of the full metadata stored in the data layer;
[0096] The third layer; the third layer is used to store the semantic information of the file; the semantic information of the file represents the logic for performing operations on the full metadata.
[0097] The fourth layer; the fourth layer is used to define the data structure of the full metadata;
[0098] The fifth layer; the fifth layer is used to encode and decode the full metadata to obtain corresponding key-value pairs;
[0099] The sixth layer; the sixth layer is used to encapsulate the interface for performing basic operations on the seventh layer for the third layer to call;
[0100] The seventh layer; the seventh layer is the disk engine, which is used to store the full metadata.
[0101] See Figure 3 , Figure 3 which is a schematic diagram of the composition structure of the data layer in the embodiment of the present application. As Figure 3 shown, the data layer is based on disk full storage and is used to store the full metadata of files, and can support a larger storage capacity at a lower cost. The data layer (or described as the full storage layer) adopts a hierarchical logical structure, and the specific implementation is as follows:
[0102] The first layer can also be described as the semantic parsing layer (represented by File Semantics Parse), which is used to receive the first request (or described as the file metadata request) sent by the cache layer. The first request is used to request writing or reading the full metadata; the first request carries a first parameter and permissions. The first parameter characterizes the attribute information of the file and the directory, and the permissions characterize the user's access permissions; the first request (or described as the file metadata request) sent by the cache layer is parsed to obtain the first parameter and permissions, and the first parameter and permissions are verified.
[0103] The second layer can also be described as the routing management layer (represented by Router Manager), which is used to maintain the routing information of the full metadata of files and directories in the entire distributed file system. The routing information characterizes the storage path of the full metadata; the storage path characterizes the mapping relationship between the full metadata and the storage node.
[0104] The third layer can also be described as the logical processing layer of the full metadata (represented by Meta Logic Process), which encapsulates the semantic information of each file; the semantic information of the file characterizes the logic for performing operations on the full metadata. The operations can include read (Read), write (Write), create (Create), link (Link), lookup (LookUp). For example, the logic of a write (Write) operation of a file can refer to the logical determination and processing of writing metadata such as Node, Chunk, and Slice into the data layer.
[0105] The fourth layer defines the data structure of the full metadata (such as Schema metadata) (represented by Meta DataStructure), including nodes (Node), edges (Edge), chunks (Chunk), and databases (Attr), which is the lowest-level abstraction of the metadata semantics.
[0106] The fifth layer is used to encode and decode the full metadata through encoding and decoding technology (represented by Encode / Decode) to obtain corresponding key-value pairs, echoing the metadata structure on the upper layer and directly docking the storage engine on the lower layer.
[0107] The sixth layer encapsulates the interfaces for performing basic operations on the storage engine of the seventh layer (the bottom layer). These operations include Put, Get, Delete, and Scan, which are directly called by the third layer (or the logical processing layer for metadata).
[0108] The seventh layer (the bottom layer) is the storage engine layer, also known as the disk engine layer (represented by Disk Engine), which is the basic storage unit for all metadata.
[0109] In actual applications, the cache layer can adopt a hierarchical logical structure.
[0110] Based on this, in some embodiments, the cache layer includes:
[0111] The first layer; the first layer is used to store the routing information of the hot metadata and the semantic information of the file. The routing information of the hot metadata represents the storage path of the hot metadata, and the semantic information of the file represents the logic for performing operations on the hot metadata; and it performs read and write interactions of the metadata with the access layer.
[0112] The second layer; the second layer is the memory engine, which is used to store the hot metadata; and it performs read and write interactions with the data layer.
[0113] It can be understood that the storage path of the hot metadata represents the mapping relationship between the hot metadata and the storage node.
[0114] It can be understood that the semantic information of the file represents the logic for performing read and write operations on the hot metadata.
[0115] It can be understood that the cache layer can involve two cache cleaning algorithms (LRU and LFU). Among them, LRU is Least Recently Used, and LFU is Least Frequently Used.
[0116] It can be understood that the cache layer can involve partitioning the memory engine. For example, it can be divided into multiple partitions, each partition is called a Paritition, one Paritition forms a Raft Group, and multiple Parititions form multiple (Multi) Raft Groups (abbreviated as Multi Raft).
[0117] It can be understood that the cache layer storing hot metadata is basically similar in structure to the data layer storing full metadata. The main difference is that the disk engine used in the data layer is replaced with a memory engine, and at the same time, a Learner role is introduced based on Multi Raft to achieve higher-performance scalability. In addition, for the complex characteristics of file metadata, various read / write methods and hot metadata replacement schemes are designed.
[0118] In some embodiments, writing the metadata into the data layer and / or the cache layer includes:
[0119] When the metadata is hot metadata, write the hot metadata into the cache layer and the data layer through a first write method.
[0120] and / or,
[0121] When the metadata is non-hot metadata, write the non-hot metadata into the data layer through a second write method.
[0122] It can be understood that the non-hot metadata may refer to other metadata except the hot metadata, that is, the metadata with an access frequency less than or equal to a first preset threshold in all the metadata corresponding to the file and the metadata with an access frequency less than or equal to a second preset threshold in all the metadata corresponding to the directory.
[0123] Here, considering the large variety and complex semantics of file and directory metadata, two write or update methods are designed for different scenarios, namely the first write method, i.e., Write through, and the second write method, i.e., Write around.
[0124] Here, Write through is mainly used for writing / updating metadata with more obvious hotness or higher read / write performance requirements. In this method, when writing the metadata into the bottom layer of the data layer, i.e., the disk engine for storing full metadata, the metadata will also be updated to the memory engine in the cache layer for storing hot metadata at the same time. While ensuring the reliability and integrity of the data through the disk engine, the access performance of hot metadata can be improved by the hot metadata replacement algorithm in the memory engine, such as LRU, LFU, etc., to replace relatively cold metadata and evict it from the cache layer.
[0125] Here, Write around is used for writing metadata with low performance requirements or less obvious hotness. This method only writes the metadata into the bottom layer of the storage architecture, i.e., the data layer (or described as the full storage layer), without synchronously updating it to the cache layer (or described as the hot cache layer).
[0126] In some embodiments, the method further includes:
[0127] Reading the hot metadata from the cache layer by a first reading method; if the hot metadata is not read from the cache layer, synchronously reading the hot metadata from the data layer;
[0128] Or,
[0129] Reading the hot metadata from the cache layer by a second reading method; if the hot metadata is not read from the cache layer, asynchronously reading the hot metadata from the data layer.
[0130] It can be understood that, for reading the metadata of files and directories, two data reading methods are provided: a first reading method, i.e., Readsync, and a second reading method, i.e., Read aysnc, to meet the requirements of different scenarios.
[0131] Here, by the first reading method, i.e., Read sync, the hot metadata is read from the cache layer. When there is a cache miss (i.e., the hot metadata is not read from the cache layer), the cache layer will read data from the data layer. After the hot metadata is read from the data layer, while returning a response to the metadata request sent by the access layer to the access layer, the read hot metadata is synchronously updated to the cache layer. This reading method is mainly used for reading critical metadata. When a cache miss occurs, the performance of the first cache replacement is relatively low, but each subsequent read after replacement is a high-performance memory read.
[0132] Here, by the second reading method, i.e., Read async, the hot metadata is asynchronously read. When there is a cache miss (i.e., the hot metadata is not read from the cache layer), an error code for the cache miss is directly returned to the access layer, and at the same time, the hot metadata is asynchronously read from the disk engine of the data layer and synchronized to the cache layer. Although the read hit rate for this time is sacrificed, the processing performance of the upper-layer logic is guaranteed. This reading method can still maintain high performance when there is a cache miss and is applicable to data reading in non-core links.
[0133] See Figure 4 , Figure 4 is a schematic diagram of reading and writing metadata in an embodiment of the present application. As Figure 4As shown, after receiving a write request sent by the access layer, the Cache Layer writes hot metadata to the Cache Layer (using a memory engine) through the first write method, i.e., Write through, and at the same time writes hot metadata to the Data Layer, and / or writes non-hot metadata to the Data Layer through the second write method, i.e., Write around. The hot metadata and non-hot metadata written to the Data Layer can form the full metadata. Similarly, after receiving a read request sent by the access layer, the Cache Layer reads hot metadata from the Cache Layer through the first read method, i.e., Read sync. In the case of a cache miss (i.e., the hot metadata is not read in the Cache Layer), the Cache Layer will read data from the Data Layer, or, through the second read method, i.e., Read async, perform asynchronous reading of hot metadata. In the case of a cache miss (i.e., the hot metadata is not read in the Cache Layer), a error code for cache miss is directly returned to the access layer, and at the same time, hot metadata is asynchronously read from the disk engine of the Data Layer and synchronized to the Cache Layer.
[0134] In some embodiments, the method further includes:
[0135] Eliminating hot metadata in the Cache Layer whose unused duration is greater than a third preset threshold through a first operation;
[0136] Or,
[0137] Eliminating hot metadata in the Cache Layer whose number of access times is less than a fourth preset threshold through a second operation.
[0138] Here, in order to meet the metadata requirements of different sub-scenarios, two cache replacement operations, i.e., the first operation LRU and the second operation LFU, are provided respectively.
[0139] Here, LRU first eliminates the metadata that has not been used for the longest time. The performance of this elimination strategy is relatively high, but there is a situation of cache pollution. For example, occasional batch reads will eliminate a large amount of hot data, resulting in a sharp drop in the LRU hit rate. LRU is a metadata elimination strategy suitable for access time sensitivity.
[0140] Here, the LFU first eliminates the metadata that has been accessed the fewest times within a certain period, which can avoid problems such as a decrease in cache hit rate caused by occasional batch reads as much as possible. However, when the access pattern is updated, due to its dependence on historical statistical data, the hit rate will decrease, and it takes a longer time to adapt to the new access pattern. LFU is a metadata elimination strategy suitable for access frequency sensitivity.
[0141] Here, when the hotspot nature of the metadata within a specific range is particularly obvious, the best way is to expand the replicas. However, based on the Multi Raft mechanism, when data is updated, the principle of majority write success needs to be satisfied. Therefore, after replica expansion, it will affect the metadata update performance.
[0142] To solve such problems, the role of Learner is introduced. Its data replication can be asynchronously replicated from a specified Leader or Follower, without affecting the metadata update performance along with the increase in the number of replicas. At the same time, it can reduce the read / write and incremental data replication pressure on the Leader and improve scalability.
[0143] Since both the storage and access of metadata require strong expansion capabilities to meet the requirements of scenarios such as intelligent computing for high performance, large capacity, and high scalability. In this application, a flexible expansion scheme is designed. On the basis of supporting independent expansion of the cache layer (or described as the data cache layer) and the data layer (or described as the full volume data layer), to meet the metadata expansion requirements of different sub-scenarios, it supports replica number expansion and horizontal split expansion.
[0144] See Figure 5 , Figure 5 is a schematic diagram of replica number expansion in the embodiment of this application. As Figure 5 shown, when it is necessary to expand the read throughput capacity of a single partition, it can be achieved by expanding the number of replicas, and the throughput is evenly distributed to more replicas to achieve the purpose of throughput expansion. For example, a partition (Paritition) in the memory engine or disk engine includes a leader node (Leader), a follower node 1 (Follower-1), and a follower node 2 (Follower-1). The replica located at the Leader node is called the Leader Replica, and the replica located at the Follower node is called the Follower Replica. To achieve replica number expansion, a new follower node 3 (Follower-3) can be added.
[0145] See Figure 6 , Figure 6 is a schematic diagram of replica number expansion in the embodiment of this application. As Figure 6As shown, for the newly added replicas during expansion, a checkpoint operation must first be performed on the leader node or a follower node to obtain Checkpoint Files. This file is used to push the existing data (including full metadata and / or hot metadata) in the disk engine and / or memory engine to the newly added replicas (corresponding to the follower node Follower-3), and perform recovery to restore the existing data on the newly added replicas (corresponding to the follower node Follower-3). Then, incremental data is replicated from the WAL of the leader until the incremental data of the leader is caught up, completing the entire process of expanding the newly added replicas.
[0146] See Figure 7 , Figure 7 is a schematic diagram of horizontal split expansion in an embodiment of the present application. As Figure 7 shown, when it is necessary to expand the data volume or simultaneously expand the read / write throughput of metadata, it can be achieved through the method of horizontal split expansion. A single partition is split into two or even more partitions, and the data and read / write throughput are evenly distributed to more partitions to achieve the purpose of simultaneously expanding the data and throughput. For example, a certain partition (denoted as Paritition-1) in the memory engine or disk engine is expanded to partition 1 (denoted as Paritition-1) and partition 2 (denoted as Paritition-2).
[0147] See Figure 8 , Figure 8 is a schematic diagram of horizontal split expansion in an embodiment of the present application. As Figure 8 shown, on the newly added partition (denoted as Paritition-2), first perform a checkpoint operation on the source partition (denoted as Paritition-1) to obtain Checkpoint Files. This file is used to push the existing data to the newly added partition (denoted as Paritition-2), and perform replication (Recovery) of the existing data and WAL incremental data catch-up operation on the newly added partition. Then, according to the new partition rules, update the access route (Update Router), and each partition accepts read / write requests within its respective range after splitting. Finally, perform a delete range operation on the two partitions to delete the existing data that no longer belongs to this partition and release the storage space, completing the entire process of horizontal split expansion.
[0148] The embodiments of the present application have the following advantages:
[0149] (1)In view of the prior art, the distributed file system either independently uses a low-performance disk engine as the storage base for metadata or uses a volatile, low-capacity, and high-cost memory engine as the storage base for metadata. Both have their respective disadvantages and cannot balance and meet the problems of high throughput, low latency, large capacity, and high scalability in scenarios such as intelligent computing.
[0150] In this application, a multi-level storage architecture is adopted, which can utilize the characteristics of high throughput and low latency of the memory engine, and at the same time utilize the advantages of large capacity, low cost, and high reliability of the disk engine to solve the demand for metadata in high-performance file storage in scenarios such as intelligent computing. On this basis, a flexible metadata reading and writing scheme is designed for the reading and writing requirements of different metadata in different scenarios. At the same time, a more flexible expansion scheme is designed for the expansion requirements of multiple scenarios.
[0151] That is, in this application, according to the requirements for high-performance distributed files in scenarios such as intelligent computing, the demand characteristics reflected in metadata are high throughput, low latency, large capacity, high reliability, and low cost. By avoiding the respective disadvantages of the disk engine and the memory engine and at the same time utilizing their respective advantages, a distributed file metadata scheme with multi-level storage is designed. It can utilize the high-performance advantages of the memory engine to store hot data with higher access frequencies, avoid its disadvantages of high cost, volatility, and small capacity, and utilize the advantages of large capacity, low cost, persistence, and high reliability of the disk engine to store all data. Since the hot data is stored in the memory engine, the performance disadvantage of the disk engine is avoided. It supports multiple reading and writing methods to meet the detailed requirements of different metadata in multiple different scenarios. It supports independent scaling of multi-level storage to meet the detailed requirements of multiple different scenarios.
[0152] To implement the data processing method of the embodiments of this application, the embodiments of this application also provide a data processing device, which is set in an electronic device. Figure 9 It is a schematic structural diagram of the composition of the data processing device of the embodiments of this application, as Figure 9 shown, the device includes:
[0153] A processing module 91, configured to write metadata into a data layer for storing all metadata and / or a cache layer for storing hot metadata included in the storage architecture; the data layer uses a disk engine, and the cache layer uses a memory engine; wherein, the metadata includes at least one of the following: all metadata; the all metadata are all metadata corresponding to files and directories in the distributed file system respectively; hot metadata; the hot metadata are metadata with an access frequency greater than a first preset threshold among all metadata corresponding to the file and metadata with an access frequency greater than a second preset threshold among all metadata corresponding to the directory.
[0154] In some embodiments, the storage architecture further includes an access layer;
[0155] Wherein, the access layer stores the routing information of the metadata and the semantic information of the file, the routing information of the metadata represents the storage path of the metadata, and the semantic information of the file represents the logic for operating on the metadata; the access layer performs read and write interactions of the metadata with the cache layer.
[0156] In some embodiments, the data layer includes at least one of the following:
[0157] The first layer; the first layer is used to receive a file metadata request sent by the cache layer, the file metadata request is used to request to write or read the full metadata; the file metadata request carries a first parameter and permissions, the first parameter represents the attribute information of the file and the directory, and the permissions represent the access permissions of the user; parse the file metadata request sent by the cache layer to obtain the first parameter and permissions, and verify the first parameter and permissions;
[0158] The second layer; the second layer is used to maintain the routing information of the full metadata stored in the data layer;
[0159] The third layer; the third layer is used to store the semantic information of the file; the semantic information of the file represents the logic for operating on the full metadata.
[0160] The fourth layer; the fourth layer is used to define the data structure of the full metadata;
[0161] The fifth layer; the fifth layer is used to encode and decode the full metadata to obtain corresponding key-value pairs;
[0162] The sixth layer; the sixth layer is used to encapsulate the interface for performing basic operations on the seventh layer for the third layer to call;
[0163] The seventh layer; the seventh layer is a disk engine for storing the full metadata.
[0164] In some embodiments, the cache layer includes at least one of the following:
[0165] The first layer; the first layer is used to store the routing information of the hot metadata and the semantic information of the file, the routing information of the hot metadata represents the storage path of the hot metadata, and the semantic information of the file represents the logic for operating on the hot metadata; and perform read and write interactions of the metadata with the access layer;
[0166] The second layer; the second layer is a memory engine for storing the hot metadata; and performing read and write interactions with the data layer.
[0167] In some embodiments, the processing module 91 is configured to:
[0168] Read the hot metadata from the cache layer by a first reading method. If the hot metadata is not read from the cache layer, synchronously read the hot metadata from the data layer;
[0169] Or,
[0170] Read the hot metadata from the cache layer by a second reading method. If the hot metadata is not read from the cache layer, asynchronously read the hot metadata from the data layer.
[0171] In some embodiments, the processing module 91 is further configured to:
[0172] Eliminate the hot metadata in the cache layer whose unused duration is greater than a third preset threshold through a first operation;
[0173] Or,
[0174] Eliminate the hot metadata in the cache layer whose number of access times is less than a fourth preset threshold through a second operation.
[0175] In practical applications, the processing module 91 can be implemented by a processor in a data processing device.
[0176] It should be noted that: when the data processing device provided in the above embodiments performs data processing, only the above division of each program module is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the data processing device provided in the above embodiments and the data processing method embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.
[0177] The embodiment of the present application also provides an electronic device, as Figure 10 shown, including:
[0178] A communication interface 101 capable of interacting with other devices;
[0179] A processor 102 connected to the communication interface 101, configured to execute the methods provided by one or more technical solutions on the electronic device side when running a computer program. And the computer program is stored on a memory 103.
[0180] It should be noted that: For the specific processing procedures of the processor 102 and the communication interface 101, please refer to the method embodiments for details and will not be elaborated here.
[0181] Of course, in practical applications, each component in the electronic device 100 is coupled together through the bus system 104. It can be understood that the bus system 104 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 104 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 10 all kinds of buses are labeled as the bus system 104.
[0182] The memory 103 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device 100. Examples of these data include: any computer program for operating on the electronic device 100.
[0183] The method disclosed in the above embodiments of the present application can be applied to or implemented by the processor 102. The processor 102 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 102 or by instructions in software form. The above-mentioned processor 102 may be a general-purpose processor, a digital data processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 103. The processor 102 reads the information in the memory 103 and combines its hardware to complete the steps of the foregoing method.
[0184] In an exemplary embodiment, the electronic device 100 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components, and is used to execute the foregoing method.
[0185] It can be understood that the memory (memory 103) in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0186] In an exemplary embodiment, the embodiments of the present invention further provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory storing a computer program. The above computer program can be executed by the processor 102 of the electronic device 100 to complete the steps described in the foregoing method on the electronic device side. The computer-readable storage medium can be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0187] Exemplarily, the embodiments of the present application further provide a computer program product, including a computer program, which can be executed by the processor 102 of the electronic device 100 to complete the steps described in any of the foregoing methods.
[0188] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0189] In addition, the technical solutions described in the embodiments of the present invention can be arbitrarily combined without conflict.
[0190] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. A data processing method, characterized in that: Applied to a storage architecture, the storage architecture includes a data layer for storing full metadata and a cache layer for storing hotspot metadata, the data layer uses a disk engine, and the cache layer uses a memory engine; the method includes: Writing metadata to the data layer and / or the cache layer; The metadata includes at least one of the following: Full metadata; the full metadata is all metadata corresponding to the files and directories in the distributed file system; Hotspot metadata; the hotspot metadata is metadata with an access frequency greater than a first preset threshold among all metadata corresponding to the file and metadata with an access frequency greater than a second preset threshold among all metadata corresponding to the directory.
2. The method according to claim 1, characterized in that The storage architecture also includes an access layer; Among them, the access layer stores the routing information of the metadata and the semantic information of the file, the routing information of the metadata represents the storage path of the metadata, and the semantic information of the file represents the logic of performing operations on the metadata; the access layer interacts with the cache layer to read and write the metadata.
3. The method according to claim 1, characterized in that The data layer includes at least one of the following: First layer; the first layer is used to receive the file metadata request sent by the cache layer, the file metadata request is used to request to write or read the full metadata; the file metadata request carries a first parameter and permission, the first parameter represents the attribute information of the file and the directory, and the permission represents the user's access rights; Parsing the file metadata request sent by the cache layer to obtain the first parameter and permission, and verifying the first parameter and permission; The second layer is used to maintain the routing information of the full metadata stored in the data layer; Third floor; The third layer is used to store the semantic information of the file; the semantic information of the file represents the logic of performing operations on the full metadata; Fourth floor; The fourth layer is used to define the data structure of the full metadata; Fifth floor; The fifth layer is used to encode and decode the full metadata to obtain corresponding key-value pairs; Sixth floor; The sixth layer is used to encapsulate the interface of the seventh layer for performing basic operations so as to be called by the third layer; Seventh floor; The seventh layer is a disk engine, which is used to store the full metadata.
4. The method according to claim 1, characterized in that The cache layer includes: The first layer; the first layer is used to store the routing information of the hotspot metadata and the semantic information of the file, the routing information of the hotspot metadata represents the storage path of the hotspot metadata, and the semantic information of the file represents the logic of performing operations on the hotspot metadata; and to interact with the access layer to read and write the metadata; The second layer is a memory engine, which is used to store the hotspot metadata and to perform read and write interactions with the data layer.
5. The method according to claim 1, characterized in that Writing metadata to the data layer and / or the cache layer includes: When the metadata is hotspot metadata, writing the hotspot metadata into the cache layer and the data layer through a first writing method; and / or, In the case where the metadata is non-hotspot metadata, the non-hotspot metadata is written into the data layer through a second writing method.
6. The method according to claim 1, characterized in that The method further comprises: Reading the hotspot metadata from the cache layer in a first reading manner, and if the hotspot metadata is not read from the cache layer, synchronously reading the hotspot metadata from the data layer; or, The hotspot metadata is read from the cache layer through a second reading method. If the hotspot metadata is not read from the cache layer, the hotspot metadata is asynchronously read from the data layer.
7. The method according to claim 1, characterized in that The method further comprises: Through the first operation, hot spot metadata in the cache layer whose unused time is greater than a third preset threshold is eliminated; or, Through the second operation, hot spot metadata in the cache layer whose access times are less than a fourth preset threshold is eliminated.
8. A data processing device, characterized in that: include: A processing module, used to write metadata into a data layer storing full metadata and / or a cache layer storing hotspot metadata included in the storage architecture; The data layer adopts a disk engine, and the cache layer adopts a memory engine; wherein the metadata includes at least one of the following: full metadata; the full metadata is all metadata corresponding to the files and directories in the distributed file system; Hotspot metadata; the hotspot metadata is metadata with an access frequency greater than a first preset threshold among all metadata corresponding to the file and metadata with an access frequency greater than a second preset threshold among all metadata corresponding to the directory.
9. An electronic device, characterized in that: comprising a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.