A data processing method, storage medium, electronic device and program product

By determining the requesting node and the home node in the metadata server, if it is different, parsing and searching for metadata from the preset memory pool, the performance problems caused by frequent data interactions between nodes are solved, and efficient metadata processing and system performance improvement are achieved.

CN120066424BActive Publication Date: 2025-07-11INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510559350.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-11
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

In the prior art, metadata files are stored on multiple nodes in the metadata server (MDS), resulting in frequent data interactions between nodes and affecting the performance of the metadata server system.

Method used

By determining the requesting node and the home node, if it is different, the request will be parsed and processed, and the metadata will be searched from the cache space of the home node in the preset memory pool, and the metadata in the cache space will be directly read to avoid data transmission between nodes.

Benefits of technology

It improves the overall throughput of the metadata server, reduces latency, ensures high availability and consistency of the system, and improves performance and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066424B_ABST
    Figure CN120066424B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, a storage medium, an electronic device, and a program product, which relate to the field of computer technologies. The method includes: in response to a processing request from a client for a target data file, determining a request node in a metadata server for receiving the processing request and an owning node corresponding to the target data file, where the owning node is the node used when creating the target data file; if it is determined that the request node is different from the owning node, parsing the processing request; in the case where the parsed processing request is to read first metadata in the target data file, searching for the first metadata in the cache space of the owning node in a preset memory pool, where the preset memory pool is composed of the cache spaces of all nodes in the metadata server; and controlling the request node to read the first metadata from the cache space according to the search result. The present application ensures the overall high availability of the metadata server, and can also improve system performance, consistency, and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, a storage medium, an electronic device, and a program product. Background Art

[0002] File system metadata management is a core component in the field of computer storage. It involves how to efficiently, reliably store, retrieve, and maintain descriptive information about files in the file system, based on which the operating system and application programs can understand and manage the data on the storage device.

[0003] Currently, relevant technical metadata files are usually stored on multiple nodes in a metadata server (MDS). During the process of reading metadata in a data file, data interaction is required among multiple nodes, which may lead to frequent data forwarding between different nodes, thereby affecting the performance of the metadata server system. Summary of the Invention

[0004] The present disclosure provides a data processing method, a storage medium, an electronic device, and a program product. Its main purpose is to solve the problem that relevant technical metadata files are usually stored on multiple nodes in a metadata server (MDS). During the process of reading metadata in a data file, data interaction is required among multiple nodes, which may lead to frequent data forwarding between different nodes, thereby affecting the performance of the metadata server system.

[0005] In a first aspect, the present application provides a data processing method, including:

[0006] Responding to a processing request of a client for a target data file, determining a request node in the metadata server for receiving the processing request and an ownership node corresponding to the target data file, where the ownership node is the node used when creating the target data file;

[0007] If it is determined that the request node is different from the ownership node, then parsing the processing request;

[0008] When it is parsed that the processing request is to read first metadata in the target data file, searching for the first metadata in the cache space of the ownership node in a preset memory pool, where the preset memory pool is composed of the cache spaces of all nodes in the metadata server;

[0009] Controlling the request node to read the first metadata from the cache space according to the search result. In a second aspect, the present application provides a data processing device, including:

[0010] A determination module, configured to determine, in response to a processing request of a client for a target data file, a request node in a metadata server for receiving the processing request and an ownership node corresponding to the target data file, where the ownership node is the node used when creating the target data file;

[0011] A parsing module, configured to parse the processing request if it is determined that the request node is different from the ownership node;

[0012] A searching module, configured to search for first metadata in a cache space of the ownership node in a preset memory pool when it is parsed that the processing request is to read the first metadata in the target data file, and the preset memory pool is composed of cache spaces of all nodes in the metadata server;

[0013] A reading module, configured to control the request node to read the first metadata from the cache space according to the search result.

[0014] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0015] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the method of the first aspect is implemented.

[0016] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0017] The data processing method, storage medium, electronic device, and program product provided by the present disclosure, wherein the method includes: in response to a processing request of a client for a target data file, determining a request node in a metadata server for receiving the processing request and an attribution node corresponding to the target data file, where the attribution node is the node used when creating the target data file; if it is determined that the request node is different from the attribution node, parsing the processing request; in the case where the parsed processing request is to read first metadata in the target data file, searching for the first metadata in the cache space of the attribution node in a preset memory pool, where the preset memory pool is composed of the cache spaces of all nodes in the metadata server; controlling the request node to read the first metadata from the cache space according to the search result. Compared with the related art, in response to the processing request of the client, the present application needs to determine whether the request node and the attribution node of the target data file to be processed are the same, and parse the specific content of the processing request when they are different. If the parsed processing request is to read the first metadata in the target data file, it is possible to directly search whether the first metadata exists in the cache space of the attribution node in the preset memory pool based on the request node, and based on the search result, control the request node to directly read the first metadata in the cache space of the attribution node, so that the present application can, during the cross-node reading process, avoid data transmission between nodes, directly read the data in the storage space of other nodes based on the preset memory pool, thereby enabling the present application to ensure the consistency of metadata among multiple nodes, reduce the latency of metadata operations, and based on the preset memory pool, the present application can also process more concurrent requests, improve the overall throughput of the metadata server, ensure the overall high availability of the metadata server, and improve system performance, consistency, and reliability.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 The flowchart shows a data processing method provided by an embodiment of the present application;

[0021] Figure 2 The schematic diagram shows a structure of an example provided by an embodiment of the present application;

[0022] Figure 3 It shows a schematic flowchart of another data processing method provided by an embodiment of the present application;

[0023] Figure 4 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0024] Figure 5 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0025] Figure 6 It shows a schematic structural diagram of an example provided by an embodiment of the present application;

[0026] Figure 7 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0027] Figure 8 It shows a schematic structural diagram of a data processing device provided by an embodiment of the present application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0029] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.

[0030] With the rapid development of information technology, the amount of data has grown exponentially, and the scale and complexity of the file system have also expanded accordingly, making the importance of metadata management more prominent. It is not only related to the performance of the file system but also directly affects the availability, security, and stability of data. Traditional file systems were not designed to anticipate the explosive growth of today's data scale, so they may encounter performance bottlenecks when dealing with a large number of small files or highly concurrent metadata operations. In addition, the rise of distributed file systems and cloud storage platforms requires that metadata management not only be efficient but also be able to maintain consistency and reliability among multiple nodes.

[0031] File system metadata management is a core component in the field of computer storage. It involves how to efficiently, reliably store, retrieve, and maintain descriptive information about files in the file system. This information is crucial for the operation of the file system because it enables the operating system and applications to understand and manage the data on the storage device. Therefore, the design and optimization of file system metadata management have become an active area of research in storage systems, and new technologies and methods are constantly being proposed and practiced to meet the growing data management needs.

[0032] The current methods for file system metadata management mainly adopt the following several solutions: Solution 1: Centralized metadata management service, which uses a centralized metadata server (MDS) to manage all metadata. The client communicates with this server to obtain and update metadata; Solution 2: Cloud-based metadata management service, which utilizes cloud computing technology to store, manage, and analyze metadata. These services are usually provided by cloud service providers, allowing organizations to manage metadata in the cloud environment; Solution 3: Distributed metadata management service, which distributes metadata to multiple nodes, and each node is responsible for managing a part of the metadata. The client accesses the corresponding node according to the distribution of the metadata.

[0033] The disadvantages of the above three solutions can specifically include: Solution 1: ① Centralized management lacks redundant protection, and a single point of failure will cause the file system to be unavailable; ② Centralized management will lead to performance bottlenecks, especially in high-concurrency access scenarios, where performance issues are more obvious; ③ Limited scalability, making it difficult to adapt to the rapidly growing data volume; Solution 2: ① Moving data to the cloud raises issues of data security and privacy; ② The reliability, stability, and performance of the system depend on the cloud service provider, with poor autonomy; ③ Higher costs, especially for large amounts of data; Solution 3: ① It requires complex metadata synchronization and consistency mechanisms; ② It needs to solve the distribution of metadata among multiple nodes and the load balancing of metadata access.

[0034] To improve the related technology, metadata files are usually stored on multiple nodes in the metadata server (MDS). During the process of reading metadata in the data file, data interaction is required among multiple nodes, which will lead to frequent data forwarding between different nodes, thus causing technical problems that affect the performance of the metadata server system. This embodiment provides a data processing method, as Figure 1 shown. This method includes the following steps:

[0035] Step 101: In response to a client's processing request for a target data file, determine the request node in the metadata server that receives the processing request and the home node corresponding to the target data file.

[0036] Among them, the home node is the node used when creating the target data file.

[0037] In the embodiments of the present application, the metadata server is a server - side component responsible for actually storing and processing metadata. It is one of the infrastructure of the metadata management system, and can be used to store metadata: maintaining a database or repository of metadata to ensure data security and availability; can be used for query services: providing interfaces for clients to query metadata information, which may involve complex search and filtering functions; can be used for transaction support: ensuring the atomicity, consistency, isolation, and durability (ACID properties) of metadata operations, especially important in a concurrent environment; can be used for synchronization and replication: to improve reliability and performance, multiple replicas of the metadata server may be deployed and a data synchronization mechanism may be implemented; can also be used to provide application programming interface (API) support: providing standardized APIs externally so that third - party applications can easily integrate with it.

[0038] In some examples, the request node can be the node that receives and processes the processing request after the client sends the processing request; correspondingly, the home node can be the node used when the target data file is first created, where the target data file can be any one of the data files stored in the metadata server.

[0039] Exemplarily, if the client uploads a processing request for data file A, it is necessary to determine the request node that receives the processing request and the home node used when data file A is first stored in the metadata server.

[0040] Step 102: If it is determined that the request node is different from the home node, then parse the processing request.

[0041] In some examples, it is necessary to determine whether the request node and the home node of the target data file are the same. Specifically, if the request node and the home node are the same, then the target data file can be directly processed based on the processing request on this node.

[0042] As an alternative, if the request node is different from the home node, then it is necessary to parse out the specific processing content corresponding to the processing request.

[0043] Step 103: In the case where the parsed processing request is to read the first metadata in the target data file, search for the first metadata in the cache space of the home node in the preset memory pool.

[0044] Among them, the preset memory pool is composed of the cache spaces of all nodes in the metadata server.

[0045] In the embodiments of the present application, as Figure 2As shown, the preset memory pool can be a pre-set storage location that enables metadata to be globally shared and cached within each node. The process of establishing the preset memory pool can include, but is not limited to: First, each node divides a memory space, and other nodes register with these memory spaces. The memory spaces of each node are abstracted into a large memory pool for all nodes to access, and the origin of each memory in the memory pool is marked as coming from each node.

[0046] In some examples, the preset memory pool includes the cache spaces of all nodes in the metadata server, and each node can read the data cached in the cache space of the preset memory pool.

[0047] As an alternative, in the embodiment of the present application, when it is parsed that the processing request is to read the first metadata in the target data file, it is necessary to first check whether the first metadata is cached in the cache space of the home node in the preset memory pool.

[0048] Step 104, control the requesting node to read the first metadata from the cache space according to the search result.

[0049] In some examples, based on the search result, if the first metadata is cached in the preset memory pool, the first metadata in the cache space of the home node can be directly read based on the requesting node; on the contrary, if the first metadata is not cached in the preset memory pool, the first metadata can be cached in the preset memory pool by the home node, and then the first metadata in the cache space of the home node can be directly read based on the requesting node.

[0050] Compared with the related art, in response to the processing request of the client, this embodiment needs to determine whether the requesting node and the home node of the target data file to be processed are the same, and parse the specific content of the processing request in the case of non-identity. If it is parsed that the processing request is to read the first metadata in the target data file, it is possible to directly check whether the first metadata exists in the cache space of the home node in the preset memory pool based on the requesting node, and based on the search result, the requesting node can be controlled to directly read the first metadata in the cache space of the home node, so that in the process of cross-node reading, this embodiment does not require data transmission between nodes, and can directly read the data in the storage space of other nodes based on the preset memory pool. Furthermore, this embodiment can ensure the consistency of metadata among multiple nodes, reduce the latency of metadata operations, and based on the preset memory pool, this embodiment can also handle more concurrent requests, improve the overall throughput of the metadata server, ensure the overall high availability of the metadata server, and also improve system performance, consistency and reliability.

[0051] Furthermore, as a refinement and extension of the above embodiment, the following methods can be adopted but are not limited to, such asFigure 3 As shown, the method includes:

[0052] Step 201, in response to a processing request from a client for a target data file, determine a request node in the metadata server for receiving the processing request and an ownership node corresponding to the target data file.

[0053] Wherein, the ownership node is the node used when creating the target data file.

[0054] Exemplarily, when processing an already created file, it is necessary to first determine whether the file ownership node is the same as the access node. If they are the same, the metadata processing is directly performed through this node. If they are different, obtain the file ownership node. For example: the ownership node of file (file) 1 is node (node) 1, and the front-end accesses the metadata of file1 through node2, then it can be determined that the ownership node of file1 is node1 and the request node is node2.

[0055] Optionally, step 201 may specifically include: determining the request node and the file name of the target data file; determining the ownership node from the node directory list according to the file name.

[0056] Wherein, the node directory list records the mapping relationship between the file name of the metadata file stored in the metadata server and the ownership node.

[0057] In the embodiments of the present application, each node jointly maintains a directory entry list, that is, the node directory list in the embodiments of the present application, which records file names, file-owned nodes, index numbers, and so on.

[0058] Exemplarily, when a file is created for the first time, it is necessary to record the file name and the file-owned node (the node where the request is sent. For example, if the client sends a request to create file1 through node1, then in the directory entry, the ownership node of file1 is node1) in the node directory list.

[0059] Exemplarily, as Figure 4 shown, each node retains a directory entry, and at the same time each node initiates a subscription to changes in the directory entry to other nodes. For example: when the client sends a request to create file1 to node1, record (file name: file1; ownership node: node1; index number: inode1) in the directory entry of node1. At the same time, other nodes receive the subscription notice of the change in the directory entry of node1 and record a copy (file name: file1; ownership node: node1; index number: inode1) in their respective node directory entries.

[0060] Exemplarily, when generating directory entries, it is determined whether the directories recorded by each node are balanced. If not, for example, when node1 creates file1 and it is found that the files attributed to node1 are not balanced with those of other nodes, and the number of files attributed to node1 is more than that of other nodes, node1 queries which of the files attributed to it have been least recently accessed. Then, the attribution of these files is transferred to other nodes.

[0061] It should be noted that such processing can enable the embodiments of the present application to not change the attribution of file1 because file1 is created through a creation request issued by node1, and the expected value of the access probability of file1 through node1 subsequently is also the largest. Keeping the attribution of file1 unchanged can avoid subsequent cross-node access to file1; it can also make the entire balancing process dominated by node1, reducing the participation of other nodes. When a client creates a file through other nodes subsequently, the dominance of the entire process is transferred to other nodes. That is: the node that creates the file dominates, avoiding excessive interaction between nodes.

[0062] Step 202: If it is determined that the requesting node is different from the attributing node, then parse the processing request.

[0063] Exemplarily, based on the example in step 201, it can be determined that the attributing node and the requesting node of file1 are different, that is, node1 and node2 are not the same node, then the processing request needs to be parsed.

[0064] Step 203: In the case where the parsed processing request is to read the first metadata in the target data file, search for the first metadata in the cache space of the attributing node in the preset memory pool.

[0065] Among them, the preset memory pool is composed of the cache spaces of all nodes in the metadata server.

[0066] Exemplarily, based on the example in step 202, if it is a read request, it is necessary to determine whether the cache space of the attributing node node1 of file1 in the preset memory pool contains the data to be read.

[0067] Optionally, after step 203, the method of this embodiment further includes: recording the target number of times the requesting node reads the first metadata based on the attributing node in the node directory list; in the case where it is determined that the target number of times is greater than or equal to the preset number threshold, determining the requesting node as the attributing node corresponding to the target data file in the node directory list.

[0068] In some examples, the present application can change the attributing node of a file in the case of frequent cross-node access to metadata.

[0069] Exemplarily, the preset number threshold can be set according to requirements or experience, and the specific value of the preset number threshold is not limited in the embodiments of the present application.

[0070] Optionally, after step 203, the embodiments of the present application further include: in the case where it is parsed that the processing request is to perform a write process on the second metadata in the target data file, controlling the request node to send the processing request to the home node; controlling the request node to perform a write process on the target metadata based on the processing request, and sending the obtained write process result to the target metadata.

[0071] For this embodiment, if it is parsed that the processing request is a write request, the write request needs to be forwarded to the corresponding file home node for processing, and the processing result is fed back to the access node; for example, if it is a write request, the request is forwarded by node2 to node1, processed by node1, and the processing result is fed back to node2; it should be noted that this can completely avoid locks between nodes, and avoid increasing the system overhead and waiting due to applying / releasing locks; usually, the read scenario of metadata access is much more than the write scenario, and avoiding locks + direct access to metadata between nodes can further improve the performance of metadata reading; on the other hand, the data volume of the metadata write request is not large, so the information interaction between nodes is not large, and thus locks can be completely avoided, bringing a higher performance improvement.

[0072] It should be noted that the write process in the embodiments of the present application includes, but is not limited to, writing, modifying, and deleting.

[0073] Optionally, when performing "controlling the request node to perform a write process on the target metadata based on the processing request, and sending the obtained write process result to the target metadata", the following method can be used but is not limited thereto, including: writing the second metadata into the cache space, generating a write process result and sending it to the target metadata; generating a write request corresponding to the home node, and performing a write process on the second metadata in the storage disk based on the write request.

[0074] Optionally, when performing "performing a write process on the second metadata in the storage disk based on the write request", the following method can be used but is not limited thereto, including: determining the target storage address of the second metadata in the storage disk, and the target logical unit corresponding to the target storage address; performing a write process on the second metadata in the storage disk based on the write request and the target logical unit.

[0075] Optionally, when performing "writing the second metadata to the storage disk based on the write request and the target logical unit", the following method can be adopted, but is not limited thereto, including: determining the target task unit group corresponding to the target logical unit, and the target task unit in the target task unit group for executing the write request, and sending the write request to the target task unit; determining the target drive process corresponding to the target task unit, and writing the second metadata to the storage disk based on the target drive process.

[0076] Exemplarily, as Figure 5 shown, each node divides multiple task unit groups according to the number of logical units. The number of task units in each task unit group corresponds one-to-one with the number of drive processes for writing data from the logical unit to the backend storage, and each task unit and drive process pair is bound to a CPU core. When the cache IO is flushed down, the logical unit ID is obtained according to the physical address of the backend storage to be flushed down, the task unit group is obtained according to the logical unit ID, and then the task units in the task unit group are polled, and the IO request is placed in the waiting queue of the task unit. The task unit sends the IO in the waiting queue to the logical unit, and the logical unit obtains the corresponding drive process according to the task unit, and the drive process is responsible for writing the IO to the backend storage, so as to ensure that the entire flushing process is completed on one CPU core.

[0077] Optionally, after executing "controlling the request node to perform write processing on the target metadata based on the processing request and sending the obtained write processing result to the target metadata", it further includes: determining the number of files corresponding to each node in the node directory list; in the case of determining that the number of target nodes is greater than or equal to the file quantity threshold, selecting the files with the number of processing times less than or equal to the preset processing threshold from the target nodes as the files to be transferred; transferring the mapping relationship between the files to be transferred in the node directory list and the target nodes to the mapping relationship between the files to be transferred and other nodes except the target nodes.

[0078] Wherein, the target node is any one node in the node directory list.

[0079] In some examples, the embodiments of the present application can check the number of IOs to be processed in the waiting queue of each node cache flushing task unit after the write processing. If the number of IOs of a certain node is significantly more than that of other nodes, the files belonging to this node need to be transferred to other idle nodes.

[0080] Step 204, when the first metadata is found, control the request node to read the first metadata from the cache space.

[0081] Exemplarily, based on the example in step 203, if the data to be read exists in the cache space of the home node node1 in the preset memory pool, that is, the cache hits, the home node memory can be directly accessed for metadata reading, avoiding the interaction between nodes. That is, directly access the cache in node1 through node2, avoiding the interaction between node1 and node2.

[0082] Step 205, in the case where the first metadata is not found, read the first metadata from the storage disk of the metadata server into the cache space through the home node, and control the requesting node to read the first metadata in the cache space.

[0083] Optionally, step 205 may specifically include: in the case where the first metadata is not found, control the requesting node to send the processing request to the home node; control the home node to read the first metadata from the storage disk into the cache space based on the processing request.

[0084] Exemplarily, based on the example in step 203, if the data to be read does not exist in the cache space of the home node node1 in the preset memory pool, that is, the cache misses, trigger the home node to issue a disk read request, read the requested metadata into the cache, and then the requesting node directly reads the cache data. That is, forward the read request to node1, node1 initiates a disk read request, and caches the data in the memory of node1. Node2 directly accesses the cache in node. During the whole process, there is only request forwarding between node1 and node2, without data forwarding. This can make the nodes in the embodiments of the present application completely lock-free, avoiding the increase of system overhead and waiting due to application / release of locks. Usually, the read scenario of metadata access is much more than the write scenario. The combination of lock-free and direct metadata access between nodes can further improve the performance of metadata reading. On the other hand, the data volume of metadata write requests is not large, so the information interaction between nodes is not large, which can be completely lock-free and bring higher performance improvement.

[0085] Optionally, when executing "control the home node to read the first metadata from the storage disk into the cache space" in the specific content of step 205, the following method can be adopted but is not limited to this, including: generate a disk read request corresponding to the home node, send the disk read request to the task unit corresponding to the home node, and determine the logical unit corresponding to the home node; based on the driver corresponding to the logical unit, read the first metadata from the disk address of the storage disk corresponding to the logical unit into the cache space.

[0086] In the embodiments of the present application, such as Figure 6As shown in the figure, the metadata backend storage medium, i.e., the storage disk in the embodiments of the present application, is globally shared among each node. Specifically, the storage disk includes multiple logical units (vLuns) divided by 512M in size, and each logical unit belongs to one node. Each logical unit corresponds to multiple driver processes for finally writing the metadata in the cache to the backend storage. Secondly, multiple task unit (fibre) processes are established between the metadata driver process and the logical unit for flushing the cached data to the logical unit. The driver process and the task unit are in one-to-one correspondence and are bound to the same central processing unit (CPU) core. In this way, it is ensured that the disk flushing of each metadata cache from the task unit to the driver process and finally writing to the backend storage are all processed within the same core.

[0087] As an optional method, the processing logic for reading data from the storage disk in the embodiments of the present application is similar to the above-mentioned processing logic for writing to the storage disk. Specifically, a disk reading request corresponding to the home node can be generated, the disk reading request is sent to the task unit corresponding to the home node, and the logical unit corresponding to the home node is determined. Based on the driver program corresponding to the logical unit, the first metadata is read from the disk address of the storage disk corresponding to the logical unit into the cache space.

[0088] Optionally, the method of this embodiment further includes: in the case of detecting a failure of a target node in the metadata server, allocating the logical unit corresponding to the target node to other nodes except the target node, so that other nodes except the target node execute the processing request based on the logical unit corresponding to the target node.

[0089] Exemplarily, when a node fails, the load balancing of metadata access is achieved through the remapping of the logical unit and the node. For example, after node1 fails, the logical unit associated with node1 is reallocated to other nodes. When the client accesses the file belonging to node1, it is found that through the index node (inode) of the file, the corresponding logical unit ID is obtained, and then the remapped node is obtained according to the logical unit ID, and the metadata is obtained through the corresponding node.

[0090] To illustrate the specific implementation process of this embodiment, the following specific application examples are given, such as Figure 7 as shown in the figure, but not limited thereto:

[0091] When the client issues a metadata read / write request through a certain node, the node obtains the file's home node by querying the directory entry. If the home node is the same as the requesting node, the read / write request is directly issued through the requesting node; if the home node is different from the requesting node, for the write request, the write data is directly flushed to the home node's cache through the inter-node network; for the read request, the remote cache is directly accessed through the inter-node network. If the cache hits, the cache is directly read and the result is returned to the requesting node; when the cache misses, the home node is triggered to access the backend storage to read the data into memory and return the data to the requesting node.

[0092] It should be noted that the node directory list in the embodiments of the present application records the home node of the file. When the file is just written, the metadata can be initially load-balanced among the nodes through the directory entry. Subsequently, based on the metadata access load of each node, secondary load balancing can be performed. The metadata cache is globally shared within each node, achieving the consistency of cached data among nodes and reducing the overhead of frequent data synchronization and application / release of locks among nodes to ensure cache consistency among nodes. The mechanism for flushing the metadata cache to the backend storage medium ensures that each I / O operation from the start of flushing to the final landing on the backend storage is processed on the same CPU core, avoiding the overhead caused by CPU core switching. At the same time, the global sharing of the backend storage by all nodes enables the logical unit to which the data belongs to be transferred to other nodes in real time when a node fails, continuing to provide metadata services and ensuring the continuity of the service.

[0093] Compared with the related art, in response to the processing request of the client, this embodiment needs to determine whether the requesting node and the home node of the target data file to be processed are the same, and parse the specific content of the processing request in the case of non-identity. If it is parsed that the processing request is to read the first metadata in the target data file, the cache space of the home node in the preset memory pool can be directly searched based on the requesting node, and based on the search result, the requesting node can be controlled to directly read the first metadata in the cache space of the home node. This enables this embodiment to directly read the data in the storage space of other nodes based on the preset memory pool without data transmission between nodes during the cross-node reading process. Furthermore, this embodiment can ensure the consistency of metadata among multiple nodes, reduce the latency of metadata operations, and based on the preset memory pool, this embodiment can also handle more concurrent requests, improve the overall throughput of the metadata server, ensure the overall high availability of the metadata server, and also improve the system performance, consistency, and reliability.

[0094] The embodiments of the present application also provide a data processing device, as Figure 8 shown. The device includes: a determination module 31, a parsing module 32, a search module 33, and a reading module 34.

[0095] A determination module 31, configured to determine, in response to a processing request of a client for a target data file, a request node in a metadata server for receiving the processing request and an ownership node corresponding to the target data file, where the ownership node is the node used when creating the target data file;

[0096] A parsing module 32, configured to parse the processing request if it is determined that the request node is different from the ownership node;

[0097] A searching module 33, configured to, when it is parsed that the processing request is to read first metadata in the target data file, search for the first metadata in a cache space of the ownership node in a preset memory pool, where the preset memory pool is composed of cache spaces of all nodes in the metadata server;

[0098] A reading module 34, configured to control the request node to read the first metadata from the cache space according to the search result.

[0099] In some examples of this embodiment, the reading module 34 is specifically configured to, when the first metadata is found, control the request node to read the first metadata from the cache space; when the first metadata is not found, read the first metadata from a storage disk of the metadata server into the cache space through the ownership node, and control the request node to read the first metadata in the cache space.

[0100] In some examples of this embodiment, the reading module 34 is specifically further configured to, when the first metadata is not found, control the request node to send the processing request to the ownership node; control the ownership node to read the first metadata from the storage disk into the cache space based on the processing request.

[0101] In some examples of this embodiment, the reading module 34 is specifically further configured to generate a disk reading request corresponding to the ownership node, send the disk reading request to a task unit corresponding to the ownership node, and determine a logical unit corresponding to the ownership node; based on a driver corresponding to the logical unit, read the first metadata from a disk address of a storage disk corresponding to the logical unit into the cache space.

[0102] In some examples of this embodiment, the parsing module 32 is further configured to, when it is parsed that the processing request is to perform a write process on second metadata in the target data file, control the request node to send the processing request to the ownership node; control the request node to perform a write process on the target metadata based on the processing request, and send the obtained write process result to the target metadata.

[0103] In some examples of this embodiment, the parsing module 32 is further specifically configured to write the second metadata into the cache space, generate a write processing result and send it to the target metadata; generate a write request corresponding to the home node, and perform a write process on the second metadata in the storage disk based on the write request.

[0104] In some examples of this embodiment, the parsing module 32 is further specifically configured to determine the target storage address of the second metadata in the storage disk, and the target logical unit corresponding to the target storage address; perform a write process on the second metadata in the storage disk based on the write request and the target logical unit.

[0105] In some examples of this embodiment, the parsing module 32 is further specifically configured to determine the target task unit group corresponding to the target logical unit, and the target task unit in the target task unit group for executing the write request, and send the write request to the target task unit; determine the target drive process corresponding to the target task unit, and perform a write process on the second metadata in the storage disk based on the target drive process.

[0106] In some examples of this embodiment, the determination module 31 is specifically configured to determine the request node and the file name of the target data file; determine the home node from the node directory list according to the file name, and the mapping relationship between the file name of the metadata file stored in the metadata server and the home node is recorded in the node directory list.

[0107] In some examples of this embodiment, the reading module 34 is further configured to record the target number of times for the request node to read the first metadata based on the home node in the node directory list; in the case where it is determined that the target number of times is greater than or equal to the preset number threshold, determine the request node as the home node corresponding to the target data file in the node directory list.

[0108] In some examples of this embodiment, the parsing module 32 is further specifically configured to determine the number of files corresponding to each node in the node directory list; in the case where it is determined that the number of target nodes is greater than or equal to the file number threshold, select the files with the processing times less than or equal to the preset processing threshold from the target nodes as the files to be transferred, and the target nodes are any nodes in the node directory list; transfer the mapping relationship between the files to be transferred and the target nodes in the node directory list to the mapping relationship between the files to be transferred and the other nodes except the target nodes.

[0109] In some examples of this embodiment, the reading module 34 is further configured to, in the case where a target node in the metadata server is detected to have a failure, allocate the logical unit corresponding to the target node to other nodes except the target node, so that the other nodes except the target node execute the processing request based on the logical unit corresponding to the target node.

[0110] It should be noted that for other corresponding descriptions of each functional unit involved in the data processing device provided in this embodiment, reference can be made to Figure 1 the corresponding description in [reference document], which will not be elaborated here.

[0111] Based on the method as shown in Figure 1 above, correspondingly, this embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as shown in Figure 1 above is implemented.

[0112] Based on the method as shown in Figure 1 above, correspondingly, this embodiment also provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method as shown in Figure 1 above is implemented.

[0113] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, and this software product can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0114] Based on the method as shown in Figure 1 above, and Figure 8 the virtual device embodiment as shown in [reference document], in order to achieve the above object, this embodiment of the application also provides an electronic device, such as a personal computer or a server, and this device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as shown in Figure 1 above.

[0115] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc., and an optional user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.

[0116] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not limit the physical device, and it may include more or fewer components, or combine some components, or have different component arrangements.

[0117] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical devices, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the related art, this embodiment responds to the processing request of the client, and needs to determine whether the request node and the belonging node of the target data file to be processed are the same. If they are not the same, the specific content of the processing request is parsed. If it is parsed that the processing request is to read the first metadata in the target data file, then it is possible to directly search in the cache space of the belonging node in the preset memory pool based on the request node to check if the first metadata exists, and based on the search result, control the request node to directly read the first metadata in the cache space of the belonging node. This enables this embodiment to not require data transmission between nodes during cross-node reading, and can directly read the data stored in the storage space of other nodes based on the preset memory pool. Furthermore, this embodiment can ensure the consistency of metadata between multiple nodes, reduce the latency of metadata operations, and based on the preset memory pool, this embodiment can also process more concurrent requests, improve the overall throughput of the metadata server, ensure the overall high availability of the metadata server, and also improve system performance, consistency, and reliability.

[0119] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0120] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A data processing method, characterized in that, including: In response to a processing request for a target data file from a client, determine a request node in the metadata server for receiving the processing request and an ownership node corresponding to the target data file, where the ownership node is the node used when creating the target data file; If it is determined that the request node is different from the ownership node, then parse the processing request; When it is parsed that the processing request is to read first metadata in the target data file, search for the first metadata in the cache space of the ownership node in a preset memory pool, where the preset memory pool is composed of the cache spaces of all nodes in the metadata server; Control the request node to read the first metadata from the cache space according to the search result.

2. The method according to claim 1, wherein The controlling the request node to read the first metadata from the cache space according to the search result includes: When the first metadata is found, control the request node to read the first metadata from the cache space; When the first metadata is not found, read the first metadata from the storage disk of the metadata server into the cache space through the ownership node, and control the request node to read the first metadata in the cache space.

3. The method according to claim 2, wherein When the first metadata is not found, reading the first metadata from the storage disk of the metadata server into the cache space through the ownership node includes: When the first metadata is not found, control the request node to send the processing request to the ownership node; Control the ownership node to read the first metadata from the storage disk into the cache space based on the processing request.

4. The method according to claim 3, characterized in that, The controlling the ownership node to read the first metadata from the storage disk into the cache space based on the processing request includes: Generate a disk read request corresponding to the ownership node, send the disk read request to the task unit corresponding to the ownership node, and determine the logical unit corresponding to the ownership node; Based on the driver corresponding to the logical unit, read the first metadata from the disk address of the storage disk corresponding to the logical unit into the cache space.

5. The method according to claim 1, wherein After the step of if it is determined that the request node is different from the ownership node, then parse the processing request, the method further includes: When it is parsed that the processing request is to write second metadata to the target data file, control the request node to send the processing request to the ownership node; Control the request node to perform a write process on target metadata based on the processing request, and send the obtained write process result to the target metadata.

6. The method according to claim 5, wherein The controlling the request node to perform a write process on the target metadata based on the processing request, and send the obtained write process result to the target metadata includes: Perform a write process on the second metadata in the cache space, generate the write process result and send it to the target metadata; Generate a write request corresponding to the attribution node, and perform a write process on the second metadata in the storage disk based on the write request.

7. The method according to claim 6, characterized in that, Performing a write process on the second metadata in the storage disk based on the write request includes: Determine the target storage address of the second metadata in the storage disk and the target logical unit corresponding to the target storage address; Based on the write request and the target logical unit, perform a write process on the second metadata in the storage disk.

8. The method according to claim 7, characterized in that The performing a write process on the second metadata in the storage disk based on the write request and the target logical unit includes: Determine the target task unit group corresponding to the target logical unit and the target task unit in the target task unit group for executing the write request, and send the write request to the target task unit; Determine the target drive process corresponding to the target task unit, and perform a write process on the second metadata in the storage disk based on the target drive process.

9. The method according to claim 5, characterized in that, The determining the request node in the metadata server for receiving the processing request and the attribution node corresponding to the target data file includes: Determine the request node and the file name of the target data file; Determine the attribution node from the node directory list according to the file name, and the mapping relationship between the file name of the metadata file stored in the metadata server and the attribution node is recorded in the node directory list.

10. The method according to claim 9, wherein After controlling the request node to read the first metadata from the cache space according to the search result, the method further includes: Record the target number of times that the request node reads the first metadata based on the attribution node in the node directory list; When it is determined that the target number of times is greater than or equal to the preset number threshold, determine the request node as the attribution node corresponding to the target data file in the node directory list.

11. The method according to claim 9, characterized in that, After controlling the request node to perform a write process on the target metadata according to the processing request and sending the obtained write process result to the target metadata, the method further includes: Determine the number of files corresponding to each node in the node directory list; When it is determined that the number of target nodes is greater than or equal to the file number threshold, select the files with the processing times less than or equal to the preset processing threshold from the target nodes as the files to be transferred, and the target nodes are any nodes in the node directory list; Transfer the mapping relationship between the file to be transferred and the target node in the node directory list to the mapping relationship between the file to be transferred and other nodes except the target node.

12. The method according to claim 1, characterized in that, The method further includes: When it is detected that a target node in the metadata server fails, allocate the logical unit corresponding to the target node to other nodes except the target node, so that other nodes except the target node execute the processing request based on the logical unit corresponding to the target node.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 12.

14. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product, having a computer program stored thereon, characterized in that, When the computer program product is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Metadata request processing method and device, equipment and medium

    CN113127420A

  • Metadata query method and device based on distributed file system and storage medium

    CN114116613A