Data processing method and device
By allowing any node to receive and process access requests in a distributed storage system, and perform data queries and transmissions within the system, the hot issues of metadata access are solved, system performance and efficiency are improved, and data up-to-dateness and consistency are ensured.
Patent Information
- Application Number
- CN202311866573.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In distributed file systems, the metadata access hotspot problem leads to performance degradation, which is difficult to effectively solve in the existing technology.
By allowing any node to receive and process access requests for objects in the directory in a distributed storage system, direct access to the home node is avoided, and data is transmitted from nodes with objects stored to nodes that process requests are achieved by using query and acquisition mechanisms to realize load balancing and data processing.
Effectively avoid or eliminate access hotspots, improve the performance of metadata operations and system processing efficiency, reduce the load pressure of the home node, and ensure the latest and consistency of data.
Smart Images

Figure CN120234313A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data storage, and in particular, to a data processing method and apparatus. Background Art
[0002] In addition to providing file storage services, a distributed file system also provides a metadata service (MDS). Among them, metadata is data that describes a file or directory, and the metadata service includes storing metadata and supporting metadata access. For example, reading metadata, modifying metadata, etc.
[0003] Distributed file systems can be divided into fully symmetric distributed file systems and asymmetric distributed file systems. In an asymmetric distributed file system, some storage nodes support metadata access, which easily leads to a metadata access bottleneck. Each storage node in a symmetric distributed file system supports metadata access, but metadata access in a symmetric distributed file system is based on directory ownership access. Specifically, the metadata of an object in a directory is stored in the home node of the directory, and an access request for the metadata of an object in the directory is only processed by the home node. When the access frequency of the metadata of an object in a certain directory is high, the home node of the directory experiences an access hotspot, and the operating performance of the metadata deteriorates. Summary of the Invention
[0004] This application provides a data processing method and apparatus in a distributed storage system, which can avoid or eliminate access hotspots.
[0005] In a first aspect, a data processing method in a distributed storage system is provided. The distributed storage system includes multiple nodes, and objects in a directory are stored in multiple nodes. An object is a sub-directory or file of the directory. The method includes: a first node among the multiple nodes receives a first access request, the first access request includes a first operation instruction and a first object instruction, the first object instruction is used to indicate a first object in the directory, and the first operation instruction is used to indicate an operation to be performed on the first object; the first node, based on the first object instruction, queries that a second node among the multiple nodes stores data of the first object; the first node obtains the data of the first object from the second node; the first node operates on the data of the first object based on the first operation instruction.
[0006] Wherein, the distributed storage system may be composed of multiple storage nodes. In this case, the above-mentioned multiple nodes are specifically storage nodes in the distributed storage system.
[0007] The distributed storage system can be composed of multiple storage nodes and at least one computing node. In this case, the above-mentioned multiple nodes can specifically be the storage nodes in the distributed storage system, or the above-mentioned multiple nodes can include storage nodes and computing nodes (that is, some of the above-mentioned multiple nodes are storage nodes and some are computing nodes).
[0008] The data of an object can be the metadata of the object. When the object is a file, the data of the object can be the metadata of the object, can be the content of the object, or can include both the metadata and the content of the object at the same time.
[0009] In addition, the processing of an access request can be divided into semantic processing and data processing. Among them, semantic processing can also be called semantic operation, which refers to obtaining the information that can be executed by a node based on the information of the upper-layer application in the access request to facilitate data processing of the data of the object. Data processing, also known as data operation, refers to the operation of the data of the object by the node, such as modification operation, reading operation, creation operation, etc. In this method, querying which nodes store the data of the object belongs to the semantic processing of the access request. Obtaining the data of the object and operating on the data of the object belong to the data processing of the access request. Through this method, the same node can complete the semantic processing and data processing of the access request, and this node can be a node other than the node storing the data of the object.
[0010] In the method provided in this application, the node that receives and processes the access request of the object and the home node of the directory to which the object belongs can be different nodes. Among them, any node among the multiple nodes can receive one or more access requests of the same object and process the received access request. That is to say, through this method, nodes other than the home node of the directory can receive and process access requests for objects in the directory (such as access requests for the metadata of the object).
[0011] More specifically, any node other than the home node of the directory can receive an access request for an object in the directory. In other words, an access request for an object can be sent to any node other than the home node of the directory, and it is not necessary to send the access request for the object in the directory to the home node of the directory. For example, it can be set that the directory includes object F1, the home node of the directory is node G1, and the above-mentioned multiple nodes also include nodes G2, G3, etc. Nodes G2, G3, etc. can receive the access request for object F1, and it is not necessary for node G1 to receive the access request for object F1. In this way, when accessing the object in the directory frequently, nodes G2 or G3, etc. can receive the access request for object F1, reducing the pressure on the home node of the directory to receive access requests.
[0012] The directory may further include multiple objects such as object F2, object F3, etc. Among them, access requests for different objects can be received by different nodes respectively, so that the pressure of receiving access requests can be shared among different nodes.
[0013] The node that receives the access request can process the access request. As described above, the node that receives the access request can be any node, not necessarily the home node of the directory. In this way, nodes other than the home node of the directory can process the access request. Returning to the above example, nodes such as node G2 or node G3 receive the access request for object F1, and then, node G2 or node G3 can process the access request for object F1. In this way, it is not necessary for the home node of the directory to process the access request for object F1, thus reducing the pressure on the home node of the directory to process access requests, and thus avoiding or eliminating the access hot spot that may be caused by the home node of the directory processing the access request for object F1.
[0014] Access requests for different objects in the directory can be received and processed by different nodes respectively, so that the pressure of processing access requests can be shared among different nodes.
[0015] In addition, as described above, any node can receive and process the access request for object F1. Therefore, the node that receives the access request for object F1 may not store the data of object F1. It can be set that node G2 receives the access request for object F1, and node G2 does not store the data of object F1. In the method provided in this application, node G2 can query the node that stores the data of object F1 and obtain the data of object F1 from that node. Thus, node G2 can process the access request for object F1 based on the data of object F1.
[0016] In short, through the method provided in this application, any node (such as a node other than the home node of the directory) in the distributed storage system can receive and process the object access request in the directory. Thus, when accessing an object in the directory, there is no need for home-based access based on the directory, thus avoiding or eliminating problems such as access hot spots caused by home-based access.
[0017] In a possible implementation, a third node among multiple nodes is used to record the nodes that store the data of the objects in the directory; the first node queries that a second node among the multiple nodes stores the data of the first object based on the first object indication, including: the first node queries, based on the first object indication, in the third node that the second node stores the data of the first object; wherein, the third node is used to record that the first node stores the data of the first object after the first node obtains the data of the first object.
[0018] In this implementation, when a node obtains the first data, it is necessary to store the data of the first object. Here, the storage can be caching (i.e., storing the data of the first object in memory). When a node obtains the first data, for example, when the first node obtains the first data from the second node, it can record in the third node that the first node stores the data of the first object. In this way, when other nodes need to obtain the data of the first object, they can query in the third node that the first node stores the data of the first object and obtain the data of the first object from the first node. Thus, in the case where the node storing the data of the first object is not fixed, it can be ensured that the nodes requiring the data of the first object can obtain the data of the first object. In addition, whenever an access request for the first object is processed, it is not necessary to obtain the data of the first object from the home node of the first object, which reduces the pressure on the home node of the first object to transmit the data of the first object.
[0019] In a possible implementation, the method further includes: the first node receives a second access request, the second access request includes a second operation indication and a second object indication, the second object indication is used to indicate a second object in a directory, and the second operation indication is used to indicate an operation to be performed on the second object; when the load of the first node is greater than or equal to the load of the third node, the first node sends the second operation indication and the second object indication to the third node; wherein, the third node is used to obtain the data of the second object based on the second object indication and operate on the data of the second object based on the second operation indication.
[0020] Among them, the first node is the node that first receives the second access request among multiple nodes, that is, the first node is the access node of the second access request. In this implementation, when the load of the access node of the second access request is not less than the load of the third node, the access node sends the second access request to the third node, so that the first access request is preferentially processed by the third node. Among them, the third node can complete the query of the node storing the data of the second object locally, which improves the query efficiency and thus can improve the efficiency of processing the second access request.
[0021] In a possible implementation, the first node obtains the data of the first object from the second node, including: when the data of the first object stored in the second node is the latest data of the first object, the first node obtains the data of the first object from the second node.
[0022] There may be multiple nodes storing the data of the first object, and the data of the first object stored in some nodes may be the historical data of the first object rather than the latest data of the first object. Generally, an access request for an object is a request to operate on the latest data of the object. In this implementation manner, when the data of the second node is the latest data of the first object, the first node obtains the data of the first object from the second node, so that it can be ensured that the data obtained by the first node is the latest data of the first object, and further ensure that the first node processes the first access request based on the latest data of the first object, so that the first access request can be correctly processed.
[0023] In a possible implementation manner, the multiple nodes further include a fourth node for persistently storing the data of the first object, and the method further includes: the first node sends the data of the first object after the operation to the fourth node, so that the fourth node persistently stores the data of the first object after the operation; wherein, the data of the first object after the operation is the data obtained by the first node operating on the data of the first object based on the first operation instruction.
[0024] In this implementation manner, when the first node completes the operation on the data of the first object, it can send the data after the operation to the fourth node, and the fourth node completes the persistent storage of the data after the operation. The fourth node is the node that performs the persistent storage operation on the data of the first object, and is not necessarily the home node of the first object. The fourth node can store the data of the first object after the operation in the local hard disk, or can store the data of the first object after the operation in the hard disk of other nodes. Exemplarily, when the available capacity of the local hard disk of the fourth node is sufficient, the fourth node can store the data of the first object in the local hard disk. When the available capacity of the local hard disk of the fourth node is insufficient, the fourth node can store the data of the first object in the hard disk of other nodes. Among them, when the fourth node stores the data of the first object after the operation in the local hard disk, the fourth node is the home node of the first object.
[0025] In a possible implementation manner, the data of the first object is located in the first page; the first node queries that the second node among the multiple nodes stores the data of the first object based on the first object indication, including: the first node queries that the second node includes the first page based on the first object indication; the first node obtains the data of the first object from the second node, including: the first node obtains the first page from the second node to obtain the data of the first object.
[0026] In this implementation manner, storing the data of an object in terms of pages facilitates the storage and management of the data of the object. And when obtaining the data of an object, query and obtain the page including the data of the object, so that the data of the object can be obtained conveniently and quickly.
[0027] In a possible implementation, the first page may be associated with the memory addresses of different nodes among multiple nodes. The memory address associated with the first page is used to store the data in the first page. The first node queries, based on the indication of the first object, that the second node includes the first page, including: The first node queries, based on the indication of the first object, that the first page is associated with the first memory address of the second node. The first node obtains the first page from the second node, including: The first node obtains the data in the first page from the first memory address. The first node associates the first page with the second memory address of the first node and stores the data in the first page into the second memory address.
[0028] Among them, the page is used to logically store the data of the object, and the data in the page is specifically stored in the memory address associated with the page. The page can be associated with the memory addresses of different nodes, so that multiple nodes can store the data of the object simultaneously, thus realizing the storage of the data of the object in multiple nodes. Among them, the fact that the memory of the node stores the data of the object can be referred to as the node caching the data of the object.
[0029] In this implementation, the first node queries that the first page is associated with the memory address of the second node and can obtain the data in the first page from the memory address associated with the first page in the second node, so as to obtain the data of the first object. The first node can associate the first page with the memory address of the first node and store the data in the first page into the memory address associated with the first page in the first node, so as to store the data of the first object in the memory of the first node, enabling the processor of the first node to operate on the data of the first object.
[0030] In a possible implementation, the method further includes: The first node receives a third access request. The third access request includes a third operation indication and a third object indication. The third object is used to indicate the third object in the directory, and the third operation indication is used to indicate the operation that needs to be performed on the third object. The first node queries, based on the third object indication, that the fifth node among multiple nodes stores the data of the third object. The first node sends the third operation indication to the fifth node, so that the fifth node operates on the data of the third object based on the third operation indication.
[0031] In this implementation, when the access node of the access request and the node storing the data of the access object of the access request are not the same node, the access node of the access request can forward the access request to the node storing the data of the access object, so that the node storing the data of the access object processes the access request, thus saving the operation of transmitting the data of the access object between different nodes and improving the processing efficiency of the access request.
[0032] In addition, in this implementation manner, the access node of the access request and the node storing the data of the access object of the access request can jointly process the access request, so that the load generated by processing the access request is borne by multiple nodes, avoiding the occurrence of access hotspots.
[0033] Among them, the access node of the access request can perform semantic processing on the access request. The node storing the data of the access object of the access request can perform data processing on the access request. Thus, the load of semantic processing of the access request can be borne by the access node of the access request, and the load of data processing of the access request can be borne by the node storing the data of the access object.
[0034] Among them, the semantic processing of the access request further includes: identifying the storage location of the data of the access object. Among them, the storage location may include: the identifier of the page where the access object is located. When the storage location of the data of the access object is identified, the storage location and the operation instruction can be sent to the node for data processing. Then, the node for data processing can directly obtain the data of the access object from the storage location, so that operations such as object identification and querying the storage location of the data of the object do not need to be performed anymore, and only operations of obtaining data and operating on the data (i.e., data processing) need to be performed.
[0035] Thus, the node for semantic processing does not need to perform data processing, and thus does not need to bear the load of data processing. The node for data processing does not need to perform semantic processing, and thus does not need to bear the load of semantic processing.
[0036] In a possible implementation manner, the first operation instruction is used to indicate the creation of a first object; the first node, based on the first object instruction, queries that the second node among multiple nodes stores the data of the first object, including: the first node, based on the first object instruction, queries that the second node includes a second page, where the second page is used to store the data of the first object; the first node obtains the data of the first object from the second node, including: the first node obtains the second page from the second node; the first node, based on the first operation instruction, operates on the data of the first object, including: generating the data of the first object and storing the data of the first object in the second page.
[0037] Among them, the first node's querying the second page for storing the data of the first object belongs to the semantic processing of the first access request. Obtaining the page, generating the data of the first object, and storing the generated data of the first object in the second page belong to the data processing of the first access request.
[0038] In a possible implementation, the first access request is sent by the sixth node in the distributed storage system, and the first operation instruction is used to indicate that a read operation needs to be performed on the first object; based on the first operation instruction, the first node operates on the data of the first object, including: the first node sends the data of the first object to the sixth node.
[0039] Among them, sending the data of the first object to the sixth node belongs to the data processing of the first access request.
[0040] In a possible implementation, the first operation instruction is used to indicate that a modification operation needs to be performed on the first object; based on the first operation instruction, the first node operates on the data of the first object, including: the first node modifies the data of the first object based on the first operation instruction.
[0041] Among them, modifying the data of the first object belongs to the data processing of the first access request.
[0042] In a second aspect, a data processing device is provided. The distributed storage system includes multiple nodes, and this device is configured in the first node among the multiple nodes; among them, the objects in the directory are stored in multiple nodes, and the object is a sub-directory or a file of the directory. The device includes: a receiving unit, configured to receive a first access request, where the first access request includes a first operation instruction and a first object instruction, the first object instruction is used to indicate the first object in the directory, and the first operation instruction is used to indicate the operation that needs to be performed on the first object; a query unit, configured to query, based on the first object instruction, that the second node among the multiple nodes stores the data of the first object; an obtaining unit, configured to obtain the data of the first object from the second node; an operation unit, configured to operate on the data of the first object based on the first operation instruction.
[0043] In a possible implementation, the third node among the multiple nodes is used to record the nodes storing the data of the objects in the directory; the query unit is configured to: query, based on the first object instruction, that the second node stores the data of the first object in the third node; among them, the third node is used to record that the first node stores the data of the first object after the obtaining unit obtains the data of the first object.
[0044] In a possible implementation, the device further includes: a sending unit; among them, the receiving unit is further configured to: receive a second access request, where the second access request includes a second operation instruction and a second object instruction, the second object instruction is used to indicate the second object in the directory, and the second operation instruction is used to indicate the operation that needs to be performed on the second object; the sending unit is configured to: when the load of the first node is greater than or equal to the load of the third node, send the second operation instruction and the second object instruction to the third node; among them, the third node is used to obtain the data of the second object based on the second object instruction, and operate on the data of the second object based on the second operation instruction.
[0045] In a possible implementation, the obtaining unit is configured to: when the data of the first object stored in the second node is the latest data of the first object, obtain the data of the first object from the second node.
[0046] In a possible implementation, the multiple nodes further include a fourth node for persistently storing the data of the first object, and the apparatus further includes: a sending unit, configured to send the data of the first object after the operation to the fourth node, so that the fourth node persistently stores the data of the first object after the operation; wherein, the data of the first object after the operation is the data obtained by the operation unit operating on the data of the first object based on the first operation instruction.
[0047] In a possible implementation, the data of the first object is located in the first page; the query unit is configured to: based on the indication of the first object, query that the second node includes the first page; the obtaining unit is configured to: obtain the first page from the second node to obtain the data of the first object.
[0048] In a possible implementation, the first page can be associated with the memory addresses of different nodes among the multiple nodes, and the memory address associated with the first page is used to store the data in the first page; the query unit is configured to: based on the indication of the first object, query that the first page is associated with the first memory address of the second node; the obtaining unit is configured to: obtain the data in the first page from the first memory address; the obtaining unit is further configured to: associate the first page with the second memory address of the first node and store the data in the first page into the second memory address.
[0049] In a possible implementation, the apparatus further includes: a sending unit; wherein, the receiving unit is configured to: receive a third access request, the third access request includes a third operation instruction and a third object indication, the third object is used to indicate a third object in the directory, and the third operation instruction is used to indicate an operation to be performed on the third object; the query unit is configured to: based on the third object indication, query that the fifth node among the multiple nodes stores the data of the third object; the sending unit is configured to: send the third operation instruction to the fifth node, so that the fifth node operates on the data of the third object based on the third operation instruction.
[0050] In a third aspect, a data processing apparatus is provided, including: a memory, configured to store an executable program; a processor, configured to execute the method provided in the first aspect by running the executable program.
[0051] In a fourth aspect, a computer-readable storage medium is provided, including: computer program instructions, when the computer program instructions are executed by a computer device, the computer device executes the method provided in the first aspect.
[0052] In a fifth aspect, there is provided a computer program product comprising instructions, characterized in that when the instructions are run on a computer device, the computer device is caused to execute the method provided in the first aspect.
[0053] The beneficial effects of the second to fifth aspects can be referred to the introduction of the beneficial effects of the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1A is a schematic diagram of a distributed storage system provided by an embodiment of the present application;
[0055] Figure 1B is a schematic diagram of a distributed storage system provided by an embodiment of the present application;
[0056] Figure 2 is a schematic diagram of a distributed storage system provided by an embodiment of the present application;
[0057] Figure 3 is a schematic diagram of the structure of a node in a distributed storage system provided by an embodiment of the present application;
[0058] Figure 4 is a schematic diagram of the structure of a node in a distributed storage system provided by an embodiment of the present application;
[0059] Figure 5 is a schematic diagram of a data storage method provided by an embodiment of the present application;
[0060] Figure 6 is a schematic diagram of a data processing method provided by an embodiment of the present application;
[0061] Figure 7 is a schematic diagram of a cache coherence protocol provided by an embodiment of the present application;
[0062] Figure 8 is a schematic diagram of a cache coherence protocol provided by an embodiment of the present application;
[0063] Figure 9 is a schematic diagram of a load balancing method provided by an embodiment of the present application;
[0064] Figure 10 is a schematic diagram of a load balancing method provided by an embodiment of the present application;
[0065] Figure 11 is a schematic diagram of the functions of a policy control node provided by an embodiment of the present application;
[0066] Figure 12 is a schematic diagram of a load balancing policy provided by an embodiment of the present application;
[0067] Figure 13 It is a schematic diagram of a data processing method provided by an embodiment of the present application;
[0068] Figure 14 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0069] Figure 15 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application. Specific embodiments
[0070] Next, the solution provided by the embodiment of the present application will be described in conjunction with the accompanying drawings. Among them, in the embodiment of the present application, "a plurality of" means two or more than two. "First", "second", etc. are only used to distinguish similar objects and do not have to be used to describe a specific order or the number of objects.
[0071] To facilitate the understanding of the solution provided by the embodiment of the present application, before introducing the solution provided by the embodiment of the present application, some technical terms that may be involved in the embodiment of the present application will be introduced first.
[0072] Distributed storage system: A network storage architecture including multiple storage nodes. In some embodiments, the distributed storage system may further include at least one computing node. Storage nodes and computing nodes can be collectively referred to as nodes. Among them, storage nodes are used to store data, and the data can be dispersed and stored in different storage nodes to achieve high-reliability storage and high-performance access of data. Computing nodes run client programs and can access the data stored by storage nodes. Among them, a distributed file system is a common distributed storage system. Different storage nodes in the distributed storage system can be linked through a network, and storage nodes and computing nodes can be linked through a network.
[0073] Objects in a directory: A general term for files and subdirectories in the directory. That is, the objects in the directory can refer to the files in the directory or the subdirectories in the directory.
[0074] Access request: Information used to request an operation on an object, such as creating an object, modifying the content or metadata of an object, reading the content or metadata of an object, etc.
[0075] Access request access node: The node that first receives the access request, and this node is the node in the distributed storage system where it is located for processing access requests. Among them, the access request access node is not necessarily the node that actually processes the access request. In other words, the access node can process the access request or forward the access request to other nodes in the distributed storage system for processing access requests, so that other nodes process the access request.
[0076] The home node of an object: It refers to the storage node that persistently stores the data of the object in a distributed storage system. That is, the hard disk of the home node of the object stores the data of the object. Usually, the home nodes of the objects in the same directory are the same node. Therefore, the home node of an object can also be called the home node of the directory to which the object belongs. The home nodes of different directories can be the same node or different nodes.
[0077] The cache node of an object: It refers to the node that stores the data of the object in memory. That is, if the memory of a node stores the data of the object, then this node is called the cache node of the object.
[0078] The management node of an object: Also known as the master node of the object, it is a node in a distributed storage system used to record the nodes that store the data of the object. Among them, the management node of a directory records the cache node and the home node of the object. Usually, the management nodes of the objects in the same directory are the same node. Therefore, the management node of an object can also be called the management node of the directory where the object is located. In addition, the management node of an object and the home node of the object can be the same node or different nodes. The management nodes of different directories can be the same node or different nodes.
[0079] The latest data of an object: It refers to the data of the latest version of the object, that is, the data obtained after the execution of the most recent modification operation on the object.
[0080] Page: A segment of virtual address or virtual storage space used to store data. A page can also be called a data block or a logical block, which is used to implement the logical storage of data. A page is associated with a physical storage address (such as a memory address), and the physical storage address associated with the page is used to physically store the data in the page. That is, the data in the page specifically refers to the data stored in the physical storage address associated with the page.
[0081] Persistent storage: Also known as disk flushing, it refers to storing the data in memory to the hard disk to achieve the permanent preservation of the data.
[0082] The metadata of an object: A kind of data used to describe an object. Among them, the metadata of a file includes the name of the file, the identifier (ID) of the file, the most recent modification time, the size, the directory where it is located, etc. The metadata of a directory includes the name of the directory, the ID of the directory, the creation time, the ID of the upper-level directory (i.e., the parent directory) where it is located, etc.
[0083] Directory entry (dentry): An object in a directory corresponds to a directory entry of the directory, which is used to record metadata such as the name of the object and the ID of the inode of the object. Among them, the directory entries of all the objects in the directory constitute the dentry table of the directory.
[0084] Inode: That is, an index node, which is a data structure used to record the metadata of an object. The inode of an object includes the inode ID and other metadata of the object, etc. Among them, the inode including the metadata of the object can be called the inode of the object. The inodes of all objects in a directory form an inode table. The inode table is a horizontal table, and one row in it records the inode of one object.
[0085] Index structure: It is a structure that records the mapping relationship between the identifier of an object and the storage location of the data of the object. The index structure takes the identifier of the object as input, and based on the identifier of the object, the storage location of the data of the object can be located. Common index structures include hash index structure, B tree, B+ tree, B-link tree, etc. The index structure usually has multiple layers, and each layer includes at least one page node. A page node represents a page. In a tree-shaped index structure, the page nodes in the bottom layer are called leaf nodes, the page nodes in the layer above the leaf nodes can be called intermediate nodes, and the page nodes at the top layer are called root nodes.
[0086] Atomicity: It means that multiple operations are either all executed successfully or all not executed, without the situation of partial execution.
[0087] Two-phase transaction: Also known as two-phase commit (2PC), it is an algorithm designed to ensure consistency when all nodes in a distributed system architecture perform transaction commits in the fields of computer networks and databases.
[0088] In one solution, to reduce the occurrence probability of access hotspots in a distributed file system, the method of splitting subdirectories is adopted. The objects in the directory of the distributed file system are divided into different subdirectories, and then the different subdirectories are stored in different storage nodes respectively. Among them, the storage node where the subdirectory is located supports access to the metadata of the objects in the subdirectory. This solution does not eliminate the access hotspot problem. For example, when the access frequency of the metadata of the objects in a certain subdirectory is relatively high, the same access hotspot will occur in the storage node where the subdirectory is located. Moreover, the subdirectory is not visible externally. When the objects in the subdirectory change, it is necessary to synchronously modify the dentry table of the subdirectory and the dentry table of the directory (that is, the directory from which the subdirectory is split). Since the dentry table of the subdirectory and the dentry table of the directory are usually located in different storage nodes, a two-phase transaction is required to ensure atomicity, which results in a decrease in the performance of the storage node supporting metadata access (such as the performance of operations such as reading metadata and modifying metadata).
[0089] The embodiments of the present application provide a distributed storage system 100 and a data processing method. Through this method, multiple nodes in the distributed storage system 100 can provide access services for the same object, that is, the access to the object can be processed by different nodes in the distributed storage system 100. Among them, the object is a file or a sub-directory in a directory. Thus, the access hotspot problem caused by processing the access to the object in the directory by a fixed node is eliminated. Moreover, in this distributed storage system, the node that processes the access request can obtain the data to be processed by the access request (i.e., the data of the access object of the access request) and process this data. That is to say, the data to be processed by the access request can be processed by the same node. Therefore, there is no cross-node modification of metadata, that is, there is no need for a two-phase transaction to ensure atomicity, which guarantees the performance of the node providing the access service.
[0090] Next, first introduce the distributed storage system 100 provided by the embodiments of the present application. The distributed storage system 100 may include multiple nodes for processing access requests, such as node 111, node 112, node 113, node 114, etc.
[0091] In some embodiments, refer to Figure 1A , the distributed storage system 100 includes storage nodes. The nodes for processing access requests in the distributed storage system 100 are specifically the storage nodes in the distributed storage system 100. A storage node is a device or equipment with data processing and data storage capabilities, such as a server. A storage node may include a processor, a memory, and one or more hard disks. Among them, the memory can cache the data of the object, the processor can operate the data of the object, and the hard disk can persistently store the data of the object. Among them, one or more hard disks of the storage node can form a storage pool. In addition, in this embodiment, the distributed storage system 100 may include computing nodes or may not include computing nodes. Among them, when the distributed storage system 100 includes computing nodes, these computing nodes are not used to process access requests.
[0092] In some embodiments, refer to Figure 1B , the distributed storage system 100 includes computing nodes and storage nodes. Among them, the nodes for processing access requests in the distributed storage system 100 include both storage nodes and computing nodes. For example, as Figure 1B shown, node 111 and node 112 are computing nodes in the distributed storage system 100, and node 113 and node 114 are storage nodes in the distributed storage system 100.
[0093] A computing node can be a device or equipment with data processing capabilities. Among them, the computing node includes a processor and a memory. The memory can cache the data of an object, and the processor can operate on the data of the object. In some embodiments, the computing node 110 can be a server or a terminal, etc. Among them, the terminal can be a mobile phone, a tablet computer, a vehicle-mounted terminal, a laptop computer, a handheld computer, a computer, etc. The computing node runs a client program and can be used as a client in the distributed storage system 100. In some embodiments, the computing node can be a private client (DPC) of the distributed storage system 100.
[0094] Among them, in the embodiments of the present application, the memory can be a dynamic random access memory (DRAM), a non-volatile random access memory (NVRAM), etc. The NVRAM can be a non-volatile dual in-line memory module (NVDIMM), a storage class memory (SCM), a persistent memory (PMEM), etc. The hard disk can be a solid state drive (SSD).
[0095] In some embodiments, the distributed storage system 100 can be a network attached storage (NAS) system. Among them, the nodes in the distributed storage system 100 can be NAS computing nodes or NAS storage nodes. For example, node 111 and node 112 can be NAS computing nodes, and node 113 and node 114 can be NAS storage nodes. Among them, the NAS storage node can be a NAS server or a NAS array, etc.
[0096] In some embodiments, the distributed storage system 100 can be a storage system based on a storage area network (SAN). Among them, different nodes are connected through a fibre channel (FC).
[0097] Such as Figure 1A Or Figure 1B As shown, different nodes in the distributed storage system 100 can be connected through a network, for example, through a wide area network (WAN) or a local area network (LAN). Among them, such asFigure 1A or Figure 1B As shown, a node has a network card through which it can communicate with other nodes over a network. Among them, the network card is also known as a network interface card (NIC). In some embodiments, the network can be a smart network interface card (SmartNIC) with a data processing unit (DPU). In some embodiments, when two nodes are computing nodes, the network between the two nodes can be a network suitable for computing, such as Ethernet. In some embodiments, when two nodes are storage nodes, the network between the two nodes can be a network suitable for storage, such as an InfiniBand network (IB). In some embodiments, when one of the two nodes is a computing node and the other is a storage node, the two nodes can be connected through a network suitable for computing and a network suitable for storage. Among them, the network on the computing node side is a network suitable for computing, and the network on the storage node side is a network suitable for storage.
[0098] Peer-to-peer access can be achieved between different nodes in the distributed storage system 100. Peer-to-peer access between nodes includes that a node can access the data stored on other nodes, such as obtaining the data stored on other nodes. Exemplarily, as Figure 2 shown, nodes 112, 113, 114, etc. can obtain the data of object 1 stored on node 111. In some embodiments, as Figure 2 shown, a node can obtain the data of an object stored on another node through a cache coherence protocol to ensure that the obtained data is the latest data of the object. The cache coherence protocol will be introduced below and will not be elaborated here.
[0099] Nodes in the distributed storage system 100 can process access requests. Among them, an access request includes an object indication and an operation indication. The object indication is used to indicate the access object of the access request. Among them, the access object of the access request is the object to be operated on by the access request. The operation indication is used to indicate the operation to be performed on the access object. Processing an access request can include: identifying the access object of the access request based on the object indication in the access request, and identifying the node storing the data of the access object, and then, the data of the access object can be obtained from this node, and the data of the access object can be operated on based on the operation indicated by the operation indication. In some embodiments, as Figure 3 shown, a node in the distributed storage system 100 includes an access processing module. The access processing module is used to process access requests, that is, the node processes access requests through the access processing module in this node.
[0100] In some embodiments, several nodes in the distributed storage system 100 can serve as management nodes, such as Figure 3 the node 113 shown. The management node is used to record the nodes storing the data of the object, that is, the nodes where the data of the object is located. In some embodiments, the node serving as the management node is specifically a storage node in the distributed storage system 100. Among them, the management node includes a management module, and the management module is used to record the nodes storing the data of the object. That is to say, the management node records the nodes storing the data of the object through the management module. In some embodiments, the management node can also process access requests. For example, Figure 3 the node 113 shown includes both a management module and an access processing module, so that the node 113 can both record the nodes storing the data of the object and process access requests.
[0101] The nodes storing the data of the object can include the cache nodes of the object, that is, the nodes storing the data of the object in the memory. The nodes storing the data of the object can also include the home nodes of the object, that is, the nodes storing the data of the object in the hard disk.
[0102] A directory in the distributed storage system 100 corresponds to one or more management nodes. The management node corresponding to the directory can be called the management node of the directory. The management node of the directory is used to record the nodes storing the data of the objects in the directory, that is, the cache nodes and home nodes of the objects in the directory. Among them, the management node of the directory is associated with the identifier of the directory. Exemplarily, the identifier of the directory can specifically be the handle of the directory. The object indication in the access request includes the identifier of the directory and the identifier of the object, where the object is the access object of the access request, and the directory is the directory where the access object of the access request is located. Exemplarily, the identifier of the object can be the name or handle of the object. The node can query the management node of the directory based on the identifier of the directory in the object indication, and then query the node storing the data of the object (i.e., the access object of the access request) in the management node based on the identifier of the object in the object indication.
[0103] In some embodiments, a node can forward the access request it receives to other nodes, so that the other nodes process the access request. Thus, when the load of the node receiving the access request (such as the access node of the access request) is relatively heavy, the node can forward the access request to a node with a lighter load. Among them, as Figure 3As shown in the figure, the nodes in the distributed storage system 100 have a load balancing module. The load balancing module in the node is used to forward access requests to other nodes. In one example, one or more nodes in the distributed storage system 100 have a policy control module. The policy control module is used to control the load balancing policy. For example, it selects a load balancing policy from multiple load balancing policies and sends the selected load balancing policy to the load balancing modules in each node, so that the load balancing modules in each node can forward access requests based on the selected load balancing policy. Among them, the node with the policy control module can be called the policy control node. The functions of the policy control node and the load balancing policy will be specifically introduced below and will not be elaborated here.
[0104] The data of an object can refer to the metadata of the object, such as the directory entry, inode, etc. of the object. If the object is a file in a directory, the data of the object can also include the content of the file. The content of the file is the data of the file itself or the data recorded by the file. Among them, the content of the file can be stored in the data space of the file. The data space of the file is a logically continuous data space.
[0105] In some embodiments, as Figure 4 shown, the data of an object can be stored in a page. When the object is a file, the metadata of the file and the content of the object can be stored in different pages respectively. Among them, the page is located in the node. The management node that records the node storing the data of the object can specifically be the node where the page where the data of the object is located is located. The node obtaining the data of the object from other nodes can specifically be the node obtaining the page where the object is located from other nodes.
[0106] As mentioned above, the page is used for logical data storage, and the data in the page is specifically stored in the physical storage address associated with the page. The page being located in the node specifically means that the page is associated with the physical storage address of the node. The current node obtaining a page from other nodes specifically means obtaining the data of the page from the physical storage address associated with the page in other nodes, associating the page with the physical storage address in the current node, and storing the obtained page data in the physical storage address associated with the page in the current node.
[0107] In one example of this embodiment, as Figure 5 shown, the physical storage address associated with the page is specifically a memory address. That is to say, the data in the page is specifically stored in the memory of the node. Among them, the memory address refers to a section of addresses in the memory. The data being stored in the memory can also be referred to as the data being cached in the memory. The page can be associated with the memory addresses of multiple nodes in the distributed storage system 100, and the page can be distributed among multiple nodes.
[0108] Among them, the identifier of a page is globally unique in the distributed storage system 100. That is to say, in the distributed storage system 100, a unique page can be determined through the identifier of the page. The identifier of the page is associated with the identifier of the node where the page is located. In this way, the node where the page is located can be identified through the identifier of the page. When multiple nodes have the page where the object's data is located, the node obtains the page where the latest data of the object is located. That is to say, when other nodes obtain the object's data, they specifically obtain the page from the node where the page where the latest data of the object is located. Among them, the identifier of the object is associated with the identifier of the page where the object's data is located, and the page where the object's data is located can be queried through the identifier of the object.
[0109] In some embodiments, a page can be associated with the hard disk address of a certain node. Among them, the hard disk address is a section of address in the hard disk. The data in the page is persistently stored at this hard disk address. Among them, this node can be called the home node of this page, or it can also be called the home node of the data in this page. This node is specifically a storage node in the distributed storage system 100.
[0110] In some embodiments, as Figure 4 shown, a node can include a log module and a sequence number module. Among them, the log module is used to generate logs for the operations performed on the object. The log can be a write-ahead log (WAL). The sequence number module is used to generate a sequence number for the log generated by the log module. As Figure 5 shown, the log generated by the log module can be stored in the hard disk to achieve persistent storage. In addition, when the memory of the node includes NVRAM, the log can be stored in the NVRAM to achieve persistent storage. The sequence number of the log can be stored in the memory to facilitate the processor to call. The log and the sequence number of the log will be specifically introduced below and will not be elaborated here.
[0111] The distributed storage system 100 provided by the embodiments of the present application has been introduced above. Next, in combination with the distributed storage system 100, the data processing method provided by the embodiments of the present application will be introduced.
[0112] Referring to Figure 6 , it can be set that in step 601, the node 111 receives the access request A1. Among them, the node 111 can be any node among the nodes in the distributed storage system 100 used to process access requests.
[0113] In some embodiments, the node 111 can be the access node of the access request A1. Among them, in Figure 1AIn the case shown, that is, when node 111 is a storage node in the distributed storage system 100 and the node for processing access requests is a storage node, node 111 can receive access request A1 from a computing node of the distributed storage system 100 or from an external computing device of the distributed storage system 100. In Figure 1B the case shown, that is, when the distributed storage system 100 includes computing nodes and storage nodes, and both the computing nodes and the storage nodes can be used to process access requests. If node 111 is a computing node, node 111 can receive an operation issued by a user and generate access request A1 based on the operation issued by the user. That is to say, when node 111 is a computing node, the fact that node 111 receives access request A1 specifically means that node 111 receives the operation of the user and generates access request A1 based on the operation of the user. Alternatively, if node 111 is a computing node or a storage node, node 111 can receive access request A1 from an external computing device of the distributed storage system 100.
[0114] Access request A1 includes object indication B1 and operation indication C1. Among them, object indication B1 indicates the access object of access request A1. It can be set that object indication B1 indicates the object B11 in directory B, then the access object of access request A1 is object B11. In some embodiments, object indication B1 may include the identifier of directory B and the identifier of object B11. Operation indication C1 indicates the operation that access request A1 needs to perform on the access object of access request A1, that is, operation indication C1 indicates the operation that needs to be performed on object B11.
[0115] Continue to refer to Figure 6 , node 111 can query the node storing the data of object B11 in step 602.
[0116] Among them, node 111 can first query the data of object B11 locally, that is, first determine whether node 111 itself stores the data of object B11. If node 111 itself stores the data of object B11, node 111 can operate on the data of object B11 stored by node 111 itself according to operation indication C1. In some embodiments, if node 111 itself stores the data of object B11 and the data of object B11 stored by node 111 itself is the latest data of object B11, node 111 can operate on the data of object B11 stored by node 111 itself according to operation indication C1. Determine whether the data of object B11 stored by node 111 itself is the latest data. This will be specifically introduced below and will not be elaborated here.
[0117] If node 111 itself does not store the data of object B11, or the data of object B11 stored by the node itself is not the latest data of object B11, node 111 can query the node storing the data of object B11 in the management node of directory B based on object indication B1. As described above, object indication B1 includes the identifier of directory B, and the identifier of directory B is associated with the management node of directory B. Thus, node 111 can obtain the management node of directory B based on the identifier of directory B.
[0118] As Figure 6 shown, the management node of directory B can be set as node 113. Node 111 can query the node storing the data of object B11 in node 113. When the node storing the data of object B11 is queried, node 111 can obtain the data of object B11 from this node. It can be set that node 112 stores the data of object B11, then node 111 can obtain the data of object B11 from node 112.
[0119] In some embodiments, node 111 can send a query request to node 113 based on object indication B1, and the query request includes the identifier of object B11. Node 113 can query the node storing object B11 based on the identifier of object B11. When node 113 queries that node 112 stores the data of object B11, in step 603, it can send a query result to node 111, and the query result includes the identifier of node 112. Exemplarily, the identifier of node 112 can be the network address of node 112, such as the Internet Protocol (IP) address. Node 111 can obtain the data of object B11 from node 112 based on this query result. Specifically, it can identify node 112 based on the identifier of the node in this query result, and obtain the data of object B11 from node 112.
[0120] As described above, the cache node of object B11 and the home node of object B11 belong to the nodes storing the data of object B11. If the nodes storing the data of object B11 queried by node 111 in node 113 include the cache node of object B11 and the home node of object B11. Then node 111 obtains the data of object B11 from the cache node of the cache node of object B11.
[0121] Since the cache node of object B11 caches the data of object B11 in memory, the data of object B11 may be lost from memory. For example, when the data life cycle of object B11 in memory ends, or when the cache node powers off and restarts, etc., the data of object B11 is lost from memory. After the data of object B11 is lost from the memory of the cache node of object B11, this cache node is no longer the cache node of object B11. Therefore, it is possible that there is no cache node of object B11 in the distributed storage system 100, but only the home node of object B11 exists. In this case, the node where node 111 queries the data of object B11 stored in node 113 is the home node of object B11, and node 111 obtains the data of object B11 from the home node of object B11.
[0122] In some embodiments, the data stored in node 112 is the latest data of object B11, and what node 111 obtains from node 112 is the latest data of object B11. That is to say, when the data of object B11 stored in node 112 is the latest data of object B11, node 111 obtains the data of object B11 from node 112.
[0123] There may be multiple cache nodes of object B11 in the distributed storage system 100. For example, recently, multiple nodes have obtained and cached the data of object B11 due to processing access requests for object B11, and these nodes have become the cache nodes of object B11. Node 111 can query the node that caches the latest data of object B11 among these multiple cache nodes.
[0124] Among them, the management node of directory B1, that is, node 113, can record the status of the data of object B11 in each cache node of object B11, and the status of the data can reflect whether the data is the latest data. Node 113 can record the status of the data of object B11 in the cache node based on the cache coherence protocol.
[0125] In some embodiments, the cache coherence protocol adopted by node 113 may be the MESI (modified exclusive shared or invalid) protocol. The MESI protocol divides the data in memory into four states, namely: modified (M) state, exclusive (E) state, shared (S) state, and invalid (I) state.
[0126] If the data of object B11 cached by a node is in the M state, it means that the data of object B11 is exclusive to this node, and the data of object B11 cached by this node is dirty (that is, the data of object B11 cached by this node is different from the data of object B11 that has been persistently stored, or in other words, the data of object B11 cached by this node has not completed persistent storage), and the data of object B11 cached by other nodes is in the I state.
[0127] If the data of object B11 cached by a node is in the E state, it means that this node exclusively owns the data of object B11, and the data of object B11 cached by this node is clean (that is, the data of object B11 cached by this node is the same as the data of object B11 that has been persistently stored, or in other words, the data of object B11 cached by this node has completed persistent storage), and this node can read and write the data of object B11. Among them, if the data of object B11 is in the E state, it can be considered that the data of object B11 is the latest data.
[0128] If the data of object B11 cached by a node is in the S state, it means that the data of object B11 cached by this node is clean, and multiple nodes share this data. Among them, if the data of object B11 is in the S state, it can be considered that the data of object B11 is the latest data.
[0129] If the data of object B11 cached by a node is in the I state, it means that the data of object B11 cached by this node has been or is being modified by other nodes and is invalid data.
[0130] The state of the data of object B11 cached by a node can change, such as Figure 7As shown, the data of object B11 cached by the current node is in the E state. The current node can modify the data of object B11 so that the data of object B11 cached by the current node enters the M state. When the data of object B11 cached by the current node is in the M state, when the persistent storage of the data of object B11 cached by the current node is completed, the data of object B11 cached by the current node enters the E state. When the data of object B11 cached by the current node is in the E state, it indicates that read and write locks are imposed on the data of object B11 cached by the node. When other nodes need to read the data of object B11, the current node releases the write lock of the data of object B11 and retains the read lock of the data of object B11, and the data of object B11 cached by the current node enters the S state. Among them, the operations of releasing the write lock of the data and retaining the read lock of the data can be called lock downgrading. If the data of object B11 cached by the current node is in the E state and other nodes need to modify the data of object B11, the current node can release the read lock and the write lock, and the data of object B11 cached by the current node enters the I state. If the data of object B11 cached by the current node is in the S state and the current node needs to modify the data of object B11, the current node can add a write lock, and the data of object B11 cached by the current node enters the E state. If the data of object B11 cached by the current node is in the I state and the current node needs to modify the data of object B11, the current node can obtain the data of object B11 cached by other nodes or obtain the data of object B11 from the hard disk, and then cache the obtained data of object B11 and add a write lock to the data of object B11, and the data of object B11 cached by the current node enters the E state. If the data of object B11 cached by the current node is in the I state and the current node needs to read the data of object B11, the current node can obtain the data of object B11 cached by other nodes or obtain the data of object B11 from the hard disk, and then cache the obtained data of object B11 and add a read lock to the data of object B11, and the data of object B11 cached by the current node enters the S state.
[0131] Each node can record the state of the data of object B11 cached by the node itself, and the node can notify the state of the data of object B11 to node 113, and node 113 can record the state of the data of object B11 cached by the node. Among them, the state of the data here refers to the states of M state, E state, S state, and I state.
[0132] When processing access request A1 at node 111, if the data of object B11 is cached at node 111 itself, and the data of object B11 at node 111 itself is in the S state or E state, it means that the data of object B11 cached at node 111 itself is the latest data of object B11, and node 111 can operate on the data of object B11 cached at itself. If the data of object B11 cached at node 111 itself is in the M state, then when the data of object B11 becomes in the E state or S state, node 111 operates on the data of object B11 cached at itself. That is to say, when the data of object B11 at node 111 itself is in the S state, E state or M state, node 111 does not need to obtain the data of object B11 from other nodes.
[0133] In the case where node 111 does not cache the data of object B11 itself, and in the case where the data of object B11 cached at node 111 itself is in the I state, node 111 queries in node 113 for nodes where the cached data of object B11 is in the E state or S state, and then obtains the data of object B11 from the queried nodes.
[0134] If it is queried in node 113 that the data of object B11 cached at a certain node (for example, node 112) is in the M state, then node 111 waits until the data of object B11 cached at node 112 enters the E state or S state, and then obtains the data of object B11 from node 112.
[0135] In addition, from the above description, it can be seen that when the data of object B11 cached at a certain node is in the M state, node 111 needs to wait for the data of object B11 in the M state to complete persistent storage before it can obtain the data of object B11. The time required for persistent storage of data is relatively long (usually in milliseconds or even seconds). For this situation, an improved MESI protocol is provided below.
[0136] Refer to Figure 8 , in the improved MESI protocol, four states are added, namely M_R state, M_F state, M_EDP state, and M_W state. These four states are obtained by splitting the M state in the original MESI protocol and can be called four sub-states of the M state. Through these four sub-states, when the current node (for example, node 112) is operating on the data of object B11, node 111 can obtain the data of object B11 from node 112. Specifically as follows.
[0137] The data of object B11 cached at node 112 being in the M_W state means that node 112 is modifying the data of object B11. At this time, the data of object B11 cannot be persistently stored, cannot be read or written by other nodes, and can only be read and written by node 112.
[0138] When node 112 finishes modifying the data of object B11, the data of object B11 cached by node 112 enters the M_R state. The data of object B11 cached by node 112 being in the M_R state means that node 112 has finished modifying the data of object B11 but has not yet persisted the data of object B11. At this time, the data of object B11 cached by node 112 is the data that has been modified by node 112, that is, the most recent modification of the data of object B11 is completed, and it is the latest data of object B11. Since the data of object B11 cached by node 112 has not been persistently stored, therefore, the data of object B11 cached by node 112 at this time belongs to dirty data.
[0139] When node 112 starts to persistently store the data of object B11, the data of object B11 cached by node 112 enters the M_F state from the M_R state. The data of object B11 cached by node 112 being in the M_F state means that node 112 is persistently storing the data of object B11.
[0140] Among them, if the persistent storage of the data of object B11 fails, the data of object B11 cached by node 112 returns from the M_F state to the M_R state.
[0141] In addition, when the data of object B11 cached by node 112 is in the M_R state, if node 112 needs to modify the data of object B11 again, the data of object B11 cached by node 112 enters the M_W state. When the modification is completed again, the data of object B11 cached by node 112 enters the M_R state.
[0142] When the data of object B11 cached by node 112 is in the M_R state, if node 111 requests to obtain the data of object B11, the data of the object cached by node 112 enters the M_EDP state. When the data of the object cached by node 112 is in the M_EDP state, node 111 can obtain the data of object B11 cached by node 112. At this time, the data of object B11 is dirty data. And, when the data of the object cached by node 112 is in the M_EDP state, node 111 can impose a write lock on the data of object B11, and node 112 releases the write lock of the data of object B11, completing the transfer of the write permission of the data of object B11 from node 112 to node 111. In addition, when the data of object B11 cached by node 112 is in the M_EDP state, node 112 cannot read or write the data of object B11. When the data of object B11 cached by node 111 is persistently stored, the data of object B11 cached by node 112 enters the I state.
[0143] In some embodiments, when the data of object B11 cached by a node is in the M_F state, that is, when node 112 is persistently storing the data of object B11, if node 111 requests to obtain the data of object B11, node 112 can cancel the persistent storage of the data of object B11, that is, cancel the disk write. After canceling the disk write, the data of object B11 in node 112 enters the M_R state, and then enters the M_EDP state from the M_R state.
[0144] In some embodiments, when the data of object B11 cached by a node is in the M_F state, that is, when node 112 is persistently storing the data of object B11, if node 111 requests to obtain the data of object B11, node 112 can copy the data of object B11 to obtain two copies of the data of object B11. One of them is used for node 112 to persistently store the data of object B11, and the other is sent to node 111 for node 111 to use.
[0145] In this way, through the MESI protocol or the improved MESI protocol, it can be ensured that node 111 obtains the latest data of object B11.
[0146] In some embodiments, the node caches the data of object B11 through a page, that is to say, the data of object B11 cached by the node is located in the page. For convenience of description, the page where the data of object B11 is located can be called page B12. Among them, page B12 can be one or more pages. Specifically, when the data volume of the data of object B11 is less than or equal to the data volume that the page can accommodate, the data of object B11 can be located in one page. Among them, when the data of object B11 is less than the data volume that the page can accommodate, in addition to accommodating the data of object B11, this page can also accommodate the data of other objects. When the data volume of object B11 is greater than the data volume that the page can accommodate, the data of object B11 can be stored by multiple pages, that is to say, the data of object B11 is located in multiple pages.
[0147] As described above, the page has a globally unique identifier in the distributed storage system 100, and the identifier of the object is associated with the identifier of the page where the data of the object is located, and the management node records the node where the page where the data of the object is located is located. Thus, in step 602, node 111 can query the node having page B12 (that is, the page where the data of object B11 is located) in node 113 based on the identifier of object B11. The query result in step 603 can be the identifier of the node having page B12. It can be set that the node having page B12 is node 112, then in step 604, node 111 can obtain page B12 from node 112 based on the identifier of node 112. After obtaining page B12, the data of object B11 can be obtained from page B12.
[0148] In some embodiments, at step 602, node 111 may, based on the identifier of object B11, query in node 113 for the identifier of the page where the data of object B11 is located (i.e., the identifier of page B12). In one example, the index structure of directory B may be used to query for the page where the data of object B11 is located. Specifically, the identifier of the page may be used as the value stored in the index structure, and the identifier of the object may be used as the key stored in the index node. Thus, the identifier of the page can be queried in the index node through the identifier of the object. Among them, the identifier of the page is associated with a key range, and the key range includes one or more keys, and the key is the identifier of the object. More specifically, the size of the key of the object is represented by the order of the lexicographical order corresponding to the identifier of the object. In the index structure of directory B, querying with the key of object B11 obtains the identifier of the page associated with the key range to which object B11 belongs. The identifier of the page associated with the key range is the identifier of the page associated with object B11, the identifier of page B12. After querying the identifier of page B12, at step 602, based on the identifier of page B12, query for the node having page B12.
[0149] There may be multiple nodes having page B12, and node 111 may query in node 113 for the node having the most recent page B12. Among them, the most recent page 12 refers to the data obtained after the execution of the most recent modification operation for this page. The modification operation for the page is specifically a modification operation for the data in the page. Among them, node 113 may record the state of page B12 in the node, and the state of page B12 in the node can reflect whether page B12 in this node is the most recent page. Node 113 may record the state of page B12 in the node based on the cache coherence protocol. In some embodiments, the cache coherence protocol may be MESI as described above or the improved MESI as described above. Among them, referring to the way in which the node makes the data of object B11 enter different states introduced above, the node may make page B12 enter different states and notify node 113 of the state. Details are not described one by one here.
[0150] In this way, it can be ensured that node 111 obtains the most recent page B12, and further obtains the most recent data of object B11 from page B12.
[0151] As described above, a page is associated with the memory address of a node, and the data in the page is specifically stored in the memory associated with the page. The management node can record the node having the page and the memory address associated with the page in the node. In step 602, node 111 can query in node 113 the node having page B12 and the memory address associated with page B12 in the node. The node having the page can be set as node 112. In step 604, node 111 can obtain the data in page B12 from the memory address associated with page B12 in node 112. After obtaining the data in page B12, node 111 can associate page B12 with the memory address in node 111 and store the obtained data in page B12 into the memory address associated with page B12 in node 111.
[0152] In some embodiments, the page where the data of an object is located can also be associated with the hard disk address in the home node of the object, and this address is used for persistently storing the data of the object. The management node can record the home node of the object and the hard disk address associated with the page in the home node of the object. When there is no cache node for object B11 in the distributed storage system 100 (i.e., no node's memory address is associated with page B12), or when the data of object B11 in the cache node of object B11 is in an invalid state (such as the I state described above), node 111 can query in node 113 the home node of object B11 and the hard disk address associated with page B12 in the home node of object B11. The home node of object B11 can be set as node 114. Node 111 can obtain the data in page B12 from the hard disk address associated with page B12 in node 114 based on the query result. After obtaining the data in page B12, node 111 can associate page B12 with the memory address in node 111 and store the obtained data in page B12 into the memory address associated with page B12 in node 111.
[0153] In the above manner, node 111 can obtain the data of object B11. When node 111 obtains the data of object B11, node 111 becomes a cache node of object B11.
[0154] Among them, after node 111 obtains the data of object B11, node 113 can record that node 111 stores the data of object B11. Exemplarily, after obtaining the data of object B11, node 111 can send a notification to node 113, and this notification indicates that node 111 has obtained the data of object B11. Node 113 can record that node 111 stores the data of object B11 based on this notification.
[0155] Next, node 111 can execute step 605 to operate on the data of operation target B11 based on operation instruction C1.
[0156] As described above, the data of object B11 can be the metadata of object B11. The operations indicated by operation instruction C1 (i.e., the operations that object B11 needs to perform) can include metadata operations, such as creating or deleting object B11 in directory B, renaming object B11, hard link, sym link, getattribute, set attribute, and so on. Metadata operations on object B11 involve operating on the dentry table and inode table of directory B. For example, when creating object B11, it is necessary to insert the directory entry of object B11 into the dentry table of directory B and insert the inode of object B11 into the inode table. For another example, when renaming object B11, it is necessary to modify the directory entry of object B11 in the dentry table of directory B and modify the inode of object B11 in the inode table, and so on.
[0157] When object B11 is a file, the operations indicated by operation instruction C1 can also include data operations, such as file reading operations and file modification operations. Among them, the file modification operation includes writing data or deleting data in the content of the file. Data operations on object B11 involve the content of object B11 and the inode of object B11. Among them, when modifying the content of object B11, it is necessary to modify the inode of object B11. For example, when modifying the content of object B11, it is necessary to modify metadata such as the modification time (mtime) and file size (size) in the inode of object B11. The modification of the content of object B11 and the modification of the inode of object B11 need to ensure atomicity. Otherwise, it may cause the actual metadata of object 11 to be inconsistent with the metadata in the inode of object B11.
[0158] Some operations may require writing data, such as modification operations, creation operations, etc. There are two execution methods for such operations, namely synchronous writing and asynchronous writing. Among them, synchronous writing means that the execution result is returned only after the persistent storage of the written data is completed. Asynchronous writing means that the execution result can be returned without completing the persistent storage of the written data. The execution method of synchronous writing needs to complete the persistent storage of data before returning the execution result, resulting in a poor user experience. Asynchronous writing can return the execution result relatively quickly, providing a good user experience. However, asynchronous writing returns the execution result before the persistent storage of data is completed. Compared with synchronous writing, asynchronous writing has poor reliability.
[0159] In some embodiments, when the operation instruction C1 indicates a metadata operation and the indicated operation requires writing data, node 111 may perform the operation indicated by the operation instruction C1 in a synchronous write manner. This is because the data written in a metadata operation is metadata, and the data volume is usually small, and the time required for persistent storage is short. Performing metadata operations using synchronous writes can balance user experience and reliability.
[0160] In some embodiments, when the operation instruction C1 indicates a data operation and the indicated operation requires writing data, node 111 may perform the operation indicated by the operation instruction C1 in an asynchronous write manner. Since the data written in a data operation includes the content of a file, the data volume is usually large, and the time required for persistent storage is long. Performing data operations using asynchronous writes can guarantee the user experience.
[0161] The operation instruction C1 may indicate a metadata operation. When the object B11 is a file, the operation instruction C1 may also indicate a data operation. Next, an example is given to introduce the process of node 111 operating on the data of the object B11 based on the operation instruction C1.
[0162] In some embodiments, the operation instruction C1 is used to indicate the creation of the object B11, that is, based on the access request A1, the operation to be performed on the object B11 is specifically to create the object B11 in the directory B.
[0163] In this embodiment, in step 602, the node where the page storing the data of the object B11 is located is queried. The page storing the data of the object B11 may also be referred to as the page B12. The data of the object B11 includes the directory entry of the object B11 and the inode of the object B11. The directory entry of the object B11 needs to be inserted into the dentry table of the directory B, and the inode of the object B11 needs to be inserted into the inode table of the directory B. The page B12 includes the page for storing the directory entry of the object B11 and the page for storing the inode of the object B11. For ease of description, the page for storing the directory entry of the object B11 may be referred to as the page B121, and the page for storing the inode of the object B11 may be referred to as the page B122.
[0164] In an example, if the dentry table of the directory B is stored in a page, then the page B121 is specifically the page where the dentry table of the directory B is located.
[0165] In one example, if the dentry table of directory B is stored in multiple pages, page B121 can be one of the multiple pages. At step 602, page B121 can be queried among the multiple pages. Generally, in the dentry table of directory B, the arrangement order of the directory entries of objects is consistent with the size order of the keys of the objects. Among them, the size of the identifier of an object can be represented by the size order of the lexicographical order corresponding to the identifier. Thus, based on the size of the key of object B11, page B121 is queried among the multiple pages. Specifically, one page corresponds to a key range, and the key range includes multiple keys. This page is used to store the directory entries of objects whose keys are within this key range. At step 602, page B121 can be queried through the index structure of directory B. Specifically, the identifier of the page can be used as a leaf node in the index structure. The identifier of the page is associated with a key range, and the key range includes one or more keys, and the key is the identifier of the object. In the index structure of directory B, querying with the key of object B11 obtains the identifier of the page associated with the key range to which object B11 belongs. The identifier of the page associated with the key range is the identifier of page B121. After querying the identifier of page B121, at step 602, node 111 can query the node having page B121 based on the identifier of page B121. Then, through step 604, node 111 can obtain page B121. For specific reference, please refer to the above introduction and will not be elaborated here.
[0166] Referring to the method of obtaining page B121, node 111 can obtain page B122, which will not be elaborated here.
[0167] At step 605, node 111 can generate the data of object B11 based on operation instruction C1 and store the data of object B11 into page 12.
[0168] Among them, metadata such as the object number, name, inode ID, access permission, creation date, and size of object B11 can be generated, and based on the metadata such as the object number, name, and inode ID of object B11 used to form the directory entry, the directory entry of object B11 is generated, and based on the metadata such as the access permission and creation date of object B11 used to form the inode, the inode of object B11 is generated. Node 111 stores the directory entry of object B11 into page B121 and stores the inode of object B11 into page B122.
[0169] In some embodiments, when node 111 is a storage node, access request A1 is received by node 111 from computing node D1 in distributed storage system 100, and operation instruction C1 is used to indicate that a read operation needs to be performed on object B11, that is, computing node D1 needs to obtain the data of object B11. When node 111 obtains the data of object B11, in step 605, node 111 sends the obtained data of object B11 to computing node D1, so that computing node D1 obtains the data of object B11.
[0170] In some embodiments, operation instruction C1 is used to indicate that a modification operation needs to be performed on object B11. In step 605, node 111 modifies the data of object B11 based on operation instruction C1. Node 111 caches the modified data of object B11.
[0171] In one example, operation instruction C1 can indicate to modify the name of object B11. The data of object B11 can specifically be the directory entry of object B11. Node 111 can modify the directory entry of object B11 based on the modification of the name of object B11. In one example, operation instruction C1 can indicate to modify the access permission of object B11. The data of object B11 can specifically be the inode of object B11. Node 111 can modify the inode of object B11 based on the modification of the permission of object B11. In one example, object B11 is a file, operation instruction C1 can indicate to modify the content of object B11. The data of object B11 includes the content of object B11 and the inode of object B11. Node 111 can modify the content of object B11 and the inode of object B11 based on operation instruction C1.
[0172] Among them, the above examples illustrate the operations of node 111 on the data of object B11, which are not exhaustive. In other embodiments, node 111 can also perform other operations on object B11, which will not be listed one by one here.
[0173] So far, node 111 can complete the operations on the data of object B11.
[0174] Among them, the above process of processing the access request can be divided into semantic processing and data processing. Among them, semantic processing can also be called semantic operation, which refers to obtaining the information that can be executed by the node based on the information of the upper-layer application in the access request to facilitate data processing on the data of the object. Specifically, semantic processing includes the node identifying the access object of the access request based on the object indication in the access request, identifying the storage location of the data of the access object, and so on. Data processing, also known as data operation, refers to the operation of the node on the data of the access object, such as the object modification operation, object read operation, object creation operation, etc. mentioned above.
[0175] In some embodiments, after the node 111 finishes operating on the data of object B11, it may initiate the persistent storage of the data of object B11. For convenience of description, the data of object B11 after being operated on by the node 111 based on the operation instruction C1 may be referred to as the data of the operated object B11. The node 114 may be set as the node for persistent storage of object B11.
[0176] In some embodiments, the node 113 records the node for persistent storage of object B11. In an example of this embodiment, the node for persistent storage of object B11 may be the home node of object B11.
[0177] When it is queried that the node 114 is the node for persistent storage of object B11, in step 606, the node 111 may send the data of the operated object B11 to the node 114. In step 607, the node 114 may perform persistent storage on the data of the operated object B11.
[0178] In some embodiments, the node 114 is the home node of object B11, that is, the hard disk of the node 114 is used to store the data of the object. In this embodiment, the node 114 stores the data of the operated object B11 into the hard disk of the node 114, realizing the persistent storage of the data of the operated object B11. In some embodiments, the node 114 is not the home node of object B11, but only the node that performs the persistent storage operation on object B11. In this embodiment, the node 114 may persistently store the data of the operated object B11 into the hard disk of the home node of object B11. Exemplarily, the node 114 may send the data of the operated object B11 to the network card of the home node of object B11, and the network of the home node of object B11 may directly store the received data of the operated object B11 into the hard disk of the home node.
[0179] In this way, the persistent storage of the data of the operated object B11 can be completed.
[0180] In some embodiments, as described above, the nodes in the distributed storage system 100 have a log module and a sequence number module. The log module is used to generate a log for an operation, and this operation may be an operation performed on an object. For convenience of description, the log generated for an operation performed on an object is referred to as the log of this object. The sequence number module is used to generate a sequence number for the log. Among them, whenever an operation is performed on an object, a log of this object can be generated. The sequence number of the log is also called the serial number of the log, indicating the generation order of the log. Among them, the sequence number of the log can be recorded into this log.
[0181] The log of an object includes the data of the object, specifically the data after an operation is performed on the object. For example, Node 111 operates on the data of Object B11 and obtains the data of Object B11 after the operation. The log module in Node 111 generates a log b1 for the data of Object B11 in this operation of Node 111. Then, log b1 includes the data of Object B11 after the operation. As described above, the cache node of an object caches the data of the object in memory. The latest data of an object may be a single copy, that is, there may be only one node caching the latest data of the object. If a node fails or a node powers off and restarts, it may cause the loss of the data of the object. The data of the object in the log is used to recover the lost data of the object.
[0182] To ensure the generation efficiency of logs, each node has a log module and a sequence number module. The log module in the node is used to generate logs for the operations performed on the object by this node, and the sequence number module in the node is used to generate sequence numbers for the logs generated by this node. Therefore, multiple nodes may store the logs of the same object. Among them, the data of the object in the most recently generated log is the latest data of the object. When recovering the lost data of the object, the log of the object most recently generated needs to be used. To identify the generation order of the logs of the same object in different nodes, the sequence numbers of the logs of the same object need to be globally incremented, that is, regardless of which node performs the most recent operation on the object, the sequence number of the log generated for the most recent operation on the object is the maximum sequence number of the logs of the object. The global increment of the sequence number can be achieved through the following two methods.
[0183] In one method, a global sequence number is used. For example, the node uses a global clock for related operations, that is, the clocks used by different nodes in the distributed storage system 100 are the same. In this way, the sequence number module in the node can generate a sequence number for the log based on the time when this node performs the operation on the object. In this way, it is ensured that the execution time of the operation on the object and the size of the sequence number of the log generated for this operation are consistent.
[0184] In another way, as described above, a node obtains the data of an object from other nodes. When the node obtains the data of the object from other nodes, it also obtains the sequence number of the log of the object. Then, based on the sequence number of the log obtained from other nodes, it generates the sequence number of the log in this node, so that the sequence number of the log in this node is greater than the sequence number of the log obtained from other nodes, thereby achieving the global increment of the sequence number. For example, when node 111 obtains the data of object B11 from node 112, node 111 obtains the sequence number of the log of object B11 in node 112 from node 112. Among them, when node 112 has multiple logs of object B11, the sequence number obtained by node 111 is the sequence number of the log with the largest sequence number among the multiple logs. Node 111 operates on the data of object B11 and generates a log b1 of object B11 for this operation. The sequence number module of node 111 generates the sequence number of log b1 based on the sequence number obtained by node 111 from node 112. Specifically, the sequence number module of node 111 can first generate a sequence number for log b1, and compare the generated sequence number with the sequence number obtained from node 112. If the generated sequence number is less than or equal to the sequence number obtained by node 111 from node 112, add 1 to the sequence number obtained from node 112, and use the result as the sequence number of log b1. If the generated sequence number is greater than the sequence number obtained by node 111 from node 112, the generated sequence number is used as the sequence number of log b1.
[0185] In this way, when the node loses the latest data of the object, the node can restore the latest data of the object based on the latest log of the object. Among them, the latest log of the object refers to the log with the largest sequence number among all the logs of the object.
[0186] In addition, the node can persistently store the logs generated by this node to achieve the permanent preservation of the logs. For example, the logs can be stored on the hard disk. Another example is that when the memory of the node includes NVRAM, the logs can be stored in the NVRAM. The node can cache the sequence numbers of the logs in the memory to facilitate the node to call (such as sending the sequence numbers of the logs to other nodes, etc.).
[0187] In some embodiments, as described above, the data of the object is stored in a page, and the operation of the node on the data of the object is to operate on the data of the object in the page. Then the log of the object is the log of the page where the data of the object is located. Specifically, when the node operates on the data of the object, it generates a log for the page where the object is located, and generates a sequence number for the log of the page. And, according to the execution time of the operation on the object, the sequence numbers of the logs of the page also maintain a global increment. The specific implementation method can refer to the above introduction and will not be elaborated here.
[0188] Among them, the log of the page includes the serial number of the log, the identifier of the page, the data of the object after the operation, and the position of the data of the object after the operation in the page. Among them, the position of the data in the page can be represented by the length of the data and the offset address of the data in the page. Thus, when a node loses the latest data of an object, the node can restore the latest data of the object based on the latest log of the page where the data of the object is located. Among them, the latest log of the page refers to the log with the largest serial number among all the logs of the page.
[0189] In some embodiments, there may be multiple accessed objects for access request A1. That is to say, in addition to object B11, there may be other objects accessed through access request A1. Thus, multiple objects can be accessed through access request A1. The processing of other objects by node 111 can be implemented with reference to the introduction of the processing of object B11 above, and will not be elaborated here one by one.
[0190] In some embodiments, as described above, the nodes in the distributed storage system 100 have a load balancing module. Through the load balancing module, a node with a smaller load can be selected to process the access request.
[0191] In an example, the storage nodes in the distributed storage system 100 are used to process access requests, and the computing nodes are used to send access requests to the storage nodes. In this example, the computing nodes have a load balancing module, can select a storage node with a smaller load in the distributed storage system 100, and send the access request to the selected storage node. For example, as Figure 9 shown, when a computing node has an access request to send to a storage node, if the computing node identifies that the load of storage node E1 is greater than the load of storage node E2, the computing node will send the access request to storage node E2, so that storage node E2 processes this method request.
[0192] In an example, the nodes in the distributed storage system 100 for processing access requests have a load balancing module. When a node for processing an access request receives an access request, if the load of the node is large, the node can select a node with a smaller load from the nodes for processing access requests through the load balancing module, and forward the access request to the selected node. For example, when node 111 receives access node A2, it can compare the load of node 111 itself with the loads of other nodes, and when the load of node 111 itself is greater than the loads of other nodes, forward access request A2 to other nodes. In an example, as Figure 10As shown, when node 111 receives access request A2, it first compares the load of node 111 itself with the load of the management node of access request A2. The management node of access request A2 is the management node of the access object of access request A2, for example, node 113. Thus, when the load of the management node of access request A2 is less than or equal to the load of the access node of access request A2, the management node of access request A2 processes this access request. When the management node processes the access request, the management node can query the node storing the data of the access object by itself, which improves the query efficiency.
[0193] In some embodiments, as described above, one or more nodes in the distributed storage system 100 are used as policy control nodes. The policy control node has a policy control module. Through the policy control module, the policy control node can control the load balancing policy to control how the nodes in the distributed storage system 100 forward access requests according to the load balancing policy. Among them, the control of the load balancing policy includes load identification and policy formulation, which are specifically as follows.
[0194] Refer to Figure 11 , each node in the distributed storage system 100, such as node 111, node 112, etc., can collect its own running metrics. Among them, the running metrics of the node are used to calculate the load of the node. Exemplarily, the metrics used to calculate the load include the utilization rate of the central processing unit (CPU), the memory utilization rate, and the information of the workload model (such as the read-write ratio, the random order degree, etc.).
[0195] Each node can report the collected running metrics to the policy control node. The policy control node can calculate the load of the node based on the running metrics of the node. In some embodiments, the policy control node can determine whether a directory access hot spot occurs in the node based on the load of the node. Specifically, if the policy control node finds that the load of a certain node is heavy, the policy control node can further analyze the reason for the heavy load of the node. For example, the policy control node can obtain the log of the node and analyze the reason for the heavy load of the node based on the log of the node. If the reason for the heavy load of the node is that the node frequently processes access requests for directories, it means that a directory access hot spot has occurred in the node. Among them, the access request for a directory refers to the access request for the objects in the directory, that is, the objects in the directory are the access objects of the access request.
[0196] In this way, the policy control node can complete load identification. Next, the policy control node can formulate a load balancing policy based on the result of load identification, that is, the load of the node and the directory access hot spot.
[0197] In some embodiments, refer to Figure 12, if the difference between the loads of each node is small, the load balancing policy formulated by the policy control node can be the management node priority policy or the home node priority policy. Exemplarily, when the difference between the load of the node with the largest load and the load of the node with the smallest load is less than the preset load threshold Y1, it can be confirmed that the difference between the loads of each node is small. Exemplarily, the load threshold can be represented as a percentage of the maximum load of the node. For example, the load threshold Y1 is 5% of the maximum load of the node.
[0198] Among them, the management node priority policy means that the access request is preferentially forwarded to the management node corresponding to the access request, so that the management node processes the access request. Among them, the management node corresponding to the access request refers to the management node of the access object of the access request. Exemplarily, when performing load balancing based on the management node priority policy, the access node of the access request can compare its own load with the load of the management node corresponding to the access request. If the load of the access node is greater than or equal to the management node corresponding to the access request, the access node forwards the access request to the management node corresponding to the access request.
[0199] The home node priority policy means that the access request is preferentially forwarded to the home node corresponding to the access request, so that the home node processes the access request. Among them, the home node corresponding to the access request refers to the home node of the access object of the access request. Exemplarily, when performing load balancing based on the home node priority policy, the access node of the access request can compare its own load with the load of the home node corresponding to the access request. If the load of the access node is greater than or equal to the home node corresponding to the access request, the access node forwards the access request to the home node corresponding to the access request.
[0200] In some embodiments, refer to Figure 12 , if there are nodes with light loads, the load balancing policy formulated by the policy control node can be the recommended node policy. Among them, the recommended node policy uses the node with a light load as the recommended node, so that when the access node of the access request has a heavy load, the access request is forwarded to the recommended node, so that the recommended node processes the access request. Among them, when the load of the node is less than the preset load threshold Y2, it can be confirmed that the load of the node is light. When the load of the node is greater than the preset load threshold Y3, it can be confirmed that the load of the node is heavy. Among them, the load threshold Y3 is greater than the load threshold Y2. Exemplarily, the load threshold Y2 is 10% of the maximum load, and the load threshold Y2 is 60% of the maximum load. In addition, the policy control node can add the identifier of the recommended node to the recommended node policy. Thus, when the access node of the access request receives the recommended node policy, it can identify the recommended node based on the identifier of the recommended node.
[0201] In some embodiments, if a directory access hotspot occurs in a node, the policy control node may formulate a load balancing policy for the access nodes of the directory, such as Figure 12 As shown, the formulated load balancing policy may be a ratio policy or a polling policy. Among them, the access nodes of the directory refer to the access nodes of the access requests whose access objects are the objects in the directory. Generally, the node where the directory access hotspot occurs is the access node of the directory.
[0202] The ratio policy means that the access nodes of the directory distribute the access requests received for the directory within a preset duration to different nodes among multiple nodes according to a preset ratio, so that different nodes process different access requests for the directory. Among them, the access nodes of the directory may belong to the multiple nodes or may not belong to the multiple nodes. For example, the preset ratio is 20%, and correspondingly, the multiple nodes are specifically 5 nodes. For another example, the preset ratio is 10%, and correspondingly, the multiple nodes are specifically 10 nodes. Among them, the access nodes of the directory can predict the number of access requests received for the directory within the preset duration based on the frequency of receiving the access requests for the directory in history, and thus can distribute the received access requests for the directory to different nodes according to the preset ratio. Among them, the policy control node may add the identifiers of the multiple nodes to the ratio policy. Thus, when receiving the ratio policy, the access nodes of the directory can distribute the access requests for the directory to different nodes among the multiple nodes based on the identifiers of the multiple nodes.
[0203] The polling policy means that the access nodes of the directory sequentially send the access requests received for the directory to different nodes among multiple nodes, so that different nodes process different access requests for the directory. That is, when receiving an access request for the directory, the access request is sent to one of the multiple nodes, and when receiving the next access request for the directory, the access request is sent to the next node among the multiple nodes. Among them, the access nodes of the directory may belong to the multiple nodes or may not belong to the multiple nodes. In addition, the policy control node may add the identifiers of the multiple nodes to the polling policy. Thus, when receiving the polling policy, the access nodes of the directory can sequentially send the access requests for the directory to different nodes among the multiple nodes based on the identifiers of the multiple nodes.
[0204] When formulating the load balancing policy, the policy control module may send the formulated load balancing policy to the corresponding node, so that the corresponding node can forward the access request based on the load balancing policy to hand over the access request to other nodes for processing, thereby achieving load balancing and avoiding or eliminating access hotspots.
[0205] In some embodiments, when the access node of the access request and the cache node of the access object of the access request are not the same node, the access node of the access request may forward the access request to the cache node of the access object of the access request, so that the cache node of the access object of the access request processes the access request, thereby saving the operation of the processing node of the access request receiving data from the cache node and improving the processing efficiency of the access request. Among them, the access node of the access request may query the cache node of the access object of the access request in the management node of the access object of the access request, so as to forward the access request to the cache node.
[0206] In some embodiments, the access node of the access request and the cache node of the access object of the access request may jointly process the access request, so that the load generated by processing the access request is borne by multiple nodes, avoiding the occurrence of access hotspots. Next, in combination with Figure 13 , the solution of this embodiment will be introduced.
[0207] As Figure 14 shown, in step 1401, node 111 receives access request A3. In some embodiments, node 111 may be the access node of access request A3.
[0208] Access request A3 includes object indication B3 and operation indication C3. Among them, object indication B3 indicates the access object of access request A3. It can be set that the access object of access request A3 is the object B31 in directory B, then object indication B3 indicates the object B31 in the directory. Operation indication C3 indicates the operation to be performed on object B3.
[0209] Node 111 may perform semantic processing on access request A3. As described above, semantic processing can also be called semantic operation, which refers to obtaining the information executable by the node based on the information of the upper-layer application in the access request to facilitate data processing of the data of the object.
[0210] The semantic processing performed by node 111 on access request A3 includes identifying the access object of access request A3 based on object indication B3. Among them, object indication B3 includes the identifier of directory B and the identifier of object B31 in the directory. Node 111 can identify that the access object of access request A3 is the object B31 in directory B based on the identifier of directory B and the identifier of object B31.
[0211] The semantic processing performed by node 111 on access request A3 further includes querying the storage location of the data of the access object (i.e., object B31) of access request A3. As Figure 13 shown, node 111 may query the storage location of the data of object B31 in node 113 through step 1302.
[0212] In one example, as described above, the data of an object is stored in a page. The storage location of the data of object B31 can be the identifier of the page where the data of object B31 is located. Among them, node 113 records the page where the data of object B31 is located, and node 111 can query the page where the data of object B31 is located in node 113 to obtain the storage location of the data of object B31.
[0213] In one example, as described above, the data of an object is stored in a page, and the data of the object is specifically stored in the memory address associated with the page. The storage location of the data of object B31 can be the memory address associated with the page where the data of object B31 is located. Among them, the memory address associated with the page where the data of object B31 is located is specifically a section of the address in the memory of the cache node of object B31. Among them, the cache node of object B31 can be node 112.
[0214] Node 111 obtains the query result from node 113 through step 1303, and the query result includes the storage location of the data of object B31 that is queried.
[0215] So far, node 111 has completed the semantic processing of access request A3.
[0216] Node 111 sends access request A3 and the storage location of the data of object B31 to the cache node of object B31, that is, node 112, through step 1304.
[0217] Node 112 can perform data processing on object B31. Specifically, in step 1305, node 112 operates on the data of object B31 based on access request A3 and the storage location of the data of object B31. Among them, when the storage location is the identifier of a page, node 112 can identify the memory address associated with the identifier of the page, and then, based on the operation instruction C3 in access request A3, operate on the data in this memory address, so as to complete the operation on the data of object B31. When the storage location is a memory address, node 112 can operate on the data in this memory address based on operation instruction C3, so as to complete the operation on the data of object B31.
[0218] In some embodiments, after node 112 finishes operating on the data of object B31, node 112 can send the operated data of object B31 to node 114 through step 1306. Node 114 persistently stores the operated data of object B31 through step 1307. Specifically, it is implemented with reference to the introduction of step 607 above and will not be elaborated here.
[0219] In summary, in the method provided by the embodiments of the present application, the node that processes the access request and the node that stores the data of the access object may not be the same node. Among them, when the node that processes the access request and the node that stores the data of the access object are not the same node, the node that processes the access request may obtain the data of the access object from the node that stores the data of the access object, so that the node that processes the access request can complete the processing of the access request. In addition, the node that processes the access request and the node that stores the data of the access object may be different, which can avoid or eliminate problems such as access hotspots caused by directory ownership access.
[0220] Referring to Figure 14 , an embodiment of the present application provides a data processing apparatus 1400. The apparatus 1400 may be configured in a first node among a plurality of nodes included in a distributed storage system; wherein, the objects in the directory are stored in the plurality of nodes, and the objects are subdirectories or files of the directory. As Figure 14 shown, the apparatus 1400 includes:
[0221] A receiving unit 1410, configured to receive a first access request, the first access request including a first operation instruction and a first object instruction, the first object instruction being used to indicate a first object in the directory, and the first operation instruction being used to indicate an operation to be performed on the first object;
[0222] A query unit 1420, configured to, based on the first object instruction, query that a second node among the plurality of nodes stores data of the first object;
[0223] An obtaining unit 1430, configured to obtain the data of the first object from the second node;
[0224] An operation unit 1440, configured to, based on the first operation instruction, operate on the data of the first object.
[0225] In some embodiments, a third node among the plurality of nodes is used to record the nodes storing the data of the objects in the directory; the query unit 1420 is configured to: based on the first object instruction, query in the third node that the second node stores the data of the first object; wherein, the third node is used to record that the first node stores the data of the first object after the obtaining unit 1430 obtains the data of the first object.
[0226] In an example of this embodiment, the apparatus 1400 further includes: a sending unit 1450; wherein, the receiving unit 1410 is further configured to: receive a second access request, the second access request including a second operation instruction and a second object instruction, the second object instruction being used to indicate a second object in the directory, and the second operation instruction being used to indicate an operation to be performed on the second object; the sending unit 1450 is configured to: in a case where the load of the first node is greater than or equal to the load of the third node, send the second operation instruction and the second object instruction to the third node; wherein, the third node is configured to obtain data of the second object based on the second object instruction, and operate on the data of the second object based on the second operation instruction.
[0227] In some embodiments, the obtaining unit 1430 is configured to: when the data of the first object stored in the second node is the latest data of the first object, obtain the data of the first object from the second node.
[0228] In some embodiments, the plurality of nodes further includes a fourth node for persistently storing data of the first object, and the apparatus further includes: a sending unit 1450, configured to send the data of the first object after the operation to the fourth node, so that the fourth node persistently stores the data of the first object after the operation; wherein, the data of the first object after the operation is data obtained by the operation unit 140 operating on the data of the first object based on the first operation instruction.
[0229] In some embodiments, the data of the first object is located in a first page; the query unit 1420 is configured to: based on the first object instruction, query that the second node includes the first page; the obtaining unit 1430 is configured to: obtain the first page from the second node to obtain the data of the first object.
[0230] In some embodiments, the first page can be associated with memory addresses of different nodes among the plurality of nodes, and the memory address associated with the first page is used to store the data in the first page; the query unit 1420 is configured to: based on the first object instruction, query that the first page is associated with a first memory address of the second node; the obtaining unit 1430 is configured to: obtain the data in the first page from the first memory address; the obtaining unit 1430 is further configured to: associate the first page with a second memory address of the first node, and store the data in the first page into the second memory address.
[0231] In some embodiments, the apparatus 1400 further includes: a sending unit 1450; wherein, the receiving unit 1410 is configured to: receive a third access request, the third access request including a third operation instruction and a third object indication, the third object being used to indicate a third object in the directory, and the third operation instruction being used to indicate an operation to be performed on the third object; the query unit 1420 is configured to: based on the third object indication, query that the fifth node among the multiple nodes stores data of the third object; the sending unit 1430 is configured to: send the third operation instruction to the fifth node, so that the fifth node operates on the data of the third object based on the third operation instruction.
[0232] The functions of the functional units of the apparatus 1400 can be implemented with reference to the operations performed by the node 111 described above. For example, Figure 6 or Figure 10 or Figure 13 the operations performed by the node 111. Additionally, the second node can be the node 112 described above, the third node can be the node 113 described above, and the fourth node can be the node 114 described above.
[0233] An embodiment of the present application provides a data processing apparatus 1500. As Figure 15 shown, the data processing apparatus 1500 includes a processor 1510 and a memory 1520. The memory 1520 is used to store an executable program. The processor 1510 is configured to execute the executable program stored in the memory 1520, so that the data processing apparatus 1500 can perform the operations performed by the node 111 in the above text, such as Figure 6 or Figure 10 or Figure 13 the operations performed by the node 111.
[0234] An embodiment of the present application further provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computer device or be stored in any available medium. When the computer program product runs on a computer device, it causes the computer device to perform the operations performed by the node 111 in the above text, such as Figure 6 or Figure 10 or Figure 13 the operations performed by the node 111.
[0235] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computer device or a data storage device such as a data center including one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid state drives), etc. The computer-readable storage medium includes instructions that direct the computer device to perform the operations performed by node 111 in the foregoing, such as Figure 6 or Figure 10 or Figure 13 the operations performed by node 111 therein.
[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method in a distributed storage system, characterized in that The distributed storage system includes multiple nodes. Among them, the objects in the directory are stored in the multiple nodes. The object is a sub-directory or file of the directory. The method includes: A first node among the multiple nodes receives a first access request. The first access request includes a first operation instruction and a first object instruction. The first object instruction is used to indicate a first object in the directory, and the first operation instruction is used to indicate the operation to be performed on the first object. Based on the first object instruction, the first node queries that a second node among the multiple nodes stores the data of the first object. The first node obtains the data of the first object from the second node. The first node operates on the data of the first object based on the first operation instruction.
2. The method according to claim 1, characterized in that A third node among the multiple nodes is used to record the nodes storing the data of the objects in the directory. The first node queries that a second node among the multiple nodes stores the data of the first object based on the first object instruction, including: The first node queries in the third node that the second node stores the data of the first object based on the first object instruction. Among them, the third node is used to record that the first node stores the data of the first object after the first node obtains the data of the first object.
3. The method according to claim 2, characterized in that, The method further includes: The first node receives a second access request. The second access request includes a second operation instruction and a second object instruction. The second object instruction is used to indicate a second object in the directory, and the second operation instruction is used to indicate the operation to be performed on the second object. In the case where the load of the first node is greater than or equal to the load of the third node, the first node sends the second operation instruction and the second object instruction to the third node; among them, The third node is used to obtain the data of the second object based on the second object instruction and operate on the data of the second object based on the second operation instruction.
4. The method according to any one of claims 1 to 3, characterized in that, The first node obtains the data of the first object from the second node, including: When the data of the first object stored in the second node is the latest data of the first object, the first node obtains the data of the first object from the second node.
5. The method according to any one of claims 1 to 4, characterized in that, The multiple nodes further include a fourth node for persistently storing the data of the first object. The method further includes: The first node sends the data of the first object after operation to the fourth node, so that the fourth node persistently stores the data of the first object after operation; among them, the data of the first object after operation is the data obtained by the first node operating on the data of the first object based on the first operation instruction.
6. The method according to any one of claims 1-5, characterized in that The data of the first object is located in the first page. The first node queries that a second node among the multiple nodes stores the data of the first object based on the first object instruction, including: The first node queries that the second node includes the first page based on the first object instruction. The first node obtains data of the first object from the second node, including: the first node obtains the first page from the second node to obtain data of the first object.
7. The method according to claim 6, wherein The first page can be associated with memory addresses of different nodes among the multiple nodes, and the memory addresses associated with the first page are used to store data in the first page; The first node, based on the indication of the first object, queries that the second node includes the first page, including: the first node, based on the indication of the first object, queries that the first page is associated with a first memory address of the second node; The first node obtains the first page from the second node, including: The first node obtains data in the first page from the first memory address; The first node associates the first page with a second memory address of the first node and stores the data in the first page in the second memory address.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: The first node receives a third access request, where the third access request includes a third operation indication and a third object indication, the third object is used to indicate a third object in the directory, and the third operation indication is used to indicate an operation that needs to be performed on the third object; The first node, based on the third object indication, queries that a fifth node among the multiple nodes stores data of the third object; The first node sends the third operation indication to the fifth node, so that the fifth node operates on the data of the third object based on the third operation indication.
9. A data processing device, characterized in that, A distributed storage system includes multiple nodes, and the device is configured in a first node among the multiple nodes; wherein, objects in a directory are stored in the multiple nodes, the objects are subdirectories or files of the directory, and the device includes: A receiving unit, configured to receive a first access request, where the first access request includes a first operation indication and a first object indication, the first object indication is used to indicate a first object in the directory, and the first operation indication is used to indicate an operation that needs to be performed on the first object; A query unit, configured to, based on the first object indication, query that a second node among the multiple nodes stores data of the first object; An obtaining unit, configured to obtain data of the first object from the second node; An operating unit, configured to operate on the data of the first object based on the first operation indication.
10. The device according to claim 9, characterized in that, A third node among the multiple nodes is used to record nodes storing data of objects in the directory; The query unit is configured to: based on the first object indication, query in the third node that the second node stores data of the first object; Wherein, the third node is configured to record that the first node stores data of the first object after the obtaining unit obtains the data of the first object.
11. The device according to claim 10, characterized in that, The device further includes: a sending unit; wherein, The receiving unit is further configured to: receive a second access request, where the second access request includes a second operation instruction and a second object instruction, the second object instruction is used to indicate a second object in the directory, and the second operation instruction is used to indicate an operation to be performed on the second object; The sending unit is configured to: when the load of the first node is greater than or equal to the load of the third node, send the second operation instruction and the second object instruction to the third node; where, The third node is configured to obtain data of the second object based on the second object instruction, and operate on the data of the second object based on the second operation instruction.
12. The device according to any one of claims 9 to 11, characterized in that The obtaining unit is configured to: when the data of the first object stored in the second node is the latest data of the first object, obtain the data of the first object from the second node.
13. The device according to any one of claims 9-12, characterized in that, The multiple nodes further include a fourth node for persistently storing data of the first object, and the apparatus further includes: A sending unit, configured to send the data of the first object after operation to the fourth node, so that the fourth node persistently stores the data of the first object after operation; where the data of the first object after operation is data obtained by the operation unit operating on the data of the first object based on the first operation instruction.
14. The device according to any one of claims 9-13, characterized in that, The data of the first object is located in a first page; The querying unit is configured to: based on the first object instruction, query that the second node includes the first page; The obtaining unit is configured to: obtain the first page from the second node to obtain the data of the first object.
15. The device according to claim 14, characterized in that, The first page can be associated with memory addresses of different nodes among the multiple nodes, and the memory address associated with the first page is used to store the data in the first page; The querying unit is configured to: based on the first object instruction, query that the first page is associated with a first memory address of the second node; The obtaining unit is configured to: obtain the data in the first page from the first memory address; The obtaining unit is further configured to: associate the first page with a second memory address of the first node, and store the data in the first page into the second memory address.
16. The device according to any one of claims 9-15, characterized in that, The apparatus further includes: a sending unit; where, The receiving unit is configured to: receive a third access request, where the third access request includes a third operation instruction and a third object instruction, the third object is used to indicate a third object in the directory, and the third operation instruction is used to indicate an operation to be performed on the third object; The querying unit is configured to: based on the third object instruction, query that a fifth node among the multiple nodes stores the data of the third object; The sending unit is configured to: send the third operation instruction to the fifth node, so that the fifth node operates on the data of the third object based on the third operation instruction.
17. A data processing device, characterized in that, Including: A memory, configured to store an executable program; A processor, configured to execute the method according to any one of claims 1-8 by running the executable program.
18. A computer-readable storage medium, characterized in that, Comprising computer program instructions which, when executed by a computer device, cause the computer device to perform the method according to any one of claims 1-8.
19. A computer program product comprising instructions, characterized in that, When the instructions are run by a computer device, cause the computer device to perform the method according to any one of claims 1-8.
Citation Information
Cited By
Method and apparatus for processing data
WO2025138810A1