Data management method, device and system, storage medium and computer program product

By monitoring and matching the timing information of IO service data in a distributed storage system, the data inconsistency caused by virtual IP drift is solved, and the consistency of IO data and system stability are achieved.

CN119945887APending Publication Date: 2025-05-06SANGFOR TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411774898.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During the virtual IP drift of the client node, the distributed storage system cannot recognize the order of the IO content, resulting in the IO data in the failed node being released. Data coverage may occur after recovery, resulting in inconsistent data.

Method used

By monitoring the IO service data output by the distributed storage node, the timing information it contains is determined, and matches it with the latest timing information. If it does not match, it will not respond to the output operation of IO service data to ensure the consistency of the data.

Benefits of technology

It ensures consistency of IO data before and after virtual IP drift, avoids data coverage problems, and improves the stability and user experience of distributed storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945887A_ABST
    Figure CN119945887A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data management method, which is applied to a distributed storage system, and comprises the following steps: if it is monitored that a first distributed storage node outputs input / output IO service data corresponding to a first virtual address, determining first time sequence information included in the IO service data; and if the first time sequence information is not matched with the latest time sequence information corresponding to the first virtual address, not responding to the output operation of the IO service data. The embodiment of the invention also discloses a data management device and system, a storage medium and a computer program product. Through the method of managing the IO data of the corresponding client node based on the virtual IP drift condition, the consistency of the IO data executed before and after the virtual IP drift is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of distributed storage technology, and in particular to a data management method, device, system, storage medium, and computer program product. Background Art

[0002] With the rapid development of cloud computing technology, distributed storage systems can automatically switch to normal distributed storage nodes when any distributed storage node in the distributed storage system to which the client node is connected fails by combining the virtual Internet Protocol (IP) with service status detection technologies such as Keepalived technology and Cluster Trivial Database (CTDB) technology. As the distributed storage nodes are automatically switched, the virtual IP corresponding to the client node will drift between the distributed storage nodes to continuously provide distributed storage services to the client through the virtual IP. This process is transparent to the client and technically implements a High Availability Cluster (HACluster) and optimizes the user experience by improving the continuity and stability of distributed storage.

[0003] However, currently, when the virtual IP corresponding to the client node drifts between distributed storage nodes, the distributed storage system cannot identify the order of the input / output (IO) content of the client node, resulting in the IO content in the faulty distributed storage node not being released. Before the faulty distributed storage node is restored, the drifted virtual IP executes the new IO content in the new distributed storage node. After the faulty distributed storage node is restored, it is easy for the IO content in the faulty distributed storage node to overwrite the new IO content, resulting in data inconsistency. Summary of the Invention

[0004] In view of this, the embodiments of the present application hope to provide a data management method, device, system, storage medium and computer program product, which solves the problem that during the current client node virtual IP drift process, the IO data stored in the corresponding distributed storage node before the virtual IP drift may overwrite the IO data in the corresponding distributed storage node after the virtual IP drift occurs, resulting in data inconsistency. A method for managing the IO data of the corresponding client node based on the virtual IP drift status is proposed, which ensures the consistency of the IO data executed before and after the virtual IP drift.

[0005] To achieve the above objectives, the technical solution of this application is implemented as follows:

[0006] The present application provides a data management method, which is applied to a distributed storage system and includes:

[0007] If it is monitored that the first distributed storage node outputs input / output IO service data corresponding to the first virtual address, determining first timing information included in the IO service data;

[0008] If the first timing information does not match the latest timing information corresponding to the first virtual address, the output operation of the IO service data is not responded to.

[0009] In the above solution, the method further includes:

[0010] Delete the IO service data.

[0011] In the above solution, the first timing information includes the first virtual address and first drift identification information for identifying a drift status of the first virtual address, and the first virtual address is used for the client node to establish a communication connection with the distributed storage system.

[0012] In the above solution, the method further includes:

[0013] If a failure of the first distributed storage node is detected, determining one or more second virtual addresses configured on the first distributed storage node; wherein the one or more second virtual addresses include the first virtual address;

[0014] Determining one or more third distributed storage nodes from one or more second distributed storage nodes included in the distributed storage system; wherein the first distributed storage node belongs to one or more of the second distributed storage nodes;

[0015] Drift one or more of the second virtual addresses to one or more of the third distributed storage nodes.

[0016] In the above solution, the method further includes:

[0017] Determining the latest historical drift identification information corresponding to each of the second virtual addresses;

[0018] Each piece of the historical drift identification information is updated to obtain second drift identification information of each second virtual address.

[0019] In the above solution, the method further includes:

[0020] Based on each piece of second drift identification information and the corresponding node identification information of the third distributed storage node, the routing policy of the corresponding second virtual address is updated.

[0021] In the above solution, the method further includes:

[0022] If an IO service write request sent by the client node based on the first virtual address is received, determining the IO service to be executed corresponding to the IO service write request;

[0023] Determining the latest current timing information of the first virtual address;

[0024] Using the latest timing information to identify the IO service to be executed, to obtain an identified IO service;

[0025] The identified IO service is written into the third distributed storage node corresponding to the first virtual address.

[0026] The present application provides a data management device, which is applied to a distributed storage system. The device includes: a determination unit and an execution unit; wherein:

[0027] The determining unit is configured to, upon receiving input / output IO service data corresponding to a first virtual address sent by a first distributed storage node, determine first timing information included in the IO service;

[0028] The execution unit is configured to not respond to the output operation of the IO service data if the first timing information does not match the latest timing information corresponding to the first virtual address.

[0029] The present application provides a distributed storage system, comprising: one or more second distributed storage nodes; wherein:

[0030] The distributed storage system is used to manage one or more of the second distributed storage nodes and to implement the steps of any of the above-mentioned data management methods.

[0031] The present application provides a storage medium having a data management program stored thereon. When the data management program is executed by a processor, the steps of any of the above-mentioned data management methods are implemented.

[0032] The present application provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned data management methods when executed by a processor.

[0033] The data management method, device, system, storage medium and computer program product provided by the embodiments of the present application are as follows: if a first distributed storage node monitors that the first distributed storage node outputs IO business data corresponding to a first virtual address, the first timing information included in the IO business data is determined, and if the first timing information does not match the latest timing information corresponding to the first virtual address, the output operation of the IO business data is not responded to. In this way, the distributed storage system analyzes the first timing information of the IO business data output by the first distributed storage node, and if the first timing information does not match the latest timing corresponding to the first virtual address, the output operation of the IO business data is not responded to, thereby ensuring that the data that the client node can obtain is the latest, achieving data consistency, and solving the problem that in the current process of virtual IP drift of the client node, the IO data stored in the corresponding distributed storage node before the virtual IP drift may overwrite the IO data in the corresponding distributed storage node after the virtual IP drift occurs, resulting in data inconsistency. A method for managing the IO data of the client node corresponding to the virtual IP drift is proposed, thereby ensuring the consistency of the IO data executed before and after the virtual IP drift. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A flowchart of a data management method provided in an embodiment of the present application;

[0035] Figure 2 A schematic diagram of the main module architecture of a distributed storage system provided in an embodiment of the present application;

[0036] Figure 3 A schematic diagram of the system structure of a distributed storage system provided in an embodiment of the present application;

[0037] Figure 4 A schematic diagram of the structure of a data management device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0039] The embodiment of the present application provides a data management method, which is applied to a distributed storage system. Figure 1 As shown, the method includes the following steps:

[0040] Step 101: If it is monitored that a first distributed storage node outputs input / output IO service data corresponding to a first virtual address, first timing information included in the IO service data is determined.

[0041] In an embodiment of the present application, the first virtual address is typically a first virtual Internet Protocol (IP) address. A virtual IP is an IP that is not assigned to a real host. When a host node server fails and cannot provide external services, the virtual IP can be dynamically switched to a backup host. In distributed storage, a cluster trivial database (CTDB) is often used to manage virtual IPs. The first timing information is used to indicate the order in which the first virtual address drifts to the first distributed storage node, that is, the first timing information can indicate the number of times the first virtual address has drifted.

[0042] The distributed storage system manages the distributed storage nodes it manages, and upon monitoring that a first distributed storage node outputs input / output (IO) service data corresponding to a first virtual address, determines first timing information included in the IO service data. The IO service data may be obtained by the first distributed storage node responding to an IO service request, obtaining corresponding request data, obtaining the first timing information corresponding to the first distributed storage node, and identifying the obtained request data using the first timing information.

[0043] Step 102: If the first timing information does not match the latest timing information corresponding to the first virtual address, do not respond to the output operation of the IO service data.

[0044] In an embodiment of the present application, the latest timing information corresponding to the first virtual address can be obtained by the distributed storage system from the cluster monitor (Monitors, MON) of the distributed storage system, that is, after the timing of each virtual address in the distributed storage system is updated, it will be backed up and updated in the MON. The first timing information is matched with the latest timing information corresponding to the first virtual address, that is, it is determined whether the first timing information is consistent with the latest timing information corresponding to the first virtual address. If the first timing information is inconsistent with the latest timing information corresponding to the first virtual address, it is determined that the first timing information does not match the latest timing information corresponding to the first virtual address. At this time, the distributed storage system does not respond to the output operation of the first distributed storage node, that is, it does not read the IO business data output by the first distributed storage node. In other words, the distributed storage system does not forward the IO business data output by the first distributed storage node to the client node corresponding to the first virtual address.

[0045] Based on the foregoing embodiment, in other embodiments of the present application, the distributed storage system is further configured to perform the following steps:

[0046] Delete IO service data.

[0047] In an embodiment of the present application, after the distributed storage system determines that it does not respond to the output operation of the first distributed storage node, it can delete the received IO business data. In this way, the storage resources of the distributed storage system can be saved and the utilization rate of the storage resources of the distributed storage system can be improved.

[0048] Based on the foregoing embodiment, in other embodiments of the present application, the first timing information includes a first virtual address and first drift identification information for identifying a drift status of the first virtual address, and the first virtual address is used to establish a communication connection between the client node and the distributed storage system.

[0049] In the embodiment of the present application, the first drift identification information can be represented by the number of drifts, which can also be called a version number, and can be identified by an epoch.

[0050] Based on the foregoing embodiment, in other embodiments of the present application, the distributed storage system is further configured to perform the following steps:

[0051] If a failure of the first distributed storage node is detected, determining one or more second virtual addresses configured on the first distributed storage node; wherein the one or more second virtual addresses include the first virtual address;

[0052] Determining one or more third distributed storage nodes from one or more second distributed storage nodes included in the distributed storage system; wherein the first distributed storage node belongs to one or more second distributed storage nodes;

[0053] Drift the one or more second virtual addresses to one or more third distributed storage nodes.

[0054] In an embodiment of the present application, one or more second virtual addresses configured on the first distributed storage node are stored in a virtual address pool to which the first distributed storage node is mounted.

[0055] The distributed storage system performs fault monitoring on one or more second distributed storage nodes it manages. When a fault is detected in a first distributed storage node, one or more second virtual addresses configured on the first distributed storage node are drifted. In this case, the one or more second virtual addresses are all virtual addresses included in the virtual address pool of the first distributed storage node. In some application scenarios, if the number of distributed storage nodes to which the drifting can be performed is insufficient, or the number of virtual addresses that can be configured on the distributed storage node to which the drifting can be performed is insufficient, the one or more second virtual addresses that are drifted may also be part of the virtual addresses configured on the first distributed storage node, for example, virtual addresses that are used more frequently in the first distributed storage node. The drift processing process may be to determine one or more third distributed storage nodes from all the second distributed storage nodes included in the distributed storage system, and then migrate and allocate the one or more second virtual addresses to the one or more third distributed storage nodes, so that the one or more third distributed storage nodes provide distributed storage services for the client nodes corresponding to the one or more second virtual addresses. In this way, it is ensured that the distributed storage system can continue to provide distributed storage services for the client nodes corresponding to the one or more second virtual addresses, thereby realizing a high-availability cluster and ensuring the user experience.

[0056] Among them, it should be noted that the one or more second distributed storage nodes included in the distributed storage system may be all storage nodes included in the distributed storage system, including available storage nodes and unavailable storage nodes. In this case, one or more third distributed storage nodes selected from the one or more second distributed storage nodes included in the distributed storage system are available storage nodes. In some application scenarios, the one or more second distributed storage nodes included in the distributed storage system may only be the currently available storage nodes included in the distributed storage system. In the above two scenarios, determining one or more third distributed storage nodes may be randomly selected from one or more second distributed storage nodes, or may be selected based on the distributed node performance such as the remaining resources of the distributed storage nodes in one or more second distributed storage nodes. For example, the first number of one or more second virtual addresses is determined, and then the available distributed storage nodes in the one or more second distributed storage nodes are sorted in order from most to least in terms of remaining resources, and the first first number of distributed storage nodes are selected to obtain one or more third distributed storage nodes.

[0057] Based on the foregoing embodiment, in other embodiments of the present application, the distributed storage system is further configured to perform the following steps:

[0058] Determining the latest historical drift identification information corresponding to each second virtual address;

[0059] Each piece of historical drift identification information is updated to obtain second drift identification information of each second virtual address.

[0060] In an embodiment of the present application, this part of the steps may be performed after the distributed storage system drifts one or more second virtual addresses to one or more third distributed storage nodes. The latest historical drift identification information corresponding to each second virtual address may be obtained from the storage area storing the latest historical drift identification information corresponding to each second virtual address, and then the latest historical drift identification information corresponding to each second virtual address obtained may be updated to obtain the second drift identification information of each second virtual address. Exemplarily, when the latest historical drift identification information of a second virtual address is 2, the latest historical drift identification information 2 of the second virtual address is updated, for example, by adding 1, and the second drift identification information of the second virtual address is obtained as 3.

[0061] Based on the foregoing embodiment, in other embodiments of the present application, the distributed storage system is further configured to perform the following steps:

[0062] Each second virtual address and the corresponding second drift identification information are stored as timing information parameters in a storage area; wherein the storage area includes at least one of the following: a protocol gateway service corresponding to the third distributed storage node corresponding to each second virtual address, and a cluster monitor of the distributed storage system. It should be noted that the above storage areas are merely exemplary, and this application does not impose any limitations on the storage areas of the timing information parameters.

[0063] In an embodiment of the present application, each second virtual address is bound to the corresponding second drift identification information to form a timing information parameter for each second virtual address, and the timing information parameter is stored in the corresponding storage area. When the storage area is a protocol gateway service corresponding to a distributed storage node, the distributed storage node stores the timing information corresponding to all configured virtual addresses. When the storage area is a cluster monitor, the cluster monitor can store the timing information parameters corresponding to all virtual addresses included in the distributed storage system. In this way, when performing timing information comparison and analysis, the distributed storage system can obtain the current and latest historical timing information from the storage area.

[0064] Based on the foregoing embodiment, in other embodiments of the present application, the distributed storage system is further configured to perform the following steps:

[0065] Based on each second drift identification information and the node identification information of the corresponding third distributed storage node, the routing policy of the corresponding second virtual address is updated.

[0066] In an embodiment of the present application, after determining the third distributed storage node corresponding to each second virtual address after drift, the routing policy for each second virtual address is updated. Thus, when data traffic corresponding to each second virtual address subsequently arrives, the corresponding data traffic is directed to the corresponding third distributed storage node via the corresponding routing policy. Simultaneously, the second drift identification information for each second virtual address is also updated in the routing policy for the second virtual address. Thus, the number of times each second virtual address has drifted can be obtained from the routing policy, allowing subsequent analysis based on the corresponding second drift identification information, and even using the corresponding second drift identification information to identify the data traffic being transmitted. This improves data processing efficiency.

[0067] Based on the foregoing embodiment, in other embodiments of the present application, the distributed storage system is further configured to perform the following steps:

[0068] If an IO service write request sent by the client node based on the first virtual address is received, determining the IO service to be executed corresponding to the IO service write request;

[0069] Determine the latest current timing information of the first virtual address;

[0070] The latest timing information is used to identify the IO business to be executed, and the identified IO business is obtained;

[0071] The identification IO service is written into the third distributed storage node corresponding to the first virtual address.

[0072] In an embodiment of the present application, an IO service write request is a request for writing the corresponding data content to the corresponding third distributed storage node, that is, a request operation for storing data, and the IO service to be executed is writing the content. The current latest timing information of the first virtual address can be obtained from the storage area, for example, it can be obtained from the protocol gateway service, or it can be obtained from the cluster monitor. The specific information can be determined by the actual application scenario and is not specifically limited here. Among them, the latest timing information is associated with the third distributed storage node, and the latest current timing information is calculated after the first virtual address drifts to the third distributed storage node.

[0073] After the client node establishes a communication connection with the distributed storage system based on the first virtual address, when the client node generates data corresponding to the IO service that needs to be stored in the distributed storage system, the client node sends an IO service write request to the distributed storage system through the first virtual address. The distributed storage system determines the corresponding pending IO service that needs to be stored and processed based on the IO service write request, that is, the corresponding IO service data that needs to be stored. At this time, the distributed storage node determines the latest timing information currently corresponding to the first virtual address, and uses the latest timing information to identify the received pending IO service, that is, uses the latest timing information to identify the received IO service data that needs to be stored, thereby obtaining an identified IO service, that is, the IO service data identified and processed using the latest timing information. Afterwards, the identified IO service obtained by the identification processing is written and solidified to realize the storage of the identified IO service in the third distributed storage node corresponding to the first virtual address. In this way, when the corresponding IO service read request is received subsequently, the identified IO service can be read from the third distributed storage node to obtain the IO service data identified with the latest timing information.

[0074] Based on the above embodiments, the present application provides a distributed storage system, the structure of which can be used as follows: Figure 2 As shown, it includes: access layer (Access Layer), storage service layer (Service Layer) and persistence layer (Persistence Layer), wherein the access layer includes the client access layer and the network high availability module, the storage service layer includes a first service module for providing metadata services and a second service module for providing data services, and the persistence layer includes a storage pool. Figure 3 As shown, the distributed storage system includes n nodes, wherein the floating IP included in each node is assigned an epoch information for indicating the drift timing of the floating IP in the distributed storage system. For example, Figure 2 The epoch information of different nodes in the current assignment is 1.

[0075] When the distributed storage system detects a failure in node 1, the network high availability module drifts the floating IP address (e.g., xx.xx.xx.1) on node 1 to node 2. At this point, the network high availability module changes the epoch of xx.xx.xx.1 to 2 and stores it in the cluster monitor (MON). Simultaneously, the module updates xx.xx.xx.1 and the epoch value of 2 to the protocol gateway service corresponding to node 2. It should be noted that node 2 is selected by the network high availability module from nodes other than node 1, specifically by analyzing the remaining resources and service capabilities of the remaining nodes. Finally, the information that xx.xx.xx.1 has been switched to node 2 is set on the corresponding network card. This ensures that subsequent I / O services reach the corresponding protocol gateway, such as the Server Message Block (SMB) service and / or the Network File System (NFS) service, through xx.xx.xx.1. The received I / O services are then identified using xx.xx.xx.1 and the epoch value of 2.

[0076] If the distributed storage system receives expired IO services, it will determine whether the current IO service has expired based on the epoch corresponding to the IP address and perform the corresponding discard operation. The process can be as follows:

[0077] 1. The client sends I / O services to the distributed storage system through the NFS and SMB protocols. The client accesses the distributed storage system through the floating IP address xx.xx.xx.1.

[0078] 2. After receiving the I / O request, the protocol gateway service of the distributed storage system identifies it with IP address xx.xx.xx.1 and an epoch value of 1 before sending it to the distributed storage system. However, an unexpected failure of Node 1 causes a floating IP switch, shifting the floating IP address xx.xx.xx.1 to Node 2 and changing the epoch value to 2.

[0079] Among them, the drift timing can be recorded as a version number. For example, when the virtual IP1 is in node A, the epoch value of the virtual IP1 is 1. When node A fails, the virtual IP1 drifts to node B, and the epoch of the virtual IP1 will be modified to 2. If the node B fails after the failure of node A is repaired, and the virtual IP1 drifts to node A again, the epoch of the virtual IP1 will be modified to 3. That is, every time each virtual IP drifts, the corresponding version number information epoch will be updated in a +1 manner to reflect the change in drift. In some application scenarios, a single change method with a gradually decreasing epoch can also be used to implement it, which can be determined by the actual implementation process.

[0080] The process of using IP xx.xx.xx.1 and an epoch value of 1 to identify the corresponding virtual IP may be implemented by a management component of the distributed storage system.

[0081] 3. The client retries the storage IO and writes new content.

[0082] 4. The distributed storage system receives the IO service previously sent by node 1.

[0083] 5. The distributed storage system identifies the epoch in the IO service request issued by node 1 and determines that the epoch is 1, which is smaller than the latest epoch 2 of the floating IP xx.xx.xx.1 currently stored in the distributed storage system. Therefore, the distributed storage system discards the IO service.

[0084] The tag information is bound / tagged on the IO service, and the drift comparison process can be as follows:

[0085] 1. The pending IO initiates an execution request to the management component of the service layer.

[0086] The management component can be divided into a metadata management component and a data management component. The metadata management component is used to respond to and manage metadata IO requests, and the data management component is used to respond to and manage data IO requests.

[0087] 2. The management component obtains the timing stamp information of the IO to be executed, and compares it based on the drift count of the virtual address obtained in MON, and decides to execute or discard the IO based on the comparison result.

[0088] Among them, the management component can obtain the drift count stored in MON from MON by subscription. Correspondingly, a specific implementation method is that MON can regularly push the latest drift count to all subscribers according to the update strategy, wherein the management component can cache the latest drift count of all virtual addresses received in the management component; in this process, when the network high availability module detects that the floating IP has drifted, it will update the corresponding drift count in MON.

[0089] In some application scenarios, another way for the management component to obtain the latest drift count is that when the network high availability module detects that the floating IP has drifted, it updates the corresponding drift count in MON and synchronizes the updated drift count to all subscribers, including the management component.

[0090] 6. Notify the protocol gateway service to discard the IO business.

[0091] Among them, when executing IO, the version number carried by the virtual IP address corresponding to the current IO can be retrieved. If it is lower than the version number corresponding to this virtual IP in the current system, it means that this IO is old IO and can be discarded.

[0092] It should be noted that during the implementation of the above method of this application, it is mainly aimed at the update operation of metadata or data. When the update operation includes adding, deleting, modifying, etc., the epoch needs to be verified to ensure the consistency of the data.

[0093] In some application scenarios, the metadata management component may be implemented by a first service module for providing metadata services, and the data management component may be implemented by a second service module for providing data services.

[0094] It should be noted that the aforementioned virtual address, floating IP and virtual IP address are all the same content.

[0095] In this way, the floating IP drift process is identified and marked in a distributed time series in the distributed storage system to solve potential data inconsistencies caused by the IP switching process of IO services in the distributed storage system, thereby ensuring data consistency.

[0096] The data management method provided by the embodiment of the present application is as follows: if the first distributed storage node monitors that the first distributed storage node outputs IO business data corresponding to the first virtual address, the first timing information included in the IO business data is determined, and if the first timing information does not match the latest timing information corresponding to the first virtual address, the output operation of the IO business data is not responded to. In this way, the distributed storage system analyzes the first timing information of the IO business data output by the first distributed storage node, and if the first timing information does not match the latest timing corresponding to the first virtual address, the output operation of the IO business data is not responded to, thereby ensuring that the data that the client node can obtain is the latest, achieving data consistency, and solving the problem that in the current client node virtual IP drift process, the IO data stored in the corresponding distributed storage node before the virtual IP drift may overwrite the IO data in the corresponding distributed storage node after the virtual IP drift occurs, resulting in data inconsistency. A method for managing the IO data of the client node corresponding to the virtual IP drift is proposed, thereby ensuring the consistency of the IO data executed before and after the virtual IP drift.

[0097] It should be noted that, at present, it is usually adopted to identify the timing of each IO business data, and to sort each IO business data according to the above timing when performing data management, so as to avoid the old IO business data covering the new IO business data. The currently commonly used scheme will bring a large workload of timing identification and timing sorting, thereby increasing the load of the distributed storage system. The data management method provided by the embodiment of the present application only needs to identify the timing information of the virtual IP, and assign the timing information of the virtual IP to the IO business data. Compared with the sequential identification timing based on each IO business data, the number of timing identifications is greatly reduced, and a lightweight timing identification scheme is realized; at the same time, by comparing the timing information of the virtual IP, the timing of the IO business data can be quickly determined, and the segmented business scenario where the IO content in the failed distributed storage node covers the new IO content is avoided. At the same time, the comparison speed of the IO timing comparison is improved, and the load of the distributed storage system is reduced.

[0098] Based on the above embodiments, the present invention provides a data management device, which is applied to a distributed storage system. The data management device can be applied to Figure 1 And corresponding embodiments, refer to Figure 4 As shown, the data management device 2 includes: a determination unit 21 and an execution unit 22; wherein:

[0099] The determining unit 21 is configured to determine first timing information included in the IO service upon receiving input / output IO service data corresponding to the first virtual address sent by the first distributed storage node;

[0100] The execution unit 22 is configured to not respond to the output operation of the IO service data if the first timing information does not match the latest timing information corresponding to the first virtual address.

[0101] In other embodiments of the present application, the device further includes: a deletion unit; wherein:

[0102] Delete IO service data.

[0103] In other embodiments of the present application, the first timing information includes a first virtual address and first drift identification information for identifying a drift status of the first virtual address, and the first virtual address is used for the client node to establish a communication connection with the distributed storage system.

[0104] In other embodiments of the present application, the device further includes: a processing unit; wherein:

[0105] The determining unit is further configured to, if a failure of the first distributed storage node is detected, determine one or more second virtual addresses configured on the first distributed storage node; wherein the one or more second virtual addresses include the first virtual address;

[0106] The determining unit is further configured to determine one or more third distributed storage nodes from the one or more second distributed storage nodes included in the distributed storage system; wherein the first distributed storage node belongs to one or more second distributed storage nodes;

[0107] The processing unit is configured to drift one or more second virtual addresses to one or more third distributed storage nodes.

[0108] In other embodiments of the present application, the apparatus further includes: an updating unit; wherein:

[0109] a determining unit, configured to determine the latest historical drift identification information corresponding to each second virtual address;

[0110] The updating unit is configured to update each piece of historical drift identification information to obtain second drift identification information of each second virtual address.

[0111] In other embodiments of the present application, the device further includes: a storage unit; wherein:

[0112] A storage unit is used to store each second virtual address and the corresponding second drift identification information as timing information parameters in a storage area; wherein the storage area includes at least one of the following: a protocol gateway service corresponding to a third distributed storage node corresponding to each second virtual address, and a cluster monitor of the distributed storage system.

[0113] In other embodiments of the present application, the updating unit is further configured to update the routing policy of the corresponding second virtual address based on each second drift identification information and the node identification information of the corresponding third distributed storage node.

[0114] In other embodiments of the present application, the device further includes: an identification unit and a writing unit; wherein:

[0115] a determining unit configured to, upon receiving an IO service write request sent by the client node based on the first virtual address, determine the IO service to be executed corresponding to the IO service write request;

[0116] a determining unit, configured to determine the latest current timing information of the first virtual address;

[0117] An identification unit, configured to identify the IO service to be executed using the latest timing information to obtain an identified IO service;

[0118] The writing unit is configured to write the identification IO service into the third distributed storage node corresponding to the first virtual address.

[0119] It should be noted that the information interaction process between the units and modules in this embodiment can refer to the information interaction process described in the aforementioned method embodiment, and will not be repeated here.

[0120] The data management device provided in the embodiment of the present application, if receiving the IO business data corresponding to the first virtual address output by the first distributed storage node and monitoring the first distributed storage node, determines the first timing information included in the IO business data, and if the first timing information does not match the latest timing information corresponding to the first virtual address, does not respond to the output operation of the IO business data. In this way, the distributed storage system analyzes the first timing information of the IO business data output by the first distributed storage node, and does not respond to the output operation of the IO business data when the first timing information does not match the latest timing corresponding to the first virtual address, thereby ensuring that the data that the client node can obtain is the latest, achieving data consistency, and solving the problem that in the current client node virtual IP drift process, the IO data stored in the corresponding distributed storage node before the virtual IP drift may overwrite the IO data in the corresponding distributed storage node after the virtual IP drift occurs, resulting in data inconsistency. A method for managing the IO data of the client node corresponding to the virtual IP drift is proposed, thereby ensuring the consistency of the IO data executed before and after the virtual IP drift.

[0121] Based on the above embodiments, the embodiments of the present application provide a distributed management system that can be applied to Figure 1 In the corresponding embodiment, the distributed management system includes at least: one or more second distributed storage nodes; wherein:

[0122] A distributed storage system is used to manage one or more second distributed storage nodes to implement the following Figure 1 The data management methods provided in the corresponding embodiments will not be described in detail here.

[0123] Based on the above embodiments, the embodiments of the present application provide a computer-readable storage medium, referred to as a storage medium, which stores one or more data management methods, and the one or more data management methods can be executed by one or more processors to implement the following. Figure 1 The data management methods provided in the corresponding embodiments will not be described in detail here.

[0124] Based on the above embodiments, the embodiments of the present application provide a computer-readable storage medium, referred to as a storage medium, which stores one or more programs, which can be executed by one or more processors to implement the reference Figure 1 The implementation process of the method provided in the corresponding embodiment will not be repeated here.

[0125] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0126] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, ..., air conditioner, or network communication link device, etc.) to execute the methods described in each embodiment of the present application.

[0128] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0129] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0131] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A data management method, characterized in that: The method is applied to a distributed storage system, and the method comprises: If it is monitored that the first distributed storage node outputs input / output IO service data corresponding to the first virtual address, determine first timing information included in the IO service data; If the first timing information does not match the latest timing information corresponding to the first virtual address, the output operation of the IO service data is not responded to.

2. The method according to claim 1, characterized in that The method further comprises: Delete the IO service data.

3. The method according to claim 1, characterized in that The first timing information includes the first virtual address and first drift identification information for identifying a drift status of the first virtual address, and the first virtual address is used for a client node to establish a communication connection with the distributed storage system.

4. The method according to claim 1, characterized in that: The method further comprises: If it is detected that the first distributed storage node fails, determine one or more second virtual addresses configured on the first distributed storage node; wherein the one or more second virtual addresses include the first virtual address; Determine one or more third distributed storage nodes from one or more second distributed storage nodes included in the distributed storage system; Drift one or more of the second virtual addresses to one or more of the third distributed storage nodes.

5. The method according to claim 4, characterized in that The method further comprises: Determine the latest historical drift identification information corresponding to each of the second virtual addresses; Each of the historical drift identification information is updated to obtain second drift identification information of each of the second virtual addresses.

6. The method according to claim 5, characterized in that The method further comprises: Based on each of the second drift identification information and the corresponding node identification information of the third distributed storage node, the corresponding routing policy of the second virtual address is updated.

7. The method according to claim 4, characterized in that The method further comprises: If an IO service write request sent by the client node based on the first virtual address is received, determining the IO service to be executed corresponding to the IO service write request; Determine the latest current timing information of the first virtual address; Using the latest timing information to identify the IO service to be executed, to obtain an identified IO service; The identified IO service is written into the third distributed storage node corresponding to the first virtual address.

8. A data management device, characterized in that: The device is applied to a distributed storage system, and comprises: a determination unit and an execution unit; wherein: The determining unit is configured to determine first timing information included in the IO service if input / output IO service data corresponding to the first virtual address sent by the first distributed storage node is received; The execution unit is configured to not respond to an output operation of the IO service data if the first timing information does not match the latest timing information corresponding to the first virtual address.

9. A distributed storage system, characterized in that: The distributed storage system comprises: one or more second distributed storage nodes; wherein: The distributed storage system is used to manage one or more of the second distributed storage nodes, and to implement the steps of the data management method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a data management program, and when the data management program is executed by the processor, the steps of the data management method according to any one of claims 1 to 7 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the data management method according to any one of claims 1 to 7.