Scheduling permission acquisition method, device, system and storage medium

CN116185589BActive Publication Date: 2026-10-09ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310147892.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2026-10-09
Estimated Expiration
2043-02-10

AI Technical Summary

Benefits of technology

[0009] In this embodiment, multiple management nodes can compete for the right to update the lease information stored on the storage service node to become the master node. As the master node, it has scheduling authority over the data in the storage cluster, thus achieving master election among multiple management nodes. This master-competition method, compared to the distributed lock method, does not require waiting for the distributed lock to expire, which helps improve the efficiency of master-slave node switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116185589B_ABST
    Figure CN116185589B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a scheduling permission acquisition method, device, system and storage medium. In the embodiments of the present application, a plurality of management and control nodes can access a storage service node to contend for the permission to update the lease information stored by the storage service node and become a master node, so that the master node has the scheduling permission for the data of the storage cluster, and the master selection of the plurality of management and control nodes is realized. Compared with the distributed lock mode, the master selection mode does not need to wait for the distributed lock to be invalid, and helps to improve the switching efficiency of the master and standby nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data storage technology, and in particular to a method, device, system and storage medium for obtaining scheduling permissions. Background Technology

[0002] To address the challenges of managing massive amounts of data, a partitioned data management scheme has been introduced. This scheme uses a control node to manage data within a partition and is responsible for scheduling partitions across multiple storage nodes. To improve system robustness, data storage management systems typically deploy multiple control nodes. To prevent split-brain scenarios in the data storage system, multiple control nodes need to elect a master node, which is responsible for partition scheduling, thereby maintaining data consistency among storage nodes. Therefore, how to elect a master node among multiple control nodes has become a pressing technical problem to be solved in this field. Summary of the Invention

[0003] This application provides a method, device, system, and storage medium for obtaining scheduling permissions, in order to solve the problem of electing a master among multiple management nodes.

[0004] This application provides a method for obtaining scheduling permissions, applicable to management nodes, including: By accessing the storage service node, it competes with other control nodes for the right to update the lease information stored on the storage service node; the storage service node has the function of ensuring data consistency of the storage cluster; If the management node wins the right to update the lease information stored on the storage service node, it sets the node status of the management node to the master node status, so that the management node can obtain the scheduling authority of the storage cluster.

[0005] This application embodiment also provides a scheduling permission acquisition method, applicable to storage service nodes, wherein the storage service node has the function of ensuring data consistency of the storage cluster and stores lease information of the master node among multiple management nodes; The method includes: Obtain the lease update request issued by the control node; the lease update request is issued by the control node in order to compete for the right to update the lease information stored by the storage service node; Obtain the target lease information carried in the lease renewal request from the lease renewal request; If the target lease information and the lease information currently stored by the storage service node are consistent, update the lease information stored by the storage service node so that the control node obtains the permission to update the lease information stored by the storage service node; The lease update success message is returned to the management node, so that the management node can respond to the lease update success message by setting the node status to master node status and obtaining scheduling authority for the storage cluster.

[0006] This application embodiment also provides a data storage system, including: multiple management nodes and storage service nodes of a storage cluster; The storage service node has the function of ensuring data consistency of the storage cluster and stores the lease information of the master node among the multiple management nodes; the master node has scheduling authority over the storage cluster. The control node is used to execute the steps in the scheduling permission acquisition method executed by the control node above; The storage service node is used to execute the steps in the scheduling permission acquisition method executed by the storage service node.

[0007] This application embodiment also provides a computing device, including: a memory, a processor, and a communication component; the memory is used to store computer programs; The processor is coupled to the memory and is used to execute the computer program to perform the steps in the scheduling permission acquisition method executed by the control node and / or the scheduling permission acquisition method executed by the storage service node.

[0008] This application also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the scheduling permission acquisition method executed by the control node and / or the scheduling permission acquisition method executed by the storage service node.

[0009] In this embodiment, multiple management nodes can compete for the right to update the lease information stored on the storage service node to become the master node. As the master node, it has scheduling authority over the data in the storage cluster, thus achieving master election among multiple management nodes. This master-competition method, compared to the distributed lock method, does not require waiting for the distributed lock to expire, which helps improve the efficiency of master-slave node switching. Attached Figure Description

[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the structure of the storage management system provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the process of multi-control node master election provided in an embodiment of this application; Figure 3This is a schematic diagram illustrating the process of master node election provided in the embodiments of this application; Figure 4 A schematic diagram illustrating the process of selecting a master for a backup node provided in this application embodiment; Figure 5 A schematic diagram illustrating the node state switching process of the control node provided in this application embodiment; Figure 6 A schematic diagram illustrating the process of a control node scheduling data across a storage cluster, as provided in an embodiment of this application. Figures 7-9 A flowchart illustrating the scheduling permission acquisition method provided in this application embodiment; Figure 10 A schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] In some embodiments of this application, to achieve master election by multiple management nodes, multiple management nodes can access the storage service node and compete for the right to update the lease information stored on the storage service node to become the master node. As the master node, it has scheduling authority over the data in the storage cluster, thus achieving master election by multiple management nodes. This master-competition method, compared to the distributed lock method, does not require waiting for the distributed lock to expire, which helps improve the efficiency of master-slave node switching.

[0013] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0014] It should be noted that the same reference numerals denote the same object in the following figures and embodiments. Therefore, once an object is defined in one figure or embodiment, it does not need to be discussed further in subsequent figures and embodiments.

[0015] Figure 1 This is a schematic diagram of the storage management system provided in an embodiment of this application. Figure 1 As shown, the storage management system S10 refers to a system that manages the storage cluster 30. This system S10 mainly includes multiple management nodes 10 and storage service nodes 20. In the embodiments of this application, "multiple" means more than one, i.e., two or more.

[0016] In this embodiment, the storage service node 20 can provide data storage-related service functions, such as ensuring data consistency of the storage cluster 30. The storage service node 20 can ensure data consistency of the storage cluster 30 based on a distributed consistency protocol (such as the Raft protocol). Of course, in addition to ensuring data consistency of the storage cluster 30, the storage service node 20 can also have other data storage-related service functions, such as storage transaction processing.

[0017] The storage service node 20 can be implemented as a device, apparatus, virtual machine (VM), container, or software module that provides data consistency services to ensure the storage cluster 30. The device providing data consistency services can be a single server device or a cloud-based server array. Alternatively, the device providing data consistency services can also refer to other computing devices with corresponding service capabilities, such as computers or other terminal devices (running service programs). In this embodiment, the storage service node 20 can be deployed in the cloud, such as in the central cloud of an edge cloud system. Of course, the storage service node 20 can also be deployed on the same physical machine as the storage node 301 in the storage cluster 30.

[0018] The aforementioned storage cluster 30 refers to a cluster used for storing data, which may include multiple storage nodes 301. Storage nodes 301 can provide storage resources. Storage nodes 301 can also be independent storage devices, or they can be storage media within a physical machine (such as one or more disks).

[0019] In this embodiment, the storage cluster 30 can store data using a centralized storage system or a distributed storage system. The distributed storage system can be a distributed key-value storage system, etc. The data stored in the storage cluster 30 can be application data, such as file data, image data, or data tables. However, the data stored in the storage cluster 30 can also be metadata of the application data. For example, in some embodiments, the storage cluster 30 implements distributed storage of data based on block storage services or partition storage services. Figure 1 The illustration uses a storage cluster of 30 based on partitioned storage services for distributed data storage as an example, but it is not intended to be limiting. Figure 1 The partition ij is shown, where i = 0, 1…M; j = 0, 1…N.

[0020] Block storage services refer to the distributed storage of data in the form of data blocks (Chunks) on different disks (such as disks D0-Dn).

[0021] Partitioning is the process of dividing a database table into its components. Each table initially has only one partition, but as data is inserted, the management node horizontally splits the table according to certain rules, creating two partitions. As the number of rows in a table increases, more and more partitions are created. Since these partitions may not be able to be stored on the same storage node, they can be distributed across multiple storage nodes. A table can be horizontally partitioned into several partitions based on its row key. Each partition has a start row and a stop row representing the data range of that partition.

[0022] In this embodiment, the data stored in storage node 301 may be application data and / or metadata; correspondingly, the data table may be an application data table and / or a metadata table. Each row in the metadata table records information about an application data partition. The row key may include the table name, the starting row key, and timestamp information, etc. The metadata table records the address mapping relationship of the application data partitions, that is, the address of the storage node of the partition.

[0023] For distributed storage services, a master-slave service is often used to improve service stability. This involves using multiple master nodes 10 to act as both master and backup nodes, providing data storage services and partition metadata management services. The persistent data on the master node 10 can be written to a distributed file system. The master node 10 is responsible for managing the storage cluster 30. For distributed storage systems with partitioned storage, the master node 10 can also manage and allocate partitions, such as allocating new partitions when a partition splits, and migrating the partitions of a departing storage node to other storage nodes when a storage node leaves the system.

[0024] Multiple management nodes 10 can compete for master control and gain scheduling authority over storage cluster 30. The management node that successfully competes for master control becomes the master node, which has scheduling authority over storage cluster 30. To prevent split-brain scenarios, some conventional solutions use distributed coordination service software (such as Zookeeper) to allow multiple management nodes 10 to compete for master control. The distributed coordination service software provides a distributed lock service. The principle is that if an abnormal process encounters an error while maintaining the lock service, it will release the distributed lock. Other processes continuously compete for the lock to detect the anomaly and begin to take over the service. Once the lock is acquired, it indicates that it has the right to provide external services and can begin providing services.

[0025] For distributed lock services, switching between primary and backup nodes requires waiting for the lock timeout before the switch can proceed. In a distributed environment, such timeouts are typically several seconds (e.g., 5 seconds) or longer. In large-scale and complex environments, in order to tolerate short-term anomalies, such timeouts are made even longer, resulting in low service switching efficiency.

[0026] In some embodiments of this application, to improve service switching efficiency, the lease information of the master node can be stored on the storage service node 20. The lease information can be implemented as lease key-value pairs (Lease KV), and the lease key can be a set identifier, etc. Updating the lease information mainly refers to updating the lease value. In some embodiments, the lease value can be implemented as an incrementing sequence, and updating the lease information can be implemented by increasing the lease value. For example, each update can increase the lease value by a set increment based on the current value, etc. The set increment can be 1.

[0027] In this embodiment, as Figure 1 As shown, any control node 10 can compete with other control nodes for the right to update the lease information stored on the storage service node 20 by accessing the storage service node 20. The control node that wins the right to update the lease information stored on the storage service node 20 becomes the master node. The master node has scheduling authority over the storage cluster 30. Accordingly, if control node 10 wins the right to update the lease information stored on the storage service node 20, it can set its node state to master node state, thereby granting it scheduling authority over the storage cluster. The scheduling authority over the storage cluster 30 may include: operation authority over the storage nodes of the storage cluster 30 and data scheduling authority over the storage nodes of the storage cluster 30 (i.e., the authority to determine the storage location of data). For example, it can write data to the storage cluster 30, update the data stored in the storage cluster 30, or delete the data stored in the storage cluster 30, etc.

[0028] For the master node in its management state, the management node 10 can perform data scheduling on the storage cluster and generate data scheduling information. For example, the management node 10 in its master node state can perform data scheduling on the storage cluster based on a load balancing strategy to obtain data scheduling information. This data scheduling information may include, for example, the storage location of the data to be scheduled.

[0029] The front-end machine 40 can obtain data scheduling information from the control node 10. Based on the storage location in the data scheduling information, the front-end machine 40 stores the data to be scheduled in the corresponding location. Here, the front-end machine 40 refers to the logical functional node that executes the data scheduling information generated by the control node 10. It can be implemented as an independent physical machine, or as a software functional module, virtual machine, or container on a physical machine.

[0030] In this embodiment, multiple management nodes can access the storage service node and compete for the right to update the lease information stored on the storage service node to become the master node. As the master node, it has the authority to schedule data within the storage cluster, thus achieving master election among multiple management nodes. This master-competition method, compared to the distributed lock method, does not require waiting for the distributed lock to expire, which helps improve the efficiency of master-slave node switching.

[0031] Specifically, such as Figure 2 As shown, when multiple control nodes 10 compete for the master, the target lease information LV can be obtained for any control node 10. Figure 2 Step 1); The target lease information can be lease information LV1 stored locally on the control node 10, or lease information LV2 (defined as the first lease information) obtained from the storage service node. Further, the control node 10 can provide a lease update request carrying the target lease information LV to the storage service node 20 (…). Figure 2 Step 2).

[0032] Specifically, control node 10 can use the target lease information LV as a precondition to generate an atomic lease update request containing that precondition. A precondition refers to the prerequisite for a node responding to a lease update request to execute the request; that is, the conditions that the node responding to the lease update request must meet to execute the request. An atomic lease update request is a lease update request executed atomically. This request will not be interrupted by the thread scheduling mechanism during execution, ensuring the atomicity of the lease update request. In this embodiment, the node responding to the lease update request is storage service node 20, and the prerequisite for storage service node 20 to respond to the lease update request and update the stored lease information is that the lease information stored by storage service node 20 is consistent with the target lease information LV.

[0033] Optionally, the control node 10 may employ a transaction mechanism, using the target lease information LV as a precondition, to generate an atomic lease update request containing that precondition. The transaction mechanism can be understood as an if-else statement. That is, if the storage service node 20 meets the precondition, the stored lease information is updated.

[0034] In some embodiments, the lease update request may further include: lease information to be updated. Lease information to be updated refers to the content to which the lease information stored by the storage service node 20 has been updated. For lease key-value pairs, the lease information to be updated may be: the lease value in the target lease information + 1, such as LV+1.

[0035] The target lease information is determined by the current node status of the control node 10. In this embodiment, the node status includes: primary node status and backup node status. The control node in the primary node status has scheduling authority over the storage cluster, while the control node in the backup node status does not have scheduling authority over the storage cluster.

[0036] like Figure 3 As shown, if the current node state of the control node 10 is the master node state, then the control node 10 can obtain the lease information LV1 stored locally as the target lease information LV ( Figure 3 Step 1). Accordingly, the control node 10 can use the locally stored lease information LV1 as a prerequisite to generate an atomic lease update request containing the prerequisite, and provide the lease update request to the storage service node 20. Figure 3 Step 2).

[0037] Optionally, the lease update request may carry the lease information to be updated as: the lease value stored locally on the control node 10 + 1, i.e., LV1+1.

[0038] like Figure 4 As shown, if the current node state of the control node 10 is a standby node state, then the control node 10 can obtain the first lease information LV2 stored by the storage service node 20 from the storage service node 20. Figure 4 Step 1). The control node 10 can compare whether the acquired first lease information LV2 is consistent with the lease information LV1 stored locally on the control node 10. Figure 4 Step 2). If the first lease information LV2 is inconsistent with the lease information LV1 stored locally by the control node 10, it indicates that another control node is continuously updating the lease information stored by the storage service node 20, meaning the master node is continuously renewing the lease. Therefore, even if the first lease information LV2 is inconsistent with the lease information LV1 stored locally by the control node 10, the control node 10 can continue to maintain its standby node status, that is, maintain the node status of the control node 10 as a standby node. Figure 4 Step 3.2). The control node in standby node status has no scheduling authority over storage cluster 30. Furthermore, control node 10 can update the lease information of its local storage to the first lease information LV2 stored by storage service node 20. Figure 4 (Not shown in the image).

[0039] Accordingly, if the first lease information LV2 is consistent with the lease information LV1 stored locally by the control node 10, it indicates that no master node is updating the lease information stored by the storage service node 20. In this case, the control node 10, which is in standby node status, can compete for the right to update the lease information stored by the storage service node 20 to become the master node. Specifically, as follows... Figure 4As shown, if the first lease information LV2 in the standby node state is consistent with the lease information LV1 stored locally by the control node 10, the control node 10 can use the first lease information LV2 stored in the storage service node 20 as the target lease information LV. Figure 4 Step 3.1).

[0040] Accordingly, the control node 10 can use the first lease information LV2 as a prerequisite to generate an atomic lease update request containing the prerequisite.

[0041] Optionally, the lease renewal request may carry the lease information to be renewed as: the lease value in the first lease information + 1, i.e., LV2+1.

[0042] Regardless of whether it's the primary node or the backup node, after generating a lease update request, the control node 10 can provide the lease update request carrying the target lease information to the storage service node 20 (corresponding to...). Figure 2 Step 2 Figure 3 Step 2 and Figure 4 Step 3.1).

[0043] In this embodiment, the control node 10 can periodically renew the lease. Specifically, each time a renewal period arrives, the control node 10 obtains the target lease information and generates a lease update request carrying that information. Then, it provides the lease update request generated for this renewal period to the storage service node. The renewal period is less than or equal to the lease's validity period. Generally, the renewal period is ≤1 / 3 of the lease's validity period.

[0044] Storage service node 20 can receive the lease update request and obtain the target lease information from the lease update request. Furthermore, storage service node 20 can compare the currently stored lease information (defined as second lease information) LV3 with the target lease information LV to see if they match (corresponding to...). Figure 2 Step 3 Figure 3 Step 3 and Figure 4 Step 4). In Figure 3 In step 3, the target lease information LV is the first lease information LV1 stored locally on the control node 10, which is in the master node state. Figure 4 In step 4, the target lease information LV is Figure 4 The first lease information LV2 obtained in step 1 is stored in the storage service node.

[0045] The second lease information LV3 currently stored in the storage service node 20 may be the same as or different from the first lease information LV2 read from the storage service node 20 by the control node in the backup node state, depending on whether other control nodes update the lease information stored in the storage service node and become the master node.

[0046] If the second lease information LV3 currently stored by storage service node 20 is consistent with the target lease information LV, it indicates that no other controlling node has become the master node and is updating the lease information stored by storage service node 20. Therefore, controlling node 10 acquires the permission to update the lease information stored by the storage service node. Accordingly, if... Figure 2 As shown, if the second lease information LV3 currently stored by the storage service node 20 is consistent with the target lease information LV, the storage service node 20 can update the second lease information LV3 to obtain the updated lease information LV4 (corresponding to...). Figure 2 Step 4.1 Figure 3 Step 4.1 and Figure 4 Step 5.1).

[0047] Specifically, for a lease update request carrying lease information to be updated, the storage service node 20 can update the locally stored second lease information to the lease information to be updated carried in the lease update request (such as LV+1). The updated lease information LV4 is the lease information to be updated. For a lease update request without carrying lease information to be updated, the storage service node 20 can add a set increment (such as +1) to the locally stored second lease information LV3 to obtain the updated lease information LV4, and so on.

[0048] After the lease information is successfully updated, storage service node 20 can return a lease update success message to control node 10 (corresponding to...). Figure 2 Step 5.1 Figure 3 Step 5.1 and Figure 4 Step 6.1). For the control node 10, upon receiving the lease update success message, it determines that it has won the right to update the lease information stored on the storage service node 20; and sets the node state of the control node 10 to the master node state, so that the control node 10 can obtain the scheduling right of the storage cluster 30 (corresponding to...). Figure 2 Step 6 Figure 3 Step 6 and Figure 4 Step 13).

[0049] Control node 10 can also respond to a successful lease update message by modifying the locally stored lease information to the updated lease information LV4 (corresponding to) in the aforementioned storage service node 20. Figure 2 Step 7 Figure 3 Step 7 and Figure 4Step 7.1).

[0050] Accordingly, if the second lease information LV3 currently stored by storage service node 20 is inconsistent with the target lease information LV, it indicates that another control node has become the master node and updated the lease information stored by storage service node 20. In this case, control node 10 has not obtained the permission to update the lease information stored by storage service node 20. Accordingly, if... Figure 2 As shown, if the second lease information LV3 currently stored by the storage service node 20 is inconsistent with the target lease information LV, the storage service node 20 may return a lease update failure message to the control node 10 (corresponding to...). Figure 2 Step 4.2 Figure 3 Step 4.2 and Figure 4 Step 5.2).

[0051] Accordingly, upon receiving a lease update failure message, the control node 10 can determine that it has not won the right to update the lease information stored by the storage service node 20; and in response to the lease update failure message, it can set the node status of the control node 10 to the standby node status (corresponding to...). Figure 2 Step 5.2 Figure 3 Step 5.2 and Figure 4 Step 6.2). The management node in standby node status has no scheduling authority over storage cluster 30.

[0052] In the embodiment where the management node 10 periodically renews its lease, the management node 10, which is in the master node state, can change its node state to the backup node state after receiving lease renewal failure messages N times consecutively. N≥1 and N≤ the quotient of (T0 / T1), where N is an integer. T0 represents the lease validity period, and T1 represents the renewal period.

[0053] Combination Figure 3 and Figure 5 For the control node 10, which is in the master node state when sending the lease update request, if the control node (master node) successfully updates the lease information stored in the storage service node 20, it maintains the master node state. Figure 5 Step 1); If the control node (primary node) fails to update the lease information stored on storage service node 20, then the primary node status will be switched to backup node status. Figure 5 Step 2), enter the master-slave process in the backup node state.

[0054] The master-slave process for standby node status can be combined with Figure 4 and Figure 5 Explanation will be provided. (In conjunction with...) Figure 4 and Figure 5 For the control node 10, which is in standby node status, the above embodiment can be referred to. Figure 4The scheduling permission acquisition method in steps 1-6.2 is used to compete for the master node. For the management node 10, which is in standby node status, it can also update the locally stored lease information LV1 in response to the lease update failure message. Figure 4 The first lease information LV2 of the storage service node 20 obtained in step 1 ( Figure 4 Step 7.2).

[0055] In an implementation where a standby node successfully seizes the master position, the original master node may not be able to detect in time that another control node has successfully seized the master position. The original master node continues to schedule storage cluster 30, leading to split-brain and double-write problems, which damage the data consistency of storage cluster 30.

[0056] For example, consider two management nodes A and B, where management node A is the primary node and management node B is the backup node. When management node B detects that the lock of management node A has been lost, it will switch to primary node mode. However, at this time, management node A may not yet be aware that management node B has switched to primary node. If management nodes A and B simultaneously perform data scheduling on storage cluster 30, a dual-write problem will occur, resulting in a split-brain scenario.

[0057] To address the aforementioned issues, in this embodiment, input / output (IO) barrier information (defined as first IO barrier information) can be configured in the storage service node 20. The controlling node that acquires the right to update the IO barrier information stored on the storage service node 20 has that right. The master node also stores IO barrier information (defined as second IO barrier information).

[0058] IO barrier information can be implemented as IO Fence Key-Value pairs (Lease KV), where the IO Fence Key can be a set identifier. Updating lease information mainly refers to updating the IO Fence Value. In some embodiments, the IO barrier value can be implemented as an incrementing sequence, and updating the IO barrier information can be implemented by increasing the IO barrier value. For example, each update of the IO barrier information can increase the IO barrier value by a set increment based on the current value. The set increment can be 1.

[0059] like Figure 6 As shown, based on the aforementioned IO barrier information, for the control node 10 in the master node state, when performing data scheduling on the storage cluster 30, the second IO barrier information of the local storage of the control node 10 can be used as a prerequisite to generate a data scheduling request for the storage cluster; and this data scheduling request will be provided to the storage service node 20. Figure 6Step 1). The storage service node 20 can compare whether the second IO barrier information carried in the data scheduling request is consistent with the first IO barrier information stored locally by the storage service node 20. Figure 6 Step 2); If the second IO barrier information is consistent with the first IO barrier information, it indicates that the control node 10 has scheduling authority over the storage cluster 30, and the control node 10 is allowed to perform data scheduling on the storage cluster 30. Figure 6 Step 3.1).

[0060] Accordingly, the control node 10 can perform data scheduling on the storage cluster 30 according to the set data scheduling policy to obtain data scheduling information. Figure 6 Step 4.1). Data scheduling information is used to indicate which storage node the data to be scheduled will be stored on. Accordingly, data scheduling data includes: the storage location of the data to be scheduled, etc. The set data scheduling strategy can be a load balancing strategy or a storage fragmentation minimization strategy, etc. Among them, for the storage fragmentation minimization strategy, the control node 10 can determine the storage node with free storage space greater than or equal to the data to be scheduled based on the free storage space of multiple storage nodes 301; then, from the storage nodes with free storage space greater than or equal to the data to be scheduled, select the target storage node with the smallest free storage space; and use the target storage node as the storage location of the data to be scheduled.

[0061] For the front-end machine, data scheduling information can be obtained from the control node 10; and based on the data scheduling information, the data to be scheduled is stored to the target storage node, etc. For the control node 10 in standby node state, if it wins the right to update the lease information of the storage service node 20, it can also request to modify the IO barrier information of the storage service node 20. In this way, when the original master node performs data scheduling on the storage cluster, because the IO barrier information of the storage service node 20 has changed, the IO barrier information stored by the original master node is inconsistent with the IO barrier information stored by the storage service node 20, making it impossible for the original master node to perform data scheduling on the storage cluster. The original master node can promptly detect that another control node has successfully seized the master position and promptly switch its node state to standby node state, which can reduce the probability of split-brain problems.

[0062] Optionally, such as Figure 6 As shown, if the second IO barrier information is inconsistent with the first IO barrier information, it indicates that the control node 10 has no scheduling authority over the storage cluster 30, and the storage service node 20 can return a no-scheduling-authority prompt message to the control node 10. Figure 6 Step 3.2). Based on the "no scheduling permission" notification, management node 10 can promptly detect that another management node has successfully seized the master position and switch its own node status to standby node status. Figure 6 Step 4.2).

[0063] The following is combined with Figure 4 The process of a control node 10, which is in standby node status, modifying the IO barrier information stored by the storage service node 20 after winning the right to update the lease information stored by the storage service node 20 is illustrated by an example.

[0064] like Figure 4 As shown, when sending a lease update request, the control node in standby node status can also obtain the IO barrier information FV1 stored by storage service node 20 in response to the lease update success message. Figure 4 Step 8); and can generate an IO barrier update request based on the IO barrier information FV1, and provide the IO barrier update request to the storage service node 20 ( Figure 4 Step 9). The IO barrier update request carries the IO barrier information FV1 stored in storage service node 20.

[0065] Optionally, the control node 10 can use the IO barrier information FV1 stored by the storage node 20 as a precondition to generate an atomic IO barrier update request containing this precondition. A description of the precondition can be found in the section on generating lease update requests above, and will not be repeated here. An atomic IO barrier update request refers to an IO barrier update request executed atomically. This request will not be interrupted by the thread scheduling mechanism during execution, ensuring the atomicity of the IO barrier update request.

[0066] Optionally, the control node 10 may employ a transaction mechanism, using the IO barrier information FV1 as a precondition, to generate an atomic lease update request containing this precondition. That is, if the storage service node 20 meets the precondition, the storage lease information is updated.

[0067] Based on the aforementioned IO barrier update request, the storage service node 20 can obtain the IO barrier information FV1 carried in the IO barrier update request; further, it can compare whether the IO barrier information FV2 currently stored locally is consistent with the IO barrier information FV1 carried in the IO barrier update request. Figure 4 Step 10). If the currently stored IO barrier information FV2 is consistent with the IO barrier information FV1 carried in the IO barrier update request, it means that no other successfully preemptive control node 10 has updated the IO barrier information stored by storage service node 20. Therefore, the control node 10 that provided the IO barrier update request has successfully preempted the node. Accordingly, storage service node 20 can update its locally stored IO barrier information FV2 to obtain the updated IO barrier information FV3, provided that the locally stored IO barrier information FV2 is consistent with the IO barrier information FV1 carried in the IO barrier request. Figure 4 Step 11.1).

[0068] Because the standby node that successfully seized the master node updated the IO barrier information stored on the storage service node, the IO barrier information stored on the original master node became inconsistent with that stored on storage service node 20. Thus, when the original master node requests data scheduling for the storage cluster, the storage service node detects that the IO barrier information carried in the data scheduling request differs from the IO barrier information in its local storage, preventing the original master node from scheduling data. Storage service node 20 can also return a "no scheduling permission" warning message to the original master node if the IO barrier information carried in the data scheduling request differs from the IO barrier information in its local storage. This allows the original master node to promptly detect that another control node has successfully seized the master node and switch its node state to standby, thereby helping to reduce the probability of a split-brain problem.

[0069] like Figure 4 As shown, for storage service node 20, after updating the IO barrier information FV2 of local storage, it can return an IO barrier information update success message to management node 10. Figure 4 Step 12.1). In response to the successful update message of the IO barrier information, the management node 10 can switch its node state to the master node state, thereby granting the management node 10 scheduling authority over the storage cluster 30. Figure 4 Step 13). In response to the successful update message for the IO barrier information, the control node 10 can also store the updated IO barrier information FV3 stored on the storage service node 20 locally, so that subsequent data scheduling of the storage cluster 30 can be performed based on the IO barrier information FV3 as a prerequisite. Figure 4 (Not shown). For a detailed implementation of the master node's data scheduling of storage cluster 30 based on IO barrier information FV3 as a prerequisite, please refer to [link to relevant documentation]. Figure 6 The relevant content will not be repeated here.

[0070] In some embodiments, the IO barrier information FV2 currently stored locally by the storage service node 20 may differ from the IO barrier information FV1 carried in the IO barrier update request. Accordingly, if the IO barrier information FV2 currently stored locally differs from the IO barrier information FV1 carried in the IO barrier update request, the storage service node 20 may return an IO barrier update failure message to the management node 10. Figure 4 Step 11.2).

[0071] For control node 10, in response to an IO barrier update failure message, the node state of control node 10 can be maintained in the standby node state. Figure 4 Step 12.2), and wait for the next master-grabbing cycle to arrive, then continue the above process of grabbing the master by accessing the storage service node.

[0072] In other embodiments, after the standby node wins the right to update the lease information, the storage service node 20 may malfunction, or the network between the standby node and the storage service node 20 may malfunction, preventing the storage service node 20 from subsequently updating the IO barrier information. In this case, if the management node 10, which is in the standby node state that has won the right to update the lease information, is changed to the master node state, the original master node may not yet be aware that another management node has switched to the master node because the IO barrier information stored by the storage service node 20 has not been updated. If the new master node and the original master node simultaneously perform data scheduling on the storage cluster 30, a double-write problem will occur, resulting in a split-brain scenario.

[0073] Therefore, to address this issue, if the control node 10, currently in standby mode, does not receive a successful IO barrier information update message within a set timeframe, its node status can be maintained as standby, awaiting the next master-claiming cycle to resume the aforementioned process of accessing the storage service node to claim master. This reduces the probability of a split-brain scenario. Combined with... Figure 4 achievable Figure 5 This diagram illustrates the node state switching of the control node regarding the backup node status. Specifically, for control node 10, which is in backup node status when sending a lease update request, if the lease information in the storage service node is continuously updated, the backup node status is maintained. Figure 5 Step 3): If the lease information in the storage service node is no longer updated, the system enters a state of vying for the right to update the lease information stored on the storage service node. Figure 5 Step 4). If the control node (backup node) fails to update the lease information stored on storage service node 20, then maintain the backup node status. Figure 5 Step 5).

[0074] If the control node (standby node) successfully updates the lease information of storage service node 20, it enters the process of competing for the right to update the IO barrier information of storage service node. Figure 5 Step 6). If the IO barrier information update fails, maintain the backup node status ( Figure 5 Step 7); If the IO barrier information is successfully updated, switch to the master node state ( Figure 5 Step 8).

[0075] In addition to the storage management system provided in the above embodiments, this application also provides a scheduling permission acquisition method. The scheduling permission acquisition method provided in this application is illustrated below from the perspectives of the management node and the storage service node. The scheduling permission acquisition method described in this application refers to a method for acquiring scheduling permissions for a storage cluster.

[0076] Figure 7 This is a flowchart illustrating a scheduling permission acquisition method provided in an embodiment of this application. This method is applicable to management nodes. Figure 7 As shown, the method mainly includes: 701. By accessing the storage service node, it competes with other management nodes for the right to update the lease information stored on the storage service node; the storage service node is used to ensure data consistency of the storage cluster.

[0077] 702. If the node wins the right to update the lease information of the storage service node, the node status of the control node is set to the master node status, so that the control node can obtain the scheduling authority of the storage cluster.

[0078] In some embodiments of this application, to improve service switching efficiency, lease information of the master node can be stored on the storage service node. A description of the lease information can be found in the relevant content of the above system embodiments, and will not be repeated here.

[0079] In this embodiment, for the control node, in step 701, it can compete with other control nodes for the right to update the lease information stored on the storage service node by accessing the storage service node. The control node that wins the right to update the lease information stored on the storage service node becomes the master node. The master node has scheduling authority over the storage cluster. Accordingly, if it wins the right to update the lease information stored on the storage service node, in step 702, the node state of the control node can be set to master node state, so that the control node obtains scheduling authority over the storage cluster. For a description of the scheduling authority over the storage cluster, please refer to the relevant content of the above system embodiment, which will not be repeated here.

[0080] In this embodiment, the control node can access the storage service node and compete with other control nodes for the right to update the lease information stored on the storage service node to become the master node. As the master node, it has scheduling authority over the data in the storage cluster, thus enabling master election among multiple control nodes. This master-competition method, compared to the distributed lock method, does not require waiting for the distributed lock to expire, which helps improve the efficiency of master-slave node switching.

[0081] The following is combined with Figure 8 and Figure 9 The method for obtaining scheduling permissions provided in this embodiment will be described in detail.

[0082] Figure 8 This is a flowchart illustrating another scheduling permission acquisition method provided in an embodiment of this application. This method is applicable to management nodes, such as... Figure 8 As shown, the method mainly includes: 801. Obtain target lease information; the target lease information is the lease information stored locally on the control node or the first lease information stored on the storage service node.

[0083] 802. Provide a lease update request carrying the target lease information to the storage service node, so that the storage service node can update the second lease information if the target lease information is consistent with the second lease information currently stored by the storage service node, so as to obtain the updated lease information and return a lease update success message.

[0084] 803. Upon receiving a successful lease update message, determine that you have won the right to update the lease information stored on the storage service node.

[0085] 804. In response to the lease update success message, modify the lease information of the local storage to the updated lease information, and set the node status of the control node to the master node status to obtain scheduling permissions for the storage cluster.

[0086] Figure 9 This is a flowchart illustrating another scheduling permission acquisition method provided in an embodiment of this application. This method is applicable to storage service nodes. Figure 9 As shown, the method mainly includes: 901. Obtain the lease update request issued by the control node; the lease update request is issued by the control node to compete for the right to update the lease information stored by the storage service node.

[0087] 902. Obtain the target lease information carried in the lease renewal request from the lease renewal request.

[0088] 903. If the target lease information and the lease information currently stored by the storage service node are consistent, update the lease information stored by the storage service node so that the control node can obtain the permission to update the lease information stored by the storage service node.

[0089] 904. Return a lease update success message to the management node so that the management node can respond to the lease update success message by setting the node status to master node status and obtaining the scheduling authority of the storage cluster.

[0090] In this embodiment, when multiple control nodes compete for master status, for any control node, in step 801, the target lease information LV can be obtained. The target lease information can be lease information LV1 stored locally by the control node, or it can be lease information (defined as first lease information) LV2 stored by the obtained storage service node. Further, in step 802, a lease update request carrying the target lease information LV can be provided to the storage service node.

[0091] Specifically, the target lease information LV can be used as a precondition to generate an atomic lease update request that includes the precondition.

[0092] Optionally, a transaction mechanism can be used to generate an atomic lease update request containing the target lease information LV as a prerequisite.

[0093] In some embodiments, the lease update request may further include: lease information to be updated. Lease information to be updated refers to the content to which the lease information stored by the storage service node has been updated.

[0094] The target lease information mentioned above is determined by the current node status of the control node. In this embodiment, the node status includes: primary node status and backup node status. The control node in the primary node status has scheduling authority over the storage cluster, while the control node in the backup node status does not have scheduling authority over the storage cluster.

[0095] If the current node state of the control node is the master node state, then 0 can obtain the locally stored lease information LV1 as the target lease information LV. Accordingly, the locally stored lease information LV1 can be used as a precondition to generate an atomic lease update request containing the precondition, and the lease update request can be provided to the storage service node.

[0096] If the current node status of the control node is a standby node, the first lease information LV2 of the storage service node can be obtained. Furthermore, the obtained first lease information LV2 can be compared with the lease information LV1 of the control node's local storage. If the first lease information LV2 is inconsistent with the lease information of the control node's local storage, it indicates that another control node is continuously updating the lease information of the storage service node 20 as the master node, meaning the master node is continuously renewing leases. Therefore, if the first lease information is inconsistent with the lease information LV1 of the control node 10's local storage, the node status of the control node can be maintained as a standby node. A control node in standby node status has no scheduling authority over the storage cluster 30. Furthermore, the lease information of the local storage can be updated to the first lease information LV2 of the storage service node.

[0097] Accordingly, if the first lease information LV2 is consistent with the lease information LV1 stored locally by the control node, it indicates that no master node is updating the lease information stored by the storage service node. In this case, the control node 10, which is in standby node status, can compete for the right to update the lease information stored by the storage service node to become the master node. Specifically, if the first lease information LV2 is consistent with the lease information LV1 stored locally by the control node 10, the control node in standby node status can use the first lease information LV2 stored by the storage service node 20 as the target lease information LV.

[0098] Accordingly, the first lease information LV2 can be used as a precondition to generate an atomic lease update request containing that precondition.

[0099] Regardless of whether it is the primary node or the backup node, after generating the lease update request, in step 802, the lease update request carrying the target lease information can be provided to the storage service node.

[0100] For the storage service node, in step 901, a lease update request issued by the control node can be obtained, and in step 902, the target lease information can be obtained from the lease update request. Further, the lease information currently stored by the storage service node 20 (defined as the second lease information) LV3 can be compared with the target lease information LV to see if they are consistent.

[0101] The second lease information LV3 currently stored by the storage service node may be the same as or different from the first lease information LV2 read from the storage service node by the control node which is in standby node status. This depends on whether other control nodes update the lease information stored by the storage service node and become the master node.

[0102] If the second lease information LV3 currently stored by the storage service node is consistent with the target lease information LV, it indicates that no other controlling node has become the master node and is updating the lease information stored by the storage service node. Therefore, the controlling node acquires the permission to update the lease information stored by the storage service node. Accordingly, if the second lease information LV3 currently stored by the storage service node 20 is consistent with the target lease information LV, in step 903, the second lease information LV3 can be updated to obtain the updated lease information LV4.

[0103] After the lease information is successfully updated, in step 904, a lease update success message can be returned to the management node. For the management node, in step 803, upon receiving the lease update success message, it determines that it has won the right to update the lease information stored on the storage service node; and in step 804, the node state of the management node is set to the master node state, so that the management node can obtain the scheduling authority of the storage cluster.

[0104] In step 804, in response to the lease update success message, the lease information stored locally can be modified to the updated lease information LV4 in the aforementioned storage service node.

[0105] Accordingly, if the second lease information LV3 currently stored by the storage service node is inconsistent with the target lease information LV, it indicates that another control node has become the master node and is updating the lease information stored by the storage service node. In this case, the control node has not obtained the permission to update the lease information stored by the storage service node. Accordingly, for the storage service node, if the second lease information LV3 currently stored by the storage service node is inconsistent with the target lease information LV, a lease update failure message can be returned to the control node.

[0106] Accordingly, upon receiving a lease update failure message, the control node can determine that it has not won the right to update the lease information stored on the storage service node; and in response to the lease update failure message, it can set the node status of control node 10 to standby node status. A control node in standby node status has no scheduling authority over the storage cluster.

[0107] For the management node 10, which is in the master node state when sending a lease update request, if the management node (master node) successfully updates the lease information stored in the storage service node 20, it maintains the master node state; if the management node (master node) fails to update the lease information stored in the storage service node 20, it switches the master node state to the backup node state and enters the master-grabbing process in the backup node state.

[0108] For a management node in standby state, it can also update the locally stored lease information LV1 to the first lease information LV2 stored by the acquired storage service node in response to a lease update failure message.

[0109] In an implementation where a standby node successfully seizes the master node, the original master node may not be able to detect in time that another control node has successfully seized the master node. The original master node continues to schedule data for the storage cluster, leading to split-brain and dual-write problems, which damage the data consistency of the storage cluster.

[0110] To address the aforementioned issues, in this embodiment, input / output (IO) barrier information (defined as first IO barrier information) can be configured in the storage service node. The controlling node that acquires the right to update the IO barrier information stored on the storage service node has that right. The master node also stores IO barrier information (defined as second IO barrier information).

[0111] Based on the aforementioned I / O barrier information, when a control node in master node mode performs data scheduling on the storage cluster, it can use the second I / O barrier information in its local storage as a prerequisite to generate a data scheduling request for the storage cluster and provide this request to the storage service node. The storage service node can compare the second I / O barrier information carried in the data scheduling request with the first I / O barrier information in its local storage. If the second I / O barrier information matches the first I / O barrier information, it indicates that the control node has scheduling authority over the storage cluster. Accordingly, the storage service node can allow the control node in master node mode to perform data scheduling on the storage cluster. Conversely, the control node in master node mode can perform data scheduling on the storage cluster according to a preset data scheduling strategy.

[0112] For a control node in standby mode, if it wins the right to update the lease information stored on the storage service node, it can also request to modify the IO barrier information stored on the storage service node. Thus, when the original master node requests data scheduling for the storage cluster, the changed IO barrier information on the storage service node will cause inconsistencies between the original master node's IO barrier information and the storage service node's IO barrier information, preventing the original master node from scheduling data for the storage cluster. The original master node can promptly detect when another control node has successfully seized master status and switch its node state to standby mode, reducing the probability of a split-brain problem.

[0113] Optionally, if the second IO barrier information is inconsistent with the first IO barrier information, it indicates that the control node lacks scheduling authority over the storage cluster 30, and the storage service node can return a no-scheduling-authority prompt to the control node. Based on the no-scheduling-authority prompt, the control node can promptly detect that another control node has successfully seized the master position and switch its own node status to standby node status.

[0114] The following example illustrates the process by which a control node in standby mode modifies the IO barrier information of a storage service node after winning the right to update the lease information stored on the storage service node.

[0115] When sending a lease update request, the control node in standby mode can also obtain the IO barrier information FV1 stored by the storage service node in response to the lease update success message; and can generate an IO barrier update request based on the IO barrier information FV1, and provide the IO barrier update request to the storage service node. This IO barrier update request carries the IO barrier information FV1 stored by the storage service node.

[0116] Optionally, an atomic IO barrier update request containing the IO barrier information FV1 stored in the storage service node is generated, using this as a prerequisite. For a description of the prerequisite, please refer to the section on generating lease update requests above; it will not be repeated here.

[0117] Optionally, a transaction mechanism can be employed, using the IO barrier information FV1 as a precondition, to generate an atomic lease update request containing this precondition. That is, if the storage service node meets the precondition, the storage lease information is updated.

[0118] Based on the aforementioned IO barrier update request, the storage service node can obtain the IO barrier information FV1 carried in the IO barrier update request. Furthermore, it can compare whether the locally stored IO barrier information FV2 is consistent with the IO barrier information FV1 carried in the IO barrier update request. If the locally stored IO barrier information FV2 is consistent with the IO barrier information FV1 carried in the IO barrier update request, the locally stored IO barrier information FV2 is updated to obtain the updated IO barrier information FV3.

[0119] Because the standby node that successfully seized the master node updated the IO barrier information stored on the storage service node, the IO barrier information stored on the original master node became inconsistent with that stored on the storage service node. Thus, when the original master node requests data scheduling for the storage cluster, the storage service node detects that the IO barrier information carried in the scheduling request differs from the IO barrier information in its local storage, preventing the original master node from scheduling data. The storage service node can also return a "no scheduling permission" warning to the original master node if the IO barrier information in the scheduling request differs from the IO barrier information in its local storage. This allows the original master node to promptly recognize that another control node has successfully seized the master node and switch its node state to standby, thereby helping to reduce the probability of a split-brain problem.

[0120] For storage service nodes, after updating the IO barrier information FV2 in local storage, they can return an IO barrier information update success message to the management node. In response to this message, the management node can switch its node state to master node state, thereby gaining scheduling authority over the storage cluster. The management node can also, in response to the IO barrier information update success message, store the updated IO barrier information FV3 stored by the storage service node locally, for subsequent data scheduling of the storage cluster based on IO barrier information FV3 as a prerequisite. For details on the specific implementation of master node data scheduling of the storage cluster based on IO barrier information FV3, please refer to the relevant content above, which will not be repeated here.

[0121] In some embodiments, the IO barrier information FV2 currently stored locally by the storage service node may be inconsistent with the IO barrier information FV1 carried in the IO barrier update request. Accordingly, if the IO barrier information FV2 currently stored locally is inconsistent with the IO barrier information FV1 carried in the IO barrier update request, the storage service node may return an IO barrier update failure message to the management node.

[0122] For the control node, in response to the IO barrier update failure message, the node status of the control node 10 can be kept in the standby node state, and when the next master-grabbing cycle arrives, the above process of grabbing the master by accessing the storage service node can continue.

[0123] In other embodiments, after a standby node gains the right to update lease information, the storage service node may malfunction, or the network between the standby node and the storage service node may become abnormal, preventing the storage service node from subsequently updating the IO barrier information. In this situation, if the managing node that gained the right to update the lease information is changed to the master node, the original master node may not yet be aware that another managing node has switched to the master node because the IO barrier information stored on the storage service node has not been updated. The new master node and the original master node simultaneously perform data scheduling on the storage cluster, leading to a double-write problem and resulting in a split-brain scenario.

[0124] Therefore, to address this issue, for a control node in standby mode, if it does not receive a successful IO barrier information update message within a set time period, its node status can be maintained as a standby node, waiting for the next master-claiming cycle to arrive before continuing the process of accessing the storage service node to claim master. This reduces the probability of a split-brain scenario. The set time period can begin from the standby node that successfully claims the right to update the lease information sending an IO barrier information update message.

[0125] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 701 and 702 can be device A; or the execution subject of step 701 can be device A, and the execution subject of step 702 can be device B; and so on.

[0126] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 701, 702, etc., are merely used to distinguish different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0127] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause one or more processors to perform the steps in the above-described scheduling permission acquisition methods.

[0128] Figure 10 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 10 As shown, the computing device includes: a memory 100a, a processor 100b, and a communication component 100c. The memory 100a is used to store computer programs.

[0129] In some embodiments, the computing device may be implemented as a management node. Accordingly, the processor 100b is coupled to the memory 100a for executing a computer program to: compete with other management nodes for the right to update lease information stored on the storage service node by accessing the storage service node using the communication component 100c; the storage service node is used to ensure data consistency of the storage cluster; and, if the management node wins the right to update the lease information stored on the storage service node, it sets the node state of the management node to the master node state, so that the management node obtains scheduling authority over the storage cluster.

[0130] In some embodiments, when the processor 100b competes with other management nodes for the right to update lease information stored on the storage service node by accessing the storage service node, it is specifically configured to: obtain target lease information; the target lease information is lease information stored locally by the management node or the first lease information obtained from the storage service node; provide a lease update request carrying the target lease information to the storage service node through the communication component 100c, so that the storage service node can update the second lease information if the target lease information is consistent with the second lease information currently stored by the storage service node, so as to obtain the updated lease information and return a lease update success message; and, upon receiving the lease update success message, determine that it has won the right to update the lease information stored on the storage service node.

[0131] The processor 100c is also used to: modify the locally stored lease information to the updated lease information in response to a lease update success message.

[0132] Accordingly, if the target lease information and the lease information stored by the storage service node are inconsistent, the storage service node returns a lease update failure message to the control node. The processor 100b is also configured to: upon receiving a lease update failure message, determine that it has not won the right to update the lease information stored by the storage service node; in response to the lease update failure message, set the control node's node status to standby node status; the control node in standby node status has no scheduling authority over the storage cluster.

[0133] In some embodiments, when a lease update request carrying target lease information is provided to a storage service node, the management node is in a master node state. Accordingly, when obtaining the target lease information, the processor 100b is specifically configured to: obtain locally stored lease information as the target lease information. Accordingly, the processor 100b is also configured to: generate an atomic lease update request containing the locally stored lease information as a prerequisite.

[0134] In other embodiments, the storage service node stores first input / output I / O barrier information; the management node in the master node state stores second I / O barrier information. The processor 100b is further configured to: generate a data scheduling request for the storage cluster based on the second I / O barrier information; provide the data scheduling request to the storage service node via the communication component 100c, so that the storage service node allows the management node in the master node state to perform data scheduling on the storage cluster if the first I / O barrier information and the second I / O barrier information are consistent; and perform data scheduling on the storage cluster according to a preset data scheduling strategy.

[0135] In some other embodiments, when a lease update request carrying target lease information is provided to a storage service node, the control node is in standby node status. Accordingly, when the processor 100b obtains the target lease information, it is specifically configured to: obtain first lease information stored by the storage service node; and if the first lease information is consistent with the lease information stored locally by the control node, use the first lease information as the target lease information.

[0136] Furthermore, the processor 100b is also configured to: generate an atomic lease update request containing the first lease information as a prerequisite.

[0137] Accordingly, the processor 100b is also configured to: maintain the node status of the control node as a standby node when the first lease information is inconsistent with the lease information stored locally on the control node; and update the lease information stored locally to the first lease information; wherein the control node in the standby node status has no scheduling authority over the storage cluster.

[0138] In some embodiments, a storage service node stores input / output I / O barrier information; a control node that wins the right to update the lease information stored by the storage service node has the right to update the I / O barrier information stored by the storage service node. Accordingly, the processor 100b is further configured to: in response to a lease update success message, obtain the first I / O barrier information stored by the storage service node; generate an I / O barrier update request based on the first I / O barrier information stored by the storage service node; provide the I / O barrier update request to the storage service node through the communication component 100c, so that the storage service node updates the second I / O barrier information if the second I / O barrier information currently stored locally is consistent with the first I / O barrier information, and returns an I / O barrier information update success message to the control node; and, in response to the I / O barrier information update success message, switch the node state of the control node to the master node state, so that the control node obtains scheduling rights over the storage cluster.

[0139] Optionally, when the processor 100b generates an IO barrier update request based on the IO barrier information stored in the storage service node, it is specifically used to: generate an atomic IO barrier update request containing the IO barrier information stored in the storage service node as a prerequisite.

[0140] Accordingly, if the IO barrier information currently stored locally is inconsistent with the IO barrier information carried in the IO barrier update request, the storage service node returns an IO barrier information update failure message to the control node. The processor 100b is also configured to: in response to the IO barrier information update failure message, maintain the control node's node state in standby node state; or, if no IO barrier information update success message is received within a set time period, maintain the control node's node state in standby node state.

[0141] The computing device provided in this embodiment, when implemented as a control node, can become the master node by competing with other control nodes for the permission to update the lease information stored on the storage service nodes. As the master node, it has scheduling authority over the data in the storage cluster, thus enabling master election among multiple control nodes. This master-competition method, compared to the distributed lock method, does not require waiting for the distributed lock to expire, which helps improve the efficiency of master-slave node switching.

[0142] The computing device provided in this application embodiment can also be implemented as a storage service node. Accordingly, the processor 100b can be used to: obtain a lease update request issued by the management node through the communication component 100c; the lease update request is issued by the management node to compete for the right to update the lease information stored by the storage service node; obtain the target lease information carried in the lease update request from the lease update request; if the target lease information is consistent with the lease information currently stored by the storage service node, update the lease information stored by the storage service node so that the management node can obtain the right to update the lease information stored by the storage service node; and return a lease update success message to the management node so that the management node can respond to the lease update success message by setting the node state to the master node state to obtain the scheduling authority of the storage cluster.

[0143] Accordingly, processor 100b is also used to: if the target lease information and the lease information currently stored by the storage service node are inconsistent, return a lease update failure message to the management node so that the management node can set the node status to standby node status; the management node in standby node status has no scheduling authority over the storage cluster.

[0144] In some embodiments, a storage service node stores input / output I / O barrier information; a control node that wins the right to update the lease information stored by the storage service node has the right to update the I / O barrier information. The processor 100b is further configured to: obtain an I / O barrier update request issued by the control node via the communication component 100c; obtain the I / O barrier information carried in the I / O barrier update request; update the I / O barrier information stored in the storage service node if the currently stored I / O barrier information matches the I / O barrier information carried in the I / O barrier update request, to obtain the updated I / O barrier information; and return an I / O barrier information update success message to the control node via the communication component 100c, so that the control node can switch its node state to the master node state in response to the I / O barrier information update success message.

[0145] Accordingly, the processor 100b is also configured to: return an IO barrier information update failure message to the management node when the IO barrier information currently stored locally is inconsistent with the IO barrier information carried in the IO barrier update request, so that the management node can respond to the IO barrier information update failure message and keep the node state of the management node in the standby node state.

[0146] The computing device provided in this embodiment, when implemented as a storage service node, can respond to lease update requests from the management node. It determines whether the management node has won the right to update the lease information stored on the storage service node based on whether the lease information carried in the lease update request is consistent with the lease information stored locally. It then returns a lease update success message to the management node that successfully responded to the lease update request and updated the lease information. In this way, the management node can set its own node status to master to obtain scheduling rights over the storage cluster, enabling master contention among management nodes.

[0147] In some alternative implementations, such as Figure 10 As shown, the computing device may also include components such as a power supply component 100d. Figure 10 The diagram only shows some components and does not mean that the computing device must contain them. Figure 10 The inclusion of all components does not imply that a computing device can only include... Figure 10 The components shown.

[0148] In this embodiment, the memory is used to store computer programs and can be configured to store various other data to support operation on its host device. The processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Electrically Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0149] In the embodiments of this application, the processor can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), or a complex programmable logic device (CPLD); or it can be an advanced RISC machine (ARM) or a system on chip (SoC), etc., but is not limited thereto.

[0150] In this embodiment, the communication component is configured to facilitate wired or wireless communication between its host device and other devices. The device housing the communication component can access wireless networks based on communication standards, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In another exemplary embodiment, the communication component may also be implemented based on Near Field Communication (NFC), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), or other technologies.

[0151] In this embodiment, a power supply component is configured to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.

[0152] It should be noted that the terms "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0153] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code.

[0154] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0155] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0156] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0157] In a typical configuration, a computing device includes one or more processors (CPU, etc.), input / output interfaces, network interfaces, and memory.

[0158] Memory may include non-persistent storage in computer-readable media, such as random-access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0159] Computer storage media are readable storage media, also known as removable media. Removable and non-removable media can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient media, such as modulated data signals and carrier waves.

[0160] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.

[0161] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for obtaining scheduling permissions, applicable to management nodes, the method comprising: By accessing the storage service node, it competes with other control nodes for the right to update the lease information stored on the storage service node; The storage service node has the function of ensuring data consistency of the storage cluster; wherein, the storage service node stores first input / output IO barrier information; if it wins the right to update the lease information stored on the storage service node, the node state of the control node is set to the master node state, so that the control node can obtain the scheduling authority of the storage cluster. Wherein, if the control node is in the master node state, the control node stores second IO barrier information, and the method further includes: Using the second IO barrier information as a prerequisite, a data scheduling request is generated for the storage cluster; The data scheduling request is provided to the storage service node so that, if the first IO barrier information and the second IO barrier information are consistent, the storage service node allows the management node at the master node to perform data scheduling on the storage cluster. Data scheduling is performed on the storage cluster according to a preset data scheduling strategy.

2. The method according to claim 1, wherein the step of competing with other control nodes for the right to update the lease information stored on the storage service node by accessing the storage service node includes: Obtain target lease information; The target lease information is the lease information stored locally by the control node or the first lease information obtained from the storage service node; A lease update request carrying target lease information is provided to the storage service node, so that if the target lease information is consistent with the second lease information currently stored by the storage service node, the storage service node updates the second lease information to obtain the updated lease information and returns a lease update success message. Upon receiving the lease update success message, it is determined that the right to update the lease information stored on the storage service node has been won. The method further includes: In response to the lease update success message, the locally stored lease information is modified to the updated lease information.

3. The method according to claim 2, wherein if the target lease information and the lease information stored by the storage service node are inconsistent, the storage service node returns a lease update failure message to the control node; The method further includes: Upon receiving the lease update failure message, it is determined that the user did not win the right to update the lease information stored on the storage service node. In response to the lease update failure message, the node status of the control node is set to standby node status; the control node in standby node status has no scheduling authority over the storage cluster.

4. The method according to claim 2 or 3, wherein when providing a lease update request carrying target lease information to the storage service node, if the control node is in a master node state, the step of obtaining the target lease information includes: Obtain locally stored lease information as the target lease information; The method further includes: Using the locally stored lease information as a prerequisite, an atomic lease update request containing the prerequisite is generated.

5. The method according to claim 2 or 3, wherein when providing a lease update request carrying target lease information to the storage service node, if the control node is in standby node status, the step of obtaining the target lease information includes: Obtain the first lease information stored by the storage service node; If the first lease information is consistent with the lease information stored locally by the control node, the first lease information shall be used as the target lease information; The method further includes: Using the first lease information as a prerequisite, generate an atomic lease update request that includes the prerequisite.

6. The method according to claim 5, further comprising: If the first lease information is inconsistent with the lease information stored locally by the control node, the node status of the control node is kept as a standby node; and the lease information of the local storage is updated to the first lease information; wherein, the control node in the standby node status has no scheduling authority over the storage cluster.

7. The control node that wins the right to update the lease information stored on the storage service node according to claim 2 or 3 has the right to update the IO barrier information stored on the storage service node; The method further includes: In response to the lease update success message, obtain the first IO barrier information stored by the storage service node; Based on the first IO barrier information stored in the storage service node, an IO barrier update request is generated; The IO barrier update request is provided to the storage service node so that the storage service node can update the second IO barrier information if the second IO barrier information currently stored locally is consistent with the first IO barrier information, and return an IO barrier information update success message to the control node; In response to the successful update message of the IO barrier information, the node state of the management node is switched to the master node state, so that the management node can obtain scheduling authority over the storage cluster.

8. A method for obtaining scheduling permissions, applicable to a storage service node, wherein the storage service node has the function of ensuring data consistency of the storage cluster and stores lease information of the master node among multiple management nodes; The storage service node stores input / output I / O barrier information; The control node that wins the right to update the lease information stored on the storage service node has the right to update the IO barrier information; The method includes: Obtain the lease update request issued by the control node; the lease update request is issued by the control node in order to compete for the right to update the lease information stored by the storage service node; Obtain the target lease information carried in the lease renewal request from the lease renewal request; If the target lease information and the lease information currently stored by the storage service node are consistent, update the lease information stored by the storage service node so that the control node obtains the permission to update the lease information stored by the storage service node; The lease renewal success message is returned to the management node, so that the management node can respond to the lease renewal success message by setting the node status to the master node status and obtaining the scheduling authority of the storage cluster; The method further includes: Obtain the IO barrier update request issued by the control node; Obtain the IO barrier information carried in the IO barrier update request from the IO barrier update request; If the IO barrier information currently stored locally is consistent with the IO barrier information carried in the IO barrier update request, update the IO barrier information stored in the storage service node to obtain the updated IO barrier information. A message indicating successful IO barrier information update is returned to the control node, so that the control node can switch its node state to master node state in response to the message.

9. A data storage system, comprising: Multiple management nodes and storage service nodes in the storage cluster; The storage service node has the function of ensuring data consistency of the storage cluster and stores the lease information of the master node among the multiple management nodes; The master node has scheduling authority over the storage cluster; The control node is used to execute the steps in the method according to any one of claims 1-7; The storage service node is used to perform the steps in the method of claim 8.

10. A computing device, comprising: Memory, processor, and communication components; The memory is used to store computer programs; The processor is coupled to the memory and the communication component to execute the computer program for performing the method executed by the control node of claim 9, and / or the method executed by the storage service node of claim 9.

11. A computer-readable storage medium storing computer instructions that, when executed by one or more processors, cause the one or more processors to perform the method executed by the control node of claim 9, and / or the method executed by the storage service node of claim 9.

Citation Information

Patent Citations

  • Master selection method and device for distributed database system

    CN114860848A