Flow control method and device for distributed database system and electronic equipment
By using the scheduler in a distributed database system to determine the data migration task and adjust the time window capacity, the problem of insufficient traffic control effect in the prior art is solved, and more efficient traffic control and system stability are achieved.
Patent Information
- Application Number
- CN202510511454.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art has insufficient traffic control effect on task scheduling in large-scale distributed database clusters, making it difficult to effectively improve the stability of the system.
When the data migration conditions are met, the scheduler determines the data migration task and determines the first capacity of the time window based on the stored data amount of the target main shard. If the current available capacity is not less than the first capacity, the data migration task is sent. At the same time, the window capacity adjustment parameters are determined based on the data migration waiting time and execution time, and the capacity of the time window is adjusted to achieve flow control.
Through dynamic adjustment of the time window, traffic control of each node is achieved, node performance bottlenecks are avoided, task execution effect is improved, and the stability and resource utilization efficiency of distributed database systems are improved.
Smart Images

Figure CN120030001A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a flow control method, device and electronic equipment for a distributed database system. Background Art
[0002] In a large-scale distributed database cluster, there are multiple storage nodes. Data migration between different storage nodes is a common operation, which can be achieved by transferring data snapshots between storage nodes. Data migration consumes computing resources, storage resources, and traffic resources, so it needs to be properly managed. Among them, the storage node that performs data migration can be called a data migration node.
[0003] In a distributed database cluster, data migration tasks are one of the various tasks that require task scheduling, which includes the initiation and execution of task scheduling. Flow control of task scheduling helps improve the stability of large-scale distributed database clusters. However, the effect of flow control of task scheduling by related technologies is still insufficient and needs to be improved. Summary of the invention
[0004] The embodiments of the present application provide a flow control method, device, electronic device, computer-readable storage medium, and computer program product for a distributed database system. In order to achieve the purpose of improving the effect of flow control in task scheduling, the technical solutions provided by the embodiments of the present application are as follows: According to one aspect of an embodiment of the present application, a flow control method for a distributed database system is provided, wherein the distributed database system includes multiple nodes and a scheduling manager, at least some of the multiple nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; The method is performed by the scheduling manager, and the method includes: In response to satisfying the data migration condition, a data migration task is determined, and according to the amount of stored data of the target primary shard corresponding to the data migration task, a first capacity of the time window occupied by the source node where the target primary shard is located for data migration is determined; when the current available capacity of the time window corresponding to the source node is not less than the first capacity, the first capacity is subtracted from the current available capacity, and the data migration task is sent to the source node, so that the source node sends the data stored in the target primary shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; wherein the target primary shard is the primary shard corresponding to the slave shard in the node that satisfies the data migration condition; receiving a window capacity adjustment parameter sent by the source node, where the window capacity adjustment parameter is determined by the source node according to a data migration waiting time and a data migration execution time of the data migration task; The window capacity of the time window corresponding to the source node is adjusted according to the window capacity adjustment parameter.
[0005] According to another aspect of an embodiment of the present application, a flow control method for a distributed database system is provided, wherein the distributed database system includes multiple nodes and a scheduling manager, at least some of the multiple nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; The method is performed by a first node, where the first node is any node among the multiple nodes, and the method includes: In response to receiving the data migration task sent by the scheduling manager, according to the node identifier of the destination node carried in the data migration task, sending the data stored in the target primary shard to the destination node for synchronization; Determine the data migration waiting time and data migration execution time of the data migration task; Determining a window capacity adjustment parameter according to the data migration waiting time and the data migration execution time; Sending the window capacity adjustment parameter to the scheduling manager, so that the scheduling manager adjusts the window capacity of the time window corresponding to the first node according to the window capacity adjustment parameter; The data migration task is sent by the scheduling manager in the following manner: In response to the second node satisfying the data migration condition, determining a data migration task, wherein the target primary shard is the primary shard corresponding to the slave shard in the second node, and the source node receiving the data migration task is the first node; Determine, according to the amount of stored data of the target primary shard in the first node, a first capacity of a time window occupied by the first node for data migration; When the current available capacity of the time window corresponding to the first node is not less than the first capacity, the data migration task is sent to the first node.
[0006] Optionally, the method further comprises: In response to receiving a successful data snapshot reception response sent by the destination node, based on the node identifier of the node that meets the migration conditions included in the data migration task, a data deletion instruction is sent to the node that meets the data migration conditions, and the data deletion instruction includes the shard identifier of the target primary shard.
[0007] According to another aspect of an embodiment of the present application, a flow control method for a distributed database system is provided, wherein the distributed database system includes multiple nodes and a scheduling manager, at least some of the multiple nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; the method includes: The scheduling manager is used to determine the corresponding data migration task when it is determined that any node among the multiple nodes meets the data migration condition, and determine the first capacity of the time window occupied by the source node where the target primary shard is located for data migration according to the storage data volume of the target primary shard corresponding to the data migration task, and when the current available capacity of the time window corresponding to the source node is not less than the first capacity, subtract the first capacity from the current available capacity, and send the data migration task to the source node; wherein the source node of the data migration task is the node where the target primary shard corresponding to the slave shard in any node that meets the data migration condition is located; The source node is used to send the data stored in the target primary shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task when receiving the data migration task sent by the scheduling manager; The source node is further used to determine a data migration waiting time and a data migration execution time, determine a window capacity adjustment parameter according to the data migration waiting time and the data migration execution time, and send the window capacity adjustment parameter to the scheduling manager; The scheduling manager is further used to receive the window capacity adjustment parameter sent by the source node, and adjust the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter.
[0008] According to another aspect of an embodiment of the present application, a flow control device for a distributed database system is provided, wherein the distributed database system includes a plurality of nodes and a scheduling manager, wherein at least some of the plurality of nodes include a primary shard, and the primary shard of each node has a corresponding secondary shard in other nodes; the device is applied to the scheduling manager, and the device includes: A data migration module, configured to determine a data migration task in response to satisfying a data migration condition, determine a first capacity of a time window occupied by a source node where the target primary shard is located for data migration according to the amount of stored data of the target primary shard corresponding to the data migration task, and when the current available capacity of the time window corresponding to the source node is not less than the first capacity, subtract the first capacity from the current available capacity, and send the data migration task to the source node, so that the source node sends the data stored in the target primary shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; wherein the target primary shard is a primary shard corresponding to a slave shard in a node satisfying the data migration condition; A window capacity adjustment parameter receiving module, configured to receive a window capacity adjustment parameter sent by the source node, wherein the window capacity adjustment parameter is determined by the source node according to a data migration waiting time and a data migration execution time of the data migration task; The window capacity adjustment module is used to adjust the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter.
[0009] Optionally, the sending of the data in the target primary shard storage of the source node to the destination node for synchronization is determined by the source node through the following method: Generate a data snapshot of the target primary shard in the source node; Sending the data snapshot to the destination node, so that the destination node synchronizes the data stored in the target primary shard of the source node according to the data snapshot; The data migration waiting time and the data migration execution time of the data migration task are determined by the source node in the following manner: Recording the first moment of receiving the data migration task; Record the second time when the data snapshot of the target primary shard in the source node starts to be generated; Determine the time difference between the second moment and the first moment as the data migration waiting time; Recording a third time when a successful response to receiving the data snapshot sent by the destination node is received; The time difference between the third moment and the second moment is determined as the execution duration of the data migration.
[0010] Optionally, the window capacity adjustment parameter is a difference between the data migration waiting time and the data migration execution time; The window capacity adjustment module can be used to increase the window capacity of the time window corresponding to the source node when the window capacity adjustment parameter is a positive value; and to reduce the window capacity of the time window corresponding to the source node when the window capacity adjustment parameter is a negative value.
[0011] Optionally, the window capacity of the time window corresponding to each node of the multiple nodes includes capacities of at least two priorities; The current available capacity of the time window corresponding to the source node is determined by the data migration module in the following manner: determining the target priority of the data migration task; and taking the available capacity in the window capacity of the time window corresponding to the source node whose priority is not higher than the target priority as the current available capacity.
[0012] Optionally, the data migration module can be used to, if the target priority is higher than the lowest priority, subtract the capacity of each priority level lower than the target priority in order from low to high, until the total subtracted capacity is equal to the first capacity.
[0013] According to another aspect of an embodiment of the present application, a flow control device for a distributed database system is provided, the distributed database system comprising a plurality of nodes and a scheduling manager, at least some of the plurality of nodes comprising a primary shard, the primary shard of each node having a corresponding secondary shard in other nodes; the device is applied to the first node, the first node being any node among the plurality of nodes, the device comprising: A data sending module, configured to, in response to receiving a data migration task sent by the scheduling manager, send the data stored in the target primary shard to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; A duration determination module, used to determine the data migration waiting time and the data migration execution time of the data migration task; A window capacity adjustment parameter determination module, used to determine the window capacity adjustment parameter according to the data migration waiting time and the data migration execution time; a window capacity adjustment parameter sending module, configured to send the window capacity adjustment parameter to the scheduling manager, so that the scheduling manager adjusts the window capacity of the time window corresponding to the first node according to the window capacity adjustment parameter; The data migration task is sent by the scheduling manager in the following manner: In response to the second node satisfying the data migration condition, determining a data migration task, wherein the target primary shard is the primary shard corresponding to the slave shard in the second node, and the source node receiving the data migration task is the first node; Determine, according to the amount of stored data of the target primary shard in the first node, a first capacity of a time window occupied by the first node for data migration; When the current available capacity of the time window corresponding to the first node is not less than the first capacity, the data migration task is sent to the first node.
[0014] Optionally, the device further comprises: A data deletion module is used to respond to a successful response to receiving the data snapshot sent by the destination node, and based on the node identifier of the node that meets the migration conditions included in the data migration task, send a data deletion instruction to the node that meets the data migration conditions, wherein the data deletion instruction includes the shard identifier of the target primary shard.
[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method provided in any optional embodiment of the present application.
[0016] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in any optional embodiment of the present application are implemented.
[0017] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the method provided in any optional embodiment of the present application are implemented.
[0018] The beneficial effects of the technical solution provided by the embodiment of the present application are: dynamic flow control of each node is realized through the time window, so as to avoid the node from reaching the performance bottleneck and improve the task execution effect. Among them, the scheduling manager can judge whether the node is capable of executing the data migration task according to the current available capacity of the time window corresponding to the node and the capacity required for the node to execute the task, so as to realize flow control. The scheduling manager can also adjust the capacity of the time window corresponding to the node according to the window capacity adjustment parameter fed back by the node. Since the window capacity adjustment parameter is determined according to the data migration waiting time and the data migration execution time of the node, and the data migration waiting time and the data migration execution time can reflect the size of the task execution pressure of the node, therefore, adjusting the window capacity corresponding to the node based on the window capacity adjustment parameter can realize real-time and dynamic adjustment based on the node pressure state, which can effectively improve the flexibility of flow control, avoid the problem of waste of flow resources caused by excessive window capacity, and avoid the problem of slow task execution caused by too small window capacity, which leads to the problem of reduced stability of the distributed database system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0020] Figure 1 A schematic diagram of a distributed database system provided in an embodiment of the present application; Figure 2 A flow chart of a flow control method for a distributed database system provided in an embodiment of the present application; Figure 3 A schematic diagram of the interaction between the scheduling manager and each node provided in an embodiment of the present application; Figure 4 A schematic diagram of a data migration process provided in an embodiment of the present application; Figure 5 Another data migration process diagram provided in an embodiment of the present application; Figure 6a , Figure 6b , Figure 6c , Figure 6d , Figure 6e and Figure 6f A schematic diagram of an effect provided by an embodiment of the present application; Figure 7 A schematic diagram of a flow control device for a distributed database system provided in an embodiment of the present application; Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0022] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, such as "A and / or B" or "A, B" indicates that it is implemented as "A", or implemented as "B", or implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" can be implemented as parameter A including A1 or A2 or A3, and can also be implemented as parameter A including at least two of the three items A1, A2, A3.
[0023] In a distributed database cluster, the same data can have multiple different copies (stored in slave shards). When adding new data or bringing nodes online or offline, it is very common to migrate data between different nodes. In addition, initiating and executing scheduling are usually asynchronous. How to balance the generation capacity of initiating scheduling and the consumption capacity of executing scheduling is directly related to the execution speed of scheduling and the impact on the online business of the cluster. In order to avoid the slow generation of scheduling affecting the stability of the cluster, and to avoid the too fast generation of scheduling affecting the online business of the cluster, it is often necessary to control the generation speed of scheduling. The current solution usually limits the flow based on fixed-size values, which is less flexible and requires manual intervention. Or it only limits the storage space without fully considering the computing resources and traffic resources of the entire cluster, or limits the flow of data migration in and out, rather than the actual sending and receiving ends, resulting in poor flow control effect.
[0024] The embodiment of the present application provides a flow control method for a distributed database system, which can migrate the data of the node to other nodes when the node meets the data migration condition. For example, the data migration condition met by the node is that the node is in a busy state, and part of the data in the node can be migrated to other nodes, so that the subsequent scheduling manager can schedule other nodes to enable other nodes to perform tasks, thereby balancing the task pressure between each node and ensuring that no single node reaches a performance bottleneck. In addition, the embodiment of the present application performs flow control through the available capacity of the time window corresponding to the source node that actually sends data, and adjusts the capacity of the time window corresponding to the source node through the window capacity adjustment parameter fed back by the source node, avoiding manual intervention and improving the flexibility of flow control.
[0025] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the embodiment of the present application and the technical effects produced by the technical solution of the present application are explained below through the description of several exemplary embodiments. It should be pointed out that the following embodiments can refer to, learn from or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0026] Figure 1 A schematic diagram of the structure of a distributed database system applicable to the method provided in the embodiment of the present application is shown in FIG. Figure 1 As shown. The distributed database system includes multiple nodes and a scheduling manager, at least some of the multiple nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes. The data stored in the secondary shards is the backup data of the data stored in the primary shards. By setting the secondary shards, the load balancing and high availability of each node in the distributed database system can be achieved.
[0027] Optionally, in order to improve utilization, the data to be stored can be divided into sub-data of no less than the number of nodes in the distributed database system according to the number of nodes, the original data of a sub-data can be stored on a node, and the original data of the sub-data can be stored on at least some of the multiple nodes in the distributed database system, and the storage space storing the original data of the sub-data in the node can be called the primary shard of the sub-data, and the backup data of each sub-data can be stored on at least one other node, and the storage space storing the backup data of the sub-data in the at least one node is called the secondary shard of the sub-data. It can be understood that the primary shard and secondary shard corresponding to the same sub-data should be in different nodes so that the system has disaster recovery capabilities.
[0028] For convenience of description, for the same piece of data, the node where the primary shard corresponding to the data is located can be referred to as the primary node of the data in the following description, and the node where the secondary shard corresponding to the data is located can be referred to as the backup node or secondary node of the data. In Figure 1 In the example shown, the distributed database system includes nodes 1 to 4, the scheduling manager includes managers 1 to 3, and the client interacts with each node through Remote Procedure Call (RPC) to send data processing requests, such as data read and write requests, etc. Nodes in the system can communicate with each other. The scheduling manager is used for task scheduling of the distributed database system. The distributed database system can include only one scheduling manager or a manager cluster composed of multiple managers. The embodiments of the present application do not limit this. For ease of explanation, the embodiments of the present application will be described with one scheduling manager hereinafter.
[0029] In Figure 1 In order to facilitate the distinction between the primary shard and the secondary shard, the shard whose name only includes letters is the primary shard, and the shard whose name includes numbers is the secondary shard. For example, shard A in node 1 is the primary shard, and shard A1 in node 2 and shard A2 in node 4 are the secondary shards of shard A. Shard C in node 1 is the primary shard, and shard C1 in node 2 and shard C2 in node 3 are the secondary shards of shard C. Shard D in node 2 is the primary shard, and shard D1 in node 1 and shard D2 in node 4 are the secondary shards of shard D. Shard B in node 3 is the primary shard, and shard B1 in node 2 and shard B2 in node 4 are the secondary shards of shard B. Shard E in node 4 is the primary shard, and shard E1 in node 3 and shard E2 in node 1 are the secondary shards of shard E.
[0030] Optionally, for each primary shard, the primary shard and its secondary shards form a consensus group (RaftGroup), and the shards in a consensus group are managed through a consensus algorithm (Raft) to ensure the consistency of the data stored in each shard. For example, shard A, shard A1, and shard A2 form a Raft Group, and node 1 synchronizes data to shard A1 and shard A2 through shard A.
[0031] Figure 2 It is a schematic flowchart of the traffic control method for the distributed database system provided by the embodiments of the present application. Among them, this method is applied to the distributed database system, and this method includes S201 to S202, which are executed by the scheduling manager.
[0032] Figure 3 It is a schematic interaction diagram between the scheduling manager and each node provided by the embodiments of the present application, as Figure 3 shown. For ease of explanation, the embodiments of the present application are combined with Figure 3The interaction process between the middle scheduling manager and each node will be described for S201~S202.
[0033] Figure 3 It includes a scheduling manager, a source node, and a destination node. Among them, the source node refers to the node that receives the data migration task sent by the scheduling manager, and the destination node refers to the node that receives the data stored in the main shard sent by the source node and synchronizes it.
[0034] S201: In response to meeting the data migration condition, determine the data migration task. According to the storage data volume of the target main shard corresponding to the data migration task, determine the first capacity of the time window occupied by the source node where the target main shard is located for data migration. When the current available capacity of the time window corresponding to the source node is not less than the first capacity, subtract the first capacity from the current available capacity, and send the data migration task to the source node, so that the source node sends the data stored in the target main shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; among them, the target main shard is the main shard corresponding to the slave shard in the node that meets the data migration condition.
[0035] In a distributed database system, each node can send heartbeat information to the scheduling manager at a preset cycle. The heartbeat information can include the shard identifier of each shard in the node, the load situation of the node itself, etc., such as the usage rate of the Central Processing Unit (CPU), the usage rate of the storage space, etc. The scheduling manager receives the heartbeat information sent by each node and judges whether the node meets the data migration condition based on the load situation of the node itself in the heartbeat information. When the node meets the data migration condition, the scheduling manager can determine the data migration task. The data migration condition can be set as needed, and the embodiments of the present application do not limit this.
[0036] Optionally, the data migration condition includes one or more sub-conditions, and each sub-condition can include one or more specific conditions. Meeting the data migration condition can be meeting any sub-condition. The embodiments of the present application do not limit specifically what conditions the data migration condition includes. The data migration condition can be pre-configured in the scheduling manager according to actual application requirements. The scheduling manager can judge whether data migration is required in the system according to the data migration condition. For example, the sub-condition can be: the node is in a busy state, the node needs to release storage space, the node has a fault, the node is a hot node, etc. Among them, the hot node can be a node frequently accessed by clients.
[0037] Among them, whether the node needs to release storage space can be judged according to the storage usage rate of the node including the slave shard. For example, when the storage usage rate of the node including the slave shard reaches 95%, the node needs to release storage space. Similarly, whether the node including the slave shard is in a busy state can be judged by the tasks to be executed of the node including the slave shard. If the tasks to be executed of the node including the slave shard exceeds the preset task amount, the node is in a busy state.
[0038] When determining a data migration task, the scheduling manager may determine a destination node, wherein the destination node may be a node that does not meet the data migration conditions, and the embodiment of the present application does not limit this. The data migration task may include: the identification and indication information of the target node, etc. The indication information may include the shard identification of the target primary shard and the node identification of the node that meets the migration conditions.
[0039] The scheduling manager can also determine whether the node can perform the data migration task, and if the determination result is that the node can perform the data migration task, send the data migration task to the node. The embodiment of the present application does not limit the execution order of determining the data migration task and determining whether the node can perform the data migration task.
[0040] For ease of explanation, the node that meets the data migration conditions is called an intermediate node. It should be noted that the data that can be migrated is the data stored in the shards in the intermediate node. The embodiment of the present application does not limit which data stored in the shards in the intermediate node needs to be migrated. At least one data stored in the shards in the intermediate node can be migrated. For example, it can be determined based on the usage frequency of the data stored in the shards. For example, if the data migration condition that is met is the storage space release condition, the usage frequency of the data stored in each shard in the intermediate node can be determined and sorted, and the data stored in the shard with the lowest usage frequency can be selected for migration.
[0041] Because the slave shard and the master shard follow the raft mechanism, when the data stored in the slave shard needs to be migrated, the actual executor of the data migration operation is the node where the target master shard corresponding to the slave shard is located. Then, after the scheduler manager determines the target slave shard that needs data migration in the intermediate node, it can determine the node where the target master shard corresponding to the target slave shard is located according to the consensus group where the target slave shard is located, and determine the node where the target master shard is located as the source node. The source node is used to send the data stored in the target master shard corresponding to the slave shard that needs data migration to the destination node.
[0042] Then, the scheduling manager determines whether the object that can perform the data migration task is the source node. Specifically, the heartbeat information may also include the amount of stored data of the primary shard. The amount of stored data of the slave shard is the same as the amount of stored data of the primary shard. Therefore, the heartbeat information may be sent by the node where the primary shard corresponding to the slave shard is located. The scheduling manager may determine the first capacity of the time window occupied by the source node for data migration based on the amount of stored data of the primary shard in the source node. Specifically, the first capacity may be determined based on the ratio between the preset amount of stored data and the capacity of the occupied time window. For example, the ratio is 10 to 1, indicating that when the amount of stored data is 100MB, the first capacity is 10.
[0043] The time window refers to a flow control mechanism that counts and limits data migration tasks within a preset time period. When the set threshold is exceeded, the scheduler will take corresponding measures, such as rejecting new data migration tasks, delaying processing, or triggering an alarm mechanism.
[0044] Window capacity refers to the number of data migration tasks that can be executed concurrently within a preset time period. Available window capacity refers to the number of data migration tasks that can be executed concurrently by the node corresponding to the time window. Current available window capacity refers to the number of data migration tasks that can be executed concurrently by the node corresponding to the time window at the current moment.
[0045] When the current available window capacity is less than the window capacity, it means that at the current moment, the node corresponding to the time window is executing a certain amount of data migration tasks. It is understandable that when the node corresponding to the time window is executing a data migration task, the scheduling manager can deduct part of the current available window capacity, that is, subtract the first capacity from the current available capacity to indicate that part of the current available window capacity is occupied to reflect the number of data migration tasks being executed by the node corresponding to the time window. When the data migration task is completed, the scheduling manager can release part of the currently available window capacity previously occupied for the next execution of the data migration task.
[0046] The scheduling manager can also determine the current available capacity of the time window corresponding to the source node. After that, by comparing the current available capacity with the first capacity, it is determined whether the source node can execute the data migration task. Specifically, when the current available capacity is not less than the first capacity, it indicates that the source node has the ability to execute the data migration task. Therefore, the scheduling manager can send a data migration task to the source node, and the data migration task includes the node identifier of the destination node and the node identifiers of the nodes that meet the data migration conditions. If the current available capacity is less than the first capacity, it means that the source node is executing more data migration tasks, occupying more traffic resources, CPU resources, and hard disk I / O (Input / Output) resources, and cannot execute the data migration task. Then, the scheduling manager can re-determine the current available capacity according to a preset time period, and compare the current available capacity with the first capacity until the current available capacity is not less than the first capacity, and then send a data migration task to the first node (source node) so that the first node can send the data stored in the target primary shard to the destination node for backup according to the node identifier of the destination node carried in the data migration task. Among them, the first node is any one of the multiple nodes in the distributed database system.
[0047] Among them, the source node that receives the data migration task is the first node. The nodes that meet the data migration conditions are the second nodes (intermediate nodes). Then, the target primary shard is the primary shard corresponding to the secondary shard in the second node. Then, in response to the second node meeting the data migration conditions, the scheduling manager determines the data migration task, and determines the first capacity of the time window occupied by the first node for data migration according to the storage data volume of the target primary shard in the first node; when the current available capacity of the time window corresponding to the first node is not less than the first capacity, send the data migration task to the first node.
[0048] For the first node, in response to receiving the data migration task sent by the scheduling manager, according to the node identifier of the destination node carried in the data migration task, send the data stored in the target primary shard to the destination node so that the destination node can synchronize the data stored in the target primary shard. The first node can generate a file to be sent by means of data compression, generating a data snapshot, etc., and send the file to be sent to the destination node.
[0049] The destination node receives the data stored in the target primary shard sent by the first node and synchronizes it.
[0050] In the above data migration process, since the data stored in the target slave shard is the same as the data stored in the target primary shard, and the target slave shard and the target primary shard follow the raft mechanism, the node that actually sends the data stored in the shard is the source node, not the intermediate node. Therefore, the object for judging whether the node has the data migration capability is the source node where the target primary shard is located, so as to limit the flow of the actual sender of the shard storage data and improve the flow control effect.
[0051] In addition, the available capacity of the time window corresponding to the source node directly represents the pressure of the data migration task of the source node. When the data migration task pressure of the source node is large, the computing resources and traffic resources used are large, and the available capacity is also reduced. Therefore, the flow control effect is better by judging the available capacity of the time window corresponding to the source node. Avoid the problem of poor flow control effect caused by limiting the flow by judging the available storage space of the source node. This is because the available storage space of the source node is not necessarily related to the amount of computing resources and traffic resources currently used by the source node. For example, the source node is performing a large number of data migration tasks, and there is a lot of available storage space, but the computing resources and traffic resources used are large. If more data migration tasks are still scheduled for the source node, the efficiency of data migration task execution is low, and the source node becomes a hidden danger of bottleneck in the distributed database system. Therefore, the effect of limiting the flow by judging the available storage space of the source node is poor.
[0052] S202: Receive a window capacity adjustment parameter sent by the source node, where the window capacity adjustment parameter is determined by the source node based on the data migration waiting time and the data migration execution time of the data migration task; and adjust the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter.
[0053] In order to unify the names and facilitate description, the first node is called the source node.
[0054] Specifically, the source node can also determine the data migration waiting time and the data migration execution time of the data migration task; determine the window capacity adjustment parameters based on the data migration waiting time and the data migration execution time; and send the window capacity adjustment parameters to the scheduling manager, so that the scheduling manager adjusts the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameters.
[0055] Among them, the data migration waiting time can be the time consumed by the source node to generate the data sending file. For example, the source node obtains the data sending file by file compression, and the data migration waiting time can be the first starting time of receiving the data migration task sent by the scheduling manager, and the first ending time of compressing the data stored in the main shard to obtain the data sending file. The time interval between the first starting time and the first ending time is the data migration waiting time. The data migration execution time is the time consumed to send the data sending file to the destination node. The time when the data sending file can be sent is the second starting time, and the time when the destination node receives the data sending file is the second ending time. Since it is difficult for the source node to determine when the destination node receives the data sending file, the time when the data snapshot received successfully from the destination node can be received as the second ending time, and the time interval between the second starting time and the second ending time is the data migration execution time.
[0056] The source node can determine the difference between the data migration waiting time and the data migration execution time to obtain the window capacity adjustment parameter. It can also assign different weights to the data migration waiting time and the data migration execution time, and subtract them to obtain a numerical value as the window capacity adjustment parameter. The embodiment of the present application does not impose any restrictions on this. Then, the scheduling manager adjusts the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameters, and can realize the flow control of the source node. This is because the data migration waiting time and the data migration execution time reflect the pressure of the source node in the two stages of data migration. Generally speaking, if the data migration waiting time is too long, it means that the source node is in a state of high computing resource consumption. At this time, if the data migration execution time is short, the data migration task can be scheduled for the source node, and the data migration delay is short. If the data migration waiting time is short, it means that the source node is in a state of low computing resource consumption. At this time, if the data migration execution time is long, the data migration task will not be scheduled for the source node, avoiding the problem of long data migration delay of the source node.
[0057] The embodiment of the present application can dynamically adjust the window capacity of the time window corresponding to the source node through the window capacity adjustment parameters, for example, increasing or decreasing the available capacity of the time window corresponding to the source node, which can avoid the situation where the window capacity is set too small, resulting in the inability to fully utilize resources and causing waste of resources, and can also avoid the problem that the window capacity is set too large, causing excessive load to be added to the distributed database system, resulting in increased latency in the distributed database system, and can also avoid the problem that the window capacity cannot be flexibly adjusted to adapt to different scenarios due to the use of a fixed value window capacity, which requires manual intervention.
[0058] It should be noted that the source node is the node that sends the data stored in the target primary shard. After the data is successfully sent, the data stored in its own target primary shard is not deleted. In other words, for the source node, data migration actually refers to data sending, that is, sending the data stored in its target primary shard to other nodes for synchronization. The actual data migration is relative to the intermediate node, that is, the intermediate node may need to delete the data stored in the target slave shard.
[0059] Specifically, if the data migration condition is to meet the storage space release condition, after the data stored in the target primary shard is sent to the destination node for synchronization, the data stored in the original slave shard corresponding to the target primary shard in the node that meets the data migration condition is deleted.
[0060] In response to receiving a successful response to receiving a data snapshot sent by the destination node, the source node sends a data deletion instruction to the node that meets the data migration conditions based on the node identifier of the node that meets the migration conditions included in the data migration task, and the data deletion instruction includes the shard identifier of the target primary shard.
[0061] The node that meets the migration conditions determines the target slave shard in the same consensus group as the target primary shard based on the shard identifier of the target primary shard included in the data deletion instruction, and deletes the data stored in the target slave shard through the consensus group.
[0062] It should be noted that the nodes participating in data migration in the embodiment of the present application include an intermediate node (a node that meets the data migration conditions), a source node (a node where the target primary shard corresponding to the target slave shard is located) and a destination node. The intermediate node does not send the data stored in the target slave shard, but may delete the data stored in the target slave shard. The source node may send the data stored in the target primary shard to the destination node for synchronization, but does not delete the data stored in the target primary shard. The destination node receives the data stored in the target primary shard and performs synchronization. The three nodes jointly complete the data migration task.
[0063] Figure 4 A data migration process diagram provided in an embodiment of the present application is as follows: Figure 4 shown.
[0064] The scheduling manager sends the data migration task to the source node, and can also subtract the first capacity from the current available capacity, that is, the part of the available capacity of the current available capacity of the time window corresponding to the source node, the specific size of the occupied part of the available capacity is the size of the first capacity, to obtain the remaining available capacity. The source node receives and sends the data stored in the target primary shard to the destination node, and the destination node performs data synchronization. After that, the data snapshot reception success response is sent to the source node. After receiving the data snapshot, the source node sends a data deletion instruction to the intermediate node, and sends a data migration success instruction to the scheduling manager. The intermediate node receives the data deletion instruction and deletes the data stored in the target slave shard corresponding to the target primary shard based on the shard identifier of the target primary shard included in the data deletion instruction. The scheduling manager receives the data migration success instruction sent by the source node, and can adjust the remaining available capacity (update the available capacity). The current available capacity obtained is the window capacity of the first capacity size released, so that the remaining available capacity is changed to the current available capacity before occupation. For example, the current available capacity before occupation is 30, the remaining available capacity is 20, and the first capacity is 10. The scheduling manager receives the data migration success instruction sent by the source node and can adjust the remaining available capacity from 20 to 30, the adjustment amount is the first capacity, and obtains the current available capacity before occupation.
[0065] Since the data migration condition is to release the storage space, the storage space of the intermediate node can be released after the intermediate node deletes the data stored in the target slave shard. The data is deleted after the data stored in the target primary shard of the source node is sent to the destination node for synchronization in order to avoid the problem of deleting the data stored in the target slave shard when the data stored in the target primary shard of the source node fails to be successfully sent to the destination node for synchronization, which reduces the data redundancy capability of the distributed database system and affects the stability of the distributed database system.
[0066] If the data migration condition is satisfied except for the storage space release condition, the intermediate node may delete or not delete the data stored in the target slave shard, which can be set according to the needs.
[0067] With respect to step S201, the source node may also perform data migration by generating and sending a snapshot to achieve data synchronization. Specifically, the source node generates a data snapshot of the target primary shard in the source node; and sends the data snapshot to the destination node, so that the destination node synchronizes the data stored in the target primary shard of the source node according to the data snapshot.
[0068] In addition, the data migration waiting time and data migration execution time of the data migration task are determined by the source node in the following manner: recording the first moment when the data migration task is received; recording the second moment when the data snapshot of the target primary shard in the source node begins to be generated; determining the time difference between the second moment and the first moment as the data migration waiting time; recording the third moment when a successful response to the data snapshot received from the destination node is received; and determining the time difference between the third moment and the second moment as the data migration execution time.
[0069] Of course, the destination node may also carry in the data snapshot reception success response sent the fourth moment when the destination node receives the data stored in the target primary shard sent by the source node, then the source node may determine the time difference between the fourth moment and the second moment as the execution duration of the data migration.
[0070] Figure 5 Another data migration process diagram provided in the embodiment of the present application is as follows: Figure 5 shown.
[0071] The sending end is the source node, the migrating end is the intermediate node, and the migrating end is the destination node. The scheduling manager sends the schedule to the sending end, and the sending end generates a data snapshot of the target primary shard, and sends the data snapshot to the migrating end. The migrating end synchronizes the data stored in the target primary shard of the sending end according to the data snapshot. The sending end can also record the first moment of receiving the schedule sent by the scheduling manager, the second moment of starting to generate the data snapshot, and the third moment of receiving the successful response to the data snapshot sent by the migrating end. Based on the first moment, the second moment, and the third moment, the window capacity adjustment parameter for traffic feedback is determined, and the window capacity adjustment parameter is sent to the scheduling manager. The scheduling manager dynamically adjusts the available capacity of the time window of the sending end based on the window capacity adjustment parameter to achieve flexible control of traffic.
[0072] The source node may determine the time difference between the second moment and the first moment as the data migration waiting time, and determine the time difference between the third moment and the second moment as the data migration execution time. Based on the data migration waiting time and the data migration execution time, a window capacity adjustment parameter is determined, and the window capacity adjustment parameter may be the difference between the data migration waiting time and the data migration execution time. The window capacity adjustment parameter is specifically as follows:
[0073]
[0074]
[0075] in, is the window capacity tuning parameter, It is the sum of the data migration waiting time and the data migration execution time. The window capacity adjustment parameter can be any one of the expressions on the right side of the above equation.
[0076] Since the migrating end is not necessarily the sending end when the target migrates data from the shard, the embodiment of the present application directly performs flow control on the snapshot of generated data and the sending snapshot, that is, it performs flow control on the task scheduling of the data sending end, and since task scheduling is performed based on the feedback mechanism of the duration of each stage of data migration at the sending end, there is no need to impose additional restrictions on the migrating end.
[0077] If the window capacity adjustment parameter is the difference between the data migration waiting time and the data migration execution time, the scheduling manager adjusts the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter, including: when the window capacity adjustment parameter is a positive value, increasing the window capacity of the time window corresponding to the source node; when the window capacity adjustment parameter is a negative value, reducing the window capacity of the time window corresponding to the source node.
[0078] When the window capacity adjustment parameter is a positive value, it means that the waiting time for data migration is longer than the execution time for data migration, which means that most of the time is spent on data migration execution. The source node is capable of sending more snapshots, and the data migration delay is small. Therefore, by increasing the window capacity of the time window, the source node can execute more data migration tasks and improve the data processing efficiency of the distributed database system. When the window capacity adjustment parameter is a negative value, it means that the waiting time for data migration is shorter than the execution time for data migration, which means that the source node spends most of its time sending data during data migration, the source node is busy, and the data migration delay is long. The source node should not be allowed to receive more data migration tasks. By reducing the window capacity of the time window corresponding to the source node, the data migration task scheduling for the source node can be slowed down.
[0079] It is understandable that the flow feedback mechanism of the embodiment of the present application is determined based on the time when the data snapshot starts to be generated and sent, and the actual pressure condition of the source node can be inferred by the time when the snapshot is generated and sent. In other words, the flow feedback mechanism is determined based on the pressure of the source node processing tasks, making the flow control more accurate. The newly added time window is used to limit the number of concurrent tasks on the source node, that is, the total size of the sliding window is dynamically adjusted through the flow feedback of the currently sent snapshot. When the time window corresponding to the source node has no available capacity, it means that the source node is executing more concurrent data migration tasks, and task scheduling to the source node is stopped.
[0080] It should be noted that reducing or increasing the window capacity refers to reducing or increasing the available capacity of the time window.
[0081] With respect to step S201, different tasks have different degrees of importance. For example, some task scheduling is to break up heat spots, that is, by redistributing the load to balance the task pressure of each node, to avoid a single node becoming a system performance bottleneck. This kind of data migration task scheduling has a greater impact on the stability of the distributed database system and needs to be executed first. Some data migration task scheduling is only to free up some space, which has little impact on the stability of the distributed database system and can be postponed. Therefore, different data migration tasks may have different priorities, and the priority of the data migration task can be determined based on the data migration conditions satisfied by the node. The correspondence between the data migration conditions and the priority of the data migration task can be set as needed.
[0082] For example, when the data migration condition satisfied by the data migration task is that the intermediate node is in a busy state, the priority of the data migration task is set to a high priority, that is, the first priority; when the data migration condition satisfied by the data migration task is a storage space release condition, the priority of the data migration task is set to a low priority, that is, the second priority. When the data migration task simultaneously satisfies multiple sub-conditions in the data migration condition, such as simultaneously satisfying the node being busy and the node storage space release condition, the priority of the task corresponding to the sub-condition with a higher priority is selected as the priority setting of the data migration task. Continuing with the above example, the first priority can be selected as the priority of the data migration task.
[0083] In addition, the window capacity of the time window also has a priority. Specifically, for each of the multiple nodes, the window capacity of the time window corresponding to the node includes at least two priority capacities; the current available capacity of the time window corresponding to the source node is determined by: determining the target priority of the data migration task; and taking the available capacity of the window capacity of the time window corresponding to the source node whose priority is not higher than the target priority as the current available capacity.
[0084] If the priorities of the window capacity of the time window include low priority and high priority, generally speaking, the low priority window capacity is higher than the high priority window capacity. This is because non-urgent data migration tasks account for a high proportion in the distributed database system, so there are more low priority data migration tasks.
[0085] For example, the total window capacity of the time window is 100, of which the window capacity of the low priority is 80 and the window capacity of the high priority is 20.
[0086] Since the data migration task has a priority, and the window capacity of the time window also has a priority, therefore, in order to match the priority of the data migration task with the priority of the current available capacity of the time window, when the scheduling manager determines the current available capacity of the time window corresponding to the source node, it can first determine the target priority of the data migration task to determine the current available capacity that matches the target priority. The scheduling manager can use the available capacity of the window capacity of the time window corresponding to the source node whose priority is not higher than the target priority as the current available capacity. That is, the scheduling manager can use the available capacity of the window capacity of the time window corresponding to the source node whose priority is equal to the target priority as the current available capacity, and can also use the sum of the available capacity of the window capacity of the time window corresponding to the source node whose priority is lower than and equal to the target priority as the current available capacity.
[0087] For example, the priority (target priority) of the data migration task is the second priority, the first capacity is 20, and the priorities of the window capacity of the time window include the first priority, the second priority, and the third priority. The first priority is higher than the second priority, and the second priority is higher than the third priority. Among them, the available capacity of the first priority is 30, the available capacity of the second priority is 20, and the available capacity of the third priority is 10.
[0088] Then, the scheduling manager can use the available capacity of the second priority of 20 as the current available capacity of the time window corresponding to the source node for subsequent comparison with the first capacity. The available capacity of the first priority of 30 can also be used as the current available capacity of the time window corresponding to the source node, and the sum of the available capacity of the second priority and the available capacity of the first priority can also be used as the current available capacity of the time window corresponding to the source node.
[0089] In actual applications, high-priority data migration tasks are more important than low-priority data migration tasks and have a greater impact on the distributed database system. Therefore, high-priority data migration tasks need to be executed first. Setting different priorities for time windows can prevent low-priority data migration tasks from occupying all the current available capacity of the time window. When higher-priority data migration tasks arrive, they can only wait, resulting in a long delay for high-priority data migration tasks. In other words, setting different priorities for time windows can ensure that high-priority data migration tasks can be processed in a timely manner, reduce processing delays, and give priority to high-priority data migration tasks, which can prevent low-priority tasks from occupying too much traffic resources, thereby improving the resource utilization efficiency of the distributed database system.
[0090] In an embodiment of the present application, optionally, the first capacity is subtracted from the current available capacity, including: if the target priority is higher than the lowest priority, for the capacities of each priority level lower than the target priority level, the capacities of each priority level are subtracted in order from low to high, until the total subtracted capacity is equal to the first capacity.
[0091] It should be noted that the currently available capacity may include currently available capacity of different priorities, and the capacity with a priority not higher than the target priority may be selected for occupation. Specifically, the capacity with the lowest priority may be occupied first, which may increase the utilization rate of the window capacity and thereby improve the resource utilization rate of the distributed database system.
[0092] Using the above example, the priority (target priority) of the data migration task is 20 for the second priority first capacity, 30 for the first priority available capacity, 20 for the second priority available capacity, and 10 for the third priority available capacity. The scheduling manager can preferentially occupy 20 of the 30 available capacity of the first priority. If the first capacity is 40, 30 of the first priority available capacity and 10 of the second priority available capacity 20 can be occupied, and the total occupied available capacity is 40, that is, the first capacity.
[0093] Figure 6a , Figure 6b , Figure 6c , Figure 6d , Figure 6e and Figure 6f A schematic diagram of an effect provided by an embodiment of the present application, such as Figure 6a~6f shown.
[0094] Figure 6a , Figure 6c and Figure 6e Schematic diagram of the number of scheduling tasks generated at different times during capacity expansion, respectively, the default setting, the method provided in the embodiment of the present application, and the method without speed limit. The horizontal axis represents time and the vertical axis represents the number of operations per minute. Figure 6b , Figure 6d ,and Figure 6fThe delay diagram of the distributed database system during expansion is respectively the default setting, the method provided in the embodiment of the present application, and the method without speed limit. The delay occurs due to the data migration in the background and the real-time load in the foreground. The horizontal axis represents the type of delay, including P999 delay, P99 delay, P95 delay, and P80 delay, and the vertical axis represents the delay time. Including 3 nodes in the distributed database system, the distributed database system is expanded to 6 nodes and then reduced to 3 nodes. Compared with the default configuration, the speed of executing specific operations is increased by nearly 3 times. Compared with not performing flow control, the delay in executing specific operations is reduced by nearly half. Among them, specific operations may include any one or more of adding, deleting, modifying and checking data.
[0095] The present application embodiment provides a flow control device for a distributed database system, such as Figure 7 As shown, the distributed database system includes multiple nodes and a scheduling manager, at least some of the multiple nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; the device is applied to the scheduling manager, and the flow control device 70 of the distributed database system includes: a data migration module 701, a window capacity adjustment parameter receiving module 702 and a window capacity adjustment module 703, wherein: The data migration module 701 is used to determine a data migration task in response to satisfying a data migration condition, determine a first capacity of a time window occupied by a source node where the target primary shard is located for data migration according to the amount of stored data of the target primary shard corresponding to the data migration task, and when the current available capacity of the time window corresponding to the source node is not less than the first capacity, subtract the first capacity from the current available capacity, and send the data migration task to the source node, so that the source node sends the data stored in the target primary shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; wherein the target primary shard is a primary shard corresponding to a slave shard in a node satisfying the data migration condition; A window capacity adjustment parameter receiving module 702 is used to receive a window capacity adjustment parameter sent by the source node, where the window capacity adjustment parameter is determined by the source node according to the data migration waiting time and the data migration execution time of the data migration task; The window capacity adjustment module 703 is used to adjust the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter.
[0096] Optionally, the sending of the data in the target primary shard storage of the source node to the destination node for synchronization is determined by the source node through the following method: Generate a data snapshot of the target primary shard in the source node; Sending the data snapshot to the destination node, so that the destination node synchronizes the data stored in the target primary shard of the source node according to the data snapshot; The data migration waiting time and the data migration execution time of the data migration task are determined by the source node in the following manner: Recording the first moment of receiving the data migration task; Record the second time when the data snapshot of the target primary shard in the source node starts to be generated; Determine the time difference between the second moment and the first moment as the data migration waiting time; Recording a third time when a successful response to receiving the data snapshot sent by the destination node is received; The time difference between the third moment and the second moment is determined as the execution duration of the data migration.
[0097] Optionally, the window capacity adjustment parameter is a difference between the data migration waiting time and the data migration execution time; The window capacity adjustment module 703 may be configured to increase the window capacity of the time window corresponding to the source node when the window capacity adjustment parameter is a positive value, and to reduce the window capacity of the time window corresponding to the source node when the window capacity adjustment parameter is a negative value.
[0098] Optionally, the window capacity of the time window corresponding to each node of the multiple nodes includes capacities of at least two priorities; The current available capacity of the time window corresponding to the source node is determined by the data migration module 701 in the following manner: determining the target priority of the data migration task; and taking the available capacity in the window capacity of the time window corresponding to the source node whose priority is not higher than the target priority as the current available capacity.
[0099] Optionally, the data migration module 701 can be used to, if the target priority is higher than the lowest priority, subtract the capacity of each priority level lower than the target priority in order from low to high, until the total subtracted capacity is equal to the first capacity.
[0100] According to another aspect of an embodiment of the present application, a flow control device of a distributed database system is provided, wherein the distributed database system includes multiple nodes and a scheduling manager, at least some of the multiple nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; the device is applied to the first node, and the first node is any node among the multiple nodes, and the device includes: A data sending module, configured to, in response to receiving a data migration task sent by the scheduling manager, send the data stored in the target primary shard to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; A duration determination module, used to determine the data migration waiting time and the data migration execution time of the data migration task; A window capacity adjustment parameter determination module, used to determine the window capacity adjustment parameter according to the data migration waiting time and the data migration execution time; a window capacity adjustment parameter sending module, configured to send the window capacity adjustment parameter to the scheduling manager, so that the scheduling manager adjusts the window capacity of the time window corresponding to the first node according to the window capacity adjustment parameter; The data migration task is sent by the scheduling manager in the following manner: In response to the second node satisfying the data migration condition, determining a data migration task, wherein the target primary shard is the primary shard corresponding to the slave shard in the second node, and the source node receiving the data migration task is the first node; Determine, according to the amount of stored data of the target primary shard in the first node, a first capacity of a time window occupied by the first node for data migration; When the current available capacity of the time window corresponding to the first node is not less than the first capacity, the data migration task is sent to the first node.
[0101] Optionally, the device further comprises: A data deletion module is used to respond to a successful response to receiving the data snapshot sent by the destination node, and based on the node identifier of the node that meets the migration conditions included in the data migration task, send a data deletion instruction to the node that meets the data migration conditions, wherein the data deletion instruction includes the shard identifier of the target primary shard.
[0102] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar and has corresponding technical effects. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.
[0103] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, and the processor executes the above computer program to implement the steps of the method provided in any optional embodiment of the present application.
[0104] In an alternative embodiment, an electronic device is provided, such as Figure 8 As shown, Figure 8 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0105] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0106] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0107] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0108] The memory 4003 is used to store the computer program for executing the embodiment of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.
[0109] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0110] The embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0111] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage in these sub-steps or stages may also be executed at different times respectively. In different scenarios of execution time, the execution order of these sub-steps or stages may be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0112] The above is only an optional implementation method for some implementation scenarios of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the scheme of the embodiments of the present application, other similar implementation methods based on the technical ideas of the present disclosure are also within the protection scope of the embodiments of the present application.
Claims
1. A flow control method for a distributed database system, characterized in that: The distributed database system includes a plurality of nodes and a scheduling manager, at least some of the plurality of nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; The method is performed by the scheduling manager, and the method includes: In response to satisfying the data migration condition, a data migration task is determined, and according to the amount of stored data of the target primary shard corresponding to the data migration task, a first capacity of the time window occupied by the source node where the target primary shard is located for data migration is determined; when the current available capacity of the time window corresponding to the source node is not less than the first capacity, the first capacity is subtracted from the current available capacity, and the data migration task is sent to the source node, so that the source node sends the data stored in the target primary shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task; wherein the target primary shard is the primary shard corresponding to the slave shard in the node that satisfies the data migration condition; receiving a window capacity adjustment parameter sent by the source node, where the window capacity adjustment parameter is determined by the source node according to a data migration waiting time and a data migration execution time of the data migration task; The window capacity of the time window corresponding to the source node is adjusted according to the window capacity adjustment parameter.
2. The method according to claim 1, characterized in that The sending the data stored in the target primary shard of the source node to the destination node for synchronization includes: Generate a data snapshot of the target primary shard in the source node; Sending the data snapshot to the destination node, so that the destination node synchronizes the data stored in the target primary shard of the source node according to the data snapshot; The data migration waiting time and the data migration execution time of the data migration task are determined by the source node in the following manner: Recording the first moment of receiving the data migration task; Record the second time when the data snapshot of the target primary shard in the source node starts to be generated; Determine the time difference between the second moment and the first moment as the data migration waiting time; Recording a third time when a successful response to receiving the data snapshot sent by the destination node is received; The time difference between the third moment and the second moment is determined as the execution duration of the data migration.
3. The method according to claim 1 or 2, characterized in that: The window capacity adjustment parameter is the difference between the data migration waiting time and the data migration execution time; The step of adjusting the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter includes: When the window capacity adjustment parameter is a positive value, increasing the window capacity of the time window corresponding to the source node; When the window capacity adjustment parameter is a negative value, the window capacity of the time window corresponding to the source node is reduced.
4. The method according to claim 1, characterized in that: The window capacity of the time window corresponding to each node of the plurality of nodes includes capacities of at least two priorities; The current available capacity of the time window corresponding to the source node is determined in the following manner: Determining a target priority of the data migration task; The available capacity whose priority is not higher than the target priority in the window capacity of the time window corresponding to the source node is used as the current available capacity.
5. The method according to claim 4, characterized in that The step of subtracting the first capacity from the current available capacity includes: If the target priority is higher than the lowest priority, the capacity of each priority lower than the target priority is subtracted in order of the corresponding priority from low to high until the total subtracted capacity is equal to the first capacity.
6. The method according to claim 1, characterized in that If the data migration condition is to meet the storage space release condition, after the data stored in the target primary shard is sent to the destination node for synchronization, the data stored in the original slave shard corresponding to the target primary shard in the node that meets the data migration condition is deleted.
7. A flow control method for a distributed database system, characterized in that: The distributed database system includes a plurality of nodes and a scheduling manager, at least some of the plurality of nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; The method is performed by a first node, where the first node is any node among the multiple nodes, and the method includes: In response to receiving the data migration task sent by the scheduling manager, according to the node identifier of the destination node carried in the data migration task, sending the data stored in the target primary shard to the destination node for synchronization; Determine the data migration waiting time and data migration execution time of the data migration task; Determining a window capacity adjustment parameter according to the data migration waiting time and the data migration execution time; Sending the window capacity adjustment parameter to the scheduling manager, so that the scheduling manager adjusts the window capacity of the time window corresponding to the first node according to the window capacity adjustment parameter; The data migration task is sent by the scheduling manager in the following manner: In response to the second node satisfying the data migration condition, determining a data migration task, wherein the target primary shard is the primary shard corresponding to the slave shard in the second node, and the source node receiving the data migration task is the first node; Determine, according to the amount of stored data of the target primary shard in the first node, a first capacity of a time window occupied by the first node for data migration; When the current available capacity of the time window corresponding to the first node is not less than the first capacity, the data migration task is sent to the first node.
8. A flow control method for a distributed database system, characterized in that: The distributed database system includes a plurality of nodes and a scheduling manager, at least some of the plurality of nodes include primary shards, and the primary shard of each node has corresponding secondary shards in other nodes; The method comprises: The scheduling manager is used to determine the corresponding data migration task when it is determined that any node among the multiple nodes meets the data migration condition, and determine the first capacity of the time window occupied by the source node where the target primary shard is located for data migration according to the storage data volume of the target primary shard corresponding to the data migration task, and when the current available capacity of the time window corresponding to the source node is not less than the first capacity, subtract the first capacity from the current available capacity, and send the data migration task to the source node; wherein the source node of the data migration task is the node where the target primary shard corresponding to the slave shard in any node that meets the data migration condition is located; The source node is used to send the data stored in the target primary shard of the source node to the destination node for synchronization according to the node identifier of the destination node carried in the data migration task when receiving the data migration task sent by the scheduling manager; The source node is further used to determine a data migration waiting time and a data migration execution time, determine a window capacity adjustment parameter according to the data migration waiting time and the data migration execution time, and send the window capacity adjustment parameter to the scheduling manager; The scheduling manager is further used to receive the window capacity adjustment parameter sent by the source node, and adjust the window capacity of the time window corresponding to the source node according to the window capacity adjustment parameter.
9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Metadata migration method and metadata migration device in distributed system
CN106034080A
Database cluster data migration method and device and electronic equipment
CN111324596A
Adaptive hierarchical storage method based on time sliding window
CN111858469A
Data processing method, device and system for distributed database
CN113392067A
Data migration method and related device
CN115480711A