Time series data fragment master copy selection method and device, equipment, medium and product

By mapping cluster nodes and shards into traffic network vertices, and optimizing the main replica distribution using the minimum cost maximum flow algorithm, the problem of unbalanced main replica distribution in the existing technology is solved, and the calculation load balancing and system performance improvement is achieved.

CN120335930APending Publication Date: 2025-07-18TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510335683.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing time-sequence data sharding master-replica selection method is difficult to meet the requirement of balancing master-replica distribution in the industrial IoT production environment to carry intensive computing loads, and greedy selection algorithms can easily lead to local optimal solutions.

Method used

Map cluster nodes and shards into vertices of the traffic network, use the minimum cost maximum flow algorithm to solve the minimum cost maximum flow results in the traffic network, and optimize the main replica distribution.

Benefits of technology

The most balanced cluster master replica distribution is achieved, meeting the computing load requirements of the industrial IoT production environment, and improving the load balancing and performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335930A_ABST
    Figure CN120335930A_ABST
Patent Text Reader

Abstract

The invention provides a time sequence data fragment master copy selection method, device, equipment, medium and product, and the method comprises the steps: mapping a cluster node set and a cluster fragment set into vertexes of a flow network, adding a source point and a sink point into the flow network, and constructing the flow network; wherein the cluster fragment set comprises a plurality of time sequence data fragments, and each time sequence data fragment comprises a master copy and a slave copy; solving a minimum cost and maximum flow result in the flow network by using a minimum cost and maximum flow algorithm; the minimum-cost maximum-flow result comprises a node where each time sequence data fragment primary copy in the cluster fragment set is located. According to the scheme provided by the invention, the most balanced cluster master copy distribution is realized, so that the demand that an industrial Internet of Things production environment bears dense computing loads is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data management, and in particular, to a method, device, equipment, medium and product for selecting the master and replica of time series data shards. Background Art

[0002] A time series database cluster is an excellent solution for processing large amounts of time series data generated in industrial Internet of Things scenarios. The time series data generated in the production environment of industrial Internet of Things has the characteristics of intensive computing load. The hot spots of time series data are concentrated in the latest time partition. Taking the real data set of a certain iron and steel engineering technology company as an example, among more than three million sampling devices deployed in the production environment, at least 70% sample once per second, which means that the write throughput of the production environment reaches at least two million points per second.

[0003] In a time series database cluster, the master replica of a time series data shard usually processes write requests from the production environment and query requests initiated by users. Therefore, a balanced distribution of master replicas within the cluster helps to achieve computational balance in the cluster. Each replica of a time series data shard generally guarantees data consistency based on a consensus protocol. Many consensus protocols are equipped with master replica election rules. However, the election process is random. Usually, the replica that initiates the vote earlier has a greater probability of being elected as the master replica, which may lead to an unbalanced distribution of master replicas in the cluster. On the contrary, actively selecting the master replica for each time series data shard can generate a more balanced distribution of master replicas, but an intuitive greedy selection algorithm is likely to make the generated distribution fall into a local optimal solution.

[0004] In some exemplary technologies, the greedy selection algorithm is strictly followed, that is, for each shard, the replica located on the node with the least number of master replicas is selected as its master replica. This method can only consider a part of the nodes in the cluster at each step, lacking a global perspective, and the resulting distribution of master replicas is always unbalanced.

[0005] Therefore, the existing methods for selecting the master and replica of time series data shards are difficult to meet the requirements of the industrial Internet of Things production environment for a balanced distribution of master replicas to carry intensive computing loads. Summary of the Invention

[0006] The present invention provides a method, device, equipment, medium and product for selecting the master and replica of time series data shards, aiming to solve the defect that the existing master replica election rules and intuitive master replica selection algorithms are difficult to meet the requirements of the industrial Internet of Things production environment for a balanced distribution of master replicas to carry intensive computing loads, and to achieve the most balanced distribution of master replicas in the cluster, thereby meeting the needs of the industrial Internet of Things production environment to carry intensive computing loads.

[0007] The present invention provides a method for selecting the master and replica of time series data shards, including: Map the set of cluster nodes and the set of cluster shards to the vertices of a traffic network, and add a source node and a sink node to the traffic network to construct the traffic network; wherein, the set of cluster shards includes multiple time-series data shards, and each time-series data shard includes a primary replica and a secondary replica. Apply the minimum-cost maximum-flow algorithm to solve the minimum-cost maximum-flow result in the traffic network; wherein, the minimum-cost maximum-flow result includes the nodes where the primary replicas of each time-series data shard in the set of cluster shards are located.

[0008] According to a method for selecting primary replicas of time-series data shards provided by the present invention, the mapping of the set of cluster nodes and the set of cluster shards to the vertices of a traffic network, and adding a source node and a sink node to the traffic network to construct the traffic network includes: Map multiple nodes in the set of cluster nodes and multiple time-series data shards in the set of cluster shards to node vertices and shard vertices in the traffic network respectively. For each shard vertex, connect an edge from the source node to the shard vertex. For each node vertex and each shard vertex, if the node corresponding to the node vertex holds a replica of the shard corresponding to the shard vertex, then connect an edge from the shard vertex to the node vertex. For each node vertex, connect multiple edges from the node vertex to the sink node. Set the capacity and cost of each edge in the traffic network according to the number of primary replicas held by each node in the set of cluster nodes.

[0009] According to a method for selecting primary replicas of time-series data shards provided by the present invention, the setting of the capacity and cost of each edge in the traffic network according to the number of primary replicas held by each node in the set of cluster nodes includes: For the edge from the source node to the shard vertex, set the capacity to 1 and the cost to 0. For the edge from the shard vertex to the node vertex, set the capacity to 1 and the cost to 0. For the edge from the node vertex to the sink node, set the capacity of each edge to 1, and the cost increases with the set serial number of the edge.

[0010] According to a method for selecting primary replicas of time-series data shards provided by the present invention, the application of the minimum-cost maximum-flow algorithm to solve the minimum-cost maximum-flow result in the traffic network includes: Use the shortest path algorithm to find the augmenting path with the minimum cost from the source node to the sink node. Increase the flow along the minimum augmenting path, and update the flow and cost of the edges in the traffic network. Determine whether there is still an augmenting path. If there is, return to execute the step of using the shortest path algorithm to find the augmenting path with the minimum cost from the source point to the sink point; otherwise, use the current flow allocation scheme as the minimum cost maximum flow result.

[0011] According to a method for selecting the master replica of time series data shards provided by the present invention, after solving the minimum cost maximum flow result in the flow network by using the minimum cost maximum flow algorithm, the method further includes: For all edges connecting node vertices and shard vertices, if there is traffic passing through in the minimum cost maximum flow result, allocate each master replica of the time series data shards in the cluster shard set to the corresponding node.

[0012] According to a method for selecting the master replica of time series data shards provided by the present invention, the method further includes: Select the required hardware configuration according to the requirements analysis, deploy the server and configure the database software to form a set of cluster nodes; Select the required sharding strategy according to the data characteristics, determine the shard size and quantity to form a set of cluster shards.

[0013] The present invention also provides a device for selecting the master replica of time series data shards, including the following modules: A construction module, configured to map the set of cluster nodes and the set of cluster shards to the vertices of the flow network, and add a source point and a sink point to the flow network to construct the flow network; wherein, the set of cluster shards includes multiple time series data shards, and each time series data shard includes a master replica and a slave replica; A solving module, configured to use the minimum cost maximum flow algorithm to solve the minimum cost maximum flow result in the flow network; wherein, the minimum cost maximum flow result includes the nodes where each master replica of the time series data shards in the set of cluster shards is located.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it implements the method for selecting the master replica of time series data shards as described in any one of the above.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for selecting the master replica of time series data shards as described in any one of the above.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for selecting the master replica of time series data shards as described in any one of the above.

[0017] The method, device, equipment, medium and product for selecting the master replica of time-series data sharding provided by the present invention transform the complex master replica selection problem into a graph theory problem by mapping cluster nodes and shards to the vertices of a traffic network, which is convenient for solving using subsequent network flow algorithms. The construction of the traffic network provides the basis for the subsequent minimum-cost maximum-flow algorithm, enabling the algorithm to find the optimal solution globally and avoiding the problem of local optimal solutions. Further, the minimum-cost maximum-flow algorithm can be used to find the maximum flow allocation scheme from the source node to the sink node while ensuring the minimum total cost, ensuring the global optimality of the master replica distribution, that is, solving the master replica distribution of time-series data sharding that satisfies computational balance. Generally speaking, the solution of the present application realizes the most balanced cluster master replica distribution, and further meets the requirements of the industrial Internet of Things production environment to carry dense computing loads. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of the method for selecting the master replica of time-series data sharding provided by the present invention.

[0020] Figure 2 It is a schematic diagram for explaining the demand for cluster computing load balancing provided by the present invention.

[0021] Figure 3 It is an example schematic diagram of the cost flow selection algorithm provided by the present invention.

[0022] Figure 4 It is a schematic structural diagram of the device for selecting the master replica of time-series data sharding provided by the present invention.

[0023] Figure 5 It is a schematic structural diagram of the electronic equipment provided by the present invention. Detailed Embodiments

[0024] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0025] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0026] The terms "first", "second", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, for example, they can be implemented in an order other than those given in the diagrams or descriptions of the embodiments of this application.

[0027] In addition, the terms "including" and "having" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a series of components is not necessarily limited to those components explicitly listed, but may include other components not explicitly listed or inherent to these products or devices. The term "module" as used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that can perform the functions associated with the element.

[0028] Time series database clusters are an excellent solution for processing large amounts of time series data generated in industrial IoT scenarios. Time series data generated in industrial IoT production environments are characterized by intensive computing loads. The hot spots of time series data are concentrated in the latest time partitions. Taking a real data set from a steel engineering technology company as an example, among the more than three million sampling devices deployed in the production environment, at least 70% sample once per second, which means that the write throughput of the production environment reaches at least two million points per second.

[0029] In a time series database cluster, the master replica of the time series data shard usually handles write requests from the production environment and query requests initiated by users, so a balanced distribution of master replicas within the cluster helps achieve balanced computing in the cluster. The replicas of the time series data shards generally ensure data consistency based on a consensus protocol. Many consensus protocols are equipped with master replica election rules. However, the election process is random. Usually, the replica that initiates voting earlier has a greater probability of being selected as the master replica, which may lead to an unbalanced distribution of master replicas in the cluster. On the contrary, actively selecting a master replica for each time series data shard can generate a more balanced distribution of master replicas, but the intuitive greedy selection algorithm easily causes the generated distribution to fall into a local optimal solution.

[0030] In some exemplary techniques, the greedy selection algorithm is strictly followed. For example, suppose the cluster consists of 4 nodes Composition, configuration replication factor , it is necessary to select a shard as the primary replica. Figure 2 It is a schematic diagram showing the requirements of load balancing for cluster computing provided by the present invention. The intuitive greedy algorithm will generate an unbalanced primary replica distribution as shown Figure 2 through the following process: Select a shard on node as the primary replica; select a shard on node as the primary replica; select a shard on node as the primary replica; select a shard on node as the primary replica.

[0031] The above steps strictly follow the greedy algorithm and the greedy selection algorithm, that is, for each shard, the replica located on the node with the least number of primary replicas is selected as its primary replica. And when selecting a primary replica for shard , whether the replica located on node or is selected, the finally generated primary replica distribution is unbalanced. This method can only consider a part of the nodes in the cluster at each step, lacking a global perspective, and the finally generated primary replica distribution is unbalanced.

[0032] Therefore, the existing method for selecting the primary replica of time-series data shards is difficult to meet the requirement of an industrial Internet of Things production environment for an evenly distributed primary replica to carry a dense computing load.

[0033] In view of the above technical problems, the present application provides a method for selecting the primary replica of time-series data shards.

[0034] The following will specifically describe the technical solution of the present application and how the technical solution of the present application solves the above technical problems in detail with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the method for selecting the primary replica of time-series data of the present invention in conjunction with Figures 1-3 Figure [Figure number not provided in the original text].

[0035] Figure 1 It is a schematic flowchart of the method for selecting the primary replica of time-series data provided by the present invention. As shown Figure 1 in the figure, the method includes step 101 to step 102.

[0036] Step 101: Map the cluster node set and the cluster shard set to the vertices of a traffic network, and add a source point and a sink point to the traffic network to construct a traffic network; wherein, the cluster shard set includes multiple time-series data shards, and each time-series data shard includes a primary replica and a secondary replica.

[0037] Step 102: Use the minimum-cost maximum-flow algorithm to solve for the minimum-cost maximum-flow result in the flow network; among them, the minimum-cost maximum-flow result includes the nodes where the primary replicas of each time-series data shard in the cluster shard set are located.

[0038] In practical applications, the execution entity of this time-series data shard primary replica selection method can be a time-series data shard primary replica selection device, and there are various implementation methods for the time-series data shard primary replica selection device. For example, it can be implemented through a computer program, such as an application software, etc.; or, for example, a chip, etc. It can also be implemented as a medium storing relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or, it can also be implemented through an entity device integrated or installed with relevant computer programs, such as a server, an intelligent device, etc.

[0039] In this implementation, first, the cluster node set and the cluster shard set are modeled into a flow network, and then the flow network is used as the input of the minimum-cost maximum-flow algorithm to solve for an equilibrium primary replica distribution.

[0040] Specifically, step 101 includes: mapping the cluster node set and the cluster shard set to the vertices of the flow network, and adding a source point and a sink point to the flow network to construct the flow network; among them, the cluster shard set includes multiple time-series data shards, and each time-series data shard includes a primary replica and a secondary replica.

[0041] Among them, the cluster node set refers to the set of all nodes in the cluster, and each node represents a physical or virtual server in the cluster. The cluster shard set refers to the set of all time-series data shards in the cluster, and each shard represents a logical division unit of the data.

[0042] Exemplarily, let represent the cluster node set , where let represent the node numbered in the cluster. Let represent the replication factor of the cluster, that is, the number of replicas included in each shard. Let R represent the cluster shard set, where , and let represent the time-series data shard numbered in the cluster, and is the set composed of the nodes where its replicas are located. Let represent the set of cluster primary replicas, where represents the node where the primary replica of the shard is located. Let represent the number of primary replicas held by the node .

[0043] For example, a distributed time series database cluster includes 4 nodes, namely node to node . Then the cluster node set N can be represented as . For another example, a distributed time series database cluster contains 4 time series data shards, namely shard to shard . Then the cluster shard set R can be represented as .

[0044] In this embodiment, a flow network refers to a directed graph that includes a node set, an edge set, capacity, and cost, and is used to simulate a flow allocation problem. The flow network is used to simulate and solve the flow allocation problem from a starting point to an ending point. For example, the minimum cost maximum flow problem, that is, minimizing the total cost on the premise of satisfying the maximum flow.

[0045] Specifically, the flow network is a directed graph . Among them, V is the node set, representing each point in the network. E is the edge set, and each edge represents a directed connection from node u to node v . Each edge has a non - negative capacity , representing the maximum flow that the edge can pass through. Each edge can also have a cost , representing the cost per unit flow passing through the edge.

[0046] In practical applications, the flow network usually contains two special nodes: the source point (Source) and the sink point (Sink). Among them, the source point is a special node in the flow network, usually represented by the symbol S . It represents the starting point of the flow, that is, the flow is injected into the network from the source point. Among them, the sink point is another special node in the flow network, usually represented by the symbol T . It represents the ending point of the flow, that is, the flow finally flows to the sink point through the edges in the network.

[0047] Optionally, in a possible implementation manner, step 101 above includes: Mapping multiple nodes in the cluster node set and multiple time series data shards in the cluster shard set to node vertices and shard vertices in the flow network respectively; For each shard vertex, connect an edge from the source point to the shard vertex; For each node vertex and each shard vertex, if the node corresponding to the node vertex holds a copy of the shard corresponding to the shard vertex, then connect an edge from the shard vertex to the node vertex; For each node vertex, connect multiple edges from the node vertex to the sink; Set the capacity and cost of each edge in the traffic network according to the number of primary replicas held by each node in the cluster node set.

[0048] Specifically, in one example, setting the capacity and cost of each edge in the traffic network according to the number of primary replicas held by each node in the cluster node set includes: For the edge from the source point to the shard vertex, set the capacity to 1 and the cost to 0; For the edge from the shard vertex to the node vertex, set the capacity to 1 and the cost to 0; For the edge from the node vertex to the sink, set the capacity of each edge to 1, and the cost increases with the set serial number of the edge.

[0049] Among them, the cluster node set N is a set containing all nodes in the cluster, and each node represents a physical or virtual server. The cluster shard set R is a set containing all time-series data shards in the cluster, and each shard represents a logical division unit of the data. Specifically, map each node to a node vertex in the traffic network , and map each time-series data shard to a shard vertex in the traffic network . It can be understood that abstracting the nodes and shards in the distributed system as vertices in the traffic network model can provide a basis for subsequent traffic allocation and optimization.

[0050] Furthermore, for each shard vertex , connect an edge from the source point S to the shard vertex , with a capacity of 1 and a cost of 0. Among them, a capacity of 1 means that each shard can only select one primary replica. A cost of 0 means that no additional cost is introduced in this step. That is, the edge from the source point to the shard vertex represents that each shard can select a primary replica node.

[0051] Furthermore, for each node vertex and each shard vertex , if the node corresponding to the node vertex holds a copy of the shard corresponding to the shard vertex, that is , then connect an edge from the shard vertex to the node vertex , with a capacity of 1 and a cost of 0.

[0052] It is understandable that if a certain node has already stored a copy of a certain shard (whether it is the primary copy or the secondary copy), an edge is connected in the traffic network. From the shard vertex to the node vertex The edge represents that a certain node can store the primary copy of a certain shard. The capacity is 1, indicating that the node can store the primary copy of the shard. The cost is 0, indicating that no additional cost is introduced in this step.

[0053] Furthermore, for each node vertex , from the node vertex to the sink T Multiple edges are connected, and the capacity of each edge is 1, and the cost increases with the set serial number of the edge. Exemplarily, the cost of the th edge is .

[0054] It is understandable that the edge from the node vertex to the sink T represents the number of primary copies that each node can store, and load balancing is achieved through cost constraints. The capacity of each edge is 1, indicating that each node can store one primary copy. The cost increases with the serial number of the edge. For example, the cost of the first edge is 1, the cost of the second edge is 3, and so on. This cost setting is used to constrain the distribution of the primary copies so that the primary copies are distributed as evenly as possible among the nodes.

[0055] In this embodiment, by mapping the cluster nodes and shards to the vertices of the traffic network, the complex problem of primary copy selection is transformed into a graph theory problem, which is convenient for solving using subsequent network flow algorithms. The construction of the traffic network provides a basis for the subsequent minimum cost maximum flow algorithm, enabling the algorithm to find the optimal solution globally and avoiding the problem of local optimal solutions. Furthermore, by setting the cost of the edges, the distribution of the primary copies tends to be more balanced, avoiding excessive load on some nodes and improving the overall performance and resource utilization rate of the system. The cost setting provides a clear optimization goal for the algorithm, that is, to minimize the total cost under the premise of meeting the maximum flow, so as to achieve the global optimum of the primary copy distribution.

[0056] Furthermore, step 102 includes: applying the minimum cost maximum flow algorithm to solve the minimum cost maximum flow result in the traffic network; where the minimum cost maximum flow result includes the nodes where the primary copies of each time series data shard in the cluster shard set are located.

[0057] Among them, the Minimum Cost Maximum Flow (MCMF) algorithm is a network flow algorithm used to find a flow allocation scheme in a given flow network such that the flow from the source to the sink reaches the maximum value while the total cost is minimized. It combines the two objectives of maximum flow and minimum cost and is widely applied in fields such as resource allocation, path planning, and load balancing.

[0058] Exemplarily, the execution steps of the minimum cost maximum flow algorithm are described as follows by way of example.

[0059] Step 1: Initialize the flow of all edges to 0 and initialize the total cost to 0.

[0060] Step 2: Use a shortest path algorithm (such as Bellman - Ford or Dijkstra) to find the minimum - cost augmenting path from the source node S to the sink node T. An augmenting path is a path from the source node S to the sink node T where each edge on the path has remaining capacity and the total cost of the path is minimized.

[0061] Step 3: Determine the minimum remaining capacity on the augmenting path (i.e., the minimum value among the remaining capacities of all edges on the path). Increase the flow along this path and update the flow values of each edge on the path. At the same time, update the remaining capacity and cost of each edge in the flow network.

[0062] Step 4: Calculate the cost of the augmenting path and accumulate it into the total cost to update the total cost.

[0063] Step 5: Repeat Steps 2 to 4 until no augmenting path can be found, and obtain the minimum cost maximum flow result.

[0064] Optionally, in a possible implementation manner, the above - mentioned Step 102 includes: Use a shortest path algorithm to find the minimum - cost augmenting path from the source to the sink; Increase the flow along the minimum augmenting path and update the flow and cost of the edges in the flow network; Judge whether there is still an augmenting path. If there is, return to execute the step of using the shortest path algorithm to find the minimum - cost augmenting path from the source to the sink; otherwise, use the current flow allocation scheme as the minimum cost maximum flow result.

[0065] It is understandable that by finding the augmenting path with the minimum cost, the traffic allocation can be optimized from a global perspective, avoiding local optimal solutions. Ensure that the traffic added each time is on the path with the lowest cost under the current network state, thus gradually approaching the global optimal solution. Each time an augmenting path is found, the path selection is dynamically adjusted according to the current network state (including traffic and cost). This enables adaptation to changes in traffic in the network and ensures that the path selected each time is optimal.

[0066] Furthermore, increase the traffic along the found augmenting path, gradually approaching the maximum traffic. The traffic increased each time is the minimum remaining capacity on the path, ensuring the feasibility of traffic allocation. Update the traffic and cost of each edge on the path to ensure the correctness of the network state. The update of the cost reflects the actual cost of traffic allocation and provides a basis for subsequent path selection. Each increase in traffic and update of cost for the augmenting path bring the network state closer to the optimal solution. Through multiple iterations, gradually approach the optimal solution of maximum traffic and minimum cost.

[0067] Furthermore, determine whether the algorithm continues to iterate by judging whether there is an augmenting path. When no augmenting path can be found anymore, it means that the traffic in the network has reached the maximum value and the algorithm terminates. When the algorithm terminates, the current traffic allocation scheme is the result of the minimum-cost maximum-flow. This result includes the traffic value of each edge and the total cost, ensuring maximum traffic and minimum cost. Through multiple iterations, the algorithm can find the globally optimal traffic allocation scheme, avoiding local optimal solutions. The final result is the optimal solution that meets the dual goals of maximum traffic and minimum cost.

[0068] In this embodiment, the minimum-cost maximum-flow algorithm can be used to find the maximum traffic allocation scheme from the source node to the sink node, while ensuring the minimum total cost, ensuring the global optimality of the master-replica distribution, that is, solving the master-replica distribution of time-series data shards that meets computational balance.

[0069] In addition, in a possible implementation, after the above step 102, the method further includes: For all edges connecting node vertices and shard vertices, if there is traffic passing through in the minimum-cost maximum-flow result, then assign each master-replica of the time-series data shards in the cluster shard set to the corresponding node.

[0070] Specifically, for all edges connecting node vertices and shard vertices if there is traffic passing through in the minimum-cost maximum-flow, then let , that is, specify the master-replica of shard as the replica located on node .

[0071] In this embodiment, the theoretical result of the minimum-cost maximum-flow algorithm is transformed into an actual primary replica allocation scheme. The location of the primary replica for each shard is clearly specified, providing specific guidance for the actual deployment of the distributed system. The minimum-cost maximum-flow algorithm optimizes the distribution of primary replicas through a cost function, making the primary replicas as evenly distributed as possible among the nodes. This step ensures that the optimization result is actually applied, improving the load balancing and performance of the system.

[0072] In addition, in a possible embodiment, the above method further includes: Select the required hardware configuration according to the requirements analysis, deploy the servers and configure the database software to form a set of cluster nodes; Select the required sharding strategy according to the data characteristics, determine the shard size and quantity to form a set of cluster shards.

[0073] In this embodiment, the required hardware configuration is selected according to the requirements analysis, the servers are deployed and the database software is configured to form a set of cluster nodes; the required sharding strategy is selected according to the data characteristics, and the shard size and quantity are determined to form a set of cluster shards. These steps provide the necessary inputs for the subsequent primary replica selection method, ensuring the rationality and efficiency of the system.

[0074] To further illustrate the time-series data sharding primary replica selection method provided by this application, the following is illustrated by a specific example. Figure 2 It is a schematic diagram for explaining the load balancing requirements of cluster computing provided by the present invention. Figure 3 It is an example schematic diagram of the cost flow selection algorithm provided by the present invention.

[0075] Specifically, taking Figure 2 as an example, first model it into a Figure 3 shown traffic network, and use the minimum-cost maximum-flow algorithm to solve the minimum-cost maximum-flow result in the traffic network.

[0076] Exemplarily, let represent the set of cluster nodes , where let represent the node numbered in the cluster. Let represent the replica factor of the cluster, that is, the number of replicas contained in each shard. Let R represent the set of cluster shards, where , let represent the time-series data shard numbered in the cluster, and , which is the set composed of the nodes where its replicas are located. Let represent the set of cluster primary replicas, where represents the shard The node where the primary copy is located. Let represent the number of primary copies held by node .

[0077] The specific steps are as follows: Step (1) Map the set of cluster nodes N and the set of cluster shards R to the vertices of a traffic network, and then add a source and a sink . In this example, construct the set of node vertices , the set of shard vertices , and add the source and the sink .

[0078] Step (2) For each shard vertex , connect an edge from the source to , with a capacity of 1 and a cost of 0. In this example, the edges corresponding to this step are as shown by Figure 3 the 4 leftmost edges.

[0079] Step (3) Enumerate each node vertex and each shard vertex . If , connect an edge from to , with a capacity of 1 and a cost of 0. In this example, the edges corresponding to this step are as shown by Figure 3 the 8 middle edges.

[0080] Step (4) Enumerate each node vertex . Starting from this vertex, connect several edges to the sink . Among them, the capacity of each edge is 1, and the cost of the th edge is . In this example, the edges corresponding to this step are as shown by Figure 3 the right - hand side edges. Among them, for each node vertex , starting from it, only 2 edges need to be connected to the sink because each node holds exactly 2 shard copies. For these 2 edges, the cost of the first edge is 1 and the cost of the second edge is 3.

[0081] Step (5) Use the minimum - cost maximum - flow algorithm to obtain a set of minimum - cost maximum - flows. In this example, use the traffic network generated in the above steps as the input of the minimum - cost maximum - flow algorithm Step (6) For all edges connecting node vertices and shard vertices , if there is flow passing through in the minimum - cost maximum - flow, then let , that is, specify the shard The primary copy is the copy located at the node above.

[0082] The cost flow selection algorithm meets the requirements of cluster computing balance because it can be proven that: In this example, the minimum cost is 4, the maximum flow is 4, and .

[0083] The method for selecting the primary copy of time series data sharding provided in this embodiment maps cluster nodes and shards to the vertices of a traffic network, transforming the complex problem of primary copy selection into a graph theory problem, which is convenient for solving using subsequent network flow algorithms. The construction of the traffic network provides a basis for the subsequent minimum cost maximum flow algorithm, enabling the algorithm to find the optimal solution globally and avoiding the problem of local optimal solutions. Further, by using the minimum cost maximum flow algorithm, a maximum flow allocation scheme from the source node to the sink node can be found while ensuring the minimum total cost, ensuring the global optimality of the primary copy distribution, that is, solving the primary copy distribution of time series data sharding that meets the computing balance. Generally speaking, the solution of this application realizes the most balanced cluster primary copy distribution, thereby meeting the needs of the industrial Internet of Things production environment to carry dense computing loads.

[0084] Next, the time series data sharding primary copy selection device provided by the present invention will be described. The time series data sharding primary copy selection device described below can be correspondingly referred to the time series data sharding primary copy selection method described above.

[0085] Figure 4 is a schematic structural diagram of the time series data sharding primary copy selection device provided by the present invention. As Figure 4 shown, the time series data sharding primary copy selection device includes: a construction module 41 and a solution module 42.

[0086] The construction module 41 is used to map the cluster node set and the cluster shard set to the vertices of a traffic network, and add a source node and a sink node to the traffic network to construct the traffic network; wherein, the cluster shard set includes multiple time series data shards, and each time series data shard includes a primary copy and a secondary copy.

[0087] The solution module 42 is used to use the minimum cost maximum flow algorithm to solve the minimum cost maximum flow result in the traffic network; wherein, the minimum cost maximum flow result includes the node where the primary copy of each time series data shard in the cluster shard set is located.

[0088] In this embodiment, first, the construction module 41 models the cluster node set and the cluster shard set into a traffic network, and then the solution module 42 uses the traffic network as the input of the minimum cost maximum flow algorithm to solve and obtain a balanced primary copy distribution.

[0089] Specifically, a construction module 41 is configured to map a set of cluster nodes and a set of cluster shards to vertices of a traffic network, and add a source point and a sink point to the traffic network to construct the traffic network; wherein, the set of cluster shards includes multiple time-series data shards, and each time-series data shard includes a primary replica and a secondary replica.

[0090] Wherein, the set of cluster nodes refers to the set of all nodes in the cluster, and each node represents a physical or virtual server in the cluster. The set of cluster shards refers to the set of all time-series data shards in the cluster, and each shard represents a logical partitioning unit of the data.

[0091] Exemplarily, let represent the set of cluster nodes , where let represent the node numbered in the cluster. Let represent the replication factor of the cluster, that is, the number of replicas included in each shard. Let R represent the set of cluster shards, where , let represent the time-series data shard numbered in the cluster, and is the set composed of the nodes where its replicas are located. Let represent the set of primary replicas of the cluster, where represents the node where the primary replica of the shard is located. Let represent the number of primary replicas held by the node .

[0092] In this embodiment, a traffic network (Flow Network) refers to a directed graph that includes a set of nodes, a set of edges, capacity, and cost, and is used to simulate a traffic allocation problem. The traffic network is used to simulate and solve the traffic allocation problem from a starting point to an ending point. For example, the minimum-cost maximum-flow problem, that is, minimizing the total cost on the premise of satisfying the maximum traffic.

[0093] Specifically, the traffic network is a directed graph . Wherein, V is the set of nodes, representing each point in the network. E is the set of edges, and each edge represents a directed connection from the node u to the node v . Each edge has a non-negative capacity , representing the maximum traffic that the edge can pass through. Each edge can also have a cost , representing the cost per unit traffic passing through the edge.

[0094] In practical applications, a flow network usually contains two special nodes: a source node and a sink node. Among them, the source node is a special node in the flow network, usually represented by the symbol S It represents the starting point of the flow, that is, the flow is injected into the network from the source node. Among them, the sink node is another special node in the flow network, usually represented by the symbol T It represents the end point of the flow, that is, the flow finally flows to the sink node through the edges in the network.

[0095] Optionally, in a possible implementation manner, the above-mentioned construction module 41 is specifically used for: Mapping multiple nodes in the cluster node set and multiple time-series data shards in the cluster shard set to node vertices and shard vertices in the flow network respectively; For each shard vertex, connect an edge from the source node to the shard vertex; For each node vertex and each shard vertex, if the node corresponding to the node vertex holds a copy of the shard corresponding to the shard vertex, then connect an edge from the shard vertex to the node vertex; For each node vertex, connect multiple edges from the node vertex to the sink node; Set the capacity and cost of each edge in the flow network according to the number of primary copies held by each node in the cluster node set.

[0096] Specifically, in an example, when the above-mentioned construction module 41 is used to set the capacity and cost of each edge in the flow network according to the number of primary copies held by each node in the cluster node set, it is specifically used for: For the edge from the source node to the shard vertex, set the capacity to 1 and the cost to 0; For the edge from the shard vertex to the node vertex, set the capacity to 1 and the cost to 0; For the edge from the node vertex to the sink node, set the capacity of each edge to 1, and the cost increases with the set serial number of the edge.

[0097] Among them, the cluster node set N is a set containing all nodes in the cluster, and each node represents a physical or virtual server. The cluster shard set R contains a set of all time-series data shards in the cluster, and each shard represents a logical division unit of the data. Specifically, each node is mapped to a node vertex in the flow network , and each time-series data shard is mapped to a shard vertex in the flow network It can be understood that abstracting the nodes and shards in a distributed system as vertices in a traffic network model can provide a basis for subsequent traffic allocation and optimization.

[0098] Furthermore, for each shard vertex , connect an edge from the source point S to the shard vertex . The capacity is 1 and the cost is 0. Here, a capacity of 1 means that each shard can only select one primary replica. A cost of 0 means that no additional cost is introduced in this step. That is, the edge from the source point to the shard vertex indicates that each shard can select a primary replica node.

[0099] Furthermore, for each node vertex and each shard vertex , if the node corresponding to the node vertex holds a replica of the shard corresponding to the shard vertex, that is , then connect an edge from the shard vertex to the node vertex . The capacity is 1 and the cost is 0.

[0100] It can be understood that if a certain node already stores a replica of a certain shard (whether it is a primary replica or a secondary replica), then an edge is connected in the traffic network. The edge from the shard vertex to the node vertex indicates that a certain node can store the primary replica of a certain shard. The capacity is 1, indicating that the node can store the primary replica of the shard. The cost is 0, indicating that no additional cost is introduced in this step.

[0101] Furthermore, for each node vertex , connect multiple edges from the node vertex to the sink T . Each edge has a capacity of 1, and the cost increases with the set serial number of the edge. Exemplarily, the cost of the th edge is .

[0102] It can be understood that the edge from the node vertex to the sink T indicates the number of primary replicas that each node can store, and load balancing is achieved through cost constraints. The capacity of each edge is 1, indicating that each node can store one primary replica. The cost increases with the serial number of the edge. For example, the cost of the first edge is 1, the cost of the second edge is 3, and so on. This cost setting is used to constrain the distribution of primary replicas so that the primary replicas are distributed as evenly as possible among the nodes.

[0103] In this embodiment, the construction module 41 transforms the complex primary replica selection problem into a graph theory problem by mapping cluster nodes and shards to the vertices of a traffic network, facilitating the solution using subsequent network flow algorithms. The construction of the traffic network provides the basis for the subsequent minimum-cost maximum-flow algorithm, enabling the algorithm to find the optimal solution globally and avoiding the problem of local optimal solutions. Further, by setting the cost of the edges, the distribution of primary replicas tends to be more balanced, avoiding excessive load on certain nodes and improving the overall performance and resource utilization of the system. The cost setting provides a clear optimization goal for the algorithm, that is, to minimize the total cost while satisfying the maximum flow, thereby achieving the global optimum of the primary replica distribution.

[0104] Further, the solution module 42 is configured to solve the minimum-cost maximum-flow result in the traffic network by using the minimum-cost maximum-flow algorithm; wherein, the minimum-cost maximum-flow result includes the nodes where the primary replicas of each time-series data shard in the cluster shard set are located.

[0105] Among them, the minimum-cost maximum-flow algorithm (MCMF) is a network flow algorithm used to find a traffic allocation scheme in a given traffic network, such that the traffic from the source point to the sink point reaches the maximum value while the total cost is minimized. It combines the two objectives of maximum flow and minimum cost and is widely applied in fields such as resource allocation, path planning, and load balancing.

[0106] Optionally, in a possible implementation manner, the above solution module 42 is specifically configured to: Use the shortest path algorithm to find the augmenting path with the minimum cost from the source point to the sink point; Increase the traffic along the minimum augmenting path, and update the traffic and cost of the edges in the traffic network; Determine whether there is still an augmenting path. If so, return to execute the step of using the shortest path algorithm to find the augmenting path with the minimum cost from the source point to the sink point; otherwise, take the current traffic allocation scheme as the minimum-cost maximum-flow result.

[0107] It can be understood that by finding the augmenting path with the minimum cost, the traffic allocation can be optimized from a global perspective, avoiding local optimal solutions. Ensure that the traffic increased each time is the path with the lowest cost under the current network state, thereby gradually approaching the global optimum. Each time an augmenting path is found, the path selection will be dynamically adjusted according to the current network state (including traffic and cost). This enables adaptation to changes in traffic in the network and ensures that the path selected each time is optimal.

[0108] Furthermore, increase the flow along the found augmenting path, gradually approaching the maximum flow. The flow increased each time is the minimum residual capacity on the path, ensuring the feasibility of flow allocation. Update the flow and cost of each edge on the path to ensure the correctness of the network state. The update of the cost reflects the actual cost of flow allocation and provides a basis for subsequent path selection. Each increase in the flow of the augmenting path and the update of the cost bring the network state closer to the optimal solution. Through multiple iterations, gradually approach the optimal solution of the maximum flow and minimum cost.

[0109] Furthermore, determine whether the algorithm continues to iterate by judging whether there is an augmenting path. When no augmenting path can be found anymore, it means that the flow in the network has reached the maximum value and the algorithm terminates. When the algorithm terminates, the current flow allocation scheme is the result of the minimum-cost maximum flow. This result includes the flow value of each edge and the total cost, ensuring maximum flow and minimum cost. Through multiple iterations, the algorithm can find the globally optimal flow allocation scheme and avoid local optimal solutions. The final result is the optimal solution that meets the dual objectives of maximum flow and minimum cost.

[0110] In this embodiment, the solving module 42 can find the maximum flow allocation scheme from the source node to the sink node by using the minimum-cost maximum flow algorithm, while ensuring the minimum total cost, and ensuring the global optimality of the main replica distribution, that is, solving the main replica distribution of the time-series data shards that meets the computing balance.

[0111] In addition, in a possible embodiment, the above device further includes: A processing module, configured to, for all edges connecting node vertices and shard vertices, if there is flow passing through in the minimum-cost maximum flow result, allocate each main replica of the time-series data shards in the cluster shard set to the corresponding node.

[0112] Specifically, for all edges connecting node vertices and shard vertices if there is flow passing through in the minimum-cost maximum flow, then let , that is, specify the main replica of the shard as the replica located on the node .

[0113] In this embodiment, the processing module converts the theoretical result of the minimum-cost maximum flow algorithm into an actual main replica allocation scheme. The location of the main replica of each shard is clearly specified, providing specific guidance for the actual deployment of the distributed system. The minimum-cost maximum flow algorithm optimizes the distribution of the main replicas through the cost function, making the main replicas evenly distributed to each node as much as possible. This step ensures that the optimization result is actually applied, improving the load balancing and performance of the system.

[0114] In addition, in a possible implementation, the above device further includes: A deployment module, configured to select required hardware configurations according to demand analysis, deploy servers, and configure database software to form a set of cluster nodes; A sharding module, configured to select required sharding strategies according to data characteristics, determine shard sizes and quantities, and form a set of cluster shards.

[0115] In this implementation, required hardware configurations are selected according to demand analysis, servers are deployed and database software is configured to form a set of cluster nodes; required sharding strategies are selected according to data characteristics, shard sizes and quantities are determined, and a set of cluster shards is formed. These steps provide necessary inputs for the subsequent master replica selection method, ensuring the rationality and efficiency of the system.

[0116] The time series data sharding master replica selection device provided in this embodiment converts the complex master replica selection problem into a graph theory problem by mapping cluster nodes and shards to vertices of a traffic network, facilitating the use of subsequent network flow algorithms for solution. The construction of the traffic network provides a basis for the subsequent minimum cost maximum flow algorithm, enabling the algorithm to find the optimal solution globally and avoiding the problem of local optimal solutions. Further, the solving module can find the maximum flow allocation scheme from the source point to the sink point using the minimum cost maximum flow algorithm while ensuring the minimum total cost, ensuring the global optimality of the master replica distribution, that is, solving the time series data sharding master replica distribution that satisfies computational balance. Generally speaking, the solution of this application realizes the most balanced cluster master replica distribution, thereby meeting the requirements of the industrial Internet of Things production environment to carry dense computing loads.

[0117] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute the time series data sharding master replica selection method, and the method includes: mapping the set of cluster nodes and the set of cluster shards to vertices of a traffic network, and adding a source point and a sink point to the traffic network to construct the traffic network; where the set of cluster shards includes multiple time series data shards, and each time series data shard includes a master replica and a slave replica. Using the minimum cost maximum flow algorithm, solve the minimum cost maximum flow result in the traffic network; where the minimum cost maximum flow result includes the nodes where the master replicas of each time series data shard in the set of cluster shards are located.

[0118] In addition, when the logical instructions in the memory 530 described above can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.

[0119] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the timing data shard master replica selection method provided by the above-mentioned various methods. The method includes: mapping the cluster node set and the cluster shard set to the vertices of a traffic network, and adding a source point and a sink point to the traffic network to construct the traffic network; where the cluster shard set includes multiple timing data shards, and each timing data shard includes a master replica and a slave replica. Using the minimum cost maximum flow algorithm, solve the minimum cost maximum flow result in the traffic network; where the minimum cost maximum flow result includes the nodes where the master replicas of each timing data shard in the cluster shard set are located.

[0120] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the timing data shard master replica selection method provided by the above-mentioned various methods. The method includes: mapping the cluster node set and the cluster shard set to the vertices of a traffic network, and adding a source point and a sink point to the traffic network to construct the traffic network; where the cluster shard set includes multiple timing data shards, and each timing data shard includes a master replica and a slave replica. Using the minimum cost maximum flow algorithm, solve the minimum cost maximum flow result in the traffic network; where the minimum cost maximum flow result includes the nodes where the master replicas of each timing data shard in the cluster shard set are located.

[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for selecting the main replica of time-series data shards, characterized in that, Including: Mapping a set of cluster nodes and a set of cluster shards to vertices of a traffic network, and adding a source point and a sink point to the traffic network to construct the traffic network; wherein, the set of cluster shards includes multiple time-series data shards, and each time-series data shard includes a primary replica and a secondary replica. Using the minimum-cost maximum-flow algorithm to solve the minimum-cost maximum-flow result in the traffic network; wherein, the minimum-cost maximum-flow result includes the nodes where the primary replicas of each time-series data shard in the set of cluster shards are located.

2. The method for selecting a main replica of time-series data shards according to claim 1, wherein The mapping of the set of cluster nodes and the set of cluster shards to vertices of the traffic network, and adding a source point and a sink point to the traffic network to construct the traffic network includes: Mapping multiple nodes in the set of cluster nodes and multiple time-series data shards in the set of cluster shards to node vertices and shard vertices in the traffic network respectively. For each shard vertex, connecting an edge from the source point to the shard vertex. For each node vertex and each shard vertex, if the node corresponding to the node vertex holds a replica of the shard corresponding to the shard vertex, connecting an edge from the shard vertex to the node vertex. For each node vertex, connecting multiple edges from the node vertex to the sink point. Setting the capacity and cost of each edge in the traffic network according to the number of primary replicas held by each node in the set of cluster nodes.

3. The method for selecting the main replica of time-series data sharding according to claim 2, wherein The setting of the capacity and cost of each edge in the traffic network according to the number of primary replicas held by each node in the set of cluster nodes includes: For the edge from the source point to the shard vertex, setting the capacity to 1 and the cost to 0. For the edge from the shard vertex to the node vertex, setting the capacity to 1 and the cost to 0. For the edge from the node vertex to the sink point, setting the capacity of each edge to 1, and the cost increasing with the set serial number of the edge.

4. The method for selecting the main replica of time series data sharding according to claim 1, wherein The using of the minimum-cost maximum-flow algorithm to solve the minimum-cost maximum-flow result in the traffic network includes: Using the shortest path algorithm to find the minimum-cost augmenting path from the source point to the sink point. Increasing the flow along the minimum augmenting path, and updating the flow and cost of the edges in the traffic network. Judging whether there is still an augmenting path. If so, returning to execute the step of using the shortest path algorithm to find the minimum-cost augmenting path from the source point to the sink point; otherwise, taking the current flow allocation scheme as the minimum-cost maximum-flow result.

5. The method for selecting the main replica of time series data shards according to claim 1, wherein After using the minimum-cost maximum-flow algorithm to solve the minimum-cost maximum-flow result in the traffic network, the method further includes: For all edges connecting node vertices and shard vertices, if there is flow passing through in the minimum-cost maximum-flow result, allocating the primary replica of each time-series data shard in the set of cluster shards to the corresponding node.

6. The method for selecting the primary replica of time-series data shards according to any one of claims 1-5, characterized in that The method further includes: Selecting the required hardware configuration according to the requirement analysis, deploying servers and configuring database software to form a set of cluster nodes. Selecting the required sharding strategy according to the data characteristics, determining the shard size and quantity to form a set of cluster shards.

7. A device for selecting the main replica of time-series data shards, characterized in that, Including: A building module, which is used to map a set of cluster nodes and a set of cluster shards into vertices of a traffic network, and add a source point and a sink point to the traffic network to construct the traffic network; wherein, the set of cluster shards includes multiple time-series data shards, and each time-series data shard includes a primary replica and a secondary replica. A solving module, which is used to solve the minimum-cost maximum-flow result in the traffic network by using the minimum-cost maximum-flow algorithm; wherein, the minimum-cost maximum-flow result includes the nodes where the primary replicas of each time-series data shard in the set of cluster shards are located.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the computer program, it implements the method for selecting the primary replica of the time-series data shard according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for selecting the primary replica of the time-series data shard according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for selecting the primary replica of the time-series data shard according to any one of claims 1 to 6.

Citation Information

Cited By

  • Graph database Leader fragment distribution method in multi-Zone scene

    CN120785742A