Distributed streaming data distribution method and device
In the distributed stream data processing system, the allocation of processing units is optimized by queuing theory and clustering algorithms, and the transmission order is adjusted based on the orderly propagation tree, the network overhead and delay problems of stream tuples are solved when transmitting and processing in the random processing unit are achieved, and more efficient data transmission and processing are achieved.
Patent Information
- Application Number
- CN202510094970.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, when stream tuples are transmitted and processed in a randomly distributed processing unit, the network overhead results in large network overhead and processing delays of stream tuples.
By randomly allocating multiple processing units in the distributed system to multiple nodes in the cluster, and calculating the first transmission delay of each node based on the queuing theory, clustering the nodes using the clustering algorithm, and finally allocating and adjusting the processing units based on the ordered propagation tree to obtain multiple target clusters.
Through the queueing theory, the transmission delay between processing units is estimated, and the computer nodes are allocated using clustering algorithms and transmission delays, the transmission order of the orderly propagation tree is optimized, network overhead is reduced, and processing delay of stream tuples is reduced.
Smart Images

Figure CN120017569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of streaming big data technology, and in particular to a distributed streaming data distribution method and device. Background Art
[0002] Stream data refers to a series of continuously generated data sequences. With the maturity of software and hardware technologies and the development of applications, stream data has exploded in growth. In order to meet the high-efficiency processing requirements of large amounts of stream data, streaming big data processing systems have emerged. In practical applications, systems usually need to provide high-quality services by processing continuous stream data in real time, such as algorithmic trading, advertising investment decisions, and online car-hailing. These applications rely on stream join operations to compare tuples from two data streams and output appropriate results. Compared with traditional database join operations, continuous high-speed real-time data streams make it more challenging to perform efficient data stream join operations. Existing big data stream join methods and devices are mainly divided into two categories: parallel stream join and distributed stream join.
[0003] The widely used distributed stream connection model, the bipartite graph model, requires large-scale broadcasting of connection tuples, and the network transmission will cause the stream tuples to be out of order, resulting in inconsistent order of tuples arriving at the processing unit, and thus incomplete connection results. An ordered propagation tree model solves the problem of tuple out of order. It is necessary to transmit tuples between different processing units in sequence, store each tuple randomly in the processing unit, connect the newly arrived tuple with the stored tuple of another stream, and send the tuple to the downstream processing unit. A large-scale distributed stream connection system will deploy a large number of processing units, which may be distributed in the same or different computer nodes. Different computers may also be distributed in different racks and different physical locations. The system randomly assigns processing units to different nodes. The transmission and processing of stream tuples in these randomly distributed processing units will result in large network overhead and increase the processing delay of stream tuples.
[0004] Therefore, there is an urgent need to propose a distributed stream data distribution method and device to solve the technical problems in the prior art that when stream tuples are transmitted and processed in randomly distributed processing units, network overhead is large and processing delays of stream tuples are caused. Summary of the invention
[0005] In view of this, it is necessary to provide a distributed stream data distribution method and device to solve the technical problems in the prior art that when stream tuples are transmitted and processed in randomly distributed processing units, network overhead is large and processing delays of stream tuples are caused.
[0006] In order to solve the above problems, the present invention provides a distributed stream data distribution method, comprising: Randomly assigning multiple processing units in a distributed system to multiple nodes in a cluster, and calculating the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node; Clustering the multiple nodes according to a clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes in each initial cluster; All processing units in the multiple initial nodes of each initial cluster are allocated and adjusted based on the ordered propagation tree to obtain multiple target clusters.
[0007] In a possible implementation, calculating the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node includes: Calculating the delay of sending tuples between the plurality of processing units based on queuing theory to obtain an initial delay of each processing unit; A first transmission delay of each node is obtained according to an initial delay of the tuple between the multiple processing units in each node.
[0008] In a possible implementation manner, clustering the multiple nodes according to a clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes in each initial cluster includes: Determine a plurality of first cluster centers according to a preset range; Clustering and calculating the multiple first cluster centers and the multiple nodes according to the clustering algorithm to obtain multiple clusters, and determining an initial cluster center of each cluster; Calculate the distance between each node of each cluster and the center of each cluster according to the first transmission delay to obtain the shortest distance of each node; The nodes in each cluster and the corresponding cluster center are updated according to the shortest distance to obtain multiple initial clusters and multiple initial nodes in each initial cluster.
[0009] In a possible implementation, clustering and calculating the multiple first cluster centers and the multiple nodes according to the clustering algorithm to obtain multiple clusters, and determining an initial cluster center of each cluster includes: Clustering the plurality of first cluster centers and the plurality of nodes according to the clustering algorithm to obtain a plurality of original clusters; Calculate the square sum of the nodes in each original cluster to get the square sum of the errors; Draw a square sum curve graph according to each original cluster and the error sum of squares; determining the number of clusters according to the elbow in the sum-of-squares curve plot; According to the number of clusters, multiple clusters are obtained, and an initial cluster center of each cluster is determined.
[0010] In a possible implementation, updating the nodes in each cluster and the corresponding cluster center according to the shortest distance to obtain multiple initial clusters and multiple initial nodes in each initial cluster includes: According to the shortest distance of each node, the node is assigned to the cluster of the corresponding initial cluster center to obtain multiple target optimal clusters after update; Calculating the distances of the nodes in the multiple target optimal clusters respectively to obtain cluster distances; Determine the updated target cluster center according to the cluster distance of each target best cluster; When the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster, the multiple target best clusters are determined to be multiple initial clusters, and the nodes in each target best cluster are determined to be multiple initial nodes of each initial cluster.
[0011] In a possible implementation manner, after determining the updated target cluster center according to the cluster distance of each target best cluster, the method further includes: When the target cluster center of each target best cluster is different from the initial cluster center, or the nodes in each target best cluster are different from the nodes of the corresponding cluster, the nodes of the multiple target best clusters are reallocated according to the first transmission delay, and the corresponding cluster distances are calculated and updated again until the condition of "when the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster" is met.
[0012] In a possible implementation, performing square sum calculation on the nodes in each original cluster to obtain the square sum of errors includes: Calculate the shortest distance between each node and other nodes based on the Dijkstra algorithm to obtain the path distance of each node; All path distances of the nodes in each original cluster are calculated to obtain the sum of square errors of each original cluster.
[0013] In a possible implementation, the allocating and adjusting all processing units in the multiple initial nodes of each initial cluster based on the ordered propagation tree to obtain multiple target clusters includes: According to the queuing theory, the delay of sending tuples between all processing units in each initial node of each initial cluster is calculated to obtain a second transmission delay of each initial node; Allocating all the processing units of each initial node to initial nodes of different initial clusters according to a preset node load, and calculating a third transmission delay of the corresponding initial node; When the second transmission delay is less than or equal to the third transmission delay, the multiple allocated initial clusters are determined to be multiple target clusters.
[0014] In a possible implementation manner, after allocating all the processing units of each initial node to initial nodes of different initial clusters according to a preset node load and calculating a third transmission delay of the corresponding initial node, the method further includes: When the second transmission delay is greater than the third transmission delay, updating the second transmission delay according to the third transmission delay to obtain an updated second transmission delay; Allocating the processing units of the initial nodes of the multiple allocated initial clusters according to the preset node load, and calculating and obtaining a fourth transmission delay of the corresponding initial node; When the updated second transmission delay is less than or equal to the fourth transmission delay, the multiple allocated initial clusters are determined to be multiple target clusters.
[0015] On the other hand, the present invention also provides a distributed stream data distribution device, comprising: A random allocation module, used to randomly allocate multiple processing units in the distributed system to multiple nodes of the cluster, and calculate the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node; A clustering division module, used for clustering the multiple nodes according to a clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes of each initial cluster; The allocation adjustment module is used to allocate and adjust all processing units in the multiple initial nodes of each initial cluster based on the ordered propagation tree to obtain multiple target clusters.
[0016] The beneficial effects of the present invention are as follows: a plurality of processing units in a distributed system are randomly distributed to a plurality of nodes in a cluster, and the delay of sending tuples between the plurality of processing units is calculated based on queuing theory to obtain a first transmission delay of each node; a plurality of nodes are clustered and divided according to a clustering algorithm and a first transmission delay to obtain a plurality of initial clusters and a plurality of initial nodes of each initial cluster; all processing units in a plurality of initial nodes of each initial cluster are allocated and adjusted based on an ordered propagation tree to obtain a plurality of target clusters; the present invention estimates the transmission delay between processing units through queuing theory, allocates computer nodes using a clustering algorithm and transmission delay to obtain a plurality of initial clusters, and accordingly allocates the best transmission order to each node of the ordered propagation tree, and adjusts the processing units in the nodes of the initial cluster, thereby saving network overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of a flow chart of an embodiment of the distributed stream data allocation method provided by the present invention; Figure 2 A schematic diagram of the structure of an embodiment of the distributed system provided by the present invention; Figure 3 For the present invention Figure 1 A schematic flow chart of an embodiment of step S103; Figure 4 For the present invention Figure 3 A schematic flow chart of an embodiment of step S302; Figure 5 For the present invention Figure 3 A schematic flow chart of an embodiment of step S304; Figure 6 For the present invention Figure 1 A schematic flow chart of an embodiment of step S104; Figure 7 A schematic diagram of the structure of an embodiment of the distributed stream data distribution device provided by the present invention; Figure 8 A schematic structural diagram of an embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0018] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0019] K-means clustering algorithm is an unsupervised learning algorithm, whose goal is to divide n observations into K clusters so that each observation belongs to the cluster center (centroid) closest to it, thereby minimizing the variance within the cluster.
[0020] Elbow Method,The K-means elbow method is an intuitive and commonly used technique for,determining the optimal number of clusters (i.e., the K value) in the K-means,clustering algorithm.
[0021] Dijkstra's Algorithm, also known as Dijkstra's Algorithm, is an algorithm used to find the shortest path from a vertex to all other vertices in a graph, and is particularly suitable for the single-source shortest path problem in a weighted graph (i.e., each edge in the graph has a weight).
[0022] A cluster is a collection of one or more nodes. The nodes in the cluster can store data and provide cross-node indexing and search capabilities.
[0023] Node: A node is a server in a cluster and can include multiple processing units.
[0024] Distributed systems, Distributed Stream Processing Systems (DSPS) decompose program logic into a series of small tasks and distribute them to each node in the cluster for timely processing, which is an ideal solution for big data processing systems.
[0025] Processing unit is the smallest unit for processing data in a distributed system. The processing unit of a stream connection distributed system can implement operations such as storage, connection, and forwarding of tuples.
[0026] Distributor, distributed systems require the processing units to work together, and the processing units need to obtain data. The distributor is responsible for distributing stream tuples to the corresponding processing units according to the data distribution strategy to perform storage or connection operations.
[0027] Binary tree structure, tree is an abstract data type used to simulate a data set with tree-like structure. It is a hierarchical set of n (n>0) finite nodes. It is called a "tree" because it looks like an upside-down tree, that is, its roots are facing up and its leaves are facing down. A binary tree is an ordered tree in which the degree of the nodes in the tree is no more than 2. It is the simplest and most important tree.
[0028] Queuing theory, or random service system theory, is to obtain the statistical laws of these quantitative indicators (waiting time, queue length, busy period, etc.) through statistical research on the arrival of service objects and service time, and then improve the structure of the service system or reorganize the service objects according to these laws, so that the service system can not only meet the needs of service objects, but also make the cost of the organization the most economical or some indicators the best. It is a branch of mathematical operations research and a discipline that studies the random laws of queuing phenomena in service systems. It is widely used in random service systems that share resources such as computer networks, production, transportation, and inventory. There are three aspects of queuing theory research: statistical inference, building models based on data; system properties, that is, the probabilistic regularity of quantitative indicators related to queuing; and system optimization problems. Its purpose is to correctly design and effectively operate various service systems to achieve the best benefits.
[0029] M / M / 1 model, M / M / 1 model is a commonly used queuing model in queuing theory, which is used to simulate and analyze the queuing system of a single service station. The following is a detailed explanation of the M / M / 1 model: 1) Model definition: (1) The M / M / 1 model represents a queuing system with a single server.
[0030] (2) The first “M” represents that the customer arrival interval follows a Poisson distribution with a parameter λ (λ represents the average number of customers arriving per unit time).
[0031] (3) The second “M” represents that the service time follows an exponential distribution with a parameter μ (μ represents the average number of services completed per unit time).
[0032] "1" means there is only one service desk in the system.
[0033] 2) Model features: (1) Customer arrivals are random and independent of each other.
[0034] (2) Service times are also random and independent of each other.
[0035] (3) The queue length is unlimited, that is, the number of customers waiting for service can be infinite.
[0036] (4) The system has only one service desk, which means it can only provide service to one customer at a time.
[0037] 3) Important parameters: (1) λ (arrival rate): the average number of customers arriving per unit time.
[0038] (2) μ (service rate): the average number of services completed per unit time.
[0039] Processing delay, here specifically refers to the processing delay of distributed systems, which is the time from when a tuple arrives at the system to when the system completes processing. The processing delay of distributed systems involves many aspects, including network communication, system architecture, data processing, and resource management.
[0040] Bipartite graph structure, bipartite graph or bipartite graph, is a special graph structure. In a bipartite graph, vertices can be divided into two disjoint sets, and each edge in the graph connects two vertices that belong to these two different sets. In other words, in a bipartite graph, there is no edge where both vertices belong to the same set.
[0041] An important property of a bipartite graph is that the vertices in the graph can be divided into two independent sets (denoted as U and V), so that each edge connects a vertex in U and a vertex in V. If it is further required that every vertex in U is connected to every vertex in V (and vice versa), then such a bipartite graph is called a complete bipartite graph.
[0042] like Figure 1 As shown, a specific embodiment of the present invention discloses a distributed stream data distribution method, including: S101, randomly assigning multiple processing units in a distributed system to multiple nodes of a cluster, and calculating the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node; S102, clustering the multiple nodes according to the clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes in each initial cluster; S103 , allocating and adjusting all processing units in multiple initial nodes of each initial cluster based on the ordered propagation tree to obtain multiple target clusters.
[0043] It should be noted that distributed systems such as Figure 2As shown in the figure, the distributed system can be composed of a distributor component, a CBF (Counting Bloom Filter), a processing unit component, and a monitor component. The distributor component is responsible for distributing tuples to the processing unit; the processing unit component randomly stores tuples and connects the received tuples with the tuples of another stream; the monitor component is responsible for counting the network transmission delay, recording the number of stored tuples and connected tuples, and performing planning calculations. For example, the bottom-up allocation method is used to allocate processing units through the transmission delay constrained path clustering algorithm. When the monitor detects that the transmission between processing units is significantly different from before, the order of the processing units needs to be adjusted. When the monitor detects that a processing unit is overloaded, it notifies the distributor component to record the ID of the overloaded processing unit in the CBF.
[0044] In a specific embodiment of the present invention, the distributed stream data distribution method proposed in the embodiment of the present invention can be applied to a distributed stream connection method and device based on an ordered propagation tree model, and can also be applicable to other distributed stream processing methods and devices based on tree models. Multiple processing units in a distributed system can be first randomly distributed to multiple nodes in a cluster, and then the delay of sending tuples between multiple processing units is calculated based on queuing theory to obtain the first transmission delay of each node, and then the multiple nodes are clustered according to the clustering algorithm and the first transmission delay, so that the transmission delay between nodes is low, and multiple initial clusters and multiple initial nodes of each initial cluster can be obtained. Then, based on the bottom-up allocation method and the clustering cluster of transmission delay, all processing units in the multiple initial nodes of the ordered propagation tree are allocated and adjusted, so that multiple target clusters after adjustment can be obtained.
[0045] Compared with the prior art, the present embodiment provides a method for randomly allocating multiple processing units in a distributed system to multiple nodes of a cluster, and calculating the delay of sending tuples between multiple processing units based on queuing theory to obtain a first transmission delay of each node; clustering and dividing multiple nodes according to a clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes of each initial cluster; allocating and adjusting all processing units in multiple initial nodes of each initial cluster based on an ordered propagation tree to obtain multiple target clusters; the present invention estimates the transmission delay between processing units through queuing theory, allocates computer nodes using a clustering algorithm and transmission delay to obtain multiple initial clusters, and accordingly allocates the best transmission order to each node of the ordered propagation tree, and adjusts the processing units in the nodes of the initial cluster, thereby saving network overhead.
[0046] In some embodiments of the present invention, step S101 includes: Based on queuing theory, the delay of sending tuples between multiple processing units is calculated to obtain the initial delay of each processing unit; According to the initial delay of the tuple between the multiple processing units in each node, the first transmission delay of each node is obtained.
[0047] In a specific embodiment of the present invention, the Apache storm open source system can be used, and each instance is a processing unit. Since each processing unit must receive the connection tuple and storage tuple sent by the upstream processing unit and send tuples to the downstream, their functions are the same and relatively independent. Therefore, this device models each processing unit as an M / M / 1 queuing model. i The calculation of the delay of processing a tuple by a processing unit is shown in formula (1): (1) In the formula, is the tuple arrival rate, It represents the average number of tuples that can be processed per unit time (referred to as processing rate).
[0048] Suppose that each tuple needs to be The processing delay of each tuple is calculated as shown in formula (2): (2) In the formula, The tuple is transferred from the previous processing unit to the i The time for each processing unit.
[0049] Because there is an upstream and downstream relationship between processing units, the processing time of the upstream processing unit will affect the waiting time of tuples in the downstream processing unit. If a processing unit is overloaded, it will cause its downstream processing units to wait idle. The distributor component can be notified to reduce the storage in the processing unit. If the tuples are too large, the distributor component is notified to reduce the storage of tuples in the processing unit from now on. A time-sliding heavy-load processing unit filter can be designed using multiple CBFs.
[0050] Therefore, the delay of sending tuples between multiple processing units can be calculated by using the queuing theory formula (1) to obtain the initial delay of each processing unit. Then, the initial delay of tuples between multiple processing units in each node can be calculated by using the formula (2) to obtain the first transmission delay of each node.
[0051] In some embodiments of the present invention, Figure 3 As shown, step S103 includes: S301, determining a plurality of first cluster centers according to a preset range; S302, clustering and calculating the multiple first cluster centers and the multiple nodes according to a clustering algorithm to obtain multiple clusters, and determining an initial cluster center of each cluster; S303, calculating the distance between each node of each cluster and the center of each cluster according to the first transmission delay, to obtain the shortest distance of each node; S304 , updating the nodes in each cluster and the corresponding cluster center according to the shortest distance to obtain multiple initial clusters and multiple initial nodes in each initial cluster.
[0052] In a specific embodiment of the present invention, a preset range can be set, and the preset range can be set according to actual conditions. For example, the preset range can be n / 20~n / 2 (where n is the number of computer nodes. The range can be set wider for the first time. When the system is running later, the value can be taken according to prior experience). Then, the K-means algorithm can be used to cluster and calculate the multiple first cluster centers and the multiple nodes. Specifically: In some embodiments of the present invention, such as Figure 4 As shown, step S302 includes: S401, clustering a plurality of first cluster centers and a plurality of nodes according to a clustering algorithm to obtain a plurality of original clusters; S402, calculating the sum of squares of the nodes in each original cluster to obtain the sum of squares of errors; S403, drawing a square sum curve graph according to each original cluster and the error sum of squares; S404, determining the number of clusters according to the elbow in the square sum curve graph; S405 . Obtain multiple clusters according to the number of clusters, and determine an initial cluster center of each cluster.
[0053] In a specific embodiment of the present invention, a K-means algorithm may be used to cluster multiple first cluster centers and multiple nodes to obtain multiple original clusters, and then a constrained path method may be used to calculate the square sum of the nodes in each original cluster to obtain the error square sum. Specifically: In some embodiments of the present invention, step S402 includes: Based on Dijkstra's algorithm, the shortest distance between each node and other nodes is calculated to obtain the path distance of each node; All path distances of the nodes in each original cluster are calculated to obtain the sum of squared errors of each original cluster.
[0054] In a specific embodiment of the present invention, according to the Dijkstra algorithm and the estimation of known data, only the shortest paths between the nodes closer to each node are calculated, so that the shortest distance between each node and other nodes can be calculated, and the shortest path distance of each node, that is, the transmission delay, can be obtained. The calculation can be ended by scanning some points. The calculation of the number of these partial nodes is , then all path distances of the nodes in each original cluster can be calculated to obtain the sum of squared errors of each original cluster, as shown in formula (3): (3) In the formula, is the sum of squared errors, C i For the i The original clusters, for x With each cluster center The transmission delay between
[0055] Each original cluster can then be plotted against The relationship diagram, that is, the square sum curve diagram, observe the square sum curve diagram, determine The position where the curve changes from a sharp drop to a gentle one is the "elbow" position. The K value corresponding to this position is the optimal number of clusters, that is, the number of clusters, so that K clusters can be obtained according to the square sum curve graph. The initial cluster center can be determined according to the data of each cluster, or the initial cluster center can be selected. The initial cluster center at this location can be obtained according to experience.
[0056] Further, the distance between each node of each cluster and the center of each initial cluster in the first transmission delay is calculated to obtain the shortest distance of each node, and then the nodes in each cluster and the corresponding initial cluster center can be updated according to the shortest distance to obtain multiple initial clusters and multiple initial nodes of each initial cluster. Specifically: In some embodiments of the present invention, Figure 5 As shown, step S304 includes: S501, according to the shortest distance of each node, assign the node to the cluster at the center of the corresponding initial cluster to obtain multiple target optimal clusters after update; S502, respectively calculating the distances of the nodes in multiple target optimal clusters to obtain cluster distances; S503, determining the updated target cluster center according to the cluster distance of each target best cluster; S504: When the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster, determine the multiple target best clusters as multiple initial clusters, and determine the nodes in each target best cluster as multiple initial nodes of each initial cluster.
[0057] In a specific embodiment of the present invention, the nodes can be assigned to the clusters represented by the corresponding initial cluster centers according to the shortest distance of each node, thereby obtaining multiple target optimal clusters after update. Then, the distances of the nodes in the multiple target optimal clusters can be calculated respectively to obtain the cluster distances, as shown in formula (4): (4) In the formula, For Node and The transmission delay between
[0058] Therefore, the initial cluster center can be updated according to the center position of the cluster distance of each target best cluster to obtain the target cluster center of each target best cluster after update. Then, the target cluster center and the target best cluster can be judged to determine whether the target cluster center of each target best cluster is the same as the initial cluster center, and whether the nodes in each target best cluster are the same as the nodes of the corresponding cluster. When the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster, it means that the update is completed, and multiple target best clusters are determined to be multiple initial clusters, and the nodes in each target best cluster are determined to be multiple initial nodes of each initial cluster. When the target cluster center of each target best cluster is different from the initial cluster center, or the nodes in each target best cluster are different from the nodes of the corresponding cluster, step S303 can be repeated for each target best cluster and the target cluster center, the cluster distance of each target best cluster can be repeated, and each target best cluster and the target cluster center can be updated again, thereby performing a cyclic update until "the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster", and the process is terminated to obtain multiple initial clusters and initial nodes of each initial cluster.
[0059] In some embodiments of the present invention, Figure 6 As shown, step S104 includes: S601, calculating the delay of sending tuples between all processing units in each initial node of each initial cluster according to queuing theory to obtain a second transmission delay of each initial node; S602, allocating all processing units of each initial node to initial nodes of different initial clusters according to a preset node load, and calculating a third transmission delay of the corresponding initial node; S603: When the second transmission delay is less than or equal to the third transmission delay, determine the multiple allocated initial clusters as multiple target clusters.
[0060] In a specific embodiment of the present invention, based on the multiple initial clusters obtained above and the allocation of nodes and processing units in each initial cluster, the nodes in the ordered propagation tree can be allocated to the corresponding initial clusters. Specifically, it can be: design a bottom-up allocation method, that is, first treat each small branch as a cluster, and finally divide the processing units at the root of the tree into a cluster, which is divided into K clusters in total. Because there is data transmission between trees, it is necessary to allocate the processing unit cluster with data transmission to the computer node cluster with low transmission delay, and then use the queuing theory formula (1) and formula (2) to calculate the delay of sending tuples between all processing units in each initial node, and obtain the second transmission delay of each initial node , and a preset node load can also be set. The preset node load can be the computing power and the allocated load of the computer node, which can be set according to the actual situation, so that all processing units of each initial node can be allocated to the initial nodes of different initial clusters according to the preset node load, and the third transmission delay of the corresponding initial node can be calculated by formula (1) and formula (2) , and then compare the second transmission delay with the third transmission delay. If , then the multiple initial clusters obtained by allocation are determined to be multiple target clusters. If , then the second transmission delay is updated according to the third transmission delay to obtain the updated second transmission delay, for example, ; Then, the adjustable processing units in the initial nodes of the allocated multiple initial clusters can be allocated again according to the preset node load, and the fourth transmission delay of the corresponding initial node can be calculated by formula (1) and formula (2); when the updated second transmission delay is less than or equal to the fourth transmission delay, the allocated multiple initial clusters are determined to be multiple target clusters, and after adjusting all possible allocation schemes, the selection is terminated, and the delay is recorded as of processing units.
[0061] The embodiment of the present invention provides a transmission delay constrained path clustering algorithm for existing reliable flow processing to classify computer nodes according to transmission delay. Provide a sequence of processing units with the shortest delay, thereby reducing tuple processing delay. Modeling is performed based on queuing theory, measuring the time it takes for a processing unit to process a tuple, and calculating the tuple processing delay of each grouping in combination with the transmission delay between processing units, so as to effectively plan transmission resources. The breadth-first algorithm is used to calculate the shortest path of transmission delay to constrain the calculation process and reduce calculation overhead. A bottom-up allocation method is adopted when allocating tree nodes, that is, each small branch is first regarded as a cluster, and the parent and child processing units are allocated in one group to reduce the calculation overhead of allocating processing units. The status of the processing unit can be monitored in real time, and when the transmission delay between some processing units is too long, the processing unit can be considered to be reallocated.
[0062] In order to better implement the distributed stream data distribution method in the embodiment of the present invention, based on the distributed stream data distribution method, the embodiment of the present invention also provides a distributed stream data distribution device, such as Figure 7 As shown, the distributed stream data distribution device 700 includes: A random allocation module 701 is used to randomly allocate multiple processing units in the distributed system to multiple nodes of the cluster, and calculate the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node; A clustering module 703 is used to cluster the multiple nodes according to the clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes of each initial cluster; The allocation adjustment module 704 is used to allocate and adjust all processing units in multiple initial nodes of each initial cluster based on the ordered propagation tree to obtain multiple target clusters.
[0063] The distributed stream data distribution device 700 provided in the above embodiment can implement the technical solution described in the above distributed stream data distribution method embodiment. The specific implementation principles of the above modules or units can refer to the corresponding contents in the above distributed stream data distribution method embodiment, which will not be repeated here.
[0064] like Figure 8 As shown, the present invention also provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802 and a display 803. Figure 8 Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0065] In some embodiments, the memory 802 may be an internal storage unit of the electronic device 800, such as a hard disk or memory of the electronic device 800. In other embodiments, the memory 802 may also be an external storage device of the electronic device 800, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the electronic device 800.
[0066] Furthermore, the memory 802 may include both an internal storage unit of the electronic device 800 and an external storage device. The memory 802 is used to store application software installed in the electronic device 800 and various data.
[0067] In some embodiments, the processor 801 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 802, such as the distributed stream data distribution method of the present invention.
[0068] In some embodiments, the display 803 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 803 is used to display information of the electronic device 800 and to display a visual user interface. The components 801-803 of the electronic device 800 communicate with each other via a system bus.
[0069] In some embodiments of the present invention, when the processor 801 executes the distributed stream data distribution program in the memory 802, the following steps may be implemented: The multiple processing units in the distributed system are randomly assigned to the multiple nodes of the cluster, and the delay of sending tuples between the multiple processing units is calculated based on the queuing theory to obtain the first transmission delay of each node; Clustering the multiple nodes according to the clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes in each initial cluster; All processing units in multiple initial nodes of each initial cluster are allocated and adjusted based on the ordered propagation tree to obtain multiple target clusters.
[0070] It should be understood that: when the processor 801 executes the distributed stream data allocation program in the memory 802, in addition to the above functions, other functions can also be implemented. For details, please refer to the description of the corresponding method embodiment above.
[0071] Furthermore, the embodiment of the present invention does not specifically limit the type of the electronic device 800 mentioned, and the electronic device 800 may be a portable electronic device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of portable electronic devices include but are not limited to portable electronic devices equipped with IOS, Android, Microsoft or other operating systems. The above-mentioned portable electronic devices may also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, the electronic device 800 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0072] Accordingly, an embodiment of the present application also provides a computer-readable storage medium, which is used to store computer-readable programs or instructions. When the program or instructions are executed by a processor, the steps or functions of the distributed stream data distribution method provided in the above-mentioned method embodiments can be implemented.
[0073] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0074] The above is a detailed introduction to the distributed stream data distribution method and device provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A distributed streaming data distribution method, characterized in that: include: Randomly assigning multiple processing units in a distributed system to multiple nodes in a cluster, and calculating the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node; Clustering the multiple nodes according to a clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes in each initial cluster; All processing units in the multiple initial nodes of each initial cluster are allocated and adjusted based on the ordered propagation tree to obtain multiple target clusters.
2. The distributed stream data distribution method according to claim 1, characterized in that: The step of calculating the delay of sending tuples between the plurality of processing units based on queuing theory to obtain a first transmission delay of each node includes: Calculating the delay of sending tuples between the plurality of processing units based on queuing theory to obtain an initial delay of each processing unit; A first transmission delay of each node is obtained according to an initial delay of the tuple between the multiple processing units in each node.
3. The distributed stream data distribution method according to claim 1, characterized in that: The clustering and dividing the multiple nodes according to the clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes of each initial cluster includes: Determine a plurality of first cluster centers according to a preset range; Clustering and calculating the multiple first cluster centers and the multiple nodes according to the clustering algorithm to obtain multiple clusters, and determining an initial cluster center of each cluster; Calculate the distance between each node of each cluster and the center of each initial cluster according to the first transmission delay to obtain the shortest distance of each node; The nodes in each cluster and the corresponding initial cluster center are updated according to the shortest distance to obtain multiple initial clusters and multiple initial nodes in each initial cluster.
4. The distributed stream data distribution method according to claim 3, characterized in that: The clustering and calculating the plurality of first cluster centers and the plurality of nodes according to the clustering algorithm to obtain a plurality of clusters, and determining an initial cluster center of each cluster, comprises: Clustering the plurality of first cluster centers and the plurality of nodes according to the clustering algorithm to obtain a plurality of original clusters; Calculate the square sum of the nodes in each original cluster to get the square sum of the errors; Draw a square sum curve graph according to each original cluster and the error sum of squares; determining the number of clusters according to the elbow in the sum-of-squares curve plot; According to the number of clusters, multiple clusters are obtained, and an initial cluster center of each cluster is determined.
5. The distributed stream data distribution method according to claim 3, characterized in that: The updating of the nodes in each cluster and the corresponding initial cluster center according to the shortest distance to obtain multiple initial clusters and multiple initial nodes in each initial cluster includes: According to the shortest distance of each node, the node is assigned to the cluster of the corresponding initial cluster center to obtain multiple target optimal clusters after update; Calculating the distances of the nodes in the multiple target optimal clusters respectively to obtain cluster distances; Determine the updated target cluster center according to the cluster distance of each target best cluster; When the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster, the multiple target best clusters are determined to be multiple initial clusters, and the nodes in each target best cluster are determined to be multiple initial nodes of each initial cluster.
6. The distributed stream data distribution method according to claim 5, characterized in that: After determining the updated target cluster center according to the cluster distance of each target best cluster, the method further includes: When the target cluster center of each target best cluster is different from the initial cluster center, or the nodes in each target best cluster are different from the nodes of the corresponding cluster, the nodes of the multiple target best clusters are reallocated according to the first transmission delay, and the corresponding cluster distances are calculated and updated again until the condition of "when the target cluster center of each target best cluster is the same as the initial cluster center, and the nodes in each target best cluster are the same as the nodes of the corresponding cluster" is met.
7. The distributed stream data distribution method according to claim 4, characterized in that: The square sum calculation is performed on the nodes in each original cluster to obtain the square sum of errors, including: Calculate the shortest distance between each node and other nodes based on the Dijkstra algorithm to obtain the path distance of each node; All path distances of the nodes in each original cluster are calculated to obtain the sum of square errors of each original cluster.
8. The distributed stream data distribution method according to claim 1, characterized in that: The allocating and adjusting all processing units in the multiple initial nodes of each initial cluster based on the ordered propagation tree to obtain multiple target clusters includes: According to the queuing theory, the delay of sending tuples between all processing units in each initial node of each initial cluster is calculated to obtain a second transmission delay of each initial node; Allocating all the processing units of each initial node to initial nodes of different initial clusters according to a preset node load, and calculating a third transmission delay of the corresponding initial node; When the second transmission delay is less than or equal to the third transmission delay, the multiple allocated initial clusters are determined to be multiple target clusters.
9. The distributed stream data distribution method according to claim 8, characterized in that: After allocating all the processing units of each initial node to initial nodes of different initial clusters according to the preset node load and calculating the third transmission delay of the corresponding initial node, the method further includes: When the second transmission delay is greater than the third transmission delay, updating the second transmission delay according to the third transmission delay to obtain an updated second transmission delay; Allocating the processing units of the initial nodes of the multiple allocated initial clusters according to the preset node load, and calculating and obtaining a fourth transmission delay of the corresponding initial node; When the updated second transmission delay is less than or equal to the fourth transmission delay, the multiple allocated initial clusters are determined to be multiple target clusters.
10. A distributed stream data distribution device, characterized in that: include: A random allocation module, used to randomly allocate multiple processing units in the distributed system to multiple nodes of the cluster, and calculate the delay of sending tuples between the multiple processing units based on queuing theory to obtain a first transmission delay of each node; A clustering division module, used for clustering the multiple nodes according to a clustering algorithm and the first transmission delay to obtain multiple initial clusters and multiple initial nodes of each initial cluster; The allocation adjustment module is used to allocate and adjust all processing units in the multiple initial nodes of each initial cluster based on the ordered propagation tree to obtain multiple target clusters.