Data flow diagram processing method and device, equipment, storage medium and program product
By introducing graph editing distance and graph editing operations into the data flow graph, it is extended to the distance measurement of directed acyclic graphs, and the problem of difficulty in clustering directed acyclic graph data flow graphs in the prior art is solved, and efficient clustering processing is achieved.
Patent Information
- Application Number
- CN202510168593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
It is difficult for the prior art to effectively cluster data flow graphs in the form of directed acyclic graphs.
By introducing the concept of graph editing distance, defining graph editing operations include modifying the operator type of node and modifying the direction of directed edges, thereby expanding the distance measurement of graph editing distance to directed acyclic graph, and using K-mean clustering algorithm, etc. for clustering processing.
Effective clustering of data flow graphs in the form of directed acyclic graphs is realized, ensuring the accuracy and effectiveness of clustering.
Smart Images

Figure CN119989035A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method, apparatus, device, storage medium and program product for processing a data flow graph. Background Art
[0002] Distributed stream processing systems usually adopt a processing method similar to pipelines. Data streams are processed by a series of operators in the system, such as filtering, mapping, aggregation, etc. These operators constitute the logical process of data stream processing and are constructed into a data flow graph in the form of a directed acyclic graph. The clustering method in related technologies is usually used for undirected graphs. Therefore, a method for clustering data flow graphs in the form of directed acyclic graphs is urgently needed. Summary of the invention
[0003] In view of this, the present disclosure provides a method, apparatus, device, storage medium and program product for processing a data flow graph to solve the clustering problem of the data flow graph.
[0004] In a first aspect, the present disclosure provides a method for processing a data flow graph, the method comprising:
[0005] Obtaining a plurality of first data flow graphs;
[0006] Based on the graph edit distance between the first data flow graphs, the first data flow graphs are clustered to obtain multiple cluster clusters, wherein the graph edit distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph, and the graph editing operations include modifying at least one of the operator type of the nodes in the first data flow graph and modifying the direction of the directed edges between the nodes.
[0007] In a second aspect, the present disclosure provides a data flow graph processing device, the device comprising:
[0008] A data acquisition module, used to acquire a plurality of first data flow graphs;
[0009] A clustering processing module is used to perform clustering processing on the first data flow graphs based on the graph edit distance between the first data flow graphs to obtain multiple cluster clusters, wherein the graph edit distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph, and the graph editing operation includes modifying the operator type of the nodes in the first data flow graph and modifying at least one of the directions of the directed edges between the nodes.
[0010] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data flow graph processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0011] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for processing a data flow graph of the above-mentioned first aspect or any corresponding embodiment thereof.
[0012] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable a computer to execute the method for processing a data flow graph of the above-mentioned first aspect or any corresponding embodiment.
[0013] The data flow graph processing method provided by the embodiment of the present disclosure adds a graph editing operation for modifying the operator type of the nodes in the first data flow graph and a graph editing operation for modifying the direction of the directed edges between the nodes in the graph editing operation for calculating the graph editing distance. Therefore, the graph editing distance applicable to the undirected graph can be extended to the distance metric of the first data flow graph, so that the graph editing distance can be used to cluster the first data flow graph and the effectiveness of the clustering of the first data flow graph can be guaranteed.
[0014] The beneficial effects of the data flow graph processing device, electronic device, storage medium and program product correspond to the beneficial effects of the data flow graph processing method and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 It is a flowchart of a method for processing a data flow graph according to an embodiment of the present disclosure;
[0017] Figure 2 is a schematic diagram of a first data flow graph according to an embodiment of the present disclosure;
[0018] Figure 3 is a flowchart of another method for processing a data flow graph according to an embodiment of the present disclosure;
[0019] Figure 4is a structural block diagram of a data flow graph processing device according to an embodiment of the present disclosure;
[0020] Figure 5 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0027] A distributed stream processing system is a computing system used to process continuous data streams. Distributed stream processing systems play a vital role in the field of data processing. From an architectural perspective, a distributed stream processing system consists of multiple computing nodes that are connected through a network and work together to process massive real-time data streams. This distributed architecture enables distributed stream processing systems to efficiently cope with large-scale data through parallel processing. For example, when processing real-time behavioral data from Internet users (such as click streams, social media message streams, etc.), different computing nodes can process part of the data at the same time, thereby improving the processing speed of the distributed stream processing system.
[0028] The flow of data in a distributed stream processing system is one of the core features of the system. Data continuously enters the distributed stream processing system in the form of streams. These data streams may come from various data sources, such as sensor networks, log files, network monitoring tools, etc. The distributed stream processing system processes these data streams in real time, which means that the data streams entering the distributed stream processing system are processed almost at the same time as they enter the system, rather than being stored first and then processed like a batch processing system.
[0029] In terms of the processing logic of distributed stream processing systems, distributed stream processing systems usually adopt a processing method similar to pipelines. Data streams are processed by a series of operators in the system, such as filtering, mapping, aggregation, etc. These operators constitute the logical process of data stream processing and are constructed into a data flow graph in the form of a directed acyclic graph. The clustering method in related technologies is usually used for undirected graphs. Therefore, a method for clustering data flow graphs in the form of directed acyclic graphs is urgently needed.
[0030] In view of this, according to an embodiment of the present disclosure, an embodiment of a method for processing a data flow graph is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0031] In this embodiment, a data flow graph processing method is provided, which can be used in a distributed stream processing system. Figure 1 is a flow chart of a method for processing a data flow graph according to an embodiment of the present disclosure, such as Figure 1 As shown, the process includes the following steps:
[0032] Step S101, obtain multiple first data flow graphs.
[0033] Among them, the first data flow graph is a logical data flow directed acyclic graph (Logical Dataflow DAG). The logical data flow directed acyclic graph represents the high-level operation workflow defined by the application. The logical data flow directed acyclic graph focuses on task relationships and abstracts execution details.
[0034] For example, Figure 2 A first data flow diagram with four operators (O1, O2, O3, O4 as shown in the figure) is shown. The first data flow diagram reflects the actual deployment of the operators on the computing resources and explains in detail how the logical plan is divided, parallelized and executed.
[0035] Furthermore, the first data flow graph is obtained based on a historical data flow graph of the distributed stream processing system.
[0036] It should be noted that the first data flow graph includes one or more nodes, each node has a corresponding operator, and the operator is used to perform specific conversion operations on the data.
[0037] Step S102, based on the graph edit distance between the first data flow graphs, clustering the first data flow graphs to obtain multiple cluster clusters, the graph edit distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph, and the graph editing operations include modifying the operator type of the nodes in the first data flow graph and modifying at least one of the directions of the directed edges between the nodes.
[0038] It is worth noting that since the first data flow graph is a directed acyclic graph of logical data flow, and the clustering method in the related art is mainly designed for undirected graphs, in order to effectively cluster the first data flow graph, it is necessary to define a distance metric between the first data flow graphs. The defined distance metric needs to ensure the effectiveness of the clustering process and satisfy the triangle inequality characteristics to avoid contradictory clustering results. Among them, the triangle inequality can ensure that the direct distance between two objects is always less than or equal to the sum of the distances through the intermediate objects. Therefore, if the triangle inequality is used to define the distance metric between the first data flow graphs, the following conditions must be met between any three first data flow graphs: d(g1,g3)<=d(g1,g2)+d(g2,g3); wherein g1 is the first first data flow graph, g2 is the second first data flow graph, and g3 is the third first data flow graph. d(g1,g3) is the distance metric between g1 and g3, d(g1,g2) is the distance metric between g1 and g2, and d(g2,g3) is the distance metric between g2 and g3.
[0039] The graph edit distance (GED) satisfies the triangle inequality and can provide an interpretable similarity measure between graphs. Although the graph edit distance is an NP-hard problem, since the first data flow graph is small in size, usually containing less than 20 nodes and directed edges, the graph edit distance is still a feasible distance measurement scheme in the clustering scenario of the data flow graph of the distributed stream system. Formally, the graph edit distance between two first data flow graphs is defined as the minimum number of graph edit operations to transform one first data flow graph into another first data flow graph. For example, the graph edit distance between g1 and g2 is recorded as ged(g1,g2), which represents the minimum number of graph edit distances to transform g1 into g2. However, in the related art, the graph editing operations used to calculate the graph edit distance only include standard operations such as node insertion, node deletion, edge insertion and edge deletion, which are only applicable to undirected graphs, and it is difficult to directly use the graph edit distance in the first data flow graph. Therefore, it is necessary to introduce additional graph editing operations for the first data flow graph into the graph editing distance, so as to extend the graph editing distance to be used in the first data flow graph.
[0040] Based on this, the present disclosure adds two graph editing operations in the graph editing operation: modifying the operator type of the nodes in the first data flow graph and modifying the direction of the directed edges between the nodes.
[0041] For example, if the operator type of the node is a filter operator, the above graph editing operation can be used to change the filter operator to a connection operator.
[0042] Alternatively, the direction of the directed edge between any two nodes in the first data flow graph is reversed.
[0043] After defining the distance metric of the first data flow graph, the first data flow graph can be regarded as a data point in the clustering space, and clustering algorithms such as the K-means clustering algorithm and the hierarchical clustering algorithm are used to cluster similar first data flow graphs to obtain multiple cluster clusters.
[0044] The data flow graph processing method provided in this embodiment adds a graph editing operation for modifying the operator type of the nodes in the first data flow graph and a graph editing operation for modifying the direction of the directed edges between the nodes in the graph editing operation for calculating the graph editing distance. Therefore, the graph editing distance applicable to the undirected graph can be extended to the distance metric of the first data flow graph, so that the graph editing distance can be used to cluster the first data flow graph and the effectiveness of the clustering of the first data flow graph can be ensured.
[0045] In this embodiment, another data flow graph processing method is provided, which can be used in a distributed stream processing system. Figure 3is a flowchart of another method for processing a data flow graph according to an embodiment of the present disclosure, such as Figure 3 As shown, the process includes the following steps:
[0046] Step S301, obtaining a plurality of first data flow graphs, which can be referred to as the above step S101, will not be described in detail here.
[0047] Step S302: clustering the first data flow graphs based on the graph edit distance between the first data flow graphs to obtain a plurality of cluster clusters, wherein the graph edit distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph, wherein the graph editing operations include modifying at least one of the operator type of the nodes in the first data flow graph and modifying the direction of the directed edges between the nodes.
[0048] Optionally, in the above step S302, a K-means clustering algorithm is used to perform clustering processing on the first data flow graph to obtain a plurality of cluster clusters.
[0049] Specifically, the above step S302 includes:
[0050] Step S3021, determining multiple cluster centers in the first data flow graph.
[0051] During the clustering initialization process, a preset number of first data flow graphs are randomly selected from a plurality of first data flow graphs as initial clustering centers.
[0052] Step S3022, determining a first graph edit distance between the first data flow graph excluding the cluster center and the cluster center.
[0053] Specifically, based on the minimum number of graph editing operations for transforming the first data flow graph excluding the cluster center to the cluster center, a first graph editing distance between the first data flow graph and the cluster center is obtained.
[0054] Step S3023: Based on the first graph edit distance, clustering processing is performed between the first data flow graph except the cluster center and the cluster center to obtain multiple cluster clusters.
[0055] The cluster corresponds to the cluster center, and the cluster includes the corresponding cluster center and the first data flow graph assigned to the cluster center.
[0056] Specifically, based on the first graph edit distance, the first data flow graph except the cluster center is assigned to the cluster center with the nearest distance to obtain a plurality of cluster clusters.
[0057] Step S3024: If the clustering stop condition is not met, a graph similarity search is performed on the first data flow graph in the cluster cluster, and the cluster center of the cluster cluster is updated according to the search result.
[0058] Optionally, the clustering stop condition is that the cluster centers converge or the number of iterations of the clustering process reaches a preset iteration number threshold.
[0059] It is worth noting that if the clustering stop condition is not met, after updating the cluster center, the process returns to the above step S3023, and iteratively executes the above step S3023 and step S3024 until the clustering stop condition is met.
[0060] It is worth noting that in the clustering of the first data flow graphs, a key challenge is that in the clustering process of the first data flow graphs, there is a lack of a direct determination method for calculating the cluster center of a set of first data flow graphs. Ideally, the cluster center should be able to capture the most central or most representative first data flow graph in the cluster cluster. Unlike numerical data, it is difficult to perform an "average" operation on the first data flow graph. Therefore, a potential solution is to use the concept of a median graph to determine the cluster center of the cluster cluster so that the total graph edit distance of all first data flow graphs in the cluster cluster is minimized. However, the implementation of the concept of a central tree graph requires the calculation of the graph edit distance between each pair of first data flow graphs in the cluster cluster, which is computationally very time-consuming. Therefore, in order to solve this problem, the present disclosure proposes the concept of an approximate median graph, which is called a similarity center. By performing a graph similarity search on all the first data flow graphs in the clustering cluster, the number of times the first data flow graph is judged as a similar data flow graph by other first data flow graphs in the clustering cluster to which it belongs is identified. Since the first data flow graph with the most times is relatively close to other first data flow graphs in the clustering cluster, it is believed that the first data flow graph with the most times can effectively approximate the central data flow graph of the clustering cluster, and then the cluster center is defined as the first data flow graph in the clustering cluster that is judged as a similar data flow graph the most times.
[0061] The data flow graph processing method provided in this embodiment performs a graph similarity search on the first data flow graph in the cluster cluster when the clustering of the first data flow graph does not meet the clustering stop condition, and updates the cluster center of the cluster cluster according to the search results. Therefore, the graph similarity search can be used to capture the most central or representative first data flow graph in the cluster cluster as the cluster center, thereby ensuring the accuracy of the clustering of the first data flow graph.
[0062] In some optional implementations, the step S3024 of performing a graph similarity search on the first data flow graph in the cluster cluster, and updating the cluster center of the cluster cluster according to the search results, includes:
[0063] Step a1, performing a graph similarity search on the first data flow graph in the cluster to determine a similar data flow graph of the first data flow graph.
[0064] Specifically, the above step a1 includes:
[0065] Step a11, determining the second graph edit distance between the first data flow graphs in the clustering clusters.
[0066] Step a12: searching for similar data flow graphs of the first data flow graph in the clustering clusters based on the second graph edit distance.
[0067] Furthermore, the above step a12 includes: if the second graph editing distance between the first data flow graph and other data flow graphs in the cluster to which it belongs is less than a preset distance threshold, then determining the other data flow graphs as similar data flow graphs of the first data flow graph.
[0068] In a specific embodiment, a query graph q and a preset distance threshold t of the graph edit distance are determined in the first data flow graph of the cluster, and the graph similarity search will find all the first data flow graphs in the cluster whose second graph edit distance with the query graph q is less than the preset distance threshold as similar data flow graphs of the query graph q. q,t = {q∈G|ged(q,g)<=t}; where Sim q,t is the set of similar data flow graphs of query graph q, G is the cluster to which query graph q belongs, g is other data flow graphs in cluster G, and ged(q,g) is the second graph edit distance between query graph q and other data flow graphs g.
[0069] Step a2: updating the cluster center of the cluster based on the number of times the first data flow graph in the cluster is determined to be a similar data flow graph.
[0070] Specifically, the above step a2 includes: based on the first data flow graph with the largest number of times in the cluster, updating the cluster center of the cluster.
[0071] In a specific embodiment, the number of times C that other data flow graphs g in the cluster are determined to be similar data flow graphs is g is defined as: the number of times other data flow graphs g are determined to be similar data flow graphs in the search results of similarity search of all first data flow graphs in cluster G using a preset distance threshold t. Formally, this number can be defined as: C g =∑q∈G∏(g∈Sim q,t ), where ∏ is the indicator function.
[0072] The cluster center G of cluster G SC It can be defined as:
[0073] The data flow graph processing method provided in this embodiment uses the second graph edit distance between the first data flow graphs in the clustering cluster, and searches for similar data flow graphs corresponding to the first data flow graph in other data flow graphs in the clustering cluster. Based on the number of times the first data flow graph in the clustering cluster is determined as a similar data flow graph, the first data flow graph with the largest number of times is selected, and the cluster center of the clustering cluster is updated. Therefore, the cluster center of each clustering cluster can be determined quickly and accurately to ensure the accuracy of the clustering of the first data flow graph.
[0074] In some optional implementations, determining the second graph edit distance between the first data flow graphs in the clusters in the above step a11 includes:
[0075] Step a111, establishing a target mapping relationship of nodes between the second data flow graph and the third data flow graph, the second data flow graph and the third data flow graph are any two different first data flow graphs in the cluster.
[0076] Specifically, starting from the starting node, a target mapping relationship of nodes is established between the second data flow graph and the third data flow graph.
[0077] Step a112, starting from the first mapping relationship in the target mapping relationship, explore the first minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph.
[0078] Specifically, the first mapping relationship is an incomplete mapping relationship between nodes of the second data flow graph and the third data flow graph in the current search state. This incomplete mapping relationship allows the matching to be gradually constructed and expanded during the search process to find the optimal node mapping relationship.
[0079] Specifically, the first minimum number is determined based on the similarity of nodes (such as the similarity of node attributes) and the similarity of directed edges (such as the existence, weight, etc. of directed edges) between the second data flow graph and the third data flow graph.
[0080] Step a113: if the first minimum number is greater than a preset number threshold, prune the first mapping relationship in the target mapping relationship to obtain a pruned mapping relationship.
[0081] Among them, the preset quantity threshold is used to determine whether the current first mapping relationship is worth continuing to expand. If the calculated first minimum quantity is greater than the preset quantity threshold, it is considered that the first mapping relationship is unlikely to be the optimal solution for the least graph editing operations. Therefore, the first mapping relationship can be pruned to improve the exploration efficiency.
[0082] Step a114, based on the mapping relationship after pruning, explore the second minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph, and obtain the second graph editing distance between the second data flow graph and the third data flow graph.
[0083] It should be noted that in the process of exploring the second minimum number based on the pruned mapping relationship, if the exploration stop condition is not met, the second mapping relationship is determined from the mapping relationship other than the first mapping relationship in the target mapping relationship. Starting from the second mapping relationship, the second minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph is explored. If the second minimum number is greater than the preset number threshold, the second mapping relationship is pruned to obtain a new pruned mapping relationship, and the cycle is repeated until the exploration stop condition is met. The minimum value of all explored first minimum numbers and second minimum numbers is used as the second graph editing distance between the second data flow graph and the third data flow graph. Among them, the exploration stop condition is to find the mapping relationship of all nodes between the second data flow graph and the third data flow graph or the unexplored nodes in the search space are empty.
[0084] It is worth noting that in order to efficiently calculate the cluster center of the first data flow graph cluster, it is crucial to accelerate the graph similarity search. Although graph similarity search has been widely studied, it is still challenging to select the most appropriate graph similarity search method based on the setting of the cluster center in this embodiment.
[0085] The graph similarity search in the related art mainly adopts the filtering and verification method to reduce the computational burden of graph edit distance calculation. This method builds an index in the cluster cluster so that irrelevant graphs can be eliminated during the cluster center update phase, and then enters the verification phase to identify qualified graphs. However, this graph similarity search relies on pre-built indexes. As the first data flow graph in the cluster cluster changes during the update iteration process, the pre-built index may become outdated, resulting in a large amount of index reconstruction overhead.
[0086] Therefore, in the present disclosure, a heuristic search algorithm can be used to perform graph similarity search on the first data flow graph in the clustering cluster, and the optimization of graph similarity search is achieved by directly optimizing the calculation of the edit distance of the second graph, thereby avoiding dependence on pre-built indexes. Optionally, the heuristic search algorithm is the AStar+-LSa algorithm.
[0087] The heuristic search algorithm uses the best-first search strategy to directly optimize the calculation of the second graph edit distance. In a specific embodiment, the heuristic search algorithm starts from the starting nodes of the second data flow graph and the third data flow graph, and establishes a target mapping relationship between the nodes of the second data flow graph and the third data flow graph, so as to calculate the second graph edit distance through the least graph editing operation. During the search process, the first mapping relationship representing the incomplete mapping relationship of the node subset, that is, the partial mapping relationship in the target mapping relationship, will be gradually explored. For each first mapping relationship, a tight lower bound (that is, the first minimum number of the above-mentioned graph editing operations) is calculated by considering the similarity of the node and the directed edge. Branches with the first minimum number greater than the preset number threshold are pruned, thereby effectively reducing the search space and memory usage. The best-first search strategy not only accelerates the graph similarity search, but also effectively improves the calculation speed of the second graph edit distance calculation itself. This dual efficiency can effectively improve the clustering of the first data flow graph, greatly speed up the calculation steps of the first graph edit distance, the second graph edit distance and the cluster center in the above-mentioned clustering process, can significantly reduce the resource loss in the clustering process, and greatly reduce the calculation time.
[0088] The data flow graph processing method provided in this embodiment establishes a target mapping relationship of nodes between the second data flow graph and the third data flow graph, and starts from the first mapping relationship in the target mapping relationship to explore the first minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph. If the first minimum number is greater than the preset number threshold, the first mapping relationship in the target mapping relationship is pruned, thereby effectively improving the calculation efficiency of the second graph editing distance between the second data flow graph and the third data flow graph, thereby improving the clustering efficiency of the first data flow graph.
[0089] In the practical application of distributed stream processing systems, data flow execution should adapt to fluctuating workload characteristics to match different workload requirements. In related technologies, the parallelism of nodes in data flow graphs is usually manually adjusted by system engineers. However, this parallelism adjustment method is limited by the experience of system engineers, and is prone to low accuracy of parallelism adjustment, and the labor cost is too high. Therefore, in other related technologies, some rule-based methods and model-based methods are proposed to automatically tune the parallelism of data flow graphs in distributed stream processing systems. However, this type of parallelism tuning method either relies on the parallelism tuning history of the target stream processing task, or is based on the prediction of the performance indicators (such as latency and throughput) of the target stream processing task. However, these indirect parallelism tuning methods often lead to poor tuning decisions, especially when the performance indicator prediction is inaccurate, the tuning decision is even more inaccurate.
[0090] Based on this, in some optional implementations, the data flow graph processing method of the present disclosure further includes:
[0091] Step b1, obtain the target data flow graph of the target stream processing task.
[0092] Among them, the target data flow graph has one or more nodes, and the nodes have corresponding operators. The operators are used to perform specific conversion operations on the data. In practical applications, the operators corresponding to the nodes of the target data flow graph can be flexibly combined according to specific business needs.
[0093] Step b2, determining the third graph edit distance between the target data flow graph and the cluster centers of the multiple cluster clusters.
[0094] The calculation steps of the third graph edit distance may refer to the calculation steps of the first graph edit distance and the second graph edit distance.
[0095] Step b3: determine the target cluster corresponding to the target data flow graph based on the third graph edit distance.
[0096] Specifically, the cluster corresponding to the smallest third graph edit distance is used as the target cluster corresponding to the target data flow graph.
[0097] Step b4, based on the node parallelism characteristics represented by the target cluster, configure the parallelism of the nodes in the target data flow graph to obtain a parallelism configuration result.
[0098] Specifically, for each cluster cluster, the first data flow graph in the cluster cluster and the parallelism configuration of the nodes in the first data flow graph can be used to train the preset model corresponding to the cluster cluster to obtain the target model of each cluster cluster. The target data flow graph is input into the target model corresponding to the target cluster cluster, the parallelism of the nodes in the target data flow graph is predicted, and the nodes in the target data flow graph are configured in parallel according to the prediction result to obtain the parallelism configuration result.
[0099] Step b5: Process the target stream processing task based on the parallelism configuration result to obtain the processing result of the target stream processing task.
[0100] The data flow graph processing method provided in this embodiment, since the distributed stream processing system will generate rich execution history records from a large number of data flow operations, these records can provide rich knowledge that can be used to optimize parallelism suggestions and speed up the tuning process. The parallelism that needs to be set between data flow graphs with similar structures is close. Therefore, the data flow graph processing method of this embodiment clusters the first data flow graphs with similar structures to obtain cluster clusters. When it is necessary to process the target data flow graph of a certain target stream processing task, the node parallelism characteristics represented by the target cluster cluster with a structure similar to the target stream processing task can be used to configure the parallelism of the nodes in the target data flow graph to improve the parallelism configuration efficiency and configuration effect.
[0101] As a specific application example, a target program is installed on a distributed stream processing system, and the target program is used to execute the data flow graph processing method disclosed in the present invention. The distributed stream processing system can select a first data flow graph in the historical data flow graph, and call the target program to execute the data flow graph processing method disclosed in the present invention to cluster the first data flow graph to obtain multiple clustering clusters. And when processing the target stream processing task, the clustering cluster is used to configure the parallelism of the nodes in the target data flow graph of the target stream processing task to obtain the parallelism configuration result. The target stream processing task is processed based on the parallelism configuration result to obtain the processing result of the target stream processing task, thereby effectively utilizing the knowledge of parallelism configuration in the historical data flow graph, adopting the parallelism configuration of the historical data flow graph with similar structure, and configuring the parallelism of the target data flow graph, thereby improving the efficiency and accuracy of the parallelism configuration.
[0102] In this embodiment, a data flow graph processing device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0103] This embodiment provides a data flow graph processing device, such as Figure 4 As shown, including:
[0104] A data acquisition module 401 is used to acquire a plurality of first data flow graphs;
[0105] The clustering processing module 402 is used to perform clustering processing on the first data flow graph based on the graph editing distance between the first data flow graphs to obtain multiple cluster clusters. The graph editing distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph. The graph editing operation includes modifying the operator type of the nodes in the first data flow graph and modifying at least one of the directions of the directed edges between the nodes.
[0106] In some optional implementations, the clustering processing module 402 includes:
[0107] A center determination unit, used to determine a plurality of cluster centers in the first data flow graph;
[0108] a distance calculation unit, used to determine a first graph edit distance between the first data flow graph excluding the cluster center and the cluster center;
[0109] A cluster processing unit, configured to perform cluster processing between the first data flow graph other than the cluster center and the cluster center based on the first graph edit distance, to obtain a plurality of cluster clusters;
[0110] The center updating unit is used to perform a graph similarity search on the first data flow graph in the cluster cluster if the clustering stop condition is not met, and update the cluster center of the cluster cluster according to the search result.
[0111] In some optional implementations, the central updating unit includes:
[0112] A search subunit, configured to perform a graph similarity search on the first data flow graph in the cluster to determine a similar data flow graph of the first data flow graph;
[0113] The updating subunit is used to update the cluster center of the cluster based on the number of times the first data flow graph in the cluster is determined to be a similar data flow graph.
[0114] In some optional implementations, the search subunit is specifically used to: determine the second graph edit distance between the first data flow graphs in the clustering cluster; and search for similar data flow graphs of the first data flow graph in the clustering cluster based on the second graph edit distance.
[0115] In some optional embodiments, the search subunit is specifically used to determine the second graph editing distance between the first data flow graphs in the clustering cluster, including: establishing a target mapping relationship of nodes between the second data flow graph and the third data flow graph, the second data flow graph and the third data flow graph being any two different first data flow graphs in the clustering cluster; starting from the first mapping relationship in the target mapping relationship, exploring the first minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph; if the first minimum number is greater than a preset number threshold, pruning the first mapping relationship in the target mapping relationship to obtain the pruned mapping relationship; based on the pruned mapping relationship, exploring the second minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph to obtain the second graph editing distance between the second data flow graph and the third data flow graph.
[0116] In some optional embodiments, the search subunit is specifically used to search for similar data flow graphs of the first data flow graph in the clustering cluster based on the second graph editing distance, including: if the second graph editing distance between the first data flow graph and other data flow graphs in the clustering cluster to which it belongs is less than a preset distance threshold, then determining the other data flow graphs as similar data flow graphs of the first data flow graph.
[0117] In some optional implementations, the updating subunit is specifically used to update the cluster center of the cluster based on the first data flow graph with the largest number of times in the cluster.
[0118] In some optional implementations, the data flow graph processing device of the present disclosure further includes:
[0119] A task acquisition module is used to obtain the target data flow graph of the target stream processing task;
[0120] A distance operation module, used for determining a third graph edit distance between a target data flow graph and cluster centers of a plurality of cluster clusters;
[0121] A flow graph clustering module, used for determining a target clustering cluster corresponding to the target data flow graph based on the edit distance of the third graph;
[0122] A parallelism configuration module is used to configure the parallelism of nodes in the target data flow graph based on the node parallelism characteristics represented by the target cluster, and obtain a parallelism configuration result;
[0123] The task processing module is used to process the target stream processing task based on the parallelism configuration result to obtain the processing result of the target stream processing task.
[0124] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0125] The processing device of the data flow graph in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0126] The present disclosure also provides an electronic device having the above Figure 4 The data flow diagram shown is a processing device.
[0127] See also Figure 5 , Figure 5 is a structural block diagram of an electronic device provided by an optional embodiment of the present disclosure, such as Figure 5As shown, the electronic device includes: one or more processors 501, a memory 502, and an interface for connecting various components, including a high-speed interface and a low-speed interface. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 501 is taken as an example.
[0128] The processor 501 may be a central processing unit, a network processor or a combination thereof. The processor 501 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0129] The memory 502 stores instructions executable by at least one processor 501 , so that the at least one processor 501 executes the method shown in the above embodiment.
[0130] The memory 502 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 502 may optionally include a memory remotely arranged relative to the processor 501, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0131] The memory 502 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 502 may also include a combination of the above types of memory.
[0132] The electronic device also includes an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503 and the output device 504 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.
[0133] The input device 503 can receive input digital or character information, and generate key signal input related to the user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick, etc. The output device 504 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0134] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0135] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0136] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for processing a data flow graph, characterized in that: The method comprises: Obtaining a plurality of first data flow graphs; Based on the graph edit distance between the first data flow graphs, the first data flow graphs are clustered to obtain multiple cluster clusters, wherein the graph edit distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph, and the graph editing operations include modifying at least one of the operator type of the nodes in the first data flow graph and modifying the direction of the directed edges between the nodes.
2. The method for processing a data flow graph according to claim 1, characterized in that: The clustering process is performed on the first data flow graphs based on the graph edit distance between the first data flow graphs to obtain a plurality of cluster clusters, including: Determining a plurality of cluster centers in the first data flow graph; determining a first graph edit distance between a first data flow graph other than the cluster center and the cluster center; Based on the first graph edit distance, clustering processing is performed between the first data flow graph except the cluster center and the cluster center to obtain a plurality of the cluster clusters; If the clustering stop condition is not met, a graph similarity search is performed on the first data flow graph in the cluster cluster, and the cluster center of the cluster cluster is updated according to the search result.
3. The method for processing a data flow graph according to claim 2, characterized in that: The performing graph similarity search on the first data flow graph in the clustering cluster, and updating the cluster center of the clustering cluster according to the search result, comprises: Performing a graph similarity search on the first data flow graph in the cluster to determine a similar data flow graph of the first data flow graph; Based on the number of times the first data flow graph in the cluster is determined to be the similar data flow graph, the cluster center of the cluster is updated.
4. The method for processing a data flow graph according to claim 3, characterized in that: The performing graph similarity search on the first data flow graph in the cluster to determine a similar data flow graph of the first data flow graph includes: determining a second graph edit distance between first data flow graphs in the clusters; Based on the second graph edit distance, a similar data flow graph of the first data flow graph is searched in the clustering clusters.
5. The method for processing a data flow graph according to claim 4, characterized in that: The determining a second graph edit distance between first data flow graphs in the clustering clusters comprises: Establishing a target mapping relationship of nodes between the second data flow graph and the third data flow graph, wherein the second data flow graph and the third data flow graph are any two different first data flow graphs in the cluster; Starting from a first mapping relationship in the target mapping relationships, exploring a first minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph; If the first minimum number is greater than a preset number threshold, pruning the first mapping relationship in the target mapping relationship to obtain a pruned mapping relationship; Based on the mapping relationship after pruning, the second minimum number of graph editing operations for transforming the second data flow graph into the third data flow graph is explored to obtain a second graph editing distance between the second data flow graph and the third data flow graph.
6. The method for processing a data flow graph according to claim 4, characterized in that: The step of searching the clustering cluster for a similar data flow graph to the first data flow graph based on the second graph edit distance comprises: If the second graph edit distance between the first data flow graph and other data flow graphs in the cluster to which it belongs is less than a preset distance threshold, the other data flow graphs are determined as similar data flow graphs of the first data flow graph.
7. The method for processing a data flow graph according to claim 3, characterized in that: The updating of the cluster center of the cluster based on the number of times the first data flow graph in the cluster is determined as the similar data flow graph comprises: Based on the first data flow graph with the largest number of occurrences in the cluster, the cluster center of the cluster is updated.
8. The method for processing a data flow graph according to any one of claims 1 to 7, characterized in that: The method further comprises: Get the target data flow graph of the target stream processing task; Determining a third graph edit distance between the target data flow graph and cluster centers of the plurality of cluster clusters; Determine a target cluster corresponding to the target data flow graph based on the third graph edit distance; Based on the node parallelism characteristics represented by the target cluster, the nodes in the target data flow graph are configured in parallel to obtain a parallelism configuration result; The target stream processing task is processed based on the parallelism configuration result to obtain a processing result of the target stream processing task.
9. A data flow graph processing device, characterized in that: The device comprises: A data acquisition module, used to acquire a plurality of first data flow graphs; A clustering processing module is used to perform clustering processing on the first data flow graphs based on the graph edit distance between the first data flow graphs to obtain multiple cluster clusters, wherein the graph edit distance is used to characterize the minimum number of graph editing operations for transforming from one first data flow graph to another first data flow graph, and the graph editing operation includes modifying the operator type of the nodes in the first data flow graph and modifying at least one of the directions of the directed edges between the nodes.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data flow graph processing method described in any one of claims 1 to 8 by executing the computer instructions.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data flow graph processing method described in any one of claims 1 to 8.
12. A computer program product, characterized in that It comprises computer instructions, and the computer instructions are used to make a computer execute the data flow graph processing method described in any one of claims 1 to 8.