Grouping Method, Device and Electronic Device for Data Processing Logic
By decomposing the large-mode data processing logic to execute multiple sub-logics in parallel, the problem of excessive calculation time of large-mode data processing logic is solved, and more efficient computing resource utilization and shorter computing time are achieved.
Patent Information
- Application Number
- CN202111586532.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-23
AI Technical Summary
In real-time data processing, when the data mode of the same type of data is too large, the calculation time is long, resulting in subsequent data processing being blocked.
By obtaining the attribute node reference relationship diagram of the data processing logic and its related overhead information, grouping operations are performed to decompose the data processing logic into multiple sub-data processing logic and execute in parallel to reduce the overall computation overhead and time.
Through reasonably grouping data processing logic, the calculation time can be significantly reduced, data processing is prevented, and the utilization rate of computing resources can be improved.
Smart Images

Figure CN114239837B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computers, and in particular, to a method, apparatus, and electronic device for grouping data processing logics. Background Art
[0002] Real-time data processing logics include the following information: information of the data source (Source) (data source, type of the data source, configuration of the data source), information of the data sink (Sink) (data sink, type of the data sink, configuration of the data sink), and information of the computing processor (Processor) (including: modes of the input data and output data (Schema, name, data type, etc.), mapping calculation logics from each attribute of the input data to each attribute of the output data, such as Figure 1 as shown), where the real-time data processing logic can be: Source→Processor1→Processor2→…→Processor N→Sink.
[0003] A general real-time data processing engine can generate real-time data processing logics according to user definitions. Specifically, it first reads the data in the data source, parses and converts the type of the original data according to the mode of the input data, then executes the mapping calculation logic to obtain the output data, and finally writes the output data into the data sink according to its mode. When performing data calculations, the real-time data processing engine generally processes data in parallel, so that data processing can be carried out simultaneously under different computing units (the smallest computing module that can only process one task at a time, such as a thread), thereby maximizing the utilization rate of computing resources.
[0004] In the above parallel processing process, data from the same data source can be distributed to different computing units for parallel calculation according to certain rules (for example, user operation records can be allocated to different computing units according to the unique user identifier). However, since there may be strict timing requirements before and after the calculation of the same type of data (for example, the operation timing of a user may be: login→browse→add to cart→place an order→logout, and the corresponding five operation logs have timing), if they are executed in parallel in different computing units, the timing of data processing cannot be guaranteed. Therefore, the same type of data is usually handed over to a specific computing unit for execution during grouping. However, when the data mode scale of the same type of data is too large, the calculation time is relatively long, which will block the processing of subsequent data.
[0005] Therefore, it has become an urgent technical problem to reasonably group the data processing logic of the large mode, and then split the input data, and place the input data in different computing units according to multiple sub-data processing logics obtained after grouping and perform calculations in parallel (since the sub-data processing logics are independent of each other, they can be executed in parallel). Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method, device and electronic device for grouping data processing logic to alleviate the technical problem of long calculation time when calculating the data processing logic of the large mode in the prior art.
[0007] In a first aspect, an embodiment of the present invention provides a method for grouping data processing logic, including:
[0008] Obtain the attribute node reference relationship graph of the data processing logic, the overhead of each attribute node in the attribute node reference relationship graph, the overhead of the root node dependency tree, the overhead of the attribute node reference relationship graph, the maximum overhead of the root node dependency tree, and the overhead upper limit of attribute grouping;
[0009] If the overhead of the attribute node reference relationship graph is not greater than the overhead upper limit of attribute grouping, then use the data processing logic as the only attribute grouping;
[0010] If the overhead of the attribute node reference relationship graph is greater than the overhead upper limit of attribute grouping, then perform the following grouping operation on the data processing logic:
[0011] Initialize the minimum total overhead of the data processing logic;
[0012] Perform the following loop on the current number of attribute groupings from the initial number of attribute groupings to the number of root nodes of the attribute node reference relationship graph:
[0013] Execute the P grouping strategy according to the current number of attribute groupings to obtain the current grouping method for grouping the attribute nodes in the attribute node reference relationship graph;
[0014] After grouping the attribute nodes in the attribute node reference relationship graph according to the current grouping method, calculate the current total overhead of the data processing logic;
[0015] If the minimum total overhead is greater than the current total overhead, then use the current total overhead as the minimum total overhead and use the current grouping method as the initial optimal grouping method, and continue to execute the above loop;
[0016] If the minimum total overhead is not greater than the current total overhead, then end the loop and use the grouping method obtained in the previous loop when ending the loop as the optimal grouping method;
[0017] Group the data processing logic according to the optimal grouping method, and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other.
[0018] Further, the current number of attribute groups is P, and P grouping strategies are executed according to the current number of attribute groups, including:
[0019] If the difference between the upper limit of the overhead of the attribute grouping and the maximum overhead is greater than a preset threshold, execute the following fast P grouping strategy:
[0020] Determine a target root node among the ungrouped root nodes, where the target root node is the root node with the largest overhead in the root node dependency tree among the ungrouped root nodes;
[0021] Among the P attribute groups, determine a first target attribute group, where the first target attribute group is the attribute group with the smallest overhead among the P attribute groups;
[0022] Add the root node dependency tree of the target root node to the first target attribute group, and update the overhead of the attribute group of the first target attribute group.
[0023] Further, executing the P grouping strategy according to the current number of attribute groups further includes:
[0024] If the difference between the upper limit of the overhead of the attribute grouping and the maximum overhead is not greater than a preset threshold, execute the following greedy grouping strategy:
[0025] Arrange the root nodes in the root node reference relationship graph of the attribute nodes in descending order of the overhead of the root node dependency tree;
[0026] Add the root node dependency tree of the target root node with the largest overhead of the root node dependency tree to the first attribute group, and update the overhead of the attribute group of the first attribute group;
[0027] For the second target attribute group that has not been assigned a root node dependency tree, execute the greedy search for grouping initialization element strategy to obtain a first target root node, and add the root node dependency tree of the first target root node to the second target attribute group, and update the overhead of the attribute group of the second target attribute group;
[0028] Traverse the remaining ungrouped root nodes in descending order of the cost of the root node dependency tree, execute a greedy grouping strategy to obtain the attribute grouping corresponding to each of the ungrouped root nodes or fail to find the attribute grouping corresponding to the ungrouped root nodes. Among them, if the attribute grouping corresponding to each of the ungrouped root nodes is obtained, add the root node dependency trees of each of the ungrouped root nodes to the corresponding attribute grouping. If the attribute grouping corresponding to the ungrouped root nodes is not found, return an empty set.
[0029] Further, for the second target attribute grouping not assigned to the root node dependency tree, execute a greedy grouping initialization element strategy to obtain the first target root node, including:
[0030] For each ungrouped root node, calculate the incremental combined cumulative calculation cost of adding the root node dependency tree of the ungrouped root node to each third target attribute grouping that has been assigned to the root node dependency tree, and then obtain the total incremental combined cumulative calculation cost.
[0031] Among all the total incremental combined cumulative calculation costs, determine the root node corresponding to the maximum total incremental combined cumulative calculation cost, and use the root node corresponding to the maximum total incremental combined cumulative calculation cost as the first target root node.
[0032] Further, traverse the remaining ungrouped root nodes in descending order of the cost of the root node dependency tree, execute a greedy grouping strategy to obtain the attribute grouping corresponding to each of the ungrouped root nodes, including:
[0033] For each grouping, calculate the incremental combined cumulative calculation cost of adding the root node dependency tree of the currently ungrouped root node to each attribute grouping.
[0034] Among all the incremental combined cumulative calculation costs, determine the attribute grouping corresponding to the minimum incremental combined cumulative calculation cost.
[0035] If the sum of the cost of the attribute grouping corresponding to the minimum incremental combined cumulative calculation cost and the minimum incremental combined cumulative calculation cost is not greater than the cost upper limit of the attribute grouping, use the attribute grouping corresponding to the minimum incremental combined cumulative calculation cost as the attribute grouping corresponding to the currently ungrouped root node.
[0036] Further, the method further includes:
[0037] Obtain a change request for the attribute node reference relationship graph, where the change request carries information about newly added attribute nodes.
[0038] If the newly added attribute node is not the root node, determine all the root nodes corresponding to the newly added attribute node, and add the newly added attribute node to the attribute groups corresponding to all its root nodes;
[0039] If the newly added attribute node is the root node, execute the greedy search for grouping strategy to determine the attribute group corresponding to the newly added attribute node;
[0040] If after executing the greedy search for grouping strategy, the attribute group corresponding to the newly added attribute node is not found, then take the newly added attribute node as a new attribute group.
[0041] Further, the initial number of attribute groups is the value obtained by rounding up the ratio of the cost of the attribute node reference graph to the upper limit of the cost.
[0042] In a second aspect, an embodiment of the present invention further provides a grouping device for data processing logic, including:
[0043] An acquisition unit, configured to acquire the attribute node reference graph of the data processing logic, the cost of each attribute node in the attribute node reference graph, the cost of the root node dependency tree, the cost of the attribute node reference graph, the maximum cost of the root node dependency tree, and the upper limit of the cost of the attribute group;
[0044] A first grouping unit, configured to, if the cost of the attribute node reference graph is not greater than the upper limit of the cost of the attribute group, take the data processing logic as the only attribute group;
[0045] A second grouping unit, configured to, if the cost of the attribute node reference graph is greater than the upper limit of the cost of the attribute group, perform the following grouping operation on the data processing logic:
[0046] Initialize the minimum total cost of the data processing logic;
[0047] Perform the following loop on the current number of attribute groups from the initial number of attribute groups to the number of root nodes of the attribute node reference graph:
[0048] Execute the P grouping strategy according to the current number of attribute groups to obtain the current grouping method for grouping the attribute nodes in the attribute node reference graph;
[0049] After grouping the attribute nodes in the attribute node reference graph according to the current grouping method, calculate the current total cost of the data processing logic;
[0050] If the minimum total cost is greater than the current total cost, then use the current total cost as the minimum total cost, use the current grouping method as the initial optimal grouping method, continue to execute the above loop, and use the current number of attribute groups as the minimum number of groups;
[0051] If the minimum total cost is not greater than the current total cost, then end the loop, and use the grouping method obtained in the previous loop when ending the loop as the optimal grouping method;
[0052] Group the data processing logic according to the optimal grouping method, and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other.
[0053] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of the above first aspects are implemented.
[0054] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and run by a processor, the machine-executable instructions cause the processor to run the method according to any one of the above first aspects.
[0055] In an embodiment of the present invention, a method for grouping data processing logic is provided, including: obtaining a reference relationship graph of attribute nodes of the data processing logic, the overhead of each attribute node in the reference relationship graph of attribute nodes, the overhead of the root node dependency tree, the overhead of the reference relationship graph of attribute nodes, the maximum overhead of the root node dependency tree, and the overhead upper limit of attribute grouping; if the overhead of the reference relationship graph of attribute nodes is not greater than the overhead upper limit of attribute grouping, then use the data processing logic as the only attribute group; if the overhead of the reference relationship graph of attribute nodes is greater than the overhead upper limit of attribute grouping, then perform the following grouping operation on the data processing logic: initialize the minimum total overhead of the data processing logic; perform the following loop on the current number of attribute groups from the initial number of attribute groups to the number of root nodes of the reference relationship graph of attribute nodes: execute the P grouping strategy according to the current number of attribute groups to obtain the current grouping method for grouping the attribute nodes in the reference relationship graph of attribute nodes; after grouping the attribute nodes in the reference relationship graph of attribute nodes according to the current grouping method, calculate the current total overhead of the data processing logic; if the minimum total overhead is greater than the current total overhead, then use the current total overhead as the minimum total overhead and use the current grouping method as the initial optimal grouping method, and continue to execute the above loop; if the minimum total overhead is not greater than the current total overhead, then end the loop and use the grouping method obtained in the previous loop when ending the loop as the optimal grouping method; group the data processing logic according to the optimal grouping method and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other. Through the above description, it can be seen that the method for grouping data processing logic of the present invention can reasonably group the data processing logic based on the calculated overhead to obtain multiple sub-data processing logics. Furthermore, when the multiple sub-data processing logics are executed in parallel, the overall calculation overhead is the smallest and the calculation time is the shortest, alleviating the technical problem of long calculation time when calculating the data processing logic of a large mode in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 Schematic diagram of a computing processor provided by an embodiment of the present invention;
[0058] Figure 2 Flowchart of a method for grouping data processing logic provided by an embodiment of the present invention;
[0059] Figure 3Schematic diagram of the attribute node reference relationship provided by the embodiment of the present invention;
[0060] Figure 4 Flowchart of the data processing logic change increment strategy provided by the embodiment of the present invention;
[0061] Figure 5 Schematic diagram of grouping input and output attributes provided by the embodiment of the present invention;
[0062] Figure 6 Schematic diagram of a grouping device for data processing logic provided by the embodiment of the present invention;
[0063] Figure 7 Schematic diagram of an electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0064] Next, the technical solutions of the present invention will be described clearly and completely in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] Currently, data of the same type is usually handed over to a specific computing unit for grouping during grouping. However, when the data pattern scale of data of the same type is too large, the calculation time is relatively long, which will block the processing of subsequent data.
[0066] Based on this, this embodiment provides a grouping method for data processing logic, which can reasonably group the data processing logic based on the calculation overhead, obtain multiple sub-data processing logics, and then when the multiple sub-data processing logics are executed in parallel, the overall calculation overhead is minimized and the calculation time is the shortest.
[0067] To facilitate the understanding of this embodiment, first, a grouping method for data processing logic disclosed in the embodiment of the present invention will be introduced in detail.
[0068] Embodiment 1:
[0069] According to the embodiment of the present invention, an embodiment of a grouping method for data processing logic is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0070] Figure 2 is a flowchart of a grouping method for data processing logic according to the embodiment of the present invention, as Figure 2As shown in the figure, the method includes the following steps:
[0071] Step S202, obtaining the attribute node reference relationship graph of the data processing logic, the overhead of each attribute node in the attribute node reference relationship graph, the overhead of the root node dependency tree, the overhead of the attribute node reference relationship graph, the maximum overhead of the root node dependency tree, and the overhead upper limit of the attribute grouping;
[0072] Specifically, first obtain the data processing logic, then parse the reference relationship of each attribute to other attributes in the calculation rules of each attribute, and then construct the attribute node reference relationship graph, and the attribute node reference relationship graph (denoted as G) can be obtained. As Figure 3 shown, it is an attribute node reference relationship graph constructed. The attribute node reference relationship graph includes: multiple attribute nodes and connection edges representing the reference relationship between attribute nodes (the attributes establish a mutual dependency relationship through the calculation rules). Among them, the node to which no other attribute node applies is the root node, and it is also an attribute node.
[0073] In addition, set the overhead of each attribute node (denoted as Current Load) according to the relative value of the actual occupied calculation time mean of the syntax structure of the calculation rules. For example, Figure 3 in, the attribute node m depends on the attribute nodes a and b. If its calculation rule is m = a + b, then the overhead of the attribute node m is the relative value of the addition structure according to the actual occupied calculation time mean. The overhead of the addition structure can be set to 1, that is, the overhead of the attribute node m is 1. The attribute node y depends on the attribute nodes n and o, and its calculation rule is the Fourier transform. Then, the overhead of the attribute node y is the relative value of the Fourier transform according to the actual occupied calculation time mean. The overhead of the Fourier transform can be set to 1000, that is, the overhead of the attribute node y is 1000. The embodiments of the present invention do not specifically limit the above specific values.
[0074] The above-mentioned overhead of the root node dependency tree (denoted as Included Load) refers to the sum of the overheads of the root node and all its referenced attribute nodes. For example, Figure 3 in, for the root node y, all its referenced attribute nodes are: the attribute node n, the attribute node o, the attribute node b, the attribute node c, and the attribute node d. Then the overhead of the root node dependency tree of the root node y is the overhead of the attribute node n + the overhead of the attribute node o + the overhead of the attribute node b + the overhead of the attribute node c + the overhead of the attribute node d. The so-called root node dependency tree refers to the reference relationship between each attribute node having a reference relationship with the root node and the root node.
[0075] The overhead of the above-mentioned attribute node reference relationship graph (denoted by Sum Current Load) refers to the sum of the overheads of all attribute nodes in the attribute node reference relationship graph. For Figure 3 in Figure 3 , the overhead of the attribute node reference relationship graph is: the overhead of attribute node x + the overhead of attribute node y + the overhead of attribute node z + the overhead of attribute node m + the overhead of attribute node n + the overhead of attribute node o + the overhead of attribute node a + the overhead of attribute node b + the overhead of attribute node c + the overhead of attribute node d.
[0076] The maximum overhead of the above-mentioned root node dependency tree (denoted by Max Included Load) refers to the maximum overhead among the overheads of all root node dependency trees.
[0077] The upper limit of the overhead of the above-mentioned attribute grouping (denoted by MAX_GROUP_CAPACITY) starts as the preset upper limit of the overhead of the attribute grouping, that is, the maximum overhead of the attribute grouping set by the user. If the maximum overhead of the root node dependency tree is greater than the preset upper limit of the overhead of the attribute grouping, then the maximum overhead of the root node dependency tree is used as the upper limit of the overhead of the attribute grouping. Generally speaking, when the root node dependency tree is very complex, its corresponding maximum overhead is very large. When it is larger than the preset upper limit of the overhead of the attribute grouping, at least the root node dependency tree has to be separated into a single attribute grouping. Then, the upper limit of the overhead of the attribute grouping needs to be adjusted upward to the maximum overhead of the root node dependency tree.
[0078] Step S204, if the overhead of the attribute node reference relationship graph is not greater than the upper limit of the overhead of the attribute grouping, then the data processing logic is used as the only attribute grouping;
[0079] That is to say, when the overhead of the attribute node reference relationship graph is not greater than the upper limit of the overhead of the attribute grouping, there is no need to group the data processing logic.
[0080] Step S206, if the overhead of the attribute node reference relationship graph is greater than the upper limit of the overhead of the attribute grouping, then the following grouping operation is performed on the data processing logic:
[0081] Initialize the minimum total overhead of the data processing logic (denoted by Min Overall Load);
[0082] Perform the following loop on the current number of attribute groupings from the initial number of attribute groupings to the number of root nodes of the attribute node reference relationship graph:
[0083] Execute the P grouping strategy according to the current number of attribute groupings to obtain the current grouping method (denoted by Current Properties Group) for grouping the attribute nodes in the attribute node reference relationship graph;
[0084] After grouping the attribute nodes in the attribute node reference relationship graph according to the current grouping method, calculate the current total overhead of the data processing logic (denoted as Current Overall Load);
[0085] If the minimum total overhead is greater than the current total overhead, then take the current total overhead as the minimum total overhead, take the current grouping method as the initial optimal grouping method (denoted as Properties Group), take the current number of attribute groups as the minimum number of groups (denoted as Min Group Number), and continue to execute the above loop;
[0086] If the minimum total overhead is not greater than the current total overhead, then end the loop, and take the grouping method obtained in the previous loop when ending the loop as the optimal grouping method, the current total overhead obtained in the previous loop as the optimal total overhead, and the current number of attribute groups obtained in the previous loop as the optimal number of groups;
[0087] Specifically, the above initial number of attribute groups is the value obtained by rounding up the ratio of the overhead of the attribute node reference relationship graph to the overhead upper limit.
[0088] The current total overhead of the above data processing logic refers to the sum of the overheads of each attribute group plus the additional calculation overhead introduced by the grouping operation. The overhead of an attribute group refers to the sum of the overheads of each attribute node in the attribute group, that is, the sum of the overheads of all calculation rules in the attribute group.
[0089] In the above grouping process, that is, to find a suitable value between the initial number of attribute groups and the number of root nodes as the optimal number of groups, and determine the optimal grouping method, so that the sum of the overheads of each divided attribute group and the additional calculation overhead introduced by the grouping operation is the smallest (that is, the lowest computational complexity).
[0090] The framework of the above grouping is:
[0091] Loop for the current number of attribute groups P from rounding up
Sum Current Load / MAX_GROUP_CAPACITY
[0092] According to the current number of groups P, execute the
P grouping algorithm
[0093] Current Properties Group; if it is an empty set, continue to loop for the next number of groups;
[0094] Calculate the Current Overall Load of Current Properties Group;
[0095] When Min Overall Load > Current Overall Load
[0096] Properties Group = Current Properties Group, Min Overall Load =
[0097] Current Overall Load, Min Group Number = P
[0098] If Current Overall Load is greater than that of the previous round, terminate the loop.
[0099] Step S208: Group the data processing logic according to the optimal grouping method, and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other.
[0100] For example, Figure 3 The optimal grouping method is: the attribute nodes m, a, b, x, n, b, c form one attribute group, and the attribute nodes y, n, o, b, c, d, z form another attribute group. Then, the data processing logic can be grouped according to the above grouping method, and correspondingly divided into two sub-data processing logics, and then the two sub-data processing logics are executed in parallel.
[0101] In an embodiment of the present invention, a method for grouping data processing logic is provided, including: obtaining a reference relationship graph of attribute nodes of the data processing logic, the overhead of each attribute node in the reference relationship graph of attribute nodes, the overhead of the root node dependency tree, the overhead of the reference relationship graph of attribute nodes, the maximum overhead of the root node dependency tree, and the overhead upper limit of attribute grouping; if the overhead of the reference relationship graph of attribute nodes is not greater than the overhead upper limit of attribute grouping, then use the data processing logic as the only attribute group; if the overhead of the reference relationship graph of attribute nodes is greater than the overhead upper limit of attribute grouping, then perform the following grouping operation on the data processing logic: initialize the minimum total overhead of the data processing logic; perform the following loop on the current number of attribute groups from the initial number of attribute groups to the number of root nodes of the reference relationship graph of attribute nodes: execute the P-grouping strategy according to the current number of attribute groups to obtain the current grouping method for grouping the attribute nodes in the reference relationship graph of attribute nodes; after grouping the attribute nodes in the reference relationship graph of attribute nodes according to the current grouping method, calculate the current total overhead of the data processing logic; if the minimum total overhead is greater than the current total overhead, then use the current total overhead as the minimum total overhead and use the current grouping method as the initial optimal grouping method, and continue to execute the above loop; if the minimum total overhead is not greater than the current total overhead, then end the loop and use the grouping method obtained in the previous loop when ending the loop as the optimal grouping method; group the data processing logic according to the optimal grouping method and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other. Through the above description, it can be seen that the method for grouping data processing logic of the present invention can reasonably group the data processing logic based on the calculated overhead to obtain multiple sub-data processing logics. Furthermore, when the multiple sub-data processing logics are executed in parallel, the overall calculation overhead is minimized and the calculation time is the shortest, alleviating the technical problem of long calculation time when calculating the data processing logic of a large mode in the prior art.
[0102] The above content briefly introduces the method for grouping data processing logic of the present invention. The following describes the specific content involved in detail.
[0103] In an alternative embodiment of the present invention, the current number of attribute groups is P, and executing the P-grouping strategy according to the current number of attribute groups specifically includes:
[0104] If the difference between the overhead upper limit of attribute grouping and the maximum overhead is greater than a preset threshold (that is, the maximum overhead is relatively small compared to the overhead upper limit of attribute grouping), then execute the following fast P-grouping strategy:
[0105] (1) Determine a target root node among the ungrouped root nodes, where the target root node is the root node with the largest overhead of the root node dependency tree among the ungrouped root nodes;
[0106] (2) Among the P attribute groups, determine the first target attribute group, where the first target attribute group is the attribute group with the smallest overhead among the P attribute groups;
[0107] (3) Add the root node dependency tree of the target root node to the first target attribute group and update the overhead of the first target attribute group.
[0108] To facilitate the understanding of the above process, the following is an illustrative metaphor:
[0109] Use 3 buckets to represent the P attribute groups, and use a water-filled water cup to represent the root node dependency tree. First, select the water cup with the most water in the water-filled water cups. Then, select the bucket with the least water among the 3 buckets. Then, pour the water in the water cup with the most water into the bucket with the least water, and then update the water volume of the bucket with the poured water; then select the water cup with the most water among the remaining water-filled water cups, select the bucket with the least water among the 3 buckets, and then pour the water in the water cup with the most water into the bucket with the least water, and then update the water volume of the bucket with the poured water. After performing such operations multiple times and completing pouring the water in all the water cups with water into the 3 buckets, the grouping is completed (the metaphor only applies here).
[0110] The above-mentioned attribute group refers to a set of a series of attribute nodes including the root node and all its indirect reference nodes, and the union of these sets contains all the root nodes of the data processing logic.
[0111] In an optional embodiment of the present invention, according to the current number of attribute groups, implementing the P-grouping strategy further includes:
[0112] If the difference between the overhead upper limit of the attribute group and the maximum overhead is not greater than the preset threshold, then implement the following greedy grouping strategy:
[0113] 1) Arrange the root nodes in the attribute node reference relationship graph in descending order of the overhead of the root node dependency tree;
[0114] 2) Add the root node dependency tree of the target root node with the largest overhead of the root node dependency tree to the first attribute group and update the overhead of the first attribute group;
[0115] 3) For the second target attribute group that has not been assigned the root node dependency tree, implement the greedy strategy for finding the grouping initialization element to obtain the first target root node, and add the root node dependency tree of the first target root node to the second target attribute group and update the overhead of the second target attribute group;
[0116] 4) Traverse the remaining ungrouped root nodes in descending order of the overhead of the root node dependency tree, perform a greedy search for the grouping strategy, and obtain the attribute grouping corresponding to each ungrouped root node or fail to find the attribute grouping corresponding to the ungrouped root node. Among them, if the attribute grouping corresponding to each ungrouped root node is obtained, add the root node dependency tree of each ungrouped root node to the corresponding attribute grouping. If the attribute grouping corresponding to the ungrouped root node is not found, return an empty set and continue to loop to the next current number of attribute groups.
[0117] The goal of the above greedy grouping strategy is: given the current number of attribute groups, find the greedy solution of the grouping method with "as much intra-group duplicate calculation as possible and as little inter-group duplicate calculation as possible" in the grouping situation under the current number of attribute groups.
[0118] The above step 2) means pouring the water in the cup with the most water into the first bucket and updating the water volume of the first bucket (specifically, the amount of water poured into it).
[0119] The above step 3) refers to finding a suitable water cup for an empty bucket and pouring the water in it as the first cup of water into the empty bucket. After performing step 3), each water bucket has the first cup of water;
[0120] The above step 4) refers to finding a suitable bucket for the remaining water cups filled with water and then pouring the water in them into the suitable bucket.
[0121] The above explanation is just a metaphorical illustration for easy understanding and is not the specific solution of the embodiments of the present invention.
[0122] In an alternative embodiment of the present invention, for the second target attribute group that is not assigned to the root node dependency tree, perform a greedy search for the grouping initialization element strategy to obtain the first target root node, including:
[0123] 1)) For each ungrouped root node, calculate the incremental merged cumulative calculation overhead of adding the root node dependency tree of the ungrouped root node to each third target attribute group that has been assigned to the root node dependency tree, and then obtain the total incremental merged cumulative calculation overhead;
[0124] The above total incremental merged cumulative calculation overhead is the sum of the incremental merged cumulative calculation overhead corresponding to an ungrouped root node.
[0125] 2)) Among all the total incremental merged cumulative calculation overheads, determine the root node corresponding to the largest total incremental merged cumulative calculation overhead, and use the root node corresponding to the largest total incremental merged cumulative calculation overhead as the first target root node.
[0126] The goal of the above-mentioned greedy strategy for finding the initial elements of grouping is: given the current number of attribute groups and the set of root node dependency trees, assign a root node dependency tree to each attribute group, so that the computational cost of each attribute group is as large as possible, and the duplicate calculations between groups are as few as possible.
[0127] The result of the above process is to find the first target root node that has the greatest difference from the third target attribute group to which the root node dependency tree has been assigned, that is, the root node dependency tree in the third target attribute group is completely irrelevant (that is, the duplicate calculations between groups are as few as possible), and then put it into the second target attribute group that has not been assigned the root node dependency tree.
[0128] The above-mentioned merged cumulative computational cost increment is the sum of the costs of the attribute nodes that do not have duplicates with the attribute nodes in the third target attribute group and belong to the root node dependency tree of the ungrouped root node among the attribute nodes in the root node dependency tree of the ungrouped root node.
[0129] Calculation framework for the merged cumulative computational cost increment:
[0130] Given two attribute groups M and N, where M.Included Load > N.Included Load (indicating that the Included Load of the M attribute group is greater than that of the N attribute group, that is, the cost of the M attribute group is greater than that of the N attribute group)
[0131] Initialize Merged Complexity Delta = 0
[0132] Traverse each attribute node n in N
[0133] If n appears in M, Merged Complexity Delta += 0
[0134] If n does not appear in M, Merged Complexity Delta += n.Current Load (indicating the Current Load of node n)
[0135] Return Merged Complexity Delta.
[0136] The framework of the greedy strategy for finding the initial elements of grouping is:
[0137] Max heap <Total Merged Cumulative Computational Cost Increment, Node> H
[0138] Traverse the remaining ungrouped root nodes N
[0139] Initialize the total merged cumulative computational cost increment of the current root node
[0140] Sum Current Merged Complexity Delta = 0
[0141] For each group B' assigned to the root node dependency tree
[0142] Calculate the
Merged Cumulative Computation Overhead Increment
[0143] Sum Current Merged Complexity Delta += Current Merged Complexity Delta
[0144] Add <Sum Current Merged Complexity Delta, N> to the max heap H and pop the top element from the max heap H for output (i.e., the first target root node).
[0145] In an alternative embodiment of the present invention, the remaining ungrouped root nodes are traversed in descending order of the overhead of the root node dependency tree, and a greedy grouping strategy is executed to obtain the attribute groups corresponding to each ungrouped root node, specifically including:
[0146] 1))) For each group, calculate the Merged Cumulative Computation Overhead Increment when the root node dependency tree of the currently ungrouped root node is added to each attribute group;
[0147] 2))) Among all the Merged Cumulative Computation Overhead Increments, determine the attribute group corresponding to the minimum Merged Cumulative Computation Overhead Increment;
[0148] 3))) If the sum of the overhead of the attribute group corresponding to the minimum Merged Cumulative Computation Overhead Increment and the minimum Merged Cumulative Computation Overhead Increment is not greater than the overhead upper limit of the attribute group, then use the attribute group corresponding to the minimum Merged Cumulative Computation Overhead Increment as the attribute group corresponding to the currently ungrouped root node.
[0149] The goal of the above greedy grouping strategy is: for a given root node, find an attribute group such that the additional overhead introduced when its root node dependency tree is added to the attribute group is minimized.
[0150] The calculation process of the above Merged Cumulative Computation Overhead Increment is similar to the introduction in the above content and will not be elaborated here.
[0151] The framework of the greedy grouping strategy is:
[0152] Initialize the minimum Merged Cumulative Computation Overhead Increment of the current root node
[0153] Min Current Merged Complexity Delta = Integer.MAX_VALUE (the largest constant value of an integer, 2,147,483,647), and the optimal grouping Min B currently assigned to the current root node
[0154] For each grouping B
[0155] Calculate the incremental cost of cumulative calculation for merging the root node dependency tree of the root node N into each attribute grouping, i.e., Current Merged Complexity Delta. If Current Merged Complexity Delta < Min Current Merged Complexity Delta and Current Merged Complexity Delta + B.Included Load (indicating the Included Load of grouping B) < MAX_GROUP_CAPACITY, then Min Current Merged Complexity Delta = Current Merged Complexity Delta and Min B = B
[0156] Place the root node dependency tree of the root node N into grouping B of attributes.
[0157] In an alternative embodiment of the present invention, refer to Figure 4 , the method further includes:
[0158] Step S401: Obtain a change request for the attribute node reference relationship graph, where the change request carries information about the newly added attribute node;
[0159] Step S402: If the newly added attribute node is not the root node, determine all the root nodes corresponding to the newly added attribute node, and add the newly added attribute node to the attribute groupings corresponding to all its corresponding root nodes;
[0160] Step S403: If the newly added attribute node is the root node, execute the greedy search for grouping strategy, determine the attribute grouping corresponding to the newly added attribute node, and add the newly added attribute node to its corresponding attribute grouping;
[0161] Step S404: If no attribute grouping corresponding to the newly added attribute node is found after executing the greedy search for grouping strategy, then use the newly added attribute node as a new attribute grouping.
[0162] Steps S401 to S404 are the incremental strategy for data processing logic changes: Considering that the modification of the calculation logic may introduce changes to the dependency graph, the change of the graph structure is very likely to cause changes in the attribute grouping attribution. In some cases, the grouping situation of an attribute needs to be fixed (for example, after selecting a calculation unit according to the group, some context-related states are stored in the calculation unit, and migrating the calculation unit will lose this state). The above incremental logic for changes, when the initial attribute grouping is fixed, makes fine-tuning to the data processing logic without affecting the existing attribute grouping.
[0163] Before submitting a processing task to the real-time data processing engine, according to the user's definition of the input-output mode of each calculation processor in the model and mapping calculation logic, the grouping method of the data processing logic of the present invention can group the input-output attributes (as Figure 5 shown, divided into 2 groups).
[0164] Grouping constraint: Attributes with calculation logic dependencies must be placed in the same attribute group so that the calculation logic before and after grouping is equivalent.
[0165] The processing task splits the input data according to the input-output definitions of different attribute groups and places it in different calculation units for parallel execution. Since the attributes between groups are independent of each other, they can be executed in parallel.
[0166] Before submitting a task to the real-time data processing engine, the calculation logic of the large mode can be split based on this method to facilitate parallelization and make full use of computing resources; in certain specific cases, once the above method is adopted, the attribute grouping must be fixed. For data processing logic changes, the incremental strategy for data processing logic changes can be used to group the newly added nodes without changing the existing attribute grouping.
[0167] Embodiment 2:
[0168] The embodiment of the present invention also provides a grouping device for data processing logic. The grouping device for data processing logic is mainly used to execute the grouping method of the data processing logic provided in Embodiment 1 of the present invention. The following makes a specific introduction to the grouping device for data processing logic provided in the embodiment of the present invention.
[0169] Figure 6 is a schematic diagram of a grouping device for data processing logic according to an embodiment of the present invention. As Figure 6 shown, the device mainly includes: an acquisition unit 10, a first grouping unit 20, and a second grouping unit 30, where:
[0170] An acquisition unit for acquiring a reference relationship graph of attribute nodes of a data processing logic, the overhead of each attribute node in the reference relationship graph of attribute nodes, the overhead of a root node dependency tree, the overhead of the reference relationship graph of attribute nodes, the maximum overhead of the root node dependency tree, and the overhead upper limit of attribute grouping;
[0171] A first grouping unit for, if the overhead of the reference relationship graph of attribute nodes is not greater than the overhead upper limit of attribute grouping, taking the data processing logic as the only attribute grouping;
[0172] A second grouping unit for, if the overhead of the reference relationship graph of attribute nodes is greater than the overhead upper limit of attribute grouping, performing the following grouping operation on the data processing logic:
[0173] Initializing the minimum total overhead of the data processing logic;
[0174] Performing the following loop on the current number of attribute groupings from the initial number of attribute groupings to the number of root nodes of the reference relationship graph of attribute nodes:
[0175] Executing a P grouping strategy according to the current number of attribute groupings to obtain the current grouping method for grouping the attribute nodes in the reference relationship graph of attribute nodes;
[0176] After grouping the attribute nodes in the reference relationship graph of attribute nodes according to the current grouping method, calculating the current total overhead of the data processing logic;
[0177] If the minimum total overhead is greater than the current total overhead, taking the current total overhead as the minimum total overhead and the current grouping method as the initial optimal grouping method, and continuing to execute the above loop;
[0178] If the minimum total overhead is not greater than the current total overhead, ending the loop and taking the grouping method obtained in the previous loop when the loop ends as the optimal grouping method;
[0179] Grouping the data processing logic according to the optimal grouping method and executing the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other.
[0180] In an embodiment of the present invention, a grouping device for data processing logic is provided, including: obtaining an attribute node reference relationship graph of the data processing logic, the overhead of each attribute node in the attribute node reference relationship graph, the overhead of the root node dependency tree, the overhead of the attribute node reference relationship graph, the maximum overhead of the root node dependency tree, and the overhead upper limit of attribute grouping; if the overhead of the attribute node reference relationship graph is not greater than the overhead upper limit of attribute grouping, then use the data processing logic as the only attribute grouping; if the overhead of the attribute node reference relationship graph is greater than the overhead upper limit of attribute grouping, then perform the following grouping operation on the data processing logic: initialize the minimum total overhead of the data processing logic; perform the following loop for the current number of attribute groupings from the initial number of attribute groupings to the number of root nodes of the attribute node reference relationship graph: execute the P grouping strategy according to the current number of attribute groupings to obtain the current grouping method for grouping the attribute nodes in the attribute node reference relationship graph; after grouping the attribute nodes in the attribute node reference relationship graph according to the current grouping method, calculate the current total overhead of the data processing logic; if the minimum total overhead is greater than the current total overhead, then use the current total overhead as the minimum total overhead and use the current grouping method as the initial optimal grouping method, and continue to execute the above loop; if the minimum total overhead is not greater than the current total overhead, then end the loop and use the grouping method obtained in the previous loop when ending the loop as the optimal grouping method; group the data processing logic according to the optimal grouping method, and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other. Through the above description, it can be seen that the grouping device for data processing logic of the present invention can reasonably group the data processing logic based on the calculated overhead to obtain multiple sub-data processing logics. Furthermore, when the multiple sub-data processing logics are executed in parallel, the overall calculation overhead is the smallest and the calculation time is the shortest, alleviating the technical problem of the long calculation time in the prior art when calculating the data processing logic of a large mode.
[0181] Optionally, the current number of attribute groupings is P, and the second grouping unit is further configured to: if the difference between the overhead upper limit of attribute grouping and the maximum overhead is greater than a preset threshold, then execute the following fast P grouping strategy: determine a target root node among the ungrouped root nodes, where the target root node is the root node with the largest overhead of the root node dependency tree among the ungrouped root nodes; determine a first target attribute grouping among the P attribute groupings, where the first target attribute grouping is the attribute grouping with the smallest overhead among the P attribute groupings; add the root node dependency tree of the target root node to the first target attribute grouping and update the overhead of the attribute grouping of the first target attribute grouping.
[0182] Optionally, the second grouping unit is further configured to: if the difference between the overhead upper limit of the attribute grouping and the maximum overhead is not greater than a preset threshold, execute the following greedy grouping strategy: arrange the root nodes in the attribute node reference relationship graph in descending order of the overhead of the root node dependency tree; add the root node dependency tree of the target root node with the largest overhead of the root node dependency tree to the first attribute grouping, and update the overhead of the attribute grouping of the first attribute grouping; for the second target attribute grouping that has not been assigned to the root node dependency tree, execute the greedy search for grouping initialization elements strategy to obtain the first target root node, and add the root node dependency tree of the first target root node to the second target attribute grouping, and update the overhead of the attribute grouping of the second target attribute grouping; traverse the remaining ungrouped root nodes in descending order of the overhead of the root node dependency tree, execute the greedy search for grouping strategy to obtain the attribute grouping corresponding to each ungrouped root node or fail to find the attribute grouping corresponding to the ungrouped root node, where if the attribute grouping corresponding to each ungrouped root node is obtained, add the root node dependency tree of each ungrouped root node to the corresponding attribute grouping, and if the attribute grouping corresponding to the ungrouped root node is not found, return an empty set.
[0183] Optionally, the second grouping unit is further configured to: for each ungrouped root node, calculate the incremental merged cumulative calculation overhead when the root node dependency tree of the ungrouped root node is added to each third target attribute grouping that has been assigned to the root node dependency tree, so as to obtain the total incremental merged cumulative calculation overhead; among all the total incremental merged cumulative calculation overheads, determine the root node corresponding to the largest total incremental merged cumulative calculation overhead, and use the root node corresponding to the largest total incremental merged cumulative calculation overhead as the first target root node.
[0184] Optionally, the second grouping unit is further configured to: for each grouping, calculate the incremental merged cumulative calculation overhead when the root node dependency tree of the currently ungrouped root node is added to each attribute grouping; among all the incremental merged cumulative calculation overheads, determine the attribute grouping corresponding to the smallest incremental merged cumulative calculation overhead; if the sum of the overhead of the attribute grouping corresponding to the smallest incremental merged cumulative calculation overhead and the smallest incremental merged cumulative calculation overhead is not greater than the overhead upper limit of the attribute grouping, use the attribute grouping corresponding to the smallest incremental merged cumulative calculation overhead as the attribute grouping corresponding to the currently ungrouped root node.
[0185] Optionally, the device is further configured to: obtain a change request for the attribute node reference relationship graph, where the change request carries information about a newly added attribute node; if the newly added attribute node is not a root node, determine all root nodes corresponding to the newly added attribute node, and add the newly added attribute node to the attribute groups corresponding to all its corresponding root nodes; if the newly added attribute node is a root node, execute a greedy search for a grouping strategy to determine the attribute group corresponding to the newly added attribute node; if no attribute group corresponding to the newly added attribute node is found after executing the greedy search for a grouping strategy, use the newly added attribute node as a new attribute group.
[0186] Optionally, the initial number of attribute groups is a value obtained by rounding up the ratio of the overhead of the attribute node reference relationship graph to the overhead upper limit.
[0187] The device provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For a brief description, for parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments.
[0188] As Figure 7 shown, an electronic device 600 provided by an embodiment of the present application includes: a processor 601, a memory 602, and a bus. The memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs, the processor 601 communicates with the memory 602 through the bus, and the processor 601 executes the machine-readable instructions to perform the steps of the grouping method of the above data processing logic.
[0189] Specifically, the above-mentioned memory 602 and processor 601 can be general-purpose memory and processor, which are not specifically limited here. When the processor 601 runs the computer program stored in the memory 602, it can execute the grouping method of the above data processing logic.
[0190] The processor 601 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 601 or the instructions in the form of software. The above-mentioned processor 601 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 602, and the processor 601 reads the information in the memory 602 and combines its hardware to complete the steps of the above method.
[0191] Corresponding to the grouping method of the above data processing logic, the embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the steps of the grouping method of the above data processing logic.
[0192] The grouping device of the data processing logic provided by the embodiments of the present application may be specific hardware on the device or software or firmware installed on the device, etc. For the device provided by the embodiments of the present application, the implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the above method embodiments, and will not be repeated here.
[0193] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0194] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0195] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0196] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0197] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the vehicle marking method described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs.
[0198] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0199] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solution of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solution described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solution deviate from the scope of the technical solution of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for grouping data processing logic, characterized in that, Including: Obtain the attribute node reference relationship graph of the data processing logic, the overhead of each attribute node in the attribute node reference relationship graph, the overhead of the root node dependency tree, the overhead of the attribute node reference relationship graph, the maximum overhead of the root node dependency tree, and the overhead upper limit of the attribute grouping; If the overhead of the attribute node reference relationship graph is not greater than the overhead upper limit of the attribute grouping, then use the data processing logic as the only attribute grouping; If the overhead of the attribute node reference relationship graph is greater than the overhead upper limit of the attribute grouping, then perform the following grouping operation on the data processing logic: Initialize the minimum total overhead of the data processing logic; Perform the following loop on the current number of attribute groupings from the initial number of attribute groupings to the number of root nodes of the attribute node reference relationship graph: Execute the P-grouping strategy according to the current number of attribute groupings to obtain the current grouping method for grouping the attribute nodes in the attribute node reference relationship graph; After grouping the attribute nodes in the attribute node reference relationship graph according to the current grouping method, calculate the current total overhead of the data processing logic; If the minimum total overhead is greater than the current total overhead, then use the current total overhead as the minimum total overhead, and use the current grouping method as the initial optimal grouping method, and continue to execute the above loop; If the minimum total overhead is not greater than the current total overhead, then end the loop, and use the grouping method obtained in the previous loop when ending the loop as the optimal grouping method; Group the data processing logic according to the optimal grouping method, and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other; The grouping of the data processing logic includes: Split the input data according to the input-output definitions of different attribute groupings for the processing tasks; The parallel execution of the multiple sub-data processing logics obtained after grouping includes: Place the split input data in different computing units for parallel execution.
2. The method according to claim 1, wherein When the current number of attribute groupings is P, executing the P-grouping strategy according to the current number of attribute groupings includes: If the difference between the overhead upper limit of the attribute grouping and the maximum overhead is greater than a preset threshold, then execute the following fast P-grouping strategy: Determine a target root node among the ungrouped root nodes, where the target root node is the root node with the largest overhead of the root node dependency tree among the ungrouped root nodes; Among the P attribute groupings, determine a first target attribute grouping, where the first target attribute grouping is the attribute grouping with the smallest overhead among the P attribute groupings; Add the root node dependency tree of the target root node to the first target attribute grouping, and update the overhead of the attribute grouping of the first target attribute grouping.
3. The method according to claim 1, wherein Executing the P-grouping strategy according to the current number of attribute groupings further includes: If the difference between the overhead upper limit of the attribute grouping and the maximum overhead is not greater than a preset threshold, then execute the following greedy grouping strategy: Arrange the root nodes in the attribute node reference relationship graph in descending order of the overhead of the root node dependency tree; Add the root node dependency tree of the target root node with the largest overhead in the root node dependency tree to the first attribute group, and update the overhead of the attribute group of the first attribute group; For the second target attribute group not assigned to the root node dependency tree, execute the greedy strategy for finding group initialization elements to obtain the first target root node, and add the root node dependency tree of the first target root node to the second target attribute group, and update the overhead of the attribute group of the second target attribute group; Traverse the remaining ungrouped root nodes in descending order of the overhead of the root node dependency tree, execute the greedy strategy for finding groups, and obtain the attribute groups corresponding to each of the ungrouped root nodes or fail to find the attribute groups corresponding to the ungrouped root nodes. Among them, if the attribute groups corresponding to each of the ungrouped root nodes are obtained, add the root node dependency trees of each of the ungrouped root nodes to the corresponding attribute groups. If the attribute groups corresponding to the ungrouped root nodes are not found, return an empty set.
4. The method according to claim 3, wherein For the second target attribute group not assigned to the root node dependency tree, execute the greedy strategy for finding group initialization elements to obtain the first target root node, including: For each ungrouped root node, calculate the incremental cost of the cumulative calculation of the merger when the root node dependency tree of the ungrouped root node is added to each third target attribute group that has been assigned to the root node dependency tree, and then obtain the total incremental cost of the cumulative calculation of the merger; Among all the total incremental costs of the cumulative calculation of the merger, determine the root node corresponding to the largest total incremental cost of the cumulative calculation of the merger, and use the root node corresponding to the largest total incremental cost of the cumulative calculation of the merger as the first target root node.
5. The method according to claim 3, characterized in that Traverse the remaining ungrouped root nodes in descending order of the overhead of the root node dependency tree, execute the greedy strategy for finding groups, and obtain the attribute groups corresponding to each of the ungrouped root nodes, including: For each group, calculate the incremental cost of the cumulative calculation of the merger when the root node dependency tree of the currently ungrouped root node is added to each attribute group; Among all the incremental costs of the cumulative calculation of the merger, determine the attribute group corresponding to the smallest incremental cost of the cumulative calculation of the merger; If the sum of the overhead of the attribute group corresponding to the smallest incremental cost of the cumulative calculation of the merger and the smallest incremental cost of the cumulative calculation of the merger is not greater than the overhead upper limit of the attribute group, use the attribute group corresponding to the smallest incremental cost of the cumulative calculation of the merger as the attribute group corresponding to the currently ungrouped root node.
6. The method according to claim 1, wherein The method further includes: Obtain a change request for the attribute node reference relationship graph, where the change request carries information about newly added attribute nodes; If the newly added attribute node is not a root node, determine all the root nodes corresponding to the newly added attribute node, and add the newly added attribute node to the attribute groups corresponding to all its corresponding root nodes; If the newly added attribute node is a root node, execute the greedy strategy for finding groups to determine the attribute group corresponding to the newly added attribute node; If the attribute group corresponding to the newly added attribute node is not found after executing the greedy strategy for finding groups, use the newly added attribute node as a new attribute group.
7. The method according to claim 1, characterized in that, The initial number of attribute groups is the value obtained by rounding up the ratio of the overhead of the attribute node reference relationship graph to the overhead upper limit.
8. A grouping device for a data processing logic, characterized in that, It includes: An acquisition unit, configured to acquire the attribute node reference relationship graph of the data processing logic, the overhead of each attribute node in the attribute node reference relationship graph, the overhead of the root node dependency tree, the overhead of the attribute node reference relationship graph, the maximum overhead of the root node dependency tree, and the overhead upper limit of the attribute group; A first grouping unit, configured to, if the overhead of the attribute node reference relationship graph is not greater than the overhead upper limit of the attribute group, use the data processing logic as the only attribute group; A second grouping unit, configured to, if the overhead of the attribute node reference relationship graph is greater than the overhead upper limit of the attribute group, perform the following grouping operation on the data processing logic: Initialize the minimum total overhead of the data processing logic; Perform the following loop on the current number of attribute groups from the initial number of attribute groups to the number of root nodes of the attribute node reference relationship graph: Execute the P grouping strategy according to the current number of attribute groups to obtain the current grouping method for grouping the attribute nodes in the attribute node reference relationship graph; After grouping the attribute nodes in the attribute node reference relationship graph according to the current grouping method, calculate the current total overhead of the data processing logic; If the minimum total overhead is greater than the current total overhead, use the current total overhead as the minimum total overhead and use the current grouping method as the initial optimal grouping method, and continue to execute the above loop; If the minimum total overhead is not greater than the current total overhead, end the loop and use the grouping method obtained in the previous loop when the loop ends as the optimal grouping method; Group the data processing logic according to the optimal grouping method, and execute the multiple sub-data processing logics obtained after grouping in parallel, where the multiple sub-data processing logics are independent of each other; The grouping of the data processing logic includes: Splitting the input data according to the input-output definitions of different attribute groups for the processing tasks; The parallel execution of the multiple sub-data processing logics obtained after grouping includes: Placing the split input data in different computing units for parallel execution.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7 above.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and run by the processor, the machine-executable instructions cause the processor to run the method described in any one of claims 1 to 7 above.
Citation Information
Patent Citations
A method for calculating and adjusting the difference between production and sales of pipe network based on distribution
CN109360025A
Incremental segmentation processing method and device, computer equipment and storage medium
CN113255264A
Minimal disclosure credential verification and revocation
US20140281525A1
Metafutures-based Graphed Data Lookup
US20180268029A1