Rule-based execution plan generation method, device, equipment and medium
By evaluating and sorting predicates in the rule set, building an execution tree, and optimizing the order of execution plans, the problem of poor execution plans in the prior art is solved, and the efficiency of rule-based entity digestion is improved.
Patent Information
- Application Number
- CN202411569287.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-05
AI Technical Summary
The execution plan generated by the existing execution plan generator is not ideal in execution, resulting in low efficiency in rule-based entity digestion.
By obtaining the rules to be executed, forming a set of rules, determining all predicates, evaluating and sorting all predicates, building an execution tree, and evaluating each connected edge according to the data set to be executed, optimizing the order of the rule execution plan.
By optimizing the execution order of predicates and rules to be executed, the efficiency of entity digestion is significantly improved and unnecessary calculation and resource consumption is reduced.
Smart Images

Figure CN119227682B_ABST
Abstract
Description
Technical Field
[0001] The present application is applicable to the field of big data mining technology, and in particular, relates to a rule-based execution plan generation method, device, equipment and medium. Background Art
[0002] Entity resolution refers to identifying all data in a dataset that points to the same real entity, given a dataset. Entity resolution methods can be mainly divided into two categories: rule-based methods and deep learning model-based methods. In comparison, rule-based entity resolution methods have the advantages of high efficiency and easy explanation. This method uses a series of rules to filter out mismatches in the dataset and identify potential matching entities. In practical applications, the efficiency of rule-based entity resolution depends largely on an effective execution plan. The execution plan determines the order in which rules are evaluated and the data processing flow during the entity resolution process, which can directly affect the efficiency of entity resolution. An effective execution plan can significantly improve the speed of rule matching, reduce unnecessary computing overhead, and thus optimize the overall performance. However, in practical applications, the execution plans generated by most existing execution plan generators are not ideal in terms of execution effect. Therefore, how to optimize the generation of execution plans to improve the efficiency of rule-based entity resolution has become an urgent problem to be solved. Summary of the invention
[0003] In view of this, embodiments of the present application provide a rule-based execution plan generation method, apparatus, device and medium to solve the problem of how to optimize the generation of an execution plan to improve the efficiency of rule-based entity resolution.
[0004] In a first aspect, an embodiment of the present application provides a rule-based execution plan generation method, the execution plan generation method comprising:
[0005] Acquire the rules to be executed to form a rule set, and determine all predicates in the rule set, wherein the rules to be executed include at least one predicate;
[0006] According to the data set to be executed, all predicates are evaluated to obtain predicate evaluation results, and all predicates are sorted according to the predicate evaluation results to obtain predicate sorting results;
[0007] According to the predicate sorting result, construct an execution tree for all the to-be-executed rules in the rule set to obtain at least one execution tree, wherein the execution tree includes a root node and at least one leaf node, a connection edge between any two nodes represents a predicate, and the to-be-executed rules correspond to N connection edges connected sequentially in the execution tree, where N is an integer greater than zero;
[0008] According to the data set to be executed, each connection edge of all execution trees is evaluated to obtain edge evaluation results, and according to the edge evaluation results, the execution order of the nodes of each execution tree is sorted to obtain a rule execution plan corresponding to each execution tree.
[0009] In a second aspect, an embodiment of the present application provides a rule-based execution plan generation device, the execution plan generation device comprising:
[0010] An initial data module, used to obtain rules to be executed to form a rule set, and determine all predicates in the rule set, wherein the rules to be executed include at least one predicate;
[0011] A predicate sorting module is used to evaluate all predicates according to the data set to be executed to obtain predicate evaluation results, and sort all predicates according to the predicate evaluation results to obtain predicate sorting results;
[0012] An execution tree construction module, used to construct an execution tree for all the to-be-executed rules in the rule set according to the predicate sorting result, to obtain at least one execution tree, wherein the execution tree includes a root node and at least one leaf node, the connection edge between any two nodes represents a predicate, and the to-be-executed rules correspond to N connection edges connected in sequence in the execution tree, where N is an integer greater than zero;
[0013] The execution plan generation module is used to evaluate each connection edge of all execution trees according to the data set to be executed, obtain edge evaluation results, sort the execution order of the nodes of each execution tree according to the edge evaluation results, and obtain the rule execution plan corresponding to each execution tree.
[0014] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the execution plan generation method described in the first aspect when executing the computer program.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the execution plan generation method described in the first aspect is implemented.
[0016] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the present application evaluates all predicates in a rule set according to a data set to be executed to obtain a predicate evaluation result, sorts all predicates according to the predicate evaluation result to obtain a predicate sorting result, constructs an execution tree for all rules to be executed in the rule set according to the predicate sorting result to obtain at least one execution tree, evaluates each connecting edge of the execution tree according to the data set to be executed to obtain an edge evaluation result, sorts the execution order of the nodes of each execution tree according to the edge evaluation result to obtain a rule execution plan corresponding to each execution tree.
[0017] Among them, by evaluating the predicates in the rule set, the predicate evaluation results are obtained, and the cost-effectiveness of each predicate in the rule set being evaluated on the data set to be executed is determined. The cost-effectiveness may include the time the predicate is checked on the data set to be executed (evaluation cost) and the probability that the predicate is matched and satisfied on the data set to be executed (evaluation effect). All predicates are sorted according to the predicate evaluation results to obtain the predicate sorting results, thereby determining the order in which all predicates in each rule to be executed are evaluated, and the order strikes a balance between the evaluation cost and the evaluation effect. According to the predicate sorting results, all the rules to be executed are constructed into an execution tree, and the execution tree is sorted according to the data set to be executed. Each connecting edge is evaluated to obtain the edge evaluation result, and the probability of each to-be-executed rule composed of the connecting edge being matched and satisfied on the to-be-executed data set is determined, so that the order in which all to-be-executed rules in the rule set are evaluated is determined according to the probability. Therefore, according to the edge evaluation result and the execution plan generated by the execution tree, the execution order of the predicates and the to-be-executed rules is optimized. It can be determined which predicates to evaluate first according to the cost-effectiveness of the evaluation, and the execution plan of which rules to be evaluated first can be determined according to the evaluation effect, so that entity resolution is performed based on the execution plan, which can more efficiently process the to-be-executed data set, reduce unnecessary calculations and resource consumption, and improve the efficiency of entity resolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0019] Figure 1 This is a schematic diagram of an application environment of a rule-based execution plan generation method provided in Example 1 of the present application;
[0020] Figure 2 It is a flowchart of a rule-based execution plan generation method provided in Example 2 of the present application;
[0021] Figure 3 is a schematic diagram of an execution tree provided in Embodiment 2 of the present application;
[0022] Figure 4 It is a flowchart of a rule-based execution plan generation method provided in Example 3 of the present application;
[0023] Figure 5 It is a flowchart of a rule-based execution plan generation method provided in Example 4 of the present application;
[0024] Figure 6 It is a flowchart of a rule-based execution plan generation method provided in Example 5 of the present application;
[0025] Figure 7 It is a flowchart of a rule-based execution plan generation method provided in Example 6 of the present application;
[0026] Figure 8 It is a structural diagram of a rule-based execution plan generation device provided in Embodiment 7 of the present application;
[0027] Fig. 9 It is a structural diagram of a computer device provided in Example 8 of the present application. DETAILED DESCRIPTION
[0028] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0029] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0030] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0031] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0032] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0033] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0034] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0035] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0036] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0037] In order to illustrate the technical solution of the present application, a specific embodiment is provided below for illustration.
[0038] The first embodiment of the present application provides a rule-based execution plan generation method, which can be applied in the following aspects: Figure 1 In the application environment, the server communicates with the client, the server provides execution plan generation service, and the client triggers the execution plan generation task to the server. The client includes but is not limited to PDA, desktop computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cloud computer equipment, personal digital assistant (PDA) and other devices. The computer equipment corresponding to the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0039] See also Figure 2 , is a flow chart of a rule-based execution plan generation method provided in Example 2 of the present application. The above execution plan generation method is applied to Figure 1 The server in the example connects to the client to obtain the rule set sent by the client. Figure 2 As shown, the execution plan generation method may include the following steps:
[0040] Step S201: Obtain rules to be executed to form a rule set, and determine all predicates in the rule set.
[0041] Step S202: evaluate all predicates according to the data set to be executed to obtain predicate evaluation results, and sort all predicates according to the predicate evaluation results to obtain predicate sorting results.
[0042] In this embodiment, the data set to be executed may refer to the data set on which entity resolution is to be performed, and the data set to be executed includes multiple tuples of record data information, each tuple consists of a single or multiple attributes and corresponding attribute values, an attribute may refer to a variable or field that describes the characteristics of the tuple, and an attribute value may refer to a specific numerical value of the corresponding attribute.
[0043] A rule may refer to a conditional expression composed of logical operators, a predicate defines the specific conditions and results of the rule, a to-be-executed rule may refer to a rule for entity resolution of a to-be-executed data set, wherein the to-be-executed rule includes at least one predicate, and a rule set may refer to a set of to-be-executed rules, which combine multiple comparison methods (equality comparison and similarity comparison) to identify tuples in the to-be-executed data set that point to the same real entity.
[0044] The predicate evaluation result may refer to the result obtained by comprehensively evaluating the time required for evaluating and checking the predicate on the data set to be executed (evaluation cost) and the probability of the predicate being matched and satisfied on the data set to be executed (evaluation effect). The predicate being matched and satisfied may refer to the predicate being satisfied when certain conditions are met. For example, when the predicate is an equality comparison, the corresponding values in the corresponding tuples are completely equal, then the predicate is matched and satisfied. When the predicate is a similarity comparison, the similarity of the corresponding values in the corresponding tuples satisfies a preset threshold, then the predicate is matched and satisfied. The time required for the predicate to be evaluated and checked may refer to the time required to check whether the predicate is matched and satisfied by all tuples on the data set to be executed. The predicate sorting result may refer to the result obtained by sorting all predicates according to the predicate evaluation result.
[0045] For example, rule-based entity resolution can be defined as: given a pattern ,in, For an attribute, is the entity identifier (ID), each A tuple of represents an entity. The dataset to be executed is a set of A tuple of patterns.
[0046] In mode Above, the rule can be defined as ,in, It is about two tuples and The set of predicates, express , For rules The prerequisite is is the result, when the tuple and tuples Satisfy the rules The prerequisite is that the tuple and tuples Represents the same entity.
[0047] rule The predicate in can be defined as ,in, and is a property, c is a constant, and Used to compare whether the A attribute value of tuple t is equal to the B attribute value of tuple s. It is used to compare the similarity between the attribute value of A of tuple t and the attribute value of B of tuple s. The similarity can be measured using various similarity metrics, such as edit distance or Jaccard similarity.
[0048] For example, for a simplified e-commerce dataset, the pattern of the tuples in the e-commerce dataset is: Products (eid, pname (product name), price (price), sname (store name), description (description), color (color), saddress (store address)), for any two tuples in the dataset and , for this tuple and Rules to be executed for entity resolution It can be: and and and ,but ,in Represents the edit distance, that is, if the tuple Colors and tuples The colors are equal, and the tuples The price and tuple The prices are equal, and the tuples The store name and tuple The store names are equal, and the tuple Product name and tuple If the product names are equal, then the tuple and tuples are the same real entity.
[0049] Specifically, for any predicate in the rule set, the total evaluation time of the predicate on the data set to be executed is calculated, and the probability of the predicate being matched and satisfied on the data set to be executed is calculated; based on the total evaluation time and the probability of being matched and satisfied, the predicate evaluation result of the predicate is determined; all predicates in the rule set are traversed to obtain the predicate evaluation result of each predicate; based on the predicate evaluation results, all predicates are sorted to obtain the predicate sorting result.
[0050] Step S203: construct execution trees for all to-be-executed rules in the rule set according to the predicate sorting result to obtain at least one execution tree.
[0051] In this embodiment, the execution plan specifies the order in which each to-be-executed rule in the rule set is evaluated, and the order in which all predicates in each to-be-executed rule are evaluated. The execution tree may refer to a tree representation of the execution plan, wherein the execution tree includes a root node and at least one leaf node, and the connecting edge between any two nodes represents a predicate. The to-be-executed rules correspond to N connecting edges connected in sequence in the execution tree, where N is an integer greater than zero.
[0052] Specifically, for any rule to be executed in the rule set , if the predicate sorting result in X is: ,in, , , , is a predicate. Then according to this rule The process of constructing the execution tree can be as follows: initialize the root node N0 in the execution tree; traverse the execution tree from the root node, and process the predicates in X in turn. At the beginning, suppose that the current traversal reaches node N and the predicate being processed is , check the child nodes of N, if there is a child node Nc, and the edge (N, Nc) represents , move to the child node Nc, and process the next predicate in X , otherwise, create a new child node Nc for N so that the edge (N, Nc) represents , then moves to this new child and processes the next predicate in X , this traversal process continues until all predicates in X are processed and the current node is set as a leaf node. The associated rule is Thus, according to the above process, the predicates of all the to-be-executed rules in the rule set are sorted to construct at least one execution tree.
[0053] For example, Figure 3 The figure shows a schematic diagram of an execution tree provided in the second embodiment of the present application.
[0054] A node in the execution tree is denoted as N, and the root node is denoted as N0; the path starting from the root node It can be expressed as , where (Ni-1,Ni) is a connecting edge in the execution tree, i∈[1,L], The length of is L, that is The number of edges above; each connecting edge represents a predicate and is associated with a score, which is the execution score of the predicate (for details, please refer to the contents in steps S701 to S704). The predicates corresponding to all the connecting edges in the execution tree are: (express )、 (express )、 (express )、 (express )、 (express )、 (express ), where ≈ED represents edit distance, ≈JD represents Jaccard distance, pname (product name), price (price), sname (store name), description (description), color (color), saddress (store address).
[0055] The execution tree includes three rules to be executed, namely , and , if the rule has been processed during the build process ,exist Figure 3 The path is created in the execution tree of , if the rule The corresponding predicate sorting results are: , , then according to the rules The process of constructing the execution tree can be as follows: starting from the root node, first process , since the root node has a label If it is a child node N1, then move to N1 and process , since there is no mark at N1 , then create a new node N5 and mark the edge (N1, N5) as ,because All predicates in have been processed, and node N5 becomes a leaf node, and its associated rule is .
[0056] When the entity resolution of the to-be-executed data set is performed based on the execution plan corresponding to the execution tree, for any tuple pair in the to-be-executed data set, , traverse the execution tree from the root node through depth-first search for evaluation, and on each internal node N of the execution tree, select a child node Nc so that the execution score of the predicate p associated with the edge (N, Nc) is the highest among all the child nodes of N. Then, check whether there is a tuple pair The predicate p matches and satisfies. If so, move to Nc and process Nc in a similar way. If not, check whether N has other unexplored child nodes and process them similarly in descending order of execution score. If all child nodes of N have been explored, return to N's parent node Np and repeat the process. The evaluation process is completed when the first leaf node of the execution tree T is reached. Assume that the rule associated with this leaf node is This means that on the path from the root node to the leaf node, the tuple pair satisfies all predicates in X, so the tuple pair By rules The match is satisfied and the remaining tree traversal can be skipped.
[0057] For example, based on Figure 3 When the execution tree shown in the figure performs entity resolution on the dataset to be executed, for any tuple pair in the dataset to be executed , start traversing from the root node N0 through depth-first search. N0 has two child nodes. First, explore the nodes connected to N0 marked as , because the execution score of this edge is higher. Tuple pair If the match is satisfied, move to node N1. First, the exploration mark connected to N1 is , because the execution score of this edge is higher. Not tuple pairs If the match is satisfied, then explore the mark connected to N1 as If Nor is it a tuple pair If the match is satisfied, then return N1's parent node N0, and explore the mark connected to N0. If Tuple pair If the match is satisfied, move to node N6 and explore the tag connected to N6. ,like Tuple pair If the match is satisfied, it is determined that on the path from the root node to the leaf node N7, the tuple pair Satisfy the rules All predicates in , so the tuple pair By rules Matches satisfy, tuple pairs For the same real entity, the remaining tree traversal can be skipped.
[0058] Step S204, based on the data set to be executed, each connection edge of all execution trees is evaluated to obtain edge evaluation results, and based on the edge evaluation results, the execution order of the nodes of each execution tree is sorted to obtain a rule execution plan corresponding to each execution tree.
[0059] In this embodiment, the edge evaluation result may refer to a result obtained by evaluating the probability that the predicate is matched and satisfied on the to-be-executed data set.
[0060] Specifically, for any execution tree, determine any rule to be executed in the execution tree, and each predicate in the rule to be executed, calculate the execution score of the rule to be executed based on the probability that each predicate in the rule to be executed is matched and satisfied in the data set to be executed, and assign the execution score to each predicate in the rule to be executed. If there is a common predicate shared by at least two rules to be executed in the execution tree, determine the highest score among the execution scores of at least two rules to be executed as the execution score of the common predicate, traverse all execution trees, and determine the edge evaluation result of the connecting edge of the corresponding predicate based on the execution score of each predicate.
[0061] In the embodiment of the present application, the predicates in the rule set are evaluated to obtain the predicate evaluation results, and the cost-effectiveness of each predicate in the rule set being evaluated on the data set to be executed is determined. The cost-effectiveness may include the time the predicate is checked on the data set to be executed (evaluation cost) and the probability that the predicate is matched and satisfied on the data set to be executed (evaluation effect). All predicates are sorted according to the predicate evaluation results to obtain the predicate sorting results, thereby determining the order in which all predicates in each rule to be executed are evaluated, and the order strikes a balance between the evaluation cost and the evaluation effect. According to the predicate sorting results, all the rules to be executed are constructed into an execution tree, and the execution tree is sorted according to the data set to be executed. Each connecting edge of the tree is evaluated to obtain the edge evaluation result, and the probability of each to-be-executed rule composed of the connecting edge being matched and satisfied on the to-be-executed data set is determined, so as to determine the order in which all to-be-executed rules in the rule set are evaluated according to the probability. Therefore, according to the edge evaluation result and the execution plan generated by the execution tree, the execution order of the predicates and the to-be-executed rules is optimized. It is possible to decide which predicates to evaluate first based on the cost-effectiveness of the evaluation, and to decide which execution plans of the to-be-executed rules to evaluate first based on the evaluation effect, so as to perform entity resolution based on the execution plan, which can more efficiently process the to-be-executed data set, reduce unnecessary calculations and resource consumption, and improve the efficiency of entity resolution.
[0062] See also Figure 4 , is a flow chart of a rule-based execution plan generation method provided in Example 3 of the present application. Figure 4 As shown, in the above step S202, all predicates are evaluated according to the data set to be executed to obtain the predicate evaluation result, which may include the following steps:
[0063] Step S401, obtaining a data set to be executed.
[0064] Step S402: for any predicate, use any two tuples in the data set to be executed to calculate the evaluation time of executing the predicate, traverse the data set to be executed, and obtain the total evaluation time.
[0065] In this embodiment, the data set to be executed includes at least two tuples, and the evaluation time may be the time required for checking whether the predicate is matched and satisfied by any two tuples in the data set to be executed. The total evaluation time may refer to the time required for checking whether the predicate is matched and satisfied by all tuples in the data set to be executed.
[0066] Specifically, given a predicate p and a dataset D to be executed, the total evaluation time of predicate p on the dataset D to be executed can be recorded as , and its calculation formula can be:
[0067]
[0068] in, represents the evaluation time of predicate p for tuples t1 and t2, that is, the time required to check whether predicate p is matched by tuples t1 and t2. The actual time may be obtained by checking the actual execution process of the predicate p on tuple t1 and tuple t2, or the estimated time may be obtained by evaluating and calculating the execution process of the predicate p on tuple t1 and tuple t2.
[0069] Step S403: obtaining the probability that the predicate is satisfied in the data set to be executed.
[0070] Step S404, determining the predicate evaluation result of the predicate according to the probability and the total evaluation time, traversing all predicates, and obtaining the predicate evaluation result of each predicate.
[0071] In this embodiment, the probability of being satisfied may refer to the probability that the predicate is matched and satisfied in the to-be-executed data set.
[0072] Specifically, 1 is subtracted from the probability to obtain a probability difference, the probability difference is compared with the total evaluation time to obtain a predicate evaluation result of the predicate whose ratio is the predicate, and all predicates are traversed to obtain a predicate evaluation result of each predicate.
[0073] Among them, the probability that the predicate p is satisfied in the to-be-executed data set D can be recorded as , then according to the probability and the total evaluation time, the calculation formula for determining the predicate evaluation result of the predicate can be: .
[0074] In the embodiment of the present application, for any predicate, using any two tuples in the data set to be executed, the evaluation time of executing the predicate is calculated, the time required to check whether the predicate is matched and satisfied by any two tuples is determined, the data set to be executed is traversed, the time required to check whether the predicate is matched and satisfied by all tuples in the data set to be executed is obtained, the time cost of evaluating the predicate is determined, and the predicate evaluation result of the predicate is obtained in combination with the probability of the predicate being satisfied in the data set to be executed. The predicate evaluation result strikes a balance between the evaluation cost and the evaluation effect. According to the predicate evaluation result, when generating an execution plan, the predicate with a lower total evaluation time and easily matched and satisfied by the data set to be executed can be evaluated first, thereby reducing unnecessary calculation and resource consumption when performing entity resolution based on the generated execution plan, and improving the efficiency of entity resolution. .
[0075] See also Figure 5 , is a flow chart of a rule-based execution plan generation method provided in Example 4 of the present application. Figure 5 As shown, in the above step S402, any two tuples in the to-be-executed data set are used to calculate the evaluation time of the execution predicate, and the to-be-executed data set is traversed to obtain the total evaluation time, which may include the following steps:
[0076] Step S501, obtaining a trained neural network.
[0077] Step S502: input any two tuples and predicates in the data set to be executed into the neural network to obtain the corresponding evaluation time, and traverse all tuples in the data set to be executed to obtain the total evaluation time.
[0078] In this embodiment, the trained neural network is obtained by offline training using a data set with the same data distribution as the data set to be executed, its input is two tuples and a predicate, and its output is the evaluation time.
[0079] In the above step S402, the total evaluation time cost of traversing all tuples in the to-be-executed data set to calculate the predicate is extremely high. For example, on a data set containing 2 million tuples, the total evaluation time cost of calculating It takes more than 100 seconds on average. To improve efficiency and accuracy, a shallow neural network can be trained using a dataset with the same data distribution as the dataset to be executed, that is, a small feedforward neural network to estimate .
[0080] Based on the trained neural network, the total evaluation time of predicate p on the to-be-executed data set D can be recorded as , and its calculation formula can be:
[0081]
[0082] in, represents the evaluation time of predicate p for tuples t1 and t2.
[0083] Corresponding to the above steps S401 to S404, based on the probability that the predicate p is satisfied in the to-be-executed data set D , and the total evaluation time of the predicate p on the dataset D to be executed based on the trained neural network , the calculation formula of the predicate evaluation result of the predicate can be determined as: .
[0084] In the embodiment of the present application, any two tuples and a predicate in the data set to be executed are input into a trained neural network, and the evaluation time of the predicate is output, thereby reducing the overhead caused by the need to actually execute the predicate to measure the time in the traditional method, and improving the evaluation efficiency. By using a data set with the same data distribution as the data set to be executed for offline training, the trained neural network can learn the potential patterns and features in the data, thereby improving the accuracy of the predicted evaluation time.
[0085] See also Figure 6 , is a flowchart of a rule-based execution plan generation method provided in Example 5 of the present application. Figure 6 As shown, obtaining the probability of the predicate being satisfied in the to-be-executed data set in the above step S403 may include the following steps:
[0086] Step S601, obtaining the attribute object in the predicate.
[0087] Step S602: Use locality sensitive hashing to hash the corresponding attribute values of the attribute objects of all tuples into k buckets, and determine the number of tuples in each bucket.
[0088] Step S603, calculating the uniformity of the hash according to the number of buckets and the number of tuples in each bucket, and obtaining a uniformity value as the probability that the predicate is satisfied in the data set to be executed.
[0089] In this embodiment, the attribute object may refer to an attribute for predicate comparison, k is a predefined parameter, and similar or identical attribute values are hashed into the same bucket with a relatively high probability.
[0090] Based on this, the probability that the predicate p is satisfied in the to-be-executed data set D can be recorded as , and its calculation formula can be:
[0091]
[0092] Among them, k is the number of buckets, is the number of tuples in the i-th bucket.
[0093] For the attribute object A compared in predicate p, by quantifying the possibility that tuples t1 and t2 have different or dissimilar values on attribute object A, if the probability that tuples t1 and t2 have different values is high, predicate p is unlikely to be satisfied, so this predicate should be evaluated first because it can determine earlier that the rule corresponding to the predicate is not matched. If all tuples are hashed into the same bucket, it means that the attribute object A values of all tuples are very similar, so predicate p may be satisfied by a large number of tuple pairs, and this type of predicate should have a lower priority.
[0094] In an embodiment of the present application, local sensitive hashing is used to determine the probability that a predicate is satisfied in a data set to be executed. Based on the probability, predicates that are more likely to be unsatisfied can be preferentially evaluated when generating an execution plan. Thus, when performing entity resolution based on the generated execution plan, a large number of unmatched tuple pairs can be excluded at an early stage, thereby reducing unnecessary calculations and resource consumption and improving the efficiency of entity resolution.
[0095] See also Figure 7 , is a flowchart of a rule-based execution plan generation method provided in Example 6 of the present application, such as Figure 7 As shown, in the above step S204, each connection edge of all execution trees is evaluated according to the data set to be executed to obtain the edge evaluation result, which may include the following steps:
[0096] Step S701, for any execution tree, obtain any to-be-executed rule in the execution tree, and determine each predicate in the to-be-executed rule;
[0097] Step S702, calculating the execution score of the rule to be executed according to the probability that each predicate in the rule to be executed is satisfied in the data set to be executed, traversing all the rules to be executed in the execution tree, and assigning the execution score to each predicate in the corresponding rule to be executed;
[0098] Step S703, if there is a common predicate shared by at least two to-be-executed rules in the execution tree, determine the highest score among the execution scores of the at least two to-be-executed rules as the execution score of the common predicate;
[0099] Step S704: traverse all execution trees, and determine the edge evaluation result of the connection edge corresponding to the predicate according to the execution score of each predicate.
[0100] In this embodiment, the execution score may refer to a score that represents the evaluation order of the rules to be executed.
[0101] Specifically, for any execution tree, obtain any rule to be executed in the execution tree, determine each predicate in the rule to be executed, multiply the probability of each predicate in the rule to be executed being satisfied in the data set to be executed, obtain the multiplication result, determine the multiplication result as the execution score of the rule to be executed, traverse all the rules to be executed in the execution tree, and assign the execution score to each predicate in the corresponding rule to be executed.
[0102] The probability that the predicate p is satisfied in the to-be-executed data set D is recorded as The rules to be executed are , rules to be executed The execution score is recorded as ,but The calculation formula can be:
[0103]
[0104] If there is a common predicate shared by at least two rules to be executed in the execution tree, the highest score among the execution scores of the at least two rules to be executed is determined as the execution score of the common predicate, all execution trees are traversed, and the edge evaluation result of the connecting edge of the corresponding predicate is determined according to the execution score of each predicate.
[0105] For example, for Figure 3 The execution tree shown in the figure shows that if the predicate The execution score is 0.4, the predicate The execution score of the rule is 0.2. Execution score =0.4 0.2=0.08, assuming that the rule is also calculated Execution score =0.048, because the rule and rules Has a common predicate , then the common predicate The execution score is max{0.08,0.048}, that is, the final The execution score is 0.08.
[0106] In an embodiment of the present application, the execution score of the execution rule is calculated according to the probability that each predicate in the rule to be executed is satisfied in the data set to be executed, and the corresponding predicate is assigned. If there is a common predicate shared by at least two rules to be executed in the execution tree, the highest score among the execution scores of the at least two rules to be executed is determined as the execution score of the common predicate. According to the execution score, the predicates that are easily matched and satisfied by the data set to be executed can be preferentially evaluated when generating an execution plan. Therefore, when performing entity resolution based on the generated execution plan, a more efficient and reasonable execution path can be selected according to the execution score, thereby reducing unnecessary calculations and resource consumption and improving the efficiency of entity resolution.
[0107] Corresponding to the rule-based execution plan generation method of the above embodiment, Figure 8 The structure block diagram of the rule-based execution plan generation device provided in the seventh embodiment of the present application is shown. The execution plan generation device is applied to Figure 1 For ease of description, only the parts related to the embodiment of the present application are shown.
[0108] See also Figure 8 , the execution plan generating device comprises:
[0109] The initial data module 81 is used to obtain the rules to be executed to form a rule set, and determine all predicates in the rule set, wherein the rules to be executed include at least one predicate;
[0110] A predicate sorting module 82 is used to evaluate all predicates according to the data set to be executed to obtain predicate evaluation results, and sort all predicates according to the predicate evaluation results to obtain predicate sorting results;
[0111] An execution tree construction module 83 is used to construct an execution tree for all the to-be-executed rules in the rule set according to the predicate sorting result, to obtain at least one execution tree, wherein the execution tree includes a root node and at least one leaf node, the connection edge between any two nodes represents a predicate, and the to-be-executed rules correspond to N connection edges connected in sequence in the execution tree, where N is an integer greater than zero;
[0112] The execution plan generation module 84 is used to evaluate each connection edge of all execution trees according to the data set to be executed, obtain edge evaluation results, and sort the execution order of the nodes of each execution tree according to the edge evaluation results to obtain the rule execution plan corresponding to each execution tree.
[0113] Optionally, the predicate sorting module 82 includes:
[0114] A data set acquisition unit, used to acquire a data set to be executed, wherein the data set to be executed includes at least two tuples;
[0115] An evaluation time calculation unit, used for calculating the evaluation time of executing any predicate using any two tuples in the to-be-executed data set, traversing the to-be-executed data set, and obtaining a total evaluation time;
[0116] A probability acquisition unit, used to acquire the probability that the predicate is satisfied in the to-be-executed data set;
[0117] The evaluation result calculation unit is used to determine the predicate evaluation result of the predicate according to the probability and the total evaluation time, and traverse all predicates to obtain the predicate evaluation result of each predicate.
[0118] Optionally, an evaluation time calculation unit comprises:
[0119] A network acquisition subunit is used to acquire a trained neural network, wherein the trained neural network is obtained by offline training using a data set having the same data distribution as the data set to be executed, and its input is two tuples and a predicate, and its output is an evaluation time;
[0120] The network prediction subunit is used to input any two tuples in the data set to be executed and the predicate into the neural network to obtain the corresponding evaluation time, traverse all tuples in the data set to be executed, and obtain the total evaluation time.
[0121] Optionally, the probability acquisition unit includes:
[0122] An attribute acquisition subunit, used to acquire an attribute object in the predicate;
[0123] A hashing subunit, configured to use locality sensitive hashing to hash the attribute values corresponding to the attribute objects of all tuples into k buckets, and determine the number of tuples in each bucket, wherein k is a predefined parameter, and similar or identical attribute values are hashed into the same bucket with a relatively high probability;
[0124] The uniformity calculation subunit is used to calculate the uniformity of the hash according to the number of buckets and the number of tuples in each bucket, and obtain a uniformity value as the probability that the predicate is satisfied in the data set to be executed.
[0125] Optionally, the evaluation result calculation unit includes:
[0126] A difference calculation subunit, used for subtracting 1 from the probability to obtain a probability difference;
[0127] The ratio calculation subunit is used to compare the probability difference with the total evaluation time to obtain a predicate evaluation result whose ratio is the predicate.
[0128] Optionally, the execution plan generation module 84 includes:
[0129] A predicate determination unit, configured to obtain, for any execution tree, any rule to be executed in the execution tree, and determine each predicate in the rule to be executed;
[0130] a score calculation unit, configured to calculate an execution score of the rule to be executed according to a probability that each predicate in the rule to be executed is satisfied in the data set to be executed, traverse all the rules to be executed in the execution tree, and assign the execution score to each predicate in the corresponding rule to be executed;
[0131] a score comparison unit, configured to, if there is a common predicate shared by at least two to-be-executed rules in the execution tree, determine the highest score among the execution scores of the at least two to-be-executed rules as the execution score of the common predicate;
[0132] The evaluation result determination unit is used to traverse all execution trees and determine the edge evaluation result of the connection edge of the corresponding predicate according to the execution score of each predicate.
[0133] Optionally, the score calculation unit includes:
[0134] A product calculation subunit, used for multiplying the probability of each predicate in the to-be-executed rule being satisfied in the to-be-executed data set to obtain a multiplication result;
[0135] The score determination subunit is used to determine the multiplication result as the execution score of the rule to be executed.
[0136] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0137] Fig. 9 This is a schematic diagram of the structure of a computer device provided in Example 8 of the present application. Fig. 9 As shown, the computer device of this embodiment includes: at least one processor ( Fig. 9 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned rule-based execution plan generation method embodiments are implemented.
[0138] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that Fig. 9This is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than those shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.
[0139] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0140] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory may be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be a hard disk of a computer device, and in other embodiments, it may also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the memory may also include both an internal storage unit of the computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program, etc. The memory may also be used to temporarily store data that has been output or is to be output.
[0141] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0142] The present application implements all or part of the processes in the above-mentioned embodiment method, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing the computer program product.
[0143] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0144] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0145] In the embodiments provided in the present application, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0146] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0147] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A rule-based execution plan generation method, characterized in that: The execution plan generation method comprises: Acquire the rules to be executed to form a rule set, and determine all predicates in the rule set, wherein the rules to be executed include at least one predicate; According to the data set to be executed, all predicates are evaluated to obtain predicate evaluation results, and all predicates are sorted according to the predicate evaluation results to obtain predicate sorting results; According to the predicate sorting result, construct an execution tree for all the to-be-executed rules in the rule set to obtain at least one execution tree, wherein the execution tree includes a root node and at least one leaf node, a connection edge between any two nodes represents a predicate, and the to-be-executed rules correspond to N connection edges connected sequentially in the execution tree, where N is an integer greater than zero; According to the data set to be executed, each connection edge of all execution trees is evaluated to obtain edge evaluation results, and according to the edge evaluation results, the execution order of the nodes of each execution tree is sorted to obtain a rule execution plan corresponding to each execution tree; The method of evaluating all predicates according to the data set to be executed to obtain predicate evaluation results includes: Acquire a data set to be executed, where the data set to be executed includes at least two tuples; For any predicate, use any two tuples in the to-be-executed data set to calculate the evaluation time of executing the predicate, traverse the to-be-executed data set, and obtain the total evaluation time; Obtaining the probability that the predicate is satisfied in the to-be-executed data set; Determine the predicate evaluation result of the predicate according to the probability that the predicate is satisfied in the to-be-executed data set and the total evaluation time, traverse all predicates, and obtain the predicate evaluation result of each predicate; Determining the predicate evaluation result of the predicate according to the probability that the predicate is satisfied in the to-be-executed data set and the total evaluation time includes: Subtract 1 from the probability that the predicate is satisfied in the to-be-executed data set to obtain a probability difference value; The probability difference is compared with the total evaluation time to obtain a predicate evaluation result of the predicate whose ratio is the predicate.
2. The execution plan generation method according to claim 1, characterized in that: The using any two tuples in the to-be-executed data set to calculate the evaluation time of executing the predicate, traversing the to-be-executed data set to obtain the total evaluation time, includes: Obtaining a trained neural network, wherein the trained neural network is obtained by offline training using a data set having the same data distribution as the data set to be executed, and its input is two tuples and a predicate, and its output is an evaluation time; Input any two tuples in the data set to be executed and the predicate into the neural network to obtain the corresponding evaluation time, and traverse all tuples in the data set to be executed to obtain the total evaluation time.
3. The execution plan generation method according to claim 1, characterized in that: The obtaining the probability that the predicate is satisfied in the to-be-executed data set includes: Get the attribute object in the predicate; Using locality sensitive hashing, hash the corresponding attribute values of the attribute objects of all tuples into k buckets, and determine the number of tuples in each bucket, where k is a predefined parameter, and similar or identical attribute values are hashed into the same bucket with a relatively high probability; The uniformity of the hash is calculated according to the number of buckets and the number of tuples in each bucket, and the uniformity value obtained is the probability that the predicate is satisfied in the data set to be executed.
4. The execution plan generation method according to claim 1, characterized in that: The step of evaluating each connection edge of all execution trees according to the to-be-executed data set to obtain edge evaluation results includes: For any execution tree, obtain any to-be-executed rule in the execution tree, and determine each predicate in the to-be-executed rule; Calculate the execution score of the rule to be executed according to the probability that each predicate in the rule to be executed is satisfied in the data set to be executed, traverse all the rules to be executed in the execution tree, and assign the execution score to each predicate in the corresponding rule to be executed; If there is a common predicate shared by at least two to-be-executed rules in the execution tree, determining the highest score among the execution scores of the at least two to-be-executed rules as the execution score of the common predicate; Traverse all execution trees and determine the edge evaluation results of the connection edges of the corresponding predicates according to the execution score of each predicate.
5. The execution plan generation method according to claim 4, characterized in that: The step of calculating the execution score of the rule to be executed according to the probability that each predicate in the rule to be executed is satisfied in the data set to be executed includes: Multiply the probability that each predicate in the to-be-executed rule is satisfied in the to-be-executed data set to obtain a multiplication result; The multiplication result is determined to be the execution score of the rule to be executed.
6. A rule-based execution plan generation device, characterized in that: The execution plan generating device comprises: An initial data module, used to obtain rules to be executed to form a rule set, and determine all predicates in the rule set, wherein the rules to be executed include at least one predicate; A predicate sorting module is used to evaluate all predicates according to the data set to be executed to obtain predicate evaluation results, and sort all predicates according to the predicate evaluation results to obtain predicate sorting results; An execution tree construction module, used to construct an execution tree for all the to-be-executed rules in the rule set according to the predicate sorting result, to obtain at least one execution tree, wherein the execution tree includes a root node and at least one leaf node, the connection edge between any two nodes represents a predicate, and the to-be-executed rules correspond to N connection edges connected in sequence in the execution tree, where N is an integer greater than zero; An execution plan generation module is used to evaluate each connection edge of all execution trees according to the data set to be executed, obtain edge evaluation results, sort the execution order of the nodes of each execution tree according to the edge evaluation results, and obtain a rule execution plan corresponding to each execution tree; The predicate sorting module comprises: A data set acquisition unit, used to acquire a data set to be executed, wherein the data set to be executed includes at least two tuples; An evaluation time calculation unit, used for calculating the evaluation time of executing any predicate using any two tuples in the to-be-executed data set, traversing the to-be-executed data set, and obtaining a total evaluation time; A probability acquisition unit, used to acquire the probability that the predicate is satisfied in the to-be-executed data set; An evaluation result calculation unit, used to determine the predicate evaluation result of the predicate according to the probability that the predicate is satisfied in the to-be-executed data set and the total evaluation time, and traverse all predicates to obtain the predicate evaluation result of each predicate; The evaluation result calculation unit comprises: a difference calculation subunit, configured to obtain a probability difference by subtracting 1 from the probability that the predicate is satisfied in the to-be-executed data set; The ratio calculation subunit is used to compare the probability difference with the total evaluation time to obtain a predicate evaluation result whose ratio is the predicate.
7. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the execution plan generation method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the execution plan generation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data entity identification method and device, computer equipment and storage medium
CN114780528A
Approximate query optimization system based on machine learning
CN114911844A