A distributed query service optimization method
Patent Information
- Application Number
- CN202411122364.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-08-15
AI Technical Summary
[0004]然而,在系统API层面,尤其是分布式查询(如GraphQL)时,在实现查询时存在多个API接口存在于不同的服务器上,如果使用上述方式进行查询,存在冗余查询的问题,导致查询延迟增加、系统响应变慢等问题
[0010]基于每个所述查询执行属性,从所述至少一个候选连接树中确定目标连接树;
Smart Images

Figure CN119003617B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer processing technology, and in particular to an optimization method for distributed query services. Background Technology
[0002] In the current field of database and query technology, query optimization techniques have made significant progress, especially in the SQL database domain. Traditional query optimization research mainly revolves around determining query strategies based on cardinality estimation and cost assessment to improve query performance on structured databases.
[0003] Existing cardinality estimation methods typically rely on database statistics and simple assumptions to quickly estimate query cardinality and query paths. Cost estimation, on the other hand, builds upon cardinality estimation and aims to predict query execution latency within the database system by comprehensively considering query statements and database structure information, and then adjust the query search path strategy based on the execution latency.
[0004] However, at the system API level, especially in distributed queries (such as GraphQL), multiple API interfaces exist on different servers when implementing queries. If the above method is used for querying, there will be redundant queries, which will lead to increased query latency and slower system response. Summary of the Invention
[0005] This invention provides an optimization method for distributed query services to improve query efficiency and enhance user experience.
[0006] According to one aspect of the present invention, an optimization method for distributed query services is provided, the method comprising:
[0007] Obtain a distributed query statement and determine a first vector of at least two first components and a second vector of a second component in the distributed query statement; wherein the first component includes a query interface; and the second component includes at least object attributes, entity type, and conditional statements;
[0008] Based on the at least two first components, a tree set is determined; wherein the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent the first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected.
[0009] For each tree to be connected, at least one candidate tree corresponding to the current tree to be connected is determined from the tree set, and the query execution attributes when the current tree to be connected is determined based on the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector;
[0010] Based on each of the query execution attributes, a target join tree is determined from the at least one candidate join tree;
[0011] Based on the current tree to be connected and the target tree, the target query path is determined, and the distributed query statement is executed based on the target query path to obtain the target data.
[0012] The technical solution of this invention involves obtaining a distributed query statement and determining the first vectors of at least two first components and the second vectors of the second components in the distributed query statement; determining a tree set based on the at least two first components; the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected; for each tree to be connected, determining at least one candidate connection tree corresponding to the current tree to be connected from the tree set, and determining the query execution attributes when the current tree to be connected is connected to the candidate connection tree based on the tree vectors of the current tree to be connected, the tree vectors of the candidate connection trees, and the second vectors; determining a target connection tree from at least one candidate connection tree based on each query execution attribute; and determining the target connection tree based on the current tree to be connected and the target connection tree. This method determines the target query path and executes a distributed query statement based on it to obtain the target data. It solves the problem of low query efficiency and slow system response caused by existing technologies that adjust path strategies based on query statements and database structure information to predict execution latency. The method constructs a tree set by integrating the first vectors of at least two first components in the distributed query statement. Then, it combines the second vector of the second component with the trees to be connected in the tree set to determine the next candidate connection tree for each tree to be connected. It predicts the query execution attributes when connecting the current tree to a candidate tree, selects the target connection tree from at least one candidate tree based on these attributes, and continues until the target query path is determined. The distributed query statement is then executed using the target query path, improving query efficiency, reducing query latency, and enhancing system response, thereby improving the user experience.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart of an optimization method for distributed query services provided according to Embodiment 1 of the present invention;
[0016] Figure 2 This is a flowchart of an optimization method for distributed query services provided according to Embodiment 2 of the present invention;
[0017] Figure 3 This is a schematic diagram of the distributed query service optimization method provided in Embodiment 2 of the present invention;
[0018] Figure 4 This is a flowchart of an optimization method for distributed query services provided according to Embodiment 3 of the present invention;
[0019] Figure 5 This is a schematic diagram of the distributed query service optimization method provided in Embodiment 3 of the present invention;
[0020] Figure 6 This is a schematic diagram of the distributed query service optimization method provided in Embodiment 3 of the present invention;
[0021] Figure 7 This is a flowchart of an optimization method for distributed query services provided according to Embodiment 4 of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] Example 1
[0025] Figure 1 This is a flowchart of an optimization method for distributed query services according to Embodiment 1 of the present invention. This embodiment is applicable to distributed query scenarios. The method can be executed by a distributed query device, which can be implemented in hardware and / or software and can be configured in a computing device. Figure 1 As shown, the method includes:
[0026] S110. Obtain the distributed query statement and determine the first vector of at least two first components and the second vector of the second component in the distributed query statement.
[0027] The distributed query statement refers to a statement used for distributed queries. For example, a distributed query statement could be a GraphQL query statement. GraphQL is a query language for APIs (Application Programming Interfaces) that allows clients to specify which data to retrieve from different servers and the format of that data. The first component includes the query interface. The query interface can be an interface used to provide data for querying applications, and these interfaces can be distributed across different servers. The second component includes at least object attributes, entity types, and conditional statements. Object attributes refer to the attribute information of the query object, which can be the entity or collection of data targeted by the query operation. For example, object attributes can include, but are not limited to, table type, data type (such as numeric, character, etc.), column and row of the object. The entity type defines the structure of data (such as tables and columns in a database), which includes object attributes. For example, in the GraphQL schema, the query statement contains an entity type named User, which defines the structure of a user, including fields such as id, name, and email. These fields are the object attributes of that entity type. Conditional statements are used to filter data and return results that meet specific conditions.
[0028] In practical applications, upon receiving a distributed query statement from a client, the fields in the distributed query statement can be parsed to obtain the components of the distributed query statement. These components include object attributes, conditional statements, entity types, and query interfaces. Then, encoding techniques can be used to extract features from each first component and each second component, converting them into feature vectors. The feature vector of the first component becomes the first vector, and the feature vector of the second component becomes the second vector.
[0029] In this embodiment, determining the first vector of at least two first components in the distributed query statement includes: for each first component, determining the third vector of the interface function name of the first component; determining the fourth vector of the interface function parameters in the first component; determining the fifth vector of the return value corresponding to the first component; and determining the first vector of the first component based on the third vector, the fourth vector, and the fifth vector.
[0030] The interface function parameters can be entity types. It should be noted that the method for determining the first vector of each first component is the same; we can take the determination of the first vector of any one of the first components as an example for explanation.
[0031] Specifically, the interface function name and parameters of the first component can be extracted, and the results and return value names returned by the interface function can be determined. Then, the extracted information can be encoded separately: the encoded vector of the interface function name becomes the third vector, and the encoded vector of the interface function parameters becomes the fourth vector. The encoded vectors of the return results and the return value names are then combined to obtain the fifth vector of the return value. Furthermore, the third, fourth, and fifth vectors can be concatenated to obtain the first vector of the first component.
[0032] For example, the way to determine the first vector R(fun) of the query interface can be represented as:
[0033] R(fun)=(F(fun name ), F(fun param ), F(fun ret ));
[0034] Among them, F(fun) name ) is the third vector of interface function names; fun param =R(t) is the fourth vector of the interface function parameters, fun param = R(t), where R(t) represents the feature vector of entity type t;
[0035] F(fun ret ) is the fifth vector returned, F(fun ret )=(F(ret val ), F(ret) name )), F(ret val F(ret) is the encoded vector of the returned result. name ) is the encoded vector of the return value name.
[0036] In this embodiment, the third vector of the interface function name can be determined using a pre-trained vector determination model. For example, the vector determination model can be trained based on multiple interface function names and the feature vectors corresponding to the interface function names, enabling the vector determination model to learn the feature vectors of the interface function names. After training, the interface function name can be input into the vector determination model to obtain the third vector of the interface function name.
[0037] It should be noted that when determining the fifth vector of the return value, if the return value is a function of type object collection, then each column of the object can be represented as a vector, and these vector representations of the columns can be integrated to obtain the fifth vector of the return value.
[0038] For example, the length of the probability distribution vector for all attributes can be set to 512. For column c1 of numerical attributes, the range of values in this column is divided into 512 intervals. The frequency of values in each interval is counted, and the frequency of values in each interval is compared with the total frequency to obtain the probability of each interval. The probability of each interval is used as an element of the distribution vector and normalized to form the probability distribution as the vector representation R(c1) of this column. The encoding vector F(ret) of the returned result is determined based on the vector representation of this column. val For columns with character attributes, you can count all character types in the column (such as fixed-length character types, variable-length character types, text data types, enumeration types, etc.), then calculate the distribution vector of all strings in each type, concatenate the distribution vectors of all strings, and then perform a pooling operation to obtain an encoding vector F(ret) of shape (1, 512) for the returned result. val The vector representations of all column names in the returned result are also concatenated and then pooled to obtain an encoded vector F(ret) of the return value name with shape (1, 512). name Furthermore, based on the encoded vector F(ret) of the returned result... val The encoded vector F(ret) of the return value name and the return value name name Determine the fifth vector of the return value.
[0039] S120. Based on at least two first components, determine a tree set; wherein the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent the first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected.
[0040] A single-node tree can be a tree structure with only one node, where the leaf node and the root node are the same node. A tree vector can be used to represent the connection between the entire tree; it is a vector representation of the connection between the entire tree.
[0041] In this embodiment, each first component can be treated as a node, and each node can be treated as a tree to be connected, thus each tree to be connected is a single-node tree. The number of trees to be connected is the same as the number of first components. Furthermore, the first vector of the first component can be used as the tree vector of the tree to be connected corresponding to the first component. All the trees to be connected form a tree set, so as to plan the query path to be executed by the query statement through the trees to be connected and their tree vectors in the tree set.
[0042] S130. For each tree to be connected, determine at least one candidate tree corresponding to the current tree to be connected from the tree set, and determine the query execution attributes when the current tree to be connected is connected to the candidate tree based on the tree vector of the current tree to be connected, the tree vector of the candidate tree, and the second vector.
[0043] It should be noted that the method for determining the target query path corresponding to each tree to be joined is the same. Taking any tree to be joined as the current tree as an example, the next tree to be joined is connected sequentially from the current tree; this next tree is the candidate tree, until the target query path starting from the current tree is obtained. Query execution attributes can be used to characterize the query execution situation when executing each first component under the connection of the current tree to the candidate tree. For example, query execution attributes can be estimated based on cumulative query duration, query rate, query latency, etc.
[0044] In this embodiment, a candidate connection tree can be selected from the tree set that is distinct from the current connection tree. Alternatively, pre-configured component connection rules can be used, which include, but are not limited to, candidate component identifiers corresponding to each target component identifier and component connection relationships between components; wherein, the component identifier can be used to identify the uniqueness of the interface function (i.e., the first component). The component connection relationship can refer to the order in which components are connected. Each target component identifier can correspond to one or more candidate component identifiers. Specifically, a target component identifier that matches the component identifier of the first component corresponding to the current connection tree can be found, and then the connection tree corresponding to the first component of the candidate component identifier of the found target component identifier can be selected as a candidate connection tree. It should be noted that the component connection relationship corresponding to the component identifier pair is related to the query interface return result, and the component connection relationship between components can be configured based on the result returned by each query interface. The advantage of this setting is that by pre-configuring the component connection rules, the order in which some first components are connected can be pre-defined, and this connection order is the order in which the query statement calls the interface, improving the efficiency of query path planning while ensuring the correctness of the returned data.
[0045] For example, in a GraphQL query environment, the enumeration of query paths begins with an initial set of trees, and each candidate tree corresponding to the current tree to be joined is considered a possible candidate action. The predicted execution latency based on joining candidate trees from the current tree to be joined is used to evaluate the cost of each candidate action, thus obtaining the query execution attributes when joining candidate trees from the current tree to be joined.
[0046] S140. Based on each query execution attribute, determine the target join tree from at least one candidate join tree.
[0047] In this embodiment, the query execution attributes when connecting each candidate connection tree to the current connection tree can be combined to determine a candidate connection tree as the target connection tree.
[0048] Optionally, based on each query execution attribute, a target connection tree is determined from at least one candidate connection tree, including: selecting the candidate connection tree corresponding to the smallest query execution attribute as the target connection tree, removing the target connection tree from the tree set, and updating the tree set.
[0049] For example, the query execution attributes are sorted. In each step of action expansion, an enumeration method is used to select the candidate join tree with the smallest query execution attribute from all candidate join trees as the optimal candidate action item, i.e., the target join tree, for expansion in the i-th step. After completing the action expansion for this step, the action candidate used for expansion (i.e., the target join tree) is deleted, and the above steps are repeated until the target query path is output, at which point the entire enumeration process ends. The advantage of this setup is that it achieves optimal query performance and improves query efficiency.
[0050] S150. Based on the current tree to be connected and the target tree, determine the target query path.
[0051] In this embodiment, the current tree to be connected and the target tree to be connected can be joined to obtain a connected tree structure, which can then be used to determine the target query path. For example, both the current tree to be connected and the target tree to be connected can be used as leaf nodes to construct a binary tree, which can then be used as the connected tree structure.
[0052] S160. Execute a distributed query statement based on the target query path to obtain the target data.
[0053] After the distributed query statement is input into the parser, the parser can query the first component of the distributed server in sequence according to the connection order between the first components in the target query path, retrieve data from the distributed server, and obtain the final target data.
[0054] The technical solution provided in this embodiment obtains a distributed query statement and determines the first vectors of at least two first components and the second vectors of the second components in the distributed query statement; based on the at least two first components, a tree set is determined; the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected; for each tree to be connected, at least one candidate connection tree corresponding to the current tree to be connected is determined from the tree set, and based on the tree vectors of the current tree to be connected, the tree vectors of the candidate connection trees, and the second vectors, the query execution attributes when the current tree to be connected is connected to the candidate connection tree are determined; based on each query execution attribute, a target connection tree is determined from at least one candidate connection tree; based on the current tree to be connected and the target connection tree... This method identifies the target query path and executes distributed query statements based on it to obtain the target data. It addresses the problem in existing technologies where path adjustment strategies rely on predicting execution latency using query statements and database structure information, leading to low query efficiency and slow system response. The method constructs a tree set by integrating the first vectors of at least two first components in the distributed query statement. This, combined with the second vector of the second component and the trees to be connected in the tree set, determines the next candidate connection tree for each tree to be connected. It also predicts the query execution attributes when connecting the current tree to a candidate tree. Based on these attributes, it selects the target connection tree from at least one candidate tree until the target query path is determined. The distributed query statement is then executed using this target query path, improving query efficiency, reducing query latency, and enhancing system response, thereby improving the user experience.
[0055] Example 2
[0056] Figure 2 This is a flowchart of a distributed query service optimization method according to Embodiment 2 of the present invention. Based on the foregoing embodiments, the method of "determining the second vector of the second component in the distributed query statement" is further refined. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0057] like Figure 2 As shown, the method specifically includes the following steps:
[0058] S210. When the second component includes object attributes, determine the second vector of the object attributes based on the first sub-attribute and the first preset dimension of the object attributes.
[0059] The first sub-attribute includes at least the attribute type, whether the attribute is empty, whether the attribute is a list type, the first sequential number of the object attribute in its corresponding entity type, and the attribute name.
[0060] In this embodiment, taking a certain object attribute as an example, the object attribute can be encoded according to each first sub-attribute corresponding to the object attribute to obtain a second vector of the object attribute under a first preset dimension.
[0061] See Figure 3 We can first encode the object attribute c into a 10+512 bit one-dimensional vector, where the one-dimensional vector F(c) = (c typ ,c non ,c list ,c ind ,c name ); where c typ Occupies two binary bits, c typ The attribute type used to represent object properties (e.g., 00 represents an object type, 10 represents a numeric scalar, and 11 represents a character scalar); c non Occupies one binary bit, c non Used to indicate whether an attribute is empty; c list Occupies one binary bit, c list Used to indicate whether an attribute is a list type; c ind Occupies 6 binary bits, c ind Used to indicate the first sequential number of an object attribute in its corresponding entity type; c name Occupies 512 binary bits, c name Used to represent attribute names. Define a (10+512)×d for object attribute c. hid The matrix M(c), d hid The dimension of the second vector representing the object's attributes (i.e., the first preset dimension). Further, the one-dimensional vector F(c) and matrix M(c) are multiplied to obtain 1×d. hid The attribute representation of the dimension object attribute c is R(c), which is the second vector of the object attribute c. R(c) can be expressed as: R(c) = F(c) × M(c).
[0062] S220. When the second component includes a conditional statement, determine the second vector of the conditional statement based on the second sub-attribute and the second preset dimension of the conditional statement.
[0063] The second sub-attribute includes the attribute type of the object attribute in the conditional statement, the second ordinal number of the object attribute in its corresponding entity type, and at least one of the operators and operands. Operators include, but are not limited to, greater than, less than, equal to, greater than or equal to, and less than or equal to.
[0064] See also Figure 3 The conditional statement `con` can be encoded into a 13+512 bit one-dimensional vector `F(con)`.
[0065] Where, F(con)=(c typ ,c ind ,c opt ,c = ,c > ,c < c val );c typ It occupies two binary bits and is used to represent the attribute type of the object involved in the conditional statement; c ind It occupies six binary bits and is used to represent the second sequential number of the object attribute in the entity type; c opt Occupying two binary bits, it is used to represent the operator in the conditional statement: 11 for equal to, 10 for greater than, 01 for less than, and 00 for other. If the character operand v appears in the conditional statement, c is provided by the pre-trained model. val value, c val The length is 512 binary bits. A (13+512)×d value can also be defined for the conditional statement `con`. con The matrix M(con), d con This is the dimension of the second vector in the conditional statement (i.e., the second preset dimension). Further, multiplying the one-dimensional vector F(con) and the matrix M(c) yields 1×d. con The dimension of the conditional statement is represented by R(con), which is the second vector of the conditional statement con. R(con) can be expressed as: R(con) = F(con) × M(con).
[0066] It should be noted that if the numeric operand v appears in the conditional statement con, formulas (1), (2), and (3) can be used to process the equal to =, greater than >, and less than < operators respectively, to obtain c in F(con). = c > c < The corresponding value.
[0067] Formula (1) can be:
[0068] Formula (2) can be:
[0069] Formula (3) can be:
[0070] It should also be noted that for conditional statements connected by "OR", which include multiple sub-statements, such as the conditional statement con = con1 OR con2, max pooling can be performed on the second vector of each sub-statement to obtain the second vector R(con) of the conditional statement.
[0071]
[0072] For conditional statements connected by "AND", such as the conditional statement con = con1 AND con2, the second vector of each sub-statement in the conditional statement can be subjected to average pooling to obtain the second vector R(con) of the conditional statement. For conditional statements with multiple nested conditions, each condition can be processed one by one according to the nesting order of the conditions in the conditional statement to generate the second vector of the conditional statement.
[0073] S230. When the second component includes an entity type, determine the second vector of the entity type based on the third sub-attribute of the entity type.
[0074] The third sub-attribute includes, but is not limited to, the object attribute of the entity type, the entity type name, and the third sequence number of the entity type.
[0075] See also Figure 3 For an entity type t with k object attributes, the second vector R(t) of its entity type can be determined by concatenating the second vectors (R(c1), R(c2), ..., R(ck)) of the k object attributes into a matrix M(t), where the matrix dimension of M(t) is k×h. s Then, the matrix M(t) is input into the average pooling layer to obtain a 1×d matrix. hid The one-dimensional vector F(t) can be used, and the kernel dimension in the average pooling layer can be k×1; a pre-trained model is used to obtain the feature vector F(t) of the entity type name. name The binary value of the third sequential number of the entity type is represented as F(t). ide ), and then, F(t) ide ) and F(t name The concatenation vector is obtained by concatenating the concatenation vector with F(t). The concatenation vector is then concatenated with the matrix M(t) and processed by the feedforward layer and the ReLU activation function to obtain the representation R(t) of entity type t. R(t) is the second vector of entity type t.
[0076] For example, the second vector R(t) of entity type t can be determined based on formula (4); formula (4) can be expressed as: R(t)=Relu([F(t),F(t)) ide ),F(t name )])×(M(t)); F(t)=
[0077] MaxPool(M(t)). ReLU represents the activation function, and MaxPool represents the max pooling function.
[0078] It should be noted that S210 to S230 can be executed sequentially or in parallel. The specific execution order is not limited. The above order is only the order in which the technical solutions in each step are explained, not the execution order of each step.
[0079] The technical solution of this embodiment determines the second vector of different second components by combining the sub-attributes of specific second components. This enables the determination of the query execution attribute when the current tree to be connected to the candidate tree by mining the second vector of the second component and the vector representation of the connection tree, thereby improving the accuracy of execution attribute prediction and achieving optimal path query execution attribute.
[0080] Example 3
[0081] Figure 4 This is a flowchart of a distributed query service optimization method according to Embodiment 3 of the present invention, which further refines "S130" based on the foregoing embodiments. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0082] like Figure 4 As shown, the method specifically includes the following steps:
[0083] S310. Obtain the distributed query statement and determine the first vector of at least two first components and the second vector of the second component in the distributed query statement.
[0084] S320. Based on at least two first components, determine a tree set; wherein the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent the first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected.
[0085] S330. Input the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector into the pre-trained cost prediction model to obtain the connection vector when the current tree to be connected is connected to the candidate tree.
[0086] It should be noted that the connection vector can be used to represent the tree vector of the tree structure formed when the current tree to be connected is connected to the candidate trees.
[0087] For example, for ml connection trees, to effectively represent the combined structure of these connection trees (i.e., a connection forest), a root node can be introduced as the representative of the connection forest. The root nodes of these ml connection trees then become the leaf nodes of the root, thus forming a hierarchical tree structure. When processing the tree vectors of these connection trees, the Child-SumTree-LSTM model can be used. By processing the multiple input connection trees T... i tree vectors By performing aggregation, we can obtain the vector representation R(S) of the root node, which can be used as the connection vector when connecting multiple connection trees.
[0088] In this embodiment, the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector are input into a pre-trained cost prediction model to obtain the connection vector when the current tree to be connected is connected to the candidate tree to be connected. This includes: processing the tree vector of the current tree to be connected and the tree vector of the candidate tree to be connected based on the second vector of the entity type according to the query unit in the cost prediction model to obtain a first result corresponding to the current tree to be connected and a second result corresponding to the candidate tree to be connected; processing the first result and the second result based on the second vector of the conditional statement according to the row filtering unit in the cost prediction model to obtain a third result corresponding to the current tree to be connected and a fourth result corresponding to the candidate tree to be connected; processing the third result and the fourth result based on the second vector of the object attribute according to the column mapping unit in the cost prediction model to obtain a fifth result corresponding to the current tree to be connected and a sixth result corresponding to the candidate tree to be connected; wherein, the object attribute includes a connection attribute; and determining the connection vector when the current tree to be connected is connected to the candidate tree according to the second vector of the connection attribute, the fifth result, and the sixth result according to the connection unit in the cost prediction model.
[0089] The join attribute can be a column or field that exists in two or more tables and is used to establish the relationship between these tables. The join attribute can be configured in the component's join rules. The filter row unit is used to filter rows that meet specific conditions. The map column unit is used to select the columns to return. The join unit is used to merge two or more input data.
[0090] In this embodiment, the tree vector of the current tree to be connected can be input to the query unit. The query unit outputs the first result corresponding to the current tree to be connected based on the input tree vector and the second vector of the entity type. The first result may include hidden state and memory information.
[0091] For example, the query unit can be based on a Tree-LSTM architecture to facilitate subsequent units receiving hidden state and memory information. See also Figure 5The query unit receives the tree vector R(fun) of the current tree to be connected, processes R(fun), outputs the hidden state h and memory information m of the current tree to be connected, and passes the hidden state h and memory information m to the filter row unit.
[0092] The calculation function for the query cell can be expressed as:
[0093] i=σ([F(fun name ),F(fun param ),F(ret name )]×W i +b i )
[0094] u=tanh([F(fun name ),F(fun param ),F(ret name )]×W u +b u )
[0095] j=σ(F(ret val )×W j +b j )
[0096]
[0097] o=σ([F(fun name ),F(t name ),F(ret name ),F(ret val )]×W o +b o )
[0098]
[0099]
[0100] Where W represents a matrix, b represents the bias parameter, and σ represents the model parameter. The third vector F(fun) of the interface function name in the tree vector R(fun) is... name The fourth vector of the interface function parameters, fun param The fifth vector returned is F(fun) ret The input gates i and j are used to compute the input gates. The encoded vector F(ret) that returns the result in R(fun) is... val ), F(fun) name ), the encoding vector F(ret) of the return value name name The feature vector F(t) of the entity type name and entity type name nameThe input gates i and j are used to calculate the gating state o of the output gate. After applying the input gates i and j to u and v respectively, the memory information m is obtained, and then the hidden state h is calculated based on m and o.
[0101] Furthermore, the first result and the second vector of the conditional statement can be input into the filter row unit to output the third result of the current tree to be connected. The third result includes the updated hidden state and memory information.
[0102] See also Figure 5 The filter row unit receives the hidden state h of the current tree to be connected from the output of the query unit. j-1 and memory information m j-1 And the second vector R(con) of the query statement, and then, based on the hidden state h j-1 Using the second vector R(con), the outputs of input gate i, input gate k, forget gate f, and output gate o are calculated. Simultaneously, the candidate value g of the memory unit can be calculated based on the second vector R(con) of the query statement. Based on the outputs of input gate i, input gate k, forget gate f, the candidate value g of the memory unit, and the memory information m... j-1 The new memory information m is obtained by calculating element-wise multiplication and element-wise addition. j Then, the output gating o and the new memory information m are applied. j Calculate the new hidden state h j New memory information m j and the new hidden state h j This is the second result output by the filter row unit. The filter row unit can be used to add different conditions to the same data source for data filtering. It can also adapt to different data sources and perform diverse condition filtering if the current tree to be joined is changed.
[0103] The calculation function for the filter row cell can be expressed as:
[0104] i = σ([h j-1 ,R(con)]×W i +b i )
[0105] g = tanh(R(con) × W) g +b g )
[0106] k=σ([h j-1 ,R(con)]×W k +b k )
[0107] f=σ([h j-1 ,R(con)]×W f +b f)
[0108] o=σ([h j-1 ,R(con)]×W o +b o )
[0109] u = tanh([h j-1 ,R(con)]×W u +b u )
[0110]
[0111]
[0112] Furthermore, the third result and the second vector of object attributes can be input into the mapping column cell to output the fifth result corresponding to the current tree to be connected. The fifth result includes the updated hidden state and memory information. The object attributes can be the object attributes of the object to be returned in the distributed query statement.
[0113] See also Figure 5 The mapping column unit receives the hidden state h sent from the filtering row unit. j-1 and memory information m j-1 and the object attribute c specified in the query statement to be returned. i ,i∈1,…,l, where l is the number of attributes, and c is the join attribute used for subsequent attribute concatenation. join During the calculation of the mapping column cells, the second vector R(c) of all object attributes can be used. i ,), perform cascading to obtain R; then, calculate R and the hidden state h j-1 The product of these two vectors yields the attention weight β, which represents the importance of each object attribute. Next, the attention weights are summed with the second vector of object attributes to obtain a weighted representation of the object attributes. Based on the hidden state h j-1 Weighted representation The input gate i, input gate k, forget gate f, and output gate o, as well as the output of memory unit u, are calculated in the mapped column unit. Then, based on the weighted representation of object attributes... The candidate value g of the memory unit is calculated. Based on the input gate i, input gate k, forget gate f, the candidate value g of the memory unit, and the memory information m... j-1 The new memory unit m is calculated by element-wise multiplication and addition. j Then, the output gate o and the new memory unit m are applied. j Calculate the new hidden state h j New memory unit m jand the new hidden state h j This is the fifth result output by the mapping column unit.
[0114] The calculation function for the mapped column cells can be expressed as:
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123] It should be noted that the method for determining the second result of the candidate connection tree is similar to that for determining the first result, the method for determining the fourth result of the candidate connection tree is similar to that for determining the third result, and the method for determining the sixth result of the candidate connection tree is similar to that for determining the fifth result, which will not be elaborated here.
[0124] Furthermore, the fifth and sixth results, along with the second vector of the connection attribute corresponding to the connection tree, can be input into the connection unit to obtain the connection vector when the current tree to be connected is connected to the candidate connection tree.
[0125] See also Figure 5 The connection unit receives two mapped column units node. j-1,0 and node j-1,1 The hidden state h sent j-1,0 and h j-1,1 and memory information m j-1,0 and m j-1,1 The hidden state is concatenated with the second vector of connection attributes, and the input gate i, two forget gates f0 and f1, the output gate o, and the memory candidate value u are obtained through linear transformations and activation functions with different parameters. The memory information m is then... j-1,0 ,m j-1,1 As candidate values for memory cells, calculate the new memory cell m. j Furthermore, output gate o and a new memory unit m are applied. j Calculate the new hidden state h j New hidden state h jIt can be used as a connection vector when connecting the current tree to the candidate tree.
[0126] The calculation function for the connection element can be expressed as:
[0127] i = σ([h j-1,0 ,R(c1)×W1,h j-1,1 ,R(c2)×W2]×W i +b i )
[0128]
[0129]
[0130] o=σ([h j-1,0 ,R(c1)×W1,h j-1,1 ,R(c2)×W2]×W o +b o )
[0131] u = tanh([h j-1,0 ,R(c1)×W1,h j-1,1 ,R(c2)×W2]×W u +b u )
[0132]
[0133]
[0134] S340. Based on the second vector and the connection vector, determine the action state vector when the current tree to be connected connects to the candidate tree.
[0135] Among them, the action state vector can be used to represent the vector representation of the current tree to be connected to the candidate tree under the condition of a distributed query statement.
[0136] In this embodiment, the method for determining the action state vector when the current tree to be connected is connected to the candidate tree based on the second vector and the connection vector can be as follows: based on the second vector, determine the statement vector corresponding to the distributed query statement; based on the connection vector and the statement vector, determine the action state vector when the current tree to be connected is connected to the candidate tree.
[0137] In this embodiment, a vector representation that can characterize the distributed query statement can be determined by fusing the second vector of the second component, and this vector is used as the statement vector. Then, the connection vector and the statement vector can be concatenated to obtain the action state vector when the current tree to be connected connects to the candidate tree.
[0138] For example, the connection vector is R(F), the statement vector of the distributed query statement is R(q), and the action state vector is R(S) = [R(F), R(q)]. R(S) reflects the joint representation of the current tree to be connected and the candidate tree to be connected.
[0139] In this embodiment, determining the statement vector corresponding to the distributed query statement based on the second vector includes: constructing a connection graph based on the second vector of each second component; determining the line graph corresponding to the connection graph; processing the connection graph and the line graph based on a graph attention network to obtain an updated connection graph; and determining the statement vector corresponding to the distributed query statement based on the feature vector of each node in the updated connection graph.
[0140] In a line graph, the vertices are formed by the edges of the connection graph, and two vertices are adjacent in the line graph if and only if they share a node in the connection graph. That is, the edges of the line graph represent the adjacency relationships of edges in the connection graph. Nodes in the connection graph represent the second vectors of the second component; edges in the connection graph represent conditional connections between conditional statements and object properties, association connections between entity types, and primary / foreign key connections between object properties. A conditional connection can refer to a conditional statement where the value of an object property determines which code to execute. Association connections are used to represent interactions or relationships between different entity types. For example, association connections include one-to-one, one-to-many, or many-to-many relationships, such as a student being able to take multiple courses, a one-to-many relationship between a "student" entity and a "course" entity. Primary / foreign key connections are a specific way to implement association connections between entities. A primary key is a unique identifier for each record in a table, used to ensure that each row in the table can be uniquely identified. A foreign key is a field in one table that is also the primary key of another table; foreign keys are used to establish a connection between two tables.
[0141] In this embodiment, a connection graph can be constructed by parsing the second component in the distributed query statement. Nodes in the connection graph represent the second vectors of entity type, object attributes, and conditional statements, and edges represent the connection relationships between the second components. A line graph of the connection graph is derived based on the nodes in the connection graph G. For example, the connection graph G and the line graph LG can be found in [reference needed]. Figure 3Furthermore, a Dual Relational Graph Attention Network (DR-GAT) can be used to capture and distinguish the impact of different connection relationships on node state updates. For each node, attention weights are calculated based on the embedding representations of its neighboring nodes (considering different connection relationships). These attention weights are then used to aggregate information from neighboring nodes, update the current node's state, and update the vector representations of nodes in the connection graph and line graph. Finally, after L-fold graph computation, the updated connection graph is obtained. The feature vectors of all nodes in the connection graph are concatenated to obtain a concatenated vector. Then, the transpose of the concatenated vector is linearly transformed to obtain the attention value. Based on the attention value and the concatenated vector, the statement vector corresponding to the distributed query statement is determined.
[0142] For example, the statement vector R(q) corresponding to the distributed query statement can be determined using formula (5). Formula (5) can be expressed as:
[0143]
[0144]
[0145]
[0146] Where R(v1) to R(vn) represent the feature vectors connecting all nodes v1 to vn in the graph. The symbol for concatenation is R, and R represents the concatenated vector. T This represents the transpose of a cascaded vector. This represents the attention value. `softmax` represents the activation function.
[0147] S350. Based on the action state vector, determine the query execution attributes when the current tree to be connected is connected to the candidate tree.
[0148] It should be noted that the cost prediction model can be trained based on training samples, where the training samples can be the tree vectors of the current tree to be joined, the tree vectors of the candidate trees to be joined, and the query execution time from the current tree to the candidate trees obtained by the executor. For example, see [link to example]. Figure 6This method can convert SQL statements into GraphQL statements (i.e., distributed query statements), collect historical data including data pairs of query statements and query execution times, and then extract feature vectors of each component in the GraphQL query statement. The model is trained using plan enumeration. After training, the tree vectors of the current tree to be connected, the tree vectors of the candidate trees, and a second vector are input into the cost prediction model to obtain the query execution attributes when the current tree to be connected is connected to the candidate trees, thus achieving query path optimization.
[0149] To train a cost prediction model and update its parameters, L2 loss can be used to define and minimize the loss function. The loss function can be expressed as Loss = (y j -Q(S j A j ,θ)) 2 ,in, The online value network Q and the objective value network Q are respectively equipped with model parameters θ and θ'. - The network uses the DQN algorithm to train the cost prediction model, and the batch data of state transitions sampled from the delay experience pool is (S). j A j ,r j ,S j+1 ), where S, A, and r represent the state (i.e., the state vector when the current tree to be connected is connected to the candidate tree), the action (i.e., the candidate tree), and the accumulated Q-value, respectively. During training, every n steps... fre At each time step, the target value model is updated once, that is, the parameters θ of the online value network are copied to the target value network, i.e., θ - =θ, thus obtaining the cost prediction model. After training, the model fine-tuning phase begins, where the executor acts as a new environment, updating the empirical dataset and cost prediction model parameters in real time.
[0150] S360. Based on each query execution attribute, determine the target join tree from at least one candidate join tree.
[0151] S370. Based on the current tree to be connected and the target tree, determine the target query path.
[0152] S380. Execute a distributed query statement based on the target query path to obtain the target data.
[0153] The technical solution of this embodiment obtains the connection vector when the current tree to be connected is connected to the candidate tree by inputting the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector into a pre-trained cost prediction model. Then, based on the second vector and the connection vector, the action state vector when the current tree to be connected is connected to the candidate tree to be connected is determined. Finally, the query execution attributes when the current tree to be connected is connected to the candidate tree to be connected are determined by the action state vector. This achieves the goal of improving the effectiveness of path planning by mining the relationships and features between various elements such as query statements, states, connection trees, different connection trees, entity types, and function APIs, so that the action state vector can represent various key information such as query path, API, attributes, and entity type.
[0154] Example 4
[0155] Figure 7 This is a flowchart of a distributed query service optimization method according to Embodiment 4 of the present invention, which further refines "S150" based on the foregoing embodiments. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0156] like Figure 7 As shown, the method specifically includes the following steps:
[0157] S410. Obtain the distributed query statement and determine the first vector of at least two first components and the second vector of the second component in the distributed query statement.
[0158] S420. Based on at least two first components, determine a tree set; wherein the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent the first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected.
[0159] S430. For each tree to be connected, determine at least one candidate tree corresponding to the current tree to be connected from the tree set, and determine the query execution attributes when the current tree to be connected is connected to the candidate tree based on the tree vector of the current tree to be connected, the tree vector of the candidate tree, and the second vector.
[0160] S440. Based on each query execution attribute, determine the target join tree from at least one candidate join tree.
[0161] S450. When the current tree to be connected is a single-node tree, remove the current tree to be connected from the tree set to update the tree set.
[0162] It should be noted that when the current tree to be connected is a single-node tree, it is removed from the tree set. If the current tree to be connected is a multi-node tree, step S450 can be skipped and step S460 executed. In other words, it is only necessary to remove the current tree to be connected from the tree set at the very beginning when selecting the target tree for the current tree to be connected in the tree set. A multi-node tree can refer to a tree structure with multiple nodes. For example, a multi-node tree may include one leaf node and one root node, or two leaf nodes and one root node, or three leaf nodes, one intermediate node, and one root node.
[0163] S460. Determine whether the updated tree set is empty. If not, proceed to step S470. If yes, proceed to step S480.
[0164] In this embodiment, the completeness of the query plan, i.e., whether the target query path meets the requirements, can be determined by whether there are still unconnected candidate connection trees in the tree set.
[0165] In other words, when determining the target query path, each candidate join tree can be checked to see if it forms a complete tree structure. If it does, the connected trees when the current tree to be joined is connected to the target tree are extracted as a complete query plan and used as the final query path. If it does not form a complete tree structure, the candidate join trees are re-determined, and tree joining continues.
[0166] S470. Determine the connected trees when the current tree to be connected is connected to the target tree, and use the connected trees as the new current tree to be connected. Based on the current tree to be connected and the updated tree set, re-execute the operations of determining at least one candidate tree corresponding to the current tree to be connected, determining the query execution attributes when the current tree to be connected is connected to the candidate tree, determining the target tree, and determining the target query path.
[0167] In this embodiment, the current tree to be connected and the target tree to be connected can be connected to obtain the connected tree structure, which is the connected tree. At this time, the connected tree is a multi-node tree. This connected tree can be used as the new current tree to be connected, replacing the original current tree to be connected. Furthermore, based on the new current tree to be connected and the updated tree set, the operations of determining at least one candidate tree corresponding to the current tree to be connected and determining the query execution attributes when connecting the current tree to the candidate tree in step S130, as well as determining the target tree in step S140 and determining the target query path in step S150, can be re-executed.
[0168] However, it should be noted that the connected tree at this time may be based on the connection of a multi-node tree and a single-node tree. In this case, the tree vector of the connected tree can be determined based on the connection vector when the current tree to be connected is connected to the target tree.
[0169] In this embodiment, the connection vector when connecting the current tree to the target tree can be determined as follows: when the current tree to be connected is a multi-node tree, based on the query unit, filter row unit, and mapping column unit in the cost prediction model, and based on the second vector and the tree vector of the target tree, a seventh result corresponding to the target tree is determined; based on the connection unit in the cost prediction model, and based on the second vector, the seventh result, and the tree vector of the current tree to be connected, a connection vector when connecting the current tree to the target tree is determined, so as to determine the tree vector of the connected tree based on the connection vector.
[0170] This can be understood as follows: When the current tree to be connected is a multi-node tree, the tree vector of the target tree can be input into the cost prediction model. Combined with the second vector, it is processed sequentially through the query unit, row filtering unit, and column mapping unit in the cost prediction model to obtain the seventh result corresponding to the target tree. The method for determining the seventh result is similar to the method for determining the fifth result of the current tree to be connected, and will not be elaborated here. Based on the connection unit in the cost prediction model, combined with the second vector, the seventh result, and the tree vector of the current tree to be connected, the connection vector when the current tree to be connected is determined to connect to the target tree is determined, and this connection vector is used as the tree vector of the connected tree.
[0171] S480. Perform a connection operation on the current tree to be connected and the target tree to be connected to obtain the connected tree, and determine the target query path based on the connected tree.
[0172] Specifically, the tree structure connecting the current tree to the target tree can be considered as the connected tree, and the connected tree is the target query path.
[0173] S490. Execute a distributed query statement based on the target query path to obtain the target data.
[0174] The technical solution provided in this embodiment removes the current tree to be connected from the tree set when it is a single-node tree to update the tree set. When the updated tree set is not empty, it determines the connected trees when the current tree to be connected is connected to the target tree and uses the connected trees as the new current tree to be connected. Based on the current tree to be connected and the updated tree set, it re-executes the operations of determining at least one candidate tree corresponding to the current tree to be connected, determining the query execution attributes when the current tree to be connected is connected to the candidate tree, determining the target tree, and determining the target query path. This process continues until the updated tree set is empty. Then, it performs a connection operation on the current tree to be connected and the target tree to obtain a connected tree. Based on the connected tree, it determines the target query path, thereby achieving optimal query execution attributes of the target query path while ensuring the correctness and effectiveness of data queries based on the target query path.
[0175] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0176] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An optimization method for distributed query services, characterized in that, include: Obtaining a distributed query statement and determining the first vector of at least two first components and the second vector of a second component in the distributed query statement includes: for each first component, determining the third vector of the interface function name of the first component; determining the fourth vector of the interface function parameters in the first component; determining the fifth vector of the return value corresponding to the first component; determining the first vector of the first component based on the third vector, the fourth vector, and the fifth vector; when the second component includes an object attribute, determining the second vector of the object attribute based on the first sub-attribute and the first preset dimension of the object attribute; the first sub-attribute includes the attribute type, whether the attribute is empty, whether the attribute is a list type, and the first order of the object attribute in its corresponding entity type. The second component includes at least one of an ID and an attribute name; when the second component includes a conditional statement, a second vector of the conditional statement is determined based on a second sub-attribute and a second preset dimension of the conditional statement; the second sub-attribute includes at least one of the attribute type of the object attribute in the conditional statement and a second sequential ID, operator, and operand of the object attribute in its corresponding entity type; when the second component includes an entity type, a second vector of the entity type is determined based on a third sub-attribute of the entity type; the third sub-attribute includes at least one of the object attribute of the entity type, the entity type name, and a third sequential ID of the entity type; wherein, the first component includes a query interface; the second component includes at least an object attribute, an entity type, and a conditional statement; Based on the at least two first components, a tree set is determined; wherein the tree set includes at least two trees to be connected; the trees to be connected correspond to single-node trees; the nodes of the trees to be connected represent the first components; the tree vectors of the trees to be connected are determined based on the first vectors of the first components corresponding to the nodes of the trees to be connected. For each tree to be connected, at least one candidate tree corresponding to the current tree to be connected is determined from the tree set, and the query execution attributes when the current tree to be connected is determined based on the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector; Based on each of the query execution attributes, a target join tree is determined from the at least one candidate join tree; Based on the current tree to be connected and the target tree, the target query path is determined, and the distributed query statement is executed based on the target query path to obtain the target data.
2. The method according to claim 1, characterized in that, The step of determining the query execution attributes when the current tree to be connected to the candidate tree, based on the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector, includes: The tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector are input into a pre-trained cost prediction model to obtain the connection vector when the current tree to be connected is connected to the candidate tree to be connected. Based on the second vector and the connection vector, determine the action state vector when the current tree to be connected is connected to the candidate tree; Based on the action state vector, the query execution attributes when the current tree to be connected is connected to the candidate tree are determined.
3. The method according to claim 2, characterized in that, The step of inputting the tree vector of the current tree to be connected, the tree vector of the candidate tree to be connected, and the second vector into a pre-trained cost prediction model to obtain the connection vector when the current tree to be connected is connected to the candidate tree includes: Based on the query unit in the cost prediction model, the tree vector of the current tree to be connected and the tree vector of the candidate tree to be connected are processed according to the second vector of the entity type, respectively, to obtain a first result corresponding to the current tree to be connected and a second result corresponding to the candidate tree to be connected; Based on the filter row unit in the cost prediction model, and based on the second vector of the conditional statement, the first result and the second result are processed respectively to obtain a third result corresponding to the current tree to be connected and a fourth result corresponding to the candidate tree; Based on the mapping column unit in the cost prediction model, and using the second vector of the object attribute to compare the third and fourth results respectively, a fifth result corresponding to the current tree to be connected and a sixth result corresponding to the candidate tree to be connected are obtained; wherein, the object attribute includes a connection attribute; Based on the connection unit in the cost prediction model, and based on the second vector of the connection attribute, the fifth result, and the sixth result, the connection vector when the current tree to be connected is connected to the candidate tree is determined.
4. The method according to claim 2, characterized in that, The step of determining the action state vector when the current tree to be connected connects to the candidate tree based on the second vector and the connection vector includes: Based on the second vector, determine the statement vector corresponding to the distributed query statement; Based on the connection vector and the statement vector, determine the action state vector when the current tree to be connected is connected to the candidate tree.
5. The method according to claim 4, characterized in that, The step of determining the statement vector corresponding to the distributed query statement based on the second vector includes: Based on the second vector of each second component, a connection graph is constructed; wherein, the nodes in the connection graph represent the second vectors of the second components; the edges in the connection graph represent the conditional connection relationship between the conditional statement and the object attribute, the association connection relationship between the entity types, and the primary and foreign key connection relationship between the object attributes; Determine the line graph corresponding to the connection diagram; The connection graph and the line graph are processed based on a graph attention network to obtain an updated connection graph; Based on the feature vector of each node in the updated connection graph, the statement vector corresponding to the distributed query statement is determined.
6. The method according to claim 1, characterized in that, The step of determining the target join tree from the at least one candidate join tree based on each of the query execution attributes includes: The candidate join tree corresponding to the smallest query execution attribute is selected as the target join tree, and the target join tree is removed from the tree set.
7. The method according to claim 1, characterized in that, The step of determining the target query path based on the current tree to be connected and the target connection tree includes: When the current tree to be connected is a single-node tree, the current tree to be connected is removed from the tree set to update the tree set; When the updated tree set is not empty, determine the connected trees when the current tree to be connected is connected to the target tree, and use the connected trees as the new current tree to be connected. Based on the current tree to be connected and the updated tree set, re-execute the operations of determining at least one candidate tree corresponding to the current tree to be connected, determining the query execution attributes when the current tree to be connected is connected to the candidate tree, determining the target tree, and determining the target query path. When the updated tree set is empty, a connection operation is performed on the current tree to be connected and the target tree to be connected to obtain a connected tree, and the target query path is determined based on the connected tree.
8. The method according to claim 7, characterized in that, The tree vector of the connected tree is determined based on the connection vector when the current tree to be connected is connected to the target tree; the method further includes: When the current tree to be connected is a multi-node tree, based on the query unit, filter row unit and mapping column unit in the cost prediction model, and based on the second vector and the tree vector of the target tree, a seventh result corresponding to the target tree is determined; Based on the connection unit in the cost prediction model, and based on the second vector, the seventh result, and the tree vector of the current tree to be connected, the connection vector when the current tree to be connected is connected to the target tree is determined, so as to determine the tree vector of the connected tree based on the connection vector.
Citation Information
Patent Citations
A caching method for a query intermediate result set of a distributed database system
CN109947796A
Distributed database table connection sequence and connection operator optimization method and system
CN118394785A