Method, device and equipment for optimizing performance of expression in graph database
By designing expression performance optimization methods in the graph database, including type derivation and metadata synchronization mechanisms, the problem of low expression performance in the graph database is solved, and more efficient computing performance and lower network communication overhead are achieved.
Patent Information
- Application Number
- CN202510651333.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The expression performance in the graph database is low, especially when processing graph function expressions, the computational complexity is high and the existing technology cannot effectively optimize, resulting in performance bottlenecks.
Through the performance optimization method of expressions in the design drawing database, it includes obtaining user query requests, parsing and converting them into typeless expressions, performing type derivation, generating typed expressions, and performing run-time calculations in the executor module. At the same time, the metadata synchronization mechanism and the prefetch mechanism of subgraph topology cache are adopted to optimize the calculation of graph function expressions.
It improves the calculation efficiency of expressions in the graph database, significantly reduces the overhead of remote process calls, optimizes the calculation performance of metadata functions, and improves the performance performance of graph databases in large-scale distributed environments.
Smart Images

Figure CN120179868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of graph database optimization, and in particular, to a method, device, and equipment for optimizing the performance of expressions in a graph database. Background Art
[0002] Expressions in a graph database are mainly used to describe data calculation methods. Common expressions include constant expressions, attribute extraction expressions, scalar function expressions, aggregation expressions, and graph function expressions (such as pattern functions, existential graph functions, and metadata functions). When processing these expressions, especially graph function expressions, due to the high computational complexity, there are performance bottlenecks, especially in a distributed graph database architecture.
[0003] Traditional relational databases have mature solutions for expression implementation, but the unique data model and query language (GQL) of graph databases lead to new problems in expression implementation.
[0004] Specific defects include:
[0005] The query expressions in a graph database need to be adapted according to the topological structure of the graph data, and the existing expression implementation in a graph database cannot effectively handle this adaptation problem, resulting in low performance.
[0006] The design of the expression life cycle in a graph database is relatively complex, especially in the distributed architecture of a graph database, and there is a lack of effective optimization mechanisms. Summary of the Invention
[0007] Object of the Invention: The object of the present invention is to solve the defects in the prior art and provide a method, device, and equipment for optimizing the performance of expressions in a graph database.
[0008] Technical Solution:
[0009] In a first aspect, the present application proposes a method for optimizing the performance of expressions in a graph database, which is used for optimizing expressions in a graph database and includes the steps of:
[0010] Obtain a query request from a user and pass the query request to a query processing module in the graph database;
[0011] The query processing module performs lexical and syntactic analysis on the query request through a parser module and converts the query request input by the user into a typeless expression;
[0012] Pass the typeless expression to a validator module for type inference and convert it into a typed expression, where the goal of the inference is to attach the correct data type to each node of the expression;
[0013] Input a typed expression into the executor module, bind and store the runtime calculation address of the expression to generate a runtime expression, where the expression contains specific execution information;
[0014] Pass the runtime expression into the executor module for execution. According to the query request, the executor module calculates and searches the graph data and returns the query result;
[0015] Among them, in the validator or executor stage, set up a metadata synchronization mechanism, including broadcasting graph metadata from the metadata management module to each cluster node through a remote procedure call heartbeat mechanism. Each node saves a local metadata snapshot for reference in the expression calculation during query;
[0016] Control the broadcast and publication of metadata with the remote procedure call heartbeat mechanism. After the node receives the confirmation signal in the publication stage, the graph metadata snapshot is used for type deduction and calculation of the expression.
[0017] Preferably, the metadata management module is also responsible for the creation, broadcast and publication of graph metadata, and synchronizes metadata updates to distributed cluster nodes through a regular remote procedure call heartbeat mechanism, including: users define entity types, edge types, node labels, and attribute type information of the graph database by executing graph metadata definition language requests;
[0018] After executing the data definition language request, the graph metadata is persisted to the metadata management module, and the persisted graph metadata is broadcast to other nodes in the cluster;
[0019] Each node confirms whether it has received the new metadata according to the heartbeat signal and returns a confirmation signal for confirmation.
[0020] Preferably, for each data partition of the storage node in each cluster node, the storage node receives and saves a graph metadata snapshot;
[0021] For the query processing node in each cluster node, a graph metadata snapshot is retained in the local service process.
[0022] Preferably, pass the runtime expression into the executor module for execution. According to the query request, the executor module calculates and searches the graph data and returns the query result, including:
[0023] The expression execution process is divided into an operator layer and an expression layer;
[0024] Among them, the operator layer is responsible for managing the input table and the output table, and the data table contains multiple batches;
[0025] The expression layer is responsible for receiving the input data from the operator layer, accessing the data in the form of views, performing expression calculations, and outputting the results as data blocks and returning them to the operator layer.
[0026] Preferably, in the expression layer, after the calculation is completed, the output data will be optimized by a writer and summarized in the form of data blocks and written into the output table of the operator layer;
[0027] The management of data tables includes the division of batches;
[0028] Under the management of the operator layer, each data table is divided into multiple batches to process data in chunks.
[0029] Preferably, during the expression calculation process, through view abstraction of data access, the expression layer operates on the data through views;
[0030] When writing output data, the data is optimized through the corresponding writer.
[0031] Preferably, it includes: pattern functions, existential graph functions;
[0032] Optimizing the pattern functions and existential graph functions includes:
[0033] When first pulling graph data from the storage node, by prefetching the corresponding topological subgraph and storing it in the cache of the local query processing node, the degree n of the prefetch topological subgraph can be configured by the user, and during the prefetch process, for each node, by requesting topological data from the storage node, gradually load the relevant topological subgraph nodes and edges into the topological subgraph of the query processing node;
[0034] When performing graph function calculations, preferentially obtain topological subgraph data from the local topological subgraph for iterative calculations;
[0035] When the cached topological subgraph hits, directly perform topological calculations in the local cache; otherwise, send a remote procedure call request to the storage node to pull the required graph data and update the local topological subgraph;
[0036] Preferably, the graph database also includes graph function expressions, which include: metadata graph functions;
[0037] Using the graph metadata synchronization mechanism to enable each node in the cluster to locally save a consistent version of the metadata snapshot;
[0038] During the calculation process of metadata function expressions, untyped expressions perform type inference in the validator module to generate typed expressions with clear types and corresponding catalog versions, including:
[0039] The typed expression after type inference internally contains the return data type and the corresponding catalog version, and is converted into a runtime expression by the executor module during the execution phase, so that the access to the preservation metadata uses the correct version during calculation. Among them, the runtime expression can generate a corresponding metadata accessor, specify the required entity type and attribute name through the property accessor, and retrieve and return the metadata result requested by the user using the local catalog snapshot.
[0040] In a second aspect, an embodiment of the present invention provides an expression performance optimization device in a graph database, including the method according to any one of the above embodiments, including:
[0041] An acquisition transfer unit for acquiring a query request of a user and transferring the query request to a query processing module in the graph database;
[0042] An analysis conversion unit for the query processing module to perform lexical and syntactic analysis on the query request through a parser module and convert the query request input by the user into an untyped expression;
[0043] A derivation conversion unit for transferring the untyped expression to a validator module for type inference and converting it into a typed expression, where the goal of the derivation is to attach the correct data type to each node of the expression;
[0044] An execution unit for inputting the typed expression into an executor module, binding and storing the runtime calculation address of the expression to generate a runtime expression, where the expression contains specific execution information; inputting the runtime expression into the executor module for execution, and according to the query request, the executor module calculates and searches for graph data and returns a query result; among them, in the validator or executor phase, a metadata synchronization mechanism is set, including broadcasting graph metadata from the metadata management module to each cluster node through a remote procedure call heartbeat mechanism, and each node saves a local metadata snapshot for reference in the expression calculation during the query;
[0045] Controlling the broadcast and release of metadata with a remote procedure call heartbeat mechanism, and after the node receives the confirmation signal in the release phase, the graph metadata snapshot is used for type inference and calculation of the expression.
[0046] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory. Among them, the memory is used to store one or more computer programs; when one or more computer programs stored in the memory are executed by the processor, the electronic device can implement the method according to any possible design of the first aspect above.
[0047] Fourthly, the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method according to any one of the above embodiments is implemented.
[0048] Fifthly, another embodiment of the present invention provides a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to execute the method of any possible design in any of the above aspects.
[0049] Beneficial effects:
[0050] Efficient implementation of graph database expressions. By improving the lifecycle management of expressions, a type deduction and runtime calculation mechanism for expressions is designed, which improves the calculation efficiency, parses graph metadata and is used for expression form conversion, stronger typed expressions are more efficient, and the lifecycle and code implementation framework of expressions are clearly defined;
[0051] Optimization of graph function expressions. For pattern functions and existential graph functions, a prefetch mechanism of subgraph topology caching is adopted, which reduces the dependence on the storage layer during each calculation, significantly reduces the remote procedure call overhead, and improves the calculation speed;
[0052] Graph metadata synchronization mechanism. A metadata synchronization scheme based on the heartbeat mechanism of remote procedure calls is designed, which ensures the metadata consistency of each node in the distributed cluster, thereby optimizing the calculation performance of metadata functions;
[0053] Implementation path of performance optimization. Through the special optimization of graph functions and the localization processing of graph metadata, the execution efficiency of complex expressions in the graph database is successfully improved, especially the performance in a large-scale distributed environment;
[0054] Greatly improves the expression calculation ability of the graph database, reduces the network communication overhead, and optimizes the overall performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic flowchart of the method provided by the present invention;
[0056] Figure 2 It is a schematic diagram of the expression lifecycle provided by the present invention;
[0057] Figure 3 It is a schematic diagram of the untyped expression data structure of the present invention;
[0058] Figure 4 It is a schematic diagram of the metadata synchronization mechanism of the present invention;
[0059] Figure 5 It is a design diagram of the expression runtime of the present invention;
[0060] Figure 6 This is the optimized design diagram of the pattern function or existence graph function of the present invention;
[0061] Figure 7 This is the schematic diagram of the implementation of the metadata function expression of the present invention;
[0062] Figure 8 This is a schematic diagram of a device provided by the present invention;
[0063] Figure 9 This is a schematic diagram of a device provided by the present invention. Detailed implementation manners
[0064] To make the technical solutions of the present invention clearer, the following further describes the present invention in detail with specific embodiments in conjunction with the accompanying drawings.
[0065] Embodiment 1
[0066] To make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0067] Aiming at the problems existing in the prior art, such as Figure 1-2 As shown, the present invention provides a method for optimizing the performance of expressions in a graph database, which is used for optimizing expressions in a graph database, and includes the steps of:
[0068] S101. Obtain the query request of the user and pass the query request to the query processing module in the graph database (graphd, query processing is one of the core modules in the Nebula graph database, mainly used to respond to the query request of the user and execute query operations);
[0069] The user sends a query request through the interface of the graph database (such as the GQL query language), and the query request first enters the query processing module of the graph database.
[0070] S102. The query processing module performs lexical and syntactic analysis on the query request through the parser module (Parser), and converts the query request input by the user into a typeless expression;
[0071] After the parser module receives a query request, it first performs lexical and syntactic analysis on it. The goal of the analysis is to transform the user's query request into a form that the graph database can understand;
[0072] After analysis, the query request is transformed into an UntypedExpr, which is an intermediate expression without a specific data type and usually appears as a tree structure.
[0073] S103. Pass the UntypedExpr to the Validator module for type inference and convert it into a TypedExpr. Among them, the goal of the inference is to attach the correct data type to each node of the expression;
[0074] The UntypedExpr is passed to the Validator module for type inference. The role of the validator is to deduce the correct data type for each node of the expression based on the metadata of the graph database (such as node types, edge types, attribute types, etc.);
[0075] After the inference, the correct data type is attached to each node of the expression to generate a TypedExpr.
[0076] S104. Input the TypedExpr into the Executor module to bind and store the runtime calculation address of the expression to generate a RuntimeExpr (which represents the actual operation of the expression in the graph database when it is executed in the computing engine. It carries all the information required during the execution process and ensures that the expression can be correctly calculated and optimized). Among them, the expression contains specific execution information;
[0077] The TypedExpr enters the Executor module, and the executor is responsible for converting the type information of the expression into a computable form that can be executed at runtime;
[0078] In the executor stage, the calculation address of the expression is bound and stored, and finally a RuntimeExpr is generated. The RuntimeExpr contains specific execution information, including functions, data types, etc. required for the calculation;
[0079] The RuntimeExpr is passed into the Executor module for execution. According to the query request, the Executor module returns the query result by calculating and searching the graph data;
[0080] The RuntimeExpr is passed to the Executor module for continued execution. The executor will return the query result by calculating and searching the graph data according to the query request. The query result will be output through the calculation of the Executor module and finally returned to the user.
[0081] Among them, in the validator or executor stage, a metadata synchronization mechanism is set up, including broadcasting graph metadata from the metadata management module (metad module, which is a key component in the Nebula graph database responsible for the persistence and management of metadata) to each cluster node through a remote procedure call heartbeat mechanism (RPC, Remote Procedure Call Heartbeat Mechanism, a mechanism for ensuring metadata consistency among nodes in a distributed graph database system). Each node saves a local metadata snapshot for reference in expression calculation during queries;
[0082] The remote procedure call heartbeat mechanism is used to control the broadcasting and publishing of metadata. After a node receives the confirmation signal in the publishing stage, the graph metadata snapshot is used for type deduction and calculation of expressions.
[0083] In the validator or executor stage, the system uses a metadata synchronization mechanism to ensure metadata consistency among different cluster nodes. This mechanism is implemented through a remote procedure call heartbeat mechanism.
[0084] Remote procedure call heartbeat mechanism: The metadata management module broadcasts graph metadata to each cluster node. Each node receives the heartbeat message and updates its local metadata snapshot (Catalog Snapshot) for use in queries;
[0085] Only after a node receives the confirmation signal (ACK) in the metadata publishing stage will it use the updated graph metadata for type deduction and calculation execution of expressions. This process ensures that in a distributed environment, all nodes can access consistent metadata, thus guaranteeing the correctness of query calculations.
[0086] In some specific embodiments, in combination with Figure 3 ;
[0087] Top layer: count expression
[0088] The outermost layer is a count expression, which is usually used in queries to calculate the number of elements that meet the conditions.
[0089] Nested: case expression
[0090] Below the count expression, a case expression is nested. This is a conditional expression that returns different values based on the truth or falsehood of the condition.
[0091] In this figure, the conditional judgment is to check whether n.score (the score of entity n) is greater than 60.
[0092] Condition (case expression)
[0093] The conditional part judges whether n.score is greater than 60.
[0094] If the condition is true (i.e., the score is greater than 60), return 1, indicating that the condition is met (possibly counting the items that meet the condition).
[0095] If the condition is false (the score is not greater than 60), return null, indicating that the condition is not met;
[0096] The attribute expression represents accessing the score attribute of entity n; this attribute expression is used in the conditional part of the case expression to compare whether n.score is greater than 60.
[0097] It is analyzed that it has five leaf nodes, and each leaf node represents an element in the expression. The following analyzes their meanings and type derivation processes in detail:
[0098] 1. Leaf node: null type: Empty type (null). In SQL or graph database queries, null usually represents a missing value or a null value, and it has no data type itself.
[0099] 2. Leaf node: 1 type: Integer type (Integer). 1 is a constant, representing the value returned when n.score > 60. Its type is integer type, representing a numerical value;
[0100] 3. Leaf node: 60 type: Integer type (Integer). 60 is also a constant, representing a comparison value in the case statement, used to compare with n.score, and its type is also integer type.
[0101] 4. Leaf node: n, type: Node type (Node Type). n represents an entity in the graph database, usually a node. In this example, n represents a student (Student) node. The type of n is Node type, that is, it is an entity node in the graph database, containing certain attributes and labels.
[0102] 5. Leaf node: score, type: String type (String). score is an attribute, usually representing a specific attribute of a certain node. Here, the score attribute is used to store the scores of students. Its type is String type, representing the attribute name.
[0103] The type inference process of the above example is as follows:
[0104] Catalog Snapshot: Catalog Snapshot contains the definitions, relationships, and attributes of all entity types in the graph database. For example, the node types Teacher and Student in the graph database, as well as the edge types between them (such as Teach), are defined in the Catalog Snapshot.
[0105] Graph Pattern: In the query, n is generated by a graph pattern (such as (v:Teacher)-[:Teach]->(n:Student)). Here, the graph pattern describes the relationships and nodes in the graph, and n represents the Student node.
[0106] Accessing Catalog Snapshot: Through the labels in the graph pattern (such as Student), the query accesses the Catalog Snapshot to infer all the attribute types of the node type Student. All the attributes and their types of this node will be listed in the Catalog Snapshot. For example, the type of the score attribute may be Integer or Float, etc.
[0107] Deriving Attribute Types: In this example, n is generated by the graph pattern:
[0108] (v:Teacher)-[:Teach]->(n:Student). By accessing the Catalog Snapshot, the query can derive the data type of the attribute score of n, and thus determine the type of the score attribute used in case when n.score > 60.
[0109] In some specific embodiments, in combination with Figure 4 , the metadata management module is also responsible for the creation (CREATE), broadcast (BROADCAST), and publication (PUBLIC) of the Catalog Snapshot, and synchronizes metadata updates to the distributed cluster nodes through a regular remote procedure call heartbeat mechanism, including: users define entity types, edge types, node labels, and attribute type information of the graph database by executing Catalog Snapshot definition language requests;
[0110] After executing the Data Definition Language request (DLL request), the Catalog Snapshot (Part1, Part2) is persisted to the metadata management module, and the persisted Catalog Snapshot is broadcast to other nodes in the cluster;
[0111] Each node confirms whether it has received the new metadata according to the heartbeat signal and returns a confirmation signal for confirmation.
[0112] In some specific embodiments, for each data partition of the storage nodes in each cluster node (storaged, the storage nodes in the cluster, responsible for storing graph data and metadata, and providing data to the query processing module), the storage node will receive and save a copy of the graph metadata snapshot (Catalog Snapshot);
[0113] For the query processing nodes in each cluster node (the graph query module, the query module of the graph database, responsible for processing user query requests (Query Request) and performing graph queries according to the query conditions. It will parse and optimize the query based on the metadata), a copy of the graph metadata snapshot is retained in the local service process.
[0114] Specifically, the user defines or modifies information such as entity types (such as node types, edge types), node labels, and attribute types in the graph database by executing DDL (Data Definition Language) requests. These requests are first processed by the metadata management module, which is responsible for the creation and persistent storage of graph metadata, and broadcasts this graph metadata to other nodes in the cluster. After executing the data definition language request, the graph metadata will be persisted to Metad and pushed to other nodes through the broadcast mechanism. Each node in the cluster (such as storage nodes and query processing) will receive the broadcast graph metadata and confirm whether it has received the updated metadata. The node will return an acknowledgement signal (ACK) message to the metadata management module, indicating that they have received and processed the updated metadata. Once the graph metadata synchronization is completed, each node can use the updated metadata to process query requests, such as Expression Pushdown (an optimization technique that pushes part of the computational tasks in the query to the storage layer for execution, thereby reducing data transmission and improving query efficiency) and Expression Evaluation (performing the computational operations of the query expression, processing and calculating the query based on the graph metadata, and returning the final query result), to execute specific query tasks; the heartbeat synchronization mechanism is used to ensure the consistency of metadata among distributed nodes, and metadata localization is used to optimize expressions to achieve performance.
[0115] In some specific embodiments, in combination with Figure 5 , the runtime expression is passed into the executor module for execution. According to the query request, the executor module calculates and searches for graph data and returns the query result, including:
[0116] The expression execution process is divided into an operator layer and an expression layer;
[0117] Among them, the operator layer is responsible for managing the input table and the output table, and the data table contains multiple batches;
[0118] The expression layer is responsible for receiving the input data from the operator layer, accessing the data in the form of views, performing expression calculations, and outputting the results as data blocks and returning them to the operator layer.
[0119] In some specific embodiments, in the expression layer, after the calculation is completed, the output data will be optimized by a writer and summarized in the form of data blocks and written into the output table of the operator layer;
[0120] The management of data tables includes the division of batches;
[0121] Under the management of the operator layer, each data table is divided into multiple batches to process data in chunks.
[0122] In some specific embodiments, during the expression calculation process, through view abstraction of data access, the expression layer operates on the data through views;
[0123] When writing output data, the corresponding writer is used to optimize the data processing.
[0124] Specifically, the operator layer:
[0125] This layer is responsible for managing and processing tables. Each table contains multiple batches. During the query execution process, the input data is organized into multiple batches, called Input Table. The output results are also stored in the Outputtable. This layer is mainly responsible for data transmission and organization.
[0126] The expression layer:
[0127] This layer is responsible for performing expression calculations. It receives data from the operator layer and accesses the data through views (such as NodeView).
[0128] During the expression calculation process, the execution function (such as Expr::BatchEval) will process the data of each batch, perform necessary calculations, generate results, and the calculation results are returned in the form of data blocks (Data Chunks) and written into the output table (Output Table) of the operator layer;
[0129] The writing process uses the corresponding type of writer (such as a node writer) to optimize data writing to ensure efficient writing of calculation results.
[0130] Data is input into the expression for calculation in batches. Batch processing can effectively reduce the number of data processing times and resource consumption, improve the efficiency of query processing. By passing data to the expression calculation module in batches, it is possible to better manage memory and computing resources, avoiding performance bottlenecks when processing large amounts of data at one time.
[0131] During the calculation process, data is accessed through a view and output is optimized through a writing mechanism (writer). This mechanism ensures efficient data reading and writing. Especially in large-scale data processing, it can avoid frequent disk I / O operations and improve the efficiency of data transmission and storage. Each data type has a corresponding view and writing optimizer (such as a node writer), making data calculation and writing more efficient.
[0132] In some specific embodiments, in combination with Figure 6 , which includes: pattern function, exists function;
[0133] Optimize the pattern function and exists function, including:
[0134] When pulling graph data from the storage node for the first time, prefetch the corresponding topological subgraph and save it in the cache of the local query processing node. The degree n of the prefetch topological subgraph can be configured by the user. And during the prefetch process, for each node, by requesting topological data from the storage node, gradually load the relevant topological subgraph nodes and edges into the topological subgraph of the query processing node;
[0135] When performing graph function calculation, preferentially obtain topological subgraph data from the local topological subgraph for iterative calculation;
[0136] When the cached topological subgraph hits, directly perform topological calculation in the local cache; otherwise, send a remote procedure call request to the storage node to pull the required graph data and update the local topological subgraph;
[0137] Specifically, the relationship between the storage layer and the graph query layer: Figure 6 On the left side in the figure is the storage node layer, which stores data of nodes and edges. The stored data is stored on the disk through disk storage. On the right side in the figure is the graph query node layer, which contains a subgraph cache for storing graph data prefetched from the storage layer. The red dots are the graph data designed for the previous query. The yellow dots and dashed lines in the figure are prefetched from the storaged to the subgraph cache of the query processing through remote procedure calls. The black dots and corresponding edges in the figure will not be prefetched.
[0138] When performing a graph query, especially for functions that require graph topology calculations such as pattern functions or existential graph functions, the graph query layer requests node and edge data from the storage layer through remote procedure calls and prefetches the n-degree topological subgraphs into the local subgraph cache.
[0139] When executing a query, the graph query layer first looks for the required topological data in the local cache. If a hit occurs in the cache, the calculation is directly performed in memory without having to send a remote procedure call request to the storage layer; by prefetching the commonly used topological subgraphs into the cache, it is possible to avoid pulling data from the storage layer every time a query is made, thus reducing the remote procedure call communication overhead with the storage layer and optimizing the query performance. If the cache is hit, the execution speed of the pattern function expression will be greatly improved because the calculation can be directly performed in memory, solving the problem in the prior art of not knowing how to optimize the performance of the pattern function expression: the pattern function is a specific type of graph query expression in a graph database and usually results in poor performance due to frequent topological calculations. By using the subgraph cache mechanism (Subgraph Cache) to prefetch topological subgraphs and reduce communication with the storage layer, the execution efficiency of the pattern function is significantly improved.
[0140] In some specific embodiments, in combination with Figure 7 , the graph database further includes graph function expressions, which include: metadata functions;
[0141] Utilize the graph metadata synchronization mechanism to enable each node in the cluster to locally store a consistent version of the metadata snapshot;
[0142] During the calculation process of the metadata function expression, the untyped expression performs type inference in the validator module to generate a typed expression with a clear type and the corresponding catalog version (Catalog Version), including:
[0143] The typed expression after type inference will internally contain the return data type and the corresponding catalog version, and is converted into a runtime expression by the executor module during the execution phase to ensure that the access to the metadata uses the correct version. Among them, the runtime expression can generate the corresponding metadata accessor, specify the required entity type and attribute name through the attribute accessor, and retrieve and return the metadata result requested by the user using the local catalog snapshot.
[0144] Specifically, for the metadata function: Untyped expression: The metadata function expression (such as hasProperty(v, "prop")) requests graph metadata (such as the property prop of node v) in the query. The expression in the initial stage is an untyped expression (UntypedExpr), indicating that the type deduction of this expression has not been performed yet;
[0145] Metadata function: Typed expression (TypedExpr):
[0146] In the validator module, the untyped expression undergoes type deduction and is converted into a typed expression (TypedExpr). During the type deduction stage, explicit type information is attached inside the expression, and it includes the version number (catalog version) of the associated graph metadata.
[0147] For example, the return value type BOOL and the entity type Teacher are deduced here.
[0148] Metadata function: Runtime expression (RuntimeExpr):
[0149] In the executor module, the typed expression is converted into a runtime expression (RuntimeExpr) to prepare for actual calculation.
[0150] The runtime expression generates accessors related to graph metadata access,
[0151] (such as PropertyGetter), and obtains the corresponding metadata through the local metadata snapshot.
[0152] The accessor specifies the required entity type (such as Teacher) and property name (such as prop), and retrieves and returns the final result from the local metadata snapshot.
[0153] The metadata function expression is usually used to access and query the metadata of the graph database. Since it needs to frequently access the graph metadata source (such as metad) to obtain information such as node attributes, the performance is often poor. This solution is optimized through the design of the expression life cycle: during the type deduction stage, the type of the expression is inferred in advance, and a suitable metadata accessor is constructed; the query is executed through the local metadata snapshot, reducing frequent remote procedure call communication with metad, thereby reducing latency and resource consumption.
[0154] In some embodiments, the present application proposes an apparatus for optimizing the performance of expressions in a graph database, combined with Figure 8 , including the method described in the above embodiments, including:
[0155] An acquisition and transfer unit 301, configured to acquire a query request of a user and transfer the query request to a query processing module in a graph database;
[0156] An analysis and conversion unit 302, configured to enable the query processing module to perform lexical and syntactic analysis on the query request through a parser module, and convert the query request input by the user into a typeless expression;
[0157] A derivation and conversion unit 303, configured to transfer the typeless expression to a validator module for type derivation and convert it into a typed expression, wherein the goal of the derivation is to attach a correct data type to each node of the expression;
[0158] An execution unit 304, configured to input the typed expression into an executor module, bind and store the runtime calculation address of the expression to generate a runtime expression, wherein the expression contains specific execution information; input the runtime expression into the executor module for execution, and according to the query request, the executor module returns a query result by calculating and searching graph data; wherein, in the validator or executor stage, a metadata synchronization mechanism is set up, including implementing the broadcast of graph metadata from a metadata management module to each cluster node through a remote procedure call heartbeat mechanism, and each node saves a local metadata snapshot for reference in the expression calculation during the query;
[0159] The broadcast and publication of metadata are controlled by a remote procedure call heartbeat mechanism, and after the node receives the confirmation signal in the publication stage, the graph metadata snapshot is used for type derivation and calculation of the expression.
[0160] All relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.
[0161] In some other embodiments of the present invention, embodiments of the present invention disclose an electronic device, as Figure 9 shown, the electronic device may include: one or more processors 401; a memory 402; a display 403; one or more applications (not shown); and one or more computer programs 404, and the above devices may be connected through one or more communication buses 405. Wherein the one or more computer programs 404 are stored in the above memory 402 and are configured to be executed by the one or more processors 401, and the one or more computer programs 404 include instructions, and the above instructions may be used to execute each step as Figures 1 to 7 and the corresponding embodiments.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0163] In each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0164] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0165] The above is only the specific implementation manner of the embodiments of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present invention should be covered by the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for optimizing expression performance in a graph database, used for optimizing expressions in a Nebula graph database, characterized in that: Includes steps: Obtain the user's query request and pass the query request to the query processing module in the graph database; The query processing module performs lexical and grammatical analysis on the query request through the parser module, and converts the query request input by the user into a typeless expression; Pass the untyped expression to the validator module for type inference and convert it into a typed expression. The inference goal is to attach the correct data type to each node of the expression. Inputting a typed expression into an executor module so that a runtime calculation address of the expression is bound and stored to generate a runtime expression, wherein the expression includes specific execution information; The runtime expression is passed to the executor module for execution. According to the query request, the executor module returns the query result by calculating and searching the graph data. In the validator or executor stage, a metadata synchronization mechanism is set up, including broadcasting graph metadata from the metadata management module to each cluster node through a remote procedure call heartbeat mechanism. Each node saves a local metadata snapshot for reference in expression calculation in the query. The remote procedure call heartbeat mechanism is used to control the broadcast and publication of metadata. After the node receives the confirmation signal of the publishing phase, the metadata snapshot is used for type deduction and calculation of expressions.
2. The method according to claim 1, characterized in that The metadata management module is also responsible for the creation, broadcasting and publishing of graph metadata, and synchronizes metadata updates to distributed cluster nodes through a regular remote procedure call heartbeat mechanism, including: users define the entity type, edge type, node label, and attribute type information of the graph database by executing graph metadata definition language requests; After executing the data definition language request, the graph metadata is persisted to the metadata management module, and the persisted graph metadata is broadcast to other nodes in the cluster; Each node confirms whether it has received new metadata based on the heartbeat signal and returns a confirmation signal.
3. The method according to claim 1, characterized in that For each data partition of the storage node in each cluster node, the storage node will receive and save a snapshot of the metadata; For each query processing node in the cluster, a snapshot of the image metadata is retained in the local service process.
4. The method according to claim 1, characterized in that: The runtime expression is passed to the executor module for execution. According to the query request, the executor module returns the query results by calculating and searching the graph data, including: The expression execution process is divided into operator layer and expression layer; Among them, the operator layer is responsible for managing the input table and the output table, and the data table contains multiple batches; The expression layer is responsible for receiving input data from the operator layer, accessing data through views, performing expression calculations, and outputting the results as data blocks back to the operator layer.
5. The method according to claim 4, characterized in that At the expression layer, after the calculation is completed, the output data will be optimized for writing through the writer and summarized into data blocks and written to the output table of the operator layer; The management of data tables includes the division of batches; Each data table is divided into multiple batches under the management of the operator layer to process data in blocks.
6. The method according to claim 5, characterized in that During expression calculation, data access is abstracted through views, and the expression layer operates on data through views; When writing output data, the data is optimized by the corresponding writer.
7. The method according to claim 5, characterized in that The graph database also includes graph function expressions, including: pattern functions, existence graph functions; Optimize the pattern function and existence graph function, including: When the graph data is pulled from the storage node for the first time, the corresponding topological subgraph is pre-fetched and saved in the cache of the local query processing node; The degree n of the prefetched topology subgraph can be configured by the user; When performing graph function calculations, topology subgraph data is obtained from the local topology subgraph for iterative calculations first; When the cached topology subgraph hits, the topology calculation is performed directly in the local cache; otherwise, a remote procedure call request is sent to the storage node to pull the required graph data and update the local topology subgraph.
8. The method according to claim 1, characterized in that The graph database also includes graph function expressions, including: metadata graph functions; Use the graph metadata synchronization mechanism to enable each node in the cluster to locally save a consistent version of the metadata snapshot; During the metadata function expression calculation process, the untyped expression is typed inferred in the validator module to generate a typed expression with a clear type and corresponding directory version, including: After type deduction, the typed expression will contain the return data type and the corresponding directory version, and will be converted into a runtime expression by the executor module during the execution phase, so that the access to the preservation data uses the correct version during calculation; Among them, the runtime expression can generate the corresponding metadata accessor, specify the required entity type and attribute name through the attribute accessor, and use the local directory snapshot to retrieve and return the metadata results requested by the user.
9. A device for optimizing expression performance in a graph database, comprising the method according to any one of claims 1 to 8, characterized in that: include: An acquisition and transmission unit is used to obtain a user's query request and transmit the query request to a query processing module in the graph database; The analysis and conversion unit is used for the query processing module to perform lexical and grammatical analysis on the query request through the parser module, and convert the query request input by the user into a typeless expression; A derivation conversion unit is used to pass the untyped expression to the validator module for type derivation to convert it into a typed expression, wherein the derivation goal is to attach a correct data type to each node of the expression; An execution unit is used to input a typed expression into an executor module, so that the runtime calculation address of the expression is bound and stored to generate a runtime expression, wherein the expression contains specific execution information; the runtime expression is passed into the executor module for execution, and according to the query request, the executor module returns the query result by calculating and searching the graph data; wherein, in the validator or executor stage, a metadata synchronization mechanism is set, including broadcasting the graph metadata from the metadata management module to each cluster node through a remote procedure call heartbeat mechanism, and each node saves a local metadata snapshot for reference in the expression calculation in the query; The remote procedure call heartbeat mechanism is used to control the broadcast and publication of metadata. After the node receives the confirmation signal of the publishing phase, the metadata snapshot is used for type deduction and calculation of expressions.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the computer program is executed by the processor, the processor implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Interactive large data analysis query processing method
CN105279286A
Graph data query method and system, computer equipment, and readable storage medium
CN113961730A
Unified SQL query method oriented to heterogeneous data sources
CN117093599A
Hybrid distributed graph data storage and calculation method
CN117112692A
Method and apparatus for channel encoding and decoding in communication or broadcasting system
US20180323807A1