Method and apparatus for optimizing query of graph database, and graph database system including same
By integrating a modular relational query optimizer with Neo4j, the method addresses the limitations of Neo4j's query optimizer, improving performance and scalability through accurate cardinality estimation and advanced optimization techniques.
Patent Information
- Application Number
- PCT/KR2024/015578
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-10-15
- Publication Date
- 2025-07-03
AI Technical Summary
Neo4j's query optimizer does not fully utilize relational database optimization strategies, leading to poor query execution performance and inefficiency in processing complex queries due to limited cardinality estimation and lack of modularity, which hinders the flexibility and scalability of graph database systems.
Integrate a modularized relational database query optimizer to convert graph queries into relational logical plans, apply advanced optimization techniques like join pushdown, and modify the cost model for graph queries, using an interface to map operators and optimize the query plan effectively.
Generates efficient and effective query plans with accurate cardinality estimation, enhancing the performance and scalability of graph database systems by leveraging advanced relational optimization strategies.
Smart Images

Figure KR2024015578_03072025_PF_FP_ABST
Abstract
Description
Method and device for optimizing queries in a graph database, and a graph database system including the same
[0001] The present invention relates to a method and device for optimizing a query of a graph database, and a graph database system including the same, and more particularly, to a Cypher query optimizer using a transformation-based relational query optimizer.
[0002] Neo4j, one of the most widely used graph databases today, features proprietary query optimization technology. Neo4j's query optimization is tailored to the characteristics of graph data and includes various strategies for efficient data access and processing. Neo4j's query optimization includes the use of indexes, understanding and analyzing execution plans, and leveraging node labels.
[0003] Query optimization technology has been developed over a long period of time in the field of relational databases. Even graph databases share many operators from existing relational databases, such as joins and aggregations. Therefore, it is crucial to leverage the diverse query optimization strategies developed in relational databases. However, Neo4j does not fully utilize the diverse query optimization strategies of relational databases. For example, Neo4j does not comprehensively support rule-based optimization. This involves optimizing queries using predefined rules, such as eliminating unnecessary operations and combining consecutive UNIONs. While these optimization strategies have been widely adopted by various relational database optimizers, Neo4j does not fully utilize them. Consequently, Neo4j can lead to poor query execution performance and inefficiencies in complex query processing.
[0004] Meanwhile, cardinality refers to the number of records obtained as a query result, and accurately estimating the cardinality after the operation of each operator within the query plan plays a crucial role in generating an effective query plan. However, Neo4j has limited statistical support and a relatively simple cardinality estimation algorithm. While Neo4j maintains basic statistics such as the number of nodes and relationships, the number of relationships between labels, and index distribution, these are insufficient to accurately estimate the cardinality of complex queries that include various constraints. Consequently, this makes it difficult to generate appropriate query plans, resulting in degraded query execution performance and inefficiency in processing complex queries.
[0005] Neo4j's query optimizer is not modular, making it difficult to integrate or update new optimization strategies or algorithms. This makes it difficult to keep pace with the rapid advancements in query optimization technology and limits the flexibility and scalability of graph database systems. Ultimately, these limitations could negatively impact the future performance of graph database systems.
[0006] The present invention aims to provide a method and device for generating an efficient and effective query plan for a graph query based on relational database query optimization technology.
[0007] The present invention provides a device for optimizing graph queries by integrating various optimization techniques such as join pushdown using a modularized relational database query optimizer, additionally introducing operators and optimization rules necessary for graph queries, and modifying a cost model appropriate for a graph query engine.
[0008] According to one aspect of the present invention, a device for optimizing a graph database query includes an interface for converting a graph query into a relational logical plan; and a relational query optimizer for optimizing the converted relational logical plan; wherein the interface can convert an optimized relational physical plan generated by the relational query optimizer into a physical plan executable in a graph database, and provide the converted physical plan to a graph query processor.
[0009] At this time, the interface stores mapping data in which operators for a relational database and operators for a graph database are mapped, and the interface can convert the optimized relational physical plan into a physical plan executable in the graph database based on the mapping data.
[0010] At this time, the interface can provide graph information for optimizing the relational logic plan to the relational query optimizer.
[0011] At this time, when converting the graph query into the relational logic plan, the interface may convert the graph query into an abstract syntax tree, and convert a portion of the abstract syntax tree representing a subgraph matching of the graph query into a join tree.
[0012] At this time, the relational query optimizer includes an operator and an optimization rule for a graph database for optimizing the relational logic plan, and the interface can convert a portion of the abstract syntax tree representing a path query of the graph query using the operator and an optimization rule for the graph database.
[0013] At this time, the relational query optimizer calculates the query processing cost for each plan using a pre-prepared cost model, and generates the plan with the lowest cost among the calculated costs as the relational physical plan, and the pre-prepared cost model may be a tuned cost model in which the cost of an operator for a relational database mapped to an operator for a graph database is changed to the cost of an operator for the graph database.
[0014] At this time, the relational query optimizer further includes an operator for a second graph database for graph queries in addition to an operator for a relational database that is mapped to an operator for the graph database, and the cost of the operator for the second graph database may be reflected in the pre-prepared cost model.
[0015] A method for optimizing a query of a graph database, according to one aspect of the present invention, may include: converting a graph query into a relational logical plan; optimizing the converted relational logical plan; converting the optimized relational physical plan into a physical plan executable in a graph database; and providing the converted physical plan to a graph query processor.
[0016] At this time, the step of converting the optimized relational physical plan into a physical plan executable in the graph database may be a step of converting the optimized relational physical plan into a physical plan executable in the graph database based on mapping data in which an operator for a relational database and an operator for the graph database are mapped.
[0017] At this time, the optimizing step may include a step of receiving graph information for optimizing the relational logic plan from the graph database.
[0018] At this time, the step of converting the graph query into the relational logic plan may include a step of the interface converting the graph query into an abstract syntax tree; and a step of the interface converting a portion of the abstract syntax tree representing a subgraph matching of the graph query into a join tree.
[0019] At this time, the step of converting the graph query into the relational logic plan may be a step of converting a portion of the abstract syntax tree representing a path query of the graph query using an operator and optimization rule for the graph database.
[0020] At this time, the optimizing step includes a step of calculating a query processing cost for each plan using a pre-prepared cost model; and a step of generating a plan with the lowest cost among the calculated costs as the relational physical plan. The pre-prepared cost model may be a tuned cost model in which the cost of an operator for a relational database mapped to an operator for a graph database is changed to the cost of an operator for the graph database.
[0021] At this time, the cost model prepared in advance may reflect the cost of an operator for a second graph database for graph queries in addition to an operator for the relational database that is mapped to an operator for the graph database.
[0022] According to one aspect of the present invention, a graph database system is provided, including: an interface for converting a graph query into a relational logical plan; a relational query optimizer for optimizing the converted relational logical plan; a graph database; and a graph query processor, wherein the interface converts an optimized relational physical plan generated by the relational query optimizer into a physical plan executable in the graph database and provides the converted physical plan to the graph query processor, and the graph query processor can obtain a result for a graph query in the graph database using the converted physical plan.
[0023] At this time, the interface stores mapping data in which operators for a relational database and operators for the graph database are mapped, and the interface can convert the optimized relational physical plan into a physical plan executable in the graph database based on the mapping data.
[0024] At this time, when converting the graph query into the relational logic plan, the interface may convert the graph query into an abstract syntax tree, and convert a portion of the abstract syntax tree representing a subgraph matching of the graph query into a join tree.
[0025] At this time, the relational query optimizer includes an operator and an optimization rule for a graph database for optimizing the relational logic plan, and the interface can convert a portion of the abstract syntax tree representing a path query of the graph query using the operator and the optimization rule for the graph database.
[0026] At this time, the relational query optimizer calculates the query processing cost for each plan using a pre-prepared cost model, and generates the plan with the lowest cost among the calculated costs as the relational physical plan, and the pre-prepared cost model may be a tuned cost model in which the cost of an operator for a relational database mapped to an operator for a graph database is changed to the cost of an operator for the graph database.
[0027] At this time, the relational query optimizer further includes an operator for a second graph database for graph queries in addition to an operator for a relational database that is mapped to an operator for the graph database, and the cost of the operator for the second graph database may be reflected in the pre-prepared cost model.
[0028] According to the present invention, a method and device for generating an efficient and effective query plan for a graph query based on a relational database query optimization technology can be provided.
[0029] According to the present invention, existing relational advanced query optimization strategies are adopted, and in particular, an excellent query plan can be quickly generated through accurate cardinality estimation.
[0030] According to the present invention, a graph query optimization device can be provided that can integrate and develop various optimization techniques, such as join pushdown, using a modularized relational database query optimizer.
[0031] Figure 1 illustrates a cipher query according to one embodiment of the present invention.
[0032] Figure 2 illustrates a property graph according to one embodiment of the present invention.
[0033] Figure 3 illustrates a homomorphic semantic table according to one embodiment of the present invention.
[0034] Figure 4 illustrates a graph database system according to one embodiment of the present invention.
[0035] Fig. 5 illustrates a graph query optimization device according to one embodiment of the present invention.
[0036] FIG. 6 is a flowchart illustrating a method for optimizing a query of a graph database according to one embodiment of the present invention.
[0037] FIG. 7 is a conceptual diagram illustrating an example of a device, graph database system, or computing system for optimizing queries of a generalized graph database capable of performing at least part of the processes of FIGS. 1 to 6.
[0038] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.
[0039] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0040] In the embodiments of the present application, “at least one of A and B” may mean “at least one of A or B” or “at least one of combinations of one or more of A and B.” Furthermore, in the embodiments of the present application, “at least one of A and B” may mean “at least one of A or B” or “at least one of combinations of one or more of A and B.”
[0041] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0042] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0043] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0044] Hereinafter, with reference to the attached drawings, preferred embodiments of the present invention will be described in more detail. In order to facilitate an overall understanding in describing the present invention, identical reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.
[0045] Figure 1 illustrates a cipher query according to one embodiment of the present invention.
[0046] Figure 2 illustrates a property graph according to one embodiment of the present invention.
[0047] Figure 3 illustrates a homomorphic semantic table according to one embodiment of the present invention.
[0048] Hereinafter, the description will be given with reference to FIGS. 1 to 3.
[0049] The present invention relates to a method for generating a query execution plan for efficiently processing a cipher query when a cipher query for graph data following a property graph model is given.
[0050] For example, in a cipher query, parentheses may represent nodes (e.g., labels: Person, Forum) of a property graph (10), and square brackets may represent edges (i.e., relationships, e.g., labels: MEMBER, KNOWS).
[0051] At this time, the MEMBER edge is directed, but the KNOWS edge does not have a specific direction, so it indicates both directions.
[0052] Cypher queries can be divided into subgraph matching, path queries (Kleene-star), and relational operators.
[0053] For example, the result of a cipher query on the property graph (10) can be represented as in Fig. 3. Looking at the MATCH clause, p1 is acquainted with p2 and p3, and p1 and p2 are members of the Forum, respectively, and p3 is not a member of the Forum. Therefore, looking at the return data corresponding to each field of the table, pl can be {B}, f can be {E}, count(p2) can be "1", and count(p3) can be "1". Meanwhile, even if p1 is {A}, since ABC are connected to each other by a KNOWS edge, it can be seen that A(p1) and C(p3) are also connected by a KNOWS edge. Therefore, the result of the second record of the table in Fig. 3 can also be obtained.
[0054] Figure 4 illustrates a graph database system according to one embodiment of the present invention.
[0055] A graph database (management) system (1000) can receive a graph query (e.g., a cypher query (Q)) from a user and provide the user with a result for the graph query (e.g., a result such as that shown in FIG. 3).
[0056] The graph database system (1000) may include a query engine (1100) and a graph database (graph-native storage) (1200).
[0057] The query engine (1100) may also be referred to as a database query execution engine or a graph query processing engine.
[0058] The query engine (1100) may include a device for optimizing a query of a graph database (hereinafter, a graph query optimization device (query optimizer) (1110)) and a graph query processor (1120) for obtaining a result for a graph query (100) from a graph database (1200) using a physical plan which is a result of graph query optimization.
[0059] A graph database (1200) may include a catalog manager (1210), a cache manager (1220), and a graphlet manager (1230).
[0060] The catalog manager (1210) manages all metadata for the database (e.g., number of total nodes / edges, number of nodes by label, attribute-related information, etc.) and can perform the role of changing or providing metadata upon request.
[0061] The cache manager (1220) may be a module that manages data stored on a disk so that it can be quickly provided in memory.
[0062] The graphlet manager (1230) may be referred to as a graph data manager. The graph data manager (1230) may be a module that processes storage and access requests for nodes / edges on a graph.
[0063] The graph query optimization device (1110) of the graph database system (1000) of the present invention can efficiently process complex graph queries (100) and support faster data analysis compared to conventional technologies.
[0064] Fig. 5 illustrates a graph query optimization device according to one embodiment of the present invention.
[0065] Hereinafter, the description will be made with reference to FIG. 1, FIG. 4, and FIG. 5.
[0066] A graph query optimization device (1110) may include a parser (e.g., a cypher parser (1111)), a transformation-based relational query optimizer (1112), a C2L module (Cypher parse tree to relational logical plan) (1113), an MDP (Metadata Provider) module (1114), and a P2S (relational physical plan to S62 (graph database) physical plan, physical plan transformer) module (1115).
[0067] At this time, the C2L module (1113), the MDP module (1114), and the P2S module (1115) may be components included in the interface. At this time, the 'module' may mean a translator.
[0068] The parser (1111) may have a similar function to a conventional parser.
[0069] The transformation-based relational query optimizer (1112) is a general-purpose query optimizer, which may be a modular optimizer. For example, the relational query optimizer may be a modular query optimizer developed in conjunction with the graph database of the present invention, or may be a commercially available conventional modular query optimizer.
[0070] For example, if the relational query optimizer (1112) is a conventional modular query optimizer that is in use, it may be an optimizer that satisfies the following conditions: 1) it must be an optimizer that has been researched and developed for a long time in the relational database field, 2) it must be highly scalable (e.g., it must be capable of adding new optimization rules, etc.), and 3) it must be capable of cost-based optimization (e.g., it must be capable of tuning by adjusting the cost model for each database system).
[0071] The present invention optimizes graph queries by tuning a relational query optimizer (1112). Furthermore, by applying advanced query optimization techniques used in relational databases, query plans can be generated more efficiently and quickly than in the prior art. Furthermore, modularity enhances the scalability of the technology. The present invention is expected to deliver superior query optimization performance in graph databases.
[0072] Below, the plan generation process of a graph query optimization device (1110) using a relational query optimizer (1112) used in a relational database is described.
[0073] The C2L module (1113) of the above interface can perform graph (i.e., cipher) query transformation.
[0074] The above graph query transformation may mean, for example, transforming a given graph query (100) into a form that can be understood by a relational query optimizer (1112).
[0075] For example, the C2L module (1113) can transform a graph query (100) into an abstract syntax tree (AST) (preprocessing step) to be combined with a relational query optimizer (1112), and then transform (reconstruct) the abstract syntax tree into a logical operator tree of the relational query optimizer (1112). At this time, the conventional ANTLR4 technology can be used for the preprocessing step.
[0076] In order to convert the above abstract syntax tree into the above logical operator tree, a process of converting the nodes of the above abstract syntax tree into logical operators of a relational query optimizer (1112) is required.
[0077] As described above, the cipher query (100) can be divided into subgraph matching, path query, and relational operator, and therefore the process of converting it into the logical operator can also be divided into three stages.
[0078] In the case of subgraph matching, the C2L module (1113) can transform the portion representing the subgraph matching within the abstract syntax tree into a join tree. This process may be a process of transforming the structure of the subgraph matching query into a join operation that can be processed by the relational query optimizer (1112).
[0079] A relational query optimizer used in a relational database may be, for example, an optimizer for supporting a SQL query language. However, since the present invention must support, for example, a Cypher query language used in a graph database, operations that are not supported in the SQL query language are required. Therefore, corresponding operators, particularly variable-length path and shortest-path query operators, may be added to the relational query optimizer (1112). That is, a set of logical / physical query operators and optimization rules for a graph query language may be added to the relational query optimizer (1112). One of the main tasks of the relational query optimizer (1112) is to convert logical operators into corresponding physical operators. For example, the relational query optimizer (1112) has rules for converting logical join operators into physical hash joins, physical merge joins, and physical index nested loop joins. For example, the above optimization rule may be a rule that converts a newly added logical graph query operator that was not previously included in the relational query optimizer (1112) into a corresponding newly added physical graph query operator.
[0080] Therefore, the relational query optimizer (1112) may be an optimizer tuned to apply the graph query language of the present invention.
[0081] In the case of path queries, the C2L module (1113) can convert them using, for example, variable-length path queries and shortest path operators newly added to the relational query optimizer (1112). This process may be a step of reinterpreting the path search characteristics of the Cypher query (100) into a form that the relational query optimizer (1112) can understand and process.
[0082] The C2L module (1113) can convert relational operators by utilizing relational logical operators that already exist in the relational query optimizer (1112). This process may be a process of mapping and processing the relational data processing portion of a Cypher query to existing operators of the relational query optimizer (1112).
[0083] According to one embodiment of the present invention, the various components of a Cypher query can be effectively converted into a form that can be understood and processed by a relational query optimizer (1112) through the above-described process of the C2L module (1113). This conversion is an important step that integrates the complex structure and characteristics of a Cypher query with the advanced query processing capabilities of the relational query optimizer (1112), enabling the generation of an optimized query plan.
[0084] The MDP module (1114) of the above interface can obtain graph information (i.e., metadata) for optimizing a graph query from a graph database (1200) (specifically, a catalog manager (1210)) and provide it to a relational query optimizer (1112). The metadata may include, for example, a schema, a histogram, various statistical information (e.g., the number of tuples in a graph), and type information. At this time, the MDP module (1114) can load the metadata into the memory of the graph database system, and then convert the format to fit the data type of the relational query optimizer (1112) and provide it to the relational query optimizer (1112). Using the statistical information, the relational query optimizer (1112) can estimate cardinality using a cardinality estimation algorithm.
[0085] The P2S module (1115) of the above interface converts the physical plan provided as an output of the relational query optimizer (1112) into a physical plan (P) that can be executed in the connected graph query processor (1120) and graph database (1200). Q ) can be converted to .
[0086] In a conventional relational query optimizer, the physical plan of the relational query optimizer is converted into a file in a specific language and provided, and then the relational database creates a query plan that can be executed using the file in the specific language. However, this process incurs additional conversion overhead, which can slow down the query optimization process.
[0087] Therefore, unlike the conventional relational query optimizer (1112) method, the present invention can increase the speed of the query optimization process by directly receiving the C++ data structure as is through an API and generating a physical plan without a conversion process into a file of the specific language.
[0088] The conversion process of converting the physical plan generated by the relational query optimizer (1112) into a plan that can be executed in a graph database through the P2S module (1115) can be performed by taking into account the differences in operators between the relational database and the graph database (1200). For example, the database query execution engine (1100) used in the present invention can support a special operator called 'adjacent index join'. The main function of the adjacent index join operator is to combine connected edges (edges) for given nodes of the property graph by utilizing the adjacency list, and plays an important role in the subgraph matching process. In the physical plan of the relational query optimizer (1112), the subgraph matching process is expressed in the form of an index join. The converter of the P2S module (1115) of the present invention recognizes the operator mapping relationship between the graph database and the relational database, and performs an appropriate conversion based on this in accordance with the requirements of the graph database. For example, the appropriate transformation may mean transforming a physical index nested loop join into an AdjacencyIndexJoin, or transforming a TableScan into a NodeScan or an EdgeScan.
[0089] Additionally, if the graph database (1200) supports filter pushdown during a scan operation, the filter pushdown function can be utilized to efficiently execute queries.
[0090] The above-described physical plan transformation process may be a key element that enables the present invention to effectively link with a relational query optimizer (1112) by taking into account the characteristics of a graph database.
[0091] FIG. 6 is a flowchart illustrating a method for optimizing a query of a graph database according to one embodiment of the present invention.
[0092] Hereinafter, the description will be given with reference to FIGS. 4 to 6.
[0093] At step (S210), the interface (C2L module (1113)) can convert a graph query into a relational logic plan.
[0094] At this time, step (S210) may include a step in which the interface converts the graph query into an abstract syntax tree, and a step in which the interface converts a portion of the abstract syntax tree representing a subgraph matching of the graph query into a join tree.
[0095] At this time, the relational query optimizer (1110) may include operators and optimization rules for a graph database for optimizing the graph query. Step (S210) may further include a step in which the interface converts a portion of the abstract syntax tree representing a path query of the graph query using operators and optimization rules for the graph database.
[0096] In step (S220), the relational query optimizer (1110) may receive and optimize the converted relational logic plan from the interface. At this time, the relational query optimizer (1110) may optimize the relational logic plan using operators and optimization rules for the graph database.
[0097] At this time, step (S220) may include a step of receiving graph information for optimizing the graph query from the graph database (1200).
[0098] At this time, in step (S220), when performing optimization, the relational query optimizer (1110) may include a step of calculating a query processing cost for each plan using a pre-prepared cost model, and a step of generating a plan with the lowest cost among the calculated costs as the relational physical plan.
[0099] At this time, the pre-prepared cost model may be a second cost model tuned by changing the cost of an operator for a relational database that is mapped to an operator for a graph database by changing the cost of the operator for the graph database to the cost of the operator for the graph database, instead of the first cost model of an existing relational query optimizer that cannot process graph queries. At this time, the process of tuning the first cost model to generate the second cost model may be performed by a processor of a separate computing device.
[0100] For example, a physical index nested loop join, which is an operator for a relational database, can be mapped to an adjacency index join, which is an operator for a graph database. The cost of an adjacency index join in the graph query processing engine (1100) is much cheaper than the cost of a physical index nested loop join. That is, the cost of a physical index nested loop join in the first cost model should be set as low as the cost of an adjacency index join.
[0101] Also, in the case of the table scan operator, the relational query optimizer (1110) scans a table with a fixed schema, but when using a graph database, especially a property graph model, the schema may not be fixed to one. In this case, a different schema means that the amount of data that must be scanned for each tuple is different. Therefore, since all schemas have a fixed size, a difference occurs between a relational table that needs to scan a large amount of data at once and a graph database that needs to scan a smaller amount of data at once than a relational table. Therefore, the cost of the table scan operator must be adjusted to be lower than the cost of the matching relational operator to a certain level.
[0102] At this time, the relational query optimizer (1110) may further include an operator for a second graph database for graph queries in addition to an operator for a relational database that is mapped to an operator for a graph database. An operator for the second graph database that is not included in the operator for the relational database may be added, and for example, since the cost for a variable-length path query (variable-length traversal) or a shortest path operator is not included in the first cost model, the cost for the newly added operator for the second graph database must be applied when tuning the first cost model. The second cost model may be a model to which the cost for the newly added operator for the second graph database is applied.
[0103] In step (S230), the interface (P2S module (1115)) can convert the optimized relational physical plan into a physical plan executable in a graph database.
[0104] The P2S module (1115) may store mapping data in which operators for a relational database and operators for a graph database are mapped.
[0105] At this time, step (S230) may be a step of converting the optimized relational physical plan into a physical plan executable in the graph database based on mapping data in which an operator for a relational database and an operator for the graph query processor are mapped.
[0106] In step (S240), the interface can provide the transformed physical plan to the graph query processor.
[0107] Thereafter, the graph query processor (1120) can obtain results for a graph query from the graph database (1200) using the converted physical plan.
[0108] According to one embodiment of the present invention, the structural flexibility that can be obtained by using a modularized relational query optimizer is a distinguishing feature from the conventional Neo4j, and has the effect of generating a better query execution plan and continuously improving the performance of a graph database system.
[0109] That is, according to one embodiment of the present invention, it has a clear differentiation from the existing approach of Neo4j through the adoption of advanced query optimization strategies, more accurate cardinality estimation, and the possibility of continuous technology development through modularization.
[0110] The present invention can be of great help to businesses that require applications utilizing graph databases, such as complex network analysis, social interaction analysis, and complex data pattern recognition. Furthermore, by improving the processing speed and efficiency of Cypher queries, data analysts can gain faster and more accurate insights, which can significantly impact business decision-making and strategy formulation.
[0111] The present invention provides faster and more efficient query processing capabilities to businesses utilizing graph databases, which may provide a competitive advantage to businesses adopting the technology.
[0112] According to the present invention, a graph query optimizer can be installed in an add-on manner to a graph database system.
[0113] FIG. 7 is a conceptual diagram illustrating an example of a device, graph database system, or computing system for optimizing queries of a generalized graph database capable of performing at least part of the processes of FIGS. 1 to 6.
[0114] At least a portion of the process of a method for optimizing a query of a graph database according to one embodiment of the present invention can be executed by the computing system (2000) of FIG. 7.
[0115] Referring to FIG. 7, a computing system (2000) according to one embodiment of the present invention may be configured to include a processor (2100), a memory (2200), a communication interface (2300), a storage device (2400), an input interface (2500), an output interface (2600), and a bus (2700).
[0116] A computing system (2000) according to one embodiment of the present invention may include at least one processor (2100) and a memory (2200) that stores instructions that instruct the at least one processor (2100) to perform at least one step. At least some steps of a method according to one embodiment of the present invention may be performed by the at least one processor (2100) loading and executing instructions from the memory (2200).
[0117] The processor (2100) may mean a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor on which methods according to embodiments of the present invention are performed.
[0118] Each of the memory (2200) and the storage device (2400) may be configured with at least one of a volatile storage medium and a non-volatile storage medium. For example, the memory (2200) may be configured with at least one of a read-only memory (ROM) and a random access memory (RAM).
[0119] Additionally, the computing system (2000) may include a communication interface (2300) that performs communication via a wireless network.
[0120] Additionally, the computing system (2000) may further include a storage device (2400), an input interface (2500), an output interface (2600), etc.
[0121] Additionally, each component included in the computing system (2000) can communicate with each other by being connected by a bus (2700).
[0122] Examples of the computing system (2000) of the present invention may include a desktop computer, a laptop computer, a notebook, a smart phone, a tablet PC, a mobile phone, a smart watch, a smart glass, an e-book reader, a portable multimedia player (PMP), a portable game console, a navigation device, a digital camera, a digital multimedia broadcasting (DMB) player, a digital audio recorder, a digital audio player, a digital video recorder, a digital video player, a PDA (Personal Digital Assistant), etc.
[0123] The operations of the method according to an embodiment of the present invention can be implemented as a computer-readable program or code on a computer-readable recording medium. A computer-readable recording medium includes any type of recording device that stores information readable by a computer system. Furthermore, a computer-readable recording medium can be distributed across network-connected computer systems, allowing the computer-readable program or code to be stored and executed in a distributed manner.
[0124] Additionally, the computer-readable recording medium may include hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, flash memory, etc. The program instructions may include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the computer using an interpreter, etc.
[0125] While some aspects of the present invention have been described in the context of a device, they may also represent a description of a corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method may also be described as a corresponding block or item or a feature of a corresponding device. Some or all of the method steps may be performed by (or using) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, at least one or more of the most important method steps may be performed by such a device.
[0126] In embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In embodiments, the field-programmable gate array may operate in conjunction with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by some hardware device.
[0127] Although the present invention has been described above with reference to preferred embodiments thereof, it will be understood by those skilled in the art that various modifications and changes may be made to the present invention without departing from the spirit and scope of the present invention as set forth in the claims below.
Claims
1. An interface that converts graph queries into relational logical plans; and A relational query optimizer that optimizes the above-mentioned transformed relational logical plan; Including, The above interface converts an optimized relational physical plan generated by the relational query optimizer into an executable physical plan in a graph database, and provides the converted physical plan to a graph query processor. A device that optimizes queries in graph databases.
2. In the first paragraph, the interface stores mapping data in which an operator for a relational database and an operator for the graph database are mapped, The above interface converts the optimized relational physical plan based on the above mapping data into an executable physical plan in the graph database. A device that optimizes queries in graph databases.
3. A device for optimizing a query of a graph database, wherein the interface provides graph information for optimizing the relational logic plan to the relational query optimizer in the first paragraph.
4. In the first paragraph, when the interface converts the graph query into the relational logic plan, Convert the above graph query into an abstract syntax tree, Converting the part of the above abstract syntax tree that represents the subgraph matching of the above graph query into a join tree, A device that optimizes queries in graph databases.
5. In the fourth paragraph, the relational query optimizer includes an operator and an optimization rule for a graph database for optimizing the relational logic plan, The above interface converts a portion of the abstract syntax tree representing a path query of the graph query using an operator for the graph database and the optimization rule. A device that optimizes queries in graph databases.
6. In the first paragraph, the relational query optimizer calculates the query processing cost for each plan using a pre-prepared cost model, and generates the plan with the lowest cost among the calculated costs as the relational physical plan. The above pre-prepared cost model is a tuned cost model in which the cost of an operator for a relational database mapped to an operator for a graph database is changed to the cost of an operator for the graph database. A device that optimizes queries in graph databases.
7. In the 6th paragraph, the relational query optimizer further includes an operator for a second graph database for graph queries in addition to an operator for a relational database mapped to an operator for the graph database, The above pre-prepared cost model reflects the cost of the operator for the second graph database. A device that optimizes queries in graph databases.
8. Step of converting graph query into relational logic plan; A step of optimizing the above-mentioned converted relational logic plan; A step of converting the above optimized relational physical plan into an executable physical plan in a graph database; and A step of providing the above transformed physical plan to a graph query processor; Including, How to optimize queries in graph databases.
9. A method for optimizing a query of a graph database, wherein the step of converting the optimized relational physical plan into a physical plan executable in the graph database in the 8th paragraph is a step of converting the optimized relational physical plan into a physical plan executable in the graph database based on mapping data in which an operator for a relational database and an operator for the graph database are mapped.
10. A method for optimizing a query of a graph database, wherein the optimizing step comprises a step of receiving graph information for optimizing the relational logic plan from the graph database.
11. In the 8th paragraph, the step of converting the graph query into the relational logic plan is: The interface comprises a step of converting the graph query into an abstract syntax tree; and The above interface converts a portion of the abstract syntax tree representing a subgraph matching of the graph query into a join tree; Including, How to optimize queries in graph databases.
12. In the 11th paragraph, the step of converting the graph query into the relational logic plan is a step of converting a part of the abstract syntax tree representing a path query of the graph query using an operator and optimization rule for a graph database. How to optimize queries in graph databases.
13. In the 8th paragraph, the optimizing step comprises: A step of calculating the query processing cost for each plan using a pre-prepared cost model; and A step of generating a plan with the lowest cost among the calculated costs as the relational physical plan; Including, The above pre-prepared cost model is a tuned cost model in which the cost of an operator for a relational database mapped to an operator for a graph database is changed to the cost of an operator for the graph database. How to optimize queries in graph databases.
14. A method for optimizing a query of a graph database in accordance with claim 13, wherein the cost model prepared in advance reflects the cost of an operator for a second graph database for graph queries in addition to an operator for the relational database mapped to an operator for the graph database.
15. Interface for converting graph queries into relational logical plans; A relational query optimizer that optimizes the above-mentioned transformed relational logical plan; Graph databases; and Graph query processor; Including, The above interface converts an optimized relational physical plan generated by the relational query optimizer into an executable physical plan in the graph database, and provides the converted physical plan to the graph query processor. The above graph query processor obtains results for a graph query from the graph database using the above transformed physical plan. Graph database system.
16. In the 15th paragraph, the interface stores mapping data in which an operator for a relational database and an operator for the graph database are mapped, The above interface converts the optimized relational physical plan based on the above mapping data into an executable physical plan in the graph database. Graph database system.
17. In the 15th paragraph, when the interface converts the graph query into the relational logic plan, Convert the above graph query into an abstract syntax tree, Converting the part of the above abstract syntax tree that represents the subgraph matching of the above graph query into a join tree, Graph database system.
18. In the 17th paragraph, the relational query optimizer includes an operator and an optimization rule for a graph database for optimizing the relational logic plan, The above interface converts a portion of the abstract syntax tree representing a path query of the graph query using an operator for the graph database and the optimization rule. Graph database system.
19. In clause 15, the relational query optimizer calculates the query processing cost for each plan using a pre-prepared cost model, and generates the plan with the lowest cost among the calculated costs as the relational physical plan. The above pre-prepared cost model is a tuned cost model in which the cost of an operator for a relational database mapped to an operator for a graph database is changed to the cost of an operator for the graph database. Graph database system.
20. In the 19th paragraph, the relational query optimizer further includes an operator for a second graph database for graph queries in addition to an operator for a relational database mapped to an operator for the graph database, The above pre-prepared cost model reflects the cost of the operator for the second graph database. Graph database system.
Citation Information
Patent Citations
Data blood relationship acquisition method and device
CN116166718A
Database capable of intergrated query processing and data processing method thereof
KR101731579B1
Optimization technique for database application
KR101919771B1
Rule-Based Extendable Query Optimizer
US20140330807A1
Performing Complex Operations in a Database Using a Semantic Layer
US20150120699A1
Cited By
Graph database query method and system based on stream-oriented computation
CN121188061A