Ontology-based query routing for distributed knowledge bases

By using the query coordinator in the distributed knowledge base to generate query blocks and identify data nodes, the problem that existing systems cannot effectively route queries is solved, and efficient query processing and response are achieved.

CN114586028BActive Publication Date: 2025-05-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080070100.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2020-09-30
Publication Date
2025-05-23
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

Existing systems cannot effectively manage and route queries in distributed knowledge bases, resulting in inefficiency and waste of computing resources.

Method used

Receive ontology queries through the query coordinator, generate query blocks, and identify data nodes based on quantifiers and operations, select appropriate data nodes and send subqueries to achieve effective query routing.

Benefits of technology

Improves the system's responsiveness and efficiency, reduces the computational overhead, and optimizes the latency of query execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114586028B_ABST
    Figure CN114586028B_ABST
Patent Text Reader

Abstract

A technique for query routing is provided. An ontology query is received by a query coordinator. One or more query blocks are generated based on the ontology query, each query block indicating one or more operations and one or more quantifiers representing data flow between the query blocks. For each query block in the one or more query blocks, at least one data node is identified based on the one or more quantifiers and the one or more operations. One or more of the identified data nodes are selected based on a predefined cost criterion, and one or more subqueries are sent to the selected one or more data nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to knowledge bases, and more particularly to efficient query routing in distributed knowledge bases using ontologies. Background Art

[0002] More and more enterprises are using knowledge bases (KBs) to enhance their analysis and improve the decision-making, efficiency and effectiveness of their systems. Typically, within the domain of an enterprise, a KB is relatively specialized. For example, financial institutions rely on KBs with important financial knowledge, such as data related to government regulations of the financial market. On the contrary, a health care enterprise may maintain a KB with a considerable amount of data collected from medical literature. There is a substantial need for deep domain specialization, effective systems and effective technologies to manage these KBs. Existing systems that do not know the domain ontology that provides an entity-centric view of the domain model cannot effectively manage and route queries to the KB.

[0003] Additionally, KBs can be distributed across multiple data sites with different capabilities and costs in order to improve operations. Existing architectures such as federated databases rely on centralized mediators to aggregate data from each such source. This is inefficient and scales poorly. Additionally, existing systems do not understand the underlying ontology and capabilities of each data source and therefore cannot route queries efficiently. This reduces the efficiency of such systems, requiring a large amount of computing resources to respond to typical queries. Summary of the invention

[0004] According to one aspect of the present invention, a method is provided, comprising: receiving an ontology query by a query coordinator, and generating one or more query blocks based on the ontology query, each query block indicating one or more operations and one or more quantifiers representing the data flow between the query blocks. The method also includes, for each query block in the one or more query blocks, identifying at least one data node based on one or more quantifiers and one or more operations. In addition, the method includes selecting one or more data nodes based on a predefined cost criterion, and sending one or more subqueries to the selected one or more data nodes. Advantageously, the method enables the query coordinator to effectively route queries in a distributed environment based on the operations and quantifiers required for each block of the query. This reduces computational overhead and improves system responsiveness.

[0005] According to an embodiment of the present disclosure, identifying at least one data node for a first query block in one or more query blocks also includes identifying a first set of data nodes configured to perform one or more operations indicated by the first query block based on predefined capability standards, and identifying a second set of data nodes configured to store one or more quantifiers indicated by the first query block based on a predefined concept mapping. In such an embodiment, the query coordinator improves on existing systems by considering the capabilities of the underlying data nodes and the location of the storage of concepts in the ontology. This again improves efficiency and reduces computational waste.

[0006] According to another embodiment of the present disclosure, the method further includes, when determining that the first data node belongs to both the first set of data nodes and the second set of data nodes, associating the first query block with an annotation indicating the first data node. Advantageously, this enables the query coordinator to intelligently route query blocks to data nodes by identifying nodes that can provide the necessary quantifiers and perform the necessary operations, which reduces or eliminates data transfers between nodes when completing queries. This thereby reduces the computational overhead of executing queries.

[0007] According to another embodiment of the present disclosure, the method further includes, when it is determined that no data node belongs to both the first set of data nodes and the second set of data nodes, (i) identifying a first data node belonging to the first set of data nodes, (ii) associating the first query block with an annotation indicating the first data node, and (iii) generating one or more additional query blocks to satisfy one or more quantifiers of the first query block. Advantageously, such an embodiment enables the query coordinator to efficiently split the query blocks to allow them to be executed within separate data nodes. This allows the appropriate nodes to be identified immediately, and reduces transmission costs and improves latency in executing queries.

[0008] According to another embodiment of the present disclosure, the method further includes, when determining that one or more operations indicated by the first query block include a join operation between two or more quantifiers, identifying a first data node belonging to the first set of data nodes, and when determining that the first data node is configured to store at least one of the two or more quantifiers, associating the first query block with an annotation indicating the first data node. This embodiment provides an advantage by enabling the query coordinator to minimize data transfer costs by intelligently selecting nodes based on data of a node repository using a predefined mapping.

[0009] According to another embodiment of the present disclosure, the method also includes associating the first query block with an annotation indicating the first data node based on further determining that the first data node is a relational data node, wherein based on determining that the second data node is not a relational data node, excluding from the annotation a second data node that is also configured to store at least one of the two or more quantifiers. Advantageously, such an embodiment enables the query coordinator to utilize the relative strength of each data node when determining an efficient routing plan. For example, because relational data stores typically perform join operations more efficiently than other types of backend repositories, the query coordinator can select a relational repository for join operations whenever possible.

[0010] According to another embodiment of the present disclosure, the method also includes selecting one or more of the data nodes based on a predefined cost criterion by identifying each possible combination of data nodes across each query block in one or more query blocks, generating a corresponding mobility descriptor for each corresponding possible combination of data nodes, the corresponding mobility descriptor defining the cost of transmitting data between the corresponding combination of data nodes; and selecting a first combination of data nodes based on determining that the mobility descriptor of the first combination of data nodes is lower than the mobility descriptors of all other possible combinations of data nodes. An advantage of this embodiment is that the computational efficiency of each potential routing plan can be evaluated before selecting a plan. This improves efficiency and reduces the cost and latency of query execution.

[0011] According to another embodiment of the present disclosure, the method further includes, for each corresponding query block in one or more query blocks, identifying a corresponding query fragment corresponding to the corresponding query block from the ontology query, determining the type of data node assigned to the corresponding query block, selecting a corresponding query translator based on the determined type of data node assigned to the corresponding query block, and generating a corresponding subquery by processing the corresponding query fragment using the corresponding query translator. Advantageously, this enables the query coordinator to translate subsections of the query based on the target data node using a predefined mapping. This improves the scalability of the system and reduces the response time of each selected backend.

[0012] According to different embodiments of the present invention, any of the above embodiments may be implemented by a computer-readable storage medium. The computer-readable storage medium contains computer program code, which, when executed by the operation of one or more computer processors, performs operations. In an embodiment, the operations performed may correspond to any combination of the above methods and embodiments.

[0013] According to another different embodiment of the present disclosure, any of the above embodiments may be implemented by a system. The system includes one or more computer processors and a memory containing a program, and when executed by one or more computer processors, the program performs operations. In an embodiment, the operations performed may correspond to any combination of the above methods and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 An architecture configured to place data of a knowledge base and route ontology queries according to one embodiment disclosed herein is shown.

[0015] Figure 2 A workflow for processing and routing ontology queries according to one embodiment disclosed herein is shown.

[0016] Figure 3A and Figure 3B Depicted is an example ontology that may be used to evaluate and route queries according to one embodiment disclosed herein.

[0017] Figure 4A and Figure 4B A workflow for parsing and routing an example ontology query according to one embodiment disclosed herein is shown.

[0018] Figure 5 is a block diagram illustrating a query processing coordinator configured to route ontology queries according to one embodiment disclosed herein.

[0019] Figure 6 is a flow chart illustrating a method for processing and routing ontology queries according to one embodiment disclosed herein.

[0020] Figure 7 is a flow chart illustrating a method for processing an ontology query to efficiently route query chunks according to one embodiment disclosed herein.

[0021] Figure 8 is a flow chart illustrating a method for evaluating potential query routing plans for efficiently routing ontology queries according to one embodiment disclosed herein.

[0022] Fig. 9 is a flow chart illustrating a method for routing ontology queries according to one embodiment disclosed herein.

[0023] Fig.10 A workflow for evaluating a workload and storing knowledge base data according to one embodiment disclosed herein is shown.

[0024] Fig.11 Depicted is an example hypergraph for evaluating workloads and storing knowledge base data according to one embodiment disclosed herein.

[0025] Fig.12 is a block diagram illustrating a data placement coordinator configured to evaluate workloads and store data according to one embodiment disclosed herein.

[0026] Fig.13 is a flow chart illustrating a method for evaluating and summarizing query workloads to inform data placement decisions according to one embodiment disclosed herein.

[0027] Fig.14 is a flow chart illustrating a method for modeling an ontology workload to inform data placement decisions according to one embodiment disclosed herein.

[0028] Fig.15 is a flow chart illustrating a method for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein.

[0029] Fig.16 is a flow chart illustrating a method for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein.

[0030] Fig.17 is a flow chart illustrating a method for mapping ontology concepts to storage nodes according to one embodiment disclosed herein. DETAILED DESCRIPTION

[0031] Embodiments of the present disclosure provide an ontology-driven architecture that uses multiple data stores or nodes to provide support for various query types, each of which can have different capabilities. Advantageously, embodiments of the present disclosure enable queries to be efficiently processed and routed to a diverse data repository based on existing ontologies, which improves system latency and reduces the computing resources required to identify and return relevant data. In an embodiment, the technology described herein can be used to support a variety of query applications, including natural language and conversational interfaces. The system described herein can provide transparent access to the underlying backend repository via an abstract ontology query language (OQL), which allows users to express their information needs for domain ontologies.

[0032] In one embodiment, in order to provide improved performance for different query types, the system first optimizes the placement of KB data into different repositories by placing subsets of the data in appropriate backend repositories according to their capabilities. In some embodiments, at runtime, the system parses and compiles OQL queries and translates them into query languages ​​and / or APIs for different backend data nodes. In one embodiment, the system described herein also employs a query coordinator that routes queries to a single or multiple backend repositories based on the placement of relevant data. This efficient routing reduces the latency required to generate results and return them to the requesting entity.

[0033] In an embodiment, the KB data to be queried may include any information, including structured, unstructured and / or semi-structured data. In order to support deep domain specialization, the embodiments disclosed herein utilize domain ontology. It is worth noting that in one embodiment, the domain ontology only defines entities and their relationships at the metadata level, without providing instance level information. That is, in one embodiment, the domain ontology provides the domain model, and the instance level data is stored in various back-end data repositories (also referred to as data sites and data nodes). In an embodiment, a data node may include any repository architecture, including (one or more) relational databases, (one or more) inverted index document repositories, (one or more) JavaScript Object Notation (JSON) repositories, (one or more) graph databases, and the like.

[0034] In some embodiments, the system can move, store and index KB data in any backend, and the backend provides the capabilities required to support query types. In other words, the system does not need to store joint data across existing operational data. This is significantly different from the joint database and the mediator-based method. In some embodiments, the system performs a data placement step based on capabilities, wherein the knowledge base data that conforms to a given ontology is stored in various backend resources. In at least one embodiment, the multi-repository architecture is configured to minimize data movement between different data repositories, because data movement not only causes data transmission costs, but also causes expensive data conversion costs. A solution to minimize data movement costs is to replicate data in all backend nodes, thereby ensuring that queries can be answered by a single repository without any data movement. However, this also requires wasteful replication, and invalidates the purpose of using a multi-repository architecture to balance the different capabilities of multiple data storage devices. In an embodiment of the present invention, KB data can be stored in any number of backend resources, and the system intelligently routes queries to select (one or more) the most appropriate data repository and minimize data movement costs.

[0035] In some embodiments described herein, the system architecture provides appropriate abstractions to query data without knowing how the data is stored and indexed in multiple data repositories. In one embodiment, to achieve this goal, an ontology query language (OQL) is introduced. OQL is expressed for the domain ontology of an enterprise. Users of the system only need to know the domain ontology that defines entities and their relationships. In an embodiment, the system understands the mapping of various ontology concepts to backend data repositories and their corresponding schemas, and provides a query translator from OQL to the target query language of the underlying system.

[0036] In an embodiment, for any given query, one or more data nodes may be involved. Embodiments of the present disclosure provide techniques for identifying the repositories involved and generating appropriate subqueries to calculate the final query results. In some embodiments, the system does not use a mediator approach, in which a global query is divided into multiple subqueries and the final result is assembled in a mediator. Instead, one or more of the data repositories are used as mediators, and (one or more) back-end data repositories finalize the query answers. For example, in one embodiment, a relational database is used to finalize query responses because the relational repository is likely to be able to quickly complete a join operation.

[0037] Figure 1 An architecture 100 configured to route ontology queries according to one embodiment disclosed herein is shown. In the illustrated embodiment, a query processing coordinator 115 receives an ontology query 105 constructed using OQL, which has a corresponding ontology schema 110. The ontology query 105 is processed by the query processing coordinator 115 to identify and return relevant data. In an embodiment, the query processing coordinator 115 routes the query to a backend 125 including various data nodes 130A-N. In the illustrated embodiment, the routing is performed at least in part based on a set of concept maps 140 and capabilities 135.

[0038] In the illustrated embodiment, ontology query 105 is formatted based on OQL, which is used to represent queries that operate on a set of concepts and relationships in ontology schema 110. In one embodiment, OQL can express queries including aggregation, union, nested subqueries, etc. In some embodiments, OQL can also express full-text and field search predicates, as well as path queries. OQL queries typically include a single query block or a combination of multiple OQL query blocks, wherein each OQL query block includes a SELECT clause and a FROM clause. In some embodiments, an OQL block may also include a WHERE clause, a GROUP BY clause, an ORDER BY clause, a FETCH FIRST clause, and / or a HAVING clause. In an embodiment, an OQL query operates on a set of tuples constructed by the Cartesian product of the concepts referenced in the FROM clause.

[0039] In an embodiment, the ontology schema 110 describes entities and their relationships in a KB at a semantic level, regardless of how the data is actually stored in the backend resources. In one embodiment, the ontology schema 110 describes entities related to the domain, attributes potentially associated with various entities, and potential relationships between different entities. The ontology schema 110 can provide a rich and expressive data model that captures various real-world relationships between entities, such as functional, inheritance, association, etc. In some embodiments, the ontology schema 110 does not include instance data. That is, the ontology schema 110 defines entities and relationships, but does not include data related to specific entities or relationships in the KB. For example, the ontology schema 110 can define that a "Company" entity can have several attributes such as "name" and "address", but the ontology schema 110 does not include data for a specific instance, such as a company with the name "Main Street Grocer" and the address "123 Main Street". Instead, this instance data is maintained separately in the backend 125.

[0040] In an embodiment, the concept map 140 indicates the correspondence between the logical schema represented by the ontology schema 110 and the physical schema of the underlying data nodes 130 in the backend 125. For example, assuming that the ontology schema 110 is defined as Where C = {c n |1≤n≤N} represents a set of concepts (also called entities), R = {r k |1≤k≤K} represents the set of relations between concepts / entities, and P = {p m|1≤m≤M} is a set of data attributes. In one embodiment, each relationship is between two or more concepts, and each data attribute corresponds to a characteristic of a concept. In the illustrated embodiment, the concept map 140 indicates the mapping between concepts, relationships, and attributes (defined in the ontology schema 110) and the underlying data nodes 130. That is, in one such embodiment, the concept map 140 indicates each data node 130 storing a given concept (e.g., entity), relationship, and / or attribute. For example, the concept map 140 may indicate that instances of the "Company" entity are stored in data nodes 130A and 130M, while data related to a specific type of relationship between the "Company" entities is stored in data node 130B.

[0041] In one embodiment, the capabilities 135 indicate the operations and capabilities provided by each data node 130. In some embodiments, the capabilities of the data node 130 are expressed as views that list all possible queries (view definitions) that can be processed / answered by the data repository. Although this approach is flexible, it is not scalable because the number of view definitions may be very large, potentially leading to the problem of query rewriting using an unlimited number of views. In order to improve scalability, some embodiments of the present disclosure describe their capabilities based on the operations supported by the backend data nodes 130 (e.g., joins, groupings, aggregations, fuzzy text matching, path expressions, etc.) rather than listing all possible queries that the data repository can answer. In addition, at least one embodiment of the present disclosure provides a more fine-grained description of each supported operation by utilizing a mechanism for expressing any associated restrictions. For example, in one such embodiment, an aggregation function of type MAX may only be supported on numeric types.

[0042] Thus, in one embodiment, the capabilities 135 indicate, for each given data node 130, a set of operations that the node can perform (and, in some embodiments, associated restrictions on the capabilities). In an embodiment, the query processing coordinator 115 evaluates the capabilities 135 and the concept maps 140 to select one or more data nodes 130 to which the received ontology query 105 should be routed. To this end, in one embodiment, the query processing coordinator 115 identifies data nodes 130 that contain the desired data (e.g., based on the concept maps 140), are capable of completing the desired operation(s) (e.g., based on the capabilities 135), or both. The query processing coordinator 115 may then generate a subquery for each selected data node 130.

[0043] In an embodiment, the query processing coordinator 115 uses a set of translators 120A-N to translate subqueries as needed. In the illustrated embodiment, each type of data node 130 has a corresponding translator 120. In some embodiments, each translator 120A-N receives all or part of an OQL query and generates an equivalent query in the language and / or syntax of the corresponding data node 130A-N. For example, translator 120A can generate a SQL query for a relational database contained in data node 130A, while translator 120B outputs a graph query for a graph repository contained in data node 130B.

[0044] In one embodiment, the translator 120 relies on a schema mapping that maps concepts and relationships represented in a domain ontology to appropriate schema objects in a target physical schema. For example, for a relational backend data node 130, a schema mapping may provide (1) correspondences between concepts in the ontology and tables in the relational schema; (2) data attributes or properties of concepts in the ontology and table columns in the physical schema; and / or (3) relationships between primary key-foreign key constraints between concepts in the ontology and tables corresponding to the concepts in the database. Similarly, for a JSON document repository, a schema mapping may map concepts, data attributes, and relationships represented in the ontology to appropriate field paths in a JSON document.

[0045] In some embodiments, the translator 120 also handles special concepts and relationships that can be represented in the ontology, such as joins, inheritance, and traversals between concepts that typically represent connection conditions between ontology concepts. Depending on the physical data layout, these are translated into appropriate operations supported by the backend 125 data nodes 130.

[0046] In an embodiment, each data node 130 receives a subquery and generates a response for the query processing coordinator 115. If the data node 130 has all the necessary data locally and is able to complete the indicated operation(s), the node executes the query and returns the result. In some embodiments, the subquery may indicate that the data node 130 should send a query to one or more other data nodes 130 to retrieve the data required to complete the query. For example, in some embodiments, it is preferred that the join operation is performed by the relational data node because the relational system tends to have low latency for join operations. In addition, some operations are only possible on certain nodes. For example, assume that data node 130A is the only node that can complete the join operation, but the data to be connected only exists on data node 130B. In one embodiment, data node 130A receives a subquery indicating that it connects to related data, as well as one or more other subqueries to be forwarded to data node 130B to retrieve data.

[0047] That is, in one such embodiment, the query processing coordinator 115 prepares the subqueries for the translation of data node 130B and sends them to data node 130A. This allows data node 130A to simply forward them to data node 130B. Data node 130B then returns the data to data node 130A, which completes the operation and returns the results to the query processing coordinator 115. In another embodiment, the query processing coordinator 115 may send the subqueries to data node 130B and forward the resulting data to data node 130A for processing. In yet another embodiment, the query processing coordinator 115 performs the join locally.

[0048] In the illustrated embodiment, the architecture 100 also includes a data placement coordinator 150 that determines which data node(s) 130A-N should be used to store data in the knowledge base. As shown, the data placement coordinator 150 receives similar indications of the capabilities 155 of the data nodes 130, as well as an indication of a query workload 160. In one embodiment, the query workload 160 indicates an average or expected set of queries for the knowledge base. For example, in one embodiment, the query workload 160 is generated by observing the user's interaction with the knowledge base over time. The system can then aggregate these interactions to determine the average, expected, or typical workload of the system. In some embodiments, the query workload 160 includes indications of which concepts to query together, which operations to apply to each concept, and the like.

[0049] In an embodiment, the data placement coordinator 150 determines, for each concept / entity in the ontology schema 110, which data node(s) 130 should store instance-level data for the concept based on the node capabilities 155 and the known query workload 160. For example, based on the operations that a given data node 130A is able to perform and further based on the query workload 160, the data placement coordinator 150 may determine that data node 130A should store all data for the "company" entity and the "public metric" entity because they are often queried together. Similarly, the data placement coordinator 150 may determine to place instances of the "document" concept in data node 130B based on determining that queries frequently include performing fuzzy matching operations on "document" data and that data node 130B supports fuzzy matching.

[0050] Such intelligent data placement can reduce subsequent data transfers during runtime. In embodiments, the data placement coordinator 150 can be utilized at startup (e.g., to provide initial placement) and / or periodically during runtime (e.g., to reformulate placement decisions based on how the query workload 160 evolves over time). Although depicted as separate components for conceptual clarity, in embodiments, the operations of the query processing coordinator 115 and the data placement coordinator 150 can be combined or distributed across any number of components.

[0051] As shown, the data placement coordinator 150 outputs a set 165 of placement decisions to the backend 125, and / or to one or more intermediate services, such as extract, transform, and load (ETL) services. Then, based on these selections, the instance-level data in the knowledge base is stored in the appropriate data nodes 130. The concept map 140 reflects the current placement of the data. In an embodiment, if the data placement coordinator 150 is used to modify the data placement based on the updated query workload 160, the concept map 140 is similarly updated. In the following, Figures 2 to 9 The query processing coordinator 115 is discussed in more detail and assumes that the data has already been placed. Figures 10 to 17 The data placement coordinator 150 and various techniques to ensure efficient data placement are discussed in greater detail.

[0052] Figure 2 A workflow 200 for processing and routing ontology queries according to one embodiment disclosed herein is shown. As shown, the workflow 200 begins when an OQL query 205 is received. The OQL query 205 is provided to an OQL parser 210, which parses the query to determine the meaning of the query. As shown by arrow 215, the parsed OQL query is then provided to a query graph model (QGM) builder 220. In one embodiment, the QGM builder 220 generates a logical representation of the query in the form of a query graph model. This reduces the complexity of query compilation and optimization. Figure 4A and Figure 4B Discussing an example query graph model in more detail. In an embodiment, QGM uses operator boxes such as SELECT, GROUP BY, SETOP, etc. to capture data flow and dependencies in a query. In an embodiment, operations within a box can be freely reordered between them, but box boundaries are respected when generating a query execution plan. That is, the query execution plan must follow the order of the query box.

[0053] In one embodiment, quantifiers are used to represent the data flow between query blocks. This format allows the system to reason about query equivalence and apply rewrite optimizations. In an embodiment, the QGM representation of a query enables the system to focus on optimizing the data flow between different data repositories during query execution, while the selection of the actual physical execution plan is deferred to the underlying data nodes 130 responsible for executing the query box or fragment. In an embodiment, the QGM 225 is used to generate an optimized multi-store execution plan that minimizes data movement and transformation across different backends.

[0054] In some embodiments, the QGM 225 includes a set of quantifiers at the bottom that provide a set of input concepts to the query, and each query block in the QGM has a header and a body. The body of each block includes a set of predicates that describe the set operations to be performed on the input concept set (such as a join), and the header expression describes how the output properties of the result concepts should be calculated. In other words, in an embodiment, the body of each box contains a set of predicates (also called operations), each of which will be applied to the input quantifiers of the box. In some embodiments, predicates / operations that reference a single quantifier are classified as local predicates, while predicates / operations that reference multiple quantifiers express connection predicates.

[0055] In the illustrated workflow 200, the QGM 225 is passed to an operator placement component 230 that continues the query routing process. As shown, the operator placement component 230 also receives a collection of concept maps 235 and node capabilities 240, and annotates the query blocks in the QGM 225 based on the capabilities and concept maps. In one embodiment, the operator placement component 230 moves the QGM from the bottom to the top and annotates each operation in the query with a collection of possible repositories that can perform the operation. In some embodiments, annotations are generated differently for local predicates and join predicates. Recall that in some embodiments, all head expressions associated with a single quantifier are processed in the same manner as local predicates.

[0056] In one embodiment, if the query block includes a local predicate / operation (e.g., a single quantifier), the operator placement component 230 determines whether the quantifier is a base concept. If so, the predicate is annotated with an indication of the data nodes 130 that (i) contain the concept and (ii) have the ability to execute the predicate. In an embodiment, if the quantifier comes from another QGM block (i.e., it is calculated by another block), the operator placement component 230 annotates the predicate with a set of data repositories that (i) complete the QGM box for the quantifier and (ii) can execute the predicate. In some embodiments, if none of the data nodes 130 that contain inputs to the predicate also have the ability to execute the predicate, the operator placement component 230 annotates it with a set of repositories that can execute the predicate.

[0057] For example, if the predicate contains a fuzzy search, but the data is only stored in the relational backend data nodes 130, the operator placement component 230 can annotate the predicate with a document repository that can compute the fuzzy search, even though the data is not stored there. Note that during query execution, data will need to be moved in this case.

[0058] In some embodiments, if the query block includes a join predicate / operation, the operator placement component 230 examines the join type and join predicate. In an embodiment, each join predicate is associated with two or more quantifiers: one for each join input. In an embodiment, the operator placement component 230 identifies the occurrences of these quantifiers based on whether the identified quantifier is an input to a calculation or a base concept and where the quantifier comes from (e.g., is it calculated and / or stored locally, or will it be received from another node).

[0059] In the case where all quantifiers cover the base concepts, the operator placement component 230 evaluates the set of data nodes 130 where any base concepts reside, and for each such repository, determines whether the repository supports the join operation. The operator placement component 230 then annotates the join operation with an indication of the repository of the type that contains the one or more required concepts and supports the join operation. In some embodiments, if one of the repositories is a relationship node, the operator placement component 230 annotates the operation with an indication of the relationship data node 130 rather than the remaining nodes.

[0060] In some embodiments, if both quantifiers are generated by other QGM query blocks, the operator placement component 230 annotates the join operation with a data node 130 that supports the join operation type (in some embodiments, only relationship repositories). In an embodiment, if the data repository that generated the quantifier is capable of performing joins (JOINs), the operator placement component 230 also annotates the operation with their indications.

[0061] In another embodiment, if one of the quantifiers is a base concept and the other is computed from another QGM query block, the join operation placement decision is similar to the above discussion when both quantifiers are computed by the other QGM block. In an embodiment, the operator placement component 230 annotates the join operation with a collection of data repositories that support the join type and are local to at least one of the join inputs.

[0062] As shown, the annotated QGM 245 is then passed to the block placement component 250. In one embodiment, the block placement component 250 utilizes the annotations generated by the operator placement component 230 to determine possible placement options for query boxes in the query.

[0063] In one embodiment, determining the placement options for a "select" query box is modeled as the problem of determining a minimum set cover. That is, the block placement component 250 determines the minimum number of data nodes 130 required to satisfy the placement of all operators within a given select box. In some embodiments, the block placement component 250 performs the predicate-repository grouping using a greedy heuristic algorithm that ensures that each predicate is placed in a single repository and the total number of repositories spanned by the query box is minimized. Once each predicate is annotated with the appropriate repository where it will be placed, the block placement component 250 determines whether the query block needs to be split.

[0064] In an embodiment, if all predicates in a select query box are placed in the same data node 130, the query box is annotated to be placed in that repository. Conversely, if the predicates in the select query box are placed in more than one data node 130 (indicating that the query box needs to be processed by multiple repositories), the block placement component 250 determines that the block must be split into multiple query boxes. In this case, the block placement component 250 then divides the block so that each result (sub)block includes predicates assigned to a single data node.

[0065] In some embodiments, for a "group by" query box, the block placement component 250 determines placement based on a repository of the type that supports aggregation operations and the repository that processed the previous query block of the query box that feeds the group by box. In an embodiment, if the repository that feeds these inputs also supports "group by" and "aggregation" functions, the block placement component 250 places the selected "group by" box in the same repository. However, if the repository does not support these operations, the block placement component 250 annotates the box with a list of data nodes 130 that can perform the group by operation. Note that this placement will require data movement from the feeding repository to the repository that can process the box.

[0066] Once the block placement component 250 has annotated each query block with all possible placements (one or more) of each query block (e.g., all data nodes 130 that can execute the block), the cost component 255 determines the cost of each combination of placements. In one embodiment, the cost model used for query routing focuses only on data transfer and data transformation costs in order to select between alternative query execution plans, because the cost of data movement and transformation may dominate the overall cost of any execution plan in a multi-repository environment. Therefore, in at least some embodiments, the cost of executing a given execution plan is determined by aggregating the cost of data movement for each source and target data repository pair in the query execution plan. Note that in some embodiments, different operations such as joins and graph operations can be executed with very different performance on different data nodes 130. For example, although join operations may be supported by various data repositories (such as JSON repositories), relational databases typically provide the best performance for join operations.

[0067] In some embodiments, the cost model also takes into account such variations in the execution costs of operations for different repositories. In other embodiments, the cost model does not include these factors. In one such embodiment, a declarative mechanism that expresses the capabilities of the underlying repository is used to handle such situations. For example, to avoid performing a connection operation on a JSON repository, the system can completely block the connection capability of the JSON repository in the capability description, or add restrictions to limit its applicability to specific data types.

[0068] In an embodiment, the cost component 255 enumerates all possible placement combinations on all query blocks, and generates an execution plan for each such combination by grouping certain query boxes together into groups. In an embodiment, each group in the query plan includes query boxes that are both (i) results in the QGM (e.g., separated by a single hop) and (ii) processable by the same data node 130. The cost component 255 then connects the groups of these blocks based on the QGM flow, defining edges that represent the data flow between the data nodes 130. In an embodiment, the cost component 255 also generates a set of data movement descriptors for each execution plan based on the edges between the groups in the generated plan.

[0069] In some embodiments, using these movement descriptors for each plan, the cost component 255 uses the above-mentioned data movement cost model to find the cost for each execution plan, and picks the one with the minimum cost. For example, in one embodiment, the cost component 255 determines the cost of each of the corresponding movement descriptors for each possible execution plan. This can include delays, computational costs, etc. These values ​​can then be aggregated within each execution plan to determine which (which) execution plans have the lowest cost. As shown, the minimum cost plan is then selected and used to generate a query 260 of one or more translations of the data node 130.

[0070] Figure 3A and Figure 3B Depicted is an example ontology 300 that can be used to evaluate and route queries according to an embodiment disclosed herein. In the illustrated embodiment, an ellipse is used to depict each concept 305A-I, while a rounded rectangle is used to depict each attribute 310A-N. In addition, arrows are used to depict associations between concepts 305 and attributes 310, while bold arrows depict relationships between concepts 305. In addition, each relationship arrow is labeled based on the type of relationship. For example, as shown, a "Company" concept 305B is a subclass of a "Public Company" concept 305D. In addition, the company concept 305B is associated with several attributes 310, including a "name" attribute 310E and an identifier attribute 310D, both of which are strings.

[0071] As described above, in an embodiment, the data in the knowledge base conforms to an ontology 300. In other words, the ontology 300 defines the concepts and entities in the knowledge base, as well as the relationships between the entities and the attributes associated with each concept. It is worth noting that the ontology 300 does not include any instance data (e.g., data about a particular company), but rather defines the structure of the data. In some embodiments, concepts 305, attributes 310, and / or relationships may be distributed across any number of data nodes 130. In one embodiment, data is placed in a data node 130 based at least in part on the type of data.

[0072] In one such embodiment, if the concept map indicates that a given concept 305 is stored in a given data node 130, then all instance data corresponding to that concept 305 is stored in the data node 130. For example, if the map indicates that data node 130A includes a "PublicMetric" concept 305F, then the query processing coordinator 115 may retrieve data about any instance of "PublicMetric" (e.g., any metric for any company) from data node 130A. Thus, if a received query would require access to "PublicMetric" data, the query processing coordinator 115 would route at least a portion of the query to data node 130A (or to another node that also serves concept 305F).

[0073] Figure 4A and Figure 4B A workflow 400 for parsing and routing example ontology queries according to one embodiment disclosed herein is shown. In the illustrated embodiment, input 405 is received and evaluated to retrieve relevant data from data nodes 130. In the illustrated embodiment, input 405 is natural language text (e.g., from a user). However, in various embodiments, input 405 may include queries or other data. In addition, input 405 may be received from any number of sources, including automated applications, user-oriented applications, directly from users, etc. In some embodiments, input 405 is received as part of a chatbot or other interactive application that allows a user to search and explore a knowledge base.

[0074] In the example shown, input 405 is the phrase "Show me the total revenue of all companies that filed technology patents in the last 5 years". In the illustrated workflow 400, input 405 is parsed and evaluated using one or more natural language processing (NLP) techniques such as semantic analysis, keyword search, sentiment analysis, intent analysis, etc. This enables the system to generate an OQL query 410 based on input 405. Of course, in some embodiments, input 405 itself is an OQL query.

[0075] As shown, the corresponding OQL query 410 for the input 405 includes a "select" operation and a "group by" operation, as well as an indication of the relevant tables or concepts, and a "where" clause indicating restrictions on the query. The OQL query 410 is then parsed by the query processing coordinator 115 to generate an OGM 435A, which is a logical representation of the query and includes multiple query blocks. In the illustrated embodiment, the OGM includes a "select" query block 415B and a "group by" query block 415A.

[0076] As shown, QGM 435A includes a set 430A-E of quantifiers at the bottom, which provides an input concept set to the query. In the illustrated embodiment, these include "PublicMetricData", "PublicMetric", "PublicCompany", "Document" and "CompanyInfo". In an embodiment, each query block 415 in QGM 435A includes a header 420 and a body 425. The body 425 generally describes the set operation to be performed on the input concept set (such as a connection), and the header 420 expression describes how to calculate the output attributes of the result concept. For example, in the illustrated embodiment, the header 420B of the "select" query block 415B specifies the output attributes "oPMD.value", "oPMD.year_calendar" and "oCl.id". The body 425B contains a set of predicates to be applied to the input quantifiers 430A-E. As described above, in an embodiment, the predicates quoting a single quantifier are local predicates, while the predicates quoting multiple quantifiers express connection predicates.

[0077] like Figure 4B As shown, because the query processing coordinator 115 recognizes that the "select" query block 415B includes predicates that must be executed by different data nodes 130. Specifically, although most of the predicates in the query block 415B can be executed by the relational data repository, the query processing coordinator 115 has determined that the predicates "oD->companylnfo=oCI" and "oD.selfMATCH('Tech Patent Filed')" correspond to operations performed by elastic search that are not supported by relational database nodes. Therefore, the query processing coordinator 115 splits the query block 415B by separating these predicates into a new query block 415C. As shown, the query block 415C serves as the input of the query block 415B.

[0078] In the illustrated embodiment, by splitting query chunks 415B, the query processing coordinator 115 has ensured that each chunk can be executed entirely within a single data node 130. For example, query chunks 415A and 415B can both be executed in a relational data node 130, while query chunk 415C will be executed by an elasticsearch-enabled data node 130. Thus, in an embodiment, the query processing coordinator 115 identifies or generates subqueries corresponding to query chunks 415A and 415B, translates the subqueries into the appropriate language and / or format for the relational node, and sends the translated subqueries to the relational node.

[0079] In the illustrated embodiment, the query processing coordinator 115 will further identify or generate a subquery to complete the query block 415C and translate it into a language and / or format supported by the elastic search node. In one embodiment, the query processing coordinator 115 additionally sends the subquery to the relational data repository, which will act as an aggregator. The relational node can then forward the subquery to the elastic node and use the results returned by the elastic node to complete its own subquery.

[0080] Figure 5 is a block diagram illustrating a query processing coordinator 115 configured to route ontology queries according to one embodiment disclosed herein. Although depicted as a physical device, in an embodiment, the query processing coordinator 115 may be implemented using (one or more) virtual devices and / or across multiple devices (e.g., in a cloud environment). As shown, the query processing coordinator 115 includes a processor 510, a memory 515, a storage device 520, a network interface 525, and one or more I / O interfaces 530. In the illustrated embodiment, the processor 510 retrieves and executes programming instructions stored in the memory 515, and stores and retrieves application data residing in the storage device 520. The processor 510 generally represents a single CPU and / or GPU, multiple CPUs and / or GPUs, a single CPU and / or GPU with multiple processing cores, etc. The memory 515 is generally included to represent a random access memory. Storage device 520 can be any combination of disk drives, flash-based storage devices, etc., and can include fixed and / or removable storage devices, such as fixed disk drives, removable memory cards, cache, optical storage devices, network attached storage devices (NAS), or storage area networks (SAN).

[0081] In some embodiments, input and output devices (such as keyboards, monitors, etc.) are coupled via the I / O interface(s) 530. In addition, via the network interface 525, the query processing coordinator 115 may be communicatively coupled with one or more other devices and components (e.g., via a network 580, which may include the Internet, local network(s), etc.). As shown, the processor 510, memory 515, storage device 520, network interface(s) 525, and I / O interface(s) 530 are communicatively coupled via one or more buses 575. In addition, the query processing coordinator 115 is communicatively coupled to a plurality of data nodes 130A-N via the network 580. Of course, in embodiments, the data nodes 130A-N may be directly coupled to the query processing coordinator 115, accessible via a local network, integrated into the query processing coordinator 115, etc. Although not included in the illustrated embodiment, in some embodiments, the query processing coordinator 115 is also communicatively coupled with the data placement coordinator 150.

[0082] In the illustrated embodiment, the storage device 520 includes an ontology 560, capability data 565, and a data map 570. Although depicted as residing in the storage device 520, in an embodiment, the ontology 560, capability data 565, and data map 570 may be stored in any suitable location and manner. In an embodiment, as described above, the ontology 560 indicates entities or concepts related to the domain in which the query processing coordinator 115 operates, as well as the potential attributes of each concept / entity and the relationship between the entities / concepts. In at least one embodiment, the ontology 560 does not include instance-level data. Instead, the actual data in the KB is stored in the data nodes 130A-N.

[0083] In an embodiment, as described above, capability data 565 indicates the capabilities of each data node 130A-N, which may include, for example, an indication of which operation(s) each data node 130A-N supports, and any corresponding restrictions on that support. For example, capability data 565 may indicate that data node 130A supports a "join" operation, but only supports integer data types. In an embodiment, as described above, data map 570 indicates the data node(s) 130 storing instance data for each entity / concept defined in ontology 560. For example, data map 570 may indicate that all instances of the "Company" concept are stored in data nodes 130A and 130B, while all instances of the "Document" concept are stored in data nodes 130B and 130N.

[0084] In the illustrated embodiment, the memory 515 includes a query application 535. Although depicted as software resident in the memory 515, in embodiments, the query application 535 may be implemented using hardware, software, or a combination of hardware and software. As shown, the query application 535 includes a parsing component 540, a routing component 545, and a set of (one or more) translators 120. Although depicted as discrete components for conceptual clarity, in embodiments, the operations of the parsing component 540, the routing component 545, and the (one or more) translators 120 may be combined or distributed over any number of components.

[0085] In an embodiment, the parsing component 540 receives OQL queries, parses them to determine their meaning, and generates a logical representation (such as a QGM) of the query. As described above, in some embodiments, the logical representation includes a collection of query blocks, each of which specifies one or more predicates that define the operations to be performed on the input quantifiers of the block. In an embodiment, these quantifiers can be (one or more) basic concepts stored in one or more data nodes 130 and / or data calculated by another query block. In one embodiment, the logical representation enables the query application 535 to understand and reason about the structure of the query and how data flows between blocks. This enables the routing component 545 to effectively route queries (or subqueries from queries).

[0086] In the illustrated embodiment, the routing component 545 receives the logical representation (e.g., the query graph model) from the parsing component 540 and evaluates it to identify one or more data nodes 130 to which the query should be routed. In some embodiments, this includes analyzing the predicates included in each query block in the logical representation. In one embodiment, the routing component 545 marks or annotates each predicate in a given query block based on the (one or more) data nodes 130 that can complete the predicate. In some embodiments, this includes (one or more) data nodes 130 that contain (one or more) relevant quantifiers and / or can perform (one or more) indicated operations. In addition, in at least one embodiment, once each predicate in the block has been processed, the routing component 545 can evaluate the annotation to determine the placement of the block. That is, the routing component 545 determines whether the query block can be placed in a single data node 130. If so, in one embodiment, the routing component 545 assigns the block to the repository. In addition, in an embodiment, if the block cannot be processed by a single data node 130, the routing component 545 splits the block into two or more query blocks and repeats the routing process for each query block.

[0087] In an embodiment, once each query block has been assigned to a single data node 130, the (one or more) translators 120 generate one or more corresponding translated queries. In an embodiment, each translator 120 corresponds to a specific data node 130 architecture and is configured to translate the OQL query (or subquery) into an appropriate query for the corresponding data storage device architecture. For example, in one such embodiment, the system may include a first translator 120 that translates OQL into a query for a relational repository (e.g., SQL), a second translator 120 that translates OQL into a JSON query, a third translator 120 that translates into a query for a document search, and a fourth translator 120 that translates OQL into a graph query.

[0088] In an embodiment, based on the data node 130 that will execute the query or subquery, the query (or subquery) is routed to the appropriate translator 120. In some embodiments, each query (or subquery) is then sent directly to that data node 130 (e.g., via an application programming interface or API). In some embodiments, if completing the query will require moving data between data nodes 130 (e.g., to connect data from two or more repositories), the query application 535 can determine the data flow based on the generated logical representation and send the query appropriately. That is, the query application 535 uses the QGM to determine which (which) repository will need to receive data from other (which) repositories and send the required queries to these connected repositories. For example, if data node 130A is to receive data from data node 130N and complete one or more operations on it, the query application 535 can send a first query to data node 130A to retrieve data from the node, as well as a second query configured for data node 130N. Then, data node 130A can query data node 130N using the provided query itself.

[0089] Figure 6 6 is a flow chart illustrating a method 600 for processing and routing ontology queries according to one embodiment disclosed herein. The method 600 begins at box 605, where the query processing coordinator 115 receives an ontology query for execution against a knowledge base. In an embodiment, the ontology query is expressed relative to a predefined domain ontology, where the ontology specifies concepts, properties, and relationships related to the domain. In box 610, the query processing coordinator 115 parses the received query and generates a logical representation thereof. In some embodiments, the query processing coordinator 115 generates a query graph schema representation of the query. In an embodiment, the logical representation includes one or more query blocks representing the basic concept(s) implied by the query, the operation(s) to be performed on the data, and the data flow.

[0090] The method 600 then continues to box 615, where the query processing coordinator 115 maps each query block in the logical representation to one or more data repositories based on predefined concept mappings and / or data storage capabilities. For example, in one embodiment, the query processing coordinator 115 determines what concepts are relevant for each query block (e.g., what data the query block will access). Using the concept mapping, the query processing coordinator 115 can then identify (one or more) data repositories that can provide the required data. In addition, in one embodiment, the query processing coordinator 115 determines what operations will be required for each query block. Using the predefined storage capabilities, the query processing coordinator 115 can then identify which (some) data nodes are capable of performing (one or more) required operations. The query processing coordinator 115 then maps (one or more) query blocks to (one or more) data nodes, striving to minimize data movement between repositories.

[0091] At box 620, the query processing coordinator 115 selects one of the mapped query blocks. The method 600 then proceeds to box 625, where the query processing coordinator 115 generates a corresponding translated query for the query block based on the mapped data repository. In one embodiment, this includes identifying or generating subqueries from the received query to execute the selected query block. The query processing coordinator 115 then determines the configuration of the data node that will execute the query block. That is, in one embodiment, the query processing coordinator 115 determines the language and / or format of the query that the node is configured to process. In another embodiment, the query processing coordinator 115 identifies a translator corresponding to the mapped data repository. The query processing coordinator 115 then uses the translator to generate an appropriate query for the repository.

[0092] The method 600 then continues to box 630, where the query processing coordinator 115 determines whether there is at least one additional query block that has not been processed. If so, the method 600 returns to box 620. Otherwise, the method 600 continues to box 635, where the query processing coordinator 115 sends one or more translated queries to one or more data nodes. In one embodiment, the query processing coordinator 115 sends each query to a corresponding node. In another embodiment, the query processing coordinator 115 sends queries based on the data flow in the logical representation. For example, the query processing coordinator 115 may send multiple queries to a single node so that the node can forward the query to (one or more) appropriate storage to retrieve data. The node can then act as a mediator to finalize the results.

[0093] Advantageously, by employing one of the data nodes as a mediator, the query processing coordinator 115 can reduce its computational overhead and extend the scalability of the system. At block 640, the query processing coordinator 115 receives the finalized query results from the data repository acting as a mediator (or from the only data repository to which the query is transmitted if the received query can be executed in a single backend resource). The query processing coordinator 115 then returns the results to the requesting entity.

[0094] Figure 7 700 is a flow chart illustrating a method 700 for processing ontology queries to efficiently route query blocks according to an embodiment disclosed herein. In some embodiments, method 700 provides additional details for routing query blocks. Method 700 begins at box 705, where the query processing coordinator 115 generates one or more query blocks to represent the received query, as described above. In one embodiment, each query block specifies one or more quantifiers indicating the input data of the block. These quantifiers can correspond to basic concepts (e.g., instance data in a knowledge base) and / or intermediate computational data (e.g., data retrieved from a repository and processed and / or transformed in some manner). In addition, in an embodiment, each query block specifies the operation to be performed on the quantifier.

[0095] At box 710, the query processing coordinator 115 selects one of the generated blocks. In addition, at box 715, the query processing coordinator 115 identifies a set of data nodes that can satisfy the (one or more) operations defined by the selected block. In one embodiment, the query processing coordinator 115 does this by accessing a predefined capability definition that defines a set of (one or more) operations that each data repository can complete and any corresponding (one or more) restrictions on the capability. The method 700 then proceeds to box 720, where the query processing coordinator 115 identifies a set of nodes that can satisfy the (one or more) quantifiers listed in the query block. That is, for each quantifier corresponding to a base concept, the query processing coordinator 115 identifies the node that stores the base concept.

[0096] In one embodiment, for each quantifier computed by another query block, the query processing coordinator 115 identifies the data node(s) that have been assigned to the query block. Recall that, in an embodiment, the query processing coordinator 115 maps query blocks to nodes by walking up from the bottom of the QGM. Thus, if a first block depends on data computed in a second block, the second block will necessarily have been evaluated and assigned one or more data nodes before the query processing coordinator 115 begins evaluating the first query block.

[0097] Method 700 then continues to block 725, where the query processing coordinator 115 determines whether there is at least one data node that can provide both the quantifier and the execution operation. In one embodiment, this includes determining whether there is overlap between the identified sets. For example, if the quantifiers are all basic concepts, the overlap between the two sets indicates a set of nodes that can perform all required operations and store all concepts. If the quantifier includes items calculated by other query blocks, the overlapping set indicates nodes that can perform operations and will (potentially) execute (one or more) other blocks.

[0098] In an embodiment, if there is an overlap between the sets, the method 700 continues to block 730, where the query processing coordinator 115 annotates the selected query block with the identified nodes in the overlap. The annotation indicates that the specified node has the potential to execute the query block, but not the final assignment of the block to the node(s). In an embodiment, if any query block has multiple nodes included in the annotation, the query processing coordinator 115 performs a cost analysis to evaluate each alternative in an attempt to minimize data transfers between repositories, as discussed in more detail below. The method 700 then proceeds to block 740.

[0099] In the illustrated embodiment, if there is no overlap between the sets, the query processing coordinator 115 determines that there is no single data node that may complete the query block. In some embodiments, if all operations can be performed by a single repository, the query processing coordinator 115 annotates the block with the repository that may perform the operations of the block. Other repository(s) may then be assigned to provide quantifiers. In the illustrated embodiment, the method 700 then continues to box 735, where the query processing coordinator 115 splits the selected query block into two or more blocks to create a query block that can be executed entirely within a single node. In one embodiment, splitting the selected query block includes identifying a subset of the quantifier(s) and / or the operation(s) that can be executed by a single storage. For example, the query processing coordinator 115 may identify a subset of specified predicates that share annotations (e.g., predicates that can be executed by the same repository). The query block may then be split by separating the predicates into corresponding boxes based on the subsets to which the predicates belong. For example, if a query block requires traditional relational database operations as well as fuzzy matching operations, the query processing coordinator 115 may determine that the fuzzy matching (or) should be split into separate boxes. This may enable the query processing coordinator 115 to assign each split block to a single data repository (e.g., the relational operations may be assigned to the relational repository, and the fuzzy matching operations may be assigned to a different backend capable of executing it).

[0100] In the illustrated embodiment, these newly generated query blocks are placed in a queue to be evaluated, just like the existing blocks. The method 700 then continues to block 740. At block 740, the query processing coordinator 115 determines whether there is at least one additional query block that has not been evaluated. If so, the method 700 returns to block 710 to iterate through the blocks. That is, the query processing coordinator 115 continues to iterate through the query blocks, splitting boxes if necessary, until all query blocks are annotated with at least one data node.

[0101] If all query blocks have been annotated, the method 700 continues to block 745, where the query processing coordinator 115 executes the query. In some embodiments, this includes evaluating alternative combinations of assignments to minimize data transfers, as discussed in more detail below. In some embodiments, if the annotation of one or more of the query blocks indicates a single data node, the query processing coordinator 115 simply assigns the block to the indicated node because there are no alternatives that can be considered.

[0102] Figure 8 800 is a flow chart illustrating a method 800 for evaluating potential query routing plans in order to efficiently route ontology queries according to one embodiment disclosed herein. In the illustrated embodiment, the method 800 begins after query blocks have been annotated with their potential assignments (e.g., with nodes that can execute the entire block and provide all required quantifiers). The method 800 begins at box 805, where the query processing coordinator 115 selects one of the possible combinations of node assignments. That is, in an embodiment, the query processing coordinator 115 generates all possible combinations of repositories for blocks based on the corresponding annotations. To this end, the query processing coordinator 115 can iteratively select different options for each block until all possible choices have been generated.

[0103] For example, assume that the first block is labeled "Node A" and "Node B", and the second block is labeled "Node A". In an embodiment, the query processing coordinator 115 will determine that the possible routing plan includes assigning the first block and the second block to "Node A", or assigning the first block to "Node B" and assigning the second block to "Node A". At block 805, the query processing coordinator 115 selects one of the identified combinations to evaluate. The method 800 then continues to block 810.

[0104] In box 810, the query processing coordinator 115 identifies the (one or more) data movements that will be required under the selected plan. In various embodiments, the movement is determined based on the repository selection and / or query graph model in the plan. Continuing with the above example, if the query processing coordinator 115 can determine that the plan to assign two blocks to "Node A" will not require movement because both query blocks can be grouped into a single node. That is, because the blocks are results in the model (e.g., directly connected / separated by a single hop, with no other blocks in between) and are assigned the same node, they are grouped together and no data movement is required. In contrast, assigning the first block to "Node B" will require data to be transferred between "Node A" and "Node B" because the QGM indicates that data flows from the second block to the first block, and the blocks are assigned to different repositories.

[0105] The method 800 then continues to block 815, where the query processing coordinator 115 selects one of the identified data transfers required for the selected query plan. At block 820, the query processing coordinator 115 determines the computational cost of the selected move. In an embodiment, the determination is made based on a predefined cost model. In some embodiments, the model may indicate the computational cost of transferring data from a first node to a second node for each ordered pair of data nodes. The cost may include, for example, the delay introduced by actually transferring the data and / or transforming the data as needed to allow the destination node to operate on it, the processing time and / or memory requirements of the transfer / transformation, etc.

[0106] In box 825, the query processing coordinator 115 determines whether the selected combination requires any additional data transfer. If so, the method 800 returns to box 815. Otherwise, the method 800 continues to box 830, where the query processing coordinator 115 calculates the total cost of the selected plan by aggregating the individual costs of each move. In box 835, the query processing coordinator 115 determines whether there is at least one alternative plan that has not yet been evaluated. If so, the method 800 returns to box 805. Otherwise, the method 800 proceeds to box 840, where the query processing coordinator 115 ranks the plans based on the aggregated costs of the plans. In an embodiment, the query processing coordinator 115 selects the plan with the lowest determined cost. The query processing coordinator 115 then executes the plan by translating and routing subqueries based on the node allocation of the minimum cost plan.

[0107] Fig. 9is a flow chart illustrating a method 900 for routing ontology queries according to one embodiment disclosed herein. The method 900 begins at box 905, where the query processing coordinator 115 receives an ontology query. At box 910, the query processing coordinator 115 generates one or more query blocks based on the ontology query, each query block indicating one or more operations and one or more quantifiers representing data flow between query blocks. The method 900 then proceeds to box 915, where the query processing coordinator 115 identifies at least one data node for each of the one or more query blocks based on the one or more quantifiers and the one or more operations. In addition, at box 920, the query processing coordinator 115 selects one or more of the identified data nodes based on a predefined cost criterion. In addition, at box 925, the query processing coordinator 115 then sends one or more sub-queries to the selected one or more data nodes.

[0108] Fig.10 A workflow 1000 for evaluating workloads and storing knowledge base data according to an embodiment disclosed herein is shown. As described above, enterprise applications typically need to support different query types depending on their query workloads. In order to support these different query types, embodiments of the present disclosure utilize multiple backend repositories, such as relational databases, document repositories, graph repositories, and the like. In embodiments of the present disclosure, the system has the ability to move, store, and index knowledge base data in any backend that provides the required capabilities for the supported query types. With this flexibility in organizing data, initial data placement across multiple backend repositories can play a key role in efficient query execution.

[0109] To achieve this efficient execution, some embodiments of the present disclosure provide an offline data preparation and loading phase, which includes intelligent data placement. Typically, data ingestion, placement, and loading into multiple back-end repositories involve a series of operations. Initially, data for a domain-specific knowledge base is ingested from a variety of different sources including structured, semi-structured, and unstructured data. In some embodiments, the first stage of data ingestion from these data sources is a data enrichment / curation process, which includes information extraction, entity resolution, data integration, and transformation. The output data generated by this step that conforms to the domain ontology is then fed to the data placement coordinator 150 to generate appropriate data placement for multiple back-end data repositories. In some embodiments, the data has already been managed and may have been used in response to queries before using the data placement coordinator 150. Once satisfactory data placement is determined, the data loading module places the instance data in the appropriate data repository according to the data placement plan.

[0110] In some embodiments, data movement can be avoided during query execution by replicating the entire set of data across all data nodes. However, this solution results in huge replication and significantly increases storage space overhead. In addition, in many embodiments, not all repositories provide all the necessary capabilities required for queries, and even full replication cannot completely eliminate data movement. In order to minimize unnecessary storage costs and data movement, some embodiments of the present disclosure provide a capability-based data placement algorithm that assigns data to data repositories while taking into account the expected workload and the capabilities of the backend data repositories (e.g., in terms of the operations they can perform on the stored data).

[0111] In an embodiment, the data placement coordinator 150 reasons about data placement at the level of query operations on concepts of domain ontologies representing patterns of data stored in the knowledge base. In some embodiments, the coordinator identifies distinct and potentially overlapping subsets of the ontology based on a given workload for the knowledge base and the capabilities of the underlying repositories, and outputs a mapping between the identified subsets of data and the target data repositories where the data should be stored.

[0112] In the illustrated embodiment, the workflow 1000 begins when a collection of OQL queries 1005 is received. In some embodiments, the OQL queries 1005 are previously submitted queries against a knowledge base. For example, in some embodiments, the knowledge base is a pre-existing corpus of data that users and applications can query and explore. In such embodiments, the OQL queries 1005 may correspond to queries previously submitted by users, applications, and other entities when interacting with the knowledge base.

[0113] In an embodiment, OQL queries 1005 generally represent average, typical, expected, and / or historical workloads of a knowledge base. In other words, OQL queries 1005 generally represent queries that a knowledge base receives (or expects to receive) during runtime operation, and can be used to identify sets of concepts that are often queried together, operations that are typically performed on each concept, etc. As shown, these representative OQL queries 1005 are provided to an OQL query analyzer 1010, which analyzes and evaluates them to generate a summarized workload 1015.

[0114] In an embodiment, the OQL query analyzer 1010 expresses the OQL query 1005 as a collection of concepts and relationships with corresponding operations performed on them. Based on this, the analyzer generates a summarized workload 1015. In other words, the summarized workload 1015 reflects a summary of the provided OQL query 1005, which enables a deeper analysis to identify patterns in the data. In one embodiment, the summarized workload 1015 includes a collection of summarized queries generated based on the OQL query 1005. It is worth noting that in some embodiments, the summarized queries are not complete queries that can be executed against the knowledge base. Instead, the summarized queries indicate clusters of concepts that may be queried together during runtime and a collection of operations that may be applied to each cluster of concepts.

[0115] In some embodiments, the OQL query analyzer 1010 takes as input a collection 1005 of OQL queries expressed against a domain ontology, and for each query, creates two collections. In one such embodiment, the first is a collection of concepts accessed by the query, and the second is a collection of operations (e.g., joins, aggregations, etc.) performed by the query on those concepts. In at least one embodiment, to generate a summarized workload 1015 representation of a given OQL query 1005, the OQL query analyzer 1010 groups queries that access the same collection of concepts into a group, and then creates a collection of associated operations that combine each query in the group. This will be discussed in more detail below.

[0116] In the illustrated embodiment, the summarized workload 1015 is then provided to a hypergraph modeler 1020, which evaluates the provided summarized query to generate a hypergraph 1025. In an embodiment, the hypergraph 1025 is a graph comprising a set of vertices and a set of edges (also referred to as hyperedges), wherein each edge may be connected to any number of vertices. In some embodiments, each vertex in the hypergraph 1025 corresponds to a concept from a domain ontology, and each hyperedge corresponds to a summarized query from the summarized workload 1015. For example, each hyperedge spans a set of concepts indicated by a corresponding summarized query. In one embodiment, each hyperedge is also annotated or labeled with an indication of a set of operations indicated by a corresponding summarized query.

[0117] As shown, the hypergraph 1025 is evaluated by the data placement component 1030 together with the ontology schema 110 and the node capabilities 155. In one embodiment, the data placement component 1030 groups the concepts and relationships in the hypergraph 1025 into potentially overlapping subsets based on query operations. Then, the data corresponding to these subsets can be placed on various backend data nodes based on the operations they support. In some embodiments, data placement decisions are made at the granularity of the identified ontology subsets, and all data for any ontology concept is placed. In other words, the placement decision 165 does not horizontally split concepts between different repositories. For example, if the data placement component 1030 places the "company" concept in the first data node, all instance-level data corresponding to the "company" concept is placed in the first data node.

[0118] In some embodiments, the ontology-based data placement component 1030 follows a two-step approach for data placement. First, the data placement component 1030 runs a graph analysis algorithm on the hypergraph 1025 representing the summarized workload 1015 to group concepts in the domain ontology schema 110 based on the similarity of the operations performed on these concepts. Next, the data corresponding to these identified groups or subsets are mapped to the underlying data repositories according to their respective capabilities while minimizing the amount of replication required. In an embodiment, the resulting capability-based data placement minimizes data movement (and data transformation) for a given workload at query processing time, thereby greatly enhancing the efficiency of query processing in a multi-repository environment. As shown, the final output of the data placement component 1030 is a concept-to-store mapping (e.g., placement decision 165) that maps the ontology concepts to appropriate data nodes.

[0119] Although not depicted in the illustrated workflow 1000, in an embodiment, the system then utilizes placement decisions 165 to store the data in various backend resources. In one embodiment, the system does this by invoking an extract, transform, and load (ETL) service that performs any transformations or conversions necessary to allow the data to be stored in the relevant data nodes 130.

[0120] In some embodiments, workflow 1000 is used to periodically re-evaluate and refine data placement in order to maintain the efficiency of the system. For example, during off-peak time (e.g., during non-business hours), the system can invoke workflow 1000 based on a set 1005 of updated OQL queries (e.g., including queries received after the last data placement decision) to determine whether placement should be updated to reflect evolving workloads. This can improve the efficacy of the system by preventing data placement from becoming stale.

[0121] Fig.11 An example hypergraph 1100 for evaluating workloads and storing knowledge base data is depicted according to one embodiment disclosed herein. In the depicted embodiment, each concept 1105A-G is depicted as an ellipse, and each hyperedge 1110A-D is depicted as a dashed line surrounding its corresponding concept 1105. For example, hyperedge 1110A connects concepts 1105A ("Company"), 1105B ("PublicCompany"), 1105C ("PublicMetric"), and 1105D ("PublicMetricData"). As shown, each hyperedge 1110 can include both unconnected subsets of concepts 1105 (e.g., hyperedges 1110A and 1110C do not overlap) and overlapping subsets (e.g., hyperedges 1110A and 1110B overlap with respect to "Company" concept 1105A).

[0122] In addition, as shown, each hyperedge 1110 is labeled with the relevant operation 1115A-C of the edge. As described above, in one embodiment, the label indicates a set of operations that can or have been applied to the concepts 1105 connected by the hyperedge 1110. For example, in the depicted example, the hyperedge 1110A is associated with operations 1115A ("connect"), 1115B ("aggregation"), and 1115C ("fuzzy" matching). It is worth noting that each operation 1115 can be associated with any number of hyperedges 1110 and / or concepts 1105. In the described embodiment, the "connect" operation 1115A is associated with hyperedges 1110A and 1110B, the "aggregation" operation 1115B is associated with hyperedges 1110A and 1110D, and the "fuzzy" operation 1115C is associated with hyperedges 1110A, 1110D, and 1110C.

[0123] Fig.121 is a block diagram illustrating a data placement coordinator 150 configured to evaluate workloads and store data according to one embodiment disclosed herein. Although depicted as a physical device, in embodiments, the data placement coordinator 150 may be implemented using (one or more) virtual devices and / or across multiple devices (e.g., in a cloud environment). As shown, the data placement coordinator 150 includes a processor 1210, a memory 1215, a storage device 1220, a network interface 1225, and one or more I / O interfaces 1230. In the illustrated embodiment, the processor 1210 retrieves and executes programming instructions stored in the memory 1215, and stores and retrieves application data residing in the storage device 1120. The processor 1210 generally represents a single CPU and / or GPU, multiple CPUs and / or GPUs, a single CPU and / or GPU with multiple processing cores, etc. The memory 1215 is generally included to represent a random access memory. Storage device 1220 can be any combination of disk drives, flash-based storage devices, etc., and can include fixed and / or removable storage devices, such as fixed disk drives, removable memory cards, cache, optical storage devices, network attached storage devices (NAS), or storage area networks (SAN).

[0124] In some embodiments, input and output devices (such as keyboards, monitors, etc.) are coupled via the I / O interface(s) 1230. Additionally, via the network interface 1225, the data placement coordinator 150 may be communicatively coupled with one or more other devices and components (e.g., via a network 1280, which may include the Internet, local network(s), etc.). As shown, the processor 1210, memory 1215, storage 1220, network interface(s) 1225, and I / O interface(s) 1230 are communicatively coupled via one or more buses 1275. Additionally, the data placement coordinator 150 is communicatively coupled to the data nodes 130A-N via the network 1280. Of course, in embodiments, the data nodes 130A-N may be directly coupled to the data placement coordinator 150, accessible via a local network, integrated into the data placement coordinator 150, etc. Although not included in the illustrated embodiment, in some embodiments, the data placement coordinator 150 is also communicatively coupled with the query processing coordinator 115.

[0125] In the illustrated embodiment, storage 1220 includes a copy of domain ontology 1260, capability data 1265, and data mapping 1270. Although depicted as residing in storage 1220, in an embodiment, ontology 1260, capability data 1265, and data mapping 1270 may be stored in any suitable location and manner. In an embodiment, ontology 1260, capability data 1265, and data mapping 1270 correspond to ontology 560, capability data 565, and data mapping 570 discussed above with reference to query processing coordinator 115.

[0126] For example, as described above, ontology 1260 may indicate entities or concepts related to a domain, as well as potential attributes of each concept / entity and relationships between entities / concepts, without including instance-level data. Similarly, as described above, capability data 1265 indicates the capabilities of each data node 130A-N, which may include, for example, an indication of which operation(s) each data node 130A-N supports, and any corresponding limitations(s) on that support.

[0127] In an embodiment, data map 1270 is generated by placement application 1235 and indicates the data node(s) 130 where instance data is stored for each entity / concept defined in ontology 1260. For example, data map 1270 may indicate that all instances of the "Company" concept are stored in data nodes 130A and 130B, while all instances of the "Document" concept are stored in data nodes 130B and 130N.

[0128] In the illustrated embodiment, the memory 1215 includes a placement application 1235. Although depicted as software residing in the memory 1215, in embodiments, the placement application 1235 may be implemented using hardware, software, or a combination of hardware and software. As shown, the placement application 1235 includes a generalization component 1240, a supergraph component 1245, and a placement component 1250. Although depicted as discrete components for conceptual clarity, in various embodiments, the operations of the generalization component 1240, the supergraph component 1245, and the placement component 1250 may be combined or distributed across any number of components.

[0129] In an embodiment, the summary component 1240 receives a previous OQL query (or sample query) that reflects a previous, current, expected, typical, average, anticipated, and / or general workload on the system. The summary component 1240 then evaluates the provided query to summarize the workload in the form of a summarized query. In an embodiment, the summarized query generally indicates a set of concepts that are queried together and operations performed on the concepts, but does not correspond to an actual query that can be executed. The functionality of the summary component 1240 is described below with reference to Fig.13 Discuss in more detail.

[0130] In an embodiment, the hypergraph component 1245 receives the summarized workload information from the summarization component 1240 and parses it to generate a hypergraph representing the query workload. As described above, in some embodiments, the hypergraph includes a vertex for each concept in the ontology 1260. In addition, in an embodiment, the concept vertices are connected by hyperedges representing the summarized query workload. The operation of the hypergraph component 1245 will be described below with reference to Fig.14 Discuss in more detail.

[0131] In one embodiment, the placement component 1250 evaluates the hypergraph to generate data placement decisions (e.g., data map 1270). These decisions can then be used to route the data to the appropriate nodes. For example, in one embodiment, the system iterates through the knowledge base to place the data. For each data item, the system can determine its corresponding concept, find the corresponding data node 130, and store the data in the indicated node(s). In some embodiments, the system also performs any transformations or conversions appropriate for the destination repository. The functionality of the placement component 1250 is described below with reference to Fig.15 and Fig.16 Discuss in more detail.

[0132] Fig.13 13 is a flow chart illustrating a method 1300 for evaluating and summarizing a query workload to inform data placement decisions according to one embodiment disclosed herein. The method 1300 begins at block 1305, where the data placement coordinator 150 receives one or more previous queries representing a knowledge base workload. At block 1310, the data placement coordinator 150 selects one of the received queries for evaluation. The method 1300 then proceeds to block 1315, where the data placement coordinator 150 identifies the concept(s) involved by the selected query. Similarly, at block 1320, the data placement coordinator 150 identifies the operation(s) invoked by the query.

[0133] In one embodiment, the data placement coordinator 150 associates the determined set of concepts specified by the query with the set of (one or more) operations invoked by the query. In some embodiments, the operation-concept pairings are determined on a granular concept / operation level basis. Notably, in some other embodiments, the data placement coordinator 150 does not determine which operations are linked to which concepts, but rather links the set of operations as a whole to the set of concepts. That is, the data placement coordinator 150 ignores the actual execution details / expected results of the query and defines the set of related concepts and operations based on the data / operations involved rather than the specific application of the transformation.

[0134] The method 1300 then proceeds to box 1325, where the data placement coordinator 150 determines whether there is at least one additional query to be evaluated. If so, the method 1300 returns to box 1310. Otherwise, the method 1300 proceeds to box 1330. In group 1330, the data placement coordinator 150 groups the received queries based on the set of corresponding (one or more) concepts for each. In one embodiment, this includes grouping or clustering all queries that specify a set of matching concepts. For example, the data placement coordinator 150 may group all queries that specify a "Company" concept and a "PublicMetric" concept into a single group. It is worth noting that in such an embodiment, queries that specify only the "Company" concept or the "PublicMetric" concept will be placed in different groups. Similarly, a query that specifies "Company" and "PublicMetric" but also includes a "Document" concept will not be placed in the first group.

[0135] That is, in an embodiment, the data placement coordinator 150 groups queries that specify a set of concepts that are exactly matched. Queries that specify additional concepts or fewer concepts are placed in other groups. In one embodiment, the set of concepts corresponding to each respective grouping is used to form a respective summarized query. That is, in the illustrated embodiment, the data placement coordinator 150 aggregates received prior queries by grouping queries that access the same concept into clusters, regardless of the operations performed by each query.

[0136] The method 1300 then proceeds to block 1335, where the data placement coordinator 150 selects one of the defined query groups. At block 1340, the data placement coordinator 150 selects one of the queries associated with the selected group. Additionally, at block 1345, the data placement coordinator 150 associates the corresponding operation specified by the selected query with the summarized query representing the selected query group. In one embodiment, if the operation is already reflected in the summarized query, the data placement coordinator 150 does not add it again. That is, in an embodiment, the set of operations associated with the summarized query is a binary value indicating the presence or absence of a given operation.

[0137] In embodiments, the operations associated with a summarized query may be at any level of granularity. For example, in some embodiments, the data placement coordinator 150 defines a summarized query at an operation level (e.g., "join") without regard to the specific type of operation (e.g., inner join), restrictions on the operation, or details of the operation (e.g., the type of data, such as string concatenation or integer concatenation). In other embodiments, these details are included in the summarized query description to provide richer details for subsequent processing.

[0138] The method 1300 then continues to block 1350, where the data placement coordinator 150 determines whether there is at least one additional query in the selected group. If so, the method 1300 returns to block 1340. If not, at block 1355, the data placement coordinator 150 determines whether there is at least one additional query group that has not been evaluated. If so, the method 1300 returns to block 1335. Otherwise, the method 1300 continues to block 1360. At block 1360, the data placement coordinator 150 stores the summarized workload so that it can be used and evaluated to generate the hypergraph.

[0139] Fig.14 14 is a flow chart illustrating a method 1400 for modeling an ontology workload to inform data placement decisions according to one embodiment disclosed herein. The method 1400 begins at block 1405, where the data placement coordinator 150 selects one of the concepts of the ontology. At block 1410, the data placement coordinator 150 generates a hypergraph vertex for the selected concept. The method 1400 then continues to block 1415, where the data placement coordinator 150 determines whether there are any additional concepts remaining in the ontology that do not yet have vertices in the hypergraph. If so, the method 1400 returns to block 1405. Otherwise, the method 1400 proceeds to block 1420.

[0140] At box 1420, the data placement coordinator 150 selects one of the summarized queries. As discussed above, in an embodiment, each summarized query indicates a set of concepts and a corresponding set of operations that have been applied to the concepts in a previous workload. At box 1425, the data placement coordinator 150 generates a hyperedge linking each of the (one or more) concepts indicated by the selected summarized query. The method 1400 then proceeds to box 1430, where the data placement coordinator 150 marks the newly generated hyperedge with an indication of the operations indicated by the selected summarized query. In this manner, the data placement coordinator 150 can subsequently evaluate the hypergraph to identify relationships and patterns of use of the knowledge base.

[0141] The method 1400 then continues to block 1435, where the data placement coordinator 150 determines whether there is at least one additional generalized query to be evaluated and merged into the hypergraph. If so, the method 1400 returns to block 1420. Otherwise, the method 1400 continues to block 1440, where the data placement coordinator 150 stores the generated hypergraph for subsequent use.

[0142] Fig.15 and Fig.16 A flow chart illustrating a method for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein is depicted. Fig.15The method 1500 discussed illustrates one embodiment of an operational-based clustering technique for generating a concept map, and the following references Fig.16 The method 1600 discussed illustrates one embodiment of a minimum coverage technique.

[0143] In one embodiment, an operation-based clustering technique groups concepts based on the operations to which they are subjected. In one such embodiment, for each operation description in a hypergraph, the data placement coordinator 150 creates a corresponding cluster. The data placement coordinator 150 then iterates over the group of operation descriptions associated with each hyperedge, and for each such operation description, all concepts spanned by the hyperedge are assigned to the cluster of the corresponding operation. In an embodiment, once the concepts are clustered together, the data placement coordinator 150 assigns each concept cluster to a set of data nodes so that each node has a capability description that matches the operation description of the cluster (e.g., capable of performing the corresponding operation of the cluster). Finally, in an operation-based system, the data placement coordinator 150 generates a mapping that maps each concept in each cluster to a corresponding set of identified data repositories.

[0144] As an example of one embodiment of an operation-based clustering technique, consider Fig.11 1100 is provided in FIG. 1111. Initially, the data placement coordinator 150 generates a hypergraph 1100 for each operation 1115A-C (e.g., C Join , C Agg and C Fuzzy ) to generate clusters. The system then determines for each hyperedge the set of operations associated with it. For each indicated operation, the system assigns the concept specified by the hyperedge 1110 to the corresponding cluster. Continuing with the example above, C Join The concepts 1105A, 1105B, 110C, and 1105D from the hyperedge 1110A, and the concepts 1105E and 1105F from the hyperedge 1110B will be included. Agg will contain concepts 1105A, 1105B, 1105C, and 1105D from hyperedge 1110A, and concept 1105G from hyperedge 1110D. Fuzzy Concepts 1105A, 1105B, 1105C, and 1105D from hyperedge 1110A will be included, as well as concept 1105H from hyperedge 1110C.

[0145] In many embodiments, these operation-based clusters have significant overlap. For example, note that concepts 1105A, 1105B, 1105C, and 1105D are included in each cluster. To finalize the mapping, in one embodiment, the data placement coordinator 150 identifies all data nodes 130 for each cluster that can perform the corresponding operation. The data placement coordinator 150 then maps all concepts 1105 in the cluster to all identified data nodes 130. In some embodiments, although the operation-based technique can minimize or reduce data movement during query processing by placing data in all storages that support the corresponding operation, it does introduce some replication overhead because clusters of the same concept can be placed at multiple repositories if they have the ability to satisfy the operation of the cluster.

[0146] In some embodiments, to further reduce replication overhead, an embodiment of a minimum cover technique is utilized. In one embodiment, the minimum cover embodiment improves upon the operation-based technique by further minimizing the amount of data replication while still minimizing data movement during query processing. In an embodiment, the technique utilizes a minimum set cover algorithm to find the minimum number of data repositories required to support the complete set of operations required for each hyperedge in the query workload hypergraph. In some embodiments, the minimum cover technique minimizes the span of each hyperedge over the set of data repositories that satisfy the set of operations required for the hyperedge.

[0147] In one example embodiment, the minimum cover technique includes, for each hyperedge in the hypergraph, the data placement coordinator 150 finds the minimum number of data nodes that cover all indicated operations. For example, if all operations can be completed by a single data node 130A, the minimum set includes only that node. If data node 130A cannot complete one or more operations, one or more other data nodes 130B-N are added to the minimum set until all operations are satisfied. Once the minimum set is determined for the hyperedge, each concept in the hyperedge is mapped to each node in the corresponding minimum set.

[0148] As an example of one embodiment of the minimum covering clustering technique, consider Fig.11, assume that the data nodes 130 include a first data node 130A configured to support operations 1115A and 1115B, a second data node 130B configured to support operations 1115B and 1115C, and a third data node 130C that only supports operation 1115B. For hyperedge 1110A, the data placement coordinator 150 can determine that no node can support all three indicated operations individually, but the set of data nodes 130A and 130B can support all three indicated operations. Because these two nodes can support the entire hyperedge, there is no need to add data node 130C to the set.

[0149] Similarly, hyperedge 1110B will be assigned to data node 130A (the only node configured to provide join operations), while hyperedge 1110C will be assigned to data node 130B (the only node configured to provide fuzzy matching). Finally, hyperedge 1110D can be assigned to data nodes 130A and 130B, or data nodes 130B and 130C. In some embodiments, the data placement coordinator 150 selects between these otherwise equivalent alternatives based on other criteria, such as the latency or computing resources of each, predefined preferences, etc. The data placement coordinator 150 then maps the concepts 1105 of each hyperedge 1110 to the assigned data node(s) 130.

[0150] Fig.15 1 is a flow chart illustrating an operator-based method 1500 for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein. The method 1500 begins at block 1505, where the data placement coordinator 150 selects one of the operations indicated by the hypergraph. At block 1510, the data placement coordinator 150 generates a cluster for the selected operation. The method 1500 then proceeds to block 1515, where the data placement coordinator 150 determines whether there is at least one additional operation reflected in the hypergraph that does not already have a cluster associated with it. If so, the method 1500 returns to block 1505. Otherwise, the method 1500 continues to block 1520.

[0151] At block 1520, the data placement coordinator 150 selects one of the hyperedges in the hypergraph for analysis. At block 1525, the data placement coordinator 150 identifies the concept(s) and operation(s) associated with the selected edge. The method 1500 then continues to block 1530, where the data placement coordinator 150 selects one of the indicated operations. Additionally, at block 1535, the data placement coordinator 150 identifies the corresponding cluster for the selected operation and adds all concepts represented by the selected hyperedge to the cluster. The method 1500 continues to block 1540, where the data placement coordinator 150 determines whether the selected edge indicates at least one additional operation to be processed. If so, the method 1500 returns to block 1530.

[0152] If no additional operations are associated with the selected hyperedge, the method 1500 continues to box 1545, where the data placement coordinator 150 determines whether the hypergraph includes at least one additional edge that has not yet been evaluated. If so, the method 1500 returns to box 1520. Otherwise, the method 1500 proceeds to box 1550, where the data placement coordinator 150 maps the (one or more) operation clusters to the (one or more) corresponding data nodes configured to process each operation. For example, in one embodiment, the data placement coordinator 150 identifies a set of data nodes that can perform the corresponding operation for each cluster. In an embodiment, the data placement coordinator 150 then maps each concept in the cluster to the set of (one or more) identified data nodes. The data placement coordinator 150 can use these mappings to distribute data in the knowledge base among various data nodes.

[0153] Fig.16 1 is a flow chart illustrating a minimum cover-based method 1600 for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein. The method 1600 begins at block 1605, where the data placement coordinator 150 selects one of the hyperedges in the hypergraph. At block 1610, the data placement coordinator 150 identifies the corresponding concepts and operations associated with the selected edge. Additionally, at block 1615, the data placement coordinator 150 determines a minimum set of data nodes that can collectively satisfy all indicated operations.

[0154] In one embodiment, the data placement coordinator 150 does this by iteratively evaluating a combination of data nodes to determine whether the combination satisfies the indicated operation. That is, whether each operation indicated by the selected edge can be performed by at least one data node in the combination. If not, the combination can be discarded (or another node can be added). Subsequently, the data placement coordinator 150 can identify the combination (one or more) with the smallest number of data nodes, because these combinations will likely result in the smallest data movement during runtime. In one embodiment, if two or more combinations are equally small, the data placement coordinator 150 can select the best combination using predefined criteria or preferences. For example, a predefined rule may indicate that a combination including at least one relational data repository is preferred over a combination without a relational data repository. As another example, a rule may indicate the weight or priority of (one or more) specific types of repositories and / or (one or more) specific data nodes 130. In such an embodiment, the data placement coordinator 150 can aggregate these weights for each combination to determine which repositories to use. The selected edge is then marked with an indication of the determined set of data nodes.

[0155] Once the minimum set of data nodes is determined, the method 1600 proceeds to box 1620, where the data placement coordinator 150 determines whether there is at least one additional hyperedge in the hypergraph. If so, the method 1600 returns to box 1605. Otherwise, the method 1600 proceeds to box 1625. In box 1625, the data placement coordinator 150 selects one of the available data nodes in the system. In box 1630, the data placement coordinator 150 identifies all hyperedges in the hypergraph that have labels that include the selected data node. The data placement coordinator 150 then groups or clusters these hyperedges (or each included concept) together to form a group / cluster of concepts to be stored in the selected node. The method 1600 then proceeds to box 1635, where the data placement coordinator 150 determines whether there is at least one additional data node in the system that has not yet been assigned a group / cluster. If so, the method 1600 returns to box 1625.

[0156] Otherwise, the method 1600 proceeds to block 1640 where the data placement coordinator 150 maps, for each data node, all concepts included in the corresponding cluster to the repository. The data placement coordinator 150 can then use these mappings to distribute data across the various data nodes in the knowledge base.

[0157] Fig.1717 is a flow chart illustrating a method 1700 for mapping ontology concepts to storage nodes according to one embodiment disclosed herein. The method 1700 begins at box 1705, where the data placement coordinator 150 determines query workload information corresponding to a domain. At box 1710, the data placement coordinator 150 models the query workload information as a hypergraph, where the hypergraph includes a set of vertices and a set of hyperedges, where each vertex in the set of vertices corresponds to a concept in an ontology associated with the domain. The method 1700 then proceeds to box 1715, where the data placement coordinator 150 generates a mapping between the concepts and a plurality of data nodes based on the hypergraph and further based on predefined capabilities of each of the plurality of data nodes. Additionally, at box 1720, the data placement coordinator 150 establishes a distributed knowledge base based on the generated mapping.

[0158] The description of various embodiments of the present disclosure has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements existing in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

[0159] In the foregoing and / or the following, reference is made to the embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the specifically described embodiments. On the contrary, any combination of the foregoing and / or the following features and elements, whether or not related to different embodiments, is intended to be used to implement and practice the intended embodiments. In addition, although the embodiments disclosed herein can achieve advantages over other possible solutions or prior art, whether a given embodiment achieves a specific advantage does not limit the scope of the present disclosure. Therefore, the foregoing and / or the following aspects, features, embodiments and advantages are merely illustrative and are not considered to be elements or limitations of the appended claims unless expressly stated in the claims. Similarly, reference to the "present invention" should not be interpreted as a generalization of any inventive subject matter disclosed herein, and should not be considered to be elements or limitations of the appended claims unless expressly stated in the claims.

[0160] Various aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may all be collectively referred to herein as a "circuit," "module," or "system."

[0161] The present invention may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform various aspects of the present invention.

[0162] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal sent by a wire.

[0163] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0164] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as "C" programming language or similar programming languages). The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider through the Internet). In certain embodiments, in order to perform various aspects of the present invention, the electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.

[0165] Various aspects of the present invention are described herein with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention. It will be understood that each frame of the flow chart and / or block diagram and the combination of frames in the flow chart and / or block diagram can be implemented by computer-readable program instructions.

[0166] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can guide the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0167] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0168] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention.In this regard, each frame in the flow chart or block diagram can represent the module, segment or part of instruction, which includes one or more executable instructions for realizing the logical function of (one or more) specification.In some alternative embodiments, the function mentioned in the frame may not occur in the order mentioned in the figure.For example, the two frames shown in succession can actually be performed substantially at the same time, or these frames can sometimes be performed in reverse order, depending on the function involved.It will also be noted that the combination of each frame of the block diagram and / or flow chart illustration and the frame in the block diagram and / or flow chart illustration can be realized by a dedicated hardware-based system that performs a specified function or action or performs a combination of special hardware and computer instructions.

[0169] Embodiments of the present invention may be provided to end users via a cloud computing infrastructure. Cloud computing generally refers to providing scalable computing resources as a service over a network. More formally, cloud computing may be defined as providing an abstracted computing capability between computing resources and their underlying technology architecture (e.g., servers, storage devices, networks), thereby enabling convenient on-demand network access to a shared pool of configurable computing resources that can be quickly provisioned and released with minimal management effort or service provider interaction. Thus, cloud computing allows users to access virtual computing resources (e.g., storage, data, applications, and even complete virtualized computing systems) in a "cloud" without regard to the underlying physical systems (or the location of those systems) used to provide the computing resources.

[0170] Typically, cloud computing resources are provided to users on a pay-per-use basis, where users are charged only for the computing resources actually used (e.g., the amount of storage space consumed by the user or the number of virtualized systems instantiated by the user). Users can access any resource residing in the cloud at any time and from anywhere on the Internet. In the context of the present invention, users can access applications (e.g., query processing coordinator 115) or related data available in the cloud. For example, the query processing coordinator 115 can be executed on a computing system in the cloud and evaluate queries and back-end resources. In this case, the query processing coordinator 115 can route queries and store back-end resources and / or capability configurations at a storage location in the cloud. Doing so allows users to access the information from any computing system attached to a network (e.g., the Internet) connected to the cloud.

[0171] While the foregoing is directed to embodiments of the present invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof, and the scope of the invention is determined by the claims that follow.

Claims

1. A computer-implemented method for processing ontology queries, include: The ontology query is received by the query coordinator; generating a logical representation of a query based on the ontology query, the logical representation of the query comprising an ordered set of one or more query blocks, wherein each respective query block indicates one or more operations to be applied to respective input data, and the ordered set of one or more query blocks indicates one or more quantifiers of data flow between query blocks; for a first query block in the logical representation of the query, identifying at least one data node based on the one or more quantifiers and the one or more operations, the one or more operations comprising evaluating capabilities of one or more data nodes, wherein the capabilities indicate a set of operations that the one or more data nodes are capable of performing; selecting one or more of the identified data nodes based at least in part on a predefined cost criterion indicating a cost of moving data between the data nodes; as well as Send one or more subqueries to the selected data node or nodes.

2. The method of claim 1 , wherein for a first query block of the one or more query blocks, at least one data node is identified include: identifying, based on predefined capability criteria, a first set of data nodes configured to perform the one or more operations indicated by the first query block; as well as Based on a predefined concept mapping, a second set of data nodes configured to store the one or more quantifiers indicated by the first query block is identified.

3. The method according to claim 2, further comprising: include: Upon determining that the first data node belongs to both the first set of data nodes and the second set of data nodes, the first query block is associated with an annotation indicating the first data node.

4. The method according to claim 2, further comprising: include: Upon determining that no data node belongs to both the first set of data nodes and the second set of data nodes: identifying a first data node belonging to a first set of data nodes; associating the first query block with an annotation indicating the first data node; as well as One or more additional query blocks are generated to satisfy the one or more quantifiers of the first query block.

5. The method according to claim 2, further comprising: include: Upon determining that the one or more operations indicated by the first query block include a join operation between two or more quantifiers: identifying a first data node belonging to a first set of data nodes; as well as Upon determining that the first data node is configured to store at least one of the two or more quantifiers, the first query block is associated with an annotation indicating the first data node.

6. The method of claim 5, wherein upon further determining that the first data node is a relational data node, associating the first query block with an annotation indicating the first data node is performed, wherein based on determining that the second data node is not a relational data node, excluding from the annotation a second data node that is also configured to store at least one of the two or more quantifiers.

7. The method of claim 1 , wherein one or more of the data nodes are selected based on a predefined cost criterion include: identifying each possible combination of data nodes across each of the one or more query blocks; For each respective possible combination of data nodes, generating a respective mobility descriptor, the respective mobility descriptor defining a cost of transferring data between the respective combination of data nodes; as well as Based on determining that the mobility descriptors of the first combination of data nodes are lower than the mobility descriptors of all other possible combinations of data nodes, the first combination of data nodes is selected.

8. The method of claim 1, further comprising, for each corresponding query block of the one or more query blocks: identifying, from the ontology query, a corresponding query fragment corresponding to the corresponding query block; Determining the type of data node assigned to the corresponding query block; selecting a corresponding query translator based on the determined type of data node assigned to the corresponding query block; and A corresponding sub-query is generated by processing the corresponding query fragment using the corresponding query translator.

9. A computer readable storage medium containing computer program code, which when executed by one or more computer processors performs an operation, the operation include: The ontology query is received by the query coordinator; generating a logical representation of a query based on the ontology query, the logical representation of the query comprising an ordered set of one or more query blocks, wherein each respective query block indicates one or more operations to be applied to respective input data, and the ordered set of one or more query blocks indicates one or more quantifiers of data flow between query blocks; for a first query block in the logical representation of the query, identifying at least one data node based on the one or more quantifiers and the one or more operations, the one or more operations comprising evaluating capabilities of one or more data nodes, wherein the capabilities indicate a set of operations that the one or more data nodes are capable of performing; selecting one or more of the identified data nodes based at least in part on a predefined cost criterion indicating a cost of moving data between the data nodes; as well as Send one or more subqueries to the selected data node or nodes.

10. The computer-readable storage medium of claim 9, wherein for a first query block of the one or more query blocks, at least one data node is identified include: identifying, based on predefined capability criteria, a first set of data nodes configured to perform the one or more operations indicated by the first query block; as well as Based on a predefined concept mapping, a second set of data nodes configured to store the one or more quantifiers indicated by the first query block is identified.

11. The computer-readable storage medium of claim 10, wherein the operation further comprises: include: Upon determining that the first data node belongs to both the first set of data nodes and the second set of data nodes, the first query block is associated with an annotation indicating the first data node.

12. The computer-readable storage medium of claim 10, wherein the operation further comprises: include: Upon determining that no data node belongs to both the first set of data nodes and the second set of data nodes: identifying a first data node belonging to a first set of data nodes; associating the first query block with an annotation indicating the first data node; as well as One or more additional query blocks are generated to satisfy the one or more quantifiers of the first query block.

13. The computer-readable storage medium of claim 10, wherein the operation further comprises: include: Upon determining that the one or more operations indicated by the first query block include a join operation between two or more quantifiers: identifying a first data node belonging to a first set of data nodes; as well as Upon determining that the first data node is configured to store at least one of the two or more quantifiers, the first query block is associated with an annotation indicating the first data node.

14. The computer-readable storage medium of claim 9, wherein one or more of the data nodes are selected based on a predefined cost criterion. include: identifying each possible combination of data nodes across each of the one or more query blocks; For each respective possible combination of data nodes, generating a respective mobility descriptor, the respective mobility descriptor defining a cost of transferring data between the respective combination of data nodes; as well as Based on determining that the mobility descriptors of the first combination of data nodes are lower than the mobility descriptors of all other possible combinations of data nodes, the first combination of data nodes is selected.

15. The computer-readable storage medium of claim 9, the operations further comprising, for each respective query block of the one or more query blocks: identifying, from the ontology query, a corresponding query fragment corresponding to the corresponding query block; Determining the type of data node assigned to the corresponding query block; selecting a corresponding query translator based on the determined type of data node assigned to the corresponding query block; and A corresponding sub-query is generated by processing the corresponding query fragment using the corresponding query translator.

16. A system for processing ontology queries, include: one or more computer processors; as well as a memory containing a program that, when executed by the one or more computer processors, performs operations including: The ontology query is received by the query coordinator; generating a logical representation of a query based on the ontology query, the logical representation of the query comprising an ordered set of one or more query blocks, wherein each respective query block indicates one or more operations to be applied to respective input data, and the ordered set of one or more query blocks indicates one or more quantifiers of data flow between query blocks; for a first query block in the logical representation of the query, identifying at least one data node based on the one or more quantifiers and the one or more operations, the one or more operations comprising evaluating capabilities of the one or more data nodes, wherein the capabilities indicate a set of operations that the one or more data nodes are capable of performing; selecting one or more of the identified data nodes based at least in part on a predefined cost criterion indicating a cost of moving data between the data nodes; and Send one or more subqueries to the selected data node or nodes.

17. The system of claim 16, wherein for a first query block of the one or more query blocks, at least one data node is identified include: identifying, based on predefined capability criteria, a first set of data nodes configured to perform the one or more operations indicated by the first query block; as well as Based on a predefined concept mapping, a second set of data nodes configured to store the one or more quantifiers indicated by the first query block is identified.

18. The system of claim 17, wherein the operation further comprises: include: Upon determining that the first data node belongs to both the first set of data nodes and the second set of data nodes, the first query block is associated with an annotation indicating the first data node.

19. The system of claim 17, wherein the operation further comprises: include: Upon determining that no data node belongs to both the first set of data nodes and the second set of data nodes: identifying a first data node belonging to a first set of data nodes; associating the first query block with an annotation indicating the first data node; as well as One or more additional query blocks are generated to satisfy the one or more quantifiers of the first query block.

20. The system of claim 17, wherein the operation further comprises: include: Upon determining that the one or more operations indicated by the first query block include a join operation between two or more quantifiers: identifying a first data node belonging to a first set of data nodes; as well as Upon determining that the first data node is configured to store at least one of the two or more quantifiers, the first query block is associated with an annotation indicating the first data node.

Citation Information

Patent Citations

  • Information access using ontologies

    WO2005008358A2