Ontology-Based Data Storage for Distributed Knowledge Bases

By using the data coordinator to model the query workload as a hypergraph in the distributed knowledge base and generating a mapping of concepts and data nodes, the problem that existing systems cannot effectively manage and route queries is solved, and efficient data storage and query response are achieved.

CN114586012BActive Publication Date: 2025-07-11INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080070094.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2020-09-30
Publication Date
2025-07-11
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

When managing distributed knowledge bases, existing systems cannot effectively understand the underlying ontology and capabilities of each data source, resulting in inefficiency in query routing, high computing resources consumption, and existing architectures rely on centralized intermediaries to lead to inefficiency and poor scalability.

Method used

Query workload information is determined through the data coordinator, modeled as a hypergraph, generate mappings between concepts and data nodes, and effectively place and store data in a distributed environment based on the hypergraph and node capabilities, and optimize the query routing process.

Benefits of technology

It reduces data storage computing overhead, improves system responsiveness and efficiency, reduces computation waste at runtime, and realizes efficient data placement and query routing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114586012B_ABST
    Figure CN114586012B_ABST
Patent Text Reader

Abstract

Techniques for distributed data placement are provided. Query workload information corresponding to a domain is determined by a data coordinator and modeled as a hypergraph, where the hypergraph includes a set of vertices and a set of hyperedges, and each vertex in the set of vertices corresponds to a concept in an ontology associated with the domain. Based on the hypergraph and further based on predefined capabilities of each of a plurality of data nodes, a mapping between concepts and the plurality of data nodes is generated. A distributed knowledge base is established based on the generated mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to knowledge bases, and more particularly, to using ontologies for efficient data placement in a distributed knowledge base. Background Art

[0002] More and more enterprises are leveraging knowledge bases (KBs) to enhance their analytics and improve the decision-making, efficiency, and effectiveness of their systems. Typically, within an enterprise's domain, a KB is relatively specialized. For example, a financial institution relies on a KB with important financial knowledge, such as data related to government regulations in financial markets. In contrast, a healthcare enterprise may maintain a KB with a substantial amount of data collected from medical literature. There is a substantial need for effective systems and technologies for managing these KBs with deep domain specialization. Existing systems that do not know the domain ontology providing a central view of the domain schema cannot effectively manage and route queries to the KB.

[0003] In addition, a KB can be distributed across multiple data sites with different capabilities and costs to improve operations. Existing architectures such as federated databases rely on a centralized mediator to aggregate data from each such source. This is inefficient and has poor scalability. Additionally, existing systems do not understand the underlying ontology and capabilities of each data source and thus cannot effectively route queries. This reduces the efficiency of such systems and requires a large amount of computing resources to respond to typical queries. Summary of the Invention

[0004] According to one aspect of the present invention, a method is provided that includes: determining, by a data coordinator, query workload information corresponding to a domain. The method further includes modeling the query workload information as a hypergraph, where the hypergraph includes a set of vertices and a set of hyperedges, and each vertex in the set of vertices corresponds to a concept in an ontology associated with the domain. Additionally, the method includes generating a mapping between concepts and a plurality of data nodes based on the hypergraph and further based on predefined capabilities of each of the plurality of data nodes, and establishing a distributed knowledge base based on the generated mapping. Advantageously, the method enables the data coordinator to effectively place and store data in a distributed environment based on the existing workload and the capabilities of each data node. This reduces the computational overhead required to store data and further improves system responsiveness by providing an effective data mapping.

[0005] According to another embodiment of the present disclosure, determining query workload information includes receiving a set of previous ontology queries. The method according to this embodiment includes generating a first set of concepts accessed by a first query in the set of previous ontology queries, and generating a second set of operations performed by the first query. The first generalized query is generated by the following steps: identifying, from the set of previous ontology queries, a set of queries having a corresponding matching first set, and determining an aggregated set of operations based on the corresponding second set of each query in the identified set of queries. Then the first generalized query is associated with the aggregated set of operations and the concepts reflected in the corresponding matching first set. In such an embodiment, the data coordinator improves the existing system by effectively generalizing previous queries to determine a good storage plan for the data, so as to meet the expected needs. This again improves efficiency and reduces computational waste at runtime.

[0006] According to yet another embodiment of the present disclosure, modeling query workload information as a hypergraph includes creating a vertex for each concept in the ontology and creating a first hyperedge for the first generalized query, where the first hyperedge connects a first set of vertices in the hypergraph, and the first set of vertices corresponds to the concepts reflected in the matching first set. In one such embodiment, the method further includes labeling the first hyperedge with the aggregated set of operations. Advantageously, such an embodiment enables the data coordinator to effectively represent the workload in graph form, which allows the coordinator to better and more effectively evaluate the data to drive improved placement decisions. This significantly improves runtime performance.

[0007] According to yet another embodiment of the present disclosure, generating a mapping includes creating a first cluster for a first operation included in the hypergraph. This embodiment then includes identifying a first set of concepts connected by the first hyperedge in the hypergraph, and identifying a first set of operations indicated by the first hyperedge. The method includes assigning the first set of concepts to the first cluster when it is determined that the first set of operations includes the first operation. One advantage of such an embodiment is that it enables efficient evaluation of the hypergraph and results in highly reliable data placement that requires minimal movement at runtime.

[0008] According to another embodiment of the present disclosure, generating a mapping further includes mapping the first set of concepts to one or more data nodes by the following steps: identifying a set of data nodes capable of performing the first operation, and mapping each concept in the first set of concepts to each data node in the identified set of data nodes. Advantageously, this enables the system to generate a data mapping based on the capabilities of the nodes while considering previous workloads. This increases the likelihood that the system will be adequately positioned to respond to future queries with minimal latency and resource consumption.

[0009] According to another embodiment of the present disclosure, generating the mapping includes identifying a first set of concepts connected by a first hyperedge in a hypergraph, identifying a first set of operations indicated by the first hyperedge, and determining a minimum set of data nodes capable of jointly performing the set of operations. The method then includes generating a cluster including the first set of concepts and labeling the cluster with the minimum set of data nodes. Such an embodiment improves existing solutions by minimizing the replication of data in the system, which reduces storage costs and additionally reduces latency and resource consumption caused by data movement in the system.

[0010] According to another embodiment of the present disclosure, generating the mapping further includes mapping each concept in the first set of concepts to each data node in the minimum set of data nodes. This similarly reduces the storage cost and waiting time of the system.

[0011] According to yet another embodiment of the present disclosure, establishing a distributed knowledge base includes, for each corresponding concept in an ontology, identifying the corresponding data node indicated by the mapping, identifying the data corresponding to the corresponding concept, and facilitating the storage of the identified data in the corresponding data node. Advantageously, such an embodiment enables a coordinator to efficiently generate the mapping and place the data in appropriate nodes in an efficient manner, which reduces the computation before and during runtime.

[0012] According to different embodiments of the present invention, any of the above embodiments can be implemented by a computer-readable storage medium. The computer-readable storage medium contains computer program code that, when executed by the operation of one or more computer processors, performs the operations. In an embodiment, the operations performed can correspond to any combination of the above methods and embodiments.

[0013] According to yet another different embodiment of the present disclosure, any of the above embodiments can be implemented by a system. The system includes one or more computer processors and a memory containing a program that, when executed by the one or more computer processors, performs the operations. In an embodiment, the operations performed can correspond to any combination of the above methods and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Illustrates an architecture configured to place knowledge base data and route ontology queries according to an embodiment disclosed herein.

[0015] Figure 2 Illustrates a workflow for processing and routing ontology queries according to an embodiment disclosed herein.

[0016] Figure 3A and Figure 3BDepicts an example ontology that can be used to evaluate and route queries according to an embodiment disclosed herein.

[0017] Figure 4A and Figure 4B Illustrates a workflow for parsing and routing example ontology queries according to an embodiment disclosed herein.

[0018] Figure 5 Is a block diagram showing a query processing coordinator configured to route ontology queries according to an embodiment disclosed herein.

[0019] Figure 6 Is a flowchart showing a method for processing and routing ontology queries according to an embodiment disclosed herein.

[0020] Figure 7 Is a flowchart showing a method for processing ontology queries to efficiently route query blocks according to an embodiment disclosed herein.

[0021] Figure 8 Is a flowchart showing a method for evaluating potential query routing plans to efficiently route ontology queries according to an embodiment disclosed herein.

[0022] Figure 9 Is a flowchart showing a method for routing ontology queries according to an embodiment disclosed herein.

[0023] Figure 10 Illustrates a workflow for evaluating workloads and storing knowledge base data according to an embodiment disclosed herein.

[0024] Figure 11 Depicts an example hypergraph for evaluating workloads and storing knowledge base data according to an embodiment disclosed herein.

[0025] Figure 12 Is a block diagram showing a data placement coordinator configured to evaluate workloads and store data according to an embodiment disclosed herein.

[0026] Figure 13 Is a flowchart showing a method for evaluating and summarizing query workloads to inform data placement decisions according to an embodiment disclosed herein.

[0027] Figure 14 Is a flowchart showing a method for modeling ontology workloads to inform data placement decisions according to an embodiment disclosed herein.

[0028] Figure 15 Is a flowchart showing a method for evaluating hypergraphs to drive data placement decisions according to an embodiment disclosed herein.

[0029] Figure 16 is a flowchart showing a method for evaluating a hypergraph to drive data placement decisions according to an embodiment disclosed herein.

[0030] Figure 17 is a flowchart showing a method for mapping ontology concepts to storage nodes according to an embodiment disclosed herein. Detailed Description

[0031] Embodiments of the present disclosure provide an ontology-driven architecture that uses multiple data stores or nodes to support various query types, where each data store or node can have different capabilities. Advantageously, embodiments of the present disclosure enable queries to be efficiently processed and routed to diverse data repositories based on an existing ontology, which improves the latency of the system and reduces the computational resources required to identify and return relevant data. In an embodiment, the techniques described herein can be used to support various query applications, including natural language and conversational interfaces. The systems described herein can provide transparent access to underlying backend repositories via an abstract ontology query language (OQL) that allows users to express their information needs against a domain ontology.

[0032] In one embodiment, to provide improved performance for different query types, the system first optimizes the placement of KB data into different repositories by placing subsets of the data in appropriate backend repositories based on their capabilities. In some embodiments, at runtime, the system parses and compiles OQL queries and translates them into the query languages and / or APIs of different backend data nodes. In one embodiment, the systems described herein also employ a query coordinator that routes queries to single or multiple backend repositories based on the placement of relevant data. This efficient routing reduces the latency required to generate and return results to the requesting entity.

[0033] In an embodiment, the KB data to be queried can include any information, including structured, unstructured, and / or semi-structured data. To support deep domain specialization, embodiments disclosed herein utilize a domain ontology. Notably, in one embodiment, the domain ontology defines entities and their relationships only at the metadata level and does not provide instance-level information. That is, in one embodiment, the domain ontology provides a domain schema, while instance-level data is stored in various backend data repositories (also referred to as data sites and data nodes). In an embodiment, the data nodes can include any repository architecture, including (one or more) relational databases, (one or more) inverted index document repositories, (one or more) JavaScript Object Notation (JSON) repositories, (one or more) graph databases, etc.

[0034] In some embodiments, the system can move, store, and index KB data in any backend, and the backend provides the capabilities required to support the query type. In other words, the system does not need to federate data across existing operational data stores. This is significantly different from federated databases and mediator-based approaches. In some embodiments, the system performs a capability-based data placement step, in which knowledge base data conforming to a given ontology is stored in various backend resources. In at least one embodiment, the multi-repository architecture is configured to minimize data movement between different data repositories, because data movement not only incurs data transfer costs but also results in expensive data transformation costs. One solution to minimize data movement costs is to replicate data across all backend nodes, thus ensuring that queries can be answered by a single repository without any data movement. However, this also requires wasteful replication and defeats the purpose of using a multi-repository architecture to balance the different capabilities of multiple data stores. In embodiments of the present invention, KB data can be stored in any number of backend resources, and the system intelligently routes queries to select the most appropriate data repository(s) and minimize data movement costs.

[0035] In some embodiments described herein, the system architecture provides appropriate abstractions to query data without knowing how the data is stored and indexed in multiple data repositories. In one embodiment, to achieve this, an Ontology Query Language (OQL) is introduced. OQL is expressed against the enterprise's domain ontology. The users of the system only need to know the domain ontology that defines the entities and their relationships. In embodiments, the system understands the mapping of various ontology concepts to backend data repositories and their corresponding schemas, and provides a query translator from OQL to the target query language of the underlying system.

[0036] In embodiments, for any given query, one or more data nodes may be involved. Embodiments of the present disclosure provide techniques for identifying the involved repositories and generating appropriate subqueries to compute the final query result. In some embodiments, the system does not use the mediator approach, in which a global query is divided into multiple subqueries and the final result is assembled in the mediator. Instead, one or more of the data repositories are used as mediators, and the (one or more) backend data repositories finalize the query answer. For example, in one embodiment, a relational database is used to finalize the query response because the relational repository is likely to be able to complete the join operation quickly.

[0037] Figure 1FIG. 100 shows an architecture configured to route ontology queries according to an embodiment disclosed herein. In the illustrated embodiment, a query processing coordinator 115 receives an ontology query 105 constructed using OQL, which has a corresponding ontology schema 110. The ontology query 105 is processed by the query processing coordinator 115 to identify and return relevant data. In an embodiment, the query processing coordinator 115 routes the query to a backend 125 that includes various data nodes 130A-N. In the illustrated embodiment, this routing is performed at least in part based on a set of concept mappings 140 and capabilities 135.

[0038] In the illustrated embodiment, the ontology query 105 is formatted based on OQL, which is used to represent queries that operate on a set of concepts and relationships in the ontology schema 110. In one embodiment, OQL can express queries that include aggregation, union, nested subqueries, etc. In some embodiments, OQL can also express full-text and field search predicates, as well as path queries. An OQL query typically consists of a single query block or a union of multiple OQL query blocks, where each OQL query block consists of a SELECT clause and a FROM clause. In some embodiments, an OQL block can also contain a WHERE clause, a GROUP BY clause, an ORDER BY clause, a FETCH FIRST clause, and / or a HAVING clause. In an embodiment, an OQL query operates on a set of tuples constructed from the Cartesian product of the concepts referenced in the FROM clause.

[0039] In an embodiment, the ontology schema 110 describes entities and their relationships at a semantic level, regardless of how the data is actually stored in the backend resources. In one embodiment, the ontology schema 110 describes domain-related entities, attributes that may be associated with various entities, and potential relationships between different entities. The ontology schema 110 can provide a rich and expressive data model that captures various real-world relationships between entities, such as functionality, inheritance, union, etc. In some implementations, the ontology schema 110 does not include instance data. That is, the ontology schema 110 defines entities and relationships, but does not include data related to specific entities or relationships in the KB. For example, the ontology schema 110 can define that a "Company" entity can have several attributes such as "name" and "address", but the ontology schema 110 does not include data for a specific instance, such as a company with the name "Main Street Grocer" and the address "123 Main Street". Instead, this instance data is maintained separately in the backend 125.

[0040] In an embodiment, the concept mapping 140 indicates the correspondence between the logical schema represented by the ontology schema 110 and the physical schema of the underlying data nodes 130 in the backend 125. For example, assume the ontology schema 110 is defined as where C = {c n | 1 ≤ n ≤ N} represents the set of concepts (also referred to as entities), R = {r k | 1 ≤ k ≤ K} represents the set of relationships between concepts / entities, and P = {p m | 1 ≤ m ≤ M} is the set of data attributes. In one embodiment, each relationship is between two or more concepts, and each data attribute corresponds to a characteristic of a concept. In the illustrated embodiment, the concept mapping 140 indicates the mapping between the concepts, relationships, and attributes (defined in the ontology schema 110) and the underlying data nodes 130. That is, in one such embodiment, the concept mapping 140 indicates each data node 130 that stores a given concept (e.g., entity), relationship, and / or attribute. For example, the concept mapping 140 may indicate that instances of the "Company" entity are stored in data nodes 130A and 130M, while data related to a specific type of relationship between "Company" entities is stored in data node 130B.

[0041] In an embodiment, the capabilities 135 indicate the operations and capabilities provided by each data node 130. In some embodiments, the capabilities of the data nodes 130 are expressed as views that enumerate all possible queries (view definitions) that can be processed / answered by the data store. While this approach is flexible, it is not scalable because the number of view definitions can be very large, potentially leading to the problem of query rewriting using an infinite number of views. To improve scalability, some embodiments of the present disclosure describe their capabilities in terms of the operations supported by the backend data nodes 130 (e.g., join, group, aggregation, fuzzy text matching, path expressions, etc.) rather than enumerating all possible queries that the data store can answer. Additionally, at least one embodiment of the present disclosure provides a more fine-grained description of each supported operation by leveraging a mechanism for expressing any associated restrictions. For example, in one such embodiment, an aggregation function of type MAX may only be supported on numeric types.

[0042] Thus, in one embodiment, the capabilities 135 indicate, for each given data node 130, the set of operations that the node can perform (and, in some embodiments, the related restrictions on that capability). In an embodiment, the query processing coordinator 115 evaluates the capabilities 135 and the concept mapping 140 to select one or more data nodes 130 to which the received ontology query 105 should be routed. To do this, in one embodiment, the query processing coordinator 115 identifies data nodes 130 that contain the required data (e.g., based on the concept mapping 140), can perform the required operation(s) (e.g., based on the capabilities 135), or both. The query processing coordinator 115 then generates subqueries for each selected data node 130.

[0043] In an embodiment, the query processing coordinator 115 uses a set of translators 120A-N to translate subqueries as needed. In the illustrated embodiment, each type of data node 130 has a corresponding translator 120. In some embodiments, each translator 120A-N receives all or part of an OQL query and generates an equivalent query in the language and / or syntax of the corresponding data node 130A-N. For example, translator 120A may generate an SQL query for a relational database contained in data node 130A, while translator 120B outputs a graph query for a graph repository contained in data node 130B.

[0044] In one embodiment, the translator 120 relies on a schema mapping that maps concepts and relationships represented in the domain ontology to appropriate schema objects in the target physical schema. For example, for a relational backend data node 130, the schema mapping may provide (1) a correspondence between the concepts in the ontology and the tables in the schema; (2) the data attributes or traits of the columns of the tables in the physical schema corresponding to the concepts in the ontology; and / or (3) the relationship between the concepts in the ontology and the primary key-foreign key constraints between the tables corresponding to the concepts in the database. Similarly, for a JSON document repository, the schema mapping may map the concepts, data attributes, and relationships represented in the ontology to appropriate field paths in the JSON document.

[0045] In some embodiments, the translator 120 also processes special concepts and relationships that may be represented in the ontology, such as unions, inheritances, and traversals between concepts that typically represent join conditions between ontology concepts. Depending on the physical data layout, these are translated into appropriate operations supported by the backend 125 data node 130.

[0046] In an embodiment, each data node 130 receives a subquery and generates a response for the query processing coordinator 115. If the data node 130 has all the necessary data locally and is able to complete the indicated operation(s), the node executes the query and returns the result. In some embodiments, the subquery may indicate that the data node 130 should send a query to one or more other data nodes 130 to retrieve the data needed to complete the query. For example, in some embodiments, preferably, a join operation is performed by a relational data node because relational systems tend to have low latency for join operations. Additionally, some operations are only possible on certain nodes. For example, assume that data node 130A is the only node that can complete a join operation, but the data to be joined only exists on data node 130B. In one embodiment, data node 130A receives a subquery instructing it to join related data, and one or more other subqueries to be forwarded to data node 130B to retrieve the data.

[0047] That is, in one such embodiment, the query processing coordinator 115 prepares subqueries for translation by data node 130B and sends them to data node 130A. This allows data node 130A to simply forward them to data node 130B. Then, data node 130B returns the data to data node 130A, which completes the operation and returns the result to the query processing coordinator 115. In another embodiment, the query processing coordinator 115 may send the subqueries to data node 130B and forward the resulting data to data node 130A for processing. In yet another embodiment, the query processing coordinator 115 performs the join locally.

[0048] In the illustrated embodiment, the architecture 100 further includes a data placement coordinator 150 that determines which of the data nodes 130A-N should be used to store data in the knowledge base. As shown, the data placement coordinator 150 receives similar indications of the capabilities 155 of the data nodes 130, as well as an indication of the query workload 160. In one embodiment, the query workload 160 indicates an average or expected query set for the knowledge base. For example, in one embodiment, the query workload 160 is generated by observing the interactions of users with the knowledge base over time. The system can then aggregate these interactions to determine the average, expected, or typical workload of the system. In some embodiments, the query workload 160 includes indications of which concepts are queried together, which operations are applied to each concept, etc.

[0049] In an embodiment, based on the node capabilities 155 and the known query workload 160, the data placement coordinator 150 determines which data node(s) 130 should store the instance-level data of each concept / entity in the ontology schema 110. For example, based on the operations that a given data node 130A can perform and further based on the query workload 160, the data placement coordinator 150 may determine that the data node 130A should store all the data of the "company" entity and the "public metric" entity because they are often queried together. Similarly, the data placement coordinator 150 may determine to place the instances of the "document" concept in the data node 130B based on determining that the queries frequently include performing fuzzy matching operations on the "document" data and the data node 130B supports fuzzy matching.

[0050] Such intelligent data placement can reduce subsequent data transfers during runtime. In an embodiment, the data placement coordinator 150 can be utilized at the start (e.g., to provide an initial placement) and / or periodically during runtime (e.g., to reform the placement decision based on how the query workload 160 evolves over time). Although depicted as discrete components for conceptual clarity, in an embodiment, the operations of the query processing coordinator 115 and the data placement coordinator 150 can be combined or distributed across any number of components.

[0051] As shown, the data placement coordinator 150 outputs a set of placement decisions 165 to the backend 125 and / or to one or more intermediate services such as extract, transform, and load (ETL) services. Then, based on these selections, the instance-level data in the knowledge base is stored in the appropriate data nodes 130. The concept mapping 140 reflects the current placement of the data. In an embodiment, if the data placement coordinator 150 is used to modify the data placement based on an updated query workload 160, the concept mapping 140 is similarly updated. Hereinafter, Figures 2 to 9 the query processing coordinator 115 is discussed in more detail, and it is assumed that the data has been placed. Refer to Figures 10 to 17 The data placement coordinator 150 and various techniques to ensure effective data placement are discussed in more detail.

[0052] Figure 2Shows a workflow 200 for processing and routing ontology queries according to an embodiment disclosed herein. As shown, the workflow 200 begins upon receipt of an OQL query 205. The OQL query 205 is provided to an OQL parser 210, which parses the query to determine the meaning of the query. As shown by arrow 215, the parsed OQL query is then provided to a Query Graph Model (QGM) constructor 220. In one embodiment, the QGM constructor 200 generates a logical representation of the query in the form of a query graph model. This reduces the complexity of query compilation and optimization. The following discusses example query graph models in more detail with reference to Figure 4A and Figure 4B Figure 4B . In an embodiment, the QGM uses operator boxes such as SELECT, GROUP BY, SETOP, etc. to capture the data flow and dependencies in the query. In an embodiment, the operations within a box can be freely reordered among themselves, but the box boundaries are respected when generating the query execution plan. That is, the query execution plan must follow the order of the query boxes.

[0053] In one embodiment, quantifiers are used to represent the data flow between query blocks. This format allows the system to reason about query equivalence and apply rewrite optimizations. In an embodiment, the QGM representation of a query enables the system to focus on optimizing the data flow between different data repositories during query execution, while the selection of the actual physical execution plan is deferred to the underlying data nodes 130 responsible for executing the query boxes or fragments. In an embodiment, a QGM 225 is used to generate an optimized multi-store execution plan that minimizes data movement and transformation across different backends.

[0054] In some embodiments, the QGM 225 includes a set of quantifiers at the bottom that provide a set of input concepts to the query, and each query block in the QGM has a head and a body. The body of each block includes a set of predicates that describe the set operations (such as joins) to be performed on the set of input concepts, and the head expression describes how the output attributes of the result concept should be computed. In other words, in an embodiment, the body of each box contains a set of predicates (also called operations), each of which will be applied to the input quantifiers of the box. In some embodiments, predicates / operations that reference a single quantifier are classified as local predicates, while predicates / operations that reference multiple quantifiers express join predicates.

[0055] In the illustrated workflow 200, the QGM 225 is passed to the operator placement component 230 that continues the query routing process. As shown, the operator placement component 230 also receives a collection of concept maps 235 and node capabilities 240, and annotates the query blocks in the QGM 225 based on the capabilities and concept maps. In one embodiment, the operator placement component 230 moves the QGM from the bottom to the top and annotates each operation in the query with a collection of possible repositories that can perform the operation. In some embodiments, the annotations are generated differently for local predicates and join predicates. Recall that in some embodiments, all head expressions associated with a single quantifier are treated in the same way as local predicates.

[0056] In one embodiment, if the query block includes a local predicate / operation (e.g., a single quantifier), the operator placement component 230 determines whether the quantifier is a primitive concept. If so, the predicate is annotated with (i) an indication of the data node 130 that contains the concept and (ii) the ability to execute the predicate. In an embodiment, if the quantifier is from another QGM block (i.e., it is computed by another block), the operator placement component 230 annotates the predicate with a collection of data repositories that (i) complete the QGM box for the quantifier and (ii) are able to execute the predicate. In some embodiments, if none of the data nodes 130 that contain the input to the predicate also have the ability to execute the predicate, the operator placement component 230 annotates it with a collection of repositories that are able to execute the predicate.

[0057] For example, if the predicate includes a fuzzy search, but the data is only stored in a relational backend data node 130, the operator placement component 230 may annotate the predicate with a document repository that can compute the fuzzy search, even though the data is not stored there. Note that during query execution, data movement will be required in such a case.

[0058] In some embodiments, if the query block includes a join predicate / operation, the operator placement component 230 examines the join type and the join predicate. In an embodiment, each join predicate is associated with two or more quantifiers: one for each join input. In an embodiment, the operator placement component 230 identifies the progression of these quantifiers based on whether the identified quantifiers are computed inputs or primitive concepts and where the quantifiers come from (e.g., whether they are computed and / or locally stored, or will be received from another node).

[0059] In all cases where the quantifiers cover the basic concepts, the operator placement component 230 evaluates the set of data nodes 130 in which any basic concept resides and determines for each such repository whether that repository supports a join operation. The operator placement component 230 then annotates the join operation with an indication of the repository type that contains one or more of the required concepts and supports the join operation. In some embodiments, if one of the repositories is a relational node, the operator placement component 230 annotates the operation with an indication of that relational data node 130 rather than the rest of the nodes.

[0060] In some embodiments, if both quantifiers are produced by other QGM query blocks, the operator placement component 230 annotates the join operation with the data nodes 130 that support the join operation type (in some embodiments, including only relational repositories). In an embodiment, if the data repositories that produce the quantifiers are capable of performing a join, the operator placement component 230 also annotates the operation with their indications.

[0061] In another embodiment, if one of the quantifiers is a basic concept and the other is computed from another QGM query block, the join operation placement decision is similar to that discussed above when both quantifiers are computed by other QGM blocks. In an embodiment, the operator placement component 230 annotates the join operation with the set of data repositories that support the join type and are local to at least one of the join inputs.

[0062] As shown, the annotated QGM 245 is then passed to the block placement component 250. In one embodiment, the block placement component 250 uses the annotation generated by the operator placement component 230 to determine the possible placement options for the query boxes in the query.

[0063] In one embodiment, determining the placement options for a "select" query box is modeled as a problem of determining a minimum set cover. That is, the block placement component 250 determines the minimum number of data nodes 130 required to satisfy placing all the operators within a given select box. In some embodiments, the block placement component 250 uses a greedy heuristic algorithm to perform predicate-repository grouping, which ensures that each predicate is placed in a single repository and the total number of repositories spanned by the query box is minimized. Once each predicate is annotated with the appropriate repository in which it will be placed, the block placement component 250 determines whether the query block needs to be split.

[0064] In an embodiment, if all the predicates in a selected query box are placed in the same data node 130, the query box is annotated to be placed in that repository. Conversely, if the predicates in a selected query box are placed in more than one data node 130 (indicating that the query box needs to be processed by multiple repositories), the block placement component 250 determines that the block must be split into multiple query boxes. In this case, the block placement component 250 then partitions the block such that each resulting (sub) block includes predicates that are assigned to a single data node.

[0065] In some embodiments, for a "group by" query box, the block placement component 250 determines placement based on the repository that supports the type of aggregation operation and the repository that processed the previous query block of the query box that feeds the group by box. In an embodiment, if the repository that feeds these inputs also supports the "group by" and "aggregation" functions, the block placement component 250 places the selected "group by" box in the same repository. However, if the repository does not support these operations, the block placement component 250 annotates the box with a list of data nodes 130 that can perform the group by operation. Note that this placement will require data movement from the feeding repository to the repository that can process the box.

[0066] Once the block placement component 250 has annotated each query block with all possible placements (e.g., all data nodes 130 that can execute the block) of each query block, the cost component 255 determines the cost of each combination of placements. In one embodiment, the cost model for query routing focuses only on data transfer and data transformation costs in order to select between alternative query execution plans because the costs of data movement and transformation can dominate the total cost of any execution plan in a multi-repository environment. Thus, in at least some embodiments, the cost of executing a given execution plan is determined by aggregating the costs of data movement for each source and target data repository pair in the query execution plan. Note that in some embodiments, different operations such as joins and graph operations can be performed with very different performance on different data nodes 130. For example, although join operations may be supported by various data repositories such as JSON repositories, relational databases typically provide the best performance for join operations.

[0067] In some embodiments, the cost model also takes into account such variations in the execution costs of operations on different repositories. In other embodiments, the cost model does not include these factors. In one such embodiment, a declarative mechanism that expresses the capabilities of the underlying repository is used to handle such situations. For example, to avoid performing a join operation on a JSON repository, the system can either completely mask the join capabilities of the JSON repository in the capabilities description or add restrictions to limit its applicability to specific data types.

[0068] In an embodiment, the cost component 255 enumerates all possible placement combinations on all query blocks and generates an execution plan for each such combination by grouping certain query boxes together into groups. In one embodiment, each group in the query plan includes query boxes that are both (i) results in the QGM (e.g., separated by a single hop) and (ii) processable by the same data node 130. The cost component 255 then defines edges representing the data flow between the data nodes 130 based on the groups of these blocks in the QGM flow join. In an embodiment, the cost component 255 also generates a set of data movement descriptors for each execution plan based on the edges between the groups in the generated plan.

[0069] In some embodiments, using these movement descriptors for each plan, the cost component 255 uses the data movement cost model described above to find the cost for each execution plan and picks the one with the minimum cost. For example, in one embodiment, the cost component 255 determines the cost of each of the corresponding movement descriptors for each possible execution plan. This can include latency, computational cost, etc. These values can then be aggregated within each execution plan to determine which (if any) execution plan has the lowest cost. As shown, this minimum cost plan is then selected and used to generate one or more translated queries 260 for the data nodes 130.

[0070] Figure 3A and Figure 3B Depicted is an example ontology 300 that can be used to evaluate and route queries according to one embodiment disclosed herein. In the illustrated embodiment, each concept 305A - I is depicted using an ellipse, while each attribute 310A - N is depicted using a rounded rectangle. Additionally, arrows are used to depict the associations between the concepts 305 and the attributes 310, and bold arrows depict the relationships between the concepts 305. Further, each relationship arrow is labeled based on the type of the relationship. For example, as shown, the "Company" concept 305B is a subclass of the "PublicCompany" concept 305D. Additionally, the company concept 305B is associated with several attributes 310, including the "name" attribute 310E and the identifier attribute 310D, both of which are strings.

[0071] As described above, in an embodiment, the data in the knowledge base conforms to ontology 300. In other words, ontology 300 defines the concepts and entities in the knowledge base, as well as the relationships between the entities and the attributes associated with each concept. Notably, ontology 300 does not include any instance data (e.g., data about a specific company), but rather defines the structure of the data. In some embodiments, concepts 305, attributes 310, and / or relationships may be distributed across any number of data nodes 130. In one embodiment, data is placed in data nodes 130 at least partially based on the type of data.

[0072] In one such embodiment, if the concept mapping indicates that a given concept 305 is stored in a given data node 130, then all instance data corresponding to that concept 305 is stored in the data node 130. For example, if the mapping indicates that data node 130A includes the "PublicMetric" concept 305F, then the query processing coordinator 115 can retrieve data about any instance of "PublicMetric" (e.g., any metric for any company) from data node 130A. Thus, if a received query will require access to "PublicMetric" data, the query processing coordinator 115 will route at least a portion of the query to data node 130A (or to another node that also serves concept 305F).

[0073] Figure 4A and Figure 4B Workflow 400 for parsing and routing example ontology queries in accordance with one embodiment disclosed herein is shown. In the illustrated embodiment, an input 405 is received and evaluated to retrieve relevant data from data nodes 130. In the illustrated embodiment, the input 405 is natural language text (e.g., from a user). However, in various embodiments, the input 405 may include a query or other data. Additionally, the input 405 may be received from any number of sources, including automated applications, user-facing applications, directly from the user, etc. In some embodiments, the input 405 is received as part of a chatbot or other interactive application that allows a user to search and explore the knowledge base.

[0074] In the example shown, the input 405 is the phrase "Show me the total revenue of all companies that filed technology patents in the last 5 years". In the workflow 400 shown, one or more natural language processing (NLP) techniques such as semantic analysis, keyword search, sentiment analysis, intent analysis, etc. are used to parse and evaluate the input 405. This enables the system to generate an OQL query 410 based on the input 405. Of course, in some embodiments, the input 405 itself is an OQL query.

[0075] As shown, the corresponding OQL query 410 for the input 405 includes a "select" operation and a "group by" operation, as well as an indication of the relevant tables or concepts, and a "where" clause indicating the restrictions on the query. The OQL query 410 is then parsed by the query processing coordinator 115 to generate an OGM 435A, which is a logical representation of the query and includes a plurality of query blocks. In the embodiment shown, the OGM includes a "select" query block 415B and a "group by" query block 415A.

[0076] As shown, the QGM 435A includes a set of quantifiers 430A-E at the bottom, which provide the input concept set to the query. In the embodiment shown, these include "PublicMetricData", "PublicMetric", "PublicCompany", "Document", and "CompanyInfo". In an embodiment, each query block 415 in the QGM 435A includes a header 420 and a body 425. The body 425 generally describes the set operations to be performed on the input concept set (such as joins), while the header 420 expression describes how the output attributes of the result concept should be calculated. For example, in the embodiment shown, the header 420B of the "select" query block 415B specifies the output attributes "oPMD.value", "oPMD.year_calendar", and "oCl.id". The body 425B contains a set of predicates to be applied to the input quantifiers 430A-E. As described above, in an embodiment, a predicate that references a single quantifier is a local predicate, while a predicate that references multiple quantifiers expresses a join predicate.

[0077] As Figure 4BAs shown, since the query processing coordinator 115 recognizes that the "select" query block 415B includes predicates that must be executed by different data nodes 130. Specifically, although most of the predicates in the query block 415B can be executed by the relational data repository, the query processing coordinator 115 has determined that the predicates "oD->companylnfo = oCI" and "oD.selfMATCH('Tech Patent Filed')" correspond to operations that are executed using Elasticsearch, which is not supported by the relational database nodes. Therefore, the query processing coordinator 115 splits the query block 415B by separating these predicates into a new query block 415C. As shown, this query block 415C serves as the input to the query block 415B.

[0078] In the illustrated embodiment, by splitting the query block 415B, the query processing coordinator 115 has ensured that each block can be executed entirely within a single data node 130. For example, both query blocks 415A and 415B can be executed in the relational data node 130, while the query block 415C will be executed by the data node 130 that supports Elasticsearch. Thus, in the embodiment, the query processing coordinator 115 identifies or generates subqueries corresponding to the query blocks 415A and 415B, translates the subqueries into the appropriate language and / or format for the relational node, and sends the translated subqueries to the relational node.

[0079] In the illustrated embodiment, the query processing coordinator 115 will further identify or generate a subquery to complete the query block 415C and translate it into the language and / or format supported by the Elasticsearch node. In one embodiment, the query processing coordinator 115 additionally sends the subquery to the relational data repository, which will act as an aggregator. Then, the relational node can forward the subquery to the Elastic node and use the results returned by the Elastic node to complete its own subquery.

[0080] Figure 5FIG. 0 is a block diagram showing a query processing coordinator 115 configured to route ontology queries according to one embodiment disclosed herein. Although depicted as a physical device, in an embodiment, the query processing coordinator 115 may be implemented using virtual device(s) and / or across multiple devices (e.g., in a cloud environment). As shown, the query processing coordinator 115 includes a processor 510, a memory 515, a storage device 520, a network interface 525, and one or more I / O interfaces 530. In the illustrated embodiment, the processor 510 retrieves and executes programming instructions stored in the memory 515, and stores and retrieves application data resident in the storage device 520. The processor 510 generally represents a single CPU and / or GPU, multiple CPUs and / or GPUs, a single CPU and / or GPU having multiple processing cores, etc. The memory 515 is typically included to represent random access memory. The storage device 520 may be any combination of disk drives, flash-based storage devices, etc., and may include fixed and / or removable storage devices such as fixed disk drives, removable memory cards, caches, optical storage devices, network-attached storage (NAS), or storage area network (SAN).

[0081] In some embodiments, input and output devices (such as keyboards, monitors, etc.) are coupled via the I / O interface(s) 530. Additionally, via the network interface 525, the query processing coordinator 115 may be communicatively coupled to one or more other devices and components (e.g., via the network 580, which may include the Internet, local network(s), etc.). As shown, the processor 510, the memory 515, the storage device 520, the network interface(s) 525, and the I / O interface(s) 530 are communicatively coupled by one or more buses 575. Additionally, the query processing coordinator 115 is communicatively coupled to a plurality of data nodes 130A-N via the network 580. Of course, in an embodiment, the data nodes 130A-N may be directly coupled to the query processing coordinator 115, accessible via a local network, integrated into the query processing coordinator 115, etc. Although not included in the illustrated embodiment, in some embodiments, the query processing coordinator 115 is also communicatively coupled to a data placement coordinator 150.

[0082] In the illustrated embodiment, the storage device 520 includes an ontology 560, capability data 565, and a data map 570. Although depicted as residing within the storage device 520, in an embodiment, the ontology 560, capability data 565, and data map 570 may be stored in any suitable location and manner. In an embodiment, as described above, the ontology 560 indicates entities or concepts related to the domain in which the query processing coordinator 115 operates, as well as potential attributes of each concept / entity and relationships between the entities / concepts. In at least one embodiment, the ontology 560 does not include instance-level data. Instead, the actual data in the KB is stored in data nodes 130A-N.

[0083] In an embodiment, as described above, the capability data 565 indicates the capabilities of each data node 130A-N, which may include, for example, an indication of which operation(s) each data node 130A-N supports and any corresponding limitations on that support. For example, the capability data 565 may indicate that data node 130A supports the "connect" operation, but only for integer data types. In an embodiment, as described above, the data map 570 indicates the (one or more) data nodes 130 that store instance data for each entity / concept defined in the storage ontology 560. For example, the data map 570 may indicate that all instances of the "Company" concept are stored in data nodes 130A and 130B, while all instances of the "Document" concept are stored in data nodes 130B and 130N.

[0084] In the illustrated embodiment, the memory 515 includes a query application 535. Although depicted as software residing within the memory 515, in an embodiment, the query application 535 may be implemented using hardware, software, or a combination of hardware and software. As shown, the query application 535 includes a parsing component 540, a routing component 545, and a collection 120 of (one or more) translators. Although depicted as discrete components for conceptual clarity, in an embodiment, the operations of the parsing component 540, routing component 545, and (one or more) translators 120 may be combined or distributed across any number of components.

[0085] In an embodiment, the parsing component 540 receives OQL queries, parses them to determine their meaning, and generates a logical representation of the query (such as a QGM). As described above, in some embodiments, the logical representation includes a set of query blocks, where each query block specifies one or more predicates that define the operations to be performed on the input quantifiers of the block. In an embodiment, these quantifiers can be (one or more) primitive concepts stored in one or more data nodes 130 and / or data computed by another query block. In one embodiment, this logical representation enables the query application 535 to understand and reason about the structure of the query and how data flows between the blocks. This enables the routing component 545 to efficiently route the query (or subqueries from the query).

[0086] In the illustrated embodiment, the routing component 545 receives the logical representation (e.g., a query graph model) from the parsing component 540 and evaluates it to identify one or more data nodes 130 to which the query should be routed. In some embodiments, this includes analyzing the predicates included in each query block in the logical representation. In one embodiment, the routing component 545 marks or annotates each predicate in a given query block based on the (one or more) data nodes 130 that can fulfill the predicate. In some embodiments, this includes the (one or more) relevant quantifiers and / or the (one or more) data nodes 130 that can perform the (one or more) indicated operations. Additionally, in at least one embodiment, once each predicate in the block has been processed, the routing component 545 can evaluate the annotations to determine the placement of the block. That is, the routing component 545 determines whether the query block can be placed in a single data node 130. If so, in one embodiment, the routing component 545 assigns the block to that repository. Additionally, in an embodiment, if the block cannot be processed by a single data node 130, the routing component 545 splits the block into two or more query blocks and repeats the routing process for each query block.

[0087] In an embodiment, once each query block has been assigned to a single data node 130, the (one or more) translators 120 generate one or more corresponding translated queries. In an embodiment, each translator 120 corresponds to a specific data node 130 architecture and is configured to translate an OQL query (or subquery) into an appropriate query for the corresponding data storage device architecture. For example, in one such embodiment, the system can include a first translator 120 that translates OQL into a query for a relational repository (such as SQL), a second translator 120 that translates OQL into a JSON query, a third translator 120 that translates into a query for document search, and a fourth translator 120 that translates OQL into a graph query.

[0088] In an embodiment, based on the data node 130 on which a query or subquery will be executed, the query (or subquery) is routed to the appropriate translator 120. In some embodiments, each query (or subquery) is then sent directly to that data node 130 (e.g., via an application programming interface or API). In some embodiments, if completing the query will require moving data between data nodes 130 (e.g., to join data from two or more repositories), the query application 535 may determine the data flow based on the generated logical representation and send the query appropriately. That is, the query application 535 uses QGM to determine which (if any) repositories will need to receive data from which other (if any) repositories and sends the required queries to these connected repositories. For example, if data node 130A is to receive data from data node 130N and perform one or more operations on it, the query application 535 may send a first query to data node 130A to retrieve data from that node, as well as a second query configured for data node 130N. Then, data node 130A can use the provided query to query data node 130N itself.

[0089] Figure 6 is a flowchart showing a method 600 for processing and routing ontology queries according to one embodiment disclosed herein. Method 600 begins at block 605, where the query processing coordinator 115 receives an ontology query for execution against a knowledge base. In an embodiment, the ontology query is expressed relative to a predefined domain ontology, where the ontology specifies concepts, attributes, and relationships related to the domain. At block 610, the query processing coordinator 115 parses the received query and generates its logical representation. In some embodiments, the query processing coordinator 115 generates a query graph pattern representation of the query. In an embodiment, the logical representation includes one or more query blocks representing the (one or more) basic concepts implied by the query, the (one or more) operations to be performed on the data, and the data flow.

[0090] Method 600 then proceeds to block 615, where query processing coordinator 115 maps each query block in the logical representation to one or more data repositories based on a predefined concept mapping and / or data storage capabilities. For example, in one embodiment, query processing coordinator 115 determines what concepts are relevant for each query block (e.g., what data the query block will access). Using the concept mapping, query processing coordinator 115 can then identify the data repository(ies) that can provide the required data. Additionally, in one embodiment, query processing coordinator 115 determines what operations will be required for each query block. Using the predefined storage capabilities, query processing coordinator 115 can then identify which data node(s) can perform the required operation(s). Query processing coordinator 115 then maps the query block(s) to the data node(s), striving to minimize data movement between repositories.

[0091] At block 620, query processing coordinator 115 selects one of the mapped query blocks. Method 600 then advances to block 625, where query processing coordinator 115 generates a corresponding translated query for the query block based on the mapped data repository. In one embodiment, this includes identifying or generating subqueries from the received query to perform the selected query block. Query processing coordinator 115 then determines the configuration of the data node that will execute the query block. That is, in one embodiment, query processing coordinator 115 determines the language and / or format of the queries that the node is configured to process. In another embodiment, query processing coordinator 115 identifies a translator corresponding to the mapped data repository. Query processing coordinator 115 then uses the translator to generate an appropriate query for the repository.

[0092] Method 600 then continues to block 630, where query processing coordinator 115 determines whether there is at least one additional query block that has not been processed. If so, method 600 returns to block 620. Otherwise, method 600 continues to block 635, where query processing coordinator 115 sends the one or more translated queries to one or more data nodes. In one embodiment, query processing coordinator 115 sends each query to the corresponding node. In another embodiment, query processing coordinator 115 sends the queries based on the data flow in the logical representation. For example, query processing coordinator 115 can send multiple queries to a single node so that the node can forward the queries to the appropriate storage(s) to retrieve data. The node can then act as a mediator to finalize the determination of the results.

[0093] Advantageously, by adopting one of the data nodes as a mediator, the query processing coordinator 115 can reduce its computational overhead and extend the scalability of the system. At block 640, the query processing coordinator 115 receives the finalized query result from the data repository acting as the mediator (or from the only data repository to which the received query is sent if the received query can be executed in a single backend resource). The query processing coordinator 115 then returns the result to the requesting entity.

[0094] Figure 7 FIG. is a flowchart of a method 700 for processing an ontology query to efficiently route query blocks according to one embodiment disclosed herein. In some embodiments, method 700 provides additional details for routing query blocks. Method 700 begins at block 705, where the query processing coordinator 115 generates one or more query blocks to represent the received query, as described above. In one embodiment, each query block specifies one or more quantifiers indicating the input data of the block. These quantifiers can correspond to basic concepts (e.g., instance data in a knowledge base) and / or intermediate computed data (e.g., data retrieved from a repository and processed and / or transformed in some way). Additionally, in an embodiment, each query block specifies an operation to be performed on the quantifiers.

[0095] At block 710, the query processing coordinator 115 selects one of the generated blocks. Additionally, at block 715, the query processing coordinator 115 identifies a set of data nodes capable of satisfying the operation(s) defined by the selected block. In one embodiment, the query processing coordinator 115 does so by accessing a predefined capabilities definition that defines the set of operations that each data repository can perform and any corresponding restrictions on that capability. Method 700 then proceeds to block 720, where the query processing coordinator 115 identifies a set of nodes capable of satisfying the quantifier(s) listed in the query block. That is, for each quantifier corresponding to a basic concept, the query processing coordinator 115 identifies the node(s) storing the basic concept.

[0096] In one embodiment, for each quantifier computed by another query block, the query processing coordinator 115 identifies the data node(s) that have been assigned to that query block. Recall that, in an embodiment, the query processing coordinator 115 maps query blocks to nodes by walking up from the bottom of the QGM. Thus, if the first block depends on data computed in the second block, the second block will necessarily have been evaluated and assigned one or more data nodes before the query processing coordinator 115 begins evaluating the first query block.

[0097] Method 700 then proceeds to block 725 where the query processing coordinator 115 determines whether there is at least one data node that can provide both the quantifiers and perform the operations. In one embodiment, this includes determining whether there is an overlap between the identified sets. For example, if the quantifiers are all basic concepts, an overlap between the two sets indicates a set of nodes that can perform all the required operations and store all the concepts. If the quantifiers include items computed by other query blocks, the overlapping set indicates nodes that can perform the operations and will (potentially) execute the other block(s).

[0098] In an embodiment, if there is an overlap between the sets, method 700 proceeds to block 730 where the query processing coordinator 115 annotates the selected query block with the identified nodes in the overlap. This annotation indicates that the specified nodes have the potential to execute the query block, but does not finalize the assignment of the block to the node(s). In an embodiment, if any query block has multiple nodes included in the annotation, the query processing coordinator 115 performs a cost analysis to evaluate each alternative in an attempt to minimize data transfer between repositories, as discussed in more detail below. Method 700 then proceeds to block 740.

[0099] In the illustrated embodiment, if there is no overlap between the sets, the query processing coordinator 115 determines that there is no single data node that can complete the query block. In some embodiments, if all the operations can be performed by a single repository, the query processing coordinator 115 annotates the block with the repository that can perform the operations of the block. Then, one or more other repositories can be assigned to provide the quantifiers. In the illustrated embodiment, method 700 then proceeds to block 735 where the query processing coordinator 115 splits the selected query block into two or more blocks to create query blocks that can be fully executed within a single node. In one embodiment, splitting the selected query block includes identifying one or more subsets of the quantifier(s) and / or operation(s) that can be performed by a single store. For example, the query processing coordinator 115 can identify subsets of the specified predicates that share an annotation (e.g., predicates that can be performed by the same repository). Then, the query block can be split by separating the predicates into corresponding boxes based on the subsets to which the predicates belong. For example, if the query block requires traditional relational database operations as well as fuzzy matching operations, the query processing coordinator 115 can determine that the fuzzy matching should be split into separate boxes. This can enable the query processing coordinator 115 to assign each split block to a single data store (e.g., the relational operations can be assigned to a relational store, and the fuzzy matching operations can be assigned to a different backend that can perform it).

[0100] In the illustrated embodiment, these newly generated query blocks are placed in a queue to be evaluated, just like existing blocks. Method 700 then proceeds to block 740. At block 740, query processing coordinator 115 determines whether there is at least one additional query block that has not been evaluated. If so, method 700 returns to block 710 to iterate through the blocks. That is, query processing coordinator 115 continues to iterate through the query blocks, splitting boxes if necessary, until all query blocks are annotated with at least one data node.

[0101] If all query blocks have been annotated, method 700 proceeds to block 745, where query processing coordinator 115 executes the query. In some embodiments, this includes evaluating alternative combinations of assignments to minimize data transfer, as discussed in more detail below. In some embodiments, if the annotation of one or more of the query blocks indicates a single data node, query processing coordinator 115 simply assigns the block to the indicated node, since there are no alternatives to be considered.

[0102] Figure 8 FIG. 8 is a flowchart showing a method 800 for evaluating potential query routing plans for efficiently routing ontology queries according to one embodiment disclosed herein. In the illustrated embodiment, method 800 begins after the query blocks have been annotated with their potential assignments (e.g., with nodes capable of executing the entire block and providing all required quantifiers). Method 800 begins at block 805, where query processing coordinator 115 selects one of the possible combinations of node assignments. That is, in an embodiment, query processing coordinator 115 generates all possible combinations of the repository of blocks based on the corresponding annotations. To do this, query processing coordinator 115 can iteratively select different options for each block until all possible selections have been generated.

[0103] For example, assume that the first block is annotated as "Node A" and "Node B", and the second block is annotated as "Node A". In an embodiment, query processing coordinator 115 will determine that the possible routing plans include assigning the first and second blocks to "Node A", or assigning the first block to "Node B" and the second block to "Node A". At block 805, query processing coordinator 115 selects one of the identified combinations for evaluation. Method 800 then proceeds to block 810.

[0104] At block 810, the query processing coordinator 115 identifies the data movement(s) that will be required under the selected plan. In embodiments, the movement is determined based on the repository selection and / or the query graph model in the plan. Continuing with the above example, if the query processing coordinator 115 determines that a plan to assign two chunks to "Node A" will not require movement because both query chunks can be grouped into a single node. That is, because the chunks are results in the model (e.g., directly joined / separated by a single hop with no other chunks in between) and are assigned to the same node, they are grouped together and no data movement is required. In contrast, assigning the first chunk to "Node B" will require data transfer between "Node A" and "Node B" because the QGM indicates that data flows from the second chunk to the first chunk, and the chunks are assigned to different repositories.

[0105] Method 800 then proceeds to block 815, where the query processing coordinator 115 selects one of the identified data transfers required by the selected query plan. At block 820, the query processing coordinator 115 determines the computational cost of the selected movement. In an embodiment, this determination is made based on a predefined cost model. In some embodiments, the model can indicate the computational cost of transferring data from a first node to a second node for each ordered data node pair. The cost can include, for example, the latency introduced by actually transferring the data and / or transforming the data on demand to allow the destination node to operate on it, the processing time and / or memory requirements for the transfer / transformation, etc.

[0106] At block 825, the query processing coordinator 115 determines whether the selected combination requires any additional data transfers. If so, method 800 returns to block 815. Otherwise, method 800 proceeds to block 830, where the query processing coordinator 115 calculates the total cost of the selected plan by aggregating the individual costs of each movement. At block 835, the query processing coordinator 115 determines whether there is at least one alternative plan that has not yet been evaluated. If so, method 800 returns to block 805. Otherwise, method 800 proceeds to block 840, where the query processing coordinator 115 ranks the plans based on the aggregated cost of the plans. In an embodiment, the query processing coordinator 115 selects the plan with the lowest determined cost. The query processing coordinator 115 then executes the plan by translating and routing subqueries based on the node assignments of the minimum cost plan.

[0107] Figure 9It is a flowchart showing a method 900 for routing ontology queries according to an embodiment disclosed herein. Method 900 begins at block 905, where a query processing coordinator 115 receives an ontology query. At block 910, the query processing coordinator 115 generates one or more query blocks based on the ontology query, each query block indicating one or more operations and one or more quantifiers representing the data flow between the query blocks. Method 900 then proceeds to block 915, where the query processing coordinator 115 identifies at least one data node for each of the one or more query blocks based on the one or more quantifiers and the one or more operations. Additionally, at block 920, the query processing coordinator 115 selects one or more of the identified data nodes based on a predefined cost criterion. Further, at block 925, the query processing coordinator 115 then sends one or more sub-queries to the selected one or more data nodes.

[0108] Figure 10 It shows a workflow 1000 for evaluating a workload and storing knowledge base data according to an embodiment disclosed herein. As described above, enterprise applications typically need to support different query types according to their query workloads. To support these different query types, embodiments of the present disclosure utilize multiple backend repositories, such as relational databases, document repositories, graph repositories, etc. In embodiments of the present disclosure, the system has the ability to move, store, and index knowledge base data in any backend that provides the required capabilities for the supported query types. With this flexibility in organizing data, the initial data placement across multiple backend repositories can play a key role in efficient query execution.

[0109] To achieve such efficient runtime execution with minimal replication overhead, some embodiments of the present disclosure provide an offline data preparation and loading phase, which includes intelligent data placement. Generally, data ingestion, placement, and loading into multiple backend repositories involve a series of operations. Initially, domain-specific knowledge base data is ingested from various different sources including structured, semi-structured, and unstructured data. In some embodiments, the first phase of data ingestion from these data sources is a data enrichment / curation process, which includes information extraction, entity resolution, data integration, and transformation. Then, the output data conforming to the domain ontology generated by this step is fed into a data placement coordinator 150 to generate an appropriate data placement for multiple backend data repositories. In some embodiments, the data has been managed and may have been used in response to queries before applying the data placement coordinator 150. Once a satisfactory data placement is determined, the data loading module places the instance data in the appropriate data repositories according to the data placement plan.

[0110] In some embodiments, by replicating the entire set of data across all data nodes, data movement can be avoided during query execution. However, this solution results in a huge replication overhead and significantly increases the storage space requirements. In addition, in many embodiments, not all repositories provide all the necessary capabilities required for querying, and even full replication cannot completely eliminate data movement. To minimize unnecessary storage costs and data movement, some embodiments of the present disclosure provide a capability-based data placement technique that assigns data to data repositories while considering the expected workload and the capabilities of the backend data repositories (e.g., in terms of the operations they can perform on the stored data).

[0111] In an embodiment, the data placement coordinator 150 reasons about data placement at the level of conceptually query operations of a domain ontology representing the data stored in the knowledge base. In some embodiments, the coordinator identifies different and potentially overlapping subsets of the ontology based on a given workload for the knowledge base and the capabilities of the underlying repositories, and outputs a mapping between the identified subsets of the data and the target data repositories in which the data should be stored.

[0112] In the illustrated embodiment, the workflow 1000 begins upon receipt of a set 1005 of OQL queries. In some embodiments, the OQL queries 1005 are queries previously submitted to the knowledge base. For example, in some embodiments, the knowledge base is a pre-existing data corpus that users and applications can query and explore. In such embodiments, the OQL queries 1005 can correspond to queries previously submitted by users, applications, and other entities when interacting with the knowledge base.

[0113] In an embodiment, the OQL queries 1005 generally represent the average, typical, expected, and / or historical workload of the knowledge base. In other words, the OQL queries 1005 typically represent the queries that the knowledge base receives (or expects to receive) during runtime operations and can be used to identify sets of concepts that are frequently queried together, the operations typically performed on each concept, etc. As shown, these representative OQL queries 1005 are provided to the OQL query analyzer 1010, which analyzes and evaluates them to generate a generalized workload 1015.

[0114] In an embodiment, the OQL query analyzer 1010 expresses the OQL query 1005 as a set of concepts and relationships with corresponding operations performed on them. Based on this, the analyzer generates a generalized workload 1015. In other words, the generalized workload 1015 reflects the generalization of the provided OQL query 1005, which enables a deeper analysis to identify patterns in the data. In one embodiment, the generalized workload 1015 includes a set of generalized queries generated based on the OQL query 1005. It is worth noting that in some embodiments, the generalized queries are not fully formatted queries that can be executed against the knowledge base. Instead, the generalized queries indicate clusters of concepts that may be queried together during runtime and a set of operations that may be applied to each cluster of concepts.

[0115] In some embodiments, the OQL query analyzer 1010 takes as input a set 1005 of OQL queries expressed against a domain ontology and, for each query, creates two sets. In one such embodiment, the first is the set of concepts accessed by the query, and the second is the set of operations (e.g., join, aggregate, etc.) that the query performs on those concepts. In at least one embodiment, to generate a representation of the generalized workload 1015 for a given OQL query 1005, the OQL query analyzer 1010 groups queries that access the same set of concepts into groups and then creates a set of associated operations that combine each query in the group. This will be discussed in more detail below.

[0116] In the illustrated embodiment, the generalized workload 1015 is then provided to the hypergraph modeler 1020, which evaluates the provided generalized queries to generate a hypergraph 1025. In an embodiment, the hypergraph 1025 is a graph that includes a set of vertices and a set of edges (also called hyperedges), where each edge can be connected to any number of vertices. In some embodiments, each vertex in the hypergraph 1025 corresponds to a concept from the domain ontology, and each hyperedge corresponds to a generalized query from the generalized workload 1015. For example, each hyperedge spans the set of concepts indicated by the corresponding generalized query. In one embodiment, each hyperedge is also annotated or labeled with an indication of the set of operations indicated by the corresponding generalized query.

[0117] As shown, the hypergraph 1025 is evaluated by the data placement component 1030 together with the ontology schema 110 and the node capabilities 155. In one embodiment, the data placement component 1030 groups the concepts and relationships in the hypergraph 1025 into potentially overlapping subsets based on query operations. Then, the data corresponding to these subsets can be placed on respective backend data nodes based on the operations they support. In some embodiments, data placement decisions are made at the granularity of the identified ontology subsets, placing all data for any ontology concept. In other words, the placement decision 165 does not horizontally split concepts between different repositories. For example, if the data placement component 1030 places the "company" concept in the first data node, all instance-level data corresponding to the "company" concept is placed in the first data node.

[0118] In some embodiments, the ontology-based data placement component 1030 follows a two-step approach for data placement. First, the data placement component 1030 runs a graph analysis algorithm on the hypergraph 1025 representing the generalized workload 1015 to group the concepts in the domain ontology schema 110 based on the similarity of the operations performed on these concepts. Next, the data corresponding to these identified groups or subsets is mapped to the underlying data repositories based on their respective capabilities while minimizing the amount of replication required. In an embodiment, the resulting capability-based data placement minimizes data movement (and data transformation) for a given workload at query processing time, thus greatly enhancing the efficiency of query processing in a multi-repository environment. As shown, the final output of the data placement component 1030 is a concept-to-store mapping (e.g., the placement decision 165) that maps ontology concepts to appropriate data nodes.

[0119] Although not depicted in the shown workflow 1000, in an embodiment, the system then uses the placement decision 165 to store data in various backend resources. In one embodiment, the system does this by invoking an extract, transform, and load (ETL) service that performs any transformations or conversions necessary to allow the data to be stored in the associated data node 130.

[0120] In some embodiments, the workflow 1000 is used to periodically re-evaluate and refine the data placement to maintain the efficiency of the system. For example, during off-peak time (e.g., during non-business hours), the system can invoke the workflow 1000 based on a set of updated OQL queries 1005 (e.g., including queries received after the last data placement decision) to determine whether the placement should be updated to reflect the evolving workload. This can improve the efficacy of the system by preventing the data locations from becoming stale.

[0121] Figure 11 FIG. 1100 depicts an example hypergraph for evaluating a workload and storing knowledge base data according to one embodiment disclosed herein. In the described embodiment, each concept 1105A-G is depicted as an ellipse, and each hyperedge 1110A-D is depicted as a dashed line enclosing its corresponding concept 1105. For example, hyperedge 1110A connects concepts 1105A (“Company”), 1105B (“PublicCompany”), 1105C (“PublicMetric”), and 1105D (“PublicMetricData”). As shown, each hyperedge 1110 may include a non-connected subset of concepts 1105 (e.g., hyperedges 1110A and 1110C do not overlap) as well as an overlapping subset (e.g., hyperedges 1110A and 1110B overlap with respect to the “Company” concept 1105A).

[0122] Additionally, as shown, each hyperedge 1110 is labeled with the associated operations 1115A-C of that edge. As described above, in one embodiment, the label indicates a set of operations that can or have been applied to the concepts 1105 joined by the hyperedge 1110. For example, in the depicted example, hyperedge 1110A is associated with operations 1115A (“Join”), 1115B (“Aggregate”), and 1115C (“Fuzzy” match). It is noted that each operation 1115 may be associated with any number of hyperedges 1110 and / or concepts 1105. In the described embodiment, the “Join” operation 1115A is associated with hyperedges 1110A and 1110B, the “Aggregate” operation 1115B is associated with hyperedges 1110A and 1110D, and the “Fuzzy” operation 1115C is associated with hyperedges 1110A, 1110D, and 1110C.

[0123] Figure 12FIG. 0 is a block diagram of a data placement coordinator 150 configured to evaluate a workload and store data in accordance with an embodiment disclosed herein. Although depicted as a physical device, in an embodiment, the data placement coordinator 150 may be implemented using one or more virtual devices and / or across multiple devices (e.g., in a cloud environment). As shown, the data placement coordinator 150 includes a processor 1210, a memory 1215, a storage device 1220, a network interface 1225, and one or more I / O interfaces 1230. In the illustrated embodiment, the processor 1210 retrieves and executes programming instructions stored in the memory 1215, as well as stores and retrieves application data resident in the storage device 1120. The processor 1210 generally represents a single CPU and / or GPU, multiple CPUs and / or GPUs, a single CPU and / or GPU having multiple processing cores, etc. The memory 1215 is generally included to represent random access memory. The storage device 1220 may be any combination of disk drives, flash-based storage devices, etc., and may include fixed and / or removable storage devices, such as fixed disk drives, removable memory cards, caches, optical storage devices, network-attached storage devices (NAS), or storage area networks (SAN).

[0124] In some embodiments, input and output devices (such as a keyboard, monitor, etc.) are coupled via one or more I / O interfaces 1230. Additionally, via the network interface 1225, the data placement coordinator 150 may be communicatively coupled to one or more other devices and components (e.g., via a network 1280, which may include the Internet, one or more local networks, etc.). As shown, the processor 1210, the memory 1215, the storage device 1220, one or more network interfaces 1225, and one or more I / O interfaces 1230 are communicatively coupled via one or more buses 1275. Additionally, the data placement coordinator 150 is communicatively coupled to data nodes 130A-N via the network 1280. Of course, in an embodiment, the data nodes 130A-N may be directly coupled to the data placement coordinator 150, accessible via a local network, integrated into the data placement coordinator 150, etc. Although not included in the illustrated embodiment, in some embodiments, the data placement coordinator 150 is also communicatively coupled to a query processing coordinator 115.

[0125] In the illustrated embodiment, the storage device 1220 includes a copy of the domain ontology 1260, the capabilities data 1265, and the data map 1270. Although depicted as residing in the storage device 1220, in an embodiment, the ontology 1260, the capabilities data 1265, and the data map 1270 may be stored in any suitable location and manner. In one embodiment, the ontology 1260, the capabilities data 1265, and the data map 1270 correspond to the ontology 560, the capabilities data 565, and the data map 570 discussed above with reference to the query processing coordinator 115.

[0126] For example, as described above, the ontology 1260 may indicate entities or concepts related to the domain, as well as potential attributes of each concept / entity and relationships between the entities / concepts, excluding instance-level data. Similarly, as described above, the capabilities data 1265 indicates the capabilities of each data node 130A-N, which may include, for example, an indication of which operations each data node 130A-N supports and any corresponding limitations on that support.

[0127] In an embodiment, the data map 1270 is generated by the placement application 1235 and indicates the data node(s) 130 in which instance data is stored for each entity / concept defined in the ontology 1260. For example, the data map 1270 may indicate that all instances of the "Company" concept are stored in data nodes 130A and 130B, while all instances of the "Document" concept are stored in data nodes 130B and 130N.

[0128] In the illustrated embodiment, the memory 1215 includes the placement application 1235. Although depicted as software residing in the memory 1215, in an embodiment, the placement application 1235 may be implemented using hardware, software, or a combination of hardware and software. As shown, the placement application 1235 includes a generalization component 1240, a hypergraph component 1245, and a placement component 1250. Although depicted as discrete components for conceptual clarity, in various embodiments, the operations of the generalization component 1240, the hypergraph component 1245, and the placement component 1250 may be combined or distributed across any number of components.

[0129] In an embodiment, the overview component 1240 receives previous OQL queries (or sample queries) that reflect the previous, current, expected, typical, average, anticipated, and / or general workload on the system. The generalization component 1240 then evaluates the provided queries to generalize the workload in the form of a generalized query. In an embodiment, the generalized query generally indicates the concepts that are queried together and the set of operations performed on the concepts, but does not correspond to the actual queries that can be executed. The functionality of the generalization component 1240 is discussed in more detail below with reference to Figure 13 more detail.

[0130] In an embodiment, the hypergraph component 1245 receives the generalized workload information from the generalization component 1240 and parses it to generate a hypergraph representing the query workload. As described above, in some embodiments, the hypergraph includes vertices for each concept in the ontology 1260. Additionally, in an embodiment, the concept vertices are joined by hyperedges representing the generalized query workload. The operation of the hypergraph component 1245 will be discussed in more detail below with reference to Figure 14 and

[0131] In one embodiment, the placement component 1250 evaluates the hypergraph to generate data placement decisions (e.g., data mapping 1270). These decisions can then be used to route data to the appropriate nodes. For example, in one embodiment, the system iterates through the knowledge base to place data. For each data item, the system can determine its corresponding concept(s), look up the corresponding data node 130, and store the data in the indicated node(s). In some embodiments, the system also performs any transformations or conversions suitable for the destination repository. The functionality of the placement component 1250 will be discussed in more detail below with reference to Figure 15 and Figure 16 and

[0132] Figure 13 is a flow chart showing a method 1300 for evaluating and generalizing a query workload to inform data placement decisions according to an embodiment disclosed herein. Method 1300 begins at block 1305, where the data placement coordinator 150 receives one or more queries representing the knowledge base workload. At block 1310, the data placement coordinator 150 selects one of the received queries for evaluation. Method 1300 then proceeds to block 1315, where the data placement coordinator 150 identifies the concept(s) involved by the selected query. Similarly, at block 1320, the data placement coordinator 150 identifies the operation(s) invoked by the query.

[0133] In one embodiment, the data placement coordinator 150 associates the set of concepts specified by the query that it has determined with the set of operation(s) invoked by the query. In some embodiments, the operation-concept pairing is determined on a granular concept / operation level basis. In some other embodiments, the data placement coordinator 150 does not determine which operations are linked to which concepts, but rather links the set of operations as a whole to the set of concepts. That is, the data placement coordinator 150 ignores the actual execution details / desired results of the query and defines the set of related concepts and operations based on the data / operations involved rather than the specific application of the transformation.

[0134] Method 1300 then proceeds to block 1325 where data placement coordinator 150 determines whether there is at least one additional query to be evaluated. If so, method 1300 returns to block 1310. Otherwise, method 1300 proceeds to block 1330. In block 1330, data placement coordinator 150 groups the received queries based on the corresponding set of concepts for each. In one embodiment, this includes grouping or clustering all queries that specify a set of concepts for a match. For example, data placement coordinator 150 can group all queries that specify both the "Company" concept and the "PublicMetric" concept into a first group. It is noted that in such an embodiment, queries that specify only the "Company" concept or the "PublicMetric" concept will be placed in different groups. Similarly, queries that specify "Company" and "PublicMetric" but also include the "Document" concept will not be placed in the first group.

[0135] That is, in an embodiment, data placement coordinator 150 groups queries that specify a set of concepts for an exact match. Queries that specify additional concepts or fewer concepts are placed in other groups. In one embodiment, the set of concepts corresponding to each respective grouping is used to form a corresponding generalized query. That is, in the illustrated embodiment, data placement coordinator 150 aggregates the previously received queries by grouping queries that access the same concepts into clusters, regardless of the operations performed by each query.

[0136] Then, method 1300 proceeds to block 1335 where data placement coordinator 150 selects one of these defined query groups. In block 1340, data placement coordinator 150 selects one of the queries associated with the selected group. Additionally, in block 1345, data placement coordinator 150 associates the corresponding operation specified by the selected query with the generalized query representing the selected query group. In one embodiment, if the operation is already reflected in the generalized query, data placement coordinator 150 does not add it again. That is, in an embodiment, the set of operations associated with the generalized query is a binary value indicating the presence or absence of a given operation.

[0137] In an embodiment, the operation associated with the generalized query can be at any level of granularity. For example, in some embodiments, data placement coordinator 150 defines the generalized query at the operation level (e.g., "join") without considering the specific type of the operation (e.g., inner join), the restrictions on the operation, or the details of the operation (e.g., the type of data, such as string join or integer join). In other embodiments, these details are included in the generalized query description to provide more rich details for subsequent processing.

[0138] Method 1300 then proceeds to block 1350, where data placement coordinator 150 determines whether there is at least one additional query in the selected group. If so, method 1300 returns to block 1340. If not, at block 1355, data placement coordinator 150 determines whether there is at least one additional query group that has not yet been evaluated. If so, method 1300 returns to block 1335. Otherwise, method 1300 proceeds to block 1360. At block 1360, data placement coordinator 150 stores the summarized workload such that it can be used and evaluated to generate a hypergraph.

[0139] Figure 14 FIG. is a flowchart showing method 1400 for modeling an ontology workload to inform data placement decisions according to one embodiment disclosed herein. Method 1400 begins at block 1405, where data placement coordinator 150 selects one of the concepts of the ontology. At block 1410, data placement coordinator 150 generates a hypergraph vertex for the selected concept. Method 1400 then proceeds to block 1415, where data placement coordinator 150 determines whether there are any additional concepts remaining in the ontology that do not yet have vertices in the hypergraph. If so, method 1400 returns to block 1405. Otherwise, method 1400 proceeds to block 1420.

[0140] At block 1420, data placement coordinator 150 selects one of the summarized queries. As discussed above, in an embodiment, each summarized query indicates a set of concepts and a corresponding set of operations that have been applied to the concepts in a previous workload. At block 1425, data placement coordinator 150 generates a hyperedge that links each of the (one or more) concepts indicated by the selected summarized query. Method 1400 then proceeds to block 1430, where data placement coordinator 150 labels the newly generated hyperedge with an indication of the operation indicated by the selected summarized query. In this way, data placement coordinator 150 can subsequently evaluate the hypergraph to identify relationships and usage patterns in the knowledge base.

[0141] Method 1400 then proceeds to block 1435, where data placement coordinator 150 determines whether there is at least one additional summarized query to be evaluated and incorporated into the hypergraph. If so, method 1400 returns to block 1420. Otherwise, method 1400 proceeds to block 1440, where data placement coordinator 150 stores the generated hypergraph for subsequent use.

[0142] Figure 15 and Figure 16 depicts a flowchart showing a method for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein. Referring Figure 15The method 1500 discussed illustrates one embodiment of an operation-based clustering technique for generating a concept map, while the method 1600 discussed below Figure 16 illustrates one embodiment of a minimum coverage technique.

[0143] In one embodiment, the operation-based clustering technique groups concepts based on the operations that the concepts undergo. In one such embodiment, for each operation description in the hypergraph, the data placement coordinator 150 creates a corresponding cluster. Then, the data placement coordinator 150 iterates over the group of operation descriptions associated with each hyperedge, and for each such operation description, assigns all the concepts spanned by the hyperedge to the cluster corresponding to the operation. In an embodiment, once the concepts are clustered together, the data placement coordinator 150 assigns each concept cluster to a set of data nodes such that each node has a capability description that matches the operation description of the cluster (e.g., is able to perform the corresponding operation of the cluster). Finally, in an operation-based system, the data placement coordinator 150 generates a map that maps each concept in each cluster to the corresponding set of identified data repositories.

[0144] As an example of one embodiment of the operation-based clustering technique, consider Figure 11 the hypergraph 1100 provided in. Initially, the data placement coordinator 150 generates clusters for each operation 1115A-C (e.g., C Join , C Agg and C Fuzzy ). Then, the system determines the set of operations associated with each hyperedge. For each indicated operation, the system assigns the concepts specified by the hyperedge 1110 to the corresponding cluster. Continuing with the above example, C Join will contain the concepts 1105A, 1105B, 110C, and 1105D from hyperedge 1110A, as well as the concepts 1105E and 1105F from hyperedge 1110B. In addition, C Agg will contain the concepts 1105A, 1105B, 1105C, and 1105D from hyperedge 1110A, as well as the concept 1105G from hyperedge 1110D. Finally, C Fuzzy will contain the concepts 1105A, 1105B, 1105C, and 1105D from hyperedge 1110A, as well as the concept 1105H from hyperedge 1110C.

[0145] In many embodiments, these operation-based clusters have significant overlap. For example, note that concepts 1105A, 1105B, 1105C, and 1105D are included in each cluster. To finalize the mapping, in one embodiment, data placement coordinator 150 identifies all data nodes 130 that are capable of performing the corresponding operations for each cluster. Then, data placement coordinator 150 maps all concepts 1105 in the cluster to all identified data nodes 130. In some embodiments, although operation-based techniques can minimize or reduce data movement during query processing by placing data in all stores that support the corresponding operations, it does introduce some replication overhead because clusters of the same concept can be placed at multiple repositories if they have the ability to satisfy the operations of that cluster.

[0146] In some implementations, to further reduce replication overhead, embodiments that utilize the minimum cover technique are employed. In one embodiment, the minimum cover embodiment improves upon operation-based techniques by further minimizing the amount of data replication while still minimizing data movement during query processing. In an embodiment, the technique utilizes a minimum set cover algorithm to find the minimum number of data repositories required to support the complete set of operations for each hyperedge in the query workload hypergraph. In some embodiments, the minimum cover technique minimizes the span of each hyperedge over the set of data repositories that satisfy the set of operations required for the hyperedge.

[0147] In an example embodiment of the minimum cover technique, data placement coordinator 150 finds the minimum number of data nodes that cover all the indicated operations for each hyperedge in the hypergraph. For example, if all operations can be completed by a single data node 130A, the minimum set includes only that node. If data node 130A cannot complete one or more operations, one or more other data nodes 130B-N are added to the minimum set until all operations are satisfied. Once the minimum set is determined for a hyperedge, each concept in the hyperedge is mapped to each node within the corresponding minimum set.

[0148] As an example of an embodiment of the minimum cover clustering technique, consider Figure 11The hypergraph 1100 provided in [description] assumes that the data node 130 includes a first data node 130A configured to support operations 1115A and 1115B, a second data node 130B configured to support operations 1115B and 1115C, and a third data node 130C that only supports operation 1115B. For hyperedge 1110A, the data placement coordinator 150 can determine that no single node can support all three indicated operations, but the set of data nodes 130A and 130B can support all three indicated operations. Since these two nodes can support the entire hyperedge, there is no need to add data node 130C to the set.

[0149] Similarly, hyperedge 1110B will be assigned to data node 130A (the only node configured to provide the join operation), and hyperedge 1110C will be assigned to data node 130B (the only node configured to provide the fuzzy match). Finally, hyperedge 1110D can be assigned to data nodes 130A and 130B, or data nodes 130B and 130C. In some embodiments, the data placement coordinator 150 selects between these additional equivalent alternatives based on other criteria, such as the latency or computational resources of each, predefined preferences, and the like. Then, the data placement coordinator 150 maps the concept 1105 of each hyperedge 1110 to the assigned data node(s) 130.

[0150] Figure 15 is a flowchart showing an operator-based method 1500 for evaluating a hypergraph to drive data placement decisions according to an embodiment disclosed herein. Method 1500 begins at block 1505, where the data placement coordinator 150 selects one of the operations indicated by the hypergraph. At block 1510, the data placement coordinator 150 generates a cluster for the selected operation. Method 1500 then proceeds to block 1515, where the data placement coordinator 150 determines whether there is at least one additional operation reflected in the hypergraph that does not yet have a cluster associated with it. If so, method 1500 returns to block 1505. Otherwise, method 1500 continues to block 1520.

[0151] At block 1520, the data placement coordinator 150 selects one of the hyperedges in the hypergraph for analysis. At block 1525, the data placement coordinator 150 identifies the concept(s) and operation(s) associated with the selected edge. The method 1500 then proceeds to block 1530, where the data placement coordinator 150 selects one of the indicated operations. Additionally, at block 1535, the data placement coordinator 150 identifies the corresponding cluster for the selected operation and adds all the concepts represented by the selected hyperedge to that cluster. The method 1500 proceeds to block 1540, where the data placement coordinator 150 determines whether the selected edge indicates at least one additional operation to be processed. If so, the method 1500 returns to block 1530.

[0152] If no additional operations are associated with the selected hyperedge, the method 1500 proceeds to block 1545, where the data placement coordinator 150 determines whether the hypergraph includes at least one additional edge that has not been evaluated. If so, the method 1500 returns to block 1520. Otherwise, the method 1500 proceeds to block 1550, where the data placement coordinator 150 maps the operation(s) cluster(s) to the corresponding data node(s) configured to process each operation. For example, in one embodiment, the data placement coordinator 150 identifies a set of data nodes capable of performing the corresponding operation for each cluster. In an embodiment, the data placement coordinator 150 then maps each concept in the cluster to each node in the set of the identified data node(s). The data placement coordinator 150 can use these mappings to distribute the data in the knowledge base among various data nodes.

[0153] Figure 16 FIG. is a flowchart showing a minimum-coverage-based method 1600 for evaluating a hypergraph to drive data placement decisions according to one embodiment disclosed herein. The method 1600 begins at block 1605, where the data placement coordinator 150 selects one of the hyperedges in the hypergraph. At block 1610, the data placement coordinator 150 identifies the corresponding concepts and operations associated with the selected edge. Additionally, at block 1615, the data placement coordinator 150 determines the minimum set of data nodes that can jointly satisfy all the indicated operations.

[0154] In one embodiment, the data placement coordinator 150 does this by iteratively evaluating each combination of data nodes to determine whether the combination satisfies the indicated operations. That is, whether each operation indicated by the selected edge can be performed by at least one data node in the combination. If not, the combination can be discarded and another node or combination can be selected for testing (or another node can be added to the current combination). The data placement coordinator 150 can determine whether each combination is complete and satisfactory as it can perform all the required operations, and identify the combination(s) with the least number of data nodes as these combinations will likely result in the least amount of data movement during runtime. In one embodiment, if two or more combinations are equally small, the data placement coordinator 150 can use predefined criteria or preferences to select the best combination. For example, the predefined rule can indicate that a combination including at least one relational data repository is preferred over a combination without a relational data repository. As another example, the rule can indicate the weight or priority of the (one or more) specific types of repositories and / or the (one or more) specific data nodes 130. In such an embodiment, the data placement coordinator 150 can aggregate or otherwise combine these weights for each combination to determine which repositories to use. Then, the selected edges are labeled with an indication of the set of data nodes determined.

[0155] Once the minimum set of data nodes is determined, method 1600 proceeds to block 1620, where the data placement coordinator 150 determines whether there is at least one additional hyperedge in the hypergraph. If so, method 1600 returns to block 1605. Otherwise, method 1600 proceeds to block 1625. At block 1625, the data placement coordinator 150 selects one of the available data nodes in the system. At block 1630, the data placement coordinator 150 identifies all hyperedges in the hypergraph that have a label including the selected data node. The data placement coordinator 150 then groups or clusters these hyperedges (or each included concept) together to form a group / cluster of concepts that will be stored in the selected node. Then, method 1600 proceeds to block 1635, where the data placement coordinator 150 determines whether there is at least one additional data node in the system that has not been assigned a group / cluster. If so, method 1600 returns to block 1625.

[0156] Otherwise, method 1600 proceeds to block 1640, where the data placement coordinator 150 maps all concepts included in the corresponding cluster to repositories for each data node. The data placement coordinator 150 can then use these mappings to distribute data across various data nodes in the knowledge base.

[0157] Figure 17It is a flowchart showing a method 1700 for mapping ontology concepts to storage nodes according to an embodiment disclosed herein. Method 1700 begins at block 1705, where the data placement coordinator 150 determines query workload information corresponding to a domain. At block 1710, the data placement coordinator 150 models the query workload information as a hypergraph, where the hypergraph includes a set of vertices and a set of hyperedges, and each vertex in the set of vertices corresponds to a concept in an ontology associated with the domain. Method 1700 then proceeds to block 1715, where the data placement coordinator 150 generates a mapping between the concepts and the plurality of data nodes based on the hypergraph and further based on predefined capabilities of each of the plurality of data nodes. Additionally, at block 1720, the data placement coordinator 150 establishes a distributed knowledge base based on the generated mapping.

[0158] The description of various embodiments of the present disclosure has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or improvements made to the technology found in the marketplace, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.

[0159] In the foregoing and / or following, reference is made to embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the specifically described embodiments. Instead, any combination of the foregoing and / or following features and elements, whether or not relating to different embodiments, is contemplated for implementing and practicing the contemplated embodiments. Moreover, although the embodiments disclosed herein may achieve advantages over other possible solutions or the prior art, whether a given embodiment achieves a particular advantage does not limit the scope of the present disclosure. Accordingly, the foregoing and / or following aspects, features, embodiments, and advantages are illustrative only and are not to be considered elements or limitations of the appended claims, unless expressly recited in the claims. Likewise, references to "the invention" should not be construed as generalizing any inventive subject matter disclosed herein and should not be considered an element or limitation of the appended claims, unless expressly recited in the claims.

[0160] Aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, which may all be collectively referred to herein as "circuitry", "module" or "system".

[0161] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.

[0162] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0163] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0164] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, in order to perform aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.

[0165] Aspects of the present invention are described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0166] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0167] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative embodiments, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0169] Embodiments of the present invention may be provided to end users via a cloud computing infrastructure. Cloud computing generally refers to the provision of scalable computing resources as a service over a network. More formally, cloud computing can be defined as the computing capability that provides an abstraction between computing resources and their underlying technical architecture (e.g., servers, storage devices, networks), enabling convenient on-demand network access to a shared pool of configurable computing resources that can be rapidly provisioned and released with minimal management effort or service provider interaction. Thus, cloud computing allows users to access virtual computing resources (e.g., storage, data, applications, even complete virtualized computing systems) in the "cloud" without regard to the underlying physical systems (or the location of those systems) used to provide the computing resources.

[0170] Typically, cloud computing resources are provided to users on a pay-per-use basis, where the user is charged only for the computing resources actually used (e.g., the amount of storage space consumed by the user or the number of virtualized systems instantiated by the user). The user can access any resources residing in the cloud at any time and from anywhere on the Internet. In the context of the present invention, the user can access the applications (e.g., query processing coordinator 115) or related data available in the cloud. For example, the query processing coordinator 115 may execute on a computing system in the cloud and evaluate queries and backend resources. In such a case, the query processing coordinator 115 may route queries and store the backend resources and / or capabilities configuration at a storage location in the cloud. Doing so allows the user to access this information from any computing system attached to a network (e.g., the Internet) that is coupled to the cloud.

[0171] While the foregoing is directed to embodiments of the present invention, other and further embodiments of the present invention may be devised without departing from the basic scope thereof, and the scope of the present invention is determined by the appended claims.

Claims

1. A method, comprising: Determining, by a data coordinator, query workload information corresponding to a domain, the query workload information including an ontology query specifying a corresponding concept in an ontology associated with the domain and a corresponding operation performed by the ontology query, the ontology query including a first set having corresponding matching concepts that are accessed, and an aggregated set of operations performed by the first set; Modeling the query workload information as a hypergraph, where the hypergraph includes a plurality of vertices and a plurality of hyperedges, where each corresponding vertex among the plurality of vertices corresponds to a corresponding concept in the ontology, and each corresponding hyperedge among the plurality of hyperedges indicates a corresponding set of operations applied to the concept associated with the corresponding hyperedge, where a first hyperedge among the plurality of hyperedges is labeled with the aggregated set of operations and connects a first vertex representing a matching concept; Generating, based on the hypergraph and further based on predefined capabilities of each of a plurality of data stores, a mapping between concepts in the ontology and the plurality of data stores, where the mapping indicates, for each corresponding concept in the ontology, a corresponding subset of the plurality of data stores on which the corresponding concept will be stored; And Establishing a distributed knowledge base based on the generated mapping.

2. The method according to claim 1, wherein determining the query workload information includes: Receiving a set of previous ontology queries; Generating a first set of concepts accessed by a first query in the set of previous ontology queries; Generating a first set of operations performed by the first query; And Generating a first generalized query by the following steps: Identifying, from the set of previous ontology queries, a first group of queries having a corresponding matching first set; Determining an aggregated set of operations based on the corresponding sets of operations of each query in the first group of queries; And Associating the first generalized query with the aggregated set of operations and the concepts reflected in the corresponding matching first set.

3. The method according to claim 2, wherein modeling the query workload information as a hypergraph includes: Creating a vertex for each concept in the ontology; Creating the first hyperedge for the first generalized query, where the first hyperedge connects a first set of vertices in the hypergraph, where the first set of vertices corresponds to the concepts reflected in the matching first set; and Labeling the first hyperedge with the aggregated set of operations.

4. The method according to claim 1, wherein generating the mapping includes: Creating a first cluster for a first operation included in the hypergraph; Identifying a first set of concepts connected by a first hyperedge in the hypergraph; Identifying a first set of operations indicated by the first hyperedge; And Assigning the first set of concepts to the first cluster when determining that the first set of operations includes the first operation.

5. The method according to claim 4, wherein generating the mapping further includes mapping the first set of concepts to one or more data stores by the following steps: Identifying a set of data stores capable of performing the first operation; and Mapping each concept in the first set of concepts to each data store in the identified set of data stores.

6. The method according to claim 1, wherein generating the mapping comprises: identifying a first set of concepts connected by a first hyperedge in the hypergraph; identifying a first set of operations indicated by the first hyperedge; determining a minimum set of data stores capable of jointly performing the first set of operations; generating a cluster comprising the first set of concepts; and labeling the cluster with the minimum set of data stores.

7. The method according to claim 6, wherein generating the mapping further comprises mapping each concept in the first set of concepts to each data store in the minimum set of data stores.

8. The method according to claim 1, wherein establishing the distributed knowledge base comprises, for each corresponding concept in the ontology: identifying the corresponding data store indicated by the mapping; identifying data corresponding to the corresponding concept; and facilitating storage of the identified data in the corresponding data store.

9. A computer-readable storage medium comprising computer program code which, when executed by the operation of one or more computer processors, performs an operation comprising the steps of: determining, by a data coordinator, query workload information corresponding to a domain, the query workload information comprising an ontology query specifying a corresponding concept in an ontology associated with the domain and corresponding operations performed by the ontology query, the ontology query comprising a first set having corresponding matching concepts accessed, the first set performing an aggregated set of operations; modeling the query workload information as a hypergraph, wherein the hypergraph comprises a plurality of vertices and a plurality of hyperedges, wherein each corresponding vertex in the plurality of vertices corresponds to a corresponding concept in the ontology, and each corresponding hyperedge in the plurality of hyperedges indicates a corresponding set of operations applied to the concept associated with the corresponding hyperedge, wherein a first hyperedge in the plurality of hyperedges is labeled with the aggregated set of operations and connects a first vertex representing a matching concept; generating a mapping between the concepts in the ontology and the plurality of data stores based on the hypergraph and further based on predefined capabilities of each of the plurality of data stores, the mapping indicating, for each corresponding concept in the ontology, a corresponding subset of the plurality of data stores on which the corresponding concept will be stored; and establishing a distributed knowledge base based on the generated mapping.

10. The computer-readable storage medium according to claim 9, wherein determining the query workload information comprises: receiving a set of previous ontology queries; generating a first set of concepts accessed by a first query in the set of previous ontology queries; generating a first set of operations performed by the first query; and generating a first generalized query by the following steps: identifying a first group of queries having a corresponding matching first set from the set of previous ontology queries; determining an aggregated set of operations based on the corresponding sets of operations of each query in the first group of queries; and associating the first generalized query with the aggregated set of operations and the concepts reflected in the corresponding matching first set.

11. The computer-readable storage medium according to claim 10, wherein modeling the query workload information as a hypergraph includes: creating vertices for each concept in the ontology; creating the first hyperedge for the first generalized query, wherein the first hyperedge connects a first set of vertices in the hypergraph, and the first set of vertices corresponds to the concepts reflected in the matching first set; and labeling the first hyperedge with an aggregated set of operations.

12. The computer-readable storage medium according to claim 9, wherein generating the mapping includes: creating a first cluster for a first operation included in the hypergraph; identifying a first set of concepts connected by a first hyperedge in the hypergraph; identifying a first set of operations indicated by the first hyperedge; and when determining that the first set of operations includes the first operation, assigning the first set of concepts to the first cluster.

13. The computer-readable storage medium according to claim 12, wherein generating the mapping further includes mapping the first set of concepts to one or more data stores by the following steps: identifying a set of data stores capable of performing the first operation; and mapping each concept in the first set of concepts to each data store in the identified set of data stores.

14. The computer-readable storage medium according to claim 9, wherein generating the mapping includes: identifying a first set of concepts connected by a first hyperedge in the hypergraph; identifying a first set of operations indicated by the first hyperedge; determining a minimum set of data stores capable of jointly performing the first set of operations; generating a cluster including the first set of concepts; and labeling the cluster with the minimum set of data stores.

15. The computer-readable storage medium according to claim 9, wherein establishing the distributed knowledge base includes, for each corresponding concept in the ontology: identifying the corresponding data store indicated by the mapping; identifying the data corresponding to the corresponding concept; and facilitating storing the identified data in the corresponding data store.

16. A system, including: One or more computer processors; and a memory containing a program that performs operations when executed by one or more computer processors, the operations including: determining, by a data coordinator, query workload information corresponding to a domain, the query workload information including an ontology query specifying a corresponding concept in an ontology associated with the domain and a corresponding operation performed by the ontology query, the ontology query including a first group having a corresponding matching concept to be accessed, the first group performing an aggregated set of operations; modeling the query workload information as a hypergraph, wherein the hypergraph includes a plurality of vertices and a plurality of hyperedges, each corresponding vertex in the plurality of vertices corresponds to a corresponding concept in the ontology, and each corresponding hyperedge in the plurality of hyperedges indicates a corresponding set of operations applied to the concept associated with the corresponding hyperedge, wherein a first hyperedge in the plurality of hyperedges is labeled with an aggregated set of operations and connects a first vertex representing a matching concept; Generate a mapping between the concepts in the ontology and the multiple data stores based on the hypergraph and further based on predefined capabilities of each of the multiple data stores, wherein the mapping indicates, for each corresponding concept in the ontology, the corresponding subset of the multiple data stores on which the corresponding concept will be stored; and Establish a distributed knowledge base based on the generated mapping.

17. The system according to claim 16, wherein determining the query workload information includes: Receiving a set of previous ontology queries; Generating a first set of concepts accessed by a first query in the set of previous ontology queries; Generating a first set of operations performed by the first query; And Generating a first generalized query by the following steps: Identifying, from the set of previous ontology queries, a first group of queries having a corresponding matching first set; Determining an aggregated set of operations based on the corresponding sets of operations of each query in the first group of queries; And Associating the first generalized query with the aggregated set of operations and the concepts reflected in the corresponding matching first set.

18. The system according to claim 17, wherein modeling the query workload information as a hypergraph includes: Creating a vertex for each concept in the ontology; Creating the first hyperedge for the first generalized query, wherein the first hyperedge connects a first set of vertices in the hypergraph, wherein the first set of vertices corresponds to the concepts reflected in the matching first set; and Labeling the first hyperedge with the aggregated set of operations.

19. The system according to claim 16, wherein generating the mapping includes: Creating a first cluster for a first operation included in the hypergraph; Identifying a first set of concepts connected by a first hyperedge in the hypergraph; Identifying a first set of operations indicated by the first hyperedge; And Assigning the first set of concepts to the first cluster when determining that the first set of operations includes the first operation.

20. The system according to claim 16, wherein generating the mapping includes: Identifying a first set of concepts connected by a first hyperedge in the hypergraph; Identifying a first set of operations indicated by the first hyperedge; Determining a minimum set of data stores capable of jointly performing the first set of operations; Generating a cluster including the first set of concepts; And Labeling the cluster with the minimum set of data stores.

Citation Information

Patent Citations

  • A cloud data processing method and system based on workload

    CN108984308A

  • Process and Framework For Facilitating Data Sharing Using a Distributed Hypergraph

    US20150347480A1