System and method for generation and application of schema-independent query templates

By generating schema-independent query templates and using schema-classification mapping technology, the high cost and inefficiency problems of query writing of different mode data sets are solved, and the query generation of different mode data sets is automatically adapted to, which significantly saves time and workload.

CN115004171BActive Publication Date: 2025-05-16GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980103516.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-16
Publication Date
2025-05-16
Estimated Expiration
2039-12-16

AI Technical Summary

Technical Problem

The prior art requires manual writing of pattern-specific queries when processing data sets in different modes, resulting in high costs, high workload and inefficiency.

Method used

By generating schema-independent query templates and modifying them with schema-categorical mapping, schema-specific queries for data sets of different schemas.

Benefits of technology

It realizes that data sets of different modes are automatically adapted to data sets without changing the query semantics, significantly reducing the need for manual queries and saving time and workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004171B_ABST
    Figure CN115004171B_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for generating query templates expressed in a general, schema-independent language. Query templates can be generated "from scratch," or can be automatically generated from existing queries, a process that can be referred to as "templating" existing queries. As an example, the generation of query templates can be performed through an iterative process that iteratively generates candidate templates over time to optimize coverage of a set of existing queries. After generating the schema-independent query templates, the systems and methods described herein can automatically convert / map the templated queries into "concrete" schema-specific queries that can be evaluated on a specific customer schema / dataset. In this way, a query template for a given semantic query (e.g., "return the names of all employees") only needs to be written once.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to database management and more particularly to systems and methods for the generation of schema-agnostic query templates and their application to data sets with different schemas. Background Art

[0002] Data sets are typically structured according to a defined schema. A schema is a description or design of the structure or format of the data records contained in a data set. For example, a schema for a data set may provide a description of the attributes that are populated by each data record (e.g., represented in different columns) and / or may provide a description of how different data records and / or data tables containing various data records may be related to each other or otherwise logically associated with each other. A schema may include all implementation details required to store data records, such as the data types of data records, constraints on data records, etc.

[0003] In many cases, and for various reasons, such as the separate development and / or collection of the datasets, the schemas of the various datasets differ from each other (often significantly.) Thus, a given dataset may be structured according to a particular schema that does not necessarily match the schemas of other potentially related datasets.

[0004] Currently, given this schema-specific nature of datasets, general queries against datasets are expressed in a specific query language (such as SQL) and are explicitly designed to take into account the specific schema of the repository. Therefore, in order to query a specific dataset, a query must be developed that takes into account the complexity of the specific schema of the dataset to be queried. This process can require a significant amount of effort and expense, including developing and testing any number of intermediate, unsuccessful attempts. For larger entities, such as companies that maintain hundreds or thousands of different datasets, writing different schema-specific queries for each different dataset poses a huge challenge and source of inefficiency.

[0005] As an example, consider a query that returns the names of all employees. In Company A, information about all employees might be stored in a table called "Employees", and each employee's name might be stored in a column called "name". Then, for Company A, this query would be expressed as:

[0006] Q1: SELECT name

[0007] FROM Employees

[0008] But for Company B, the employee names might be stored in two columns named "first_name" and "last_name" in a table named "Personnel". Then, evaluating essentially the same query for Company B would require us to completely rewrite the query and write something like this:

[0009] Q2: SELECT first_name, last_name

[0010] FROM Personnel

[0011] What the simple example above illustrates is that the same semantically equivalent query needs to be tailored every time to the vocabulary of the specific schema against which the query is executed. This is actually an expectation enforced by the SQL standard (in this example). Because schemas vary widely between customers and databases, this task is almost always done manually, which again constitutes a significant source of cost, effort, and inefficiency. Summary of the invention

[0012] Various aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.

[0013] An example aspect of the present disclosure relates to a computer-implemented method for applying a schema-independent query template to a data set. The method includes obtaining a query template by a computing system including one or more computing devices, the query template including one or more references to one or more classification tags. Each of the one or more classification tags is a schema-independent representation of a data group. The method includes accessing, by the computing system, a schema-classification mapping associated with a data set stored in a database and structured according to a schema, wherein the schema-classification mapping defines a mapping between one or more classification tags and one or more components of a schema of the data set. The method includes modifying, by the computing system, the query template based on the schema-classification mapping to generate a schema-specific query. The method includes executing, by the computing system, a schema-specific query against a data set stored in a database to generate a query result. The method includes providing, by the computing system, the query result as an output.

[0014] Another example aspect of the present disclosure relates to a computing system that utilizes a query template that is independent of a pattern. The computing system includes a database storing a data set; one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operation includes obtaining a query template by the computing system, the query template including one or more references to one or more classification tags, each of the one or more classification tags being a pattern-independent representation of a data group. The operation includes accessing, by the computing system, a pattern-classification mapping associated with the data set and structured according to a pattern, wherein the pattern-classification mapping defines a mapping between one or more classification tags and one or more components of a pattern of the data set. The operation includes modifying, by the computing system, the query template based on the pattern-classification mapping to generate a pattern-specific query. The operation includes executing, by the computing system, a pattern-specific query against a data set stored in a database to generate a query result. The operation includes providing, by the computing system, the query result as an output.

[0015] Other aspects of the disclosure relate to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.

[0016] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] A detailed discussion of embodiments for those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings, in which:

[0018] Figure 1A-1B A block diagram of an example computing system is depicted according to an example embodiment of the present disclosure.

[0019] Figure 2 Depicted is a block diagram of an example query generation and application platform according to an example embodiment of the present disclosure.

[0020] Figure 3 A flow chart of an example method of generating a query template according to an example embodiment of the present disclosure is depicted.

[0021] Figure 4 Depicted is a flow diagram of an example method of applying a query template to a data set according to an example embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] Overview

[0023] In general, the present disclosure relates to systems and methods for generating and / or applying query templates expressed in a general, schema-independent language. Query templates can be generated "from scratch" or can be automatically generated from existing queries, a process that can be referred to as "templating" existing queries. As an example, the generation of query templates can be performed by an iterative process that iteratively generates candidate templates over time to optimize coverage of a set of existing queries. After generating schema-independent query templates, the systems and methods described herein can automatically convert / map the templated queries into "concrete" schema-specific queries that can be evaluated on a specific customer schema / dataset. In this way, a query template for a given semantic query (e.g., "return the names of all employees") only needs to be written once as a query template. The proposed system and method can then convert the query template into any number of schema-specific queries that appropriately consider any number of specific schemas provided as input (e.g., can be evaluated against them). On the other hand, given a query template and a description of a specific schema, the proposed system and method can evaluate the feasibility of generating a schema-specific query from a query template for a specific schema. Thus, example aspects of the present disclosure relate to generating and / or providing templated queries to a user, evaluating the feasibility of the templated queries against a user's patterns, generating pattern-specific queries from the templated queries (e.g., which may be parameterized), and executing the pattern-specific queries against the user's patterns to generate query results for the user.

[0024] More particularly, aspects of the present disclosure provide users with the ability to evaluate queries regardless of the underlying mode they are using. To this end, the systems and methods of the present disclosure can generate and apply query templates. In some embodiments, a query template can be a (query) specification that describes a general request such as "return the names of all employees".

[0025] As an example, a query template could look like this:

[0026] Query Template: SELECT<name_columa_list> , OPTIONAL( <age>)

[0027] FROM<employee_table_name>

[0028] OPTIONAL(WHERE <age> > <value>)

[0029] The example query template does not rely on a specific schema (i.e., it is "schema-agnostic"). Given such a template, and with reference to the background section, aspects of the present disclosure enable both Company A and Company B to evaluate the query template without being concerned with how the query template maps to their underlying schema. In fact, the proposed system utilizes a mechanism whereby, given: (a) a query template (such as the one above); and (b) the schema of the underlying data, the template is automatically converted into a specific query that can be run on the corresponding schema. Thus, for Company A's schema, the conversion mechanism would convert the above template into query Q1, while for Company B's schema, query Q2 would be generated. The above queries and query templates are provided as simplified examples, and more complex queries can be performed.

[0030] Compared to the automatic generation and application of templates proposed in this paper, past work on query templating has been limited to providing simple run-time arguments for queries. For example, the scope of SQL templating in the past was limited to supporting value parameters, as described below:

[0031] Q3: SELECT name

[0032] FROM Employees

[0033] WHERE age>"?"

[0034] The "?" in the WHERE clause can be dynamically replaced with a value provided by the user at runtime (for example, 25). However, the key structure of the query (such as table and column names) is always fixed. In addition, all parts of the query are required, and there is no way to handle missing data in any part.

[0035] To work around these limitations, users often rely on query views to replace real query templates. This way, the user writes a query to the view only once (which would happen if a real templated query mechanism existed), and the query is transformed into a query to the base schema (e.g., for Company A or Company B) via a set of mappings. An obvious limitation of the query view approach is that these mappings must be known a priori and are often created manually. Furthermore, in order for the query view mechanism to work, it assumes that all structures in the view (tables and columns in the case of a relational database) are mapped to the base schema.

[0036] In contrast, the templated query mechanism proposed in this article takes into account parts of the template that are optional and may or may not be mapped when the template is converted. As an example, a template that returns the names and possibly the ages (if present) of all employees may return results without an age column (for companies whose underlying schema does not store the ages of their employees) or with an age column (for companies that store this information). Similarly, a query template can have an optional filter that selects only employees who are, for example, 25 years old or older. If Company A stores the ages of its employees, then this filter may apply to employees of Company A, but may not apply to Company B when this information is not available. Such dynamic customization of query templates is a substantial advance over the basic capabilities of current mechanisms (including the ability to query views, which, as shown, cannot handle optionality).

[0037] More particularly, the query generation and application platform can be executed by a computing system including one or more computing devices. The query generation and application platform can generate one or more query templates and / or apply query templates to data sets with different schemas.

[0038] In some embodiments, query templates can be designed from scratch. As one example, a collection of predefined query templates can be developed (e.g., by developers of the platform itself) and provided to users as a package. As another example, a user can directly design a query template by entering the query template into a computing system (e.g., via a graphical user interface). As another example, a user can modify an existing mode-specific query or an existing query template to generate a new query template that can be saved for later use.

[0039] In other examples discussed in more detail below, the computing system can implement one or more algorithms to automatically generate new query templates. In one example, given a set of existing query templates (e.g., provided by a user), the computing system can automatically synthesize a query template that is a semantic equivalent of as large a portion of the set of existing query templates as possible (e.g., provides the same result set).

[0040] According to one aspect of the present disclosure, query templates can utilize taxonomy as a method of establishing a common language that describes the data sets brought into the platform by users and the query templates that the platform attempts to instantiate on these data sets. Each taxonomy may include one or more taxonomy labels. Each taxonomy label can be a schema-independent representation of a data group. For example, a taxonomy label can be associated with a semantic concept, and a taxonomy label can be an alternative reference to any schema-specific data (e.g., a column) that also references such a semantic concept or is associated with such a semantic concept (e.g., using a schema-specific vocabulary). Taxonomy labels can be manually defined and / or a default set of taxonomy labels can be provided. Thus, taxonomy can be flexible and can be bound to any different use cases or semantic query types.

[0041] For example, referring to the example query templates and queries Q1 and Q2 provided above, the classification label "name_column_list" is a schema-independent representation of the semantic concept of employee name. This classification label can be mapped to the schema-specific vocabulary used by the corresponding schemas of the datasets of Company A and Company B that reference the same semantic concept. In particular, for the semantic concept of employee name, Company A's schema contains a column named "name" in a table named "Employees", while Company B's schema contains two columns named "first_name" and "last_name" in a table named "Personnel". Therefore, the schema-independent classification label "name_column_list" can be mapped to these schema-specific vocabulary entries.

[0042] As described above, in some cases, the query template may be created manually, while in other cases, the computing system may automatically generate the query template from a set of one or more existing queries (e.g., the one or more existing queries are respectively associated with one or more existing data sets and can be executed on the one or more existing data sets). Thus, in one example, the computing system may receive a set of one or more existing mode-specific queries that are respectively associated with one or more existing data sets, and may generate a query template based on the set of one or more existing mode-specific queries.

[0043] As an example, generating a query template based on a set of existing mode-specific queries may include: iteratively generating a candidate query template based on a set of one or more existing mode-specific queries, and applying the candidate query template to one or more existing data sets respectively associated with the one or more existing mode-specific queries to obtain one or more candidate result sets. For example, executing the candidate query template against the existing data sets may include, for each existing data set, generating a corresponding mode-specific query from the candidate query template, and evaluating the corresponding mode-specific query on the existing data set to obtain a corresponding candidate result set of the existing database.

[0044] In each iteration, the computing system may compare one or more candidate result sets to one or more existing result sets generated by executing one or more pattern-specific queries against one or more existing data sets. For example, if a candidate result set generated for a particular data set matches (e.g., completely matches) an existing result set generated by executing a particular existing pattern-specific query against the particular data set, then the corresponding candidate query template may be indicated as having positive coverage for the particular existing pattern-specific query. In other words, if applying a particular candidate query template to an existing data set provides the same results as applying the corresponding existing pattern-specific query, then the candidate query template may be said to have coverage for such an existing pattern-specific query.

[0045] The computing system may iteratively search for candidate query templates that provide maximum coverage of a set of queries specific to an existing pattern. Various techniques may be performed in each iteration to generate candidate query templates to be evaluated in such iterations. As an example, existing queries and / or associated data sets may be represented using trees and / or graphs. An automated process may analyze available trees and / or overlaps / overlays of trees between candidate query templates and existing queries. Such analysis may reveal which parts of the query are necessary and which parts are optional, and may help understand how changes to the candidate query templates will change the semantics of the corresponding result set.

[0046] As another example, an iterative search technique (such as evolutionary search) can be performed to generate candidate query templates. For example, random mutations can be applied to the query templates in each iteration, and it can be evaluated whether such random mutations increase or decrease coverage. As yet another example, reinforcement learning techniques can be used to learn an agent model that generates candidate query templates. For example, the amount of coverage (or its relative increase) can be used to reward / train the agent model for generating candidate query templates. In one example, the agent model can be a recurrent neural network.

[0047] Thus, candidate query templates can be iteratively generated and evaluated until a certain stopping condition is met, with the goal of maximizing coverage over a set of existing queries. The process can be conceptually viewed as reverse engineering a set of existing queries (e.g., provided by a user) to generate query templates (e.g., which the user can then apply to other different and / or new datasets). The process can significantly save time and effort. For example, an organization such as a large company may have thousands of different datasets / databases. The organization can generate a few pattern-specific queries, which can then be used to generate (e.g., automatically generate) query templates that can be applied to a larger number of datasets. Thus, the amount of pattern-specific queries that need to be manually generated can be greatly reduced, saving time and effort.

[0048] Query templates (e.g., automatically generated query templates) can be expressed according to many different database structure languages. One example language is the SQL language. Other example languages ​​are graph-based languages, such as SPARQL. Thus, in some examples, a data set may include structured data (e.g., the data may be SQL data, and the query language may be SQL). In other examples, a data set may include semi-structured data (e.g., the data may be XML data, and the query language may be XPATH). In other examples, a data set may include graph data (e.g., RDF data) and the query language may be a graph-based language (e.g., SPARQL). Once query templates are generated, imported, or otherwise included in a query generation and application platform, these query templates may be applied to data sets (e.g., data sets uploaded, linked, or otherwise made accessible to the platform by a user). As an example, a data set accessible by a user may include a web analytics data set generated by a web analytics service such as Google Analytics.

[0049] According to one aspect of the present disclosure, the query generation and application platform can convert query templates into any number of pattern-specific queries that appropriately consider (eg, can be evaluated against) any number of specific patterns provided as input.

[0050] As an example, for a particular pair of a data set and a query template or classification, the query generation and application platform can access a schema-classification mapping associated with the data set. The schema-classification mapping can define a mapping between one or more classification labels included in the template or classification and one or more components of the schema of the data set. The schema-classification mapping can be generated manually, or can be automatically generated. Each mapping in the schema-classification mapping can be one-to-one, one-to-many, or many-to-one.

[0051] As an example, characteristics of pattern components can be matched against expected characteristics associated with classification labels to generate a mapping between pattern components and classification labels. For example, expected character lengths associated with classification labels can be used to map classification labels to pattern components.

[0052] For example, with some well-known exceptions, SKUs are typically alphanumeric strings that are between 6-8 characters in length. Therefore, a pattern component (e.g., column) containing a data entry that is also an alphanumeric string that is 6-8 characters long can be inferred to map to a SKU classification label. To give another example, UPC is uniformly assigned and is 12 digits in the United States and 13 digits in Europe. Therefore, a pattern component that matches these characteristics can be mapped to a UPC classification label. Another pattern that can be exploited is that SKU and UPC entries are typically unique and are not typically repeated in their corresponding columns.

[0053] More generally, patterns exhibited by data features associated with different pattern components can be exploited to automatically determine mappings between one or more classification labels and pattern components. In one example, a machine learning model (e.g., a neural network) can be trained to infer mappings given a pattern and / or underlying data set and a query template and / or classification as input. For example, a machine learning model can be trained in a supervised manner using existing mappings between classification labels and pattern components. As another example, in a reinforcement learning approach, an agent can be trained to generate mappings between patterns and classification labels. In particular, an agent can be trained using a reward function that evaluates how well the mapping generated by the agent enables a query template to be applied to a data set to obtain the same coverage as an existing query for the data set.

[0054] As another further example, the techniques for automatically determining a mapping between two schemas described in International Application No. PCT / US19 / 60010, filed on November 6, 2019, may be adapted to replace automatically determining a mapping between a particular schema and a schema-independent query or classification. International Application No. PCT / US19 / 60010 is incorporated herein by reference in its entirety.

[0055] After obtaining the schema-classification mapping between the data set and the classification or template, the query generation and application platform can modify the query template based on the schema-classification mapping to generate schema-specific queries.

[0056] As an example, modifying a query template based on a schema-classification mapping to generate a schema-specific query may include replacing each corresponding reference to one of the classification labels in the query template with a replacement reference to one or more components of the schema to which the classification label is mapped via the schema-classification mapping. In other words, each classification label in the query template may be replaced with a reference to a schema component mapped to such a classification label.

[0057] In some implementations, the query template may be parameterized. In such implementations, modifying the query template based on the pattern-classification mapping to generate a pattern-specific query may include: identifying one or more user-specifiable parameters included in the query template; providing a user interface that enables a user to enter values ​​for the one or more user-specifiable parameters; and constructing the pattern-specific query using the values ​​entered by the user for the one or more user-specifiable parameters via the user interface. Thus, as an example, a user may be requested to provide specific data for a filter to be applied to the query.

[0058] In some embodiments, one or more machine learning models can be used to modify query templates based on the pattern-classification mapping to generate pattern-specific queries. As an example, the machine learning model can be a sequence-to-sequence model that can be trained to rewrite query templates into pattern-specific queries, such as using the pattern-classification mapping (or its embedding) as context. For example, this approach can be conceptually similar to using a sequence-to-sequence model for machine translation.

[0059] The query generation and application platform can then execute the schema-specific query against the data set stored in the database to generate query results and provide the query results as output. For example, the query results can be displayed to a user. In another example, the schema-specific query can be generated and applied, and the query results can be provided as a service to other applications, devices, or entities (e.g., receiving a request for the query application via an application programming interface (API) and returning the query results via the API).

[0060] According to another aspect of the present disclosure, given a description of a query template and a particular pattern, the query generation and application platform can evaluate the feasibility of generating a pattern-specific query from the query template of the particular pattern. For example, when a user first introduces a data set into the platform, the platform can automatically evaluate which query templates included in the platform are feasible for running against the new data set and its associated patterns. The platform can provide certain visual indications of which query templates are feasible for execution against the data set. For example, queries that are not applicable can be "grayed out" or otherwise rendered as unselectable.

[0061] Many different techniques may be performed to evaluate the feasibility of applying a query template to a data set and its associated schema. As an example, evaluating the feasibility of generating a schema-specific query from a query template may include identifying one or more basic classification labels and one or more optional classification labels included in the query template, and determining whether, for each of the one or more basic classification labels, the schema-classification mapping is limited to a mapping to at least one component of the schema. In other words, in some embodiments, generating a schema-specific query from a query template may be feasible only when each classification label in the query template has at least one component of the schema mapped to it. In some embodiments, in principle, the query may still be evaluated in the absence of a schema component mapping to a classification label included in the query template, with NULL values ​​returned for all values ​​of the classification label. An alternative to these NULL values ​​is logic that allows the query template to be triggered only when all classification labels are based on specific relationship elements.

[0062] Furthermore, it should be noted that generating pattern-specific queries from query templates is feasible even when no pattern components are mapped to classification labels as optional classification labels. This ability to ignore the absence of data for optional labels is a significant improvement over existing pattern mapping systems, which require a one-to-one mapping of all elements of a query.

[0063] As another example technique for assessing feasibility, the amount of data associated with a pattern component involved in a query template (e.g., a pattern template that maps to a classification label included in the query template) can be compared to an average or total amount of data associated with other pattern components and / or the entire data set. This can ensure that the presence of significantly incomplete data does not give a false positive to the feasibility of applying the query template to the data set. Thus, in some embodiments, assessing the feasibility of generating a pattern-specific query from a query template can include: for at least one of the classification labels, determining the number of non-null data entries included in the data set associated with such a classification label and the total number of data entries included in the data set.

[0064] Additional aspects of the present disclosure focus on queries that have not yet been activated. For example, if a user notices that there is a very interesting query template, but the query has not been activated (actionable). In such a case, the platform can analyze the available data sets and can provide the user with an indication of what type of data must be imported to activate the query. This setup creates a useful feedback loop, where new data entering the system activates the category of queries, and the category of inactivated queries guides the client what type of data must be imported next.

[0065] Thus, example aspects of the present disclosure relate to generating and / or providing templated queries to a user, evaluating the feasibility of the templated queries against a user's patterns, generating pattern-specific queries from the templated queries (e.g., which may be parameterized), and executing the pattern-specific queries against the user's patterns to generate query results for the user.

[0066] The systems and methods of the present disclosure provide many technical effects and benefits. In particular, the proposed systems and methods do not simply map between two different existing schemas, but flexibly respond to real structural changes. As an example, the proposed systems and methods can indicate whether a given data set has sufficient data to enable the application of a query template (this concept is referred to as "feasibility" elsewhere in this document). As another example, the proposed systems and methods can enable and process optional functions within a query template. For example, applying an optional portion of a query template (e.g., an optional request for data associated with a classification label) can include determining whether a given data set has sufficient data to provide results for the optional portion of the query template.

[0067] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in greater detail.

[0068] Example devices and systems

[0069] Figure 1A A block diagram of an example computing system 100 is depicted in accordance with an example embodiment of the present disclosure. The computing system 100 includes a user computing device 102 and a database management system 130 that communicate (e.g., in a client-server relationship) over a network 180. The database management system 130 manages or otherwise accesses one or more databases 150. Communications between any of the devices 102, the system 130, and / or the databases 150 may be formatted or structured in accordance with one or more application programming interfaces (APIs).

[0070] The user computing device 102 may be any type of device, including a personal computing device and / or a client server device. The personal computing device may include a laptop computer, a desktop computer, a smart phone, a tablet computer, an embedded computing device, a game console, etc.

[0071] The user computing device 102 may include one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 executed by the processor 112 to cause the user computing device 102 to perform operations.

[0072] The database management system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or a plurality of processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 executed by the processor 132 to cause the database management system 130 to perform operations.

[0073] In some implementations, database management system 130 includes or is implemented by one or more server computing devices. Where database management system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0074] The database management system 130 may include and execute a query generation and application platform 140. The query generation and application platform 140 may provide core services for storing, processing, and protecting data to / from the database 150. In some implementations, the query generation and application platform 140 may include a query parser, a query optimizer, an execution engine, a metadata manager, an API handler, a query result returner, and / or other components.

[0075] According to one aspect of the present disclosure, the query generation and application platform 140 may generate and / or apply one or more query templates expressed in a generic schema-independent language. The query templates may be generated "from scratch," or may be automatically generated from existing queries, a process which may be referred to as "templating" existing queries. As an example, the query generation and application platform 140 may generate query templates through an iterative process that iteratively generates candidate templates over time to optimize coverage of a set of existing queries. After generating the schema-independent query templates, the query generation and application platform 140 may automatically convert / map the templated queries into "concrete" schema-specific queries that may be evaluated on a specific customer schema / data set. In this manner, a query template for a given semantic query (e.g., "return the names of all employees") need only be written once as a query template. Thus, the query generation and application platform 140 can convert the query template into any number of pattern-specific queries that appropriately consider (e.g., can be evaluated against) any number of specific patterns provided as input. On the other hand, given a query template and a description of a specific pattern, the query generation and application platform 140 can evaluate the feasibility of generating a pattern-specific query from a query template for a specific pattern. Operation Figure 2 An example query generation and application platform 140 is described.

[0076] Still referring to FIG. 1 , database 150 may store one or more data sets, for example, a data set may include a data table. Data tables (or individual rows, columns, or entries thereof) may be related to each other. In particular, each data set stored in database 150 may be organized or structured according to a particular schema. A data set may be associated with a particular user, and in some embodiments, a user must log in to query generation and application platform 140 to access its associated data set.

[0077] Database 150 may be one database or may be a plurality of physically distributed databases. Database 150 may be backed up by a plurality of storage devices (such as, for example, storage disks).

[0078] The query generation and application platform 140 (or any of its subsystems, such as Figure 2 ) can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some embodiments, the query generation and application platform 140 (or any of its subsystems) includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other embodiments, the query generation and application platform 140 (or any of its subsystems) includes one or more computer executable instruction sets stored in a tangible computer readable storage medium such as a RAM hard disk or an optical or magnetic medium.

[0079] Figure 1A An example computing system that can be used to implement the query generation and application platform 140 is shown. Other computing systems can also be used. As an example, Figure 1B An alternative arrangement is shown in which the query generation and application platform 140 is stored at and executed by the user computing device 102 .

[0080] Figure 2 2 depicts a graphical representation of an example query generation and application platform 200. The query generation and application platform 200 may be implemented by a computing system including one or more computing devices (e.g., such as Figure 1A and 1B The query generation and application platform 200 may generate one or more query templates 206 and / or apply the query templates 206 to data sets with different schemas (eg, the user data set 202 stored in the database 150).

[0081] In some implementations, query templates 206 may be designed from scratch. As one example, a collection of predefined query templates 206 may be developed (e.g., by the platform's own developers) and provided to users as a package. As another example, a user may directly design a query template 206 by entering the query template 206 into a graphical user interface of the platform 200. As another example, a user may modify an existing schema-specific query 203 or an existing query template to generate a new query template that may be saved for later use.

[0082] In other examples, such as Figure 2 , the platform 200 can implement a query template generator 205 to automatically generate a new query template 206. In one example, given a set of existing query templates 203 (e.g., provided by a user), the query template generator 205 can automatically synthesize a query template 206 that is a semantic equivalent of as large a portion of the set of existing query templates 203 as possible (e.g., provides the same result set).

[0083] According to one aspect of the present disclosure, query templates 206 can utilize taxonomies 204 as a method of establishing a common language that describes data sets (e.g., 202) brought into the platform by users and query templates 206 that the platform 200 attempts to instantiate on these data sets. Each taxonomy 204 may include one or more taxonomy labels. Each taxonomy label may be a schema-independent representation of a data group. For example, a taxonomy label may be associated with a semantic concept, and a taxonomy label may be an alternative reference to any schema-specific data (e.g., a column) that also references such a semantic concept or is associated with such a semantic concept (e.g., using a schema-specific vocabulary). Taxonomy labels may be manually defined and / or a default set of taxonomy labels may be provided. Thus, taxonomy 204 may be flexible and may be bound to any different use cases or semantic query types.

[0084] An example query template 206 for retrieving profit by product is as follows:

[0085]

[0086] In the example query template 206 provided above, unknown columns or relationships are represented by classification labels. In addition, the last line of the query template includes two special indicators marked with @user_para1 and @user_param2. These indicators can be used to trigger a UI that requires the user to fill in two values ​​in order to limit the search for profitable products to products released between a pair of dates.

[0087] As described above, in some cases, the query template 206 may be manually created, while in other cases, the query template generator 205 may automatically generate the query template 206 from a set of one or more existing queries 203 (e.g., the one or more existing queries are respectively associated with one or more existing data sets and can be executed on the one or more existing data sets). Thus, in one example, the platform 200 may receive a set of one or more existing mode-specific queries 203 that are respectively associated with one or more existing data sets, and may generate the query template 206 based on the set of one or more existing mode-specific queries.

[0088] As an example, generating the query template 206 based on the set of existing pattern-specific queries 203 may include: the query template generator 205 iteratively generates candidate query templates based on the set of one or more existing pattern-specific queries 203, and applies the candidate query templates to one or more existing data sets respectively associated with the one or more existing pattern-specific queries 203 to obtain one or more candidate result sets. For example, executing the candidate query template on the existing data sets may include, for each existing data set, generating a corresponding pattern-specific query from the candidate query template, and evaluating the corresponding pattern-specific query on the existing data set to obtain a corresponding candidate result set of the existing database.

[0089] In each iteration, the computing system may compare one or more candidate result sets to one or more existing result sets generated by executing one or more pattern-specific queries 203 against one or more existing data sets. For example, if a candidate result set generated for a particular data set matches (e.g., completely matches) an existing result set generated by executing a particular existing pattern-specific query 203 against the particular data set, then the corresponding candidate query template may be indicated as having positive coverage for the particular existing pattern-specific query 203. In other words, if applying a particular candidate query template to an existing data set provides the same results as applying the corresponding existing pattern-specific query 203, then the candidate query template may be said to have coverage for such an existing pattern-specific query.

[0090] The query template generator 205 can iteratively search for candidate query templates that provide maximum coverage of the set of existing pattern-specific queries 203. Various techniques can be performed in each iteration to generate candidate query templates to be evaluated in such iterations. As an example, the existing queries and / or associated data sets can be represented using trees and / or graphs. The automated process can analyze the available trees and / or overlaps / overlays of trees between the candidate query templates and the existing queries. Such analysis can reveal which parts of the query are necessary and which parts are optional, and can help understand how changes to the candidate query templates will change the semantics of the corresponding result set.

[0091] As another example, an iterative search technique (such as evolutionary search) can be performed to generate candidate query templates. For example, random mutations can be applied to the candidate query templates in each iteration, and it can be evaluated whether such random mutations increase or decrease coverage. As yet another example, reinforcement learning techniques can be used to learn an agent model that generates candidate query templates. For example, the amount of coverage (or its relative increase) can be used to reward / train the agent model for generating candidate query templates. In one example, the agent model can be a recurrent neural network.

[0092] Thus, candidate query templates can be iteratively generated and evaluated until a certain stopping condition is met, with the goal of maximizing coverage over a set of existing queries. The process can be conceptually viewed as reverse engineering a set of existing queries 203 (e.g., provided by a user) to generate a query template 206 (e.g., which the user can then apply to other different and / or new data sets). The process can significantly save time and effort. For example, an organization such as a large company may have thousands of different data sets / databases. The organization can generate a few pattern-specific queries, which can then be used to generate (e.g., automatically generate) query templates that can be applied to a larger number of data sets. Therefore, the amount of pattern-specific queries that need to be manually generated can be greatly reduced, thereby saving time and effort.

[0093] Query templates 206 (e.g., automatically generated query templates) can be expressed according to many different database structure languages. One example language is the SQL language. Other example languages ​​are graph-based languages, such as SPARQL. Thus, in some examples, a data set may include structured data (e.g., the data may be SQL data, and the query language may be SQL). In other examples, a data set may include semi-structured data (e.g., the data may be XML data, and the query language may be XPATH). In other examples, a data set may include graph data (e.g., RDF data) and the query language may be a graph-based language (e.g., SPARQL). Once query templates 206 are generated, imported, or otherwise included in the query generation and application platform 200, these query templates may be applied to data sets (e.g., data sets uploaded, linked, or otherwise made accessible to the platform by a user, such as data set 202). As an example, a data set accessible by a user may include a web analytics data set generated by a web analytics service such as Google Analytics.

[0094] According to one aspect of the present disclosure, query generation and application platform 200 can convert query templates 206 into any number of pattern-specific queries 214 that appropriately consider (eg, can be evaluated against) any number of specific patterns provided as input.

[0095] As an example, for a particular pair of a data set 202 and a query template 206 or a taxonomy 204, the query generation and application platform 200 can access a schema-taxonomy mapping 210 associated with the data set 202. The schema-taxonomy mapping 210 can define a mapping between one or more taxonomy labels included in the query template 206 or the taxonomy 204 and one or more components of the schema of the data set 202. The schema-taxonomy mapping 210 can be generated manually, or can be automatically generated ( Figure 2 A schema-classification map 210 is shown automatically generated by the schema-classification map generator 208).

[0096] As an example, characteristics of pattern components of patterns of data set 202 can be matched against expected characteristics associated with classification labels included in classification 204 to generate mappings between components of patterns and classification labels included in pattern-to-classification mapping 210. For example, expected character lengths associated with classification labels can be used to map classification labels to components of patterns.

[0097] For example, with some well-known exceptions, SKUs are typically alphanumeric strings that are between 6-8 characters in length. Therefore, a pattern component (e.g., column) containing a data entry that is also an alphanumeric string that is 6-8 characters long can be inferred to map to a SKU classification label. To give another example, UPC is uniformly assigned and is 12 digits in the United States and 13 digits in Europe. Therefore, a pattern component that matches these characteristics can be mapped to a UPC classification label. Another pattern that can be exploited is that SKU and UPC entries are typically unique and are not typically repeated in their corresponding columns.

[0098] More generally, patterns exhibited by data features associated with different pattern components can be exploited to automatically determine a mapping in a pattern-to-classification mapping 210 between a classification 204 or one or more classification labels of a query template 206 and pattern components of a pattern of a data set 202. In one example, a machine learning model (e.g., a neural network) can be trained to infer a mapping given a pattern and / or underlying data set and a query template 206 and / or classification 204 as input. For example, a machine learning model can be trained in a supervised manner using existing mappings between classification labels and pattern components. As another example, in a reinforcement learning approach, an agent can be trained to generate mappings between patterns and classification labels. In particular, an agent can be trained using a reward function that evaluates how well the mapping generated by the agent enables a query template to be applied to a data set to obtain the same coverage as an existing query for the data set.

[0099] As another further example, the techniques for automatically determining a mapping between two schemas described in International Application No. PCT / US19 / 60010, filed on November 6, 2019, may be adapted to replace automatically determining a mapping between a particular schema and a schema-independent query or classification. International Application No. PCT / US19 / 60010 is incorporated herein by reference in its entirety.

[0100] As an example, the sample imported data table can be as follows:

[0101] <![CDATA[ product ]]> <![CDATA[ Product Name ]]> <![CDATA[ release date ]]> <![CDATA[ price ]]> <![CDATA[ profit ]]> GXP4100 Widget 1 8 / 23 / 2016 $300 $10 GVE2300 Largest model widget 8 / 23 / 2016 $450 $20 TGB4245 Super Widget 8 / 15 / 2017 $700 $30 Pix-NEW60 Widget Camera 8 / 15 / 2017 $850 $40

[0102] Continuing with the example, an example mode-classification mapping 210 is as follows:

[0103]

[0104] As another related example, the sample template classification profile is as follows:

[0105]

[0106] In this example, the set of category tags in the profile in the template ("Profit" and "SKU") are contained in the category tags of the profile, and the template can therefore be activated. For example, before actually activating, for each col_name associated with the template, a query is issued to check if the column is actually populated with data. Since the optional check_query is missing in the profile, this means that the check queries for both columns are trivial and they only need to check if a non-NULL value exists.

[0107] Even when all the category labels associated with a query template can be based on the schema-category mapping, it may be the case that the imported relation does not actually provide data for the imported columns. For example, suppose that in the imported dataset, the user chose to hide a certain column of data (for example, all values ​​in the "Profit" column are NULL). To address such issues, a simple query can be issued to determine whether certain important columns in the imported data are correctly populated. For example, the following query counts tuples that have profit values ​​populated, and its result can be compared to the number of tuples in the entire relation.

[0108] SELECT count(*)

[0109] FROM Product

[0110] WHERE profit IS NOT NULL;

[0111] Another consideration is that, in addition to the templated (classification-annotation) part of the query, in order to activate the query template, it may also be necessary to check whether the data corresponding to the non-templated part of the query is available in the corresponding database. For example, a check can be performed to determine whether the user's dataset has indeed been imported into a database accessible to the platform. This last check can be performed even for trivial queries (e.g., queries that do not contain classification labels but rely only on predetermined and well-known relationships). For such queries, it is also useful to determine the point in time when the user has imported all necessary data into the platform to activate the query.

[0112] Still refer to Figure 2 After having obtained the schema-classification mapping 210 between the data set 202 and the classification 204 , the schema-specific query generator 212 may modify the query template 206 based on the schema-classification mapping 210 to generate a schema-specific query 214 .

[0113] As one example, modifying the query template 206 to generate the schema-specific query 214 based on the schema-classification mapping 210 can include replacing each corresponding reference to one of the classification labels in the query template 206 with a replacement reference to one or more components of the schema to which the classification label is mapped via the schema-classification mapping 210. In other words, each classification label in the query template 206 can be replaced with a reference to a schema component that is mapped to such a classification label.

[0114] For example, the replacement step can be performed by parsing the AST of the template (e.g., using the Google SQL parser API) and replacing each reference to the label with the corresponding relationship and column mentioned in the associated TemplateTaxonomyProfile. Thus, in some embodiments, the schema-specific query generator 212 can use both the schema-classification mapping that associates the user dataset 202 with the classification 204 and the template classification profile that associates the query template 206 with the classification 204.

[0115] In some implementations, the query template 206 can be parameterized. In such implementations, modifying the query template 206 based on the pattern-classification mapping 210 to generate the pattern-specific query 214 can include: identifying one or more user-specifiable parameters included in the query template 206; providing a user interface that enables a user to enter values ​​for the one or more user-specifiable parameters; and constructing the pattern-specific query 214 using the values ​​entered by the user for the one or more user-specifiable parameters via the user interface. Thus, as an example, a user can be requested to provide specific data for a filter to be applied to the query.

[0116] In some implementations, one or more machine learning models can be used to modify the query template 206 based on the pattern-classification mapping 210 to generate a pattern-specific query 214. As an example, the machine learning model can be a sequence-to-sequence model that can be trained to rewrite the query template 206 into a pattern-specific query 214, e.g., using the pattern-classification mapping 210 (or its embedding) as context. For example, this approach can be conceptually similar to using a sequence-to-sequence model for machine translation.

[0117] The query generation and application platform 200 may then execute the schema-specific query 214 against the data set 202 stored in the database 150 to generate query results 216. The query results 216 may be provided as output. For example, the query results 216 may be displayed to a user. In another example, the schema-specific query 214 may be generated and applied, and the query results 216 may be provided as a service to other applications, devices, or entities (e.g., receiving a request for query application via an application programming interface (API) and returning the query results 216 via the API).

[0118] Example Method

[0119] Figure 3 Depicted is a flow diagram of an example method 300 of generating a query template according to an example embodiment of the present disclosure.

[0120] At 302, the computing system may receive a set of one or more existing schema-specific queries respectively associated with one or more existing data sets. In particular, each existing data set may be structured according to a specific schema, and each schema-specific query may be written to take such a specific schema into account.

[0121] At 304, the computing system may generate a candidate query template based on a set of one or more existing mode-specific queries. For example, the candidate query template may be randomly initialized, or an existing query template for a related task may be used as a starting point for the candidate query template.

[0122] At 306, the computing system may apply the candidate query template to one or more existing data sets to obtain one or more candidate result sets. For example, applying the candidate query template to one or more existing data sets may include performing a Figure 4 Method 400.

[0123] Still refer to Figure 3 At 308, the computing system may compare the one or more candidate result sets to one or more existing result sets generated by executing one or more pattern-specific queries against one or more existing data sets. For example, the computing system may evaluate the coverage of the candidate query template relative to a set of existing pattern-specific queries.

[0124] At 310, the computing system can determine whether additional iterations should be performed. For example, iterations can be performed until one or more stopping criteria are met. The stopping criteria can be any number of different criteria, including, for example, a loop counter reaching a predetermined maximum value, an iteration-to-iteration variation in template adjustment falling below a threshold, an iteration-to-iteration variation in coverage falling below a threshold, a gradient of a loss function falling below a threshold, and / or various other criteria.

[0125] If it is determined at 310 that additional iterations should be performed, method 300 may return to 304 and generate another candidate query template (eg, by mutating or otherwise modifying a previous candidate query template).

[0126] However, if it is determined at 310 that additional iterations should not be performed, method 300 may proceed to 312. At 312, the computing system may save the current candidate query template as a query template that may be used later (eg, later applied to the attachment data set).

[0127] Figure 4 Depicted is a flow diagram of an example method 400 of applying a query template to a data set according to an example embodiment of the present disclosure.

[0128] At 402, a computing system may receive a data set structured according to a schema. For example, the data set may be uploaded, linked, or otherwise accessible by a user.

[0129] At 404, the computing system can access a schema-classification mapping associated with the data set. For example, the schema-classification mapping can be manually generated or automatically generated. The schema-classification mapping can map each classification label in the classification and / or query template to a component of the schema of the data set.

[0130] At 406, the computing system can evaluate the feasibility of applying one or more query templates to the data set. At 408, the computing system can provide an indication to a user (eg, in a user interface) of which query templates are feasible.

[0131] At 410, the computing system can receive a user selection of one of the query templates. At 412, the computing system can modify the query template based on the pattern-classification mapping to generate a pattern-specific query.

[0132] At 414, the computing system may perform a schema-specific query on the data set to generate query results. At 416, the computing system may provide the query results as output (eg, display the results to a user or provide the results to another system (eg, via an API)).

[0133] Additional disclosure

[0134] The technology discussed herein relates to servers, databases, software applications and other computer-based systems, as well as the actions taken by such systems and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for multiple possible configurations, combinations and divisions of tasks and functions between two components and between more components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in conjunction. Databases and applications can be implemented on a single system, or can be distributed on multiple systems. Distributed components can operate sequentially or in parallel.

[0135] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art, upon gaining an understanding of the foregoing, may easily produce changes, variations, and equivalents to such embodiments. Therefore, the present disclosure does not exclude modifications, variations, and / or additions to the subject matter that would be apparent to those of ordinary skill in the art. For example, a feature shown or described as part of one embodiment may be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to cover such changes, variations, and equivalents.

[0136] In particular, although for the purpose of illustration and discussion, certain drawings depict steps performed in a specific order, the method of the present disclosure is not limited to the order or arrangement of the specific description. The various steps of the method described herein may be omitted, rearranged, combined and / or adjusted in various ways without departing from the scope of the present disclosure.< / value> < / age> < / age>

Claims

1. A computer-implemented method for applying a schema-independent query template to a data set, the method comprising: obtaining, by a computing system comprising one or more computing devices, a query template comprising one or more references to one or more classification labels, each of the one or more classification labels being a schema-independent representation of a data set; accessing, by the computing system, a schema-to-classification mapping associated with a data set stored in a database and structured according to a schema, wherein the schema-to-classification mapping defines a mapping between the one or more classification labels and one or more components of a schema of the data set; modifying, by the computing system, the query template based on the pattern-classification mapping to generate a pattern-specific query; executing, by the computing system, the schema-specific query against the data set stored in the database to generate a query result; and providing the query result as output by the computing system, Wherein, obtaining the query template by the computing system includes: Receiving, by the computing system, a set of one or more existing schema-specific queries respectively associated with one or more existing data sets; and automatically generating, by the computing system, the query template based on the set of one or more existing schema-specific queries, The automatically generating, by the computing system, the query template based on the set of one or more existing schema-specific queries comprises: For each of multiple iterations: generating, by the computing system, a candidate query template based on the set of one or more existing schema-specific queries; applying, by the computing system, the candidate query template to the one or more existing data sets respectively associated with the one or more existing schema-specific queries to obtain one or more candidate result sets; and comparing, by the computing system, the one or more candidate result sets with one or more existing result sets generated by executing the one or more schema-specific queries against the one or more existing data sets, Whether additional iterations should be performed is determined based on whether one or more stopping criteria are met.

2. The computer-implemented method of claim 1 , wherein: Receiving, by the computing system, the set of one or more existing schema-specific queries respectively associated with the one or more existing data sets includes: receiving, by the computing system, user input from a user of the computing system, the user input identifying the set of one or more existing schema-specific queries and requesting automatic generation of the query template.

3. The computer-implemented method of claim 1 , wherein: The query template includes at least one basic part and at least one optional part.

4. The computer-implemented method of claim 1 , further comprising, before modifying, by the computing system, the query template based on the schema-classification mapping to generate the schema-specific query: The feasibility of generating the schema-specific query from the query template is evaluated, by the computing system and based at least in part on the schema-classification mapping.

5. The computer-implemented method of claim 4, wherein: Evaluating, by the computing system and based at least in part on the schema-classification mapping, the feasibility of generating the schema-specific query from the query template includes: identifying, by the computing system, one or more basic classification tags and one or more optional classification tags included in the query template; and A determination is made by the computing system as to whether, for each of the one or more base classification labels, the schema-to-classification mapping defines a mapping to at least one component of the schema.

6. A computer-implemented method according to claim 4 or 5, wherein: Evaluating, by the computing system and based at least in part on the schema-classification mapping, the feasibility of generating the schema-specific query from the query template includes: A number of non-null data entries included in the data set associated with such classification label and a total number of data entries included in the data set are determined by the computing system and for at least one of the classification labels.

7. The computer-implemented method of claim 1 , wherein: Obtaining the query template by the computing system includes obtaining a plurality of different query templates by the computing system; and The method further comprises: evaluating, by the computing system and based at least in part on the schema-classification mapping, feasibility of generating a corresponding schema-specific query from each of the plurality of query templates; and An indication is provided, by the computing system, to a user of the computing system as to which of the plurality of query templates has been evaluated as feasible for generating the corresponding schema-specific query.

8. The computer-implemented method of claim 7, further comprising: receiving, by the computing system, a user input, the user input selecting one of the plurality of query templates for which generating the corresponding mode-specific query has been evaluated as feasible; Wherein, modifying the query template by the computing system, executing the schema-specific query on the data set by the computing system, and providing the query result as output by the computing system are performed in response to the user input selecting the one of the plurality of query templates.

9. The computer-implemented method of claim 1 , wherein: Modifying the query template by the computing system based on the schema-classification mapping to generate the schema-specific query includes replacing, by the computing system, each corresponding reference to one of the classification labels in the query template with a replacement reference to one or more components of the schema to which the classification label is mapped via the schema-classification mapping.

10. The computer-implemented method of claim 1, wherein: Modifying, by the computing system, the query template based on the schema-classification mapping to generate the schema-specific query comprises: identifying, by the computing system, one or more user-specifiable parameters included in the query template; providing, by the computing system, a user interface that enables a user to input values ​​for the one or more user-specifiable parameters; and The mode-specific query is constructed by the computing system using the values ​​entered by the user via the user interface for the one or more user-specifiable parameters.

11. The computer-implemented method of claim 1 , wherein: Modifying, by the computing system, the query template based on the schema-classification mapping to generate the schema-specific query comprises: The query template and the pattern-classification mapping are processed by the computing system using a machine learning model to generate the pattern-specific query as an output of the machine learning model.

12. The computer-implemented method of claim 1, wherein: The data set stored in the database comprises a SQL data set stored in a relational database, and wherein the schema-specific query comprises a SQL query.

13. The computer-implemented method of claim 1, wherein: The data set stored in the database includes semi-structured data or graphic data.

14. A computing system utilizing a schema-independent query template, the computing system comprising: A database to store data sets; one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to implement the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Data query and location through a central ontology model

    US20050240606A1