Accelerating SPARQL queries over RDF graph sets using pattern-based result tables
Patent Information
- Application Number
- CN202580012180.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-29
- Publication Date
- 2026-08-28
Smart Images

Figure CN122663571A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to database systems, and more particularly, to using schema-based result tables to accelerate SPARQL queries on RDF atlases. Background Technology
[0002] The Resource Definition Framework (RDF) is a framework for representing relationships between "resources." Resources can be, for example, data interconnected on a network. RDF statements express relationships between resources such as documents, physical objects, people, concepts, data objects, etc. A collection of RDF statements forms a directed graph, where nodes are resources and edges are relationships between the resources connected. Additional information about RDF can be found at https: / / www.w3.org / TR / rdf11-concepts / RDF / , the full text of which is incorporated herein by reference.
[0003] Information about the resources and relationships of an RDF graph can be stored in tables within a database system. The table storing this information is referred to as an RDF table in this document. Typically, within an RDF table, graph information is stored in the form of "triples." A triple consists of a subject (the subject resource), a predicate (the relationship between the subject resource and the object resource), and an object (the object resource). The relation values within a triple are also called "properties." RDF graphs do not have any redundant edges; therefore, each triple stored in an RDF table must be unique.
[0004] Once an RDF table has been populated with information about an RDF graph, users can run queries against that RDF table to obtain information about the resources and / or relationships represented in the graph. A specialized query language called SPARQL has been developed for retrieving and manipulating data stored in RDF format. SPARQL is described in detail at www.w3.org / TR / sparql11-query / and https: / / www.w3.org / TR / sparql11-update / , the entire contents of which are incorporated herein by reference. U.S. Patent Application No. 11 / 822,531, entitled “METHOD AND SYSTEM FORUSING AUXILIARY TABLES FOR RDF DATA STORED IN A RELATIONAL DATABASE”, filed December 7, 2021 and granted November 21, 2023 (the entire contents of which are incorporated herein by reference), describes techniques for using auxiliary tables to improve the performance of SPARQL queries targeting individual RDF graphs.
[0005] Databases typically contain many RDF tables, each storing information about a distinctly different RDF graph. SPARQL supports the ability to retrieve / manipulate data from multiple RDF tables in a single query. When a SPARQL query operates on multiple RDF tables, it is said that the SPARQL query targets an "RDF atlas".
[0006] While there is no redundancy between triples within a single RDF graph, redundancy can exist between triples within an RDF graph set. That is, the triple (subject1, property1, object1) may not repeat within any individual RDF graph (e.g., RDF tables X, Y, or Z), but it can repeat across multiple RDF tables (e.g., it can appear in each of RDF tables X, Y, and Z). Therefore, a SPARQL query targeting an RDF graph set that includes RDF tables X, Y, and Z may result in retrieving the same triple (subject1, property1, object1) from each of RDF tables X, Y, and Z.
[0007] Based on SPARQL semantics, SPARQL queries must produce results based on deduplicated datasets. Therefore, all duplicate triples must be removed from the data retrieved by SPARQL queries targeting RDF atlases. Unfortunately, the process of removing duplicate triples (generally referred to as "deduplication") is computationally expensive and consumes significant memory. Thus, the need to perform deduplication adds a substantial amount of overhead to the execution of SPARQL queries targeting RDF atlases. Therefore, techniques are needed to improve the performance of SPARQL queries targeting RDF atlases.
[0008] The methods described in this section are possible methods, but not necessarily methods that have been previously conceived or adopted. Therefore, unless otherwise indicated, no method in this section should be considered prior art simply because it is included in this section. Furthermore, it should not be assumed that any method described in this section is fully understood, routine, or conventional merely because it is included in this section. Attached Figure Description
[0009] In the attached diagram: (The acronym "PRT" used in this document represents the pre-calculated results table)
[0010] Figure 1 This is a diagram depicting an example of a star pattern; Figure 2 This is a diagram illustrating an example of a chain pattern; Figure 3 It is a diagram depicting an example of a chain pattern that includes reverse characteristics; Figure 4It is a diagram depicting the RDF tables belonging to the "Family" RDF atlas; Figure 5 It is a diagram that graphically illustrates the contents of the Family RDF atlas; Figure 6 This is a block diagram of a star-shaped PRT according to one implementation method, which can be used to improve the performance of SPARQL queries targeting RDF atlases. Figure 7 This is a block diagram of a chained PRT according to one implementation method, which can be used to improve the performance of SPARQL queries targeting RDF atlases. Figure 8 This is a block diagram illustrating, according to one implementation method, how a query can be decomposed into sub-patterns that match several PRTs. Figure 9 This is a block diagram illustrating how the PRT can be maintained in response to the insertion of repeated triples according to one implementation. Figure 10 This is a block diagram illustrating how the PRT can be maintained in response to the insertion of a unique triple, according to one implementation. Figure 11 This is a block diagram illustrating how the PRT can be maintained in response to the insertion of a unique triplet to create a new chain, according to one implementation. Figure 12 This is a block diagram illustrating how a PRT can be maintained in response to the deletion of a repeating triple, according to one implementation. Figure 13 This is a block diagram illustrating how a PRT can be maintained in response to the deletion of a unique triple, according to one implementation method; and Figure 14 This is a block diagram illustrating a computer system on which the techniques described herein can be implemented. Detailed Implementation
[0011] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent that the invention may be practiced without these specific details. In other instances, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring the invention.
[0012] SPARQL Query Overview
[0013] SPARQL queries often specify a "triple schema," which indicates an "anchor" and the resource with a specified relationship to the anchor. An example of a single triple schema is: Q1: {?mom :motherOf ?kid} This single triplet pattern selects all triples whose property is "motherOf". In the Q1 triplet pattern, "mom" is the anchor variable, "motherOf" specifies the property, and "kid" is the variable for which "mom" has the "motherOf" property.
[0014] The attribute "motherOf" is a "forward" attribute because it proceeds in the same direction as the corresponding edge in the RDF directed graph. Triple schemas can use the prefix "^" to specify a "reverse" relation, which proceeds in the opposite direction to the corresponding edge in the RDF directed graph. For example, the reverse of the motherOf attribute "^:motherOf" is semantically equivalent to ":hasMother". By using the reverse attribute prefix with the attribute "motherOf", the anchor of Q1 can be changed from "mom" to "kid" as follows: Q2: {?kid ^:motherOf ?mom} Q1 and Q2 select exactly the same triples (all triples with the "motherOf" relation), but Q2 specifies the reverse property and thus changes which variable is used as the anchor for the triple pattern.
[0015] A slightly more complex triplet pattern can specify multiple relations for the same anchor. For example, consider the triplet pattern: Q3: {?s :firstName ?fnm . ?s :lastName ?lnm} In Q3, the anchor is "s". Q3 assigns two properties to the anchor "s": "firstName" and "lastName", where "fnm" is a variable of the object of the "firstName" property, and "lnm" is a variable of the object of the "lastName" property. Therefore, for any resource s, the triplet pattern Q3 will return the first name (firstname) and the last name (lastname) in the variables fnm and lnm, respectively.
[0016] A pattern that assigns multiple properties to an anchor is called a "star pattern". In a diagram representing a star pattern, the anchor is in the middle, the forward property is shown pointing away from the anchor, and the reverse property is shown pointing towards the anchor.
[0017] SPARQL patterns can specify both forward and reverse characteristics for anchors. Consider the following star pattern: Q4: {?s :firstName ?fnm . ?s :lastName ?lnm . ?s ^:fatherOf ?dad . ?s^:motherOf ?mom} Similar to Q3, Q4's anchor is "s", and the triplet pattern specifies the forward attributes "firstName" and "lastName". However, Q4 additionally specifies the reverse attributes ^:fatherOf and ^:motherOf. Therefore, for any resource s, Q4's star schema will return first name, last name, father, and mother in the variables fnm, lnm, dad, and mom, respectively. Figure 1 It is a diagram depicting the star pattern of Q4.
[0018] In addition to the star schema, SPARQL queries can also specify a chain schema. A chain schema specifies a sequence of two or more triplet schemas that form a chain. Therefore, a chain schema selects nodes that are two or more edges away from the anchor. For example, consider the following example: Q5: {?gma :motherOf ?mom . ?mom : motherOf ?kid} In this example, the pattern requires retrieval of: gma as a mother's child, and gma as a mother's child's child That is, in any chain that matches the pattern Q5, "gma" is the mother of the mother of "kid". Figure 2 This is a diagram that graphically depicts the chain pattern of Q5. As illustrated by this example, multiple edges of a chain forming a chain pattern can correspond to the same property (i.e., "motherOf").
[0019] Chained patterns can also include reverse characteristics, as shown in the following example: Q6: {?gma :motherOf ?mom . ?mom : motherOf ?kid . ?kid ^:fatherOf ?dad . ?dad ^:fatherOf ?gpa} Figure 3 This is a diagram depicting the chained patterns of Q6. In Q6, the anchor is "kid", and the pattern selects "mom" which is one "motherOf" property away from "kid", and "gma" which is two "motherOf" properties away from "kid". Additionally, the pattern selects "dad" which is one reverse "fatherOf" property away from "kid", and "gpa" which is two reverse "fatherOf" properties away from "kid".
[0020] General Overview
[0021] Techniques are provided for generating and maintaining pre-computed result tables (“PRTs”), which database servers use to improve the performance of SPARQL queries targeting RDF atlases. The PRTs used to improve the performance of SPARQL queries targeting RDF atlases are stored in the database, separate from the RDF tables belonging to the RDF atlas, but populated based on the contents of the RDF tables belonging to the RDF atlas.
[0022] In one implementation, each PRT corresponds to a specific schema. A "Simple Triple PRT" is the PRT corresponding to a simple triple schema and is used to improve the performance of SPARQL queries that include the simple triple schema corresponding to that PRT. A "Star PRT" is the PRT corresponding to a star schema and is used to improve the performance of SPARQL queries that include the star schema corresponding to that PRT. A "Chained PRT" is the PRT corresponding to a chained schema and is used to improve the performance of SPARQL queries that include the chained schema corresponding to that PRT.
[0023] It is necessary to "maintain" the PRT to ensure that using the PRT to speed up SPARQL queries will still produce accurate results. This maintenance involves ensuring that the PRT reflects all changes made to the RDF table by Data Manipulation Language (DML) statements. Therefore, techniques for maintaining the result table in both incremental and batch update modes are also described below.
[0024] Example RDF table
[0025] For the purposes of explanation, it will be assumed that the database stores a 400 error. Figure 4 The RDF table shown is for reference. Figure 4 Six RDF tables (tables 410, 412, 414, 416, 418, and 420) are stored in database 400. Within each table, triples are unique. However, when considered collectively as an RDF atlas, duplicate triples exist.
[0026] For illustrative purposes, it will be assumed that "Family" is the name of the RDF atlas created from the union of Tables 410, 412, 414, 416, 418, and 420. Within the Family RDF atlas, the following triples appear two or more times: :beth :motherOf :sue (in Tables 410 and 412) :sue :motherOf :rick (in Tables 410 and 414) :dan :fatherOf :rick (in Tables 410 and 418) :bob :fatherOf :dan (in Tables 410 and 420) :rick :lname“Smith” (in Tables 414 and 416) :rick :plays :soccer (in Tables 414 and 416) Figure 5 The contents of the Family RDF atlas are illustrated in the form of a directed graph, where “(2)” next to an edge indicates that the triple represented by that edge is repeated within the Family RDF atlas. For edges that are unique in the Family RDF atlas, no number is shown next to the edge (in such cases, “1” is implied).
[0027] In this example, no edge in the Family RDF atlas has more than two duplicates. However, the same triple may be repeated in every RDF graph / table of the RDF atlas. Therefore, in an RDF atlas with n RDF graphs / tables, the same triple may be repeated at most n times.
[0028] Create PRT
[0029] As mentioned above, for any given RDF atlas, a PRT can be created to accelerate the processing of SPARQL queries targeting that RDF atlas. According to one implementation, each PRT has a schema (its "PRT schema") and a target RDF atlas. Each PRT is populated with deduplicated results generated by executing a SPARQL query specifying its PRT schema against its RDF atlas. Examples of PRTs for each of three schema types will now be given: simple triple schema, star schema, and chain schema.
[0030] PRT for Simple Triple Pattern
[0031] For the purposes of explanation, it will be assumed that the PRT named "RT_triple" has the following attributes: PRT pattern = {?x :plays :?game} RDF graph set = Family In this example, RT_triple is associated with a simple triple pattern. In these cases, before deduplication, it is done by targeting the Family RDF atlas (in... Figure 5 The result of executing PRT mode (illustrated graphically) will be:
[0032] According to one implementation, when performing deduplication, the database server counts the number of instances of each edge / triple existing in the RDF graph and records this information in a separate column of PRT. Therefore, after deduplication, RT_triple will be filled as follows:
[0033] The “C<:plays>” column of each row stores the count of instances of the triple represented by that row. Since the triple “:rick :plays :soccer” appears twice in the Family RDF atlas, the “C<:plays>” column stores the value “2” in the row for that triple. Such duplicate tracking columns are used to maintain the PRT during DML operations that affect the contents of RDF tables belonging to the RDF atlas, as will be described in more detail below.
[0034] PRT for star pattern
[0035] For the purposes of explanation, assume that a PRT named "RT_star" has the following properties: PRT pattern = {?x :firstName ?fnm . ?x :lastName ?lnm . ?x ^:fatherOf?dad . ?x ^:motherOf ?mom} RDF graph set = Family In this example, RT_star is related to... Figure 1 The same star schema is associated with the data in a graphical representation. In SPARQL, results produced by a star schema are allowed to have missing values. Therefore, after deduplication, the result of performing the PRT schema {?x :firstName ?fnm . ?x :lastName ?lnm . ?x ^:fatherOf ?dad . ?x ^:motherOf ?mom} on the Family RDF atlas will be... Figure 6 Table 600 is shown in the diagram. As illustrated in Table 600, the rows for :beth are missing both the ^:fatherOf and ^:motherOf edges. On the other hand, the rows for :rick have all the edges specified in the star pattern.
[0036] In Table 600, column "C<:fname>" indicates the number of instances of the "fname" edge represented in that row within the Family RDF atlas for each row. Column "C<:lname>" indicates the number of instances of the "lname" edge represented in that row within the Family RDF atlas for each row. Column "C<:fatherOf>" indicates the number of instances of the "^:fatherOf" edge represented in that row within the Family RDF atlas for each row. Column "C<:motherOf>" indicates the number of instances of the "^:motherOf" edge represented in that row within the Family RDF atlas for each row.
[0037] PRT for chained mode
[0038] For the purposes of explanation, assume that a PRT named "RT_chain" has the following properties: PRT pattern = {?gma :motherOf ?mom . ?mom : motherOf ?kid . ?kid ^:fatherOf ?dad . ?dad ^:fatherOf ?gpa} RDF graph set = Family In this example, RT_chain is related to... Figure 3 The chained patterns shown are related. In SPARQL, results produced by chained patterns are not permitted to have any empty links. Therefore, after deduplication, the results produced by performing PRT patterns on an RDF atlas will be... Figure 7 The single row contained in Table 700 shown.
[0039] In Table 700, column "C<:motherOf>" indicates the number of instances of the ":motherOf" edge represented in that row within the Family RDF atlas for each row. Column "C<:motherOf>#2" indicates the number of instances of the ":motherOf#2" edge (the second instance of the ":motherOf" property) represented in that row within the Family RDF atlas for each row. Column "C<:fatherOf>" indicates the number of instances of the "^:fatherOf" edge represented in that row within the Family RDF atlas for each row. Column "C<:fatherOf>#2" indicates the number of instances of the "^:fatherOf#2" edge represented in that row within the Family RDF atlas for each row.
[0040] Improve query performance by using PRT
[0041] Once a PRT has been created for a specific schema / RDF atlas combination, it can be used to improve the performance of SPARQL queries such as: Targeting the RDF atlas associated with PRT, and For any type of PRT, specify a pattern that is the same as or a superset of the PRT pattern. For star PRTs, specify the pattern that is a subset of the PRT patterns of that PRT. The work of joining and deduplicating edges from RDF tables derived from RDF atlases is performed during PRT creation. Therefore, when using PRT to answer SPARQL queries, it is not necessary to repeat the join and deduplication that would otherwise be required to obtain the results of the PRT pattern during the execution of a SPARQL query containing the PRT pattern.
[0042] Complex SPARQL queries can include many schemas. Therefore, a single SPARQL query may be a superset of several PRT schemas. Figure 8 The image below illustrates an example of this query. (Reference) Figure 8 Figure 800 illustrates a query 800 requesting the names, surnames, and games played by the (maternal) grandmother and (paternal) grandfather for each individual represented in the Family RDF atlas. Figure 802 shows how query 800 can be broken down into sub-patterns that match several PRTs already created for the Family RDF atlas.
[0043] exist Figure 8 In the example shown, query 800 is broken down into the following sub-patterns: · ?gma :plays ?gmaGame · ?gpa :plays ?gpaGame · ?gma :motherOf ?mom . ?mom :motherOf ?kid . ?kid ^:fatherOf ?dad .?dad ^:fatherOf ?gpa · ?gma :fname ?gmaFnm . ?gma :lname ?gmaLnm · ?gpa :fname ?gpaFnm . ?gpa :lname ?gpaLnm Instead of performing the joins and deduplication required for the subpattern "?gma :plays ?gmaGame" against the Family RDF atlas, the pre-computed results for the subpattern "?gma :plays ?gmaGame" can be obtained from the PRT table named RT_triple mentioned above. Similarly, instead of performing the joins and deduplication required for the subpattern "?gpa :plays ?gpaGame" against the Family RDF atlas, the pre-computed results for the subpattern "?gpa :plays ?gpaGame" can also be obtained from RT_triple.
[0044] The subpattern "?gma :motherOf ?mom . ?mom :motherOf ?kid . ?kid ^:fatherOf ?dad . ?dad ^:fatherOf ?gpa" matches the PRT pattern of the RT_chain PRT mentioned above. Therefore, instead of performing the join and deduplication required for subpattern X against the Family RDF atlas, the pre-computed results for this subpattern can be obtained from the RT_chain.
[0045] The subpattern "?gma :fname ?gmaFnm . ?gma :lname ?gmaLnm" is a subset of the PRT patterns of the aforementioned RT_star PRT. Therefore, instead of performing the join and deduplication required for the Family RDF atlas for the subpattern "?gma :fname ?gmaFnm . ?gma :lname ?gmaLnm", the pre-computed results for this subpattern can be obtained from RT_star. Similarly, the subpattern "?gpa :fname ?gpaFnm . ?gpa :lname ?gpaLnm" is a subset of the PRT patterns of the aforementioned RT_star PRT. Therefore, instead of performing the join and deduplication required for the Family RDF atlas for the subpattern "?gpa :fname ?gpaFnm . ?gpa :lname ?gpaLnm", the pre-computed results for this subpattern can be obtained from RT_star.
[0046] Because the results of query 800 for each subschema can be obtained from the PRTs of the Family RDF atlas instead of the RDF tables of the Family RDF atlas, the database server is able to answer query 800 by performing joins between the PRTs. The overhead required to join the PRTs to answer query 800 (four joins) is significantly less than the overhead required to perform joins between the underlying RDF tables (nine joins). Furthermore, those PRTs can be used even when the schema involves multiple instances of the same property (e.g., :motherOf) and both forward and reverse properties (e.g., ":motherOf" and "^:fatherOf").
[0047] Maintain PRT
[0048] PRT can only be used to speed up SPARQL queries accessing RDF atlases if it would produce accurate results. For PRT to produce accurate results, it must be maintained in a way that reflects the current contents of the RDF tables belonging to the corresponding RDF atlas. For example, changing any RDF table belonging to the Family RDF atlas (in...) Figure 4 DML operations on the content (as shown in the diagram) may require changes to one or more PRTs (i.e., RT_triple, RT_star, and RT_chain) built for the Family RDF atlas.
[0049] In one implementation, if the PRT includes the property :p (or the reverse property ^:p) of the triple { :s :p :o} that is being deleted or inserted, then the PRT is maintained as follows: Star-shaped PRT: ○ Delete: Decrease the COUNT cell corresponding to the triplet (if the current value is 1, then set it to NULL). ○ Insert: If no row exists for this triple, add a row. Increment COUNT (set to 1 if currently NULL). Chain-type PRT: ○ Delete: Decrease the COUNT cell corresponding to the triplet (if the current value is 1, then delete the row). ○ Insert: Adds a row for each new chain created with the new triplet. Increments COUNT (sets to 1 if currently NULL). Simple triplet PRT: ○ Delete: Decrease the COUNT cell corresponding to the triplet (if the current value is 1, then delete the row). ○ Insert: If no row exists for this triple, add a row. Increment COUNT (set to 1 if currently NULL). Now, referring to its initial content... Figure 5 The Family RDF atlas is illustrated graphically, and an example of PRT maintenance is given.
[0050] Insertion of repeating triples
[0051] As explained above, duplicate triples are not allowed in RDF tables. However, it is permissible to insert a triple that is exactly the same as a triple residing in another RDF table belonging to the same RDF atlas into an RDF table. For illustrative purposes, it will be assumed that the triple ":bob :fatherOf :dan" is inserted into table Dan 418 (in... Figure 4 (See diagram in the middle). Inserting this triple into Dan table 418 is not prohibited because an instance of the triple does not currently exist in Dan table 418. However, two instances of this triple already exist in the Family RDF atlas: one in Relations table 410 and one in Bob table 420.
[0052] In response to the insertion of a duplicate triplet, the database server finds all PRT rows containing triples for which duplicates have been inserted. For the purpose of finding matching triples in the PRT, semantically equivalent triples are considered identical. Therefore, in the current case, the triple “:dan ^:fatherOf :bob” is considered identical to “:bob :fatherOf :dan”.
[0053] Once a PRT row containing a triplet that semantically matches the newly inserted triplet is identified, the database server increments the count associated with the matching triplet in those rows. In this example, both RT_chain and RT_star have rows containing the semantically equivalent triple ":bob :fatherOf :dan". Therefore, in response to the insertion of an instance of ":bob :fatherOf :dan" into table 418 of Dan, the database server updates the relevant rows of RT_chain and RT_star to increment the count associated with the matching triplet, as shown below. Figure 9 As shown in the image.
[0054] Insertion of the new triplet
[0055] The insertion of new triples (triples that do not yet exist in any RDF table in the RDF atlas) is handled differently from the insertion of duplicate triples (triples that already have at least one instance in the RDF atlas). For the purpose of illustrating how the PRT is maintained in response to the insertion of new triples, it will be assumed that the triple ":sue :motherOf :sam" is inserted into table Sue 414. In this example: The triple “:sue :motherOf :sam” did not previously exist in any RDF table of the Family RDF atlas, and The new triplet involves a new anchor (":sam") that was not previously present in the Family RDF atlas. The database server responds to the insertion of a new triplet that introduces a new anchor by inserting a row for the new anchor in all simple triplet PRTs and star PRTs that include the following: The same properties as the new triple (e.g., ":motherOf"), or The inversion of properties in the new triplet (e.g., "^:motherOf") In this example, RT_star includes the attribute "^:motherOf". Therefore, in response to ":sue :motherOf :sam" being inserted into table Sue 414, a new row is added to RT_star. In the new row, all attribute columns used for the new anchor are NULL except for the attributes specified in the new triplet. Figure 10 This is a diagram illustrating the row that will be inserted into RT_star in response to the insertion of ":sue :motherOf :sam" into table 414 of Sue.
[0056] The insertion of a new triple may also require the database server to update the chained PRT. Specifically, in response to the insertion of a triple that creates a new chain matching the chained pattern of the chained PRT, the database server adds a new row to the chained PRT.
[0057] For illustrative purposes, it will be assumed that after adding the triple “:sue :motherOf :sam”, the triple “:dan fatherOf :sam” is added to table 418 of Dan. The addition of the triple “:dan fatherOf :sam” does not involve a new anchor (“dan” and “sam” are already in the Family RDF atlas).
[0058] The addition of the triple “:dan fatherOf :sam” creates a new chain that matches the chain pattern {?gma :motherOf ?mom . ?mom : motherOf ?kid . ?kid ^:fatherOf ?dad . ?dad ^:fatherOf?gpa} in RT_chain. Therefore, the insertion of the triple “:dan fatherOf :sam” causes the database server to add a row to RT_chain for this new chain.
[0059] In addition to any updates to the linked list, adding a new triplet may also require updating existing rows in the star PRT. In this example, the insertion of the triple ":dan fatherOf :sam" creates a triple associated with the "^:fatherOf" property that was previously missing in the RT_star row for the anchor ":sam". Therefore, the insertion of the triple ":dan fatherOf :sam" causes the RT_star row associated with the anchor "sam" to be updated to reflect the new object for the "^:fatherOf" property of ":sam". Updates to RT_chain and RT_star in response to the insertion of the triple ":dan fatherOf :sam" are made in... Figure 11 The diagram in the middle is shown.
[0060] Deletion of duplicate triples
[0061] In response to the removal of a duplicate triple from an RDF atlas, the PRTs in the atlas containing rows reflecting that triple must be updated to decrement the triple's count. For example, suppose an instance of the triple ":beth :motherOf :sue" is removed from the Family RDF atlas. The triple ":beth :motherOf :sue" is reflected in two rows of RT_chain and one row of RT_star. Therefore, removing an instance of the triple ":beth :motherOf :sue" from the Family RDF atlas will result in the counts of three PRT rows being updated, as shown below. Figure 12 As shown in the image.
[0062] Deletion of the unique triple
[0063] Unlike the deletion of duplicate triples, the deletion of unique triples can break links in the chains reflected in a chained PRT. For example, suppose that after deleting one instance of the triple ":beth :motherOf :sue", the last remaining instance of the triple ":beth :motherOf :sue" is also deleted. The triple ":beth :motherOf :sue" is part of two chains reflected in the RT_chain (see [link]). Figure 12 Since chained PRTs cannot have chains with missing links, the two RT_chain lines based on the ":beth :motherOf :sue" triple must be deleted after deleting the last instance of the ":beth :motherOf :sue" triple.
[0064] Furthermore, when the last instance of a triple is deleted, any rows in the star PRT that reflect that triple must be updated so that the columns of the star PRT table rows that reflect that triple no longer reflect it. Because star PRT rows are allowed to contain null values, the rows reflecting the deleted triple are not deleted. Instead, the values of the relevant columns are set to NULL, and the count value associated with that triple is set to zero or NULL. Figure 13 The illustration shows the changes performed by the database server in response to the deletion of the last instance of the ":beth:motherOf :sue" triple. If setting the relevant column to NULL would cause any star PRT table row to contain only NULL values, then (optionally) that row can be deleted from the star PRT table.
[0065] Maintain PRT during batch update operations
[0066] Performing necessary PRT maintenance in response to the insertion / deletion of each individual triple in a bulk load operation adds a significant amount of overhead to the bulk load operation. Therefore, in one implementation, the database server ignores PRTs during the bulk load operation and rebuilds them once the bulk load operation has completed. Specifically, in one embodiment, at the start of the bulk load operation, any PRTs affected by the bulk load operation are made invisible to the SPARQL query translator. Thus, the metadata of the PRTs remains intact within the database, but the insertion / deletion during the bulk load does not result in any maintenance of the PRTs.
[0067] Because no maintenance was performed on the PRT, its contents no longer reflected the contents of the affected RDF atlas. Therefore, after the bulk load operation completes, the PRT is truncated and then repopulated. After the PRT has been repopulated to reflect the current contents of the RDF atlas, it becomes visible to the SPARQL query translator again.
[0068] Hardware Overview
[0069] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to execute the techniques, or may include digital electronic devices permanently programmed to execute the techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to execute the techniques according to program instructions in firmware, memory, other storage devices, or combinations thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement the techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hardwired and / or program logic to implement the techniques.
[0070] For example, Figure 14 This is a block diagram illustrating a computer system 1400 on which embodiments of the present invention may be implemented. The computer system 1400 includes a bus 1402 or other communication mechanism for transmitting information and a hardware processor 1404 coupled to the bus 1402 for processing information. The hardware processor 1404 may be, for example, a general-purpose microprocessor.
[0071] Computer system 1400 also includes main memory 1406, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 1402 for storing information and instructions to be executed by processor 1404. Main memory 1406 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 1404. When such instructions are stored in non-transitory storage media accessible to processor 1404, such instructions make computer system 1400 a dedicated machine customized to perform the operations specified in the instructions.
[0072] Computer system 1400 also includes a read-only memory (ROM) 1408 or other static storage device coupled to bus 1402 for storing static information and instructions of processor 1404. Storage device 1410, such as a disk, optical disk, or solid-state drive, is provided and coupled to bus 1402 for storing information and instructions.
[0073] Computer system 1400 may be coupled via bus 1402 to a display 1412, such as a cathode ray tube (CRT), for displaying information to a computer user. Input device 1414, including alphanumeric keys and other keys, is coupled to bus 1402 for transmitting information and command selections to processor 1404. Another type of user input device is cursor control 1416, such as a mouse, trackball, or cursor arrow keys, for transmitting directional information and command selections to processor 1404 and for controlling cursor movement on display 1412. Such input devices typically have two degrees of freedom on two axes (a first axis (e.g., x) and a second axis (e.g., y)) to allow the device to specify a position in a plane.
[0074] Computer system 1400 may implement the techniques described herein using custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic, which, in combination with the computer system, make computer system 1400 a special-purpose machine or program the computer system 1400 as a special-purpose machine. According to one embodiment, the techniques herein are executed by computer system 1400 in response to processor 1404 executing one or more sequences of one or more instructions contained in main memory 1406. These instructions may be read into main memory 1406 from another storage medium, such as storage device 1410. Execution of the sequence of instructions contained in main memory 1406 causes processor 1404 to perform the processing steps described herein. In alternative embodiments, hard-wired circuitry may be used instead of software instructions or in combination with software instructions.
[0075] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, or solid-state drives, such as storage device 1410. Volatile media include dynamic memory, such as main memory 1406. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a triple-hole pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips, or magnetic tape cassettes.
[0076] Storage media are distinct from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including the conductors that form bus 1402. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0077] Various forms of media can involve carrying one or more sequences of one or more instructions to processor 1404 for execution. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 1400 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 1402. Bus 1402 carries the data to main memory 1406, from which processor 1404 retrieves and executes the instructions. The instructions received by main memory 1406 may optionally be stored on storage device 1410 before or after execution by processor 1404.
[0078] Computer system 1400 also includes a communication interface 1418 coupled to bus 1402. Communication interface 1418 provides bidirectional data communication coupling to network link 1420, which is connected to local network 1422. For example, communication interface 1418 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem providing data communication connectivity to a corresponding type of telephone line. As another example, communication interface 1418 may be a LAN card providing data communication connectivity to a compatible local area network (LAN). A wireless link may also be implemented. In any such implementation, communication interface 1418 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0079] Network link 1420 typically provides data communication to other data devices via one or more networks. For example, network link 1420 may provide a connection to host computer 1424 or to data devices operated by Internet Service Provider (ISP) 1426 via local network 1422. ISP 1426, in turn, provides data communication services via a global packet data communication network now commonly referred to as the "Internet" 1428. Both local network 1422 and Internet 1428 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks, as well as signals on network link 1420 and through communication interface 1418, are example forms of transmission media that carry digital data to or from computer system 1400.
[0080] Computer system 1400 can send messages and receive data, including program code, through one or more networks, network links 1420, and communication interfaces 1418. In the Internet example, server 1430 can transmit the requested code of the application through the Internet 1428, ISP 1426, local network 1422, and communication interface 1418.
[0081] The received code can be executed by processor 1404 when it is received, and / or stored in storage device 1410 or other non-volatile storage device for later execution.
[0082] cloud computing
[0083] The term "cloud computing" is generally used in this article to describe a computing model that enables on-demand access to a shared pool of computing resources, such as computer networks, servers, software applications, and services, and allows for the rapid provisioning and release of resources with minimal management effort or service provider interaction.
[0084] Cloud computing environments (sometimes called cloud environments or the cloud itself) can be implemented in a variety of different ways to best meet different requirements. For example, in a public cloud environment, the underlying computing infrastructure is owned by an organization that makes its cloud services available to other organizations or the public. In contrast, private cloud environments are generally used only by a single organization or within a single organization. Community clouds are designed to be shared by several organizations within a community; while hybrid clouds include two or more types of clouds (e.g., private, community, or public clouds) that are bound together by data and application portability.
[0085] Generally, the cloud computing model enables functions that were previously provided by an organization's own IT department to be delivered as a service layer in the cloud environment for consumer use (depending on the public / private nature of the cloud, or whether it is within or outside the organization). Depending on the specific implementation, the precise definition of the components or functions provided by or within each cloud service layer can vary, but common examples include: Software as a Service (SaaS). SaaS (This refers to the use of software applications by consumers that run on cloud infrastructure, while...) SaaS Providers manage or control the underlying cloud infrastructure and applications. Platform as a Service (PaaS) allows consumers to develop, deploy, and otherwise control their own applications using software programming languages and development tools supported by the PaaS provider, while the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything below the runtime execution environment). Infrastructure as a Service (IaaS) allows consumers to deploy and run arbitrary software applications and / or provision processing, storage, networking, and other basic computing resources, while the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS) allows consumers to use database servers or database management systems running on cloud infrastructure, while the DBaaS provider manages or controls the underlying cloud infrastructure, applications, and servers, including one or more database servers.
[0086] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary between implementations. Therefore, the description and drawings should be considered illustrative rather than restrictive. The unique and exclusive indication of the scope of the invention, and what the applicant intends to be the scope of the invention, is the literal and equivalent scope of the set of claims in their specific form at the time of this application's grant, including any subsequent corrections.
Claims
1. A method comprising: Store multiple RDF tables in the database; Each of the plurality of RDF tables is populated with information representing the RDF graph; For a specific RDF atlas, generate and store one or more pre-computed result tables (PRTs) in the database. The specific RDF atlas includes RDF graphs represented by information from the plurality of RDF tables; Each of the one or more PRTs is associated with a distinctly different pattern; Each PRT is populated based on information from the specific RDF atlas and distinctly different pattern matching associated with each of the one or more PRTs; Filling each PRT involves: Perform a join between two or more RDF tables in the plurality of RDF tables, and Remove duplicate rows from the join result resulting from duplicate triples in the two or more RDF tables; Perform an update to the one or more PRTs such that the one or more PRTs reflect the changes made by the DML operation to the information in the plurality of RDF tables; Receive SPARQL queries targeting the specific RDF atlas; Identify the specific PRT associated with the pattern that matches the subpattern of the SPARQL query in one or more PRTs; Obtain the results of the sub-schema for the SPARQL query from the specific PRT; and The results for the SPARQL query are generated at least in part based on the results of the sub-patterns for the SPARQL query obtained from the specific PRT; The method is performed by one or more computing devices.
2. The method of claim 1, wherein the particular PRT is a chained PRT associated with a chained pattern.
3. The method of claim 2, wherein the chain pattern includes one or more reverse features.
4. The method of any one of claims 2-3, wherein the chained pattern comprises multiple instances of a particular characteristic.
5. The method of claim 1, wherein the particular PRT is a star PRT associated with a star pattern.
6. The method of claim 5, wherein the star pattern includes one or more reverse characteristics.
7. The method as described in any one of claims 1-6, wherein: The one or more PRTs include multiple PRTs; The method further includes identifying one or more additional PRTs associated with a pattern that matches a subpattern of the SPARQL query among the plurality of PRTs; as well as Generating the results of a SPARQL query includes generating the results of the SPARQL query based on information obtained from the one or more additional PRTs.
8. The method of any one of claims 1-6, wherein performing an update to the one or more PRTs comprises, in response to the insertion of a specific triplet whose instance already exists in the specific RDF graph, incrementing the count of the specific triplet in each row of the semantic equivalent of the specific triplet in each of the one or more PRTs.
9. The method of any one of claims 1-6, wherein performing an update to the one or more PRTs comprises inserting a specific triplet in response to an instance of that specific RDF graph not yet existing: The insertion of the specific triple creates a new chain that matches the chain pattern of the chained PRTs belonging to the one or more PRTs; and Insert a row for the new chain in the chained PRT.
10. The method of any one of claims 1-6, wherein performing an update to the one or more PRTs comprises inserting a specific triplet in response to an instance of the specific RDF graph not yet existing: The insertion of the specific triple was determined to have added a new anchor to the specific RDF atlas; and Insert the row associated with the new anchor in the star-shaped PRT that belongs to one or more of the PRTs.
11. The method of any one of claims 1-6, wherein performing an update to the one or more PRTs comprises, in response to the deletion of a specific triple unique in the particular RDF atlas: Deletion of the specific triple disrupts the link in the chain represented by a specific row in a chained PRT belonging to the one or more PRTs; and In response to determining that the deletion of the specific triplet breaks the link, the specific row is removed from the chained PRT.
12. The method of any one of claims 1-11, wherein performing an update to the one or more PRTs comprises, in response to the deletion of a specific triple that currently has multiple instances in the particular RDF graph, decrementing the count of the specific triple in each row of the semantic equivalent of the specific triple in each of the one or more PRTs.
13. The method of any one of claims 1-11, wherein performing an update to the one or more PRTs comprises, in response to a bulk load operation: Without deleting the metadata of the one or more PRTs, make the one or more PRTs invisible to the SPARQL query translator; Perform a batch loading operation when one or more PRTs are not visible; After the batch loading operation is complete: Truncate one or more PRTs; Refill the one or more PRTs; and This makes the one or more PRTs visible.
14. The method of any one of claims 1-13, wherein the pattern that matches the subpattern of the SPARQL query is a subpattern of a pattern that is distinctly different from the pattern associated with the particular PRT.
Citation Information
Patent Citations
Computer readable recording medium stored with control program for controlling tab sheet insertion apparatus and control method thereof
US20080199200A1