Performing SPARQL queries on RDF datasets stored in distributed storage

By using file storage and directories to implement ACID attributes in the virtual RDF graph database, the problem of updating and deleting triplets when performing SPARQL queries on heterogeneous distributed storage in the prior art is solved, and efficient RDF data management and query are achieved.

CN120196600APending Publication Date: 2025-06-24DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411881364.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to perform SPARQL queries under ACID guarantee, especially to update and delete triplets on heterogeneous distributed storage.

Method used

Provides a computer-implemented method to implement ACID properties by using file storage and directories in a virtual RDF graph database, ensuring consistent writes, and applying batch tuples on a second read-only RDF graph database to form a read-only and updating snapshot.

Benefits of technology

It realizes the ability to execute SPARQL queries under ACID guarantee, including updating and deleting triples, and improves the management and query efficiency of RDF data in heterogeneous distributed storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196600A_ABST
    Figure CN120196600A_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for updating a virtual RDF graph database. The method comprises the following steps of: providing file storage which has durability D in ACID attributes and ensures consistent writing; providing a virtual RDF graph database including a first RDF graph database updatable by the stream and a second read-only RDF graph database stored on the file store and updated in batches; a directory is provided for storing metadata describing a second read only (RDF) graph database on a file store, the directory conforming to the ACID. The method further includes obtaining a tuple stream and a batch of tuples, applying the tuple stream to the first RDF graph database, applying the batch of tuples to the second RDF graph database by calculating a snapshot of the second RDF graph database, the snapshot including the batch of tuples, calculating a sequence of concurrency control operations appropriately performed by using consistent writes stored by the file to ensure ACID attributes of the snapshot; storing the computed snapshot on a file store; and registering the computed snapshot in the directory, thereby obtaining an updated description of the virtual RDF graph database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly to methods, systems, and programs for updating an RDF graph database partitioned on heterogeneous storage. Background Art

[0002] The RDF specification (https: / / www.w3.org / TR / rdf11-concepts / ) has been published by the W3C, which represents information as a graph using triples: "The core structure of the abstract syntax is a set of triples, each triple consisting of a subject, a predicate, and an object. A collection of such triples is called an RDF graph. An RDF graph can be visualized as a graph of nodes and directed arcs, where each triple is represented as a node-arc-node link". The RDF graph stores RDF data, which, according to the previous definition, refers to the triples within the RDF graph.

[0003] SPARQL is a query language for RDF data (https: / / www.w3.org / TR / sparql11-query / ): "RDF is a directed, labeled graph data format for representing information on the Web. This specification defines the syntax and semantics of the SPARQL query language for RDF. SPARQL can be used to express queries across various data sources, whether the data is natively stored in RDF format or viewed in RDF format via intermediate software. SPARQL includes query-required and optional graph patterns and their ability to be conjoined and disjoined. SPARQL also supports aggregation, subqueries, negation, creation of values through expressions, extensible value testing, and querying by constraining the source RDF graph. The result of a SPARQL query can be a result set or an RDF graph."

[0004] RDF knowledge graphs have grown to billions of triples. To be able to scale to this data volume, computing and storage must be allocated for SPARQL queries on RDF data. In a cloud environment, computing and storage have different cost models; therefore, similar to Snowflake for SQL data, they are both distributed and scaled separately. Computing and storage are separated and distributed differently.

[0005] Now discuss storage distribution.

[0006] In the field of RDF, several techniques can be used to determine how to partition data in a distributed storage. This is discussed, for example, in Kalogeros, E., Gergatsoulis, M., Damigos, M. and Nomikos, C., 2023, Efficient query evaluation techniques over large amount of distributed linked data, Information Systems, 115, p. 102194.

[0007] A common approach such as in SPARQLGFX is disclosed in Graux, D., Jachiet, L., Genevès, P. and Layaida, N., 2016, Sparqlgx: Efficient distributed evaluation of sparql with apache spark, In The Semantic Web - ISWC 2016: 15th International Semantic Web Conference, Kobe, Japan, October 17 - 21, 2016, Proceedings, Part II 15 (pp. 80 - 87). SPARQLGFX uses a distributed file system such as the Hadoop File System (HDFS) or Amazon's Simple Storage Service (S3). A comparison between the two can be found here: https: / / www.integrate.io / blog / storing - apache - hadoop - data - cloud - hdfs - vs - s3 / . Due to the reasons listed in the article, the industry pays more attention to S3 than HDFS.

[0008] Many distributed RDF databases are read-only or do not consider how to modify partitions. "Partout", discussed in Galárraga, L., Hose, K., and Schenkel, R., April 2014, Partout: a distributed engine for efficient RDF processing, In Proceedings of the 23rd International Conference on World Wide Web (pp. 267-268), accepts updates to the graph but has no ACID guarantees. "Partout" accepts updates, but only batch updates, making it impossible to make modifications in streaming and no modifications are immediately visible to queries.

[0009] The ACID properties (ACID is an acronym for Atomicity, Consistency, Isolation, and Durability) are a set of properties of database transactions designed to ensure that data remains valid in the presence of errors, power failures, and other contingencies. For example, the ACID properties make it possible to start a transaction and, during that transaction, all SPARQL queries see the same version of the data even if the data is updated concurrently in one or more partitions. For example, the ACID properties are discussed here: https: / / en.wikipedia.org / wiki / ACID.

[0010] The literature Zou, L. and M.T., 2017, Graph-based RDF data management, Data Science and Engineering, 2, pp. 56-70 and the literature Ali, W., Saleem, M., Yao, B., Hogan, A., and Ngomo, A.C.N., 2022, A survey of RDF stores & SPARQL engines for querying knowledge graphs, The VLDB Journal, pp. 1-26 disclose recent approaches for executing SPARQL queries on distributed storage and can be regarded as a comparison of NoSQL approaches with traditional SQL approaches, as in the literature Valduriez, P., Jiménez-Peris, R., and Defined in M.T., 2021, Distributed database systems: The case for newSQL, In Transactions on Large-Scale Data-and Knowledge-Centered Systems XLVIII: Special Issue, In Memory of Univ.Prof.Dr.Roland Wagner (pp. 1-15).

[0011] In the area of relational SQL distributed storage, similar requirements have emerged, with the difference that relational databases deal with schematized tables rather than schema-less graphs in RDF. Thus, the techniques used in the area of relational databases cannot be applied as such to the area of graph databases; in fact, relational databases and graph databases rely on different paradigms.

[0012] Apache Iceberg “is an open table format for large analytical datasets. Iceberg adds tables to a compute engine using a high-performance table format that works just like an SQL table [...]”. As disclosed in the Iceberg Table Spec of the online Apache Iceberg version 1.3.0 (https: / / iceberg.apache.org / spec / ), “the Iceberg table format” aims to manage large, slowly changing collections of files in a distributed file system or key-value store in the form of a table”. In short, an Iceberg table consists of several binary files stored in a distributed storage and is accessed via metadata files that give a list of files corresponding to a snapshot of the state of the table at a given time. This concept is similar to that in the literature by Zou, L. and M.T., 2017, Graph-based RDF data management, DataScience and Engineering, 2, pp. 56-70, and the concept seen in the literature Ali, W., Saleem, M., Yao, B., Hogan, A., and Ngomo, A.C.N., 2022, A survey of RDF stores & SPARQL engines for querying knowledge graphs, The VLDB Journal, pp. 1-26. However, Iceberg provides ACID guarantees due to metadata and manifest files. However, since Iceberg is completely file-based only, Iceberg does not support streaming modifications and only supports batch updates.

[0013] Snowflake is discussed in Snowflake Key Concepts & Architecture, 2023 Snowflake Inc., https: / / docs.snowflake.com / en / user-guide / intro-key-concepts. Snowflake provides an ability called Snowpipe Streaming. It combines in a table Snowpipe, which loads data from files in a micro-batch manner similar to Iceberg, and a streaming API that loads streaming data rows with low latency. It seems that Snowflake has ACID transaction semantics, but Snowflake is a commercial product and little detail is available. The limitation of Snowflake is that it only supports the insertion of rows.

[0014] In this case, there is still a need for an improved RDF heterogeneous distributed storage mainly composed of non-ACID compatible stores that can execute SPARQL queries, including updating / deleting triples, under ACID guarantees. Summary of the Invention

[0015] Accordingly, a computer-implemented method for updating a virtual RDF graph database is provided, the virtual RDF graph database including tuples. The method includes:

[0016] - Providing:

[0017] -- A file storage having durability D among the ACID properties, the file storage being distributed or not distributed, and the file storage guaranteeing consistent writes;

[0018] -- A virtual RDF graph database, the virtual RDF graph database including:

[0019] ---A first RDF graph database that can be updated by a stream of tuples to be added to and / or removed from the first RDF graph database;

[0020] ---A second read-only RDF graph database stored on a file store and can be updated by multiple batches of tuples to be added to and / or removed from the second RDF graph database, thereby forming a read-only and updatable snapshot;

[0021] --A directory for storing metadata describing the second read-only RDF graph database on the file store, and the directory conforms to ACID;

[0022] -Obtain a stream of tuples to be added to and / or removed from the virtual RDF graph database, and obtain a batch of tuples to be added to and / or removed from the virtual RDF graph database;

[0023] -Apply the tuple stream to the first RDF graph database of the virtual RDF graph database;

[0024] -Apply the batch of tuples to the second RDF graph database of the virtual RDF graph database by:

[0025] --Calculate a snapshot of the second RDF graph database, where the snapshot includes the batch of tuples, and calculate a sequence of concurrent control operations appropriately executed by using consistent writes of the file store to ensure the ACID properties of the snapshot;

[0026] --Store the calculated snapshot on the file store; and

[0027] --Register the calculated snapshot in the directory, thereby obtaining an update description of the virtual RDF graph database.

[0028] The method may include one or more of the following:

[0029] - The ALTER REFRESH command executes a sequence of appropriate concurrent control operations. The ALTER REFRESH command ensures the ACID properties of the updates on the second RDF graph database of the virtual RDF graph database. The sequence of the ALTER REFRESH's concurrent control operations includes: uniquely identifying the most recent snapshot in the second RDF graph database, where the most recent snapshot represents the state of the snapshot formed by the second RDF graph database at a given time; verifying that no other ALTER REFRESH commands are being executed simultaneously, otherwise stopping the ALTER REFRESH command and preserving the ALTER REFRESH command that is already being executed; obtaining a list of files in the phase that includes all snapshots including the most recent snapshot; filtering the files in the list of files in the phase of the most recent snapshot, so as to obtain, at a given time, a snapshot list naming the files in the most recent snapshot, and the files in the most recent snapshot are available for subsequent access to the database; writing the names of the files in the snapshot list, so as to obtain a new snapshot that is not registered in the directory; uploading the files of the new snapshot to the phase; registering the new snapshot in the directory after the directory has identified the new snapshot as the most recent snapshot to which the batch of tuples is to be applied in an ACID transaction;

[0030] - Obtaining, at a given time, a snapshot list naming the files in the most recent snapshot, where the files in the most recent snapshot are available for subsequent access to the database, includes: obtaining a list of the names of the valid partitions; for each partition named in the list of valid partitions: verifying whether the current update operation on the partition is still pending, and if so, adding the partition to the list of valid partitions for which the current update operation is still pending; verifying whether the past update operation on the partition has not been successfully implemented, and if so, verifying in the directory whether a snapshot for which an unsuccessful update operation has been executed is registered in the directory: if not, ignoring the unsuccessful update operation; if so, removing the partition from the list of valid partitions for which the update operation is pending; creating an empty snapshot list file and completing the snapshot list file by adding the following items: the names of the files of the valid partitions in the list of valid partitions for which the current update operation is still pending; otherwise, the names of the files of the valid partitions in the list of the names of the valid partitions in the case where the list of valid partitions for which the previous update operation is still pending is empty; so as to obtain, at a given time, a snapshot list that names the files in the most recent snapshot that will be available for subsequent access to the database; and verifying that each file listed in the snapshot list exists in the phase.

[0031] - Obtaining a list of names of valid partitions includes: for each partition, searching in the file store for files named with the partition name and a specific extension called VALID_EXTENSION, where the VALID_EXTENSION files store the names of the files of the partition, and the VALID_EXTENSION files have been uploaded to the file store when the partition was created; and / or obtaining a list of names of valid partitions includes: for each partition, searching in the file store for files named with the partition name and a specific extension called TOMBSTONE_EXTENSION, where the TOMBSTONE_EXTENSION files store the names of the files in the partition that are the objects of the current deletion operation, and the TOMBSTONE_EXTENSION files have been uploaded to the file store when the deletion operation of the partition was executed; and / or for each partition named in the list of valid partitions, verifying whether the current update operation for the partition is still pending includes: for each partition, searching in the file store for files named with the partition name and a specific extension called PENDING_EXTENSION, where the PENDING_EXTENSION files store the names of the files in the partition that are the objects of the current update operation, and the PENDING_EXTENSION files have been uploaded to the file store when the current update of the partition was executed; and / or for each partition named in the list of valid partitions, verifying whether a past update operation for the partition has not been successfully implemented includes: for each partition, searching in the file store for files named with the partition name and a specific extension called CONSUMED_EXTENSION, where the CONSUMED_EXTENSION files store the names of the files in the partition that are the objects of the past update operation, and the CONSUMED_EXTENSION files have been uploaded to the file storage device when the past update of the partition was executed;

[0032] - For each partition named in the list of valid partitions, verifying whether a past update operation for the partition has not been successfully implemented further includes: verifying that the content of the PENDING_EXTENSION file is consistent with the content of the VALID_EXTENSION file, and if not, considering the list of files of the most recent snapshot to be the content of the VALID_EXTENSION file; and wherein, for each partition named in the list of valid partitions, if the current update operation for the partition is still pending, then the PENDING_EXTENSION file is ignored for this verification;

[0033] Uniquely identifying the most recent snapshot in the second RDF graph database includes: obtaining the most recent dataset snapshot identifier from the catalog and saving the most recent dataset snapshot identifier in memory as the previous dataset snapshot; and wherein registering a new snapshot in the catalog further includes: validating, by the catalog in an ACID transaction, that the most recent snapshot described in the catalog is the previous dataset snapshot; if so, registering the new snapshot in the catalog as the most recent snapshot to which the batch of tuples is applied; if not, determining that a concurrent ALTER REFRESH command has been executed and removing the new snapshot file from that stage;

[0034] - The ALTER PARTITION command performs a further sequence of concurrent control operations for updating the partitions of the second read-only RDF graph database. The ALTER PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrent control operations for ALTER PARTITION includes: downloading the VALID_EXTENSION file of the valid partition to be updated from that stage to obtain the name of the file of the valid partition; validating that no ALTER PARTITION command is being executed concurrently on the valid partition, otherwise stopping the ALTER PARTITION command and leaving the ALTER PARTITION command that is already in progress; creating a PENDING_EXTENSION file that stores the name of the file that is the object of the current update operation in the valid partition, where the name of the file in the valid partition is the object of the current update operation and is obtained from the batch of tuples; and uploading the PENDING_EXTENSION file to that stage;

[0035] - A snapshot of the second RDF graph database includes a set of batch partitions, and one or more batch partitions are split into fragments. Creating a PENDING_EXTENSION file includes: determining whether the batch of tuples is a change data capture (CDC) file and performing the following operations: if the batch of tuples is not a CDC file, generating fragments of the batch of tuples to obtain a list of fragment files; otherwise if the batch of tuples is a CDC file, for each add / delete triple of the CDC file, locating the fragment of one batch partition in one or more batch partitions to which the add / delete operation will be applied to obtain a list of fragments with an additional list of delta files, if any; and wherein creating a PENDING_EXTENSION file includes: creating a PENDING_EXTENSION that stores the list of fragments with an additional list of delta files, if any;

[0036] - The ADD PARTITION command executes a sequence of further concurrency control operations for adding a partition to the second read-only RDF graph database. The ADD PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrency control operations for ADD PARTITION includes: uploading the VALID_EXTENSION file of the partition to be added to the storage. A successful upload of the VALID_EXTENSION file indicates that there is no partition with the same name in the file storage; generating fragments of the batch of tuples for adding the partition to the second read-only RDF graph database, thereby obtaining a list of fragment files; storing the list of fragments in the VALID_EXTENSION file and uploading the VALID_EXTENSION file to this stage;

[0037] - The REMOVE PARTITION command executes a sequence of further concurrency control operations for removing a partition of the second read-only RDF graph database. The REMOVE PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrency control operations for REMOVE PARTITION includes: verifying at this stage whether there is a partition to be removed. If not, the REMOVE PARTITION command is stopped; uploading the TOMBSTONE_EXTENSION file to this stage. A successful upload confirms the existence of the partition to be removed in the second read-only RDF graph database;

[0038] - Latch the second read-only RDF graph database stored on the file storage. Latching is performed before obtaining the list of files of the most recent snapshot at this stage; unlatch the new snapshot after it is registered in the directory;

[0039] - Latching the second read-only RDF graph database includes: uploading the file named with the most recent dataset snapshot name and a specific extension called START_EXTENSION to the file storage; verifying the existence of the START_EXTENSION file in the file storage when uploading the file of the new snapshot at this stage and in each subsequent step; and wherein, unlatching the new snapshot after it is registered in the directory includes: verifying the existence of the START_EXTENSION file in the file storage and removing the START_EXTENSION file;

[0040] - The file storage is a distributed file storage and / or the first RDF graph database is stored on an in-memory data structure;

[0041] - The directory for storing metadata is stored on a database that ensures ACID properties and strong consistency.

[0042] Therefore, a computer-implemented method is provided for using isolation to ensure the execution of triple patterns of SPARQL queries on a previously updated virtual RDF graph database, the method comprising:

[0043] - Obtaining the most recent registered snapshot in the virtual RDF graph database described in the catalog;

[0044] - For the obtained triple pattern, executing the triple pattern on the first RDF graph database of the virtual RDF graph database and on the obtained most recent snapshot, thereby ensuring isolation of the execution of the triple pattern.

[0045] A computer program is also provided, which includes instructions for executing the method for update and / or the method for execution.

[0046] A computer-readable storage medium having the computer program recorded thereon is also provided.

[0047] A system including a processor coupled to a memory having the computer program recorded thereon is also provided.

[0048] An apparatus including a data storage medium having the computer program recorded thereon is also provided. The apparatus may form or be used as a non-transitory computer-readable medium, such as on SaaS (Software as a service), or other servers or cloud-based platforms, etc. The apparatus optionally includes a processor coupled to the data storage medium. Thus, the apparatus may wholly or partly form a computer system (e.g., the apparatus is a subsystem of the whole system). The system may also include a graphical user interface coupled to the processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Non-limiting examples will now be described with reference to the accompanying drawings, wherein:

[0050] - Figure 1 A flowchart showing an example of a method for updating a virtual RDF graph database is shown;

[0051] - Figure 2 A flowchart showing an example of a method for querying an updated virtual RDF graph database is shown;

[0052] - Figure 3 An example of a distributed architecture is shown;

[0053] - Figure 4 A table showing an example of instructions is shown;

[0054] - Figure 5 An example of how a modification of a triple can be represented is shown; and

[0055] - Figure 6An example of the system is shown. Detailed implementation

[0056] Refer to Figure 1The flowchart presents a computer-implemented method for updating a virtual RDF graph database. A "virtual RDF graph database" refers to a situation where at least two graphs are regarded as a logically equivalent RDF graph database, excluding relational databases that can be adjusted to expose the content of an arbitrary relational database as a knowledge graph, such as https: / / github.com / ontop / ontop. The virtual RDF graph database includes RDF tuples (or simply tuples). The method includes providing a file storage with the durability D among the ACID properties. The file storage can be distributed or non-distributed. The file storage guarantees consistent writes. "Consistent writes" relate to a consistency model that describes the conditions under which write operations from one client are visible to other clients. Consistent writes can also be referred to as "Strong Read-After-Write Consistency", which is discussed here: https: / / aws.amazon.com / fr / blogs / aws / amazon-s3-update-strong-read-after-write-consistency / . In the present invention, "consistent writes" means ensuring that readers see all previously committed writes. The method also includes providing a virtual RDF graph database. The virtual RDF graph database includes a first RDF graph database and a second read-only RDF graph database, thus forming a read-only and updatable snapshot. The first RDF graph database is updatable through a stream of tuples to be added to and / or removed from the first RDF graph database. The second read-only RDF graph database is stored on the file storage and is updatable through a batch of tuples to be added to and / or removed from the second RDF graph database. The method also includes providing a directory for storing metadata describing the second read-only RDF graph database on the file storage, and the directory conforms to ACID. Then, the method further includes obtaining a stream of tuples to be added to and / or removed from the virtual RDF graph database, and obtaining a batch of tuples to be added to and / or removed from the virtual RDF graph database. The method also includes applying the tuple stream to the first RDF graph database of the virtual RDF graph database. In addition, the method includes applying the batch of tuples to the second RDF graph database of the virtual RDF graph database. Applying the batch of tuples is performed by: calculating a snapshot of the second RDF graph database including the batch of tuples, and the calculation guarantees the ACID properties of the snapshot through a sequence of concurrent control operations appropriately executed using the consistent writes of the file storage; storing the calculated snapshot on the file storage; and registering the calculated snapshot in the directory, thereby obtaining an updated description of the virtual RDF graph database.A catalog refers to a collection of metadata and information related to the resources of a virtual RDF graph database.

[0057] The advantages of the present method will be provided below.

[0058] Using the terms of the W3C from https: / / www.w3.org / TR / rdf11-concepts / , "RDF data" is defined as the triples within an RDF graph. According to https: / / www.w3.org / TR / sparql11-query / #rdfDataset, an "RDF dataset" is a collection of such RDF graphs, and an RDF database is a collection of RDF datasets. Still using the terms known in the art (https: / / en.wikipedia.org / wiki / Partition_(database)), "partitioning is the division of a logical database or its constituent elements into distinct, independent parts". The case of having overlapping partitions will be discussed later.

[0059] As discussed in the background art section, a database can be considered to consist of computing (for query processing) and database storage. In the present invention, computing and database storage are considered as separate resources of different scales.

[0060] The present invention focuses on database storage and considers the way of managing computing to be outside the scope of the present invention. The present invention realizes a way to obtain an improved RDF heterogeneous distributed storage mainly composed of non-ACID compatible storage, and the improved RDF heterogeneous distributed storage can execute SPARQL queries, including updating / deleting triples, under ACID guarantees, regardless of the way of managing computing. Incidentally, the computing can be distributed or not. For simplicity, it is considered that the computing part is not distributed, but the computing part can be distributed without changing the present invention, because the present invention is independent of the way of managing computing. In the case where the computing is distributed, the catalog is also distributed.

[0061] In addition, the RDF data of a given RDF dataset is partitioned in a distributed storage using a given partitioning rule, and the given partitioning rule can be, for example, hashing the subject. There are many ways to partition RDF data, which is not the subject of the present invention. The only requirement is to know which quad is in which partition. Therefore, it is considered that there are ways to probe partitions during query execution using statistical data, for example, to know whether the partition can be used in the execution process.

[0062] Now discuss the concept of quads. In an example, the RDF tuples of an RDF dataset can be RDF quads. RDF quads can be obtained by adding a graph label to an RDF triple. In such an example, the RDF tuple includes an RDF graph. The W3C has published a standard specification to specify RDF Quads (also known as N-Quads), see for example "RDF 1.1 N-Quads, A line-based syntax for RDF dataset", W3C Recommendation February 25, 2014. RDF quads can be obtained by adding a graph name to an RDF triple. The graph name can be empty (i.e., for the default or unnamed graph) or an IRI, i.e., a graph IRI. In the example, the predicate of the graph can have the same IRI as the graph IRI. The graph name of each quad is the graph to which the quad is part of in the corresponding RDF dataset. As is known per se, an RDF dataset (e.g., see https: / / www.w3.org / TR / rdf-sparql-query / #rdfDataset) represents a collection of graphs. For simplicity, a quad can be generalized as a triple that includes a reference to a graph.

[0063] Return reference Figure 1 , the method updates a virtual RDF graph database. As previously mentioned, the virtual RDF graph database includes at least two graphs, and these at least two graphs are regarded as a logically equivalent RDF graph database. Therefore, updating the virtual RDF graph database means updating the logical database.

[0064] Now discuss the provided step S100.

[0065] File storage

[0066] The method provides a file storage with the durability D among the ACID properties. As is known per se, "file storage" means storing and organizing data in the form of files within a file system. The data is grouped into separate files, each file is identified by a unique name or path, and these files are hierarchically organized within a directory or folder. As is known per se, the "durability property" of the ACID properties ensures that once a transaction is committed, the changes made to the data will persist and survive any subsequent failures (such as system crashes or power outages). This means that once a file write or file update operation is confirmed as successful, the data is securely stored and will not be lost even in the event of a failure. "The transaction is committed" means that the changes made by the transaction are completed, and the system ensures the durability of these changes and their visibility to other transactions.

[0067] File storage is either distributed or not. In a distributed file storage system, files and data are distributed across multiple servers or storage nodes. The distribution of files is typically for scalability, fault tolerance, and improved performance. For example, the Hadoop Distributed File System (HDFS), Google File System (GFS) are examples of distributed file systems. Many cloud storage solutions such as Amazon S3, Google Cloud Storage, and Microsoft Azure Storage are essentially distributed file storage. In a non - distributed file storage system (also known as centralized file storage), all files and data are stored on a single server or a limited set of servers. Users and clients access files by connecting to this central server. All file requests and data retrievals pass through this central point.

[0068] File storage ensures consistent writes. As is known per se, "consistent writes" refer to write operations that ensure data consistency, that is, ensuring that readers see the content of all previously committed writes.

[0069] Virtual RDF graph database

[0070] Still at step S100, a virtual RDF graph database is provided. As is known per se, a "virtual RDF graph database" refers to a database system that provides a virtual way to query and access RDF. The term "virtual" means that the database is able to query and retrieve data from multiple distributed sources or endpoints. It is to be understood that a virtual RDF graph database includes at least two RDF graph databases that are regarded as one logically equivalent RDF graph database; this qualification does not include relational databases that may be adapted to expose the content of any relational database as a knowledge graph, such as https: / / github.com / ontop / ontop. Without the need to merge all data into a single physical database, a virtual RDF graph database allows querying data from various distributed sources as if they were part of a single integrated database.

[0071] The virtual RDF graph database includes the two RDF graph databases now being discussed.

[0072] First RDF graph database

[0073] The virtual RDF graph database includes a first RDF graph database. The first RDF graph database is updatable via a stream of tuples to be added to and / or removed from the first RDF graph database. As is known per se, streaming modification is the way in which an RDF dataset extracts RDF data (i.e., for updating the RDF dataset). Updates are performed with low latency using standard SPARQL update queries; the addition and deletion of triples and / or the addition and deletion of graphs are performed on the same dataset and are immediately available. In other words, it can be seen that one or more partitions provide the same modification capabilities as a standard graph database without distributed storage. These partitions are also referred to as "dynamic partitions". It should be reminded that the term "partition" has been defined as "dividing a logical database or its constituent elements into different independent parts". In the case of streaming modification, standard SPARQL updates to the RDF dataset are applied to the partitions of the database, and the partitions can be modified like any other graph database.

[0074] Second RDF graph database

[0075] The virtual RDF graph database includes a second RDF graph database. The second RDF graph database is stored on a file storage. The second RDF graph database is read-only, which means that update operations cannot be directly performed on the dataset; thus, the second RDF graph database is a snapshot, i.e., a static point-in-time representation of the RDF data of the second RDF graph database. The second RDF graph database is updatable via a batch of tuples to be added to and / or removed from the second RDF graph database. Thus, the second RDF graph database forms a read-only snapshot but can also be updated in batches. In the initial case (i.e., the second RDF graph database has never been updated by a batch of tuples to be added to and / or removed from the second RDF graph database), the second RDF is a snapshot. For batch updates, batch partitions will accompany the second RDF graph database (e.g., in the form of Delta files as discussed below), such that the snapshot can include two or more snapshots. In the following discussion, the term snapshot can include the initial case (a snapshot of the second RDF graph database) or the case where the updated second RDF graph database includes several snapshots. Now, the batch partitions are discussed.

[0076] As is known per se, batch update refers to batch loading. Batch loading means loading data from an RDF file into partitions (also called archive partitions) called "batch partitions". The RDF file includes tuples to be added to and / or removed from the second RDF graph database. The RDF file can be in any format, for example, Turtle format (discussed here: https: / / www.w3.org / TR / turtle / ), or TriG (discussed here: https: / / www.w3.org / TR / trig / ), or an equivalent format. Thus, a batch stores the updates to be applied to the second RDF graph database, and the batch partition stores the updates that have been applied to the second RDF graph. In both cases, the snapshot itself is not directly affected by the modification because the snapshot is read-only. The applied modifications are only supported by the batch partition. The batch partition can be added to, removed from, or changed in the dataset.

[0077] The second RDF graph database is a read-only snapshot. As is known per se, a dataset snapshot consists of the locations of all files in a distributed storage, their metadata, and optional delta files. The locations of all files means that the snapshot includes information about the locations of all files in the dataset. This information indicates where each file is stored within the distributed storage infrastructure. This can involve references to specific storage nodes, addresses, or paths. The metadata of the snapshot refers to additional information about the characteristics, origin, and structure of the RDF graph database. For example, the metadata can include, but is not limited to, bibliographic information of the RDF graph database (e.g., title, author, creation date, modification date...), RDF data encoding format (e.g., Turtle, TriG...), number of triples, SPARQL endpoint, storage location.... Thus, the metadata can provide context and facilitate the organization and management of the dataset. The delta file represents the changes or differences between two versions of data. Including optional delta files in the dataset snapshot allows for effectively tracking changes over time. These delta files can contain information about the updates, additions, or deletions made to the dataset since the last snapshot. Including optional delta files when creating a snapshot can be particularly useful for minimizing the amount of data that needs to be transferred or stored because it only captures the changes since the previous snapshot. For explanatory purposes only, it is noted that when calculating a new snapshot, it is not necessary to repeat the entire previous snapshot because the algorithm can access references (e.g., paths) to the old data files that are still valid as well as references to the new data files. Additionally, when only a few triples have been modified, the Delta file does not need to regenerate the entire data file. These two mechanisms work separately to avoid replicating the entire snapshot.

[0078] In an example, and as is known per se, a batch partition can be split into n fragments to minimize the impact of an update, where n is a positive integer (n ≥ 1). Splitting the partition into fragments means dividing the partition into a set of subgraphs, where each subgraph is called a fragment. Here, a subgraph means a subset of the triples of the partition. A batch partition includes tuples added to and / or removed from a second RDF graph database, and the batch partition is decomposed into subgraphs, i.e., fragments. This is equivalent to saying that a SPARQL query that results in a batch of tuples to be added to and / or removed from a second RDF graph database is decomposed (e.g., in a random manner) into subgraphs (i.e., into subsets of triples), and then these subgraphs are executed on the second RDF graph database. This splitting of the batch partition into fragments is an optimization and not essential to the present invention.

[0079] In cases where the total size of the batch partition is too large (e.g., the size of the batch partition is greater than a predetermined size (e.g., greater than one gigabyte)), fragments can be used. The number of fragments can be determined according to a fragmentation strategy (e.g., the maximum size of a fragment, parallel processing of fragments, etc.).

[0080] In an example, small partitions can be merged into a larger fragment to control resource usage, such as file handles. The number of small partitions to be merged can be determined according to a merging strategy (e.g., the minimum size of a partition). This merging of small partitions into a larger fragment is an optimization and not essential to the present invention.

[0081] Still in these examples, a fragment consists of an RDF data file, where the RDF data file can be a binary file or not a binary file. A fragment also includes metadata of the RDF data file (e.g., statistical data) and an optional delta file.

[0082] Thus, the virtual RDF graph database is divided into two logical parts: a first RDF graph database that can be updated using standard SPARQL updates, called a dynamic partition; and a second RDF graph database that is a read-only snapshot and can be updated by loading a batch file, called a batch partition. As described above, the second RDF graph database can include one or more snapshots.

[0083] Update of the first RDF graph database

[0084] Standard SPARQL updates to the virtual RDF dataset are applied to the first RDF graph database (also known as the dynamic partition), which can be modified like any other graph database. The same RDF graph can appear in both the batch partition and the dynamic partition. A given triple can exist in only one partition, either the batch partition or the dynamic partition, otherwise the partition rules are violated. In fact, partitioning is defined as "dividing a logical database or its constituent elements into distinct, independent parts" (https: / / en.wikipedia.org / wiki / Partition_(database)). This constraint can be lifted at the cost of higher complexity either inside the query engine or during post-processing of the results outside the database. In fact, if a triple exists in more than one partition, this will lead to the following two drawbacks:

[0085] - For read queries: If some results are generated by some triples that exist in multiple partitions (and thus are generated multiple times), then these results will be duplicated;

[0086] - For write queries: If triples that exist in multiple partitions must be deleted, then the deletion must be propagated to all partitions. Otherwise, the triples will still be visible to subsequent queries. This constraint can be lifted without changing the present invention. This constraint can be implemented as known in the art.

[0087] Update of the second RDF graph database

[0088] The second read-only RDF graph database is stored on a file store and is updatable by batches of tuples to be added to and / or removed from the second RDF graph database. The second read-only RDF graph database forms an initial snapshot. Batches of tuples arrive gradually (i.e., for each new update to the initial snapshot) and together with the initial snapshot form a new snapshot. This new snapshot is defined as the state of the batch partition of the second RDF graph database at a given time. This state is defined by all batch partitions, which means all fragments (if one or more partitions have been split into fragments); that is, all RDF files and their metadata are stored in a distributed store. Then, the snapshot consists of the locations of all files in the file store, all file metadata, and optional delta files. As discussed, delta files represent the changes or differences between two versions of the data.

[0089] Table of Contents

[0090] This method provides a directory for step S100 to store metadata describing a second read-only RDF graph database on file storage. The directory conforms to ACID, which means that i) operations involving the addition, modification, or deletion of metadata entries are executed as atomic transactions. If any part of the operation fails, the entire operation is rolled back to maintain a consistent state. ii) The directory system enforces consistency rules (in the sense of consistent writes) on the metadata. This ensures the uniqueness of the dataset snapshot defined as the "most recent one" at a given time. iii) Multiple transactions that modify the metadata simultaneously do not interfere with each other. iv) Any changes made to the metadata of the directory are persistent. This typically involves writing the changes to non-volatile memory. "Most recent snapshot" means the new snapshot discussed previously.

[0091] The directory can register the metadata of the dataset in an ACID manner in any available language because the metadata are just key / value pairs and have a strong consistency model. For example, strong consistency is discussed here: https: / / en.wikipedia.org / wiki / Strong_consistency. "Consistency model" means "consistency" in the consistency model of the CAP theorem as in this link: https: / / apple.github.io / foundationdb / consistency.html.

[0092] Regarding the consistency provided by the directory, several dataset snapshots can be stored at a given time, but only one is called the "most recent one". Most recent means the snapshot that includes all batch partitions received up to now.

[0093] In the example, the directory can be implemented by, but not limited to, a distributed key / value store with ACID transactions and strong consistency, such as FoundationsDB (https: / / apple.github.io / foundationdb / index.html), or a NewSQL database, such as MariaDB in the open source world (https: / / mariadb.com / ) (MariaDB can be distributed with Xpand presented at the following link: https: / / mariadb.com / products / enterprise / xpand / ), or any RDF database.

[0094] In the example, all dataset snapshot metadata can be stored in the directory. Optionally, all dataset snapshot metadata can be saved as files in distributed storage, and the directory only keeps the location of the files.

[0095] Now discuss the operating principle of the directory. When a transaction starts, the directory extracts the metadata of the most recent dataset snapshot from the directory. This most recent snapshot is used for all its queries until the transaction ends.

[0096] When adding, removing, or changing partitions, new partitions or segments (if splitting is used) are ultimately added and / or previous partitions or segments that may have delta files (if splitting is used) are removed. When all these operations are successfully performed on the distributed storage, a new dataset snapshot can be created. Interestingly, the file storage ensures consistent writes and thus ensures the success of the update operation.

[0097] When all these operations are successfully performed on the distributed storage device, these operations are atomically registered as the "most recent one" within the directory. Therefore, all new transactions will use the new dataset snapshot, and previous transactions will continue to use their previous dataset snapshots. Due to the ACID guarantees of atomicity, durability, and strongly consistent writes of the directory and the distributed storage, the ACID properties are ensured. The old dataset snapshots and the corresponding unused partitions (or segments and potential delta files (if splitting is used)) can be removed from the distributed storage at the end of all transactions that use them, or removed asynchronously during the garbage collection process.

[0098] As can be seen, the directory only stores metadata and not data. The directory will not store data of the same order of magnitude as in the dataset (i.e., billions of triples); therefore, in terms of computing resources and storage, the cost of maintaining this directory will be much smaller. The directory can be shared by several datasets.

[0099] Refer to Figure 3 , an example of the scheme of the architecture of the provided step S100 is shown. The architecture includes two layers. The top layer is the database layer that manages the RDF dataset. The layer below is the storage layer that stores the RDF dataset and batch partitions.

[0100] On the database layer, a first RDF graph database that can be updated by stream is represented and can be queried by a SPARQL query engine. A directory that stores the metadata of a second read-only RDF graph database on the file storage can also be queried by the SPARQL query engine; for explanatory purposes only, the SPARQL query engine does not necessarily rely on SPARQL queries to use the directory. For example, if the directory is a key / value store or an SQL database. Assume that the SPARQL query engine knows which interface / query language to use to interact with the directory. As shown, the directory stores information about the snapshots that can be accessed; the directory lists the available dataset information. The directory is also responsible for maintaining the metadata information of the snapshots.

[0101] At the storage layer, snapshots of a dataset (e.g., a second read-only RDF graph database) are stored on a file storage that ensures strongly consistent writes. For at least the most recent snapshot referenced on a directory, a list of data files of that snapshot is stored. It is to be understood that the file storage can store several lists of data files, each list corresponding to a respective snapshot. Still at the storage layer, batch partitions are stored, representing updates applied to the snapshots.

[0102] The batch partitions and the dynamic partitions (which are, respectively, the second read-only RDF graph database and the first RDF graph database stored on the file storage) constitute an RDF dataset (a virtual RDF graph database that includes at least two RDF graph databases regarded as a logically equivalent RDF graph database). The SPARQL query engine executes queries on all partitions and treats them as a single logical database. Read queries are executed on all partitions with ACID guarantees. SPARQL update queries are executed only on the dynamic partitions, without restrictions on the type of updates.

[0103] Still referring Figure 3 to the example of, the dynamic partitions are stored on an in-memory data structure, while the batch partitions are stored on a file storage that ensures consistent writes, which can be distributed or not. In this example, the SPARQL query engine executes queries on heterogeneous distributed storage.

[0104] When executing a SPARQL query, the execution of the query adheres to two main constraints: 1. The RDF dataset is dynamic (data can be added or removed); and 2. The SPARQL query is executed in an ACID manner, especially regarding isolation, because the RDF dataset is dynamic.

[0105] It is to be understood that the batch partitions may be located on a standard file system (e.g., Unix file system) rather than on a distributed object storage.

[0106] Returning to Figure 1 , at S110, a stream of tuples to be added to and / or removed from the virtual RDF graph database and a batch of tuples to be added to and / or removed from the virtual RDF graph database are obtained. The order in which they are obtained is immaterial. They can even be received simultaneously.

[0107] Next, at S120, the tuple stream is applied to the first RDF graph database of the virtual RDF graph database. This is performed in a manner known in the art.

[0108] Then, at S130, the batch of tuples is applied to the second RDF graph database. In other words, as a result of the application, batch partitions will be created. To this end, the following steps are performed.

[0109] First, calculate a snapshot of the second RDF graph database. This calculation includes the batch of tuples to obtain a new snapshot. During the computational operation that results in the generation of the new snapshot, the ACID properties are ensured by a sequence of concurrent control operations that are appropriately executed by using consistent writes to the file store on which the second RDF graph database is stored.

[0110] "Concurrent control operations that ensure the ACID properties" means that, for each transaction, the partition of the dataset snapshot taken at the start of the transaction remains available throughout the transaction and is thus not modified by concurrent additions / removals / changes. The concurrent control operations rely on the properties of a directory with ACID transactions, and the file store has read-after-write consistency after the creation / overwrite / deletion of an object, which means that "any subsequent read request immediately receives the latest version of the object", and the external storage can upload a file, but will definitely fail if the file already exists. Examples of the implementation of the algorithm that ensures "concurrent control operations that ensure the ACID properties" will be described below.

[0111] Once computed, the snapshot is stored on the file store.

[0112] Finally, the snapshot is registered in the directory such that the update description of the virtual RDF graph database is stored in the directory; thus, the update description of the virtual RDF graph database includes the update description of the first RDF graph database and the update description of the second RDF graph database. This new snapshot can be used to apply a new batch of tuples.

[0113] At this step, with ACID guarantees, SPARQL queries can be executed on RDF data partitioned on a heterogeneous distributed storage mainly consisting of non-ACID compliant storage, where the SPARQL queries can include update queries. The data for the update queries can be ingested from a bulk data file from the distributed storage or directly ingested as a streaming modification of triples onto an ACID database.

[0114] In an example, the present invention can be accessed via an API or can be accessed via SPARQL commands as a specification extension. Both are possible, but for clarity only, as discussed with reference to Figure 1 Some commands will now be discussed as extensions to SPARQL to manipulate RDF datasets. It is to be understood that the present invention is not limited to these commands; in particular, the actual syntax is for discussion purposes only. Figure 4 An overview of the commands to be discussed hereafter is provided.

[0115] In an example, ALTER REFRESH commandA sequence of concurrent control operations to be performed can be executed. The ALTER REFRESH command ensures the ACID properties of the updates to the second RDF graph database of the virtual RDF graph database. The purpose of ALTER REFRESH is to create a new snapshot of the dataset based on the current state of the dataset and to refresh the list of available dataset snapshot definitions.

[0116] As an introduction to the ALTER REFRESH command, it should be noted that the ALTER REFRESH command is responsible for ensuring the ACID properties of the manipulation of the dataset. The manipulation of the dataset can include, but is not limited to, for example, adding partitions to the dataset using the ADD PARTITION command, removing partitions from the dataset using the REMOVE PARITION command, and changing (i.e., modifying) partitions using the ALTERPARTITION command. This is summarized in Figure 4 this.

[0117] As already discussed, a dataset snapshot can be defined as the state of the dataset at a given time. This state is defined by all partitions (which means both dynamic partitions and batch partitions). As previously mentioned, batch partitions (read-only RDF graph databases stored on file storage) and dynamic partitions (the first RDF graph database that can be updated via a stream) are modified by separate commands. Since dynamic partitions are modified in a streaming manner by definition, the modifications to dynamic partitions are immediately visible. Batch partitions are also modified in batches by definition. When a transaction starts on the dataset, the transaction extracts the most recent snapshot from the catalog. This snapshot is used for all its queries until the transaction ends.

[0118] When the ALTER REFRESH command is executed, all partitions that make up all batch partitions can be listed in the stage of the dataset. It should be reminded that for the purpose of improving the performance of SPARQL queries, partitions can be segmented; in this case, all segments that make up all batch partitions are listed in the stage of the dataset. As is known per se, processing the dataset in stages is a common approach; the stage of the dataset models the location of the files that make up the dataset. For example, the definition of "stage" is provided at: https: / / docs.snowflake.com / en / sql-reference / sql / create-stage.

[0119] The partition list is saved in a so-called dataset snapshot identified by a unique name. The dataset snapshot is saved on a directory by an ACID transaction and registered as the "most recent snapshot" or "most recently available snapshot". Thus, all new transactions will use the new dataset snapshot loaded at this stage, and previous transactions will continue to use their previous dataset snapshots. This guarantees the ACID properties due to the ACID guarantees of atomicity, durability, and consistent writes for the directory and distributed storage.

[0120] Due to the dataset snapshot extracted at the start of a transaction, each transaction has its own list of partitions (or fragments, depending on the case). To be able to have ACID transactions on the dataset, these partitions (or fragments, depending on the case) are available during the transaction and are thus not modified by concurrent partition additions / removals / changes, nor are they modified by concurrent dataset refresh changes. To achieve this, the ALTER REFRESH command relies on three capabilities: 1 / the directory has ACID transactions; 2 / the external storage has read-after-write consistency after the creation / overwriting / deletion of an object, meaning that "any subsequent read request immediately receives the latest version of the object"; and 3 / the file storage can upload a file and will definitely fail if the file already exists.

[0121] In the example, the first control operation of the ALTER REFRESH command is to uniquely identify the most recent snapshot of the second RDF graph database. As already discussed above, the most recent snapshot is the snapshot that includes all batch partitions received at the time when the ALTER REFRESH command starts its execution. Thus, the most recent snapshot represents the snapshot formed by the second RDF graph database and its batch partitions. It is to be understood that the most recent snapshot may not include batch partitions, for example, at the time when the ALTER REFRESH command starts its execution, no batch tuples to be added to and / or removed from the virtual RDF graph database are received. Uniquely identifying means retrieving or obtaining the most recent dataset snapshot identifier from the directory.

[0122] The next control operation of the ALTER REFRESH command can include verifying that no other ALTER REFRESH commands are being executed concurrently. If so, the ALTER REFRESH command is stopped, and the ALTER REFRESH commands that have already been executed are retained. This helps to ensure ACID transactions on the most recent snapshot.

[0123] The next control operation of the ALTER REFRESH command may include obtaining a list of files in the phase of all snapshots. This operation is completed once, and the entire list can be copied in the memory of the process. This list is used for the rest of the ALTER REFRESH command. It should be reminded that the operation of listing files in this phase is strongly consistent according to the consistent write attribute of file storage. This means that the files listed in this phase are visible to any or all processes using the data; any subsequent read operation on the most recent snapshot reflects this data.

[0124] The next control operation of the ALTER REFRESH command may include filtering files in the list of files in the phase of the most recent snapshots. Filtering includes checking whether the files in the list are valid, where valid means that no other operations are running on these files. Filtering ensures the isolation property of the ACID property, and thus ensures that no concurrent operations will interfere with the files in the list. In other words, the view on the most recent snapshot of the database is consistent and reliable. At a given time, the result of filtering is to obtain a "snapshot list", which is a file that lists the file names of the most recent snapshots provided for subsequent access to the database. Therefore, the "snapshot list" includes partitions that have not been updated (i.e., not modified) and updates that have been executed since the last ALTER REFRESH command; the updates that have been executed since the last ALTER REFRESH command in the "snapshot list" are not visible for any subsequent actions on the database. Subsequent access includes but is not limited to SPARQL queries, tuples to be added and / or removed. Here, the given time is the moment when the ALTER REFRESH command has started its execution.

[0125] The next control operation of the ALTER REFRESH command may include writing the name of the file of the "snapshot list". Here, the term "name" means that the snapshot of the data set identified by the unique name is loaded in this phase.

[0126] Writing the name of the file of the "snapshot list" means that at a given time, a new snapshot is obtained from the previous snapshot, where all files of the previous snapshot will be available for subsequent access to the database. In other words, the names of all files in the "snapshot list" are written into the file of the new data set snapshot.

[0127] Any known technology can be used to perform the writing. For example, this can be done using real semantics (e.g., using the RDF format, TriG format in Turtle files). Optionally, any known technology for performing the writing can be used, such as a simple list of file names.

[0128] The next control operation of the ALTER REFRESH command can include uploading the files of the new snapshot to the stage. The uploaded new snapshot includes the files named in the "snapshot list" file. This is performed in a manner known in the art.

[0129] The new snapshot has not been registered yet. This means that the directory is not aware of the new snapshot and thus may not execute SPARQL queries (e.g., SELECT) on the new snapshot that includes the latest data. The registration is performed within an ACID transaction. The ACID transaction on the directory thus ensures the integrity and reliability of the registration. The new snapshot is registered as the most recent snapshot in the directory.

[0130] The ALTER REFRESH command has been executed, and the obtained batch of tuples (S110) can be applied to the registered new snapshot (i.e., the most recent snapshot) for performing the calculation of the snapshot of the second RDF graph database (S130), where the new snapshot includes this batch of tuples.

[0131] Now other examples implemented by the ALTER REFRESH command are discussed.

[0132] In an example, the following sequence of concurrency control operations can be utilized to obtain the "snapshot list", which names the files of the most recent snapshots available for subsequent access to the database at a given time.

[0133] A list of the names of the valid partitions can be obtained. This list can include one or more (but for all partitions) valid partitions of the second RDF graph database. Here, a "valid partition" means that the partition has been correctly created and created in an ACID manner.

[0134] In a first example, obtaining the list of the names of the valid partitions can include searching the file storage for the name of the control file that stores the files of the valid partitions for each partition. It is understood that in practice, a partition includes several files, so the control file stores the names of multiple files that constitute the valid partition.

[0135] As is known per se, the control file plays an important role in managing the structure and metadata of the database. The control file serves as a repository for basic metadata related to the database (e.g., database name, file location, log file details, timestamp, and the structural integrity of the current database).

[0136] Returning to the present invention, the control file can be stored on the file storage when the partition is batched on the file storage (i.e., when the partition is created).

[0137] In the example of the first example, the control file can be named with the partition name and a specific extension called VALID_EXTENSION. Thus, the control file is called the VALID_EXTENSION file. The VALID_EXTENSION control file is designed to identify the records of valid partitions. The VALID_EXTENSION file stores the names of the files of valid partitions. Naming the file with the partition name and the specific extension makes it easy to identify the VALID_EXTENSION file of the partition.

[0138] As a replacement or supplement to the first example, in the second example, obtaining a list of the names of valid partitions can include searching the file store for a control file for each partition, which stores the names of the files that are the objects of the current deletion operation in the partition. Similarly, a single file name can be stored, which is not a typical case. The control file can be stored on the file store when the deletion operation of the partition has been performed on the partition.

[0139] In the example of the second example, the control file can be named with the partition name and a specific extension called TOMBSTONE_EXTENSION. The TOMBSTONE_EXTENSION control file is designed to identify the records that are the objects of the current deletion operation in the partition. The TOMBSTONE_EXTENSION file can store the names of the files that are the objects of the current deletion operation in the partition. When the deletion operation of the partition starts, the TOMBSTONE_EXTENSION file can be uploaded to the file store.

[0140] The execution of the control operation to list the names of valid partitions has been implemented. Further control operations can include, for each partition named in the list of valid partitions, verifying whether the current update operation on the partition is still pending. If the current update operation on the partition is still pending, then add the partition to the list of valid partitions whose current update operations are still pending. In terms of "current update operation", it means that the update operation is considered pending because it has not been fully processed and / or applied on the system.

[0141] In the example, verifying that the current update operation on the partition is still pending for each partition can include: searching the file store for a control file for each partition, which identifies the records that are the objects of the current update operation in the partition; the control file stores the names of the files of valid partitions that are the objects of the current update operation. Here, a single file name can also be stored, which is not the usual case.

[0142] In an example, a control file storing the name of a file of a valid partition that is an object of a current update operation can be named with the partition name and a specific extension called PENDING_EXTENSION. The PENDING_EXTENSION control file is intended to identify records in the partition that are objects of the current update operation. The PENDING_EXTENSION file stores the names of files in the partition that are objects of the current update operation. The PENDING_EXTENSION file can be uploaded to the file store when the update of the partition has started.

[0143] After verification has been implemented (for each partition, whether the current update operation on the partition is still pending), further control operations can include verifying whether past update operations on the partition have not been successfully implemented. "Past update operations on the partition have not been successfully implemented" means all of the following "pending" files, where all of the "pending" files were used by a previous successful ALTER_REFRESH to flush but failed to clear the "pending" files from the stage; this is typically the case when the ALTER_REFRESH crashes before the end of the ALTER_REFRESH algorithm. If past update operations on a partition have not been successfully implemented, the control operations can also verify whether a snapshot for which an unsuccessful update operation has been performed is registered in the directory. The snapshot can be a past snapshot, or it can be the most recent snapshot.

[0144] If the snapshot is not recorded (or registered) in the directory, this means that the previous change flush was unsuccessful because the new snapshot should have been registered in the directory after the directory had identified the new snapshot as the most recent snapshot to which the batch of tuples was to be applied in an ACID transaction. Accordingly, the unsuccessful update operation is ignored, and the "pending" files are not pending because the previous ALTER_REFRESH has not been successfully implemented.

[0145] If the snapshot is recorded, this means that the previous ALTER_REFRESH has been successfully implemented, but information that past update operations on the partition have not been successfully implemented has been retained and it should be deleted. Accordingly, the partition can be removed from the list of valid partitions for which the update operation is pending.

[0146] In an example, verifying whether past update operations on a partition have not been successfully implemented can include searching the file store for a control file for each partition that identifies records in the partition that are objects of past update operations; the control file can store the names of files in the valid partition that are objects of past update operations. Here too, a single file name can be stored, which is not the normal case.

[0147] In an example, a control file storing the names of files that are objects of past update operations in a valid partition can be named with the partition name and a specific extension called CONSUMED_EXTENSION. The CONSUMED_EXTENSION files can have been uploaded to a file store at the start of the past update of the partition.

[0148] In an example, for each partition named in a list of valid partitions, verifying whether past update operations on the partition have not been successfully implemented can include further control operations. These further control operations use two control files: i) a control file storing the names of files that are objects of current update operations in a valid partition; and ii) a control file storing the names of files in a valid partition. Check the coherence between these two files.

[0149] For clarity only, the coherence check is discussed with reference to an example implementing PENDING_EXTENSION files and VALID_EXTENSION files. The control operations can include verifying that the content of the PENDING_EXTENSION file is consistent with the content of the VALID_EXTENSION file. The VALID_EXTENSION file stores the names of files of valid partitions, and the PENDING_EXTENSION file stores the names of files that are objects of current update operations in a partition. For a valid partition, at the start of the execution of the ALTER_REFRESH command, the names listed in the VALID_EXTENSION file should not be found in the PENDING_EXTENSION file, because the names listed in the VALID_EXTENSION file should be available for subsequent access to the database. If this is not the case, then for the verification, the PENDING_EXTENSION file is ignored. In other words, only the VALID_EXTENSION file is considered.

[0150] After it has been verified whether past update operations on a partition have not been successfully implemented, further control operations of the ALTER_REFRESH command can create an empty "snapshot list" file and complete the "snapshot list" file by adding the names of files of valid partitions for which the current update operation is still pending, or otherwise, if the list of valid partitions for which the previous update operation is still pending is empty, by adding the names of files of valid partitions for which the list of valid partitions is given. Thus, at a given time, the obtained snapshot list names the files of valid partitions so that they are available for subsequent access to the database.

[0151] For clarity only, the completion of the "snapshot list" file is now discussed with reference to an example implementing the PENDING_EXTENSION file and the VALID_EXTENSION file. After verifying whether past update operations on the verified partitions have not been successfully implemented, further control operations of the ALTER_REFRESH command can create an empty "snapshot list" file and complete the "snapshot list" file as follows. For each partition, if there is a file named in the PENDING_EXTENSION file, the list of file names constituting the partition is given by the content of the PENDING_EXTENSION file and added to the "snapshot list" file. Otherwise, the list of file names is given by the content of the VALID_EXTENSION file and added to the "snapshot list" file. Now the "snapshot list" file naming the files of the most recent snapshot available for subsequent access to the database at a given time is completed.

[0152] Then, further control operations can include verifying that each file listed in the "snapshot list" exists in the stage. This can be performed by comparing the "snapshot list" with the list of files in the stage of the most recent snapshot.

[0153] In the example, uniquely identifying can include obtaining the most recent dataset snapshot identifier from the directory and maintaining the most recent dataset snapshot identifier in memory as the previous dataset snapshot. Registering a new snapshot in the directory can also include verification by the directory in an ACID transaction that the most recent snapshot described in the directory is the previous dataset snapshot (already saved in memory). The comparison between the two identifiers of the file is straightforward. If the most recent snapshot described in the directory is the previous dataset snapshot; then the new snapshot is registered in the directory as the most recent snapshot to which the batch of tuples is applied. In the opposite case, a concurrent ALTER REFRESH command is being executed and the new snapshot file from the stage is removed.

[0154] In the example, the ALTER_REFRESH command can further include control operations for providing a latch. As is known per se, a latch is a synchronization mechanism for controlling access to shared resources to ensure consistency and avoid conflicts in concurrent operations.

[0155] In the example, the control operations can include latching a second read-only RDF graph database stored on the file storage. In this case, the latching is performed before obtaining the list of files of the most recent snapshot in the stage. Once the new snapshot (identified as the most recent snapshot) has been registered in the directory, the new snapshot is unlatched.

[0156] In an example, a latch may include using a control file, the presence of which on a file storage confirms that the second read-only RDF graph database has not been changed during the execution of the ALTER_REFRESH command. After the most recent snapshot has been uniquely identified, the control file can be uploaded to the file. Then, the control operation can verify that the control file still exists on the file storage at each subsequent step executed by the control operation. Finally, after the new snapshot has been registered in the directory, the presence of the control file is checked one last time, and the control file is removed from the file storage.

[0157] A further discussion of latches is now provided. Since the present invention relies on an RDF heterogeneous distributed storage mainly consisting of non-ACID compliant storage, as Figure 3 illustrated above, an algorithm may be needed to manage concurrent execution and ensure that the execution of the algorithm is always correct. A latch is a synchronization mechanism used to control access to shared resources to ensure consistency and avoid conflicts in concurrent operations. In an example, a file storage may provide a similar mechanism, for example, object locking in S3 (https: / / docs.aws.amazon.com / AmazonS3 / latest / userguide / object-lock.html), but not all storage provides such a mechanism. In an example, the concurrent control operation may execute a latch algorithm; object locking does not depend on the underlying infrastructure capabilities.

[0158] In an example, a dedicated START_EXTENSION control file may be used. The presence of this file will be used as a latch, as discussed in different examples. If the process that creates the START_EXTENSION file crashes before it removes the file from the stage, the file should be removed to avoid deadlocks. This is remedied by the following example of a latch algorithm.

[0159] - Add the ip address of the host (or any other means of identifying the host) and the process id to the START_EXTENSION control file;

[0160] - If a process encounters a START_EXTENSION file (referred to as the "start" file for short) when it wants to create its own "start" file, read the ip address in the file:

[0161] -- If the ip address is the ip address of another existing host within the cluster of the data set, a concurrent problem is indeed detected: the process cannot create its own "start" file;

[0162] -- Otherwise, if the ip address is the ip address of its own host, check the process id:

[0163] ---If the process id is not the same as itself, a restart of the process occurs: it is safe to replace the "start" file with a new one; / / For application deployment, ensuring one process per host may be mandatory. Otherwise, the policy may be changed accordingly.

[0164] ---Otherwise, a concurrency problem is indeed detected;

[0165] -- Otherwise, the ip address is unknown: the host that wrote the "start" file is down; / / It is safe to replace the "start" file with a new one.

[0166] There are still cases where a host goes down and then recovers (also known as "split brain" in distributed systems, see https: / / medium.com / nerd-for-tech / split-brain-in-distributed-systems-252b0d4d122e). To prevent this, all algorithms need to check that the "start" file is still its own file before making any modifications to the stage. In the detailed algorithm, this operation will be summarized by the sentence "check the'start'file is still its own file";

[0167] If this check fails, split-brain occurs. Each algorithm then exits and needs to clean up the files it has uploaded (if any). If this cleanup is not done by the algorithm, it will be done by garbage collection (see dedicated section). The algorithms are designed to ensure atomicity of operations (i.e., if an algorithm fails before the end, the files uploaded at that stage are not considered).

[0168] In the example, the control file is a file named with the name of the most recent dataset snapshot and a specific extension called START_EXTENSION. Thus, for each snapshot stored on the file store, there is a unique control file.

[0169] Several examples of the ALTER_REFRESH command have been discussed. One or more of these examples may be combined. An example of the ALTER_REFRESH command in pseudo-code algorithm form is now discussed. The terminology used in this pseudo-algorithm is similar to the terminology used in the example of the ALTER_REFRESH command. Interestingly, this pseudo-code algorithm shows an example where a partition is segmented. It is to be understood that the segmentation of the partition does not change the examples already discussed. This will be apparent in the following examples. Comments begin with / / .

[0170] Input the ALTER_REFRESH command: The data set stored on the file storage

[0171] Obtain the most recent dataset snapshot of the dataset :

[0172] - Obtain the most recent data set snapshot identifier from the directory;

[0173] - Keep this information as the "previous data set snapshot";

[0174] - If the data set snapshot itself is a file in progress (the directory only redirects to it), extract the data set snapshot from this progress. Otherwise, obtain all information from the directory;

[0175] - Connect to the data set snapshot using the cluster of the data set;

[0176] / / This means, for example, registering all segments as parts of the data set and optionally pre-extracting the metadata file.

[0177] / / Note that the connection to the data set snapshot can be kept in the cache and reused for subsequent calls. Then, when "DISCONNECT" is executed and / or any other cache replacement policy (such as LRU) is used, this cache can be cleared.

[0178] / Note that saving the data set snapshot in the cache limits the use of the garbage collection algorithm executed by the GARBAGE_COLLECT command; an example garbage collection algorithm will be provided later; any trade-off between the cache and garbage collection can be selected without changing the present invention.

[0179] Will be named with the data set name and a specific extension called START_EXTENSION (e.g., ".start") Latch the control file and upload it to storage ;

[0180] - If this call is successful, the algorithm can determine that such a file did not exist previously, and thus there is no other "ALTER REFRESH" or "GARBAGE_COLLECT" running concurrently for this data set;

[0181] - If this call fails, a concurrent "alter refresh" or "garbagecollect" is detected for this data set: exit and report an error;

[0182] Generate a new dataset snapshot name , for example, dataset_snapshot_ <uuid>”, may have a dedicated extension, e.g., if the turtle format is selected, the extension is ".ttl”

[0183] List all files in this stage

[0184] / / This operation is done once and the entire list is copied into the memory of the process. We use this copy for the rest of the algorithm.

[0185] / / Depending on the properties of the object store, the list file operation is strongly consistent.

[0186] Filter files from this list of all files in this stage :

[0187] - First, list all valid partitions by listing all files ending with VALID_EXTENSION. The partition name is given by the name of the VALID_EXTENSION file without the extension;

[0188] - For each partition name, if there is a file with the partition name followed by the extension TOMBSTONE_EXTENSION, filter it from the list of valid partitions. / / This gives us all valid partition names.

[0189] - For each valid partition name, check if there is a file starting with the partition name and ending with the extension PENDING_EXTENSION;

[0190] -- If so, keep in memory the list of all PENDIND_EXTENSION files found here; / / These files correspond to "ALTER_PARTITION" commands that have not yet been considered, see the dedicated algorithm discussed below;

[0191] - For each file with the CONSUMED_EXTENSION extension / / Note that these files only exist for ALTER_REFRESH commands that crashed before the end of the algorithm; there should not be any files in the nominal case:

[0192] -- Get the name of the dataset snapshot from the file name;

[0193] -- Check in the directory that the dataset snapshot is registered:

[0194] --- If not: Ignore the CONSUMED_EXTENSION file / / It can be removed here or garbage collected);

[0195] --- If so:

[0196] ----Read the content of the CONSUMED_EXTENSION file / / The CONSUMED_EXTENSION file gives all the "pending" files used by the snapshot of this dataset;

[0197] ----Check that the content of PENDING_EXTENSION is consistent with the content of the VALID_EXTENSION file. If not, consider that the VALID_EXTENSION file has not been updated with the content of the "dataset snapshot" in the previous extension, and treat the content of the "dataset snapshot" as the content of the VALID_EXTENSION file;

[0198] ----Remove these files from the list of all PENDIND_EXTENSION files found in the previous step / / They were used by a previous successful change but failed to clear the "pending" files from the stage;

[0199] --Create an empty list called the "snapshot list" file, which will list the names of the files of the dataset snapshot;

[0200] -For each valid partition name, add the following files to the "snapshot list" file:

[0201] --If there is a file with the extension PENDING_EXTENSION in the file list, the list of files that make up the partition is given by the content of this file and added to the "snapshot list" file; otherwise

[0202] --The list of files is given by the content of the file with the partition name and the VALID_EXTENSION extension, and added to the "snapshot list" file;

[0203] -Check that all files in the "snapshot list" exist in the listed files of this stage; / / The list of files for this stage is the list copied in memory at the start of the algorithm; no actual check of this stage is done here.

[0204] -Check that for all segments, both an RDF file and a corresponding metadata file exist; / / Note that there may be delta files (checked for their existence in the previous step of checking all files in the list); / / Note that using the VALID_EXTENSION and PENDING_EXTENSION files to list files gives the algorithm isolation against concurrent modifications of partitions;

[0205] Write the names of all files in the "snapshot list" to the file of the new dataset snapshot / / Any format can be used to write these names. For example, writing these names can be done using the real semantics of the RDF format in a Turtle file, or if no more semantics are needed, writing these names can simply be done as a list.

[0206] Check that the "start" file is still its own file and upload the dataset snapshot file in the stage / / Note that this dataset snapshot is not yet visible for subsequent transactions because it has not been registered in the directory to preserve atomicity.

[0207] Check that the "start" file is still its own file and register the new dataset snapshot in the directory as the most recent dataset snapshot Dataset snapshot , give the retention information of the "previous dataset snapshot":

[0208] - The directory checks in the ACID transaction that the "nearest dataset" it knows is the "previous dataset" given as the input to the check;

[0209] -- If so, write two pieces of information:

[0210] --- Register the nearest dataset snapshot in the directory;

[0211] --- Mark all "pending" files as consumed:

[0212] ---- Write the names of all files with the PENDING_EXTENSION extension to a file named with the dataset snapshot file name and the dedicated CONSUMED_EXTENSION (e.g., "<dataset_snapshot_name>.consumed");

[0213] ---- Check that the "start" file is still its own file and upload the CONSUMED_EXTENSION file / / Note that the CONSUMED_EXTENSION file has been used at the start of the algorithm;

[0214] --- Register this new nearest dataset snapshot in the directory, linking it to the "previous dataset" to maintain lineage / / Lineage is optional;

[0215] / / This call also checks that the "previous dataset" is correct to avoid a race condition (or when the directory has previously checked it, the directory may have acquired the lock);

[0216] / / Registering it in the directory means registering the dataset snapshot file name. Optionally, all information saved in the file can be registered in the directory instead of the file itself.

[0217] / / Note that even if the process crashes after this, since the "CONSUMED_EXTENSION” file has been created, subsequent "ALTER_REFRESH” will not reapply the pending operations, as seen at the start of the algorithm;

[0218] --- Replace the content of the VALID_EXTENSION file with the list of files included in the "dataset snapshot” for further alter refresh operations;

[0219] / / Note that atomically replacing the content of a file can be done by creating a new file with the desired content instead of overwriting the existing file.

[0220] --- For all PENDING_EXTENSION files, delete them from this stage;

[0221] --- Delete the file with the CONSUMED_EXTENSION extension;

[0222] -- If not:

[0223] --- The "most recent dataset” does not know the "previous dataset” given as input to the check. / / A concurrent "ALTER_REFRESH” command has been executed.

[0224] --- Roll back the changes made in the stage:

[0225] ---- Remove the dataset snapshot file from the stage;

[0226] ---- Remove the file with the CONSUMED_EXTENSION extension;

[0227] ---- Return an error: A concurrent "ALTER_REFRESH” command has been executed;

[0228] - Check that the "start” file is still its own file and remove it; / / For all "exit” statements of this algorithm, and if any exception was raised in the previous step, this file must be removed.

[0229] In the example, ALTER PARTITION command A sequence of concurrent control operations to be executed can be performed. The ALTER PARTITION command is executed after the ALTER REFRESH command that ensures the ACID properties of the update to the second RDF graph database of the virtual RDF graph database. The purpose of ALTER PARTITION is to update the partitions of the second read-only RDF graph database.

[0230] Now discuss examples of the ALTER PARTITION command. These examples are illustrated using a control file with a specific naming and extension scheme, understanding that the control file is not limited to these examples.

[0231] The control operations of the ALTER PARTITION command can include downloading the VALID_EXTENSION file of the valid partitions to be updated from the stage. Thus, obtain the names of the files of the valid partitions. The valid VALID_EXTENSION files have been created during the initial creation of the second read-only RDF graph database, or further updated for a successful update operation of the second read-only RDF graph database, or created during the ALTER REFRESH command, as already discussed above.

[0232] Then, the control operations of the ALTER PARTITION command can verify that the ALTER PARTITION command is not being executed concurrently on the valid partitions. If not, stop the ALTER PARTITION command and retain the ALTER PARTITION commands that have been executed.

[0233] Next, the PENDING_EXTENSION file can be created. The PENDING_EXTENSION file stores the names of the files that are the objects of the current update operation in the valid partitions. The names of the files are obtained from this batch of tuples.

[0234] Once created, the PENDING_EXTENSION file can be uploaded to the stage so that it can be accessed by further commands (if any).

[0235] Now discuss examples of the ALTER PARTITION command where one or more batch partitions of the second read-only RDF graph database are split into fragments. All batch partitions can be split. The control operations of the ALTER PARTITION command can also include determining whether the second read-only RDF data set is a Change Data Capture (CDC) file.

[0236] As is known per se, CDC (the acronym for Change Data Capture) uses delta files to identify and track the data that has changed so that actions can be taken using the changed data. A discussion of CDC can be accessed here: https: / / en.wikipedia.org / wiki / Change_data_capture. Thus, CDC for SPARQL updates can be conceived as being able to represent only the changes as constraints for extended-defined SPARQL 1.1 updates. Figure 5 An example of how a modification of a triple can be represented is shown. Any other representation can be used without changing the present invention. In the example, information about which modification of which triple(s) in which figure(s) can be provided.

[0237] If the batch of triples is not a CDC file, fragments of the batch of triples can be generated, thereby obtaining a list of fragment files. The generation of fragments is discussed with reference to the ADD_PARTITION command.

[0238] If the batch of triples is a CDC file, the following operations can be performed for each of the add / delete triples (the batch of triples) in the CDC file. Locate a fragment of a batch partition in one or more batch partitions. The located fragment of the snapshot is the fragment to which the add / delete operation will be applied. As a result of the location, if one or more delta files are associated with one or more fragments in the fragment, a list of snapshot fragments and a list of additional delta files are obtained. Then, the control operation of the ALTER PARTITION command creates a PENDING_EXTENSION file, and the list of fragments with the list of additional delta files (if any) is stored. The created (or obtained) PENDING_EXTENSION file is uploaded to the stage and thus accessible for further commands (if any) as already discussed.

[0239] A partition can be changed by using the ALTER PARTITION command. It should be reminded that for the first RDF graph database (also called dynamic partition), modifications are made in a streaming manner, so only SPARQL UPDATE queries are used.

[0240] For the second read-only RDF graph database, a batch file is used as the input for updates, for example, for creating a partition. There are several options for the input to the ALTER PARTITION command. The first way can be to give again the entire file used for creating the partition but modified as needed (i.e., deprecated triples are deleted and new triples are added). In this way, the fragments are regenerated (regenerated as a whole or incrementally). In this solution, the input can be a file in a standard format such as TriG. Another solution can be to give only the desired modifications (triples to be deleted, triples to be added) as the input. There is no standard format to formalize this scenario, so we will define one. CDC can be used for this purpose, as discussed with reference to Figure 5 as discussed in. One of the two solutions to be used can be selected by detecting the input file format (the batch of triples). For example, for full regeneration:

[0241] ALTER PARTITION [Name] INTO DATASET [dataset Name] {CONTENT = [file.trig.gz], DESCRIPTION = "description"}

[0242] For increasing input:

[0243] ALTER PARTITION [Name] INTO DATASET [dataset Name] {CONTENT = [file.cdc.gz], DESCRIPTION = "description"}

[0244] Examples of the ALTER PARTITION command have been discussed. One or more of these examples can be combined. Now, an example of the ALTER PARTITION command in the form of a pseudocode algorithm is discussed. The terms used in this pseudocode algorithm are similar to those used in the examples of the ALTER PARTITION command. Interestingly, this pseudocode algorithm shows an example where a partition is split. It is to be understood that the segments of the partition do not change the examples being discussed. This will be apparent in the following examples. Comments start with / / .

[0245] Input of the ALTER PARTITION command :

[0246] - Input RDF file, including the bulk tuples to be added and / or removed; / / The input RDF file can be a standard RDF file such as TriG or a CDC file;

[0247] - The most recent snapshot of the second read-only RDF graph database to be modified;

[0248] - The partition name.

[0249] Check if the input file exists ;

[0250] Upload the START_EXTENSION control file for latch control to storage , for example, a START_EXTENSION control file named with the partition name and a specific extension called START_EXTENSION

[0251] - If the call is successful, the algorithm definitely knows that such a file did not exist previously, and thus there is no other "add / remove / change partition" running for that partition name;

[0252] - If the call fails, a concurrent "add / remove / change partition" with that partition name is detected: exit and report an error;

[0253] Download the VALID_EXTENSION control file for partition validity control from the stage , for example, a VALID_EXTENSION control file named with the partition name and a specific extension called VALID_EXTENSION

[0254] - If such a file does not exist, there is no partition with that name. Exit and report an error.

[0255] - If so, execute the next step of the algorithm.

[0256] Check if there is a PENDING_EXTENSION control file in this stage , for example, a PENDING_EXTENSION control file named with the partition name and a specific extension called PENDING_EXTENSION (e.g., ".pending")

[0257] - If so, there is already a pending partition change operation for that partition. Exit and report an error.

[0258] - If not, execute the next step of the algorithm.

[0259] Check if there is a TOMBSTONE_EXTENSION control file in this stage , for example, a TOMBSTONE_EXTENSION control file named with the partition name and a specific extension called TOMBSTONE_EXTENSION

[0260] - If so, that partition has already been removed. Exit and report an error.

[0261] - If not, execute the next step of the algorithm.

[0262] Verify if the input file is a CDC file;

[0263] - If the input file is not a CDC file:

[0264] -- Then generate fragments for the input RDF file. Examples of the generated ones are discussed in the "ADD PARTITION" command section.

[0265] -- Obtain a list of the fragment files of the input RDF file;

[0266] - If the input file is a CDC file:

[0267] -- For each add / delete triple of the CDC file:

[0268] --- Locate the fragment file of the nearest snapshot to which the add / delete operation must be applied;

[0269] / / Generate a new version of the fragment file of the nearest snapshot (e.g., if the file can be modified in-place), or add a delta file with the same name but with a specific extension (e.g., ".delta").

[0270] / / The delta file can be textual and in this case it is identical to the input CDC, or it can be a binary file - the binary file avoids parsing the text file -.

[0271] / / The delta file can be, for example, an HDT file for all the triples to be added and an HDT file for all the triples to be deleted, but other strategies can be used without changing the invention. In addition to the existing files, a delta file needs to be read to be able to answer the BGP of a SPARQL query;

[0272] --- Obtain a list of fragments with a list of additional delta files;

[0273] -- Create a new PENDING_EXTENSION control file, for example, a PENDING_EXTENSION control file with a partition name and a specific extension called PENDING_EXTENSION;

[0274] -- Add to this new PENDING_EXTENSION control file the list of fragments with potential additional delta files from the previous step;

[0275] / / All the files that make up the partition thus given, not only the modified deltas;

[0276] -- Check that the "start" file is still its own file and upload the file at this stage;

[0277] -- Check that the "start" file is still its own file and remove it from storage. / / For all the "exit" statements of the ALTERPARTITION command algorithm and if any exception was raised in the previous step, the file must be removed.

[0278] In the example, ADD PARTITION command A sequence of concurrent control operations to be performed can be executed. The ADDPARTITION command can be executed before the ALTER REFRESH command, which ensures the ACID properties of the update to the second RDF graph database of the virtual RDF graph database. The purpose of ADD PARTITION is to add a partition to the second read-only RDF graph database. For example, a typical command sequence can be the creation of the database, the execution of the ADD PARTITION command, and then the execution of the ALTER REFRESH command, which makes the new partition visible for further processing (e.g., SPARQL queries).

[0279] It is to be understood that new partitions can be added to a given dataset. The manner of addition is different if the new partition is a dynamic partition or a batch partition. A short discussion of adding a new partition to the first RDF graph database is provided.

[0280] When creating a virtual RDF graph database, a dynamic RDF graph database can be created by default. The first RDF graph database created is an RDF graph database that will be modified by SPARQL UPDATE queries. The default partition has no name. For example, by using the following command, a name and description can be given to the partition, specifying that it is a dynamic partition:

[0281] ADD PARTITION[Partition Name]INTO DATASET[dataset Name]{CONTENT=’DYNAMIC’,DESCRIPTION="description"}

[0282] Having a default dynamic partition is a choice made for the convenience of using the present invention: this avoids having queries that do not know the dataset based on an external location. It is possible to enforce naming for all partitions.

[0283] A short discussion of adding a new partition to the second read-only RDF graph database is now provided.

[0284] For batch partitioning, the input created is a batch of tuples. The data source for the batch can be an RDF file (e.g., in TriG or HDT (discussed below) format or any other format, optionally compressed with e.g., gzip or not), or a storage format like e.g., dynamic partitioning (detailed later). The input batch can then be split into one or more fragments. If the input batch is too large, there can be several fragments: indeed, large files can be expensive to fetch over the network, fetch from object storage (cost per GB retrieved), and navigate to answer BGP, which is discussed below. Experiments have shown that a maximum size of 10 million triples per fragment RDF file can be a good compromise. Once the input batch has been split into one or more RDF files, the metadata for the fragments can be generated. Note that the choice of the final format for the RDF files used for the fragments is independent of the format of the input. The format chosen can be the one best optimized for use, e.g., choosing a binary format instead of a human-readable format to reduce size. When all this is done, the fragments (i.e., the RDF files and their metadata) are uploaded to a stage associated with a snapshot of a read-only RDF graph database. The data of the RDF files and their metadata (fragments) will be visible only after a successful ALTER REFRESH command has been executed. Partitions can be named; in this case, the name must be unique and can be an IRI or a literal value. An optional description can be added. The command can be, for example:

[0285] ADD PARTITION [Partition Name] INTO DATASET [dataset Name] {CONTENT = [file.ttl.gz], DESCRIPTION = "description"}

[0286] Now, an example of the ADD PARTITION command is discussed. This example is illustrated using a control file with a specific naming and extension scheme, understanding that these examples are not limited to these example control files. The control operations of the ADD PARTITION command can include uploading the VALID_EXTENSION control file for the partition to be added to the storage. The presence of the uploaded VALID_EXTENSION on the storage indicates that there is no partition with the same name already in the file storage.

[0287] Then, the control operations of the ADD PARTITION command can generate fragments of the batch of tuples for adding the partition on a second read-only RDF graph database. This generation can be performed as discussed in the example of the ALTER PARTITION command. Then a list of the fragment files is obtained.

[0288] Next, the control operation of the ADD PARTITION command can store the fragment list in the VALID_EXTENSION control file and upload the VALID_EXTENSION control file to the stage. It should be understood that the VALID_EXTENSION control file can be used for subsequent commands, such as the ALTER REFRSH command, the ALTER PARTITION command...

[0289] Now discuss an example of the ADD PARTITION command in the form of a pseudocode algorithm. The terms used in this pseudo-algorithm are similar to those used in the example of the ALTER PARTITION command. The comments in the pseudocode start with / / .

[0290] Input of the ADD PARTITION command :

[0291] - Input RDF file, including the batch tuples of the partition to be added; / / The input RDF file can be a standard RDF file such as TriG, or a CDC file;

[0292] - Dataset, e.g., the most recent snapshot of the second read-only RDF graph database to be modified by adding a partition;

[0293] - Name of the partition to be added;

[0294] - Reserved character as a delimiter, e.g., "_", called DELIMITER

[0295] Initiation of the ADD PARTITION command

[0296] - Check if the input RDF file name exists;

[0297] - Check that the partition name does not already have the DELIMITER character;

[0298] Upload the START_EXTENSION control file for latch control to storage , The START_EXTENSION control file is named, for example, with the partition name and a specific extension called START_EXTENSION (e.g., ".start"), and is called the "start" file;

[0299] - If the call is successful, the algorithm definitely knows that such a file did not exist previously, and thus for this partition name, no other "add / remove / change partition" is running concurrently;

[0300] - If the call fails, a concurrent "add / remove / change partition" with that partition name is detected: exit and report an error;

[0301] Upload the VALID_EXTENSION control file for partition validity control to the stage , for example, a VALID_EXTENSION control file named with the partition name and a specific extension called VALID_EXTENSION (e.g., ".valid") - check that the "start" file still exists before uploading

[0302] - If the call is successful, the algorithm definitely knows that such a file did not exist previously, so there is no partition with that name;

[0303] - If the call fails, a partition with the input partition name already exists: exit and report an error, indicating that "change partition" should be done alternatively

[0304] Generate fragments of the input RDF file:

[0305] - Parse (or read if it is binary) the input RDF file and split it into smaller files. The smaller files can have a maximum number of triples of SPLIT_SLICE (e.g., 10 million triples), potentially preserving the Skolemization of blank nodes (Skolemization is discussed later); / / Each smaller file will be part of the corresponding fragment, where the fragment includes the smaller file and the associated metadata.

[0306] - For each of these smaller files, generate an RDF file in the selected format. The format can be chosen to optimize the use case, such as a binary RDF format like HDT;

[0307] - If the option is selected, generate a metadata file for the smaller files; / / Fragments are identified by the following pair: the RDF file of the small file and the metadata file;

[0308] - Check that the "start" file is still its own file, and upload all fragments to the stage:

[0309] -- The partition name can be set in the fragment name for integrity checking. For example, all fragments can be named as follows: "fragment_ <uuid> _ <partitionname> . <extension>”, wherein, " <uuid>” is the unique identifier generated, <partitionname>Is the input partition name, <extension>is the file extension (which is different for RDF files (e.g., ".hdt" for ".trig.gz") and metadata files (e.g., ".metadata"))

[0310] -- If one of the uploads fails:

[0311] --- Exit and report an error;

[0312] --- The uploaded files will not be visible. They can be removed here or removed during a later garbage collection process;

[0313] Create a new VALID_EXTENSION control file called "validity file" e.g., a VALID_EXTENSION control file named with the partition name and a specific extension called VALID_EXTENSION;

[0314] - Write a list of all fragment files in this "validity file";

[0315] - Check that the "start" file is still its own file and upload the "validity file" to this stage; / / All fragment files uploaded before this "validity file" will not be considered by the "ALTER REFRESH" command. Thus, this "validity file" ensures the atomicity of the "ADD PARTITION" command.

[0316] Check that the "start" file is still its own file and remove it 。

[0317] / / For all "return" statements of this ADD PARTITION algorithm, and if any exceptions are raised in previous steps, the START_EXTENSION control file must be removed.

[0318] Still referring to the ADD partition command, now discuss partitions. There are multiple strategies for partitioning. Partitions are provided to the command, which means the command is not involved in partitioning. This disclosure is independent of the strategy for creating partitions.

[0319] There may also be overlapping partitions (i.e., triples are repeated in more than one partition). In this case, the management of the overlap will be done both at write queries (e.g., triples must be deleted from all partitions that have it) and read queries (e.g., knowing potential duplicates in the query result). Having or not having overlapping partitions is left as a choice for the software responsible for determining partitions.

[0320] In the example, the REMOVE PARTITION command can execute a sequence of concurrency control operations to be performed. The REMOVE PARTITION command can be executed before the ALTER REFRESH command to ensure the removal of ACID properties on the second RDF graph database of the virtual RDF graph database. The purpose of REMOVE PARTITION is to remove partitions on the second read-only RDF graph database. For example, a typical command sequence can be the creation of the database, the execution of the REMOVE PARTITION command, and then the execution of the ALTER REFRESH command, which makes the deleted partition invisible for further processing (e.g., SPARQL queries).

[0321] Removing a partition is straightforward: it makes all data associated with it invisible, understanding that past transactions can continue to view the data. This means all data for a batch partition, or all segments for a segmented batch partition, or all in-memory data structures for a dynamic partition. The command can be:

[0322] REMOVE PARTITION[NAME]FROM DATASET[Dataset Name]

[0323] This change will only be visible after a successful GARBAGE_COLLECT command, as the remove partition command does not remove files from storage. Examples of the GARBAGE_COLLECT command will be discussed later.

[0324] In the example, the control operations of the REMOVE PARTITION command can include verifying at the stage whether there is a partition to be removed, and if not, stopping the REMOVE PARTITION command.

[0325] Next, the control operations of the REMOVE PARTITION command can upload the TOMBSTONE_EXTENSION control file to the stage. The TOMBSTONE_EXTENSION control file is designed to identify records in the partition that are the objects of the current deletion operation. A successful upload confirms the existence of the partition to be removed in the second read-only RDF graph database, i.e., the records of the partition have not been deleted yet. The TOMBSTONE_EXTENSION control file can store all records to be deleted from the partition, as all these records should be deleted. The TOMBSTONE_EXTENSION control file is now uploaded in this stage; the partition is recorded as deleted.

[0326] Now discuss an example of the DELETE PARTITION command in the form of a pseudocode algorithm. The terms used in this pseudocode are similar to those in the example of the ALTER PARTITION command. The comments in the pseudocode start with / / .

[0327] Input of the DELETE PARTITION command:

[0328] - A dataset, e.g., the most recent snapshot of a second read-only RDF graph database to be modified by deleting a partition;

[0329] - The name of the partition to be deleted;

[0330] Upload the START_EXTENSION control file for latch control to storage , e.g., a START_EXTENSION control file named with the partition name and a specific extension called START_EXTENSION (e.g., ".start", called the "start" file);

[0331] - If the call is successful, the algorithm definitely knows that such a file did not exist previously, and thus there is no other "add / remove / change partition" operation running for that partition name simultaneously;

[0332] - If the call fails, a concurrent "add / remove / change partition" with that partition name is detected: exit and report the error;

[0333] - Check the VALID_EXTENSION control file for partition validity control in this stage, e.g., a VALID_EXTENSION control file named with the partition name and a specific extension called VALID_EXTENSION (e.g., ".valid"):

[0334] -- If not, there is no partition with that name: exit and report the error;

[0335] -- If so, execute the next step of the algorithm;

[0336] Upload the TOMBSTONE_EXTENSION control file to storage , e.g., a TOMBSTONE_EXTENSION control file named with the partition name and a specific extension called TOMBSTONE_EXTENSION (e.g., ".removed");

[0337] If it has been checked that the "start" file is still its own file, perform the upload;

[0338] -- If the call is successful, the algorithm definitely knows that such a file did not exist previously, so the partition has not been removed;

[0339] -- If the call fails, the partition with the input partition name has been removed: exit and report the error;

[0340] Check that the "start" file is still its own file and remove it;

[0341] / / For all "return" statements of the ADD PARTITION algorithm, and if any exception is thrown by the previous step, the START_EXTENSION control file must be removed.

[0342] In the example, to preserve blank node identifiers, the input RDF file can skolemize the blank nodes to identify them as uniquely defined by their identifiers. This is discussed in Tomaszuk, D. and Hyland-Wood, D., 2020, RDF 1.1: Knowledge representation and data integration language for the Web. Symmetry, 12(1), page 84. The ADD / REMOVE / ALTER PARTITION commands can have a "SKOLEMIZED" option to warn them to preserve the identity of blank nodes.

[0343] In the example, a dataset snapshot of the second RDF graph database can be saved in the cache. Additionally or alternatively, when all transactions using the old dataset snapshot and the corresponding unused fragments or partitions are completed, the old dataset snapshot and the corresponding unused fragments or partitions can be removed from the distributed storage. In these examples, the GARBAGE_COLLECT command can be used to remove unused data during the removal phase. As an example, the data that can be used is:

[0344] - Data records;

[0345] - Fragments, i.e., data files: RDF files (e.g., <<trig.gz>>, <<.hdt>>, etc.; depending on the choice made), corresponding metadata (e.g., ".metadata" if this option is selected), delta files (e.g., ".delta");

[0346] - Control files indicating existing partitions, such as VALID_EXTENSION control files and TOMBSTONE_EXTENSION control files;

[0347] - Control files indicating modifications to partitions, such as PENDING_EXTENSION control files and CONSUMED_EXTENSION control files.

[0348] Now discuss an example of the GARBAGE_COLLECT command algorithm. Comments start with / / .

[0349] Input of the GARBAGE_COLLECT command :

[0350] - A data set, such as the most recent snapshot of a second read-only RDF graph database;

[0351] Upload the example START_EXTENSION control file to storage , such as a START_EXTENSION control file named with the data set name and a specific extension called START_EXTENSION (e.g., ".start");

[0352] - If the call is successful, the algorithm definitely knows that such a file did not exist previously, and thus no other "ALTER REFRESH" or "GARBAGE COLLECT" commands are running concurrently for this data set;

[0353] - If the call fails, a concurrent "alter refresh" or "garbage collect" is detected for this data set: exit and report the error

[0354] List all dataset snapshots from the directory and obtain the "most recent dataset snapshot" ; / / All data set snapshots other than the "most recent data set snapshot" are candidates for removal.

[0355] - Ask the cluster of the data set to obtain all data set snapshots currently in use by the connection. Remove these from the candidate list of "data set snapshots to be removed"; / / This request to the directory can be done in several ways. For example, each time a connection is created on a given data set snapshot and each time such a connection is closed, the directory can be notified. / / Note that this is affected by the connection cache;

[0356] - Remove all "data set snapshots to be removed" from the directory;

[0357] - No new connections to the data set will be able to start with these "data set snapshots to be removed"; / / This does not affect existing connections to the data set using these "data set snapshots to be removed";

[0358] - For all data set snapshot candidates to be removed:

[0359] -- Remove the data set snapshot file from this stage;

[0360] -- Remove all CONSUMED_EXTENSION files starting with these data set snapshot names from this stage;

[0361] - List the dataset snapshots that are not candidates for removal to obtain the "active dataset snapshot list";

[0362] List all files in this stage ; / / This operation only needs to be performed once, and the entire list will be copied into the memory of the process. The algorithm uses this copy for the rest of the algorithm. According to the consistent write property of the object store, the list file operation is strongly consistent.

[0363] For all files in this list, take all files ending with the TOMBSTONE_EXTENSION extension 。 / / The corresponding partition name is given by the name of the file without the extension. These are the partition names to be removed.

[0364] - If there are files with the same partition name and VALID_EXTENSION extension, and optionally files with the same partition name and PENDING_EXTENSION extension:

[0365] -- Read the content of this / these files, which gives us the list of files called "list to be deleted" to be deleted from this stage;

[0366] -- Perform the cleaning operation:

[0367] --- Check that the start file is still its own file and delete the VALID_EXTENSION file (if there is one) from this stage;

[0368] --- Delete the PENDING_EXTENSION files from this stage;

[0369] --- Delete all files in the "list to be deleted" from this stage; / / This will delete data files: RDF files, metadata files, delta files;

[0370] --- Delete the TOMBSTONE_EXTENSION files from this stage;

[0371] Check the files to be deleted, even if some inconsistencies are found :

[0372] - TOMBSTONE_EXTENSION files without corresponding VALID_EXTENSION or PENDING_EXTENSION files: Remove them from the stage;

[0373] - CONSUMED_EXTENSION files without corresponding dataset snapshot files: Remove them from the stage;

[0374] - PENDING_EXTENSION files without VALID_EXTENSION or TOMBSTONE_EXTENSION files:

[0375] --List all the files given by the content of the PENDING_EXTENSION file and, if they exist, delete them from this stage;

[0376] --Delete the PENDING_EXTENSION file from this stage;

[0377] --If there is a file with the same partition name and TOMBSTONE_EXTENSION extension, delete it from the stage;

[0378] Execute the optional "full pass"; / / "Full pass" is a more thorough cleaning process : / / In the previous step, the algorithm obtained a list of snapshots of the active data set;

[0379] -Create an empty list called "List of data files in use": / / The "List of snapshots of the active data set" stores snapshots of data sets that are not candidates for removal, active data set snapshots;

[0380] --For each of these active data set snapshot files:

[0381] ---Read the content of the active data set snapshot file and obtain a list of data files for the data set snapshot / / RDF files, metadata files (if any), delta files (if any);

[0382] ---Add this list to the "List of data files in use"

[0383] --Create an empty list called "List of garbage files";

[0384] --Read the list of all files in the stage created at the start of the GARBAGE_COLLECT command algorithm. For each of these data files (i.e., RDF files, metadata files, delta files):

[0385] ---If the file is not also in the "List of data files in use", add it to the "List of garbage files";

[0386] --Remove all files in the "List of garbage files" from this stage;

[0387] Check that the "start" file is still its own file and remove it ;

[0388] / / For all "exit" statements of this algorithm, and if any exceptions were raised by the previous steps, the file must be removed.

[0389] In the example, the CREATE DATASET command can be used to create a new dataset. This command takes as input the location of the phase of the dataset to be created. Here is an example of a potential command to create a dataset:

[0390] CREATE DATASET <name>{

[0391] STAGE = <stage_name>

[0392] }

[0393] As a result of executing the CREATE DATASET command, the association between the logical name of the catalog-registered dataset and the location of the files modeled by the stage.

[0394] In an example, the CREATE STAGE command can be executed to create a new stage. As already defined, the stage of a dataset models the location where the files that make up the dataset are stored. For example, a "stage" is defined here: https: / / docs.snowflake.com / en / sql-reference / sql / create-stage. It is to be understood that the object of the present invention is not to describe the location within the distributed object storage. In other words, the present invention can rely on any known method for describing such a location. In an example, the files can be stored outside the database, for example, inside an external cloud storage device. The storage location can be private or public. In an example, to avoid duplicate information for each stage, storage integration can be used to delegate the authentication responsibility. The document https: / / docs.snowflake.com / en / sql-reference / sql / create-storage-integration describes an example of storage integration. It is to be understood that any known method for storage integration can be used in conjunction with the present disclosure. A single storage integration can support multiple external stages.

[0395] All this information (location of the files, internal or external storage, authentication, etc.) is metadata that describes the dataset. This metadata is directly stored in the catalog, and the metadata is similar to key / value pairs. When creating a storage integration, only general information is given; this information is specific to the cloud storage service. In an example, the following CREATE STORAGE INTEGRATION command can be used to create a dataset:

[0396]

[0397] And for the creation of a stage that uses this storage integration, the following CREATE STAGE command can be used:

[0398]

[0399] Now discuss the concept of a cluster. The concept of a cluster has been introduced in the pseudo-code algorithm of the ALTER_REFRESH command. In this example, the algorithm can use the cluster of the database to connect to the dataset snapshot. The dataset that the command wants to access represents heterogeneous storage. It may be necessary to access physical resources to be able to present the dataset to the SPARQL computing layer in a way that the SPARQL computing layer can use the dataset. Therefore, it may be necessary to link the dataset to a cluster representing the physical resources available for the dataset. As is known per se, a cluster is a set of computing resources called nodes.

[0400] The computing resources can be allocated statically or dynamically, where static means that all resources are available when the cluster starts; and dynamic means that resources automatically increase or decrease the available resources according to the workload. The computing resources can be embedded in the memory of the parent process or as an independent process.

[0401] The computing resources can be used to answer the requirements of the SPARQL computing layer. Answering the requirements of the SPARQL query engine means answering eight triple patterns (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O) and (?S,?P,?O); adding graphs that can be variable, which are the basic graph patterns (BGP). Therefore, the resources of the cluster can be used to answer these BGP for each partition.

[0402] How partitions respond to these BGPs is not specific to this disclosure, and many solutions can be chosen through various trade - offs. The virtual RDF graph database has both dynamic partitioning and batch partitioning. Dynamic partitioning can be implemented by any RDF database, such as Blazegraph described here: https: / / github.com / blazegraph / database, for example, an in - memory embedded RDF database. Batch partitioning can include a set of fragments, where each fragment is an RDF file and a metadata file. The RDF file can be in a human - readable format, such as TriG (https: / / www.w3.org / TR / trig / ) or a binary format, such as HDT; HDT is described in Fernández, J.D., Martínez - Prieto, M.A., Guyérrez, C., Polleres, A. and Arias, M., 2013, Binary RDF representation for publication and exchange (HDT). Journal of Web Semantics, 19, pp. 22 - 41. The metadata includes, but is not limited to, information that helps to easily and quickly evict in cases where the RDF file should be used to answer triple patterns (such as which graphs, predicates, and / or data types to use). Thus, batch partitioning answers a BGP by first asking the metadata if the file is useful, and if so, accessing the file to answer the BGP. For this purpose, the file may require an in - memory database to answer (such as for TriG) or not (such as for HDT). The format of the file can vary without changing the invention.

[0403] Now turning to the connection to a data set. Once a data set is created by the CREATE DATASET command, it is possible to connect to the data set by specifying a cluster (i.e., connecting to a set of computing resources called nodes). For example, the following CONNECT TO DATASET command can be used:

[0404] CONNECT TO DATASET <dataset_name> USING CLUSTER 'EMBEDDED‘

[0405] A process can have several connections to different data sets. A cluster can be used for several data sets.

[0406] In the example of the CONNECT TO DATASET command, an embedded cluster can be used. It is possible to define clusters of different nodes; in this case, the definition depends on the architecture and cloud provider used. In this case, the name of the cluster can be provided instead of the 'EMBEDDED' keyword above.

[0407] After the CONNECT TO DATASET command, the SPARQL query engine can use the dataset to answer its queries. For example, the cluster can then download the metadata of the fragments locally to be able to quickly evict the fragments without accessing external storage. Experiments have shown that the metadata is about 2% to 5% of the total size of the RDF file and is thus a good candidate for prefetching from external storage.

[0408] It is possible to disconnect from the dataset using, for example, the following DISCONNECT FROM DATASET command:

[0409] DISCONNECT FROM DATASET<dataset_name>

[0410] Now discuss the possibility of describing a phase, external storage, or dataset by using a variation of the DESCRIBE command discussed here: https: / / www.w3.org / TR / sparql11-query / #describe. For example, the DESCRIBE command can be:

[0411] DESCRIBE STORAGE INTEGRATION <name>

[0412] DESCRIBE STAGE <name>

[0413] DESCRIBE DATASET <name>

[0414] The parameters given at the CREATE command with potential constraints will be listed. For example, for confidentiality reasons, the credentials given at the creation storage command can be omitted. For the dataset, it can also list all the files of the "most recent dataset snapshot" extracted from the directory.

[0415] Now discuss the possibility of relinquishing (dropping or, drop) a stage, external storage, or dataset by using a variation of the DROP command discussed here: https: / / www.w3.org / TR / sparql11-update / #drop. For example, the DROP command can be:

[0416] DROP STORAGE INTEGRATION <name>

[0417] DROP STAGE <name>

[0418] DROP DATASET <name>

[0419] In the example, the following algorithm can be used to execute the DROP DATASET command. It is understood that this algorithm does not remove files from this stage. A "DROP STAGE" operation is required to complete it.

[0420] Input: Dataset

[0421] Remove all information required to create a connection on the dataset from the directory ; / / Therefore, a new connection cannot be created

[0422] Remove the most recent dataset definition information from the directory ; / / Therefore, a new transaction cannot be created.

[0423] In the example, the following algorithm can be used to execute the DROP STAGE command.

[0424] Input : Stage

[0425] Determine from the catalog that the stage exists and no dataset is using it

[0426] - Otherwise, return the failure;

[0427] Perform a "full" garbage collection operation ;

[0428] - If a currently running transaction is detected (i.e., the "live dataset snapshot" in the algorithm of the GARBAGE_COLLECT command), then return the error; / / The drop stage cannot be completed correctly and will need to be re-run;

[0429] - Otherwise, delete all stage information from the catalog.

[0430] In the example, the following algorithm can be used to execute the DROP STORAGE INTEGRATION command.

[0431] Input : Storage integration

[0432] Determine from the catalog that the storage integration exists and no stage is using it ;

[0433] - Otherwise, return the failure;

[0434] Remove all information about the storage integration from the catalog .

[0435] It has been discussed that computing resources can be used to answer the requirements of the SPARQL computing layer. Answering the requirements of the SPARQL query engine means answering eight triple patterns (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O) and (?S,?P,?O); by adding graphs that can be variable, they are basic graph patterns (BGP). It has also been discussed that the virtual RDF graph database includes dynamic partitioning and batch partitioning; dynamic partitioning can be implemented by any RDF database, and batch partitioning can include a set of fragments, where each fragment is an RDF file and a metadata file. Batch partitioning answers BGP by first asking the metadata whether the file is likely to be useful, and if so, accessing the file to answer BGP.

[0436] Therefore, answering a stored SPARQL query means answering a basic graph pattern (BGP). "Compound transaction" means a transaction that uses two storages (i.e., uses dynamic partitioning and batch partitioning) and thus performs all transaction operations (begin / commit / abort / end) on both storages.

[0437] As seen in the previous discussion, the ACID properties are verified by both the storage that receives streaming modifications and the storage that receives batch modifications. When starting a compound transaction that requires two storage devices, a dataset snapshot (e.g., the most recent dataset snapshot) is extracted and joined with it. This dataset snapshot will be used for all queries of the transaction. The opening of the join can prevent the GARBAGE COLLECT command from deleting files linked from the dataset definition. According to the previous example, the concurrent algorithms on the ADD PARTITION, REMOVE PARTITION commands or ALTER REFRESH, GARBAGE COLLECT commands will not affect the batch partitioning used by the dataset definition, thus ensuring isolation. The same is true for dynamic partitioning according to the "standard" ACID properties of dynamic partitioning. Therefore, the isolation of the compound transaction is ensured. Similarly, atomicity, consistency, and durability are also ensured. Therefore, the compound transaction is actually an ACID transaction.

[0438] Refer to Figure 2 , a computer-implemented method for using isolation to ensure the execution of triple patterns of SPARQL queries on a virtual RDF graph database is proposed. It should be noted that the triple patterns of SPARQL queries can include at least one of the eight triple patterns (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O), and (?S,?P,?O). The method includes obtaining the most recent registered snapshot of the virtual RDF graph database described in the directory (S200). Here, the term snapshot includes a first RDF graph database and a second RDF graph database. When the first virtual RDF graph database is updatable through a tuple stream, the state of the dataset snapshot is obtained for a given time (e.g., the time when the SPARQL query is to be executed). This means that modifications to the first RDF graph database after the given snapshot will not be considered for answering the SPARQL query. For each triple pattern obtained, the triple pattern is executed on the first RDF graph database of the virtual RDF graph database (S210). As already discussed, this is executed as known in the art. Subsequently or simultaneously, the triple pattern is executed on the most recent snapshot of the second RDF graph database of the virtual RDF graph database registered in the directory. As discussed, the most recent snapshot batch partition answers the BGP by first asking the metadata referenced in the directory whether the file might be useful, and if so, accessing the file to answer the BGP. In an example, the batch partition of the second RDF graph database can be segmented. In this example, each segment is accompanied by metadata describing the segment and may be accompanied by a delta file storing updates to the segment. The metadata and the delta file are queried to obtain an answer to the query. For example, for each triple pattern of the query, the tuples of the triple pattern answering the query can be found in the most recent snapshot, thus obtaining a first set of results, and for each triple pattern of the query, it is determined in the delta file whether tuples answering the triple pattern are to be deleted and / or added, and if tuples are to be deleted, the tuples are removed from the first set of results, and if tuples are to be added, the tuples are added to the first set of results.

[0439] The method is computer-implemented. This means that the steps (or substantially all steps) of the method are performed by at least one computer or any system, etc. Therefore, the steps of the method are performed by a computer, possibly fully automatically or semi-automatically. In an example, at least some of the steps of the method can be triggered through user-computer interaction. The required level of user-computer interaction can depend on the level of automation envisioned and be balanced with the need to fulfill the user's wishes. In an example, the level can be user-defined and / or pre-defined.

[0440] A typical example of the computer implementation of a method is to execute the method using a system suitable for that purpose. The system may include a processor coupled to a memory and a graphical user interface (GUI), on which a computer program including instructions for executing the method is recorded. The memory may also store a database. The memory is any hardware suitable for such storage and may include several physically distinct parts (e.g., one for the program and possibly one for the database).

[0441] Figure 6 An example of the system is shown, where the system is a client computer system, such as the user's workstation.

[0442] The client computer of this example includes a central processing unit (CPU) 1010 connected to an internal communication BUS 1000, and a random access memory (RAM) 1070 also connected to the BUS. The client computer may also be provided with a graphics processing unit (GPU) 1110 associated with a video random access memory 1100 connected to the BUS. The video RAM 1100 is also known as a frame buffer in the art. A mass storage device controller 1020 manages access to a mass storage device, such as a hard disk drive 1030. Mass storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks. Any of the foregoing may be supplemented or incorporated by a specially designed ASIC (application specific integrated circuit). A network adapter 1050 manages access to a network 1060. The client computer may also include a tactile device 1090, such as a cursor control device or a keyboard, etc. A cursor control device is used in the client computer to allow the user to selectively position the cursor at any desired position on a display 1080. In addition, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a plurality of signal generating devices for input control signals to the system. Generally, the cursor control device may be a mouse, and the buttons of the mouse are used to generate signals. Optionally or additionally, the client computer system may include a touch pad and / or a touch screen.

[0443] A computer program may include instructions executable by a computer, the instructions including means for causing the above system to perform the method. The program may be recorded on any data storage medium, including the memory of the system. The program may be implemented, for example, in digital electronic circuitry or in computer hardware, firmware, software, or combinations thereof. The program may be implemented as an apparatus, such as a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. The steps of the method may be performed by a programmable processor executing an instruction program to perform the functions of the method by operating on input data and generating output. Accordingly, the processor may be programmable and coupled to receive data and instructions from, and to send data and instructions to, a data storage system, at least one input device, and at least one output device. If desired, the application program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language. In any case, the language may be a compiled or interpreted language. The program may be a full installation program or an update program. The application of the program on the system in any case results in instructions for performing the method. The computer program may optionally be stored and executed on a server in a cloud computing environment that communicates with one or more clients via a network. In such a case, the processing unit executes the instructions included in the program, thereby causing the method to be executed on the cloud computing environment.< / name> < / name> < / name> < / name> < / name> < / name> < / name> < / extension> < / partitionname> < / uuid> < / extension> < / partitionname> < / uuid> < / uuid>

Claims

1. A computer-implemented method for updating a virtual RDF graph database, the virtual RDF graph database comprising tuples, the method comprising: -supply: --File storage, having persistence D in the ACID property, the file storage is distributed or not, and the file storage guarantees consistent writing; --A virtual RDF graph database, the virtual RDF graph database comprising: ---a first RDF graph database capable of being updated by a stream of tuples to be added to and / or removed from the first RDF graph database; ---a second read-only RDF graph database stored on the file storage and capable of being updated by batches of tuples to be added to and / or removed from the second RDF graph database, thereby forming a read-only and updateable snapshot; -- a directory for storing metadata describing the second read-only RDF graph database on the file storage, the directory being ACID compliant; - obtaining a stream of tuples to be added to and / or removed from the virtual RDF graph database, and obtaining a batch of tuples to be added to and / or removed from the virtual RDF graph database; - applying the tuple stream on the first RDF graph database of the virtual RDF graph database; - applying the batch of tuples on the second RDF graph database of the virtual RDF graph database by: -- computing a snapshot of the second RDF graph database, the snapshot including the batch of tuples, the computing ensuring ACID properties of the snapshot by a sequence of concurrency control operations appropriately executed using consistent writes of the file store; --storing the computed snapshot on the file storage; and --Register the computed snapshot in the directory to obtain an updated description of the virtual RDF graph database.

2. The computer-implemented method of claim 1 , wherein: The ALTER REFRESH command executes the sequence of appropriately executed concurrency control operations, the ALTER REFRESH command ensuring updated ACID properties on the second RDF graph database of the virtual RDF graph database, the sequence of concurrency control operations of the ALTER REFRESH comprising: - uniquely identifying a most recent snapshot in the second RDF graph database, the most recent snapshot representing a state of a snapshot formed by the second RDF graph database at a given time; - Verify that no other ALTER REFRESH commands are being executed simultaneously, otherwise stop the ALTER REFRESH command and keep the ALTER REFRESH command that is already being executed; - get a list of files in the stages of all snapshots including the most recent snapshot; - filtering the files in the list of files in the phase of the most recent snapshot, thereby obtaining, at the given time, a list of snapshots naming the files in the most recent snapshot that are available for subsequent access to the database; - write the name of the file in the snapshot list, thereby obtaining a new snapshot not registered in the directory; - uploading the file of the new snapshot to the stage; - registering the new snapshot in the directory after the directory has identified the new snapshot as the most recent snapshot to which the batch of tuples is to be applied in an ACID transaction.

3. The computer-implemented method of claim 2, wherein: Obtaining the snapshot list naming files in the most recent snapshot at the given time, the files in the most recent snapshot being available for subsequent access to the database, comprises: - Get a list of valid partition names; -For each partition named in the list of valid partitions: --Verify whether the current update operation for this partition is still pending, and if so, add this partition to the list of valid partitions for which the current update operation is still pending; --Verify whether a past update operation to the partition has not been successfully implemented, and if so, verify whether a snapshot on which an unsuccessful update operation was performed is registered in the directory: ---If no, ignore the unsuccessful update operation; ---If yes, remove the partition from the list of valid partitions for which the update operation is pending; - Create an empty snapshot list file and complete it by adding the following entry: --The name of the file of a valid partition in the list of valid partitions where the current update operation is still pending; otherwise -- the name of the file of the valid partition in the list of names of valid partitions in case the list of valid partitions still pending from a previous update operation is empty; thereby obtaining, at said given time, said snapshot list naming the files in said most recent snapshot that will be available for subsequent accesses to said database; and - Verify that each file listed in the snapshot list exists in the stage.

4. The computer-implemented method of claim 3, wherein: - obtaining the name list of the valid partitions comprises: for each partition, searching the file storage for a file named with the partition name and a specific extension called VALID_EXTENSION, the VALID_EXTENSION file storing the names of the files of the partition, the VALID_EXTENSION file having been uploaded to the file storage when the partition was created; and / or - obtaining the name list of the valid partitions comprises: for each partition, searching the file storage for a file named with the partition name and a specific extension called TOMBSTONE_EXTENSION, wherein the TOMBSTONE_EXTENSION file stores the names of the files in the partition that are the objects of the current deletion operation, and the TOMBSTONE_EXTENSION file has been uploaded to the file storage when the deletion operation of the partition is performed; and / or - for each partition named in the list of valid partitions, verifying whether the current update operation for the partition is still pending comprises: for each partition, searching the file store for a file named with the partition name and a specific extension called PENDING_EXTENSION, the PENDING_EXTENSION file storing the names of the files in the partition that are the object of the current update operation, the PENDING_EXTENSION file having been uploaded to the file store when the current update of the partition was performed; and / or -For each partition named in the valid partition list, verifying whether a past update operation on the partition has not been successfully implemented includes: for each partition, searching the file storage for a file named with the partition name and a specific extension called CONSUMED_EXTENSION, the CONSUMED_EXTENSION file storing the name of the file in the partition that is the object of the past update operation, and the CONSUMED_EXTENSION file has been uploaded to the file storage when performing the past update of the partition.

5. The computer-implemented method of claim 4, wherein: For each partition named in the list of valid partitions, verifying that a past update operation to the partition has not been successfully implemented further comprises: - verifying that the contents of the PENDING_EXTENSION file are consistent with the contents of the VALID_EXTENSION file, and if inconsistent, considering the file list of the most recent snapshot to be the contents of the VALID_EXTENSION file; and Wherein, for each partition named in the valid partition list, if a current update operation for the partition is still pending, the PENDING_EXTENSION file is ignored for the verification.

6. A computer-implemented method according to any one of claims 2 to 5, wherein: - uniquely identifying the most recent snapshot in the second RDF graph database comprises: obtaining a most recent dataset snapshot identifier from the directory, and storing the most recent dataset snapshot identifier in a memory as a previous dataset snapshot; And wherein registering the new snapshot in the directory further comprises: - Verify, by the catalog in an ACID transaction, that the most recent snapshot described in the catalog is the previous dataset snapshot: --If yes, registering the new snapshot in the directory as the most recent snapshot to which the batch of tuples is applied; --If not, determine that a concurrent ALTER REFRESH command is executed and remove the new snapshot file from the stage.

7. A computer-implemented method according to any one of claims 4 to 6, wherein: The ALTER PARTITION command performs a sequence of further concurrency control operations for updating the partition of the second read-only RDF graph database, the ALTER PARTITION command being executed before the ALTER REFRESH command, the sequence of concurrency control operations of the ALTER PARTITION comprising: - downloading the VALID_EXTENSION file of the valid partition to be updated from the stage, thereby obtaining the name of the file of the valid partition; - Verify that no ALTER PARTITION command is being executed simultaneously on the valid partition, otherwise stop the ALTER PARTITION command and keep the ALTER PARTITION command that is already being executed; - creating a PENDING_EXTENSION file storing the name of the file that is the object of the current update operation in the valid partition, the name of the file in the valid partition being the object of the current update operation and obtained from the batch of tuples; and -Upload the PENDING_EXTENSION file to the stage.

8. The computer-implemented method of claim 7, wherein: The snapshot of the second RDF graph database includes a set of batch partitions, one or more batch partitions are divided into fragments, and creating the PENDING_EXTENSION file includes: -Determine if the batch of tuples is a change data capture (CDC) file and do the following: --If the batch of tuples is not a CDC file, generate a fragment of the batch of tuples to obtain a list of fragment files; otherwise --If the batch tuple is a CDC file, then for each add / delete triplet of the CDC file, locate the fragment of one of the one or more batch partitions to which the add / delete operation is to be applied, thereby obtaining a list of fragments with an additional list of delta files, if any; And wherein creating the PENDING_EXTENSION file comprises: creating the PENDING_EXTENSION storing a list of fragments with an additional list of delta files, if any.

9. A computer-implemented method according to any one of claims 4 to 6, wherein: The ADD PARTITION command executes a sequence of further concurrency control operations for adding a partition on the second read-only RDF graph database, the ADD PARTITION command being executed before the ALTER REFRESH command, the sequence of concurrency control operations of the ADD PARTITION comprising: - Upload the VALID_EXTENSION file of the partition to be added to the storage, the successful upload of the VALID_EXTENSION file indicating that there is no partition with the same name in the file storage; - generating fragments of the batch of tuples for adding partitions on the second read-only RDF graph database, thereby obtaining a list of fragment files; - Storing the list of fragments in the VALID_EXTENSION file and uploading the VALID_EXTENSION file to the stage.

10. A computer-implemented method according to any one of claims 4 to 6, wherein: A REMOVE PARTITION command executes a sequence of further concurrency control operations for removing the partition of the second read-only RDF graph database, the REMOVE PARTITION command being executed before the ALTER REFRESH command, the sequence of concurrency control operations of the REMOVE PARTITION comprising: - verifying at said stage whether said partition to be removed exists, and if not, stopping said REMOVEPARTITION command; -Upload the TOMBSTONE_EXTENSION file to the stage, and a successful upload confirms the existence of the partition to be removed in the second read-only RDF graph database.

11. The computer-implemented method of any one of claims 2 to 10, further comprising: - latching the second read-only RDF graph database stored on the file storage, the latching being performed before obtaining the file list of the most recent snapshot in the phase; - unlocking the new snapshot after it is registered in the directory.

12. The computer-implemented method of claim 11, wherein: The latching of the second read-only RDF graph database comprises: - uploading a file named with the name of the most recent dataset snapshot and a specific extension called START_EXTENSION to the file storage; - verifying the presence of said START_EXTENSION file in said file storage when uploading said file of said new snapshot in said stage and in each subsequent step; And wherein, unlocking the new snapshot after the new snapshot is registered in the directory comprises: verifying the existence of the START_EXTENSION file in the file storage and removing the START_EXTENSION file.

13. A computer-implemented method according to any one of claims 1 to 12, wherein: The file storage is a distributed file storage and / or the first RDF graph database is stored on an in-memory data structure.

14. A computer-implemented method according to any one of claims 1 to 13, wherein: The directory for storing metadata is stored on a database that ensures ACID properties and strong consistency.

15. A computer-implemented method for executing a SPARQL query on a virtual RDF graph database previously updated according to any one of claims 1 to 14 using isolation guarantees for triple patterns, comprising: - obtaining the most recently registered snapshot in the virtual RDF graph database described in the directory; - For the obtained triple pattern, executing the triple pattern on the first RDF graph database and the obtained latest snapshot of the virtual RDF graph database, thereby ensuring isolation of the execution of the triple pattern.