Execution of sparql query over RDF dataset stored in distributed storage

The method addresses the challenge of executing SPARQL queries with ACID guarantees in heterogeneous RDF databases by using ACID-compliant storage and metadata catalog to ensure data integrity during updates and deletions in distributed environments.

JP2025100403APending Publication Date: 2025-07-03DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024213082
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-06
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing RDF databases lack a method to execute SPARQL queries with ACID guarantees for triple updates and deletions in a heterogeneous distributed storage environment, where non-ACID compliant storage is prevalent.

Method used

A computer-implemented method providing a file storage with ACID durability, a stream-updatable first RDF graph database, a batch-updatable second read-only RDF graph database, and an ACID-compliant metadata catalog, ensuring consistent writing and concurrent control operations to maintain ACID properties during updates.

Benefits of technology

Enables execution of SPARQL queries with ACID guarantees across heterogeneous distributed storage, ensuring data integrity and consistency during triple updates and deletions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100403000001_ABST
    Figure 2025100403000001_ABST
Patent Text Reader

Abstract

To provide a method to be implemented by a computer for updating a virtual RDF graph (directed, labeled graph data format for representing information in the Web) database comprising tuples.SOLUTION: A method comprises: providing a file storage having a durability D property of ACID (Atomicity, Consistency, Isolation and Durability) property and guaranteeing consistent write; and providing a virtual RDF graph database comprising a first RDF graph database updatable by streams and a second read-only RDF graph database stored on the file storage and updatable by batches, and a catalog for storing metadata describing the second read-only RDF graph database on the file storage, the catalog being compliant with ACID.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly, to a method, system, and program for updating an RDF graph database split on storage in a heterogeneous environment.

Background Art

[0002] The RDF specification (https: / / www.w3.org / TR / rdf11-concepts / ) for representing information as a graph using triples has been published by the W3C: "The core structure of the abstract syntax is a set of triples, each consisting of a subject, a predicate, and an object. Such a set of triples is called an RDF graph. An RDF graph can be visualized as a diagram of nodes and directed arcs, and each triple is represented as a node-arc-node link." An RDF graph stores RDF data. Here, RDF data means triples within an RDF graph as defined above.

[0003] SPARQL is a query language for RDF data (https: / / www.w3.org / TR / sparql11-query / ): "RDF is a directed labeled graph data format for representing information on the Web. This specification defines the syntax and semantics of the SPARQL query language for RDF. SPARQL can be used to express queries across various data sources, regardless of whether the data is stored natively as RDF or presented as RDF via middleware. SPARQL has the ability to query essential and optional graph patterns along with their conjunctions and disjunctions. SPARQL also supports aggregation, subqueries, negation, value creation by expressions, test of extensible values, and constraint of queries by the source RDF graph. The result of a SPARQL query is a set of results or an RDF graph."

[0004] RDF knowledge graphs have grown to billions of triples. To be able to scale to this data volume, it is necessary to distribute both the computing and storage of SPARQL queries over RDF data. In cloud environments, the cost models for computing and storage are different. Therefore, like Snowflake for SQL data, both are distributed and scaled separately. Computing and storage are separated and distributed separately.

[0005] Here, we will explain the distribution of storage.

[0006] In the field of RDF, several techniques can be used to determine how to partition data within distributed storage. This is described, for example, in Kalogeros, E., Gergatsoulis, M., Damigos, M., and Nomikos, C. (2023), "Efficient query evaluation techniques over large amount of distributed linked data", Information Systems, 115, p. 102194.

[0007] Common methods such as SPARQLGFX are published in Graux, D., Jachiet L., Geneves, P., and Layaida, N., 2016. "Sparqlgx: Efficient distributed evaluation of sparql with apache spark", Semantic Web - ISWC 2016: The 15th International Semantic Web Conference, Kobe, Japan, October 17 - 21, 2016, Proceedings, Part II 15 (pp. 80 - 87). SPARQLGFX uses distributed file systems such as Hadoop File System (HDFS) and Amazon's Simple Storage Service (S3). A comparison between the two is available at https: / / www.integrate.io / blog / storing - apache - hadoop - data - cloud - hdfs - vs - s3 / . For the reasons described in this article, S3 is the target in the industry rather than HDFS.

[0008] Many distributed RDF databases are read - only or do not consider how to modify partitions. "Partout" is described in Galarraga, L., Hose, K., Schenkel, R., April 2014, Partout: a distributed engine for efficient RDF processing, Proceedings of the 23rd International World Wide Web Conference (pp. 267 - 268). "Partout" accepts graph updates but does not provide ACID guarantees. "Partout" accepts updates, but only for batch updates, so modifications are not made in streaming or immediately visualized for queries.

[0009] The ACID properties (ACID stands for Atomicity, Consistency, Isolation, and Durability) are a set of properties of database transactions aimed at ensuring the validity of data even in case of errors, power outages, and other accidents. Thanks to the ACID properties, for example, it becomes possible for all SPARQL queries to refer to the same version of data even when a transaction is started and data is simultaneously updated in one or more partitions during this transaction. For example, the ACID properties are explained at https: / / en.wikipedia.org / wiki / ACID.

[0010] The literature Zou, L., Ozsu, M. T., 2017, Graph-based RDF data management., Data Science and Engineering, 2, pp. 56 - 70, and Ali, W., Saleem, M., Yao, B., Hogan, A., Ngomo, A. C. N., 2022, A survey of RDF stores & SPARQL engines for querying knowledge graphs, VLDB Journal, pp. 1 - 26 are recent approaches to executing SPARQL queries on distributed storage and can be seen as NoSQL approaches to traditional SQL approaches as defined in Valduriez, P., Jimenez-Peris, R. and Ozsu, M. T., 2021. Distributed database systems: The case for newSQL. In Transactions on Large-Scale Data-and Knowledge-Centered Systems XLVIII: Special Issue. In Memory of Univ. Prof. Dr. Roland Wagner (pp. 1 - 15).

[0011] Similar needs also arise in the field of relational SQL distributed storage, but there is a difference in that relational databases operate on schemaful tables rather than schema - less graphs of RDF. Therefore, the techniques used in the relational world cannot be directly applied to the graph database world. In fact, relational databases and graph databases rely on different paradigms.

[0012] As disclosed in the Iceberg table specification, Apache Iceberg 1.3.0, online, at https: / / iceberg.apache.org / spec / , Apache Iceberg is "an open table format for large analytical datasets. Iceberg uses a high - performance table format that functions like an SQL table to add a table to a computing engine." The Iceberg table format is "designed to manage a large, slowly changing collection of files within a distributed file system or key - value store as a table." That is, an Iceberg table consists of multiple binary files stored in distributed storage and is accessed via a metadata file that provides a list of files corresponding to a snapshot of the table's state at a particular point in time. This idea is similar to that found in the literature Zou,L., Ozsu,M.T., 2017, Graph - based RDF data management., Data Science and Engineering, 2, pp.56 - 70, and Ali,W., Saleem,M., Yao,B., Hogan,A., Ngomo,A.C.N., 2022, A survey of RDF stores&SPARQL engines for querying knowledge graphs, VLDB Journal, pp.1 - 26, but Iceberg provides ACID guarantees through metadata and manifest files. However, since Iceberg is completely file - based, it does not support streaming changes and only supports batch updates.

[0013] Snowflake is described in Snowflake Key Concepts&Architecture, 2023 Snowflake Inc., https: / / docs.snowflake.com / en / user-guide / intro-key-concepts, and Snowflake provides a feature called Snowpipe Streaming. This combines both Snowpipe, which loads data from files in micro-batches similar to Iceberg, and a streaming API that loads streaming data rows with low latency, into a single table. Although Snowflake seems to have ACID transaction semantics, since Snowflake is a commercial product, few details have been made public. The limitation of Snowflake is that only row insertions are supported.

Summary of the Invention

Problems to be Solved by the Invention

[0014] In such a context, there is still a need for an improved RDF heterogeneous environment distributed storage mainly consisting of non-ACID compliant storage that can execute SPARQL queries including triple updates / deletions with ACID guarantees.

Means for Solving the Problems

[0015] The present invention provides a computer-implemented method for updating a virtual RDF graph database containing tuples. The method of the present invention is to provide a file storage having the durability D property of ACID characteristics, which is a distributed or non-distributed file storage that guarantees consistent writing, and a virtual RDF graph database, A first RDF graph database that can be updated by a stream of tuples added and / or deleted to the first RDF graph database, and A second read-only RDF graph database stored in the file storage, which can be updated by a batch of tuples added and / or deleted to the second RDF graph database, thereby forming a read-only and updatable snapshot, and A virtual RDF graph database including A catalog for storing metadata describing the second read-only RDF graph database on the file storage, the catalog being ACID-compliant, and Providing Obtaining a stream of tuples added and / or deleted to the virtual RDF graph database, and obtaining a batch of tuples added and / or deleted to the virtual RDF graph database, and Applying the stream of tuples to the first RDF graph database of the virtual RDF graph database, and Applying the batch of tuples to the second RDF graph database of the virtual RDF graph database, where Calculating a snapshot of the second RDF graph database including the batch of tuples, and guaranteeing the ACID properties of the snapshot by a series of concurrent control operations appropriately executed using consistent writing to the file storage, and Storing the calculated snapshot in the file storage, and Registering the calculated snapshot in the catalog, thereby obtaining an updated description of the virtual RDF graph database, and Including applying Including

[0016] This method may include one or more of the following.

[0017] The ALTER REFRESH command executes a series of concurrent control operations that are properly executed as described above. The ALTER REFRESH command guarantees the ACID properties of the update of the second RDF graph database in the virtual RDF graph database. The series of concurrent control operations of the ALTER REFRESH uniquely identifies the last snapshot of the second RDF graph database, which represents the state of the snapshot formed by the second RDF graph database at a specific point in time, and checks that no other ALTER REFRESH command is being executed simultaneously. If one is being executed, the ALTER REFRESH command is stopped and the already executed ALTER REFRESH command is held. It obtains a list of files at a certain stage of all snapshots (groups) including the last snapshot, filters the files in the list of files at the stage of the last snapshot, thereby obtaining a snapshot list indicating the names of the files of the last snapshot that can be used for subsequent access to the database at the specific point in time, writes the names of the files in the snapshot list, thereby obtaining a new snapshot not registered in the catalog, uploads the files of the new snapshot to the stage, and after the catalog identifies the new snapshot as the last snapshot to which the batch of the tuples is applied in an ACID transaction, registers the new snapshot in the catalog.

[0018] Obtaining the snapshot list indicating the name of the file of the last snapshot that will be available for subsequent access to the database at the above-specified time point involves obtaining a list of names of valid partition(s), for each partition named in the list of the valid partition(s), checking whether the current update operation for the partition is still pending, and if it is, adding the partition to the list of valid partition(s) for which the current update operation is still pending, checking whether the past update operation for the partition has not been executed successfully, and if not, checking in the catalog whether a snapshot for which the failed update operation was executed is registered in the catalog, and if not, ignoring the failed update operation, and if it is, deleting the partition from the list of valid partition(s) for which the update operation(s) is / are pending, creating an empty snapshot list file, and adding the name of the file(s) of the valid partition(s) from the list of valid partitions for which the current update operation(s) is / are pending, or, if the list of valid partition(s) for which the previous update operation(s) is / are still pending is empty, adding the name of the file(s) of the valid partition(s) from the list of names of the valid partition(s), thereby obtaining the snapshot list indicating the name of the file(s) of the last snapshot that will be available for subsequent access to the database at the above-specified time point, and checking that each file listed in the snapshot list exists in the above stage.

[0019] Obtaining the list of names of the above-mentioned valid partition(s) involves, within the above file storage, for each partition, searching for a file named with the partition name and a specific extension called VALID_EXTENSION, where the VALID_EXTENSION file stores the name of the file(s) of the partition and the VALID_EXTENSION file was uploaded to the file storage when the partition was created, and / or obtaining the list of names of the above-mentioned valid partition(s) involves, for each partition, searching for a file named with the partition name and a specific extension called TOMBSTONE_EXTENSION within the above file storage, where the TOMBSTONE_EXTENSION file stores the name of the file(s) of the partition that is the target of the current deletion operation and the TOMBSTONE_EXTENSION file was uploaded to the file storage when the deletion operation of the partition was executed, and / or for each partition whose name is indicated in the list of the above-mentioned valid partition(s), checking whether the current update operation for the partition is still pending involves, within the above file storage, for each partition, searching for a file named with the partition name and a specific extension called PENDING_EXTENSION, where the PENDING_EXTENSION file stores the name of the file(s) of the partition that is the target of the current update operation and the PENDING_EXTENSION file was uploaded to the file storage when the current update of the partition was executed, and / or for each partition whose name is indicated in the list of the above-mentioned valid partition(s), checking whether the past update operation for the partition was not executed properly involves, within the above file storage, for each partition,Searching for a file named with the partition name and a specific extension called CONSUMED_EXTENSION, where the CONSUMED_EXTENSION file stores the name(s) of the file(s) of the partition that were the target of past update operations, and the CONSUMED_EXTENSION file is uploaded to the file storage when past updates of the partition are executed, including the searching.

[0020] For each partition named in the list of the valid partition(s), to check whether past update operations for the partition have been executed properly, it is necessary to confirm that the content of the PENDING_EXTENSION file matches the content of the VALID_EXTENSION file. If they do not match, it further includes considering the list of files of the last snapshot as the content of the VALID_EXTENSION file. For each partition named in the list of the valid partition(s), if the current update operation for the partition is still pending, the PENDING_EXTENSION file is ignored during the above confirmation.

[0021] Uniquely identifying the last snapshot of the second RDF graph database involves obtaining the last dataset snapshot identifier from the catalog and retaining the last dataset snapshot identifier in memory as the previous dataset snapshot. Registering the new snapshot in the catalog involves the catalog confirming, in an ACID transaction, that the last snapshot described in the catalog is the previous dataset snapshot. If so, registering the new snapshot in the catalog as the last snapshot to which the batch of tuples is applied. If not, it further includes determining that the ALTER REFRESH command has been executed and deleting the new snapshot file from the stage.

[0022] The ALTER PARTITION command executes a further sequence of concurrent control operations for updating the partitions of the second read-only RDF graph database. The ALTER PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrent control operations for ALTER PARTITION involves downloading the VALID_EXTENSION file of the valid partition to be updated from the stage, thereby obtaining the name(s) of the file(s) of the valid partition, and confirming that the ALTER REFRESH command is not being concurrently executed on the valid partition. If it is being executed, stopping the ALTER REFRESH command and retaining the already executed ALTER REFRESH command, and creating a PENDING_EXTENSION file that stores the name(s) of the file(s) of the valid partition that is the target of the current update operation, where the name(s) of the file(s) of the valid partition is (are) obtained from the batch of tuples and is (are) the target of the current update operation, and uploading the PENDING_EXTENSION file to the stage.

[0023] The snapshot of the second RDF graph database consists of a set of batch partitions, one or more batch partitions are split into fragments, and the creation of the PENDING_EXTENSION file determines whether the batch of the tuples is a Change Data Capture (CDC) file. If the batch of the tuples is not a CDC file, fragments of the batch of the tuples are generated, thereby obtaining a list of fragment files. Or if the batch of the tuples is a CDC file, for each add / delete triple of the CDC file, one fragment of one or more batch partitions to which the add / delete operation is applied is identified, thereby obtaining a list of fragments, along with an additional list of delta files if any. The creation of the PENDING_EXTENSION file includes creating the PENDING_EXTENSION that stores the list of the fragments, along with an additional list of delta files if any.

[0024] The ADD PARTITION command executes a further sequence of concurrent control operations for adding partitions to the second read-only RDF graph database. The ADD PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrent control operations for ADD PARTITION includes uploading the VALID_EXTENSION file of the partition to be added to the storage, where the success of the upload of the VALID_EXTENSION file indicates that there is no partition with the same name in the file storage; generating a batch fragment of the tuples for adding the partition on the second read-only RDF graph database, thereby obtaining a list of fragment files; storing the list of fragments in the VALID_EXTENSION file and uploading the VALID_EXTENSION to the stage.

[0025] The REMOVE PARTITION command executes a further sequence of concurrent control operations for deleting partitions from the second read-only RDF graph database. The REMOVE PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrent control operations for REMOVE PARTITION includes checking on the stage whether the partition to be deleted exists, and if it does not exist, stopping the REMOVE PARTITION command; uploading a TOMBSTONE_EXTENSION file to the stage, where the success of the upload verifies that the partition to be deleted exists in the second read-only RDF graph database.

[0026] Latch the second read-only RDF graph database stored in the file storage, which is executed before obtaining the list of files of the last snapshot in the stage. Unlatch after registering the new snapshot in the catalog.

[0027] Latching the second read-only RDF graph database includes uploading to the file storage a file named with the last dataset snapshot name and a specific extension called START_EXTENSION, and confirming that the START_EXTENSION file exists in the file storage when uploading the file of the new snapshot in the stage and in each subsequent step. Unlatching the new snapshot after registering it in the catalog includes confirming that the START_EXTENSION file exists in the file storage and deleting the START_EXTENSION file.

[0028] The file storage is distributed file storage, and / or the first RDF graph database is stored in an in-memory data structure.

[0029] The catalog for storing metadata is stored in a database that guarantees ACID properties and strong consistency.

[0030] Therefore, a computer-implemented method is provided for executing triple patterns of SPARQL queries with guaranteed independence against a previously updated virtual RDF graph database. This method includes obtaining the last registered snapshot of the virtual RDF graph database described in the catalog, For the obtained triple pattern, execute the triple pattern against the first RDF graph database of the virtual RDF graph database and the obtained last snapshot, thereby guaranteeing the independence of the execution of the triple pattern.

[0031] Furthermore, a computer program including instructions for executing the above update method and / or the above execution method is provided.

[0032] Furthermore, a computer-readable storage medium on which the computer program is recorded is provided.

[0033] Furthermore, a system including a processor connected to a memory on which the computer program is recorded is provided.

[0034] Furthermore, an apparatus including a data storage medium on which the computer program is recorded is provided. This apparatus may form or serve as a non-transitory computer-readable medium, for example, in SaaS (Software as a Service) or other servers, or cloud-based platforms. Alternatively, the apparatus may include a processor connected to the data storage medium. Therefore, the apparatus may form a computer system wholly or partially (for example, the apparatus is a subsystem of the entire system). The system may further include a graphical user interface connected to the processor. Here, non-limiting examples will be described with reference to the accompanying drawings.

Brief Description of the Drawings

[0035]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0036] Referring to the flowchart of FIG. 1, a computer-implemented method for updating a virtual RDF graph database is proposed. The "virtual RDF graph database" means that at least two graphs are recognized as one logically equivalent RDF graph database, excluding relational databases that may be adapted to expose the content of any relational database, such as https: / / github.com / ontop / ontop, as a knowledge graph. The virtual RDF graph database consists of RDF triples (or simply triples). This method includes providing a file storage with the Durability D property of ACID characteristics. The file storage may or may not be distributed. The file storage guarantees consistent writes. "Consistent writes" is related to the consistency model, which describes the conditions under which write operations from one client become visible to other clients. Consistent writes are also called "Strong Read-After-Write Consistency" and are described at https: / / aws.amazon.com / fr / blogs / aws / amazon-s3-update-strong-read-after-write-consistency / . In the present invention, "consistent writes" means that it is guaranteed that the reading side can refer to all previously committed writes. This method further includes providing a virtual RDF graph database. The virtual RDF graph database consists of a first RDF graph database that can be updated by a stream of triples added and / or deleted to the first RDF graph database, and a second read-only RDF graph database stored in the file storage that can be updated by a batch of triples added and / or deleted to the second RDF graph database, thereby forming a read-only and updatable snapshot.The method is a catalog for storing metadata that describes the second read-only RDF graph database on the file storage, and further includes providing a catalog compliant with ACID. Next, the method includes obtaining a stream of tuples to be added and / or deleted to / from the virtual RDF graph database, and obtaining a batch of tuples to be added and / or deleted to / from the virtual RDF graph database. The method also includes applying the stream of tuples to the first RDF graph database of the virtual RDF graph database. The method further includes applying the batch of tuples to the second RDF graph database of the virtual RDF graph database. The application is calculating a snapshot of the second RDF graph database including the batch of tuples, which is guaranteed the ACID properties of the snapshot by a series of concurrent control operations properly executed using consistent writing to the file storage, storing the calculated snapshot in the file storage, registering the calculated snapshot in the catalog, thereby obtaining an updated description of the virtual RDF graph database, and is executed by. A catalog refers to a collection of metadata and information about the resources of a virtual RDF graph database.

[0037] The advantages of this method will be described below.

[0038] Using the terms of the W3C at https: / / www.w3.org / TR / rdf11-concepts / , "RDF data" is defined as triples within an RDF graph. According to https: / / www.w3.org / TR / sparql11-query / #rdfDataset, an "RDF dataset" is a collection of such RDF graphs, and an RDF database is a collection of RDF datasets. Continuing, using terms known in the art (https: / / en.wikipedia.org / wiki / Partition_(database)), a "partition is the act of dividing a logical database or its components into separate, independent parts". The situation where partitions overlap will be described later.

[0039] As described in the Background section, a database can be regarded as consisting of computing (for query processing) and database storage. In the present invention, computing and database storage are regarded as separate resources of different scales.

[0040] The present invention focuses on database storage, and a method of managing computing is considered outside the scope of the present invention. The method of the present invention to obtain an improved RDF heterogeneous environment distributed storage mainly consisting of ACID-non-compliant storage that can execute SPARQL queries including triple updates / deletions with ACID guarantees is independent of the method of managing the above computing. Note that the computing may or may not be distributed. For simplicity, the computing part is considered not to be distributed, but since the present invention is independent of the method of managing computing, the computing part can be distributed without changing the present invention. When the computing is distributed, the catalog is also distributed.

[0041] Furthermore, the RDF data of a specific RDF dataset is split into distributed storage using specific partitioning rules that can be, for example, hashing of the subject. There are many ways to split RDF data, which is not the subject of the present invention. The only requirement is to know which quad is in which partition. Therefore, it is considered that there are means, such as using statistical information, to probe the partitions during query execution to know whether a partition can be used in the execution flow.

[0042] Here, the concept of a quad will be explained. In one example, the RDF triples of an RDF dataset may be RDF quads. An RDF quad may be obtained by adding a graph label to an RDF triple. In such an example, the RDF triple includes an RDF graph. A standard specification for specifying RDF Quads (also called N-Quads) has been published by the W3C. See, for example, "RDF 1.1 N-Quads, A line-based syntax for RDF datasets" (W3C Recommendation of February 25, 2014). An RDF quad may be obtained by adding a graph name to an RDF triple. The graph name may be empty (i.e., in the case of a default or unnamed graph) or an IRI (i.e., a graph IRI). In one example, the predicate of the graph may have the same IRI as the graph IRI. The graph name of each quad is the graph to which the quad belongs within each RDF dataset. An RDF dataset represents, as is known (e.g., https: / / www.w3.org / TR / rdf-sparql-query / #rdfDataset), a collection of graphs. For simplicity, a quad can be summarized as a triple that includes a reference to a graph.

[0043] Returning to FIG. 1, the method updates the virtual RDF graph database. As described above, the virtual RDF graph database consists of at least two graphs that are logically recognized as one equivalent RDF graph database. Therefore, updating the virtual RDF graph database means updating the logical database.

[0044] Here, the above-provided (S100) will be described.

[0045] · File storage The method provides file storage with the durability D property of the ACID property. As is known, "file storage" means storing and organizing data in the form of files within a file system. The data is grouped into individual files, each identified by a unique name or path, and these files are hierarchically organized within directories or folders. As is known, the "durability property" of the ACID property ensures that when a transaction is committed, the changes made to the data are retained and can withstand subsequent failures such as system crashes or power outages. That is, when the success of a file write or file update operation is confirmed, the data is safely stored and will not be lost even if a failure occurs. "A transaction is committed" means that the changes made by the transaction are finalized and the system ensures its durability and visibility to other transactions.

[0046] File storage is either distributed or not. In a distributed file storage system, files and data are distributed across multiple servers or storage nodes. File distribution is often done to improve scalability, fault tolerance, and performance. For example, Hadoop Distributed File System (HDFS) and Google File System (GFS) are examples of distributed file systems. Many cloud storage solutions such as Amazon S3, Google Cloud Storage, and Microsoft Azure Storage are essentially distributed file storage. A non-distributed file storage system, also called centralized file storage, stores all files and data on a single server or a limited number of servers. Users and clients connect to this central server to access files. All file requests and data retrievals go through this central point.

[0047] File storage guarantees consistent writes. As is known, "consistent writes" refer to write operations that guarantee data consistency, that is, the reader is guaranteed to be able to reference all previously committed writes.

[0048] · Virtual RDF graph database Furthermore, in S100, a virtual RDF graph database is provided. As is known, a "virtual RDF graph database" refers to a database system that provides a virtual approach to RDF queries and access. The term "virtual" indicates that the database can query and retrieve data from multiple distributed sources or endpoints. It should be understood that a virtual RDF graph database consists of at least two RDF graph databases that are logically recognized as one equivalent RDF graph database. In this definition, relational databases that may be adapted to expose the content of any relational database, such as https: / / github.com / ontop / ontop, as a knowledge graph are excluded. Instead of integrating all data into a single physical database, using a virtual RDF graph database allows data from various distributed sources to be referenced as if they were part of a single integrated database.

[0049] The virtual RDF graph database consists of the two RDF graph databases described here.

[0050] · The first RDF graph database The virtual RDF graph database includes a first RDF graph database. The first RDF graph database is updatable by a stream of tuples added and / or deleted to / from the first RDF graph database. As is known, streaming revision is a way for an RDF dataset to incorporate RDF data, i.e., a way to update an RDF dataset. The update is performed with low latency using standard SPARQL update queries. The addition and deletion of triples, and / or the addition and deletion of graphs are performed on the same dataset and become immediately available. In other words, one or more partitions appear to provide the same revision capabilities as a standard graph database without distributed storage. These partitions are also called "dynamic partitions". It is recalled that the term "partition" is defined as "dividing a logical database or its components into separate independent parts". In the context of streaming revision, the standard SPARQL update of an RDF dataset is applied to the partitions of the database and can be revised in the same way as other graph databases.

[0051] · A second RDF graph database The virtual RDF graph database includes a second RDF graph database. The second RDF graph database is stored in a file storage. The second RDF graph database is read-only, which means that direct update operations cannot be performed on the dataset. Therefore, the second RDF graph database is a snapshot, that is, it represents a static point in time of the RDF data of the second RDF graph database. The second RDF graph database can be updated by batches of tuples that are added and / or deleted to / from the second RDF graph database. Therefore, the second RDF graph database forms a read-only but batch-updatable snapshot. In the initial state, that is, the state where the second RDF graph database has not been updated by batches of tuples that are added and / or deleted to / from the second RDF graph database, the second RDF graph database is a snapshot. In the case of batch updates, since batch partitions are associated with the second RDF graph database (for example, in the form of a delta file as described below), the snapshot may include two or more snapshots. In the following description, the term "snapshot" may include the initial state (one snapshot of the second RDF graph database) or a situation where the updated second RDF graph database consists of multiple snapshots. Here, the batch partition will be described.

[0052] As is known, batch update refers to batch loading. Batch loading means loading data from an RDF file into partitions called "batch partitions" (also called archive partitions). The RDF file consists of tuples to be added and / or deleted to / from a second RDF graph database. The format of the RDF file is arbitrary, for example, the Turtle format (described at https: / / www.w3.org / TR / turtle / ), the TriG format (described at https: / / www.w3.org / TR / trig / ), or an equivalent format. Thus, the batch stores updates to be applied to the second RDF graph database, and the batch partition stores updates applied to the second RDF graph. In both cases, since the snapshot is read-only, the snapshot itself is not directly affected by modifications. Applied modifications are supported only in the batch partition. Batch partitions can be added, deleted, or modified for a dataset.

[0053] The second read-only RDF graph database is a read-only snapshot. As is known, a dataset snapshot consists of the locations of all files in distributed storage, their metadata, and optional delta files. By all file locations, it means that the snapshot contains information about the locations of all files within the dataset. This information indicates where each file is stored within the distributed storage infrastructure. This may include references to specific storage nodes, addresses, or paths. The metadata of the snapshot refers to additional information regarding the characteristics, origin, and structure of the RDF graph database. For example, the metadata consists of, but is not limited to, bibliographic information of the RDF graph database (e.g., title, author, creation date, modification date, etc.), the format in which the RDF data is encoded (e.g., turtle, TriG, etc.), the number of triples, SPARQL endpoints, storage locations, etc. Thus, the metadata provides context and facilitates the configuration and management of the dataset. Delta files represent the changes or differences between two versions of the data. Including optional delta files in the dataset snapshot enables efficient tracking of changes over time. These delta files may contain information regarding updates, additions, or deletions made to the dataset since the last snapshot. Including optional delta files is particularly useful for minimizing the amount of data that needs to be transferred or stored when creating a snapshot, as only the changes since the previous snapshot are captured. Note that, for illustrative purposes only, when a new snapshot is calculated, the algorithm can access references (e.g., paths, etc.) to old data files that are still valid and references to new data files, so it should be noted that there is no need to replicate the entire previous snapshot. Also, if only a few triples have been modified, the delta file does not need to regenerate the entire data file. These two mechanisms function separately to avoid duplication of the entire snapshot.

[0054] In one example, as is known, a batch partition may be divided into n fragments in order to minimize the impact of updates, where n is a positive integer (n≥1). Dividing a partition into fragments means dividing the partition into a set of subgraphs, and each subgraph is called a fragment. Here, a subgraph means a subset of the triples of the partition. A batch partition consists of tuples to be added and / or deleted to / from a second RDF graph database and is decomposed into subgraphs that are fragments. This means that a SPARQL query that causes the generation of a batch of tuples to be added and / or deleted to / from a second RDF graph database is decomposed (e.g., randomly) into subgraphs (i.e., subsets of triples) and then executed on the second RDF graph database. Dividing a batch partition into fragments is an optimization and not essential to the present invention.

[0055] Fragments may be used in situations where the overall size of the batch partition is very important, for example, when the size of the batch partition is larger than a predetermined size (e.g., larger than 1 gigabyte). The number of fragments may be determined according to a fragmentation strategy, such as the maximum size of the fragments or parallel processing of the fragments.

[0056] In one example, in order to control the use of resources such as file handles, small partitions may be merged into one large fragment. The number of small partitions to be merged may be determined according to a merging strategy, such as the minimum size of the partitions. Merging small partitions into one large fragment is an optimization and not essential to the present invention.

[0057] Even in these examples, the fragments consist of RDF data files, which may or may not be binary files. The fragments also include metadata of the RDF data files (e.g., statistics, etc.) and optional delta files.

[0058] Therefore, the virtual RDF graph database is divided into two logical parts. That is, a first RDF graph database called a dynamic partition that can be updated by standard SPARQL updates, and a second RDF graph database that is a read-only snapshot and is called a batch partition that can be updated by loading batches of files. As described above, the second RDF graph database may consist of one or more snapshots.

[0059] · Update of the first RDF graph database The standard SPARQL update of the virtual RDF dataset is applied to the first RDF graph database (also called the dynamic partition) and can be modified in the same way as other graph databases. The same RDF graph can appear in both the batch partition and the dynamic partition. A particular triple can exist in only one partition, whether in the batch or dynamic case; otherwise, it violates the partitioning rules. In fact, a partition is defined as "dividing a logical database or its components into separate independent parts" (https: / / en.wikipedia.org / wiki / Partition_(database)). This restriction may be lifted at the cost of increased complexity, either inside the query engine or as post-processing of the results outside the database. In fact, when a triple exists in multiple partitions, the following two problems may occur.

[0060] For read queries: If some results are generated by triples existing in multiple partitions, they may be duplicated (and thus generated multiple times).

[0061] In the case of a write query: If it is necessary to delete triples existing in multiple partitions, the deletion needs to be propagated to all partitions. Otherwise, the triples will remain visible in subsequent queries. This limitation can be lifted without changing the present invention. This limitation can be implemented as known in the art.

[0062] · Update of the second RDF graph database The second read-only RDF graph database is stored in a file storage and can be updated by batches of tuples added to and / or deleted from the second RDF graph database. The second read-only RDF graph database forms an initial snapshot. Batches of tuples arrive gradually (i.e., with each new update of the initial snapshot) and form a new snapshot together with the initial snapshot. This new snapshot is defined as the state of the batch partitions of the second RDF graph database at a specific point in time. This state is defined by all batch partitions, i.e., all fragments (if one or more partitions are split into fragments), i.e., all RDF files including the metadata stored in the distributed storage. A snapshot consists of all files in the file storage and the location of their metadata and optional delta files. As explained, the delta files represent the changes or differences between two versions of the data.

[0063] · Catalog This method provides a catalog for storing metadata that describes the second read-only RDF graph database on the above file storage (S100). The catalog complies with ACID, and i) operations involving the addition, modification, or deletion of metadata entries are executed as indivisible transactions. If part of an operation fails, the entire operation is rolled back to maintain a consistent state. ii) The catalog system applies consistency rules (the meaning of consistent writing) to the metadata. This ensures the uniqueness of a particular point in time of the dataset snapshot defined as the "last snapshot". iii) Multiple transactions that modify metadata simultaneously do not interfere with each other. iv) Changes added to the catalog's metadata are persistent. Usually, the changes are written to non-volatile storage. The "last snapshot" means the aforementioned new snapshot.

[0064] These metadata are just key / value pairs and have a strong consistency model. Therefore, this catalog can register the metadata of the dataset in an ACID manner using any available language. Strong consistency is explained, for example, at https: / / en.wikipedia.org / wiki / Strong_consistency. The "consistency model" means the "consistency" in the consistency model of the CAP theorem at the following link: https: / / apple.github.io / foundationdb / consistency.html.

[0065] Regarding the consistency provided by the catalog, there may be multiple dataset snapshots stored at a specific point in time, but only one of them is called the "last snapshot". The last snapshot means a snapshot that includes all batch partitions received so far.

[0066] In one example, the catalog may be implemented by, but is not limited to, a distributed key / value store with ACID transactions and strong consistency like FoundationsDB (https: / / apple.github.io / foundationdb / index.html), or a NewSQL database like MariaDB in the open source world (https: / / mariadb.com / ) (MariaDB may be distributed with Xpand as presented at https: / / mariadb.com / products / enterprise / xpand / ), or any RDF database.

[0067] In one example, all dataset snapshot metadata may be stored in the catalog. Alternatively, all dataset snapshot metadata may be stored as files in distributed storage, and only the location of this file may be retained in the catalog.

[0068] Here, the operating principle of the catalog will be described. When a transaction is started, the catalog retrieves the last dataset snapshot metadata from the catalog. This last snapshot is used for all queries until the transaction ends.

[0069] When a partition is added, removed, or changed, a new partition (fragment if splitting is used) is added and / or a previous partition (fragment that may include a delta file if splitting is used) is removed. If all these operations are successfully executed in the distributed storage, a new dataset snapshot can be created. Interestingly, file storage guarantees consistent writes, thus ensuring the success of update operations.

[0070] When all these operations are successfully executed on the distributed storage, they are inseparably registered as the "last snapshot" in the catalog. Therefore, all new transactions use the new dataset snapshot, and previous transactions continue to use the previous dataset snapshot. This ensures the ACID properties due to the catalog and the ACID guarantees for the inseparability, durability, and strongly consistent writes of the distributed storage. Snapshots of old datasets and their corresponding unused partitions (or fragments and optionally delta files if partitioning is used) may be deleted from the distributed storage when all transactions using them have completed or asynchronously on the garbage collection pass.

[0071] In this way, only metadata is retained in the catalog, no data is retained, and data of the same scale as the dataset (i.e., billions of triples) is not retained. Therefore, the computing resources and memory costs for maintaining this catalog are much lower. The catalog may be shared among multiple datasets.

[0072] Referring to Figure 3, an example of the architecture scheme of the offering (S100) is shown. The architecture consists of two layers. The top layer is a database layer that manages RDF datasets. The layer below it is a storage layer that stores RDF datasets and batch partitions.

[0073] In the database layer, a first RDF graph database that can be updated by a stream is represented and can be queried by a SPARQL query engine. A catalog that stores metadata describing a second read-only RDF graph database on file storage can also be queried by the SPARQL query engine. For the sake of convenience of explanation, the SPARQL query engine does not necessarily depend on SPARQL queries to use the catalog, for example, when the catalog is key / value-based or an SQL database. The SPARQL query engine is assumed to recognize the interface / query language used to perform interactive operations on the catalog. As shown, the catalog stores information about the accessible snapshots and lists the available dataset information. The catalog also plays a role in maintaining the metadata information of the snapshots.

[0074] In the storage layer, snapshots of the dataset (e.g., the second read-only RDF graph database) are stored in a file storage that guarantees strongly consistent writes. For the last snapshot referenced in the catalog, at least a list of the data files of the snapshot is stored. Note that it should be understood that the file storage may store a list of multiple data files, and each list may correspond to a respective snapshot. Furthermore, the storage layer stores batch partitions representing the updates applied to the snapshots.

[0075] Batch partitions and dynamic partitions (a second read-only RDF graph database and a first RDF graph database stored in file storage respectively) constitute one RDF dataset (a virtual RDF graph database composed of at least two RDF graph databases recognized as one logically equivalent RDF graph database). The SPARQL query engine executes queries against all partitions and recognizes them as one logical database. Read queries are executed against all partitions with ACID guarantees. SPARQL update queries are executed only against dynamic partitions, and there is no restriction on the types of updates to be executed.

[0076] Referring further to the example of Figure 3, the dynamic partition is stored in an in-memory data structure, and the batch partition is stored in file storage that guarantees consistent writing regardless of whether the file storage is distributed. In this example, the SPARQL query engine executes queries on heterogeneous environment distributed storage.

[0077] When a SPARQL query is executed, the execution of the query mainly follows two constraints. That is, 1 / the RDF dataset is dynamic (data is added or deleted), and 2 / because the RDF dataset is dynamic, the SPARQL query is executed in an ACID manner especially regarding independence.

[0078] It should be understood that the batch partition may be placed on a standard file system such as the Unix file system instead of a distributed object storage.

[0079] Returning to Figure 1, in S110, a stream of tuples to be added and / or deleted to the virtual RDF graph database and a batch of tuples to be added and / or deleted to the virtual RDF graph database are obtained. The order in which these are obtained is not important. They may be received simultaneously.

[0080] Next, in S120, the stream of tuples is applied to the first RDF graph database of the virtual RDF graph database. This is performed as known in the art.

[0081] Next, in S130, a batch of tuples is applied to the second RDF graph database. In other words, as a result of the application, batch partitions are created. For this purpose, the following steps are performed.

[0082] First, a snapshot of the second RDF graph database is calculated. This calculation includes a batch of tuples to obtain a new snapshot. During the calculation operation leading to the generation of the new snapshot, the ACID properties are guaranteed by a series of concurrent control operations appropriately performed using consistent writing of the file storage where the second RDF graph database is stored.

[0083] By "concurrent control operations that guarantee the ACID properties" it means that for each transaction, the partition(s) of the dataset snapshot obtained at the start of the transaction are always kept available during the transaction and are not modified by concurrent addition / removal / modification. The concurrent control operations rely on the properties of the catalog with ACID transactions, and the file storage has read consistency after writing, i.e., "in subsequent read requests, the latest version of the object is immediately received". Also, the external storage can upload files and will surely fail if the file already exists. An implementation example of the algorithm that guarantees the "concurrent control operations that guarantee the ACID properties" will be described below.

[0084] Once calculated, the snapshot is stored in the file storage.

[0085] Finally, the snapshot is registered in the catalog, and the updated description of the virtual RDF graph database is stored in the catalog. Therefore, the updated description of the virtual RDF graph database includes the updated descriptions of the first and second RDF graph databases. This new snapshot can be used to apply a new batch of tuples.

[0086] In this step, a SPARQL query may be executed. Here, the SPARQL query may include an update query, and there is an ACID guarantee for RDF data split across heterogeneous environment distributed storage mainly composed of non-ACID compliant storage. The data of the update query may be taken in as a batch of data files from the distributed storage, or directly imported into the ACID database as a streaming modification of triples.

[0087] In one example, the present invention may be accessed through an API or through SPARQL commands as a standard extension. Although both are possible, for the purpose of clarity only, some commands will be described as extensions of SPARQL for operating on RDF datasets as described with reference to FIG. 1. It should be understood that the present invention is not limited to these commands. In particular, the actual grammar is for illustrative purposes only. FIG. 4 shows an overview of the commands described later.

[0088] In one example, the ALTER REFRESH command may execute a sequence of concurrent execution control operations. The ALTER REFRESH command guarantees the ACID properties for the update of the second RDF graph database of the virtual RDF graph database. The purpose of ALTER REFRESH is to create a new data set snapshot based on the current state of the data set and update the list of available data set snapshot definitions.

[0089] As the introduction of the ALTER REFRESH command, it should be noted that the ALTER REFRESH command is responsible for ensuring the ACID properties of dataset operations. Dataset operations may include, but are not limited to, adding partitions to the dataset by, for example, the ADD PARTITION command, deleting partitions of the dataset by, for example, the REMOVE PARITION command, and changing (i.e., modifying) partitions by, for example, the ALTER PARTITION command. This is summarized in Figure 4.

[0090] As already explained, a snapshot of a dataset can be defined as the state of the dataset at a specific point in time. This state is defined by all partitions, i.e., dynamic partitions and batch partitions. As mentioned above, batch partitions (read-only RDF graph databases stored in file storage) and dynamic partitions (the first RDF graph database that can be updated by a stream) are modified by separate commands. Since dynamic partitions are, by definition, modified in a streaming manner, their modifications become immediately visible. Batch partitions are, by definition, changed in batches. When a transaction is started on a dataset, the last snapshot is retrieved from the catalog. This snapshot is used for all queries until the transaction ends.

[0091] When the ALTER REFRESH command is executed, all partitions that make up a batch partition may be listed in the stage of the dataset. Note that the partitions may be fragmented for the purpose of improving the performance of SPARQL queries. In this case, all fragments that make up the batch partition are listed in the stage of the dataset. As is known, it is a common approach to process the dataset by stage. The stage of the dataset models where the files that make up the dataset are stored. For example, the definition of "stage" is provided at: https: / / docs.snowflake.com / en / sql-reference / sql / create-stage.

[0092] This list of partitions is saved in a so-called dataset snapshot, identified by a unique name. This dataset snapshot is saved on the catalog by ACID transactions and registered as the "last snapshot" or "last available snapshot". Thus, all new transactions use the new dataset snapshot loaded into the stage, and previous transactions continue to use the previous dataset snapshot. This ensures the ACID properties due to the catalog and the ACID guarantees for indivisibility, persistence, and consistent writing of distributed storage.

[0093] Each transaction has its own list of partitions (or in some cases, fragments), thanks to the dataset snapshot obtained at the start of the transaction. To enable ACID transactions on the dataset, these partitions (or in some cases, fragments) are available during the transaction and are not modified by concurrent addition / removal / modification of partitions or concurrent modified dataset updates. To achieve this, the ALTER REFRESH command relies on the following three functions: 1 / The catalog has ACID transactions; 2 / After creating / overwriting / deleting an object, it has read consistency after writing to external storage (i.e., "in subsequent read requests, the latest version of the object is immediately received"); 3 / It can upload files to file storage and will definitely fail if the file already exists.

[0094] In one example, the first control operation of the ALTER REFRESH command is to uniquely identify the last snapshot of the second RDF graph database. As already explained above, the last snapshot is the snapshot that includes all batch partitions received at the start of the execution of the ALTER REFRESH command. Therefore, the last snapshot represents the snapshot formed by the second RDF graph database and its batch partitions. It should be understood that the last snapshot may not include batch partitions. For example, at the start of the execution of the ALTER REFRESH command, no batch of tuples added and / or deleted to the virtual RDF graph database may be received. To uniquely identify means that the last dataset snapshot identifier is retrieved or obtained from the catalog.

[0095] The next control operation of the ALTER REFRESH command may include verifying that no other ALTER REFRESH commands are being executed simultaneously. If any are, the ALTER REFRESH command is stopped and the ALTER REFRESH commands already in progress are retained. This ensures ACID transactions at the last snapshot.

[0096] The next control operation of the ALTER REFRESH command may include obtaining a list of files in the stage of all snapshots. This operation is performed once and the entire list may be copied into the process's memory. This list is used for the remainder of the ALTER REFRESH command. Note that this operation of listing the files in the stage is strongly consistent due to the write consistency characteristics of the file storage. This means that the files in the listed stage are visible to all processes using this data and this data is reflected in any subsequent read operations on the last snapshot.

[0097] The following control operation of the ALTER REFRESH command may include filtering the files in the file list at the stage of the last snapshot. The filtering consists of checking that the files in the list are valid in the sense that no other operations are being performed. By filtering, the independence among the ACID properties is satisfied, thus ensuring that there are no concurrent operations interfering with the files in the list. In other words, the view on the last snapshot of the database is consistent and reliable. As a result of the filtering, a "snapshot list", which is a file indicating the names of the files of the last snapshot that will be available for subsequent access to the database at a specific point in time, is obtained. Therefore, the "snapshot list" includes partitions that have not been updated (i.e., modified) and updates that have been performed since the last ALTER REFRESH command; the updates performed since the last ALTER REFRESH command "snapshot list" are invisible to any subsequent actions on the database. Subsequent accesses include, but are not limited to, SPARQL queries, tuples to be added and / or deleted. Here, the above specific point in time is the time when the ALTER REFRESH command started execution.

[0098] The following control action of the ALTER REFRESH command may include writing the name of the file of the "snapshot list". Here, the term "name" implies that a dataset snapshot identified by a unique name is loaded into the stage.

[0099] Writing the name of the "Snapshot List" file means that a new snapshot is obtained from the previous snapshot, and all files of the previous snapshot become available for subsequent access to the database at the above-specified point in time. In other words, the names of all files in the "Snapshot List" are written to the files of the new dataset snapshot.

[0100] Any known technique may be used to perform the writing. For example, this can be done using actual semantics, such as the RDF format or Trig format of turtle files. Alternatively, any known technique for performing the writing, such as a simple list of file names, may be used.

[0101] The next control operation of the ALTER REFRESH command may include uploading the files of the new snapshot to the stage. The uploaded new snapshot includes the files specified in the "Snapshot List" file. This is performed in a manner known in the art.

[0102] The new snapshot is not yet registered. This means that since the new snapshot is unknown to the catalog, a SPARQL query (e.g., SELECT) cannot be executed on the new snapshot containing the latest data. The registration is performed as an ACID transaction. The ACID transaction on the catalog thus guarantees the integrity and reliability of the registration. The new snapshot is registered as the last snapshot of the catalog.

[0103] When the ALTER REFRESH command is executed, the batch of obtained tuples (S110) can be applied to the registered new snapshot (i.e., the last snapshot) to calculate the snapshot (S130) of the second RDF graph database containing the batch of tuples.

[0104] Here, a further example of the implementation of the ALTER REFRESH command will be described.

[0105] In one example, at the above specific point in time, the acquisition of the "snapshot list" of the file name of the last snapshot that becomes available for subsequent access to the database may be performed using the following sequence of concurrent control operations.

[0106] A list of the names of valid partition(s) may be obtained. The list may include one or more (but not necessarily all) valid partitions of the second RDF graph database. Here, "valid partition" means that the partition is correctly created in the ACID manner.

[0107] In the first example, obtaining a list of the names of valid partition(s) may include, for each partition, searching for a control file in the file storage that stores the name(s) of the file(s) of the valid partition. Although one file name may be stored, in practice, since a partition consists of several files, it is understood that the control file stores the names of the files that constitute the valid partition.

[0108] As is known, control files play an important role in managing the structure and metadata of the database. They function as a repository for important metadata regarding the database, such as the database name, file location, log file details, timestamp, and current structural integrity of the database.

[0109] Returning to the present invention, the control file may be stored in the file storage when the partition is batch-processed into the file storage, that is, when the partition is created.

[0110] In a first example, the control file may be named with a partition name and a specific extension called VALID_EXTENSION. This control file is called the VALID_EXTENSION file. The VALID_EXTENSION control file is intended to identify records of valid partitions. The VALID_EXTENSION file stores the names of the file(s) of the valid partition(s). By naming the file with the partition name and the specific extension, the VALID_EXTENSION file of the partition can be easily identified.

[0111] Instead of or in addition to the first example, in a second example, obtaining a list of names of valid partition(s) may include searching the file storage for a control file that stores the names of the files of the partition that is the target of the current deletion operation for each partition. Here too, one file name may be stored, but this is not a common situation. This control file may be stored in the file storage when a partition deletion operation is executed on the partition.

[0112] In a second example, the control file may be named with a partition name and a specific extension called TOMBSTONE_EXTENSION. The TOMBSTONE_EXTENSION control file is intended to identify records of the partition that is the target of the current deletion operation. The TOMBSTONE_EXTENSION file may store the names of the file(s) of the partition that is the target of the current deletion operation. The TOMBSTONE_EXTENSION file may be uploaded to the file storage when a partition deletion operation is started.

[0113] When the execution of the control operation to list the names of valid partition(s) is completed, further control operations may include checking, for each partition named in the list of valid partition(s), whether the current update operation on the partition is still pending. If the current update operation on a partition is still pending, that partition is added to the list of valid partition(s) for which the current update operation is still pending. By "current update operation" it is meant that the update operation has not been fully processed and / or is pending because it has not been applied to the entire system.

[0114] In one example, checking, for each partition, whether the current update operation on that partition is still pending may include searching, within the file storage, for a control file that identifies the record of the partition that is the subject of the current update operation for each partition. The control file stores the name of the file(s) of the valid partition(s) that are the subject of the current update operation. Here too, only one file name may be stored, but this is not the normal situation.

[0115] In one example, the control file that stores the name of the file(s) of the valid partition(s) that are the subject of the current update operation may be named with a specific extension called PARTITION_NAME and PENDING_EXTENSION. The PENDING_EXTENSION control file is intended to identify the record of the partition that is the subject of the current update operation. The PENDING_EXTENSION file stores the name of the file(s) of the partition that is the subject of the current update operation. The PENDING_EXTENSION file may be uploaded to the file storage when the update of the partition is started.

[0116] After confirmation (checking whether the current update operation on each partition is still pending) is complete, further control operations may include checking whether past update operations on the partition completed successfully. "Past update operations on the partition did not complete successfully" means all "pending" files that were used in the previous normal ALTER_REFRESH refresh but could not be cleaned up from the stage. This typically applies when ALTER_REFRESH crashed before the ALTER_REFRESH algorithm finished. If past update operations on the partition were not executed successfully, the control operation may further check whether the snapshot at which the failed update operation was executed is registered in the catalog. The snapshot may be a past snapshot or the last snapshot.

[0117] If the snapshot is not recorded (or registered) in the catalog, since the catalog should have registered the new snapshot as the last snapshot to which the batch of tuples was applied in an ACID transaction after identifying it, it means the previous alter refresh failed. As a result, since the previous ALTER_REFRESH was not executed successfully, the failed update operation is ignored and the "pending" files are no longer pending.

[0118] If the snapshot is recorded, it means that although the previous ALTER_REFRESH was executed successfully, the information that past update operations on the partition were not executed successfully should have been deleted but is being retained. As a result, the partition can be removed from the list of valid partitions (groups) for which the update operation (s) is pending.

[0119] In one example, checking whether past update operations on a partition were not executed properly may include searching, within a file storage, for a control file that identifies, for each partition, the record of the partition that was the subject of past update operations. The control file may store the name(s) of the file(s) of the valid partition(s) that were the subject of past update operations. Here too, one file name may be stored, but this is not the normal situation.

[0120] In one example, the control file that stores the name(s) of the file(s) of the valid partition(s) that were the subject of past update operations may be named with a partition name and a specific extension called CONSUMED_EXTENSION. The CONSUMED_EXTENSION file may be uploaded to the file storage at the start of past updates to the partition.

[0121] In one example, for each partition named in the list of the said valid partition(s), checking whether past update operations on the partition were not executed properly may include further control operations. In these further control operations, two control files are used. That is, i) a control file that stores the name(s) of the file(s) of the valid partition(s) that are the subject of the current update operation, and ii) a control file that stores the name(s) of the file(s) of the valid partition(s). Consistency between these two files is checked.

[0122] For the purpose of clarification only, this consistency check will be described with reference to an example where PENDING_EXTENSION files and VALID_EXTENSION files are implemented. The control operation may include checking whether the content of the PENDING_EXTENSION file matches the content of the VALID_EXTENSION file. The VALID_EXTENSION file stores the names of the file(s) of the valid partition(s), and the PENDING_EXTENSION file stores the names of the file(s) of the partition(s) that are the target of the current update operation. In the case of a valid partition, the names listed in the VALID_EXTENSION file are not found in the PENDING_EXTENSION file. This is because when the execution of the ALTER_REFRESH command is started, the names listed in the VALID_EXTENSION file need to be available for subsequent access to the database. Otherwise, the PENDING_EXTENSION file is ignored in the check. That is, only the VALID_EXTENSION file is considered.

[0123] After checking whether past update operations on the partition have been executed successfully, an empty "snapshot list" file may be created by further control operations of the ALTER_REFRESH command, and the "snapshot list" file is completed by adding the names of the file(s) of the valid partition(s) among the list of valid partitions for which the current update operation(s) are still pending. Otherwise, if the list of valid partitions for which the previous update operation(s) are still pending is empty, the names of the file(s) of the valid partition(s) among the list of names of the valid partition(s) are added. The obtained snapshot list indicates the names of the file(s) of the valid partition(s) that are available for subsequent access to the database at a specific point in time.

[0124] For the sole purpose of clarification, an example where the PENDING_EXTENSION file and the VALID_EXTENSION file are implemented here will be referred to for the completion of the "Snapshot List" file. After checking whether past update operations on the partition have been executed normally, an empty "Snapshot List" file may be created by further control operations of the ALTER_REFRESH command, and the "Snapshot List" file is completed as follows. For each partition, if there is a file named in the PENDING_EXTENSION file, the list of file names that make up this partition is given by the content of the PENDING_EXTENSION file and added to the "Snapshot List" file. Otherwise, the list of file names is given by the content of the VALID_EXTENSION file and added to the "Snapshot List" file. At a specific point in time, the "Snapshot List" file indicating the file name of the last snapshot that can be used for subsequent access to the database is completed in this way.

[0125] Subsequently, further control operations may include verifying that each file listed in the "Snapshot List" exists in the stage. This may be performed by comparing the "Snapshot List" with the list of files in the stage of the last snapshot.

[0126] In one example, uniquely identifying may include obtaining the last data set snapshot identifier from the catalog and retaining the last data set snapshot identifier in memory as the previous data set snapshot. Registering a new snapshot in the catalog may further include, by an ACID transaction catalog, verifying that the last snapshot described in the catalog is the previous data set snapshot (held in memory). Comparing two identifiers of a file is straightforward. If the last snapshot described in the catalog is the previous data set snapshot, the new snapshot is registered in the catalog as the last snapshot to which a batch of tuples was applied. In the reverse situation, an ALTER REFRESH command is executed concurrently and the new snapshot file is deleted from the stage.

[0127] In one example, the ALTER_REFRESH command may further include a control operation that provides a latch. As is known, a latch is a synchronization mechanism used to control access to shared resources to ensure consistency and avoid conflicts of concurrent operations.

[0128] In one example, the control operation may include latching a second read-only RDF graph database stored in file storage. In this case, the latch is executed before obtaining the list of files of the last snapshot at the stage. When a new snapshot (recognized as the last snapshot) is registered in the catalog, the new snapshot is released from the latch.

[0129] In one example, the latch may include using a control file that, by existing on file storage, verifies that the second read-only RDF graph database has not been modified during the execution of the ALTER_REFRESH command. After the last snapshot has been uniquely identified, the control file may be uploaded to the file. Subsequently, the control operation may verify that the control file still exists on the file storage at each subsequent step executed by the control operation. Finally, after the new snapshot has been registered in the catalog, the existence of the control file is last verified and the control file is deleted from the file storage.

[0130] Here, the latch will be further described. Since the present invention relies on an RDF heterogeneous environment distributed storage mainly composed of ACID non-compliant storage as illustrated in FIG. 3, an algorithm may be required to manage concurrency and ensure that the execution of the algorithm is always correct. A latch is a synchronization mechanism used to control access to shared resources to ensure consistency and avoid conflicts in concurrent operations. In one example, file storage may provide a similar mechanism, such as S3 object lock https: / / docs.aws.amazon.com / AmazonS3 / latest / userguide / object-lock.html, etc., but not all storage provides it. In one example, the concurrency control operation may execute a latch algorithm. Object lock does not depend on the capacity of the underlying infrastructure.

[0131] In one example, a dedicated START_EXTENSION control file may be used. As described in various examples, the existence of this file functions as a latch. If the process that created the START_EXTENSION file crashes before deleting this file from the stage, it is necessary to delete this file to avoid deadlocks. This is solved by the following example of the latch algorithm.

[0132] Add the host's IP address (or other means to identify the host) and the process ID to the START_EXTENSION control file; When the process attempts to create its own START_EXTENSION file (simply called the "start" file), if the process encounters the "start" file, read the IP address in the file: If it is one of another existing host within the cluster of the dataset, a concurrency problem is actually detected: that is, the process cannot create its own "start" file.

[0133] Otherwise, if the IP address is that of its own host, check the process ID: If it is not the same as its own, a process restart is occurring: it is safe to replace the "start" file with a new file; / / In the deployment of the application, it may be essential to ensure one process per host. Otherwise, this strategy may be adjusted as appropriate.

[0134] Otherwise, a concurrency problem is actually detected; - Otherwise, the IP address is unknown: the host that wrote the "start" file is down; / / It is safe to replace the "start" file with a new file.

[0135] There is still a case where the host shuts down and then restarts again (also known as "split brain" in a distributed system. See https: / / medium.com / nerd-for-tech / split-brain-in-distributed-systems-252b0d4d122e). To prevent this, before making modifications to the stage, it is necessary to check in all algorithms that the "start" file is still unique. This operation is summarized in the details of the algorithm by the sentence "Check that the'start' file is still unique".

[0136] If this check fails, split brain has occurred. Each algorithm should terminate and, if there are already uploaded files, clean them up. If this cleanup is not performed by the algorithm, it will be performed by garbage collection (see dedicated items). The algorithm is designed to guarantee the atomicity of the operation (i.e., if the algorithm fails before completion, the files uploaded in the stage are not considered).

[0137] In one example, the control file is a file named with the last dataset snapshot name and a specific extension called START_EXTENSION. Therefore, there is a unique control file for each snapshot stored in the file storage.

[0138] Some examples of the ALTER_REFRESH command were described. One or more of these examples may be combined. Here, an example of the ALTER_REFRESH command will be described in the form of a pseudo-code algorithm. The terms used in this pseudo-algorithm are the same as those used in the examples of the ALTER_REFRESH command. Interestingly, this pseudo-code algorithm shows an example where the partition is fragmented. Note that it should be understood that the fragmentation of the partition does not change the examples already described. This will become clear in the following example. Comments start with / / .

[0139] Enter the ALTER_REFRESH command: the dataset stored in the file storage

[0140] Obtain the last dataset snapshot of the dataset: Obtain the last dataset snapshot identifier from the catalog; Save this information as the "snapshot of the previous dataset"; If the snapshot of the dataset itself is a file in the stage (when there is only a redirect to that file in the catalog), obtain it from the stage. Otherwise, obtain all the information from the catalog.

[0141] Connect to the snapshot of the dataset using the cluster of the dataset; / / This means, for example, registering all the fragments as part of the dataset and optionally prefetching the metadata file. / / Note that the connection to the dataset snapshot is held in the cache and can be reused in subsequent calls. This cache can be cleared when "DISCONNECT" is executed and / or by other cache replacement policies such as LRU. / / Note that if a snapshot of the dataset is held in the cache, the use of the garbage collection algorithm executed by the GARBAGE_COLLECT command is restricted; an example of the garbage collection algorithm is shown later; the trade-off between the cache and garbage collection can be selected without changing the present invention.

[0142] Upload a latch control file to storage: This is named with a specific extension called the dataset snapshot name and START_EXTENSION, for example, ".start". If this call is successful, the algorithm is sure to recognize that such a file did not exist previously, and thus, it means that other "ALTER REFRESH" or "GARBAGE_COLLECT" is not being executed simultaneously for this dataset. If this call fails, "alter refresh" or "garbage collect" being executed simultaneously for this dataset is detected: End with this error.

[0143] Generate a new dataset snapshot name: For example, "dataset_snapshot_ <uuid>」, etc. When the turtle format is selected, it may have a dedicated extension (such as.ttl).

[0144] List all files in the stage / / This operation is performed once, and the entire list is copied into the process memory. This copy is used in the rest of the algorithm. / / The file listing operation has strong consistency due to the characteristics of the object storage.

[0145] Filter files from the list of all files in the stage: First, list all valid partitions by listing all files ending with VALID_EXTENSION. The partition name is given by the name of the VALID_EXTENSION file without the extension.

[0146] For each partition name, if there is a file with the extension TOMBSTONE_EXTENSION appended to the partition name, exclude it from the list of valid partitions. / / This way, all valid partition names are obtained.

[0147] For each valid partition name, check if there is a file that starts with the partition name and ends with the extension PENDING_EXTENSION; If so, save the list of all PENDIND_EXTENSION files found here in memory; / / These files correspond to the "ALTER_PARTITION" commands that have not yet been considered. See the dedicated algorithm described later; For each file with the CONSUMED_EXTENSION extension / / Note that these files only exist for the ALTER_REFRESH commands that crashed before the end of the algorithm; normally, these files should not exist:

[0148] Get the dataset snapshot name from the file name;

[0149] Check in the catalog that this dataset snapshot is registered: If not registered: Ignore this CONSUMED_EXTENSION file / / Can be deleted here, or deleted as garbage collection. If registered: Read the content of the CONSUMED_EXTENSION file / / The CONSUMED_EXTENSION file shows all the "pending" files used in this dataset snapshot. Check that the content of the PENDING_EXTENSION file matches the content of the VALID_EXTENSION file. If not, consider that the VALID_EXTENSION file has not been updated with the content of the previous extension's "dataset snapshot", and consider the content of the "dataset snapshot" as the content of the VALID_EXTENSION file Delete these files from the list of all PENDIND_EXTENSION files found in the previous step / / These were used in a previous successful alter refresh but could not clean up the "pending" files from the stage; as its content.

[0150] Create an empty list called the "snapshot list" file that lists the file names of the dataset snapshot;

[0151] For each valid partition name, add the following files to the "snapshot list" file: If there is a file with the extension PENDING_EXTENSION in the list of files, the list of files that make up this partition is given by the content of this file and added to the "snapshot list" file. Otherwise, The list of files is provided by the content of the files having the partition name and the VALID_EXTENSION extension, and is added to the "Snapshot List" file.

[0152] Check that all files in the "Snapshot List" exist in the list file of the stage; / / The list file of the stage is the list copied to memory at the start of the algorithm. Here, no actual check for the stage is performed.

[0153] Check that for all fragments, both the RDF file and the corresponding metadata file exist; / / Note that there may be a difference file (its existence is checked in a previous step before checking all files in the list); / / Note that listing files using the VALID_EXTENSION and PENDING_EXTENSION files gives the algorithm independence with respect to simultaneous modification of partitions.

[0154] Write the names of all files in the "Snapshot List" to the file of the new dataset snapshot / / Any format can be used for writing these names, for example, execute with the actual semantics using the RDF format of the turtle file, or simply use a list if no further semantics are required.

[0155] Check that the "start" file is still unique and upload the dataset snapshot file in the stage. / / Note that to maintain indivisibility, this dataset snapshot has not yet been registered in the catalog and is thus invisible to subsequent transactions.

[0156] Check that the "start" file is still unique, register the new dataset snapshot as the last dataset snapshot in the catalog, and provide retention information for the "previous dataset snapshot":

[0157] In an ACID transaction by the catalog, check that the "last dataset" it recognizes is the "previous dataset" given as the input to the check; If so, fill in the following two pieces of information: Register the last dataset snapshot in the catalog; Mark all "pending" files as used: Write the names of all files with the PENDING_EXTENSION extension to a file named with the dataset snapshot file name and a dedicated CONSUMED_EXTENSION, e.g., "<dataset_snapshot_name>.consumed". Check that the "start" file is still unique and upload the CONSUMED_EXTENSION file. / / Note that this CONSUMED_EXTENSION file was used at the start of the algorithm; Register this new last dataset snapshot in the catalog and link it to the "previous dataset" to maintain the system / / The system is optional; / / In this call, also check whether the "previous dataset" is correct to avoid race conditions (or it may have been locked by the catalog when previously checked); / / Registration in the catalog means registering the dataset snapshot file name. Alternatively, all information held in the file rather than the file itself may be registered in the catalog.

[0158] / / Even if the subsequent process crashes later, since the "CONSUMED_EXTENSION" file has been created, as seen at the beginning of the algorithm, the operations on hold will not be reapplied in the subsequent "ALTER_REFRESH";

[0159] For further alter refresh operations, replace the content of the VALID_EXTENSION file with the list of files included in the "dataset snapshot"; / / Note that to replace the content of the file atomically, instead of overwriting the existing file, a new file with the necessary content needs to be created.

[0160] For all PENDING_EXTENSION files, delete them from the stage; Delete the file with the CONSUMED_EXTENSION extension; Otherwise: It is known that the "last dataset" is not the "previous dataset" given as input to the check. / / At the same time, the "ALTER_REFRESH" command is being executed.

[0161] Roll back the modifications made in the stage: Delete the dataset snapshot file from the stage. Delete the file with the CONSUMED_EXTENSION extension; Return an error: At the same time, the "ALTER_REFRESH" command is being executed;

[0162] Check and delete if the "start" file is still unique; / / For all "exit" statements of this algorithm and if an exception occurred in the previous step, this file needs to be deleted.

[0163] In one example, the ALTER PARTITION command may execute a sequence of concurrent control operations. The ALTER PARTITION command is executed after the ALTER REFRESH command that guarantees the ACID properties of the update of the second RDF graph database of the virtual RDF graph database. The purpose of ALTER PARTITION is to update the partitions of the second read-only RDF graph database.

[0164] Here, examples of the ALTER PARTITION command will be described. These examples are shown with control files having a specific naming and extension scheme, but it is understood that the control files are not limited to these examples.

[0165] The control operation of the ALTER PARTITION command may include downloading the VALID_EXTENSION file of the valid partition to be updated from the stage. Thereby, the name(s) of the file(s) of the valid partition is / are obtained. The valid VALID_EXTENSION file was created at the initial creation of the second read-only RDF graph database, or after a successful update operation in the second read-only RDF graph database, or has been updated during the execution of the ALTER REFRESH command already described above.

[0166] Next, the control operation of the ALTER PARTITION command may verify that the ALTER PARTITION command is not being executed concurrently on the valid partition. Otherwise, the ALTER PARTITION command is stopped and the ALTER PARTITION command already in execution is held.

[0167] Next, a PENDING_EXTENSION file may be created. The PENDING_EXTENSION file stores the name(s) of the file(s) of the partition that is the target of the current update operation. The name(s) of the file(s) is / are obtained from a batch of tuples.

[0168] The PENDING_EXTENSION file may be uploaded to the stage when created, thereby making it accessible by subsequent commands (if any).

[0169] Here, an example of the ALTER PARTITION command in which one or more batch partitions of the second read-only RDF graph database are split into fragments will be described. All batch partitions may be split. The control operation of the ALTER PARTITION command may include determining whether the read-only second RDF dataset is a Change Data Capture (CDC) file.

[0170] As is known, CDC (abbreviation for Change Data Capture) discriminates and tracks data changed using a differential file, enabling actions to be executed using the changed data. For discussions on CDC, refer to https: / / en.wikipedia.org / wiki / Change_data_capture. Therefore, CDC for SPARQL updates may be considered as a limitation of SPARQL 1.1 updates that enables changes to be expressed only as extended definitions. Figure 5 shows an example of how triple modifications can be expressed. It is also possible to use any other expression without changing the present invention. For example, information regarding which triple(s) in which graph(s) are modified may be provided.

[0171] If the batch of tuples is not a CDC file, fragments of the batch of tuples are generated, thereby obtaining a list of fragment files. The generation of fragments will be described with reference to the ADD_PARTITION command.

[0172] When the batch of tuples is a CDC file, the following operations may be performed for each add / delete triple of the CDC file (of the batch of tuples). One or more fragments of one of the batch partitions are placed. The fragment where the snapshot is placed is the fragment to which the add / delete operation is applied. As a result of the placement, a list of snapshot fragments is obtained, and if one or more delta files are associated with one or more of the fragments, an additional list of delta files is also obtained. Then, by the control operation of the ALTER PARTITION command, a PENDING_EXTENSION file is created that stores the list of fragments and the additional list of delta files (if any). The PENDING_EXTENSION file is uploaded to the stage when it is created (or obtained), thereby making it accessible by subsequent commands (if any), as already explained.

[0173] Partitions can be changed using the ALTER PARTITION command. Note that in the case of the first RDF graph database (also called the dynamic partition), since the modification is performed streaming, only SPARQL UPDATE queries are used.

[0174] In the case of a read-only second RDF graph database, as in the case of partition creation, a batch file is used as input for updates. There are several options for the input to the ALTER PARTITION command. The first method may be to provide the entire file used for partition creation again, which may be modified as necessary (i.e., delete deprecated triples and add new triples). In this way, the fragments are regenerated (either in whole or incrementally). In this solution, the input may be a file in a standard format such as TriG. As another solution, only the necessary modifications (triples to delete, triples to add) may be provided as input. Since there is no standard format to formalize this, it is defined here. As described with reference to FIG. 5, CDC may be used for this purpose. One of the two solutions to use may be determined by detecting the input file format (batch of tuples). For example, in the case of a complete regeneration: ALTER PARTITION[Name]INTO DATASET[dataset Name]{CONTENT=[file.trig.gz],DESCRIPTION="description”} In the case of incremental input: ALTER PARTITION[Name]INTO DATASET[dataset Name]{CONTENT=[file.cdc.gz],DESCRIPTION="description”} Examples of the ALTER PARTITION command were described. One or more of these examples may be combined. Here, an example of the ALTER PARTITION command is described in the form of a pseudo-code algorithm. The terms used in this pseudo-algorithm are the same as those used in the examples of the ALTER PARTITION command. Interestingly, this pseudo-code algorithm shows an example where the partition is fragmented. Note that it should be understood that the fragmentation of the partition does not change the examples described. This will become clear in the following example. Comments start with / / .

[0175] · Input to the ALTER PARTITION command An input RDF file containing a batch of tuples to be added and / or deleted; / / The input RDF file can be a standard RDF file such as TriG, or a CDC file; The last snapshot of a read-only second RDF graph database to be modified; Partition name.

[0176] Check whether the input file exists.

[0177] Upload to storage a START_EXTENSION control file for latch control. This is named, for example, with a specific extension called the partition name and START_EXTENSION.

[0178] If this call is successful, the algorithm is sure to recognize that such a file did not exist previously, and thus that no other "partition add / delete / change" is being executed simultaneously for this partition name; If this call fails, "partition add / delete / change" being executed simultaneously is detected with this partition name: end with this error;

[0179] Download the VALID_EXTENSION control file from the stage to control the validity of the partition. This is named, for example, with a specific extension called the partition name and VALID_EXTENSION.

[0180] If such a file does not exist, the partition of this name does not exist and ends with this error; If it exists, the next step of the algorithm is executed; Check whether the PENDING_EXTENSION control file exists in the stage. This is named, for example, with a specific extension called the partition name and PENDING_EXTENSION, for example ".pending". If it exists, there is already a pending partition change operation for this partition; end with this error; If it does not exist, the next step of the algorithm is executed;

[0181] Check whether the TOMBSTONE_EXTENSION control file exists in the stage. This is named, for example, with a specific extension called the partition name and TOMBSTONE_EXTENSION. If it exists, this partition has already been deleted; end with this error; If it does not exist, the next step of the algorithm is executed;

[0182] Check whether the input file is a CDC file; If the input file is not a CDC file: Generate fragments of the input RDF file. Examples of generation are described in the "ADD PARTITION" command.

[0183] Get the list of fragment files of the input RDF file; If the input file is a CDC file: For each add / delete triple in the CDC file: Place the fragment file of the last snapshot. Add / delete operations need to be applied to this; / / Generate a new version of the fragment file of the last snapshot (e.g., if the file can be modified in place), or add one delta file with a specific extension (e.g., ".delta") with the same name. / / The delta file may be a text file, in which case it is the same as the input CDC, or it may be a binary file, in which case the parsing of the text file is avoided. / / The delta file may be, for example, one HDT file for all triples to be added and one HDT file for all triples to be deleted, but other strategies may be used without changing the present invention. All that is needed is one delta file to read in addition to the existing files in order to be able to respond to the BGP of the SPARQL query;

[0184] Obtain a list of fragments that includes the add list of the delta file; Create a PENDING_EXTENSION control file. This is named, for example, with a partition name and a specific extension called PENDING_EXTENSION; Add to this new PENDING_EXTENSION control file the list of fragments that may include the delta file added in the previous step; / / Not only the delta of the modification, but also all files that make up the partition; Check that the "start" file is still its own and upload this file to the stage; Check if the "start" file is still unique and delete it from storage. / / For all "exit" statements of this ALTER PARTITION command algorithm, and if an exception occurred in the previous step, this file needs to be deleted.

[0185] In one example, the ADD PARTITION command may execute a sequence of concurrent control operations to be performed. The ADD PARTITION command may be executed before the ALTER REFRESH command. The ALTER REFRESH command guarantees the ACID properties of the update of the second RDF graph database of the virtual RDF graph database. The purpose of ADD PARTITION is to add a partition to the second read-only RDF graph database. For example, a typical sequence of commands may be the creation of the database, the execution of the ADD PARTITION command, and then the execution of the ALTER REFRESH command that visualizes the new partition and enables further processing (e.g., SPARQL query).

[0186] Note that it should be understood that new partitions may be added to a given dataset. How this is done differs between the case of dynamic partitions and the case of batch partitions. A method of adding a new partition to the first RDF graph database will be briefly described.

[0187] When a virtual RDF graph database is created, a dynamic RDF graph database may be created by default. The created first RDF graph database is the database that is modified by SPARQL UPDATE queries. The default partition has no name. For example, the following command may be used to give the partition a name and a description and indicate that it is a dynamic partition: ADD PARTITION[Partition Name]INTO DATASET[dataset Name]{CONTENT=‘DYNAMIC’,DESCRIPTION="description”}

[0188] Having default dynamic partitions is a matter of choice to facilitate the use of the present invention: this avoids having queries that do not recognize the dataset based on an external location. It may be possible to enforce naming all partitions.

[0189] Here, a method for adding a new partition to a read-only second RDF graph database will be briefly described.

[0190] In the case of batch partitioning, the input for creation is a batch of tuples. The data source for this batch can be an RDF file (e.g., TriG, HDT (described later), or any other format, which may or may not be compressed, e.g., with gzip), or a memory format such as, for example, dynamic partitioning (described in detail later). The input batch may then be split into one or more fragments. If the input batch is too large, there may be several fragments: in fact, large files can be costly in terms of acquisition via the network, acquisition from object storage (cost per GB acquired), and movement to the response to BGP; BGP will be described later. Experiments have shown that it may be a reasonable compromise to set the maximum size per fragment RDF file to 10 million triples. When the input batch is split into one or more RDF files, fragment metadata may be generated. Note that the final format selection for the fragment RDF files is independent of the input format. The selected format may be the one that optimizes usage the most. For example, to reduce size, a binary format rather than a human-readable format may be selected. Once all of these are completed, the fragments (i.e., the RDF files and their metadata) are uploaded to a stage associated with a read-only RDF graph database snapshot. The data of the RDF files and their metadata (fragments) will only become visible after the ALTER REFRESH command has been executed successfully. Partitions may be named. In this case, the name needs to be unique and can be an IRI or a literal value. Optionally, an explanation may be added. The commands are, for example, as follows: ADD PARTITION[Partition Name]INTO DATASET[dataset Name]{CONTENT=[file.ttl.gz],DESCRIPTION="description”} Here, an example of the ADD PARTITION command will be described. This example is shown with control files having a specific naming and extension scheme, but it is understood that these examples are not limited to these control file examples. The control operation of the ADD PARTITION command may include uploading the VALID_EXTENSION control file of the partition to be added to storage. The fact that the uploaded VALID_EXTENSION exists in storage indicates that there is no partition with the same name in the file storage yet.

[0191] Next, the control operation of the ADD PARTITION command may generate a fragment of a batch of tuples for adding a partition to a second read-only RDF graph database. The generation may be performed as described in the example of the ALTER PARTITION command. Then, a list of fragment files is obtained.

[0192] Next, by the control operation of the ADD PARTITION command, the list of fragments may be saved to the VALID_EXTENSION control file, and the VALID_EXTENSION control file may be uploaded to the stage. Note that it should be understood that the VALID_EXTENSION control file may be used for subsequent commands such as the ALTER REFRSH command and the ALTER PARTITION command.

[0193] Here, an example of the ADD PARTITION command will be described in the form of a pseudo-code algorithm. The terms used in this pseudo-algorithm are the same as those used in the example of the ALTER PARTITION command. Comments in the pseudo-code start with / / .

[0194] · Input of the ADD PARTITION command Input RDF file containing a batch of tuples of partitions to be added; / / The input RDF file can be a standard RDF file such as TriG or a CDC file; Dataset, e.g., the last snapshot of a read-only second RDF graph database that is changed by adding partitions; Name of the partition to be added; Characters reserved as delimiters, e.g., "_", which is called DELIMITER.

[0195] · Start of the ADD PARTITION command Check whether the input RDF file name exists; Check that the partition name does not already contain the DELIMITER character;

[0196] Upload to storage the START_EXTENSION control file for latch control. This is named, for example, with the partition name and a specific extension called START_EXTENSION, e.g., ".start", and is called the "start" file; If this call is successful, the algorithm is sure to recognize that such a file did not exist before, and thus, for this partition name, no other "partition add / delete / change" is being executed simultaneously; If this call fails, a "partition add / delete / change" being executed simultaneously is detected with this partition name: end with this error;

[0197] Upload to the stage the VALID_EXTENSION control file for controlling the validity of the partition. This is named, for example, with the partition name and a specific extension called VALID_EXTENSION, e.g., ".valid".

[0198] Check that the "start" file has not been uploaded yet.

[0199] If this call is successful, the algorithm is sure to recognize that such a file did not exist before, and thus, there will be no partition with this name; If this call fails, a partition with the entered partition name already exists: state that it is necessary to execute "alter partition" instead and end with this error;

[0200] Generate fragments of the input RDF file. Parse (read in the case of binary) the input RDF file and split it into small files. The small files may have a maximum number (e.g., 10 million triples) of triples of SPLIT_SLICE, and the scolemization of blank nodes may be retained (scolemization will be described later); / / The small files each become part of a fragment, and the fragment is composed of the small file and the associated metadata.

[0201] For each of these small files, generate an RDF file in the selected format. This format may be selected to optimize the use case, such as a binary RDF format like HDT; If this option is selected, generate a metadata file for the small files; / / Fragments are identified by pairs of RDF files and metadata files of small files; Check that the "start" file is still unique and upload all fragments to the stage;

[0202] For integrity checking, the partition name may be set in the fragment name. For example, the names of all fragments can be set as follows: "fragment_ <uuid> _ <partitionname> . <extension>". Here, " <uuid>" is the generated unique identifier, <partitionname>is the input partition name, <extension>is the file extension (different for RDF files, e.g., ".hdt", ".trig.gz", etc., and for metadata files, e.g., ".metadata", etc.); If one of the uploads fails: End with an error; The uploaded files will not be visible. These may be deleted here or in a later garbage collection pass;

[0203] Create a new VALID_EXTENSION control file. This is called the "validity file" and is named, for example, with a partition name and a specific extension called VALID_EXTENSION; Write a list of all fragment files to this "validity file"; Check that the "start" file is still unique and upload the "validity file" to the stage; / / All fragment files uploaded before this "validity file" are not considered in the "ALTER REFRESH" command. Therefore, this "validity file" guarantees the indivisibility of the "ADD PARTITION" command.

[0204] Check that the "start" file is still unique and delete it. / / For all "return" statements of this ADD PARTITION algorithm and if an exception occurred in the previous step, it is necessary to delete the START_EXTENSION control file.

[0205] Here, continue to refer to the ADD partition command to explain partitioning. There are multiple strategies for partitioning. Partitioning is provided to the command, which means that the command is not involved in partitioning. The present disclosure is independent of the strategies used for partitioning.

[0206] In one example, the REMOVE PARTITION command may execute a sequence of concurrent control operations when executed. The REMOVE PARTITION command may be executed before the ALTER REFRESH command that guarantees the ACID properties of deletion in the second RDF graph database of the virtual RDF graph database. The purpose of REMOVE PARTITION is to delete the partitions of the second read-only RDF graph database. For example, a typical command sequence may be the creation of the database, the execution of the REMOVE PARTITION command, and the execution of the ALTER REFRESH command that makes the deleted partitions invisible to enable further processing (e.g., SPARQL queries).

[0207] There may also be duplicate partitions (i.e., triples are replicated across multiple partitions). In that case, duplicate management is performed for both write queries (e.g., triples need to be deleted from all partitions where they exist) and read queries (e.g., being aware of the possibility of duplicate query results). Whether there are duplicate partitions is left to the choice of the software that determines the partitions.

[0208] Deleting a partition is straightforward. Make all data related to the partition invisible, although the data may still be visible to past transactions. This means all data of a batch partition, or fragments of a fragmented batch partition, or the in-memory data structure of a dynamic partition. The command may be something like: REMOVE PARTITION[NAME]FROM DATASET[Dataset Name] Since files are not deleted from storage by the partition deletion command, this change will only become visible after the GARBAGE_COLLECT command has succeeded. An example of the GARBAGE_COLLECT command will be described later.

[0209] In one example, the control operation of the REMOVE PARTITION command may include, on the stage, checking whether the partition to be deleted exists, and if it does not exist, stopping the REMOVE PARTITION command.

[0210] Next, the control operation of the REMOVE PARTITION command may upload the TOMBSTONE_EXTENSION control file to the stage. The TOMBSTONE_EXTENSION control file is intended to identify the records of the partition that is the target of the current deletion operation. When the upload is successful, it is confirmed that the second read-only RDF graph database has the partition to be deleted, that is, the records of the partition have not yet been deleted. The TOMBSTONE_EXTENSION control file may store all the records of the partition to be deleted. This is because all these records need to be deleted. Here, the TOMBSTONE_EXTENSION control file is uploaded to the stage: the partition is recorded as deleted.

[0211] Here, an example of the DELETE PARTITION command will be described in the form of a pseudo-code algorithm. The terms used in this pseudo-algorithm are the same as those used in the example of the ALTER PARTITION command. Comments in the pseudo-code start with / / .

[0212] · Input of the DELETE PARTITION command The last snapshot of a read-only second RDF graph database that is modified by deleting a dataset, e.g., a partition; The name of the partition to be deleted;

[0213] Upload a START_EXTENSION control file for latch control to storage. This is named, for example, with the partition name and a specific extension called START_EXTENSION, e.g., ".start", and is called the "start" file; If this call is successful, the algorithm ensures that such a file did not exist before, and thus, for this partition name, no other "partition add / delete / change" is being executed simultaneously; If this call fails, a "partition add / delete / change" being executed simultaneously is detected with this partition name: end with this error;

[0214] At the stage, check the VALID_EXTENSION control file for controlling the validity of the partition. This is named, for example, with the partition name and a specific extension called VALID_EXTENSION, e.g., ".valid".

[0215] If it does not exist, the partition of this name does not exist, and end with this error; If it exists, the next step of the algorithm is executed;

[0216] Upload a TOMBSTONE_EXTENSION control file to storage. This is named, for example, with the partition name and a specific extension called TOMBSTONE_EXTENSION, e.g., ".removed".

[0217] This upload is executed if it is checked that the "start" file is still unique; If this call is successful, the algorithm is certain to recognize that such a file did not exist previously, and thus, this partition has not yet been deleted; If this call fails, the partition with the entered partition name has already been deleted: end with this error;

[0218] Check that the "start" file is still unique and delete it. / / For all "return" statements of this ADD PARTITION algorithm, and if an exception occurred in the previous step, it is necessary to delete the START_EXTENSION control file.

[0219] In one example, to retain the identification of blank nodes, the input RDF file may skolemize the blank nodes and identify them as uniquely defined by identifiers. This is described in Tomaszuk, D. and Hyland-Wood, D., (2020) "RDF1.1: Knowledge representation and data integration language for the Web" Symmetry, 12(1), p. 84. The ADD / REMOVE / ALTER PARTITION commands can have a "SKOLEMIZED" option that warns to retain the identification information of blank nodes.

[0220] In one example, a dataset snapshot of a second RDF graph database may be retained in the cache. In addition to or as an alternative to this, old dataset snapshots and their corresponding unused fragments or partitions may be deleted from the distributed storage when all transactions using them have ended. In these examples, the GARBAGE_COLLECT command may be used to delete unused data of the stage. As an example, the data used may be as follows: · Data record; · Fragments, i.e., data files: RDF files (e.g., "trig.gz", ".hdt", etc.; different depending on the selection), corresponding metadata (e.g., ".metadata" (if this option is selected)), delta files (e.g., ".delta"). · Control files indicating existing partitions (e.g., VALID_EXTENSION control file and TOMBSTONE_EXTENSION control file). · Control files indicating changes to partitions, e.g., PENDING_EXTENSION control file and CONSUMED_EXTENSION control file.

[0221] Here, an example of the GARBAGE_COLLECT command algorithm will be described. Comments start with / / .

[0222] · Input of the GARBAGE_COLLECT command Dataset, e.g., the last snapshot of a read-only second RDF graph database;

[0223] Upload the START_EXTENSION control file to storage. This is named, for example, with a specific extension called the dataset name and START_EXTENSION, e.g., ".start"; If this call is successful, the algorithm ensures that such a file did not exist previously, and thus, no other "ALTER REFRESH" or "GARBAGE_COLLECT" commands are being executed simultaneously for this dataset; If this call fails, "alter refresh" or "garbage collect" being executed simultaneously for this dataset is detected: end with this error.

[0224] List all dataset snapshots from the catalog and obtain the "last dataset snapshot"; / / All dataset snapshots except the "last dataset snapshot" are candidates for deletion.

[0225] Request that all dataset snapshots currently in use in the connection be obtained for the dataset cluster. Remove these from the list of candidates for the "dataset snapshots to be deleted"; / / This request to the catalog can be executed in several ways. For example, the catalog may be notified each time a connection is created with a specific dataset snapshot and each time such a connection is closed. / / Note that this is affected by the connection cache;

[0226] Delete all "dataset snapshots to be deleted" from the catalog; New connections to the dataset cannot be started using these "dataset snapshots to be deleted"; / / This does not affect existing connections to the dataset using these "dataset snapshots to be deleted";

[0227] For dataset snapshots that are candidates for deletion: Delete the dataset snapshot files from the staging area; Delete all CONSUMED_EXTENSION files starting with these dataset snapshot names from the staging area; List dataset snapshots that are not candidates for deletion and obtain the "list of valid dataset snapshots";

[0228] List all files in the staging area. / / This operation is executed once and the entire list is copied into the process memory. The algorithm uses this copy for the rest of the algorithm. The file listing operation has strong consistency due to the consistent write characteristics of the object storage.

[0229] For all files in this list, obtain all files ending with the TOMBSTONE_EXTENSION extension. / / The corresponding partition names are specified by the file names without the extension. These are the partition names of the deleted partitions.

[0230] If there is a file with the same partition name and the VALID_EXTENSION extension, and optionally, if there is a file with the same partition name and the PENDING_EXTENSION extension: Reading the content of this file or these files will obtain a list of files to be deleted from the stage, called the "delete target list";

[0231] Execute the cleanup operation; Check that the starting file is still unique and delete the VALID_EXTENSION file from the stage (if it exists); Delete the PENDING_EXTENSION file from the stage; Delete all files in the "delete target list" from the stage; / / This will delete the data files, i.e., RDF files, metadata files, delta files; Delete the TOMBSTONE_EXTENSION file from the stage;

[0232] Even if any inconsistency is found, check the files to be deleted: TOMBSTONE_EXTENSION files or PENDING_EXTENSION files without the corresponding VALID_EXTENSION: Delete from the stage; CONSUMED_EXTENSION file without corresponding dataset snapshot file: Delete from the stage; PENDING_EXTENSION file or TOMBSTONE_EXTENSION file without VALID_EXTENSION: List all files given by the content of the PENDING_EXTENSION file and delete them from the stage if they exist; Delete the PENDING_EXTENSION file from the stage; If there is a file with the same partition name and TOMBSTONE_EXTENSION extension, delete it from the stage;

[0233] Execute the optional "full pass"; / / "full pass" is a more thorough cleanup pass: / / In the previous step, the algorithm obtained a list of valid dataset snapshots; Create an empty list called "list of used data files": / / The "list of valid dataset snapshots" stores dataset snapshots that are not candidates for deletion, i.e., valid dataset snapshots.

[0234] For each of these valid dataset snapshot files: Read the content of the valid dataset snapshot file and obtain a list of data files for the dataset snapshot / / RDF files, metadata files (if any), delta files (if any); Add this list to the "list of used data files".

[0235] Create an empty list called "list of garbage files"; Read the list of all files of the stage created at the start of the GARBAGE_COLLECT command algorithm. For each data file in it (i.e., RDF file, metadata file, delta file): If the file is not in the "list of used data files", add it to the "list of garbage files"; Delete all files in the "list of garbage files" from the stage;

[0236] Check that the "start" file is still unique and delete it; / / For all "exit" statements of this algorithm and if an exception occurred in the previous step, this file needs to be deleted.

[0237] In one example, the CREATE DATASET command may be used to create a new dataset. This command takes as input the location of the stage of the dataset to be created. An example of a command to create a dataset is shown below: CREATE DATASET <name>{ STAGE = <stage_name> } As a result of executing the CREATE DATASET command, the catalog registers the association between the logical name of the dataset and the location of the files modeled by the stage.

[0238] In one example, you may execute the CREATE STAGE command to create a new stage. As already defined, the stage of a dataset models the location where the files that make up the dataset are stored. For example, "stage" is defined here: https: / / docs.snowflake.com / en / sql-reference / sql / create-stage. Note that the purpose of the present invention is not to describe a location within a distributed object storage. In other words, the present invention may rely on any known method for describing such a location. In one example, the files may be stored outside the database, for example, in an external cloud storage. The storage location may be private or public. In one example, to avoid information duplication at each stage, storage integration may be used to delegate the authentication responsibility. The document https: / / docs.snowflake.com / en / sql-reference / sql / create-storage-integration describes an example of storage integration, and it is understood that any known method of storage integration can be used in the present disclosure. A single storage integration may support multiple external stages.

[0239] All of this information (file location, internal or external storage, authentication, etc.) is metadata that describes the dataset (stored in the metadata). These metadata are stored directly in the catalog and the metadata resembles key / value pairs. When a storage integration is created, only general information is provided; this information is specific to the cloud storage service. In one example, the following CREATE STORAGE INTEGRATION command may be used to create a dataset: CREATE STORAGE INTEGRATION <name>{ TYPE=EXTERNAL_STAGE, STORAGE_PROVIDER='...', URL='storage_url', CREDENTIALS_KEY_ID='keyid', CREDENTIALS__KEY_VALUE='keyvalue' }

[0240] And to create a stage using this storage integration, the following CREATE STAGE command may be used: CREATE STAGE <external_stage_name>{ STORAGE_INTEGRATION=<storage_integration_name>, TENANT='tenant', CLUSTER_ID='clusterid', BUCKET='bucket', PATH=' / <path> / ' }

[0241] Here, the concept of a cluster will be explained. The concept of a cluster is introduced in the pseudo-code algorithm of the ALTER_REFRESH command. In this example, the algorithm may connect to a dataset snapshot using a database cluster. The dataset that the command wants to access represents heterogeneous environment storage. To present this dataset to the SPARQL computing layer so that the SPARQL computing layer can use it, access to physical resources may be required. Therefore, it may be necessary to link the dataset to a cluster that represents the physical resources available to the dataset. As is known, a cluster is a group of computing resources called nodes.

[0242] The computing resources may be statically allocated or dynamically allocated. Here, static means that all resources are available at the startup of the cluster, and dynamic means that the available resources automatically increase or decrease according to the workload. The computing resources may be embedded in the memory of the parent process or embedded as a stand-alone process.

[0243] The computing resources may be used to respond to the requirements of the SPARQL computing layer. Responding to the requirements of the SPARQL query engine means responding to eight triple patterns (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O), (?S,?P,?O); adding a graph that can be a variable to this results in a Basic Graph Pattern (BGP). Therefore, the resources of the cluster may be used to respond to these BGPs for each partition.

[0244] How partitions respond to these BGPs is not specific to this disclosure, and many solutions with various trade - offs may be selected. The virtual RDF graph database has both dynamic partitions and batch partitions. Dynamic partitions may be implemented by any RDF database, for example, Blazegraph as described at https: / / github.com / blazegraph / database, for example, an embedded in - memory RDF database. Batch partitions consist of a set of fragments, and each fragment is an RDF file with a metadata file. The RDF file may be in a human - readable format such as TriG (https: / / www.w3.org / TR / trig / ) or a binary format such as HDT; for HDT, see Fernandez, J.D., Martinez - Prieto, M.A., Gutierrez, C., Polleres, A., and Arias, M., 2013, Binary RDF presentation for publication and exchange (HDT), Journal of Web Semantics, 19, pp. 22 - 41. The metadata includes, but is not limited to, information useful for quickly eliminating an RDF file if it is necessary to use the RDF file to respond to triple patterns such as which graphs, predicates, and / or data types are being used. Thus, a batch partition first asks the metadata whether the file is useful to respond to the BGP, and if it is useful, accesses the file to respond to the BGP. For this reason, the file may or may not require an in - memory database to respond (such as in the case of TriG and not in the case of HDT). The format of the file may be changed without changing the present invention.

[0245] Here, the connection to the dataset will be described. When a dataset is created with the CREATE DATASET command, it is possible to connect to the dataset by specifying a cluster, that is, a group of computing resources called nodes. For example, the following CONNECT TO DATASET command may be used: CONNECT TO DATASET <dataset_name> USING CLUSTER ‘EMBEDDED’ The process may have several connections to different datasets. A cluster may be used for several datasets.

[0246] In an example of the CONNECT TO DATASET command, an embedded cluster may be used. Different node clusters may be defined: in that case, the definition depends on the architecture and cloud provider used. In such a case, the name of the cluster may be provided instead of the above “EMBEDDED” keyword.

[0247] After the CONNECT TO DATASET command, the SPARQL query engine can respond to queries using the dataset. For example, the cluster may download the metadata of the fragments locally and quickly eliminate the fragments without accessing external storage. Experiments have shown that the metadata is about 2% - 5% of the total size of the RDF file and is suitable as a candidate for external storage or prefetching.

[0248] For example, it may be possible to disconnect from the dataset using the following DISCONNECT FROM DATASET command. DISCONNECT FROM DATASET <dataset_name>

[0249] Here, we will explain the possibility of describing a stage, external storage, or dataset using a variant of the DESCRIBE command described at https: / / www.w3.org / TR / sparql11-query / #describe. For example, the DESCRIBE command could be something like: DESCRIBE STORAGE INTEGRATION <name> DESCRIBE STAGE <name> DESCRIBE DATASET <name>

[0250] List the parameters given with the CREATE command with potential restrictions. For example, the credentials given at the time of creating a storage command can be omitted for confidentiality reasons. In the case of a dataset, all files of the "last dataset snapshot" obtained from the catalog may be listed.

[0251] Here, the possibility of deleting a stage, external storage, or dataset using a variant of the DROP command described at https: / / www.w3.org / TR / sparql11-update / #drop is described. For example, the DROP command may be something like: DROP STORAGE INTEGRATION <name> DROP STAGE <name> DROP DATASET <name>

[0252] In one example, the DROP DATASET command may be executed using the following algorithm. It should be understood that this algorithm does not delete files from the stage. To do this, a "DROP STAGE" operation is required.

[0253] Input: Dataset Delete all information from the catalog necessary to create a connection to the dataset; / / Therefore, a new connection cannot be created.

[0254] Delete the last dataset definition information from the catalog; / / Therefore, a new transaction cannot be created.

[0255] In one example, the DROP STAGE command may be executed using the following algorithm.

[0256] Input: Stage Verify from the catalog that the stage exists and there are no datasets using it.

[0257] If not, return with this failure; Execute a "complete" garbage collection operation; If a currently running transaction is detected (i.e., the "live dataset snapshot" of the GARBAGE_COLLECT command's algorithm), return with this error; / / The drop stage could not complete successfully, so it needs to be re-executed; If not, delete all stage information from the catalog.

[0258] In one example, the DROP STAGE INTEGRATION command may be executed using the following algorithm.

[0259] Input: Storage integration Verify from the catalog that there is storage integration and no stage is using it; If not, return on this failure; Delete all information regarding storage integration from the catalog.

[0260] It has been explained that computing resources may be used to respond to the requirements of the SPARQL computing layer. Responding to the requirements of the SPARQL query engine means responding to eight triple patterns (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O), (?S,?P,?O); adding a graph that can be a variable to this results in a Basic Graph Pattern (BGP). It has also been explained that the virtual RDF graph database includes both dynamic partitions and batch partitions; the dynamic partitions may be implemented by any RDF database, and the batch partitions may consist of a set of fragments, where each fragment is an RDF file with a metadata file. The batch partitions first query the metadata as to whether the file is useful and respond to the BGP, and if it is useful, access the file and respond to the BGP.

[0261] Therefore, responding to a SPARQL query on storage means responding to a Basic Graph Pattern (BGP). A "compound transaction" means a transaction that uses both storages (i.e., a transaction that uses dynamic partitions and batch partitions), and therefore, all transaction operations (start / commit / abort / end) are executed on both storages.

[0262] As described in the previous explanations, the ACID properties are verified by both the storage that receives streaming changes and the storage that receives batch changes. When a composite transaction that requires both storages is initiated, a data set snapshot (e.g., the last data set snapshot) is obtained and connected to it. This data set snapshot is used for all queries of the transaction. When the connection is opened, the GARBAGE COLLECT command may be prevented from deleting the files linked from this data set definition. From the above example, the ADD PARTITION and REMOVE PARTITION commands, or the simultaneous algorithms of the ALTER REFRESH and GARBAGE COLLECT commands, do not affect the batch partitions used in this data set definition, so independence is guaranteed. This also applies to dynamic partitions due to the "standard" ACID properties of dynamic partitions. Therefore, the independence of the composite transaction is guaranteed. Similarly, indivisibility, consistency, and durability are also guaranteed. Therefore, the composite transaction is exactly an ACID transaction.

[0263] Referring to FIG. 2, a computer-implemented method is proposed for executing triple patterns of a SPARQL query against a virtual RDF graph database with guaranteed independence. Note that the triple patterns of the SPARQL query may include at least one of eight triple patterns: (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O), and (?S,?P,?O). This method includes obtaining (S200) the last registered snapshot of the virtual RDF graph database described in the catalog. Here, the term snapshot includes the first and second RDF graph databases. Since the first virtual RDF graph database is updatable by a tuple stream, the state of the dataset snapshot is obtained for a specific time, such as the time when the SPARQL query is executed. This implies that changes to the first RDF graph database after a given snapshot are not considered for the response to the SPARQL query. For each obtained triple pattern (S210), the triple pattern is executed against the first RDF graph database of the virtual RDF graph database. As already explained, this is executed as known in the art. Next, or simultaneously, the triple pattern is executed against the last snapshot of the second RDF graph database of the virtual RDF graph database registered in the catalog. As explained, the batch partition of the last snapshot first asks the metadata referred to in the catalog whether the file is useful and responds to the BGP. If it is useful, it accesses the file and responds to the BGP. In one example, the batch partition of the second RDF graph database may be fragmented. In this example, each fragment may be accompanied by metadata describing the fragment and may also be accompanied by a delta file storing updates to the fragment. To obtain an answer to the query, the metadata and the delta file are queried.For example, for each triple pattern of the query, a tuple responding to the triple pattern of the query may be found in the last snapshot, whereby a first set of results is obtained. Also, for each triple pattern of the query, it is determined in the delta file whether to delete and / or add a tuple responding to the triple pattern. When deleting a tuple, the tuple is deleted from the first set of results, and when adding a tuple, the tuple is added to the first set of results.

[0264] This method is implemented by a computer. This means that the steps (or substantially all steps) of this method are executed by at least one computer or any similar system. Thus, the steps of this method may be executed by a computer completely automatically or semi-automatically. In one example, at least some of the steps of this method may be triggered via an interactive operation between a user and a computer. The required level of interactive operation between the user and the computer may be according to the assumed level of automation and may be balanced with the need to implement the user's requirements. In one example, this level may be defined by the user and / or may be pre-defined.

[0265] A typical example of the implementation of this method by a computer is to execute this method using a system suitable for this purpose. The system may include a processor connected to a memory storing a computer program containing instructions for executing this method and a graphical user interface (GUI). The memory may store a database. The memory is any hardware suitable for such storage and may optionally include several physically distinguishable parts (for example, one for the program and optionally one for the database).

[0266] FIG. 6 shows an example of the system, and the system is a client computer system, for example, a user's workstation.

[0267] The client computer in this example includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and a random access memory (RAM) 1070 also connected to the bus. The client computer may further include a graphics processing unit (GPU) 1110 associated with a video random access memory 1100 connected to the bus. The video RAM 1100 is also known as a frame buffer in the art. A mass storage controller 1020 manages access to a mass storage device such as a hard drive 1030. Mass memory devices suitable for specifically implementing the instructions and data of a computer program include, by way of example, all forms of non-volatile memory including semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, and magneto-optical disks. Any of the foregoing may be supplemented or incorporated by a specially designed application specific integrated circuit (ASIC). A network adapter 1050 manages access to a network 1060. The client computer may also include tactile devices 1090 such as a cursor control device and a keyboard. The cursor control device is used within the client computer to enable a user to selectively position a cursor at any desired location on a display 1080. Further, the cursor control device enables a user to select various commands and input control signals. The cursor control device includes a number of signal generating devices for inputting control signals to the system. Typically, the cursor control device may be a mouse, and the buttons of the mouse are used to generate signals. Alternatively, or in addition, the client computer system may include a sensing pad and / or a sensing screen.

[0268] This computer program may include instructions executable by a computer, and the instructions include means for causing the above system to execute this method. The program may be recordable on any data storage medium including the system's memory. The program may be implemented, for example, in digital electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. The program may be implemented as an apparatus, such as a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. The steps of the method may be executed by a programmable processor executing a program of instructions to perform the functions of this method by operating on input data and generating output. Accordingly, the processor may be programmable to receive data and instructions from a data storage system, at least one input device, and at least one output device, and also to transmit data and instructions to them and be so connected. The application program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language as required. In any case, the language may be a compiler-type language or an interpreter-type language. The program may be a full-installation program or an update program. Applying the program to the system in any case results in instructions for executing this method. Alternatively, the computer program may be stored and executed on a server in a cloud computing environment that communicates with one or more clients via a network. In such a case, the processing unit executes the instructions included in the program so that this method is executed on the cloud computing environment.< / name> < / name> < / name> < / name> < / name> < / name> < / path> < / name> < / name> < / extension> < / partitionname> < / uuid> < / extension> < / partitionname> < / uuid> < / uuid>

Claims

1. A computer-implemented method for updating a virtual RDF graph database including tuples, the method comprising: providing: a file storage having a persistence D property of ACID properties, the file storage being distributed or non-distributed and guaranteeing consistent writes; a virtual RDF graph database, a first RDF graph database that can be updated by a stream of tuples added and / or deleted to the first RDF graph database; a second read-only RDF graph database stored in the file storage, the second read-only RDF graph database being updatable by a batch of tuples added and / or deleted to the second RDF graph database, thereby forming a read-only and updatable snapshot; a virtual RDF graph database including: a catalog for storing metadata describing the second read-only RDF graph database on the file storage, the catalog being compliant with ACID; providing; obtaining a stream of tuples added and / or deleted to the virtual RDF graph database and obtaining a batch of tuples added and / or deleted to the virtual RDF graph database; applying the stream of tuples to the first RDF graph database of the virtual RDF graph database; applying the batch of tuples to the second RDF graph database of the virtual RDF graph database, calculating a snapshot of the second RDF graph database including the batch of tuples, the calculation being performed appropriately using consistent writes to the file storage and guaranteeing the ACID properties of the snapshot by a series of concurrent control operations; storing the calculated snapshot in the file storage; registering the calculated snapshot in the catalog, thereby obtaining an updated description of the virtual RDF graph database; including; applying; including the method.

2. The ALTER REFRESH command executes the series of concurrent control operations that are properly executed, and the ALTER REFRESH command guarantees the ACID properties of the update of the second RDF graph database of the virtual RDF graph database. The series of concurrent control operations of the ALTER REFRESH are to uniquely identify the last snapshot of the second RDF graph database, which represents the state of the snapshot formed by the second RDF graph database at a specific point in time. to confirm that no other ALTER REFRESH commands are being executed simultaneously, and if so, to stop the ALTER REFRESH command and hold the ALTER REFRESH command that has already been executed. to obtain a list of files at a certain stage of all snapshots (groups) including the last snapshot. to filter the files in the list of files at the stage of the last snapshot, thereby obtaining a snapshot list indicating the names of the files of the last snapshot that can be used for subsequent access to the database at the specific point in time. to write the names of the files in the snapshot list, thereby obtaining a new snapshot not registered in the catalog. to upload the files of the new snapshot to the stage. After the catalog identifies the new snapshot as the last snapshot to which the batch of the tuples is applied in an ACID transaction, to register the new snapshot in the catalog. including A method implemented by a computer according to claim 1, characterized in that.

3. Obtaining the snapshot list indicating the names of the files of the last snapshot that can be used for subsequent access to the database at the specific point in time is to obtain a list of names of valid partition(s). For each partition named in the list of the valid partition(s), Check whether the current update operation for the partition is still pending. If it is pending, add the partition to the list of valid partition(s) for which the current update operation is still pending, and Check whether a past update operation for the partition was not executed successfully. If it was not executed successfully, check in the catalog whether a snapshot for which the failed update operation was executed is registered in the catalog. If it is not registered, ignore the failed update operation. If it is registered, remove the partition from the list of the valid partition(s) for which the update operation(s) is / are pending, and Create an empty snapshot list file, and add the name of the file(s) of the valid partition(s) from the list of the valid partitions for which the current update operation(s) is / are pending, or, if the list of the valid partition(s) for which a previous update operation(s) is / are still pending is empty, the name of the file(s) of the valid partition(s) from the list of the names of the valid partition(s), thereby obtaining the snapshot list indicating the name of the file(s) of the last snapshot that can be used for subsequent access to the database at the specific point in time. Check that each file listed in the snapshot list exists in the stage including A method implemented by a computer according to claim 2, characterized by the above.

4. Obtaining the list of the names of the valid partition(s) includes searching, in the file storage, for a file named with the partition name and a specific extension called VALID_EXTENSION for each partition, where the VALID_EXTENSION file stores the name of the file(s) of the partition, and the VALID_EXTENSION file was uploaded to the file storage when the partition was created, and / or Obtaining a list of names of the effective partition(s) involves, for each partition, searching within the file storage for a file named with the partition name and a specific extension called TOMBSTONE_EXTENSION, where the TOMBSTONE_EXTENSION file stores the name(s) of the file(s) of the partition that is the target of the current deletion operation, and the TOMBSTONE_EXTENSION file is uploaded to the file storage when the deletion operation for the partition is executed, and / or For each partition named in the list of the effective partition(s), checking whether a current update operation for the partition is still pending involves, within the file storage, searching for a file named with the partition name and a specific extension called PENDING_EXTENSION for each partition, where the PENDING_EXTENSION file stores the name(s) of the file(s) of the partition that is the target of the current update operation, and the PENDING_EXTENSION file is uploaded to the file storage when the current update for the partition is executed, and / or For each partition named in the list of the effective partition(s), checking whether a past update operation for the partition was not executed successfully involves, within the file storage, searching for a file named with the partition name and a specific extension called CONSUMED_EXTENSION for each partition, where the CONSUMED_EXTENSION file stores the name(s) of the file(s) of the partition that was the target of the past update operation, and the CONSUMED_EXTENSION file is uploaded to the file storage when the past update for the partition was executed, and includes searching. A method implemented by a computer according to claim 3, characterized by the above.

5. For each partition named in the list of the effective partition(s), checking whether a past update operation on the partition has been executed properly is further including checking that the content of the PENDING_EXTENSION file matches the content of the VALID_EXTENSION file, and if not, assuming that the list of files of the last snapshot is the content of the VALID_EXTENSION file, for each partition named in the list of the effective partition(s), if a current update operation on the partition is still pending, the PENDING_EXTENSION file is ignored at the time of the check A method implemented by a computer according to claim 4, characterized in that

6. Uniquely identifying the last snapshot of the second RDF graph database includes obtaining the last dataset snapshot identifier from the catalog and holding the last dataset snapshot identifier in memory as the previous dataset snapshot, Registering the new snapshot in the catalog includes, by the catalog, checking in an ACID transaction that the last snapshot described in the catalog is the previous dataset snapshot, if so, registering the new snapshot in the catalog as the last snapshot to which the batch of the tuples is applied, if not, determining that the ALTER REFRESH command has been executed simultaneously and deleting the new snapshot file from the stage further including A method implemented by a computer according to any one of claims 2 to 5, characterized in that

7. The ALTER PARTITION command executes a further sequence of concurrent control operations for updating the partition of the second read-only RDF graph database, and the ALTER PARTITION command is executed before the ALTER REFRESH command, the sequence of the concurrent control operations of the ALTER PARTITION is Download the VALID_EXTENSION file of the valid partition to be updated from the stage, thereby obtaining the name(s) of the file(s) of the valid partition. Verify that the ALTER REFRESH command is not being executed simultaneously on the valid partition. If it is being executed, stop the ALTER REFRESH command and retain the ALTER REFRESH command that has already been executed. Create a PENDING_EXTENSION file that stores the name(s) of the file(s) of the valid partition that is the target of the current update operation, where the name(s) of the file(s) of the valid partition is / are the target of the current update operation and is / are obtained from the batch of the tuple(s). Upload the PENDING_EXTENSION file to the stage. including A method implemented by a computer according to any one of claims 4 to 6, characterized by the above.

8. The snapshot of the second RDF graph database consists of a set of batch partitions, one or more batch partitions are divided into fragments, and the creation of the PENDING_EXTENSION file Determine whether the batch of tuples is a Change Data Capture (CDC) file. If the batch of tuples is not a CDC file, generate fragments of the batch of tuples, thereby obtaining a list of fragment files. If the batch of tuples is a CDC file, for each add / delete triple of the CDC file, identify one fragment of one or more batch partitions to which the add / delete operation is applied, thereby obtaining a list of fragments, and if any, together with a list of additional differential files. including The creation of the PENDING_EXTENSION file includes creating the PENDING_EXTENSION that stores the list of fragments, and if any, together with the list of additional differential files. A method implemented by a computer according to claim 7, characterized by the above.

9. The ADD PARTITION command executes a further sequence of concurrent control operations for adding a partition to the second read-only RDF graph database, and the ADD PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrent control operations for the ADD PARTITION is uploading the VALID_EXTENSION file of the partition to be added to the storage, where the success of uploading the VALID_EXTENSION file indicates that there is no partition with the same name in the file storage, uploading generating a batch fragment of the tuples for adding the partition on the second read-only RDF graph database, thereby obtaining a list of fragment files storing the list of fragments in the VALID_EXTENSION file and uploading the VALID_EXTENSION to the stage including A method implemented by a computer, as described in any one of claims 4 to 6, characterized in that.

10. The REMOVE PARTITION command executes a further sequence of concurrent control operations for deleting a partition from the second read-only RDF graph database, and the REMOVE PARTITION command is executed before the ALTER REFRESH command. The sequence of concurrent control operations for the REMOVE PARTITION is checking on the stage whether the partition to be deleted exists, and if not, stopping the REMOVE PARTITION command uploading a TOMBSTONE_EXTENSION file to the stage, where the success of the upload corroborates the existence of the partition to be deleted from the second read-only RDF graph database, uploading including A method implemented by a computer, as described in any one of claims 4 to 6, characterized in that.

11. Latching the second read-only RDF graph database stored in the file storage, which is executed before obtaining the list of files of the last snapshot at the stage, and unlatching after registering the new snapshot in the catalog further comprising A method implemented by a computer according to any one of claims 2 to 10, characterized in that

12. Latching the second read-only RDF graph database includes uploading a file named in the file storage with the name of the last data set snapshot and a specific extension called START_EXTENSION, and confirming that the START_EXTENSION file exists in the file storage when uploading the file of the new snapshot at the stage and in each subsequent step including Unlatching the new snapshot after registration in the catalog includes confirming that the START_EXTENSION file exists in the file storage and deleting the START_EXTENSION file A method implemented by a computer according to claim 11, characterized in that

13. The file storage is a distributed file storage, and / or the first RDF graph database is stored in an in-memory data structure A method implemented by a computer according to any one of claims 1 to 12, characterized in that

14. The catalog for storing metadata is stored in a database that guarantees ACID properties and strong consistency A method implemented by a computer according to any one of claims 1 to 13, characterized in that

15. A method implemented by a computer that executes triple patterns of a SPARQL query with guaranteed independence on a virtual RDF graph database previously updated according to any one of claims 1 to 14, comprising obtaining the last registered snapshot of the virtual RDF graph database described in the catalog For the obtained triple pattern, execute the triple pattern on the first RDF graph database of the virtual RDF graph database and the obtained last snapshot, thereby guaranteeing the independence of the execution of the triple pattern including A method characterized by this.