Read-only data structure for archiving snapshots of RDF datasets into files

By introducing a read-only data structure on a large RDF knowledge graph, using index lists, predicate index lists, dictionaries and adjacency matrix, the efficient compression and query problems of RDF data sets in the prior art are solved, and the near-real-time support for fast response to SPARQL queries and incremental updates are realized.

CN120179642APending Publication Date: 2025-06-20DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411887280.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to implement efficient data compression and querying on large RDF knowledge graphs, especially in supporting SPARQL queries and storage modification history.

Method used

Provides a read-only data structure that stores snapshots of RDF datasets through index lists, predicate index lists, dictionaries, and adjacency matrices, and supports all possible triple patterns of the SPARQL query engine. The data structure includes an index list of multiple graphs of the RDF dataset, a predicate index list of each graph, a dictionary that maps RDF terms to the index, and an adjacency matrix of each predicate.

Benefits of technology

It realizes efficient compression and query of RDF datasets, supports fast read access and near real-time incremental updates, and can efficiently answer SPARQL queries in all possible triple patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179642A_ABST
    Figure CN120179642A_ABST
Patent Text Reader

Abstract

The present disclosure relates, inter alia, to a computer-implemented read-only data structure for archiving snapshots of an RDF dataset into a file, the RDF dataset comprising one or more graphs, the read-only data structure being able to be queried directly using all possible triple patterns of a SPARQL query engine. The data structure includes: an indexed list of one or more graphs of an RDF data set; for each graph of the RDF dataset, an indexed list of predicates; mapping each RDF lexical item of the data set to a corresponding index and a dictionary mapped in reverse; and for each predicate indexed in the indexed list of predicates of each graph, representing an adjacency matrix of a set of tuples of the RDF dataset, thereby obtaining a complete representation of the RDF dataset that can be queried directly using all possible triple patterns of the SPARQL query engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly, to a read-only data structure for archiving snapshots of an RDF dataset into a file, modifying the snapshots, and querying the snapshots or the modified snapshots. Background Art

[0002] The W3C has published the RDF specification for representing information as a graph using triples. The publication can be found here: https: / / www.w3.org / TR / rdf11-concepts / . The core structure of the abstract syntax is a collection of triples, each triple consisting of a subject, a predicate, and an object. A set of such triples is called an RDF graph. An RDF graph can be visualized as a graph of nodes and directed arcs, where each triple is represented as a node-arc-node link.

[0003] SPARQL is a query language for RDF data. Version 1.1 of SPARQL can be found here: https: / / www.w3.org / TR / sparql11-query / . RDF is a directed labeled graph data format for representing information in the Web. The specification defines the syntax and semantics of the SPARQL query language for RDF. SPARQL can be used to express queries across diverse data sources, whether the data is stored natively as RDF or is treated as RDF via middleware. SPARQL includes the ability to query required and optional graph patterns and their conjunctions and disjunctions. SPARQL also supports aggregation, subqueries, negation, creating values via expressions, extensible value testing, and constraining queries by the source RDF graph. The result of a SPARQL query can be a result set or an RDF graph.

[0004] RDF knowledge graphs are formed by joining triples to create information networks or information graphs. These knowledge graphs can be very large, up to hundreds of billions of triples, e.g., Uniprot (100 billion triples), Wikidata (17 billion triples), 3DS SIEM (up to 500 billion triples). Thus, it is impossible or very expensive to keep the entire graph in memory.

[0005] Experiments have shown that most of these RDF knowledge graphs are read-only, or at least rarely modified and not modified in a highly concurrent manner. In the world of relational databases, this is known as an OLAP use case. OLAP allows for fast and interactive access to aggregated, summarized, and computed data, thus allowing users to easily explore and analyze complex data sets. OLAP stands for Online Analytical Processing and is discussed, for example, here: https: / / en.wikipedia.org / wiki / Online_analytical_processing.

[0006] In real-world scenarios, such as OLAP in SQL as seen in the well-known TPC-H benchmark which can be found here https: / / www.tpc.org / tpch / default5.asp, there are still small write modifications that occur concurrently with read queries. A typical use case in the industry is that sparse write requests can be sent to large RDF knowledge graphs storing chemical formulas, while most of the stored formulas do not need to be changed.

[0007] In the realm of relational databases, Apache Parquet (https: / / parquet.apache.org / ) is known. It is an open-source, column-oriented data file format designed for efficient data storage and retrieval. It provides efficient data compression and encoding schemes, with enhanced performance for batch processing of complex data. Apache Parquet was built to support very efficient compression and encoding schemes. Multiple projects have demonstrated the performance impact of applying the right compression and encoding schemes to data. Parquet allows the specification of compression schemes at the per-column level and is not obsolete to allow for the addition of more encodings as they are invented and implemented. The target of Apache Parquet is relational data rather than graph data; however, it highlights the need for very efficient data compression and encoding schemes.

[0008] Still in the realm of relational databases, Apache Arrow (https: / / arrow.apache.org / docs / index.html) is also known. Compared to Apache Parquet, no compression scheme is used, but instead it relies on zero-copy shared memory.

[0009] In the field of graph databases, RDF HDT (Header - Dictionary - Triples) is a W3C recommendation (http: / / www.w3.org / Submission / 2011 / 03 / ) for a binary format (also known as the HDT file format) for the large - scale publication and exchange of RDF data. Its compact representation allows RDF to be stored in less space while providing direct access to the stored information. This is achieved by depicting the RDF graph from the perspective of three main components: the header, the dictionary, and the triples. The header includes extensible metadata needed to describe details of the RDF dataset and its internals. The dictionary organizes the vocabulary of strings present in the RDF graph by assigning a numeric ID to each distinct string. The triples component depicts the structure of the underlying graph in a compressed form.

[0010] HDT has not yet passed beyond the submission stage. HDT was initially envisioned as a binary serialization format for transporting RDF data, but it has been used as an RDF compressor due to its compactness. As a binary format, HDT only allows basic retrieval and has been extended later with various changes that trade off read performance / write performance / memory.

[0011] HDT++ is a variant of HDT and improves the HDT compression ratio, but it does not retain the original HDT retrieval capabilities.

[0012] iHDT++ is an extension of HDT and is discussed in: Hernández - Illera, A., Martínez - Prieto, M.A., Fernández, J.D. and A., 2020, iHDT++: improving HDT for SPARQL triple pattern resolution (iHDT++: improving HDT for SPARQL triple pattern resolution), Journal of Intelligent & Fuzzy Systems, 39(2), pp. 2249 - 2261, and is also discussed in: Hernández - Illera, A., Martínez - Prieto, M.A. and Fernández, J.D., April 2015, Serializing RDF in compressed space, 2015 Data Compression Conference (pp. 363 - 372), IEEE[5]. iHDT++ implements support for triple patterns for SPARQL parsing.

[0013] HDTFoQ is another extension of HDT and is discussed in the following literature: Martínez Prieto, M.A., Arias Gallego, M., and Fernández, J.D., 2012, Exchange and consumption of huge RDF data. The Semantic Web: Research and Applications: 9th Extended Semantic Web Conference, ESWC 2012, Heraklion, Crete, Greece, May 27 - 31, 2012. HDTFoQ is a variation on complete SPARQL queries, but its main drawback is that the read algorithm is too costly and does not scale to large RDF knowledge graphs, e.g., RDF knowledge graphs that include billions of triples.

[0014] Both HDT++ and iHDT++ are based on the assumption of organizing "class objects". For example, each subject must belong to only one predicate "family". The "basic" HDT does not support all access patterns and only supports access by subject. This is a significant limitation.

[0015] Thus, HDT was mainly designed as a binary format for large - scale publishing and exchanging RDF data rather than answering SPARQL queries. This is why there is this initial assumption that all triples must be organized by subject. Subsequent attempts to incorporate HDT into SPARQL were restrictive and made it incompatible with querying very large RDF knowledge datasets (e.g., billions of triples). HDT optimizes the memory / storage device used at the cost of read (and write / generation) performance: efficient data storage rather than retrieval. The so - called RUM conjecture (Read Update Memory) is discussed in the following literature: M. Athanassoulis, M.S. Kester, L.M.M. Maas, R. Stoica, S. Idreos, A. Ailamaki, M. Callaghan, Designing Access Methods: The RUM Conjecture, Proceedings of the 19th International Conference on Extending Database Technology (EDBT), March 15 - 18, 2016, Bordeaux, France, and this literature concludes that since some optimizations and design choices are mutually exclusive, creating a single optimal access method is not feasible. Another compromise is needed that uses a less compressed format but enables fast read access to SPARQL queries that are not limited to "retrieval by subject".

[0016] Known solutions in the art propose data compression techniques (https: / / en.wikipedia.org / wiki / Data_compression) to reduce the storage size for RDF graph databases (e.g., in memory or on disk). Such compression techniques are used to reduce storage costs (e.g., hardware costs) and environmental footprint in various applications, especially in cloud deployments.

[0017] The following literature discloses an RDF indexing technique that supports SPARQL solutions in a compressed space: S. et al., "Compressed vertical partitioning for efficient RDF management.", Knowledge and Information Systems, 2015, Vol. 44, No. 2, pp. 439 - 474. The disclosed technique (referred to as k2 triples) vertically partitions the dataset into subsets of disjoint pairs (subject, object) using predicates, one for each predicate. These subsets are represented as binary matrices of subject × object, where a 1 bit means the corresponding triple exists in the dataset. This model produces very sparse matrices that can be efficiently compressed using k2 trees.

[0018] The following literature discloses an algorithm that takes into account the characteristics of the graph and performs compression based on quadtrees: CHATTERJEE, A. et al., "Exploiting topological structures for graph compression based on quadtrees", The 2nd International Conference on Research in Computational Intelligence and Communication Networks (ICRCICN), 2016, IEEE, 2016, pp. 192 - 197. In addition, techniques for compressing data and performing queries on the compressed data itself are also introduced and discussed in detail.

[0019] The following literature discloses the use of a novel data structure for streaming graphs: NELSON, M. et al., "Queryable compression on streaming social networks", 2017 IEEE International Conference on Big Data (Big Data), IEEE, 2017, pp. 988 - 993, which is based on an indexed matrix of compressed binary trees that directly constructs the graph without using any temporary storage structures. This data structure provides fast access methods for edge existence (Is there an edge between two nodes?), neighbor queries (List the neighbors of a node), and streaming operations (Add / Delete nodes / edges).

[0020] Pelgrin, O., Taelman, R., Galárraga, L., and Hose, K., in the literature of GLENDA: Querying over RDF Archives with SPARQL, mentioned that "RDF datasets on the Internet are constantly evolving. The simplest way to track the history of RDF data is to store each revision of the dataset as a separate copy. However, for large RDF datasets with a long history, this can be daunting. This observation led to the emergence of more efficient methods for managing (i.e., storing and querying) large RDF archives."

[0021] The purpose of an RDF archive is different from that of a standalone HDT: to track the history of RDF data and be able to query its history by tracing the change history of RDF graphs. As learned in Pelgrin, O., Taelman, R., Galárraga, L., and Hose, K., in February 2023, Scaling Large RDF Archives To Very Long Histories, Proceedings of the 17th International Conference on Semantic Computing (ICSC) 2023 by IEEE (pp. 41 - 48), the applications are different, for example, version control, mining of temporal and provenance patterns, temporal data analysis, or querying the past and studying the evolution of a given knowledge domain. An RDF archive needs to address very large knowledge graphs that can be updated in batches of millions of triples or from continuous updates. The technique used to capture, store, and query long revision histories on very large RDF graphs is a combination of dataset snapshots and sequences of aggregated change sets, which is called a delta chain in the literature. The proposed solution is based on multiple delta chains, where full snapshots are inserted between change sets when the history becomes too long. All of this only involves one dataset: the dataset whose history is being tracked.

[0022] In the document "Scaling Large RDF Archives To Very Long Histories": "Snapshots are stored as HDT files, while the delta chains are materialized into two stores: one for adds and one for deletes. Each store consists of three indexes in different triple component orders (i.e., SPO, OSP, and POS), implemented as B+ trees. The keys in these indexes are individual triples linked to version metadata, i.e., the revisions of triples that exist and do not exist. In addition to the change stores, there is an index with add and delete counts for all possible triple patterns (e.g., <?s,?p,?o> or <?s, cityIn,?o>), which can be used to efficiently compute cardinality estimates, especially useful for SPARQL engines. As is common in RDF stores, RDF terms are mapped to an integer space for efficient storage and retrieval. Two disjoint dictionaries are used in each delta chain: the snapshot dictionary (using HDT) and the delta chain dictionary. Thus, the multi-snapshot approach uses D×2 (possibly disjoint) dictionaries, where D is the number of delta chains in the archive. The fetch routine depends on whether the revision will be stored as an aggregated delta or a snapshot. For version 0, the fetch routine takes the full RDF graph as input to build the initial snapshot. For subsequent revisions, a standard change set is taken as input, and OSTRICH is used to build an aggregated change set of the form us, k, where the revision s = snapshot(k) is the most recent snapshot in history. When the snapshot policy decides to materialize a revision as a snapshot, the aggregated change set is used to efficiently compute the snapshot."

[0023] The method of the document "Scaling Large RDF Archives To Very Long Histories" has two main drawbacks. First, materializing the delta chains requires querying the snapshot dictionary, which is not conducive to generating snapshots near real-time. Second, the snapshots are based on HDT, which is not conducive to efficiently answering SPARQL queries on large datasets, as discussed above.

[0024] In this case, there is still a need for fast read access to very large compressed RDF archives that can be incrementally updated (e.g., adding and / or deleting triples). This read access must efficiently answer SPARQL queries for all possible triple patterns. Regardless of the size of the RDF knowledge graph, the incremental updates must be near real-time. Summary of the Invention

[0025] Accordingly, a computer-implemented read-only data structure for archiving a snapshot of an RDF dataset into a file is provided, the RDF dataset including one or more graphs, the read-only data structure being capable of directly querying using all possible triple patterns of a SPARQL query engine. The data structure includes: an indexed list of one or more graphs of the RDF dataset; for each graph of the RDF dataset, an indexed list of predicates; a dictionary that maps each RDF term of the dataset to a corresponding index and vice versa; for each predicate indexed in the indexed list of predicates of each graph, an adjacency matrix representing a set of tuples of the RDF dataset, thereby obtaining a complete representation of the RDF dataset that can be directly queried using all possible triple patterns of a SPARQL query engine.

[0026] The method may include one or more of the following:

[0027] - When each value of the dictionary has a length equal to or less than a predetermined threshold, the dictionary has a fixed length. Preferably, the predetermined threshold is 2, 4, or 8 bytes;

[0028] - When at least one value of the dictionary has a length greater than the predetermined threshold, the dictionary has a variable length, and each value having a length greater than the predetermined threshold is indexed to and stored in an overflow data structure, which is part of the dictionary;

[0029] - At least part of the dictionary is encoded. Preferably, the values are encoded using the IEEE floating-point arithmetic standard (IEEE 754-2019 or a previous version), or the prefixes of the values are encoded using a hexadecimal key;

[0030] - The dictionary includes two or more dictionaries. Preferably, each dictionary maps each RDF term of an RDF data type of the dataset to a corresponding index and vice versa, or each dictionary maps each RDF term of an IRI prefix of the dataset to a corresponding index and vice versa;

[0031] - The adjacency matrix for each indexed predicate is obtained by performing vertical partitioning for each predicate of the RDF database. Preferably, the vertical partitioning is performed using a technique selected from: preferably, standard K2 triples constructed according to dynamic K2; two static B+ trees, where the first static B+ tree represents the subject-to-object (S, O) correspondence and the second represents the object-to-subject (O, S) correspondence; a static XAM tree carved from an XAMTree and being a read-only XAMTree;

[0032] - Each entry of the data structure has a maximum size between 32 bits and 256 bits, preferably between 32 bits or 64 bits. More preferably, the size of each entry is encoded using 32 bits or 64 bits;

[0033] - A read-only data structure is a binary file including the following items: a header including an offset to the end of the position of the page footer record; a section between the header and the page footer; a page footer including an offset to the beginning of the position of the header record; wherein the section stores: a dictionary; for each graph, an indexed list of the predicates of the graph and an adjacency matrix for each predicate indexed in the indexed list of the predicates of the graph;

[0034] A computer-implemented method for storing modifications to be applied to the above read-only data structure is also provided. The method includes: obtaining a first list of added and / or deleted tuples of an RDF data set; calculating the first read-only data structure defined above, wherein for each predicate of the added and / or deleted tuples of the RDF data set in the first list: a first adjacency matrix represents the added tuples of the RDF data set including the same predicate; and / or a second adjacency matrix represents the deleted tuples of the RDF data set including the same predicate; and storing the calculated read-only data structure into a first file.

[0035] The method may further include:

[0036] - obtaining a second list of added and / or deleted tuples of the RDF data set; calculating the second read-only data structure defined above, wherein for each predicate of the added and / or deleted tuples of the RDF data set in the first list and the second list: a first adjacency matrix represents the added tuples of the RDF database including the same predicate; and / or a second adjacency matrix represents the deleted tuples of the RDF database including the same predicate; and storing the calculated read-only data structure into a second file, thereby forming the modifications to be applied to the read-only data structure;

[0037] - obtaining further includes: obtaining a list of added and / or deleted tuples of the RDF data set for each graph of the RDF set; and wherein calculating includes: calculating a first adjacency matrix and a second adjacency matrix for each graph of the RDF set;

[0038] - calculating a mapping between the index of the dictionary of the RDF data set and the index of each dictionary of each list of added and / or deleted tuples of the RDF data set; and storing the calculated mapping into a file.

[0039] There is also provided a computer-implemented method for performing SPARQL queries on a read-only data structure stored in a file and a computed read-only data structure stored in a first file, the computer-implemented method comprising: obtaining, by a SPARQL query engine, a SPARQL request, the SPARQL query comprising at least one triple pattern; obtaining the read-only data structure and the computed read-only data structure stored in the file; for each triple pattern of the query, finding in the data structure according to any one of claims 1 to 8 a tuple that matches the triple pattern of the query, thereby obtaining a first set of results; and for each triple pattern of the query, in the computed read-only data structure in the file: determining whether to delete and / or add a tuple that matches the triple pattern; if a tuple is to be deleted, removing the tuple from the first set of results; if a tuple is to be added, adding the tuple to the first set of results.

[0040] There is also provided a computer program comprising instructions for performing a method of storing modifications to be applied to the above-mentioned read-only data structure and / or for performing a method of SPARQL query.

[0041] There is also provided a computer-readable storage medium having a computer program recorded thereon.

[0042] There is also provided a system comprising a processor connected to a memory and a graphical user interface, the memory having a computer program stored thereon.

[0043] There is also provided a device comprising a data storage medium having a computer program recorded thereon. The device may form or be used as a non-transitory computer-readable medium (e.g., on a SaaS (Software as a Service) or other server) or a cloud-based platform, etc. The device may alternatively comprise a processor connected to the data storage medium. The device may thus form, in whole or in part, a computer system (e.g., the device is a subsystem of the entire system). The system may also comprise a graphical user interface connected to the processor. Description of the Drawings

[0044] Non-limiting examples will now be described with reference to the drawings, in which:

[0045] Figure 1 is an example of a complete representation of a read-only data structure;

[0046] Figure 2 shows an example of a flowchart of a method for storing modifications to be applied to a read-only data structure;

[0047] Figure 3 shows an example of a flowchart of a method for performing a SPARQL query on a read-only data structure;

[0048] Figure 4 are examples of the system;

[0049] Figures 5 to 10 are commands for executing Figure 3 the method; and

[0050] Figure 11 shows a performance comparison with a sample query on a read - only data structure. Detailed Description

[0051] Referring Figure 1 , an object of the present invention is to propose a read - only data structure for archiving a snapshot of an RDF dataset into a file. The RDF dataset includes one or more graphs. All possible triple patterns of a SPARQL query engine can be used to directly query the read - only data structure. The data structure includes an index list (130) of one or more graphs of the RDF dataset. For each graph of the RDF dataset, the data structure also includes an index list (140) of predicates. The data structure also includes dictionaries (150, 160) that map each RDF term of the dataset to a corresponding index and vice versa. Additionally, for each predicate indexed in the predicate index list of each graph, the data structure includes an adjacency matrix (170) representing a set of tuples of the RDF dataset, whereby a complete representation of the RDF dataset that can be directly queried using all possible triple patterns of a SPARQL query engine is obtained.

[0052] This data structure provides an improved read - only data structure for archiving a snapshot of an RDF dataset into a file. By "improved archiving", it means that the data structure stores the RDF dataset in a compressed format that reduces the overall size of the RDF dataset. In fact, the archive uses a succinct data structure for its graph data representation and can be stored on disk in a very efficient manner. It should be reminded that, unlike other compressed representations, a succinct data structure is a data structure that uses an amount of space "close" to the information - theoretic lower bound but still allows efficient query operations. Thus, the use of the predicate index list for each graph, the dictionaries, and the adjacency matrix for each predicate contribute to reducing the overall size of the archive. Additionally, the compression can be improved by using a dedicated encoder for one or more of these succinct data structures. For example, the dictionary can be compressed.

[0053] In addition, the archive provided by the read-only data structure defines an RDF file / memory format designed to enable efficient retrieval and data storage. By supporting zero-copy reads, lightning-fast and resource-efficient data access is allowed without serialization overhead. The archive is generated only once and can be read multiple times. The archive is self-describing and can be mapped in one pass. The archive file describes an RDF dataset and can be regarded as an indexed trig file, where trig files are discussed here: https: / / www.w3.org / TR / trig / -. Therefore, the archive provided by the read-only data structure can be used persistently over a long period, and its organization allows data to be read without any transformation (e.g., no serialization is required).

[0054] In addition, the data structure of the archive is based on a state-of-the-art graph representation that implements full SPARQL support. Thus, read access can efficiently answer SPARQL queries for all possible triple patterns: the eight different triple patterns possible in a SPARQL query. These eight triple patterns include (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O), and (?S,?P,?O), where the variable is preceded by the symbol?. The variable is the output of the triple pattern and can be the output of a SPARQL query.

[0055] Last but not least, the read-only data structure can include one or more graphs, and thus, the read-only data structure according to the present invention complies with the provisions of the RDF 1.1 recommendations, which are discussed here: https: / / www.w3.org / TR / rdf11-concepts / –. In addition to compliance with the recommendations, this also allows the read-only data structure to be implemented and used on current RDF datasets that are mostly multi-graphs over time. In other words, the read-only data structure solves the industry problem of archiving read-only multi-graph RDF datasets.

[0056] Now refer to Figure 2, a computer-implemented method for storing modifications to be applied to the read-only data structures discussed above is also proposed. As introduced above, one requirement is to obtain a read-only data structure capable of obtaining incremental updates in near real time. For this purpose, the method for storing modifications to be applied to a read-only data structure according to the present invention includes obtaining S200 a first list of added and / or deleted tuples of an RDF dataset; thus, the tuples are modifications to be applied to the archive. Then, S210 calculates a first read-only data structure. This first read-only data structure is similar to the structure discussed above, but when constructing the first read-only data architecture, for each predicate of the added and / or deleted tuples of the RDF dataset in the first list, the first adjacency matrix represents the added tuples of the graph database including the same predicate, and / or the second adjacency matrix represents the deleted tuples of the graph database including the same predicate. The method also includes storing S220 the calculated read-only data structure in a first file. A snapshot of the first list of added and / or deleted tuples of the RDF dataset is archived into the file as a read-only data structure.

[0057] This method allows the read-only data structure of the present invention to be updated. The update depends on calculating a first read-only data structure that reflects the changes (added and / or deleted tuples) to be applied to the read-only data structure (archive). This first read-only data structure is functionally similar to the delta file of the RDF archive: the archive (a snapshot of the RDF dataset that becomes a file) is followed by one or more delta files, where one or more delta files form a delta chain. The snapshot and the delta chain are constructed using the same read-only data structure. Interestingly, one or more delta files inherit the characteristics of the read-only data structure discussed previously. It should be noted that the read-only data structure enables fast read access to SPARQL queries. Additionally, the calculation of the delta file is in near real time because it does not require access to the snapshot file and is thus independent of the size of the RDF dataset (the full graph).

[0058] Now refer to Figure 3, a computer-implemented method for performing SPARQL queries on read-only data structures and computed read-only data structures (first file or delta file) stored in a file (archive) is also proposed. The method includes obtaining S300 a SPARQL request by a SPARQL query engine, the SPARQL query including at least one triple pattern. Next, obtaining S310 the read-only data structure for archiving a snapshot of the RDF dataset into a file, and the computed read-only data structure stored in the first file. In other words, obtaining the snapshot file of the RDF dataset and also obtaining the delta file as the read-only data structure according to the present invention. For each triple pattern S320 of the query, finding tuples in the read-only data structure (snapshot file of the RDF dataset) that match the triple pattern of the query, thereby obtaining a first set of results. For each triple pattern S330 of the query, in the computed read-only data structure in the first file, perform the following steps: i) determine whether to delete and / or add tuples that match the triple pattern; ii) if a tuple is to be deleted, remove the tuple from the first set of results; iii) if a tuple is to be added, add the tuple to the first set of results.

[0059] Providing the combination of the snapshot file of the RDF dataset and the delta file to the SPARQL query engine in a unified logical view so as to be able to answer SPARQL queries with good native performance, which is contrary to HDT discussed in the background section, where additional data structures were added later in an attempt to mitigate poor read performance. Thus, fast read access can be performed on a compressed RDF archive that is incrementally updated near real-time.

[0060] The read-only data structure and method are computer-implemented. This means that the steps (or substantially all steps) of the read-only data structure and method are performed / used by at least one computer or any similar system. Thus, the steps of the method may be performed by the computer either fully automatically or semi-automatically. In an example, triggering of at least some steps of the method may be performed through user-computer interaction. The level of required user-computer interaction may depend on the level of automation foreseen and be balanced with the need to fulfill the user's wishes. In an example, the level may be user-defined and / or predefined.

[0061] A typical example of the computer implementation of the method is to use a system suitable for the purpose to perform the method. The system may include a processor connected to a memory and a graphical user interface (GUI), and a computer program including instructions for performing the method is recorded on the memory. The memory may also store a database, for example, by using memory mapping. The memory is any hardware suitable for such storage, which may include several physically distinct parts (e.g., one part for the program and possibly one part for the database).

[0062] Figure 4 An example of the system is shown, where the system is a client computer system, e.g., the user's workstation. The system can be used to build, maintain, store, use read-only data structures, and / or the system can be used to execute methods for storing modifications to be applied to the read-only data structures and / or methods for performing SPARQL queries on the read-only data structures and the computed read-only data structures stored in a first file.

[0063] The client computer of this example includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and a random access memory (RAM) 1070 also connected to the bus. The client computer is also provided with a graphics processing unit (GPU) 1110, which is associated with a video random access memory 1100 connected to the bus. The video RAM 1100 is also known as a frame buffer in the art. A mass storage device controller 1020 manages access to a mass storage device (such as a hard disk drive 1030). Mass storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, and magneto-optical disks. Any of the foregoing may be supplemented by or incorporated into a specially designed ASIC (application specific integrated circuit). A network adapter 1050 manages access to a network 1060. The client computer may also include a haptic device 1090, such as a cursor control device, a keyboard, etc. A cursor control device is used in the client computer to allow the user to selectively position the cursor at any desired location on a display 1080. Additionally, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a plurality of signal generating devices for inputting control signals to the system. Generally, the cursor control device may be a mouse, and the buttons of the mouse are used to generate signals. Alternatively or additionally, the client computer system may include a touch pad and / or a touch screen.

[0064] A computer program may include instructions executable by a computer system, the instructions including components for causing the above system to perform the construction, maintenance, storage, and use of read-only data structures, and / or components for executing methods for storing modifications to be applied to the read-only data structures, and / or components for executing methods for performing SPARQL queries on the read-only data structures and the computed read-only data structures stored in a first file.

[0065] The program can be recorded on any data storage medium (including the system's memory). For example, the program can be implemented in digital electronic circuits, or in computer hardware, firmware, software, or in a combination thereof. The program can be implemented as an apparatus, for example, a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. Method steps can be performed by a programmable processor that executes a program of instructions to perform the functions of the method by operating on input data and generating output. Thus, the processor can be programmable and connected to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to send data and instructions to them. The application program can be implemented in a high-level procedural language or an object-oriented programming language, or, if desired, in assembly language or machine language. In any case, the language can be a compiled or interpreted language. The program can be a full installation program or an update program. In any case, the application program on the system produces instructions for performing the method. The computer program can alternatively be stored and executed on a server in a cloud computing environment that communicates with one or more clients via a network. In this case, the processing unit executes the instructions contained in the program, thereby causing the method to be executed on the cloud computing environment.

[0066] Now refer to Figure 1 an example showing one of these examples to discuss examples of read-only data structures. In the remainder of the specification, the read-only data structure will equivalently be referred to as a ROSA file, an archive. ROSA stands for Read-Only Semantic Archive. A ROSA file (or ROSA archive) consists of only one file. A ROSA file includes several sections, and the collection of these sections defines the organization of the ROSA file. It should be understood that Figure 1 the ROSA file represented above shows a non-limiting example, and the positions of the sections can be changed without modifying the present invention. In Figure 1 another example, section 130 can be placed between sections 180 and 190, and section 140 can be located after section 180. Thus, when generating a ROSA file, section 140 is generated after each set of adjacency matrices related to the diagram of section 180 ( Figure 1 … diagram n), and section 130 is generated before generating the footer section 190. This order allows the ROSA file to be generated in one go.

[0067] Dividing the ROSA file into parts realizes the following fact: the memory mapping can be used to have the same layout on disk and in memory, and memory-map the file in memory when opened. Memory mapping is well-known and a discussion of memory mapping can be found here: https: / / en.wikipedia.org / wiki / Memory_map. It should be understood that memory mapping is an optimization and any other memory layout can be used without changing the present invention. Once the ROSA file is memory-mapped, only some handles and necessary internal data structures can be loaded: the codec list, the graph index list, and the mapping type. The ROSA file is designed to be mapped, but it may be interesting to be able to manipulate such a file in a streaming manner to perform filtering. For example, discard some predicates or graphs. Some tools can be made that can transform some ROSA archives to add or remove some parts. Writing a utility to remove predicates or graphs and generate a new ROSA archive is quite straightforward. This can be done without changing the present invention.

[0068] In an example, the ROSA file at least includes the following data structures:

[0069] i) A list of indexes 130 of one or more graphs of the RDF dataset;

[0070] ii) For each graph of the RDF dataset, a list of indexes 140 of the predicates of the graph;

[0071] iii) A dictionary 150 that maps each RDF term of the dataset to a corresponding index and vice versa 160;

[0072] iv) For each predicate indexed in the predicate index list of each graph, an adjacency matrix 180 representing a set of tuples of the RDF dataset.

[0073] Now discuss these data structures.

[0074] As is known per se, a dictionary is a data structure that maps each RDF term of a data set to a corresponding index. Conversely, the dictionary also maps each index to the corresponding RDF term of the data set. Thus, a "dictionary" for a graph database is a mapping, i.e., an encoding or an index, that maps each stored ID in the database (e.g., an ID stored as an RDF triple) to the corresponding content. The dictionary is used to provide an index to the RDF triple store to help optimize the persistence of potentially large amounts of duplicate information. In the context of RDF (Resource Description Framework), a "term" refers to the basic information unit or element (triple) used in an RDF statement. Each term in RDF can be used as the subject, predicate, or object of a triple. There are three types of terms in RDF: subject, predicate, and object. There can be three types of nodes in an RDF graph: IRIs, literals, and blank nodes, e.g., as discussed here: https: / / www.w3.org / TR / rdf11-concepts / . Each term can be identified by an IRI (Internationalized Resource Identifier) that permits a wider range of Unicode characters or literal values such as strings or numbers. The subject is the resource about which the statement is made. It is the "thing" or "entity" described by the RDF triple. The predicate is the property or relationship that connects the subject to the object. It defines the nature of the relationship between the subject and the object. And, the object is the value or target of the statement; it can be another resource or a literal value such as a string, number, or date. Thus, an RDF triple is formed by the combination of these terms (subject, predicate, object).

[0075] The dictionary is the most complex data structure in ROSA. In fact, the terms being indexed can vary greatly in size, which has a significant impact on the organization of the dictionary data structure. Thus, the dictionary can be implemented in a way that adapts to the terms found in the RDF data set; this will improve the performance of the dictionary.

[0076] In an example, the dictionary can have a fixed length when each value of the dictionary has a length equal to or less than a predetermined threshold or less than a predetermined threshold. A fixed-size dictionary can be implemented when the dictionary fits the threshold (e.g., number of bytes), where all data conforms to a uniform data structure. The fixed-size dictionary saves memory usage by allocating a fixed amount of space for each entry of the dictionary. This reduces the need for dynamic memory allocation and resizing, and the memory efficiency can be higher. The fixed-size dictionary also saves the indirect cost, i.e., the overhead associated with indirectly accessing data through pointers or references.

[0077] A fixed-size dictionary can consist of a vector of values called "master", where each value has a maximum length fixed to a threshold. Thus, the vector is a matrix. The position of a value within the vector is the index of that value. The value must still have a size below the threshold, as the value remains in place within the vector. To build such a dictionary, the dictionary can be built in memory and then written in a ROSA file without any conversion.

[0078] In an example, the predetermined threshold can be 2, 4, or 8 bytes, meaning that the terms indexed in the dictionary have a size not exceeding (or equal to) 2, 4, or 8 bytes. The choice of value can depend on one or more parameters, e.g., the characteristics of the data and the requirements of the system. Considered but not limited to are the characteristics of the storage medium (e.g., hard disk drive, solid state drive), the characteristics of the system's memory, future growth, and changes in the data pattern...

[0079] In an example, the dictionary can have variable length, e.g., when at least one value of the dictionary has a length greater than the predetermined threshold. A variable-length dictionary has an additional data structure: each value in the dictionary with a length greater than the predetermined threshold is indexed to and stored in an overflow data structure. The overflow data structure is part of the dictionary, meaning that the overflow data structure can be regarded as a sub-data structure of the dictionary's data structure. In this example, and as in the example regarding the fixed-length dictionary, the predetermined threshold can be 2, 4, or 8 bytes. The choice of value can also depend on one or more parameters.

[0080] For a variable-length dictionary, the same "master" vector can be used. Each time the value is above the threshold, instead of writing all values to their proper positions, a reference is made to another vector (called "blob (binary large object)") where the overflow is written. The index is still given by the position in the "master" vector, but the indirection must be checked to obtain the full value. In an example of creating such a dictionary, in the first step, the "master" vector and the "blob" vector can be built in memory, and in the second step, they can be copied to the ROSA file together with an offset matrix calculated from the cumulative size of all "blob" vectors. For example, if the maximum value fits in an unsigned 32-bit integer, it can be dumped as a matrix of unsigned 32-bit integers; otherwise, a matrix of 64-bit unsigned integers can be used, and the 64-bit unsigned integer matrix can be referenced in the main part of the dictionary's entries.

[0081] In an example, a dictionary can be encoded fully or partially, e.g., the index and / or the value are encoded. Any type of encoding can be used. In an example, the value can be encoded using the IEEE floating-point arithmetic standard (IEEE 754-2019 or a previous version): double precision or floating-point numbers can be represented with a normalized memory layout, and the normalization may depend on IEEE 754-2019 or a previous version (https: / / en.wikipedia.org / wiki / IEEE_754) rather than the value remaining as a string of characters. Such encoding can be used to optimize queries, e.g., filtering all values that match "x > 5". For example, if the dictionary is dedicated to all floating-point values encoded with IEEE 754, then by definition, the dictionary will be a fixed-size dictionary. In an example, one or more (even all) prefixes of the dictionary can be encoded using the corresponding hexadecimal key instead of duplicating the prefix in all values that start with that prefix. The prefix defines an abbreviation of a long IRI. Since the IRI mapping codec is a special codec dedicated to prefixes, the IRI mapping codec can be used. The mapping is a table that contains a list of the defined prefixes and provides a code associated with the type.

[0082] As discussed until now, a dictionary can be used to archive all values of all triples of an RDF dataset. In an example, a dictionary can include two or more dictionaries, that is, the dictionary can be split into several smaller dictionaries. In an example, each dictionary can map each RDF term of an RDF datatype of the dataset to a corresponding index and vice versa; in other words, one dictionary per RDF datatype (https: / / www.w3.org / TR / rdf11-concepts / #section-Datatypes). Additionally, or alternatively, each dictionary can map each RDF term of an IRI (acronym for Internationalized Resource Identifier) prefix of the dataset to a corresponding index and vice versa; in other words, one dictionary for IRIs (https: / / www.w3.org / TR / rdf11-concepts / #dfn-iri) can be additionally sharded with one dictionary according to the prefix (https: / / www.w3.org / TR / rdf11-concepts / #dfn-namespace-prefix). In an example of using multiple dictionaries in a ROSA file, each dictionary can be a fixed-size dictionary or a variable-length dictionary. The case where the dictionary includes both a fixed-size dictionary and a variable-length dictionary will be discussed below.

[0083] It should be understood that the case where the dictionary includes two or more dictionaries does not change the operating principle of the dictionary; in other words, the examples of the dictionary discussed above still apply.

[0084] Reference Figure 1 150, 160, 170 of [reference], now discuss an example of the implementation of the dictionary. Any implementation of the dictionary can be envisioned when the dictionary maps each RDF term of the RDF dataset to a corresponding index. The dictionary consists of dictionary entries. The dictionary entries are defined by several elements: i) its type, e.g., fixed-size or variable-length; ii) the data corresponding to its type, e.g., the "main" vector (or part) or the "main" vector and the "blob" vector (or part). This has been discussed.

[0085] In the example, the dictionary can include entries with optional indices, e.g., for access control security. The adjustment part can exist as a matrix of unsigned integers (e.g., a matrix of unsigned 16-bit integers). The adjustment part is used to remap the internal dictionary index to the dataset index in the form of (64K)*(64K).

[0086] As represented on Figure 1 150 of [reference], the dictionary has N dictionary entries (Entry1DictionaryEntry…EntryN DictionaryEntry) that make up the data of the dictionary.

[0087] In the example, metadata 152 describing the dictionary can be added. The metadata can include a type part indicating the type of the dictionary layout (fixed-size or variable-length). The metadata can also include (if any) a codec part indicating the name of the codec used in the dictionary. The metadata can also include an index_remapping part and a DictionaryRange matrix, where the DictionaryRange matrix contains the type_map_id associated with the slice and the index remapping in the associated index for each 64K index. The purpose of the index remapping is to ensure that the dictionary has continuous indices. The metadata can also include an IRI mapping codec blob that contains a list of all IRI-encoded prefixes, which is referenced by uri_prefix_mapping in the dictionary footer. uri_prefix_mapping can be stored indifferently in the metadata part 152 or the dictionary footer 170.

[0088] As in Figure 1As represented on 160, the inverse dictionary includes its own data structure. The data structure of the inverse dictionary is similar to the data structure described above. In an example, the inverse dictionary can be attached to each (non-inverse) dictionary. For example, it can be implemented as an S-Tree<hash(e), index(e)>, where the S-Tree includes a hash, index pair for each element (e). The S-Tree data structure is discussed here: https: / / en.algorithmica.org / hpc / data-structures / s-tree / . In an example, the inverse dictionary can include a matrix of pairs <hash(e), index(e)> sorted by hash.

[0089] The dictionary can include a footer 170. The footer of the dictionary can be stored in a dedicated section 170 of the ROSA file, as represented in Figure 1 as shown.

[0090] The footer of the dictionary can include a matrix of DictionaryEntry referenced by the type_map field in the DictionaryFooter 170. For each of these dictionary entries, there is, for example, an unsigned integer of the 32-bit index type_id from the type section referenced by the DictionnaryFooter, and an unsigned integer of the 32-bit index codec_id from the codec section referenced by the DictionnaryFooter. The footer of the dictionary can also be a bitmask called a flag, indicating how to interpret the entries, thus retaining one or more of the following values:

[0091] - FIXED_SIZE_DICTIONARY, specifying that only the master is used because all blobs in the dictionary have the same length. In this case, the length field indicates the width;

[0092] - VARLENGTH_DICTIONARY, specifying that the master section contains the offset of the blob in the blob section, and the length specifies the width (4 or 8) of the index;

[0093] - SHARDED_BY_HASH and SHARDED_BY_URI, indicating that the dictionary is sharded by the hash of the decoded Blob and the RDFDataType function of the PrefixId, respectively, thus allowing the dictionary to be organized into sub-dictionaries;

[0094] - Optional values can be added, for example, if optional indexes are also used.

[0095] The DictionaryFooter 170 may also include a codec section that references a matrix of the codec names used, a type section that references a matrix of the type names used, a uri_prefix section with a list of prefixes, together with an index_remapping section that references a DictionaryRange matrix used to identify types based on an index or an internal index in the corresponding dictionary, a type_map section that references the matrix detailed above, and a count field that has the number of entries in the dictionary.

[0096] When the dictionary is a variable-length dictionary, "master" and "blob" vectors can be used. Generally, about two-thirds of the storage size is related to these two parts. In practice, it may not be the best choice to select only one layout, and for this reason, the dictionary can contain both fixed-size records and variable-size records.

[0097] As described, a fixed-size dictionary can be modeled using a contiguous matrix that is very simple to generate. Generating a variable-length table may be slightly more complex. As shown in Figure 1 150, the first part with an index is followed by an offset in the Blob part (labeled index#1BlobOffset in Figure 1 ), and the last part is followed by the end of the offset (labeled index#NN EndBlobOffset in Figure 1 ). The nth Blob size is calculated by subtracting the index#N BlobOffset from the index#N+1BlobOffset.

[0098] Now refer to Figure 1 180 in to discuss the data structure of the adjacency matrix. It should be reminded that for each indexed predicate of each graph, an adjacency matrix representing a set of tuples of the RDF dataset is obtained. As Figure 1 shown, for the Figure 1 of the RDF dataset, n predicates have been indexed 140, and each indexed predicate has an adjacency matrix.

[0099] As is known per se, an adjacency matrix is a binary matrix whose size is related to the number of subjects, predicates, and / or objects (i.e., a matrix with elements having two values (e.g., 0 and 1)), where a 1 bit means that the corresponding triple (e.g., of the corresponding predicate of the adjacency matrix) exists in the RDF dataset. The size of the adjacency matrix of a predicate can be the size of the subject times the object of the RDF tuples of the predicate.

[0100] Thus, an adjacency matrix is obtained for each predicate of each graph of the RDF dataset. As is known in the art, there are several known data structures for obtaining an adjacency matrix. Any data structure of the adjacency matrix can be used to obtain the read-only data structure of the present invention. The only requirement is that the read-only data structure can implement the concept of the adjacency matrix and the read-only data structure is suitable for a fixed memory layout size, which is the case for most read-only data structures.

[0101] In an example, the adjacency matrix for each indexed predicate can be obtained by performing a vertical partitioning on each predicate of the RDF database. As is known per se, vertical partitioning is a database design technique for partitioning RDF data into separate datasets or partitions based on the properties or predicates used in the triples. This technique is typically used to optimize the storage, retrieval, and querying of RDF data, especially in cases where certain subsets of the data are accessed more frequently than others.

[0102] Vertical partitioning can be performed as is known in the art, for example, as described in the following document: "Scalable semantic web data management using vertical partitioning" by Abadi, D.J. et al., Proceedings of the 33rd International Conference on Very Large Data Bases, September 2007, pages 411 - 422.

[0103] In an example, the vertical partitioning can be a standard K2 triple partitioning, as described in the following document: "Compressed vertical partitioning for efficient RDF management" by S. et al., Knowledge and Information Systems, 2015, Volume 44, Issue 2, pages 439 - 474.

[0104] In an example, the dk 2 tree, a dynamic variant of the k2 tree partitioning, as described in the following document: "Compressed Representation of Dynamic Binary Relations with Applications" by Nieves R. Brisaboaa et al., arXiv:1707.02769, can be used to generate the K2 triple partitioning.

[0105] In an example, vertical partitioning can be performed by using two static B+ trees, where the first static B+ tree represents the subject-to-object (S, O) correspondence and the second represents the object-to-subject (O, S) correspondence. Two S-trees are obtained; the S-Tree data structure is discussed here: https: / / en.algorithmica.org / hpc / data-structures / s-tree / .

[0106] In an example, vertical partitioning can be performed by using a static XAM tree carved from an XAMTree as a read-only XAMTree. The XAMTree is discussed in the document EP22306928.7 filed on December 16, 2022. This document is incorporated herein by reference. An example implementation of the XAMTree is discussed on pages 20 to 47 of the document EP22306928.7 and is incorporated herein by reference.

[0107] Still referring to Figure 1 180, now discuss the PredicateEntry record for each predicate. Once the adjacency matrix is obtained, it is necessary to identify the predicate it refers to. This is the purpose of the PredicateEntry record.

[0108] In an example, the PredicateEntry record can include:

[0109] - An unsigned integer (e.g., 64-bit), which is called an index and identifies the predicate name in the dictionary;

[0110] - An unsigned integer (e.g., 64-bit), which is called a count and represents the number of different SO (subject, object) pairs in the predicate. This statistic can be used by the SPARQL query engine to optimize the join ordering of the query. This field is optional in the PredicateEntry;

[0111] - An unsigned integer (e.g., 64-bit) offset, which indicates the position of the adjacency matrix data in the set of adjacency matrices of the graph in the RDF dataset;

[0112] - An unsigned integer (e.g., 8-bit), which is called a type and indicates how to interpret the matrix according to the method used to construct the adjacency matrix;

[0113] - An unsigned integer (e.g., 8-bit), which is called a flag. If a predicate has special semantics that can be used by a SPARQL query engine for query optimization, this optional field is used to flag the predicate. For example, it can participate in RDFS entailment (https: / / www.w3.org / TR / rdf11-mt / #rdfs-entailment) or be part of National Language Support (NLS, https: / / en.wikipedia.org / w / index.php?title=National_Language_Support&redirect=no). Other specialized semantics can be found;

[0114] - Optionally reserve 14 bytes for future use and / or padding purposes.

[0115] In the example, the read-only data structure can also include a list of indices 130 for one or more graphs of the RDF dataset. The graph part of the read-only data structure can include a list of indices of the graphs of the RDF dataset and also specify the GraphEntry matrix. Each graph of the index list corresponds to a GraphEntry. Each GraphEntry describes the corresponding graph in order to identify the graph names in the dictionaries 150, 160, 170.

[0116] In the example, a GraphEntry can include:

[0117] - An unsigned integer (e.g., 64-bit), which is called an index for identifying the graph name in the dictionary;

[0118] - An unsigned integer (e.g., 64-bit), which is called an offset for specifying the position of the PredicateEntry matrix discussed previously. This allows knowing where the predicate exists in the graph;

[0119] - An unsigned integer (e.g., 32-bit), which is called nb_predicate, specifying the number of PredicateEntries in the graph;

[0120] - Optionally reserve 32 bits for padding and / or possible future extensions.

[0121] In an example, the read-only data structure may further include a header 100 at the beginning of the read-only data structure (ROSA file) and a footer 190 at the end of the read-only data structure. The read-only data structure may be a binary file, and / or the header may be a fixed-size header, and / or the footer may be a fixed-size footer. The ROSA carving algorithm may be one-pass, so most references are backward. Alternatively, the references may be forward, but at the cost of multiple passes to create the ROSA file. Interestingly, backward references and thus random access in the memory-mapped file are consistent with modern computers using SSDs (solid state drives), where random access is inexpensive. It should be understood that SSDs are not a requirement of the present invention, and the ROSA file may be stored on any type of memory. The header is read when opening the ROSA file to identify the file content, and the footer at the end is read with backward references to navigate within the parts of the file. As is known in the art, navigation of the file is performed using backward or forward references.

[0122] In an example, the header 100 may include a basic record. In an example, the header may include a "magic" record, which is an unsigned integer (e.g., 32 bits) for uniquely identifying the type of the ROSA file. Alternatively or additionally, the header of the ROSA file may include a version record, which is an unsigned integer tag (e.g., 32 bits) for identifying the current version of the read-only data structure format. Alternatively or additionally, the header may include an "end" record, which is an unsigned integer (e.g., 64 bits); the "end" record is the offset from the end of the footer section (i.e., from the footer record position) to know where the ROSA file ends and thus where the memory layout ends.

[0123] Now discussing examples of the footer at the end in the case of backward references, it should be understood that the footer may be organized differently in other examples. In these examples, the footer record is more complex than the header. If minor changes have to be added in a future revision, fields are added backward to keep the footer upward compatible. Like the header record, the footer contains a magic version for identifying the footer and similarly contains an offset from the start position of Rosa.

[0124] In these examples of the footer, extensions may be added to a normal ROSA file. For example, a dedicated data structure may be added to index RDF data with respect to access control security. Markers may be used in the footer, and the markers indicate whether such optional extensions exist. According to these markers, additional records may exist in the footer to address them, e.g., a security entry record having a section representing the additional data structure and the entry count in the data structure.

[0125] In these examples at the footer, the footer may have the SHA-256 of the ROSA payload; the payload of the read-only data structure includes the part (and thus the data) between the end of the header and the start of the footer. SHA-256 is known per se. SHA-256 is an acronym for Secure Hash Algorithm 256-bit, which is a cryptographic hash function belonging to the SHA-2 hash function family. The output of SHA-256 is a fixed-size 256-bit (32-byte) hash value that is typically represented as a hexadecimal number. SHA-256 can be used to check the integrity of ROSA files. Additionally or alternatively, this SHA-256 can be used to provide a reliable way to detect the difference between an outdated SHA-256 output of a summary of an RDF dataset and an invalid summary of the RDF dataset. The summary of the RDF dataset can be constructed as discussed in the document EP22306798.4 filed on December 6, 2022.

[0126] In the examples, each entry of the read-only data structure can have a maximum size between 32 bits and 256 bits. The expression "entry of the read-only data structure" means that each individual element (or entry) in the read-only data structure cannot have a size exceeding 256 bits and the minimum size is 32 bits. In case the minimum size is not reached, padding can be used to conform to the minimum size. In the examples, the minimum size can be 32 bits and the maximum size can be 64 bits. In another example, each individual element (or entry) can be encoded with 32 bits. In yet another example, each individual element (or entry) can be encoded with 64 bits. The fixed size makes it easier to construct the read-only data structure.

[0127] In the examples, the read-only data structure (ROSA archive) may be limited to 2 32 entries. This improves the compression of the read-only data structure. In fact, instead of having a huge ROSA file, multiple files can be queried, thus eliminating the need for a huge ROSA file. Additionally, cloud storage access bandwidth is not conducive to downloading huge flat files. Therefore, a read-only data structure limited to 2 32 entries adapts to common cloud storage limitations.

[0128] It should be understood that without changing the present invention, the read-only data structure can be limited to 2 64 entries, or even 2 128 entries, or even 225 6 entries: This is a trade-off between consuming more memory and accepting more entries. As already discussed, the read-only data structure can be memory-mapped, and thus the overall size of the read-only data structure may be limited by the memory available on the system.

[0129] Figure 1The read-only data structure above shows an example representing all the parts discussed above. It should be understood that the read-only data structure according to the present invention is not limited to this specific example. In particular, the minimum requirements for the read-only data structure are the existence of an index list 130 of one or more graphs of the RDF dataset; for each graph in the RDF dataset, there is an index list 140 of the predicates of the graph, a dictionary 150 that maps each RDF term in the dataset to a corresponding index and vice versa 160; and for each predicate indexed in the predicate index list of each graph, there is an adjacency matrix 180 representing a set of tuples of the RDF dataset.

[0130] Now discuss the comparison between the read-only data structure of the present invention and Apache Parquet. It should be reminded that Apache Parquet belongs to the field of relational databases, while the present invention belongs to the field of graph databases. Even though the relational world and the graph world represent two different paradigms that have no relation except for the purpose of data storage, this comparison shows the improvement of the read-only data structure of the present invention. Parquet files are organized in a specific way such that it is an enumeration of column slabs. Each slab contains its own metadata, some statistics, the slab has its own dictionary, and can be manipulated by itself. It is intended to be manipulated without any joins, and if a table contains many columns and a query uses only a few columns, it is quite efficient. Parquet files are not indexed, but if the query clause criteria can evict a large number of slabs by using statistics, a large amount of data can be skipped (e.g., time range query and slabs are organized by time). For ROSA files, 80% of the RDF graph disk usage is in the storage of the dictionary, and SPARQL queries perform a large number of joins because the query engine uses dictionary indexing to perform joins without any access in the dictionary. Incidentally, most of the expensive time is spent when accessing the adjacency matrix, which only takes 20% of the memory usage. If a query uses only 10% of the predicates (in terms of size), then in the worst case, a query on the ROSA archive will use 2% of the corresponding size.

[0131] Refer to Figure 2 , now discuss the computer-implemented method for storing the modifications to be applied to the read-only data structure (i.e., for incrementally updating the ROSA file). As already mentioned, the present invention aims to provide a compressed, updatable data structure for fast read access to the RDF dataset; in other words, the RDF archive can be updated near real-time regardless of the size of the RDF knowledge graph.

[0132] The update of the ROSA file is performed through the W3C standard SPARQL query language. SPARQL 1.1 updates enable changes to be represented as intensional or extensional definitions. Intensional definitions assign meaning to terms by specifying the necessary and sufficient conditions for when a term is to be used. This is the opposite approach to extensional definitions, which define by listing all the things that belong to the definition; listing all the triples of the RDF dataset to be updated. Updating the RDF dataset covers writing or modifying or deleting at least one triple in the RDF dataset.

[0133] In the example, the CDC technique can be used to perform updates to a read-only data store. As is known per se, CDC (the acronym for Change Data Capture) is a technique for identifying and capturing changes made to data in a database. The main purpose of CDC is to identify and track changes so that other systems or processes can react accordingly. This is particularly useful in scenarios where multiple data sources must be kept in sync or changes must be propagated to downstream systems.

[0134] Figure 5 An example of a limitation of CDC for SPARQL updates is defined as a SPARQL 1.1 update that can only represent changes as extensional definitions. It should be understood that this is an example of how a modification of a triple can be represented.

[0135] Without changing the present invention, any other representation can be used as long as information is given about which triples in which graphs are to be modified. This is now discussed.

[0136] Now refer to Figure 2 , and obtain a first list of added and / or deleted tuples of the S200 RDF dataset. "Obtain a first list of added and / or deleted tuples of the RDF dataset" means, for this method, to provide a first list of added and / or deleted tuples of the RDF dataset.

[0137] Then, at S210, a first read-only data structure is calculated. The calculated first read-only data structure is defined as discussed above, which means that the (first) ROSA file is calculated based on the first list of added and / or deleted tuples to be applied to the RDF dataset. The (first) ROSA file that constructs the first list of added and / or deleted tuples of the RDF dataset is such that for each predicate of the added and / or deleted tuples of the RDF dataset in the first list, the first adjacency matrix represents the added tuples of the RDF database that include the same predicate. Additionally, the second adjacency matrix represents the deleted tuples of the RDF database that include the same predicate.

[0138] Then, at S220, the calculated read-only data structure is stored in a first file.

[0139] The stored first file is also referred to as "Delta ROSA". Delta ROSA is a representation of CDC information. For example, it can be the binary representation of the CDC representation. By construction, Delta ROSA is based on the ROSA file format. Thus, the Delta ROSA file inherits all the properties of the ROSA file, mainly the zero-copy read property and the compactness property. Generating Delta ROSA depends only on the size of the changes and not on the size of the ROSA file (i.e., the snapshot file of the RDF dataset, or the archive); Delta ROSA is independent of the ROSA file: the ROSA file does not need to be accessed when constructing Delta ROSA. This involves an obvious advantage in terms of the performance of generating the Delta ROSA file, because although the ROSA file compression scheme is adopted, the ROSA file can still be huge.

[0140] The Delta ROSA file is a new file format similar to ROSA. Like ROSA, it contains a list of read-only adjacency matrices and a read-only dictionary. The adjacency matrix in the ROSA file describes the set of triples (ROSA contains one adjacency matrix for each predicate and each graph), while Delta ROSA contains two adjacency matrices for each predicate and each graph. One adjacency matrix describes the set of triples created by the operations described in the CDC file, and the other adjacency matrix describes the set of triples deleted by the same operations. This data structure describes the total result of these operations: for example, if a triple is added and then deleted by an operation in the CDC file, then it does not appear in the Delta ROSA file at all.

[0141] After Figure 2 S220, a second list of added and / or deleted tuples of the RDF dataset can be obtained. A second read-only data structure is calculated, and this second read-only data structure is defined as discussed above, which means that the (second) ROSA file is calculated. Once stored, this (second) ROSA file will be referred to as the second Delta ROSA file. The (second) ROSA file is calculated based on the first list of added and / or deleted tuples to be applied to the RDF dataset and the second list of added and / or deleted tuples to be applied to the RDF dataset. For the first ROSA file calculated based on the first list of added and / or deleted tuples, this second ROSA file is constructed such that for each predicate of the added and / or deleted tuples of the RDF dataset in the first list and the second list, the first adjacency matrix represents the added tuples (of the first list and the second list) of the graph database including the same predicate. Additionally, the second adjacency matrix represents the deleted tuples (of the first list and the second list) of the graph database including the same predicate.

[0142] Then, the calculated read-only data structure is stored in a second file, thus forming a second incremental ROSA file.

[0143] Each time a new list of added and / or deleted tuples of the RDF dataset is obtained (e.g., the third list, the fourth list... the nth list of added and / or deleted tuples on the RDF dataset), the same process can be repeated. For example, the calculated nth incremental ROSA file is calculated based on the first list, the second list to the nth list of added and / or deleted tuples, thereby obtaining a modification chain. This means that in order to obtain the complete history before the nth list, only the nth incremental ROSA file is needed, which is important for performance.

[0144] It is not necessary to retain all the history of the RDF knowledge graph. Only a snapshot of the RDF knowledge graph is retained as a ROSA file, which has an incremental ROSA file chain that has been received and represents the modifications to be applied to the ROSA file.

[0145] Thanks to the adjacency matrix of the incremental ROSA file or the incremental ROSA file chain of the incremental ROSA file, it is easy to list the triples added or deleted through a series of operations. Then, the incremental ROSA file can be used to change the result of the BGP operation on the ROSA file (archive) by adding or deleting triples according to the content of the incremental ROSA file from the result. The use of BGP and incremental ROSA files will be discussed in more detail below.

[0146] In the example, the ROSA file (archive) can include two or more graphs, as already discussed. In this case, obtaining a list of added and / or deleted tuples of the RDF dataset can include obtaining a list of added and / or deleted tuples of the RDF database for each graph of the RDF set, and calculating the read-only data structure for the obtained list of added and / or deleted tuples can also include calculating a first adjacency matrix and a second adjacency matrix for each graph of the RDF set.

[0147] In the example, once the chain exceeds a threshold, a new full snapshot can be regenerated as the ROSA file (archive). Regenerating means calculating a new archive that includes an update to the incremental ROSA chain. After regeneration, the previous archive is swapped with the new archive. When updates have been entered into the "original" RDF knowledge dataset, the incremental ROSA chain can be suppressed. The threshold is a trade-off between the CPU cost for regeneration and the performance cost of read queries (the memory cost of incremental ROSAs and the CPU cost of using them) that can be parameterized according to application requirements. The threshold can also be based on the size of the ROSA file chain, the lifespan of the ROSA chain, etc.

[0148] In an example, a mapping can be calculated between the indices of the dictionaries of the RDF dataset and the indices of each dictionary of each list of added and / or deleted tuples of the RDF dataset. The calculated mapping is then stored in a file. This mapping allows avoiding looking up values in the dictionary of ROSA and then looking up the index of each triple in the dictionary of the incremental ROSA based on the result of future BGP operations. In fact, the BGP operations on the ROSA file return the indices of the RDF terms defined in the ROSA dictionary, rather than the values of the triples. Although the incremental ROSA file has a local dictionary, there is no reason for this local dictionary to contain the same indices as the dictionary of the ROSA file. For this purpose, a new mapping file format is defined. It contains data structures (e.g., an adjacency matrix) that allow efficient lookup of ROSA indices based on incremental ROSA indices and vice versa. Essentially, this mapping can only be calculated based on the ROSA file and the incremental ROSA file. This file is called DROSAM, which stands for Incremental ROSA Mapping. DROSAM can be calculated each time a new incremental ROSA file is calculated. Alternatively, DROSAM can be calculated each time a new incremental ROSA file is used.

[0149] Thus, the modifications to be applied to the archive of the RDF knowledge dataset are stored in one or more incremental ROSA files, and this set of incremental ROSA archives can form a chain of modifications to be applied to the archive. The mapping can be calculated and stored in a file called DROSAM, with the aim of improving the identification of ROSA indices based on incremental ROSA indices and the reverse operation, thereby enhancing the efficiency of future BGP operations.

[0150] Reference Figure 3 , now discuss the method of the present invention for performing SPARQL queries on read-only data structures. The query is executed on the ROSA file (archive) and the latest calculated incremental ROSA file.

[0151] At S300, a SPARQL query is obtained. SPARQL is a W3C recommendation for querying RDF data and is a graph matching language built on top of RDF tuple patterns. "RDF tuple pattern" means a pattern / template formed by an RDF graph. In other words, an RDF tuple pattern is an RDF graph (i.e., a set of RDF triples), where the subject, predicate, object, or label of the graph can be replaced by variables (for querying). SPARQL is a query language for RDF data that is capable of expressing queries across diverse data sources, whether the data is stored natively as RDF or is treated as RDF via middleware. SPARQL is mainly based on graph homomorphism. Graph homomorphism is a mapping between two graphs that respects their structure. More specifically, it is a function between the vertex sets of two graphs that maps adjacent vertices to adjacent vertices.

[0152] SPARQL includes the ability to query required and optional graph patterns and their conjunctions and disjunctions. SPARQL also supports aggregation, subqueries, negation, creating values through expressions, extensible value testing, and constraining queries by the source RDF graph. This means that SPARQL queries need to conform to eight different triple patterns possible in SPARQL. These eight triple patterns include (S, P, O), (S,?P, O), (S, P,?O), (S,?P,?O), (?S, P, O), (?S,?P, O), (?S, P,?O), and (?S,?P,?O), where the variable is preceded by the symbol?. A variable is the output of a triple pattern and can be the output of a SPARQL query. In some examples, a variable can be the output of a SELECT query. The output of a SPARQL query can be constructed using variables (e.g., aggregators such as sum). Variables in a query can be used to construct a graph homomorphism (i.e., the intermediate nodes necessary to obtain the result of the query). In some examples, the variables in a query may be neither used for output nor for intermediate results.

[0153] As is known per se, BGP (an acronym for basic graph pattern) refers to a simple form of RDF query pattern used in RDF query languages such as SPARQL (SPARQL Protocol and RDF Query Language). BGP is essentially a triple set pattern, and thus BGP allows expressing triple patterns to match in RDF data.

[0154] A SPARQL query can be obtained by a SPARQL search engine. As is known per se, a SPARQL query engine is a system designed to process SPARQL queries and retrieve information from an RDF data store. The SPARQL query engine takes a SPARQL request as input and processes it against the RDF data. It is responsible for interpreting the query, optimizing the query for efficient execution, and retrieving relevant information from the RDF graph.

[0155] After obtaining the SPARQL query S300 by the SPARQL query engine, the system obtains the S310 ROSA file and the incremental ROSA file (i.e., the incremental ROSA file chain). "Obtaining the ROSA file and the incremental ROSA file" for this method means providing a database. In an example, such obtaining or providing can mean or include downloading the ROSA file and the incremental ROSA file (e.g., from an online database or an online cloud), or retrieving the database from a memory (e.g., a persistent memory), and / or memory mapping the ROSA file and the incremental ROSA file.

[0156] Next, in S320, for each triple pattern in the query, find the tuples in the read-only data structure (ROSA file) that match the triple pattern of the query. This is performed by a query engine known in the art. As a result of the query, a first set of results is obtained.

[0157] Then, in S330, perform the following steps for each triple pattern of the query in the incremental ROSA file chain. First, determine whether tuples matching the triple pattern are to be deleted and / or added.

[0158] If the answer is negative, the result of the query is those found in the first set of results.

[0159] If the answer is positive, two cases may occur:

[0160] i) If it is inferred from the incremental ROSA file chain that tuples are to be deleted, remove the tuples from the first set of results; and / or

[0161] ii) If it is inferred from the incremental ROSA file chain that tuples are to be added, add the tuples to the first set of results.

[0162] Once step S330 is implemented, the first result includes the tuples to be added and no longer includes the tuples to be removed. Thus, a SPARQL query has been performed on the modified archive. The mapping stored in the DROSAM file can also be used to improve the efficiency of BGP operations (in terms of resource consumption and execution speed).

[0163] Figure 3 An example of a method performed by a so-called ROSACE system is shown, where ROSACE stands for ROSA Composition Engine. The ROSACE engine system is capable of defining an RDF data set as the union of the state of an RDF knowledge data set materialized by ROSA files and a change set applied to that state materialized by (the latest) incremental ROSA files, and then applying read statements to that data set, where the (latest) incremental ROSA file is a DROSAM file that can be calculated from the ROSA file and the incremental ROSA file when the incremental ROSA file is first used. The ROSACE working domain can be when a set of changes is small compared to the state of the RDF data, such that it is not necessary to regenerate the ROSA file when the changes are small compared to the size of the ROSA file.

[0164] Reference Figures 6 to 10, an example of the usage scenario of the ROSACE system is discussed. In this scenario, an error is introduced into an RDF dataset called ESCO, which is a multilingual classification of European skills, competences, and occupations. ESCO is part of the Europe 2020 strategy. Esco homepage (europa.eu). The error is that "Datalogue" is written in English instead of "Datalog", as Figure 6 shown. The goal of this scenario is to correct this problem without regenerating the ROSA file representing the ESCO dataset. Therefore, the starting point is the ROSA file containing version 1.0.8 of ESCO and the error, and two SPARQL requests:

[0165] - One counts the number of instances of skosxl:label; and

[0166] - One shows the value of the skosliteralForm of the uri..<…e8>.

[0167] As Figure 7 shown, a correction in CDC form is obtained. Then, as a result of importing the CDC, an incremental ROSA file is calculated, as Figure 8 shown. Next, on Figure 9 , an incremental ROSAM file is calculated to map the index of the incremental ROSA file to the index of the ROSA file. Finally, on Figure 10 , a hybrid dataset is obtained, where the hybrid dataset includes the ROSA file, the incremental ROSA file, and the incremental ROSAM file. Therefore, queries can be executed on the ROSA file, and the queries will take into account the modifications stored on the incremental ROSA file.

[0168] Figure 11 shows the results of the experiments conducted with the implementation of the present invention. It should be understood that the response time of the queries can vary depending on the queries that have been executed. Importantly, from these results, it is noted that the queries (Datalog queries and counting SKOSXL labels) are of the same order of magnitude.

[0169] One or more examples of read-only data structures can be combined. Examples of read-only data structures can apply RDF triples or RDF quads indifferently. For ease of explanation, the concept of quads is now discussed. In an example, the RDF tuples of an RDF dataset can be RDF quads. An RDF quad can be obtained by adding a graph label to an RDF triple. In such examples, the RDF tuples include an RDF graph. The W3C has published standard specifications to specify RDF quads (also known as N-Quads), for example, see "RDF 1.1 N-Quads, A line-based syntax for RDF datasets", W3C Recommendation, February 25, 2014. An RDF quad can be obtained by adding a graph name to an RDF triple. The graph name can be empty (i.e., for the default or unnamed graph) or an IRI (i.e., a graph IRI). In an example, the predicate of the graph can have the same IRI as the graph IRI. The graph name of each quad is the graph to which the quad belongs in the corresponding RDF dataset. As is known per se (for example, see https: / / www.w3.org / TR / rdf-sparql-query / #rdfDataset), an RDF dataset represents a collection of graphs. In all the examples discussed above, the term RDF tuple (or tuple) indifferently refers to an RDF triple or an RDF quad, unless it is explicitly mentioned to use one or the other. Thus, a knowledge RDF dataset can include RDF triples or RDF quads. When a knowledge RDF dataset includes RDF quads, a BGP can be a quad pattern that additionally takes the label of the graph as a query variable. In a specific example where this method obtains one or more adjacency matrices (each adjacency matrix is a representation of multiple sets of tuples), the subject and object can be queried on one adjacency matrix.

Claims

1. A computer-implemented read-only data structure for archiving a snapshot of an RDF dataset into a file, the RDF dataset comprising one or more graphs, the read-only data structure being directly queryable using all possible triple patterns of a SPARQL query engine, the data structure comprising: an indexed list of the one or more graphs of the RDF dataset; For each graph of the RDF dataset, an indexed list of predicates; a dictionary mapping each RDF term in the dataset to a corresponding index and vice versa; For each predicate indexed in the indexed list of predicates of each graph, an adjacency matrix of a set of tuples of the RDF dataset is represented, thereby obtaining a complete representation of the RDF dataset that can be directly queried using all possible triple patterns of a SPARQL query engine.

2. The computer-implemented read-only data structure of claim 1, wherein: When each value of the dictionary has a length equal to or less than a predetermined threshold, the dictionary has a fixed length, preferably, the predetermined threshold is 2, 4 or 8 bytes.

3. The computer-implemented read-only data structure of claim 2, wherein: When at least one value of the dictionary has a length greater than the predetermined threshold, the dictionary has a variable length, and each value having a length greater than the predetermined threshold is indexed into and stored in an overflow data structure that is part of the dictionary.

4. A computer-implemented read-only data structure according to any one of claims 1 to 3, at least part of the dictionary being encoded, preferably, the values ​​being encoded using the IEEE Standard for Floating-Point Arithmetic (IEEE 754-2019 or a previous version), or the prefixes of the values ​​being encoded using a hexadecimal key.

5. A computer-implemented read-only data structure according to any one of claims 1 to 4, wherein: The dictionary includes two or more dictionaries, preferably, each dictionary maps each RDF term of the RDF data type of the data set to a corresponding index and vice versa, or each dictionary maps each RDF term of the IRI prefix of the data set to a corresponding index and vice versa.

6. A computer-implemented read-only data structure according to any one of claims 1 to 5, wherein: The adjacency matrix of each indexed predicate is obtained by performing vertical partitioning for each predicate of the RDF database, preferably, the vertical partitioning is performed by a technique selected from the following: - Standard K2 triples, preferably based on dynamic K2 construction; - two static B+ trees, wherein the first static B+ tree expresses a subject to object (S, O) correspondence and the second static B+ tree expresses an object to subject (O, S) correspondence; - A static XAM tree carved from a XAMTree as a read-only XAMTree.

7. A computer-implemented read-only data structure according to any one of claims 1 to 6, wherein: Each entry of the data structure has a maximum size between 32 bits and 256 bits, preferably between 32 bits or 64 bits, more preferably the size of each entry is encoded with 32 bits or 64 bits.

8. A computer-implemented read-only data structure according to any one of claims 1 to 7, wherein: The read-only data structure is a binary file that includes the following items: - a header including an offset from the end of the end-of-page record location; - part, between the header and the footer; - the page footer, including an offset from the start of the header record location; The section stores: - said dictionary; - for each graph, an indexed list of Predicates for the graph and the adjacency matrix for each Predicate indexed in the indexed list of Predicates for the graph.

9. A computer-implemented method for storing modifications to be applied to a read-only data structure according to any one of claims 1 to 8, comprising: Obtaining a first list of added and / or deleted tuples of the RDF dataset; computing a first read-only data structure as defined in any one of claims 1 to 8, wherein, for each predicate of the added and / or deleted tuples of the RDF dataset of the first list: - a first adjacency matrix represents said added tuples of said RDF dataset comprising the same predicate; and / or - a second adjacency matrix represents the deleted tuples of the RDF dataset including the same predicate; as well as The calculated read-only data structure is stored in a first file.

10. The computer-implemented method of claim 9, further comprising: Obtaining a second list of added and / or deleted tuples of the RDF dataset; computing a second read-only data structure as defined in any one of claims 1 to 8, wherein, for each predicate of the added and / or deleted tuples of the RDF datasets of the first list and the second list: - a first adjacency matrix represents said added tuples of said RDF database comprising the same predicate; and / or - a second adjacency matrix representing said deleted tuples of said RDF database comprising the same predicate; as well as The calculated read-only data structure is stored in a second file, thereby forming a modification to be applied to the read-only data structure.

11. The computer-implemented method of claim 9 or 10, wherein: The obtaining further comprises: obtaining, for each graph of the RDF set, a list of added and / or deleted tuples of the RDF data set; and wherein the calculating comprises: calculating, for each graph of the RDF set, a first adjacency matrix and a second adjacency matrix.

12. The computer-implemented method of any one of claims 9 to 11, further comprising: computing a mapping between an index of the dictionary of the RDF dataset and an index of each dictionary of each list of the added and / or the deleted tuples of the RDF dataset; as well as Store the computed mapping to a file.

13. A computer-implemented method for performing a SPARQL query on a read-only data structure stored in a file according to any one of claims 1 to 8 and a computed read-only data structure stored in a first file according to any one of claims 9 to 12, comprising: Obtaining, by a SPARQL query engine, a SPARQL request, wherein the SPARQL query includes at least one triple pattern; Obtaining a read-only data structure according to any one of claims 1 to 8 and a calculated read-only data structure stored in a file according to any one of claims 9 to 12; For each triple pattern of the query, find a tuple that matches the triple pattern of the query in the data structure according to any one of claims 1 to 8, thereby obtaining a first set of results; as well as for each triple pattern of the query, in a computed read-only data structure in a file according to any one of claims 9 to 12: - determining whether to delete and / or add tuples matching the triple pattern; - if a tuple is to be deleted, removing the tuple from the first set of results; - If a tuple is to be added, adding the tuple to the first set of results.

14. A computer program comprising instructions for executing the method according to any one of claims 9 to 12 and / or the method according to claim 13.

15. A system comprising a processor connected to a memory having recorded thereon a computer program according to claim 14.