Read-only data structure for archiving snapshot of RDF dataset to file

The proposed read-only data structure for RDF datasets, featuring an index list, dictionaries, and adjacency matrices, addresses inefficiencies in archiving and querying large RDF graphs, providing fast and efficient access with real-time updates.

JP2025108374APending Publication Date: 2025-07-23DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024217123
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-11
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing solutions for archiving and querying large RDF knowledge graphs are inefficient, particularly in handling incremental updates and supporting fast read access while maintaining compression, as they often compromise on read/write performance or scalability.

Method used

A read-only data structure that includes an index list of graphs, dictionaries mapping RDF terms to indices, and adjacency matrices for each predicate, enabling direct querying of all triple patterns using a SPARQL engine, with support for incremental updates.

Benefits of technology

Enables fast, resource-efficient access to large, compressed RDF archives with real-time incremental updates, supporting all possible triple patterns for SPARQL queries without significant overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108374000001_ABST
    Figure 2025108374000001_ABST
Patent Text Reader

Abstract

To provide a read-only data structure and a computer-implemented method for archiving a snapshot of an RDF dataset to a file.SOLUTION: An RDF dataset includes one or more graphs. A data structure includes: an index list of one or more graphs in an RDF dataset; an index list of predicates for each graph in the RDF dataset; a dictionary mapping each RDF term in the dataset to respective indexes; and an adjacency matrix representing the group of tuples in the RDF dataset for each predicate conversely indexed in index list of predicates in each graph. This provides a complete representation of the RDF dataset that can be directly queried using all possible triple patterns in a SPARQL query engine.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly, to a read-only data structure for archiving a snapshot of an RDF dataset into a file, modifying the snapshot, and querying the snapshot or the modified snapshot.

Background Art

[0002] The RDF specification was published by the W3C and represents information as a graph using triples. The publication is disclosed at the following link (https: / / www.w3.org / TR / rdf11-concepts / ). The core structure of the abstract syntax is a set of triples, each consisting of a subject, a predicate, and an object. Such a set of triples is called an RDF graph. An RDF graph can be visualized as a diagram of nodes and directed arcs, and each triple is represented as a node-arc-node link.

[0003] SPARQL is a query language for RDF data. See this link for version 1.1 of SPARQL (https: / / www.w3.org / TR / SPARQL11-query / ). RDF is a directed graph data format for representing information on the web. This specification defines the syntax and semantics of the SPARQL query language for RDF. SPARQL can be used to express queries across a variety of data sources, whether the data is stored natively as RDF or presented as RDF via middleware. SPARQL has the ability to query essential and optional graph patterns and their conjunctions and disjunctions. SPARQL also supports aggregation, subqueries, negation, value creation by expressions, test of extensible values, and constraint of queries by the source RDF graph. The result of a SPARQL query is a result set or an RDF graph.

[0004] An RDF knowledge graph is formed by connecting triples to create a network or graph of information. It can grow extremely large, up to hundreds of billions, such as Uniprot (100 billion triples), Wikidata (17 billion triples), 3DSSIEM (500 billion triples), etc. Therefore, it is impossible or very costly to store the entire graph in memory.

[0005] According to experiments, most of these RDF knowledge graphs have been shown to be read-only or at least rarely modified and not highly concurrent. In the world of relational databases, this is called an OLAP use case. OLAP provides fast and interactive access to aggregated, summarized, and computed data, enabling users to easily explore and analyze complex datasets. OLAP is short for Online Analytical Processing. See, for example, https: / / en.wikipedia.org / wiki / Online_analytical_processing.

[0006] In real-world scenarios such as SQL OLAP as seen in the well-known TPC-H benchmark (see https: / / www.tpc.org / tpch / default5.asp), there are still small write modifications that occur simultaneously with read queries. A typical use case in the industry is that a large amount of write requests can be sent to a large RDF knowledge graph storing chemical formulas, but most of the stored formulas do not need to be modified.

[0007] In the field of relational databases, Apache Parquet (https: / / parquet.apache.org / ) is well-known. Apache Parquet is an open-source, column-oriented data file format designed for efficient data storage and retrieval. It provides efficient data compression and encoding methods, improving performance to handle large amounts of complex data. Apache Parquet is built to support very efficient compression and encoding methods. Multiple projects have demonstrated the impact on performance by applying appropriate compression and encoding methods to data. In Parquet, the compression method can be specified on a per-column basis, and it has future-proofing in that the encoding can be added whenever an encoding is devised and implemented. The target of Apache Parquet is relational data, not graph data. However, this highlights the need for very efficient data compression and encoding schemes.

[0008] In the field of relational databases, Apache Arrow (https: / / arrow.apache.org / docs / index.html) is also well-known. Compared to Apache Parquet, no compression scheme is used, and it relies on zero-copy shared memory.

[0009] In the field of graph databases, RDFHDT (Header-Dictionary-Triples) is a W3C submission (http: / / www.w3.org / Submission / 2011 / 03 / ) of a binary format (also called the HDT file format) for the publication and exchange of large-scale RDF data. Its compact representation enables storing RDF in less space while providing direct access to the stored information. This is achieved by representing the RDF graph with three main components: a header, a dictionary, and triples. The header contains extensible metadata necessary to describe the RDF dataset and its details. The dictionary organizes the vocabulary of strings present in the RDF graph by assigning numerical IDs to each distinct string. The triples component depicts the underlying graph structure in a compressed form.

[0010] HDT has not passed beyond the submission stage. HDT was originally devised as a binary serialization format for transferring RDF data, but due to its compactness, it is being used as an RDF compressor. HDT, which was just a binary format, could only perform basic searches, but later it was extended by various variations, changing the trade-off of read performance / write performance / used memory.

[0011] HDT++ is a variant of HDT that improves the compression rate of HDT but does not retain the search functionality of the original HDT.

[0012] iHDT++ is an extension of HDT, as discussed in Hernandez-Illera, A., Martinez-Prieto, M.A., Fernandez, J.D. and Farina, A., 2020. iHDT++: improving HDT for SPARQL triple pattern resolution. Journal of Intelligent & Fuzzy Systems, 39(2), pp.2249-2261, and Hernandez-Illera, A., Martinez-Prieto, M.A. and Fernandez, J.D., 2015, April. Serializing RDF in compressed space. In 2015 Data Compression Conference (pp. 363-372).

[0013] HDTFoQ is another extension of HDT, as discussed in Martinez-Prieto, M.A., Arias Gallego, M. and Fernandez, J.D., 2012. Exchange and consumption of huge RDF data. HDTFoQ is a variation regarding complete SPARQL queries, but its main drawback is that the reading algorithm is costly and does not scale to large RDF knowledge graphs (e.g., an RDF knowledge graph composed of billions of triples).

[0014] Both HDT++ and iHDT++ assume an "object-like" organization. For example, each subject must exist only once within a "family" of predicates. The "basic" HDT does not support all access patterns, only by subject. This is a strong limitation.

[0015] Therefore, HDT was designed as a binary format for the large-scale publication and exchange of RDF data, rather than for answering SPARQL queries. That's why there is a major premise that all triples must be composed by the subject. Subsequently, attempts were made to introduce HDT into SPARQL, but there were limitations in that it was not compatible with querying very large RDF knowledge datasets (e.g., billions of triples). HDT significantly optimizes the memory / storage usage, but at the cost of reduced read (and write / generation) performance. Regarding the so-called RUM conjecture (Read Update Memory), it is described in the document M. Athanassoulis, M. S. Kester, L. M. Maas, R. Stoica, S. Idreos, A. Ailamaki, M. Callaghan. Designing Access Methods: The RUM Conjecture. In proceedings of the 19th International Conference on Extending Database Technology (EDBT), March 15-18, 2016 - Bordeaux, France. This document concludes that it is impossible to create the ultimate access method because certain optimization and design choices are mutually exclusive. Another trade-off is needed. It is a format that enables fast read access to SPARQL queries that are not highly compressed but are not limited to "subject-by-subject search".

[0016] Known solutions in the art propose data compression techniques (https: / / en.wikipedia.org / wiki / Data_compression) to reduce the size of the storage (e.g., on memory or disk) of RDF graph databases. Such compression techniques help reduce storage costs (such as hardware costs) and the environmental footprint in various applications, especially in cloud deployments.

[0017] ALVAREZ-GARCIA, S., et al., “Compressed vertical partitioning for efficient RDF management.”, Knowledge and Information Systems, 2015, vol. 44, no 2, p. 439-474 discloses an RDF indexing technique that supports SPARQL solutions in compressed space. The disclosed technique, called k2-triple, vertically partitions the dataset using predicates and divides it into discontinuous subsets of one pair (subject, object) per predicate. These subsets are represented as a subject × object binary matrix, where 1 bit means that the corresponding triple exists in the dataset. As a result of this model, the matrix becomes very sparse and is efficiently compressed using a k2-tree.

[0018] The literature CHATTERJEE, A., et al., “Exploiting topological structures for graph compression based on quadtrees.”, In: 2016 Second International Conference on Research in Computational Intelligence and Communication Networks (ICRCICN). IEEE, 2016. p. 192-197 discloses an algorithm that performs compression based on quadtrees considering the characteristics of graphs. Furthermore, a technique for executing queries on the compressed data itself while compressing the data is introduced and explained in detail.

[0019] The document NELSON, M., et al., “Queryable compression on streaming social networks.”, In: 2017 IEEE International Conference on Big Data (Big Data). IEEE, 2017. p. 988-993 discloses the use of a novel data structure for streaming graphs. This data structure is based on an indexed array of compressed binary trees and constructs the graph directly without using a temporary storage structure. The data structure provides fast access methods for edge existence (whether an edge exists between two nodes), neighborhood queries (a list of the neighborhood of a node), and streaming operations (addition / removal of nodes / edges).

[0020] The document Pelgrin, O., Taelman, R., Galarraga, L. and Hose, K., GLENDA: Querying over RDF Archives with SPARQL states that "RDF datasets on the web are constantly evolving." The simplest way to track the history of RDF data is to save each revision of the dataset as an independent copy. However, this can be an exorbitant burden for large-scale RDF datasets with a long history. This observation has led to the emergence of a more efficient way to manage, i.e., store and query, large-scale RDF archives.

[0021] The purpose of an RDF archive is different from that of an HDT only. It is to track the history of RDF data by tracking the change history of RDF graphs and make that history queryable. As seen in Pelgrin, O., Taelman, R., Galarraga, L. and Hose, K., 2023, February. Scaling Large RDF Archives To Very Long Histories. In 2023 IEEE 17th International Conference on Semantic Computing (ICSC) (pp. 41-48), the applications are diverse. For example, version management, mining temporal and revision patterns, temporal data analysis, or studying past queries and the evolution of a given knowledge domain. An RDF archive needs to deal with very large knowledge graphs updated by batches of millions of triples or continuous updates. The techniques used to capture, store, and query the long revision history of very large RDF graphs are a combination of snapshots of the dataset and a sequence of aggregated change sets called delta chains and the literature. The proposed solution is based on multiple delta chains that insert full snapshots between change sets when the history becomes too long. All of these refer to only one dataset, the dataset whose history is being tracked.

[0022] In the literature "Scaling Large RDF Archives To Very Long Histories", "Snapshots are stored as HDT files, while the delta chain is materialized in two stores. One is for additions and the other is for deletions. Each store is composed of three indexes in different triple component orders of SPO, OSP, and POS, and is implemented as a B+ tree. The keys of these indexes are individual triples linked to version metadata, that is, the revisions in which the triples exist and the revisions in which they do not. In addition to the change store, there are indexes for the counts of additions and deletions for all possible triple patterns, such as <?s,?p,?o> and <?s,cityIn,?o>. As is common in RDF stores, RDF terms are mapped to an integer space to achieve efficient storage and retrieval. There are a snapshot dictionary (using HDT) and a delta chain dictionary. Therefore, in our multi-snapshot approach, we use a dictionary of D×2 (possibly disconnected), where D is the number of delta chains in the archive. The ingestion routine depends on whether the revision is saved as an aggregated delta or as a snapshot. At revision 0, the ingest routine receives the complete RDF graph as input to build the first snapshot. In subsequent revisions, it takes a standard change set as input and uses OSTRICH to build an aggregated change set in the form of us,k. When the snapshot policy decides to materialize revision s' as a snapshot, an aggregated change set is used to calculate the snapshot efficiently.

[0023] The approach of the document "Scaling Large RDF Archives To Very Long Histories" has two main drawbacks. First, to materialize the delta chain, it is necessary to query the snapshot dictionary, which is not suitable for generating in a near-real-time state. Second, the snapshot is based on HDT and, as mentioned above, is not suitable for efficiently answering SPARQL queries on large datasets.

[0024] In such a situation, there remains a need for fast read access to a very large and compressed RDF archive that can be incrementally updated, such as adding or deleting triples. This read access must efficiently answer all possible triple patterns for SPARQL queries. Incremental updates must be almost in real-time regardless of the size of the RDF knowledge graph.

Summary of the Invention

[0025] Therefore, a computer-implemented read-only data structure for archiving snapshots of an RDF dataset into a file, where the RDF dataset has one or more graphs, and the read-only data structure can be directly queried using all possible triple patterns of a SPARQL query engine, the data structure including an index list of one or more graphs of the RDF dataset, an index list of predicates for each graph of the RDF dataset, a dictionary that associates each RDF term of the dataset with its respective index and vice versa, and for each predicate indexed in the index list of predicates of each graph, an adjacency matrix representing a group of tuples of the RDF dataset, whereby a complete representation of the RDF dataset that can be directly queried using all possible triple patterns of a SPARQL query engine is obtained, is provided.

[0026] This method may include one or more of the following. · When each value of the dictionary has a length equal to or less than a predetermined threshold, the dictionary has a fixed length. Preferably, the predetermined threshold is 2, 4, or 8 bytes. · When at least one value of the dictionary has a length greater than the predetermined threshold, the dictionary has a variable length, and each value having a length greater than the predetermined threshold is indexed and stored in an overflow data structure, and the overflow data structure is part of the dictionary. · At least a part of the dictionary is encoded. Preferably, the values are encoded using the IEEE floating-point arithmetic standard (IEEE 754-2019 or a previous version), or the prefix of the values is encoded using a hexadecimal key. · The dictionary consists of two or more dictionaries. Preferably, each dictionary associates each RDF term of the RDF data type of the dataset with a respective index and vice versa, or each dictionary associates each RDF term of the IRI prefix of the dataset with a respective index and vice versa. · The adjacency matrix for each indexed predicate is obtained by performing vertical partitioning for each predicate of the RDF database. Preferably, the vertical partitioning is performed by a technique selected from the following. Standard K2 - triples, preferably those made from dynamic K2. Two static B+ - trees. The first static B+ - tree represents the subject - object (S, O) correspondence, and the second static B+ - tree represents the object - subject (O, S) correspondence. A static XAM tree obtained by loading an XAM tree and imprinting it as a read - only XAM tree. · Each entry of the data structure has a maximum size configured between 32 bits and 256 bits, preferably between 32 bits and 64 bits. More preferably, the size of each entry is encoded in 32 bits or 64 bits. · The read-only data structure is a binary file including the following: a header composed of an offset to the end of the recording position of the footer, a section composed between the header and the footer, the footer composed of an offset with respect to the recording start position of the header, and the section stores, for each of the dictionary and each graph, an index list of the predicates of the graph and an adjacency matrix of each predicate indexed by the index list of the predicates of the graph.

[0027] Furthermore, a computer-implemented method for storing a modification applied to any of the above read-only data structures, the method comprising: obtaining a first list of added and / or deleted tuples of an RDF data set; calculating a first read-only data structure defined according to any one of claims 1 to 8, the calculating being performed for each predicate of the added and / or deleted tuples of the RDF data set of the first list, wherein the first adjacency matrix represents the added tuples of the RDF data set consisting of the same predicate, and / or the second adjacency matrix represents the deleted tuples of the RDF data set consisting of the same predicate; and storing the calculated read-only data structure in a first file.

[0028] This method may further include the following. · Obtaining a second list of added and / or deleted tuples of the RDF data set; calculating a second read-only data structure defined as above, the calculating being performed for each predicate of the added and / or deleted tuples of the RDF data set of the first and second lists, wherein the first adjacency matrix represents the added tuples of the RDF database consisting of the same predicate, and / or the second adjacency matrix represents the deleted tuples of the RDF database consisting of the same predicate; and storing the calculated read-only data structure in a second file, thereby forming a modification applied to the read-only data structure. · The obtaining step further includes, for each graph of the RDF set, obtaining a list of added and / or deleted tuples of the RDF dataset, and the calculating step includes, for each graph of the RDF set, calculating a first adjacency matrix and a second adjacency matrix. · Calculating a mapping between the index of the dictionary of the RDF dataset and the index of each dictionary of each list of added and / or deleted tuples of the RDF dataset, and saving the calculated mapping to a file.

[0029] Furthermore, a computer-implemented method for SPARQL queries on the read-only data structure stored in the above file and the calculated read-only data structure stored in the above first file, the method comprising: obtaining a SPARQL query by a SPARQL query engine, the SPARQL query including at least one triple pattern; obtaining the read-only data structure according to the above; storing the calculated read-only data structure in a file according to the above; for each triple pattern of the query, finding, according to any one of claims 1 to 8, a tuple in the data structure that responds to the triple pattern of the query, thereby obtaining a first set of results; for each triple pattern of the query, converting the calculated read-only data structure into a file according to the method described in any one of claims 9 to 12; determining whether a tuple that answers the triple pattern is deleted or added; if the tuple is deleted, deleting the tuple from the first result set; and if the tuple is added, adding the tuple to the first result set. A computer-implemented method is provided.

[0030] Furthermore, a computer program including the above method and / or instructions for executing the above method is provided.

[0031] Furthermore, there is provided a computer-readable recording medium storing the above computer program.

[0032] Furthermore, there is provided a system including a processor coupled to a memory, wherein the computer program is recorded in the memory.

[0033] Furthermore, there is provided a device including a data storage medium storing a computer program. The device may form, or function as, a non-transitory computer-readable medium in, for example, SaaS (Software as a Service) or other servers, or cloud-based platforms, or the like. Alternatively, the device may include a processor coupled to the data storage medium. Thus, the device may form all or part of a computer system (for example, the device is a subsystem of an overall system). The system may further include a graphical user interface coupled to the processor.

Brief Description of the Drawings

[0034]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0035] Referring to FIG. 1, an object of the present invention is to propose a read-only data structure for archiving a snapshot of an RDF dataset to a file. The RDF dataset includes one or more graphs. A read-only data structure that can be directly queried using all possible triple patterns of a SPARQL query engine. The data structure includes an index list (130) of one or more graphs of the RDF dataset. For each graph of the RDF dataset, the data structure further includes an index list (140) of predicates. The data structure also includes dictionaries (150, 160) that map each RDF term of the dataset to its respective index and vice versa. Further, the data structure includes an adjacency matrix (170) representing a group of tuples of the RDF dataset for each predicate indexed in the index list of predicates of each graph, whereby a complete representation of the RDF dataset that can be directly queried using all possible triple patterns of the SPARQL query engine is obtained.

[0036] Such a data structure provides a read-only data structure that improves the archiving of snapshots of RDF datasets to files. "Improving the archive" means that the data structure stores the RDF dataset in a compressed format and reduces the overall size of the RDF dataset. In fact, this archive uses a concise data structure for the representation of graph data and can be stored in a very efficient way on disk. Recall that a concise data structure is a data structure that uses "near" the information-theoretic lower bound space, unlike other compressed representations, and still enables efficient query operations. Therefore, the use of an index list of predicates for each graph, a dictionary and an adjacency matrix for each predicate, contributes to reducing the size of the entire archive. Furthermore, compression can be improved by using a dedicated encoder for one or more of these concise data structures. For example, a dictionary (or dictionaries) can be compressed.

[0037] Furthermore, the archive provided by the read-only data structure defines an RDF file / memory format for efficient search and data storage. By supporting zero-copy reads, it enables fast and resource-efficient data access without the overhead of serialization. The archive is generated only once and can be read back as many times as needed. The archive is self-describing and can be mapped in one go. The archive file describes the RDF dataset and can be seen as a kind of indexed triG file. For triG files, see https: / / www.w3.org / TR / trig / . Thus, the archive provided by the read-only data structure can be used for long-term persistence and, due to its structure, can be read without converting the data (e.g., without the need for serialization).

[0038] Furthermore, the archive's data structure is based on state-of-the-art technology for graph representations that enable full support for SPARQL. Therefore, read access can efficiently answer eight different triple patterns possible with SPARQL queries, i.e., all triple patterns. Such eight triple patterns are (S,P,O), (S,?P,O), (S,P,?O), (S,?P,?O), (?S,P,O), (?S,?P,O), (?S,P,?O), and (?S,?P,?O), where the symbol? is prefixed to variables in these patterns. A variable is the output of a triple pattern and may also be the output of a SPARQL query.

[0039] Finally, the read-only data structure can consist of one or more graphs. As a result, the read-only data structure according to the present invention complies with the provisions of the RDF 1.1 recommendation (see https: / / www.w3.org / TR / rdf11-concepts / ). Not only does it comply with the recommendation, but the read-only data structure can be implemented and used in the current RDF dataset which is a multi-graph. In other words, the read-only data structure addresses the industrial problem of archiving read-only multi-graph RDF datasets.

[0040] Next, referring to FIG. 2, a computer-implemented method for storing modifications applied to the read-only data structure described above is also proposed. As introduced above, the read-only data structure can incorporate incremental updates almost in real time. For that purpose, the method for storing modifications applied to the read-only data structure according to the present invention includes step S200 of obtaining a first list of added and / or deleted tuples of the RDF dataset. The tuples are the modifications applied to the archive. Next, the first read-only data structure is calculated (step S210). This first read-only data structure is the same as that described above in this specification. However, when constructing the first read-only data structure, for each predicate of the added and / or deleted tuples of the RDF dataset in the first list, the first adjacency matrix represents the added tuples of the graph database consisting of the same predicate, and / or the second adjacency matrix represents the deleted tuples of the graph database consisting of the same predicate. Also, the method includes step S220 of storing the calculated read-only data structure in a first file. A snapshot of the first list of added and / or deleted tuples of the RDF dataset is archived in the file that is the read-only data structure.

[0041] In such a way, the read-only data structure of the present invention can be updated. The update depends on the calculation of the first read-only data structure that reflects the changes (add and / or delete tuples) to be applied to the read-only data structure (archive). This first read-only data structure is functionally similar to the delta file of the RDF archive. One or more delta files follow the archive (snapshot of the RDF dataset file), and one or more delta files form a delta chain. The snapshot and the delta chain are made of the same read-only data structure. Interestingly, one or more delta files inherit the characteristics of the read-only data structure described above. It should be noted that the read-only data structure enables fast read access to SPARQL queries. Furthermore, the calculation of the delta file is almost real-time because it does not require access to the snapshot file and does not depend on the size of the RDF dataset (full graph).

[0042] Next, referring to FIG. 3, a computer-implemented method for executing a SPARQL query against a read-only data structure stored in a file (archive) and a calculated read-only data structure (first file or delta file) is also proposed. This method includes step S300 of obtaining a SPARQL query by a SPARQL query engine, where the SPARQL query includes at least one triple pattern. Next, a snapshot of the RDF dataset is archived to a file, and a read-only data structure for storing the calculated read-only data structure in the first file is obtained (step S310). In other words, a snapshot file of the RDF dataset is obtained, and a delta file which is the read-only data structure according to the present invention is also obtained. For each triple pattern of the query (step S320), a tuple responding to the triple pattern of the query is searched in the read-only data structure (snapshot file of the RDF dataset), thereby obtaining a first set of results. For each triple pattern of the query (step S330), the following steps are executed from the calculated read-only data structure to the first file: i) Determine whether a tuple answering the triple pattern is to be deleted and / or added. ii) If a tuple is deleted, delete the tuple from the first result set. iii) If a tuple is added, add the tuple to the first result set.

[0043] The combination of the snapshot file of the RDF dataset and the delta file is provided to the SPARQL query engine in one unified logical view. Thus, fast read access is enabled for a compressed RDF archive that is updated incrementally almost in real time.

[0044] The read-only data structure and the present method are implemented on a computer. This means that the steps of the read-only data structure and the present method (or substantially all steps) are executed / used by at least one computer, or any system. Thus, the steps of the present method are executed by a computer, in some cases completely automatically or semi-automatically. In an example, the trigger for at least some of the steps of the present method is executed by the interaction between the user and the computer. The required level of interaction between the user and the computer may depend on the balance between the expected level of automation and the need to fulfill the user's wishes. In an example, this level can be user-defined and / or pre-defined.

[0045] A typical example of the computer implementation of the method is to execute the method on a system adapted for this purpose. This system may include a processor coupled to a memory and a graphical user interface (GUI), and the memory stores a computer program including instructions for executing the present method. Also, the memory can store a database, for example, by using a memory map. The memory is any hardware adapted for such storage and may in some cases be composed of a plurality of physically different parts (for example, one for the program and one for the database).

[0046] Figure 4 shows an example of a system, which is a client computer system, for example, the user's workstation. The system can be used to construct, maintain, store, and use a read-only data structure, and / or the system can be used to execute the present method for storing modifications applied to the read-only data structure, and / or the present method of SPARQL query on the read-only data structure and the calculated read-only data structure stored in the first file.

[0047] The client computer of this embodiment is composed of a central processing unit (CPU) 1010 connected to an internal communication bus 1000 and a random access memory (RAM) 1070 also connected to the bus. The client computer further includes a graphical processing unit (GPU) 1110 related to 1100 connected to the bus. The video RAM 1100 is also known as a frame buffer in the art. The mass storage device controller 1020 manages access to mass memory devices such as the hard drive 1030. A large number of memory devices suitable for specifically embodying computer program instructions and data include semiconductor memory devices such as EPROM, EEPROM, flash memory devices, magnetic disks such as built-in hard disks and removable disks, magneto-optical disks, etc., all forms of non-volatile memory. Any of the above may be complemented or incorporated by a specially designed ASIC (application-specific integrated circuit). The network adapter 1050 manages access to the network 1060. The client computer can also include a contact device 1090 such as a cursor control device and a keyboard. A cursor control device is used in the client computer so that a cursor can be selectively placed at any position on the display 1080. Also, various commands can be selected and control signals can be input by the cursor operating device. The cursor control device includes a number of signal generators for inputting control signals to the system. Generally, the cursor control device is a mouse, and the buttons of the mouse are used to generate signals. Alternatively or additionally, the client computer system may include a touch pad and / or a touch screen.

[0048] A computer program may include instructions executable by a computer system, the instructions including means for causing the system to perform the execution of the present method for constructing, maintaining, storing, using, and / or storing modifications applied to a read-only data structure, and / or for causing the system to execute SPARQL queries on a read-only data structure and a calculated read-only data structure stored in a first file.

[0049] The program can be recorded on any data storage medium including the system's memory. The program can be implemented, for example, in digital electronic circuitry, computer hardware, firmware, software, or combinations thereof. The program can be implemented as a product implemented in a machine-readable storage device for execution by, for example, a programmable processor. Method steps can be executed by a programmable processor executing a program of instructions that operate on input data and generate output to perform the functions of the present method. Accordingly, the processor can be programmably coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to transmit data and instructions. An application program can be implemented in a high-level procedural or object-oriented programming language, or, if desired, in assembly or machine language. In any case, the language can be a compiled language or an interpreted language. This program can be a full-installation program or an update program. When this program is applied to the system, instructions for performing the present method are displayed in any case. The computer program may be stored and executed on a server in a cloud computing environment. The server communicates with one or more clients via a network. In such a case, the present method is executed on the cloud computing environment by a processing device executing instructions configured by the program.

[0050] Next, an example of a read-only data structure will be described with reference to FIG. 1 showing an example thereof. In the following description, the read-only data structure is equivalently referred to as a ROSA file, that is, an archive. ROSA is an abbreviation of Read-Only Semantic Archive. A ROSA file (or ROSA archive) consists of only one file. The ROSA file is composed of several sections, and the set of these sections defines the configuration of the ROSA file. The ROSA file shown in FIG. 1 shows a non-limiting example, and it should be understood that the positions of each part can be changed without changing the present invention. In another example of FIG. 1, section 130 may be arranged between sections 180 and 190, and section 140 may be arranged immediately after section 180. Therefore, when generating a ROSA file, section 140 is generated after each set of adjacency matrices (Graph1...Graphn) related to the graph of section 180, and section 130 is generated immediately before generating the footer section 190. According to this rule, one-pass generation of the ROSA file becomes possible.

[0051] By splitting the ROSA file into sections, a memory map can be used to have the same layout on disk and in memory, and the file can be memory-mapped when opened. Memory mapping is well-known, see https: / / en.wikipedia.org / wiki / Memory_map. It should be understood that memory mapping is an optimization and other memory layouts may be used without changing the present invention. When the ROSA file is memory-mapped, only some handles and the necessary internal data structures, such as the codec list, graph index list, type of map, etc., are loaded. Although the ROSA file is intended to be mapped, it may be interesting that this kind of file can be operated on by streaming to perform filtering. For example, deleting predicates or graphs. It is possible to create tools to convert the ROSA archive and add or remove sections. It is very easy to write a utility to delete predicates and graphs and generate a new ROSA archive. This can be done without changing the present invention.

[0052] In an embodiment, the ROSA file includes at least the following data structures. i) An index list 130 of one or more graphs of the RDF dataset. ii) For each graph of the RDF dataset, an index list 140 of the predicates of that graph. iii) A dictionary 150 that maps each RDF term of the dataset to its respective index, and a dictionary 160 that maps in the reverse direction. iv) For each predicate indexed in the index list of the predicates of each graph, an adjacency matrix 180 representing a group of tuples of the RDF dataset.

[0053] These data structures will be described.

[0054] As is known per se, a dictionary is a data structure that associates each RDF term in a dataset with each index. Conversely, the dictionary also associates each index with each RDF term in the dataset. Therefore, the "dictionary" of a graph database maps each storage ID stored in the database (e.g., an ID stored as an RDF triple) to its respective content, i.e., it is an encoding or indexing. The dictionary is used to provide an index to the RDF triple store and helps optimize the persistence of information that may be repeated extensively. In the context of RDF (Resource Description Framework), a "term" refers to the basic unit of information or an element used in an RDF sentence (triple). Each term in RDF can play the role of the subject, predicate, or object of a triple. There are three types of terms in RDF: subject, predicate, and object. An RDF graph has three types of nodes: IRI, literal, and blank node. See, for example, https: / / www.w3.org / TR / rdf11-concepts / . Each term can be identified by an IRI (Internationalized Resource Identifier) that allows a wider range of Unicode characters, or a literal value (such as a string or a number). The subject is the resource about which a statement is made. It is the "thing" or "entity" described by the RDF triple. The predicate is the property or relationship that connects the subject and the object. It defines the nature of the relationship between the subject and the object. And the object is the value or target of the statement. It can be another resource or a literal value (string, number, date, etc.). Thus, an RDF triple is formed by a combination of these terms (subject, predicate, and object).

[0055] The dictionary is the most complex data structure in ROSA. In fact, the size of the terms indexed there can vary quite a lot, which strongly affects the composition of the dictionary data structure. Therefore, the implementation of the dictionary can be adapted to the terms found in the RDF dataset. This improves the performance of the dictionary.

[0056] In an example, the dictionary can have a fixed length when each value of the dictionary has a length equal to or less than a predetermined threshold. When the dictionary fits within the threshold (e.g., a certain number of bytes), a fixed-size dictionary can be implemented where all data fits within a unified data structure. The fixed-size dictionary saves memory usage by allocating a fixed amount of space for each entry in the dictionary. This reduces the need for dynamic memory allocation and resizing, improving memory efficiency. The fixed-size dictionary further saves indirect costs (the overhead associated with accessing data indirectly through pointers or references).

[0057] The fixed-size dictionary may consist of one vector of values called "master", where the maximum length of each value is fixed to the threshold. Thus, the vector is an array. The position of a value within the vector is the index of this value. Since the value is held in-place within the vector, it is still essential that the value is of a size below the threshold. To construct such a dictionary, it may be possible to construct it in memory and describe it in a ROSA file without the need for conversion.

[0058] In an example, the predetermined threshold can be 2, 4, or 8 bytes. This means that the size of the terms indexed in the dictionary does not exceed (or is equal to) 2, 4, or 8 bytes. The choice of value may depend on one or more parameters such as the characteristics of the data and the requirements of the system. Considerations can include, but are not limited to, the characteristics of the storage medium (hard disk drive, solid-state drive, etc.), the characteristics of the system's memory, future growth, changes in data patterns, etc.

[0059] In an example, the dictionary may be of variable length. For example, at least one value of the dictionary may have a length greater than a predetermined threshold. An additional data structure is attached to the variable-length dictionary. Each value of the dictionary having a length greater than the predetermined threshold is indexed and stored in an overflow data structure. The overflow data structure is part of the dictionary and can be regarded as a sub-data structure of the data structure of the dictionary. In this example, similar to the example regarding the fixed-length dictionary, the predetermined threshold can be 2, 4, or 8 bytes. The selection of the value may depend on one or more parameters.

[0060] In the case of a variable-length dictionary, the same "master" vector can be used. Each time a value exceeds the threshold, instead of writing all the values in-place, a reference to another vector called a "blob" where the overflowed value is written is written. The index is given by the position within the "master" vector, but indirection has to be checked to obtain the whole value. In an example of creating this kind of dictionary, in the first step, the "master" and "blob" vectors are constructed in memory, and in the second step, they can be copied to a ROSA file together with an array of offsets calculated by the cumulative size of all the "blob" vectors. For example, if the maximum value fits within an unsigned 32-bit integer, it can be dumped as an array of unsigned 32-bit integers. Otherwise, an array of 64-bit unsigned integers can be used, and this array of 64-bit unsigned integers can be referenced in the master section of the dictionary entries.

[0061] For example, the dictionary may be encoded in whole or in part, for example, the index and / or value may be encoded. The type of encoding does not matter. In the example, the value may be encoded using the IEEE standard for floating-point arithmetic (IEEE754-2019 or previous versions). Double or floating-point numbers can be represented in a normalized memory layout, and the normalization may depend on IEEE7542019 or previous versions (https: / / en.wikipedia.org / wiki / IEEE_754) instead of a character string that holds the value as a character. This encoding can be used to optimize queries, for example, filtering all values that match "x>5". For example, if the dictionary is dedicated to all floating-point values encoded in IEEE754, by definition, the dictionary is a fixed-size dictionary. In the example, one or more (or all) of the prefixes of the dictionary can be encoded using their respective hexadecimal keys instead of copying the prefix for all values of the dictionary that start with this prefix. The prefix defines a short abbreviation for a long IRI. Since the IRI mapping codec is a special codec for prefixes only, the IRI mapping codec can be used. The mapping is a table that contains a list of defined prefixes and provides codes related to the type.

[0062] As described so far, it is possible to have one dictionary for all values of all triples of the RDF dataset to be archived. In the example, the dictionary may be composed of two or more dictionaries, that is, the dictionary may be split into several small dictionaries. As an example, each dictionary can associate each RDF term of the RDF data type of the dataset with its respective index and vice versa. Otherwise, there is one dictionary for each RDF data type (https: / / www.w3.org / TR / rdf11-concepts / #section-Datatypes). Additionally, alternatively, each dictionary can associate each RDF term of the IRI (acronym for Internationalized Resource Identifier) prefix of the dataset with its respective index or vice versa. Otherwise, it is sharded one for the IRI (https: / / www.w3.org / TR / rdf11-concepts / #dfn-iri) and one for each prefix (https: / / www.w3.org / TR / rdf11-concepts / #dfn-namespace-prefix). In an example where multiple dictionaries are used in the ROSA file, each dictionary can be a fixed-size dictionary or a variable-length dictionary. The case where the dictionary is composed of both fixed-size and variable-length dictionaries will be described later.

[0063] It should be understood that even if the dictionary is composed of two or more dictionaries, the operating principle of the dictionary remains the same. In other words, the examples of the dictionary described in this specification apply.

[0064] Next, with reference to 150, 160, and 170 in FIG. 1, an example of the dictionary will be described. The dictionary maps each RDF term of the RDF dataset to its respective index. The dictionary is composed of dictionary entries. The dictionary entries are defined by several elements. i) Its type. For example, fixed-length or variable-length. ii) Data corresponding to its type, such as a "master" vector (or section), or a "master" and "blob" vector (or section). This has already been described.

[0065] In an example, the dictionary may include entries with any index, such as for access control security. The adjust section may exist as an array of unsigned integers, such as an array of unsigned 16-bit integers. The adjust section is used to remap the dictionary internal index to the dataset index in the form of (64K)*(64K).

[0066] As shown at 150 in FIG. 1, the dictionary has N dictionary entries (Entry1 DictionaryEntry... EntryN DictionaryEntry) that make up the data of the dictionary.

[0067] In an example, metadata 152 for describing the dictionary may be added. The metadata may be composed of a type section indicating the type of the dictionary layout (fixed size or variable length). Also, the metadata may constitute a codec section indicating the name of the codec if there is a codec used in this dictionary. The metadata may further include an index_remapping section, and for each 64K index, a type_map_id associated with the slice, and a DictionaryRange array including the index remapping in the associated index. The index remapping functions for the purpose of ensuring that the dictionary has consecutive indexes. The metadata may further include an IRI mapping code blob including a list of all IRI encoding prefixes referred to by the uri_prefix_mapping of the dictionary footer. The uri_prefix_mapping is stored irregularly in the metadata 152 or the dictionary 170.

[0068] As shown at 160 in FIG. 1, the inverse dictionary constitutes its own data structure. The data structure of the inverse dictionary is the same as the foregoing. In the example, the inverse dictionary may be added to each (non-inverted) dictionary. For example, it may be implemented as S-Tree<hash(e), index(e)>. Here, the S-tree contains a pair of the hash and index of each element (e). For the data structure of the S-Tree, refer to https: / / en.algorithmica.org / hpc / data-structures / s-tree / . For example, the inverse dictionary includes an array of pairs <hash(e),index(e)> sorted by hash.

[0069] The dictionary includes a footer 170. The footer of the dictionary may be stored in the dedicated part 170 of the ROSA file as shown in FIG. 1.

[0070] The footer of the dictionary includes an array of DictionaryEntry referenced by the type_map field of the dictionary footer 170. Each entry of these dictionaries has an unsigned integer from the type section referenced by the DictionnaryFooter, for example, a 32-bit index type_id, and an unsigned integer from the codec section referenced by the DictionnaryFooter, for example, a 32-bit index codec_id. Also, the footer of the dictionary may be a bitmask called flags indicating how to interpret the entry, and holds one or more of the following values. ·FIXED_SIZE_DICTIONARY. This specifies that only the master is used because all blobs of the dictionary are of the same length. In this case, the length field indicates the width. ·VARLENGTH_DICTIONARY. This specifies that the master section includes the offsets of the blobs in the blob section, and the length specifies the width of the index (4 or 8). · SHARDED_BY_HASH and SHARDED_BY_URI. These indicate that the dictionary is sharded by the hash of the decoded blob and by the RDFDataType function of the PrefixId respectively. · Optional values. Optional values can also be added, such as when using the index of the option.

[0071] The dictionary footer 170 may include a codec section that references an array of codec names used, a type section that references an array of type names used, a uri_prefix section that has a list of prefixes, an index_remapping section that references an array of DictionaryRanges used to identify types from the index, an internal index within the dictionary, a type_map section that references the aforementioned arrays, and a count field that indicates the number of entries in the dictionary.

[0072] If the dictionary is a variable-length dictionary, "master" and "blob" vectors can be used. Generally, about two-thirds of the memory capacity is associated with these two sections. In practice, it may not be the best choice to select only one layout, so the dictionary may contain both fixed-size and variable-size records.

[0073] Thus, a fixed-size dictionary can be modeled using a continuous array and is very easy to generate. The generation of a variable-length table may be a bit more complex. As shown in 150 of Figure 1, in the first part with an index, there are offsets within the blob part (denoted as index#1BlobOffset in Figure 1), and in the last part, there are the ends of the offsets (denoted as index#NN EndBlobOffset in Figure 1). The size of the nth blob is calculated by subtracting index#N BlobOffset from index#N+1 BlobOffset.

[0074] Next, the data structure of the adjacency matrix will be described with reference to 180 in FIG. 1. Note that for each indexed predicate of each graph, an adjacency matrix representing a group of tuples of the RDF data set is obtained. As shown in FIG. 1, in graph 1 of the RDF data set, n predicates are indexed with 140, and each indexed predicate has an adjacency matrix attached to it.

[0075] As is known per se, the adjacency matrix is a binary matrix (i.e., a matrix having elements with two values, for example, 0 and 1) related to the number of subjects, predicates, and / or objects, and 1 bit means that the corresponding triple (for example, of each predicate of the adjacency matrix) exists in the RDF data set. The size of the adjacency matrix for a certain predicate may be the size of the subject × object of the RDF tuple for that predicate.

[0076] Therefore, an adjacency matrix is obtained for each predicate of each graph of the RDF data set. As is known in the art, several data structures for obtaining an adjacency matrix are known. In order to obtain the read-only data structure of the present invention, any data structure for the adjacency matrix can be used. The condition is simply that the read-only data structure can implement the concept of the adjacency matrix and the read-only data structure fits into a fixed layout size of the memory.

[0077] In the example, the adjacency matrix of each indexed predicate may be obtained by performing vertical partitioning for each predicate of the RDF database. As is known per se, vertical partitioning is a database design technique in which RDF data is divided into separate data sets or partitions based on the properties or predicates used in the triples. This technique is generally used to optimize the storage, search, and query of RDF data, and is particularly used in scenarios where a specific subset of the data is accessed more frequently than others.

[0078] Vertical partitioning may be performed as is known in the art, for example, as described in the document Abadi, D.J., et al., “Scalable semantic web data management using vertical partitioning.”, In Proceedings of the 33rd international conference on Very large databases, September 2007, pp. 411-422.

[0079] In an example, vertical partitioning may be a standard K2-triple split as described in the document ALVAREZ-GARCIA, S., et al., “Compressed vertical partitioning for efficient RDF management.”, Knowledge and Information Systems, 2015, vol. 44, no 2, p. 439-474.

[0080] In an example, a dk2-tree (a dynamic variation of the k2-tree partitioning described in the document Nieves R. Brisaboaa, et al., “Compressed Representation of Dynamic Binary Relations with Applications”, arXiv:1707.02769.) may be used to generate k2-triple partitioning.

[0081] In an example, vertical partitioning may be performed using two static B+ trees. Here, the first static B+ tree represents the subject-object correspondence (S, O), and the second static B+ tree represents the object-subject correspondence (O, S). Two S trees are obtained. For the data structure of the S tree, refer to https: / / en.algorithmica.org / hpc / data-structures / s-tree / .

[0082] In an example, vertical partitioning may be performed using a static XAM tree carved from the XAM tree as a read-only XAM tree. The XAM tree is described in EP22306928.7 filed on December 16, 2022. This document is incorporated herein by reference. An implementation example of the XAM tree is described from page 20 to page 47 of document EP22306928.7, which is incorporated herein by reference.

[0083] Referring to 180 in FIG. 1, the PredicateEntry record for each predicate will be described. Once the adjacency matrix is obtained, it is necessary to identify which predicate it refers to. This is the purpose of the PredicateEntry record.

[0084] As an example, the PredicateEntry record includes the following. · Index. An unsigned integer, for example a 64-bit unsigned integer, that identifies the predicate name in the dictionary. · count. An unsigned integer, for example a 64-bit unsigned integer, representing the number of SO (Subject, Object) pairs in the predicate. This statistic can be used by the SPARQL query engine to optimize the join order of the query. This field is optional for PredicateEntry. · An unsigned integer, for example a 64-bit offset, indicating the position of the adjacency matrix data in the set of adjacency matrices of the graph of the RDF dataset. · type. An unsigned integer, for example an 8-bit unsigned integer, indicating the method of interpreting the matrix according to the method used to construct the adjacency matrix. · Flag. For example, an 8-bit unsigned integer. This optional field is used to tag predicates when the SPARQL query engine can use special semantics for query optimization. For example, it participates in RDFS entailment (https: / / www.w3.org / TR / rdf11-mt / #rdfs-entailment) or is part of national language support (NLS, https: / / en.wikipedia.org / w / index.php?title=National_Language_Support&redirect=no). There may also be other proprietary semantics. · 14 bytes reserved optionally for future use and padding.

[0085] In an example, the read-only data structure may further include an index list 130 of one or more graphs of the RDF dataset. The graph section of the read-only data structure is composed of an index list of the graphs of the RDF dataset and also specifies an array of GraphEntry. Each graph in the index list corresponds to a GraphEntry. Each GraphEntry describes the respective graph for the purpose of identifying the graph name in the dictionaries 150, 160, 170.

[0086] As an example, a GraphEntry includes the following. · Index. An unsigned integer for identifying the graph name in the dictionary. · Offset. An unsigned integer such as 64 bits for specifying the position of the PredicateEntry array. This allows one to know that a predicate exists within the graph. · nb_predicate. A 32-bit unsigned integer indicating the number of PredicateEntry in the graph. · 32 bits reserved optionally for padding and future extensions.

[0087] In an example, the read-only data structure may further include a header 100 at the beginning of the read-only data structure (ROSA file) and a footer 190 at the end of the read-only data structure. The read-only data structure may be a binary file, the header may be a fixed-size header, and the footer may be a fixed-size footer. Since the carving algorithm of ROSA is a one-pass algorithm, most references are backward references. Alternatively, the references may be forward, but multiple passes are required to create the ROSA file. Interestingly, backward references in the memory map file, i.e., random access, are compatible with modern computers using SSDs (solid state drives) where random access is inexpensive. SSDs are not a requirement of the present invention, and it should be understood that the ROSA file can be stored in any type of memory. The header is read at the opening of the ROSA file to identify the contents of the file, and the footer is read last as a backward reference for moving within the sections of the file. Navigation within the file using backward or forward references is performed as known in the art.

[0088] In an example, the header 100 includes a basic record. As an example, the header can constitute a "magic" record that is, for example, a 32-bit unsigned integer value used to uniquely identify the type of the ROSA file. Alternatively or additionally, the header of the ROSA file may include a version record that is, for example, a 32-bit unsigned integer tag used to identify the current version of the read-only data structure format. Alternatively or additionally, the header may include an "end" record that is, for example, a 64-bit unsigned integer. The "end" record is the offset to the end of the footer section, i.e., the footer record position, and is for knowing the end of the ROSA file and thus the end of the memory layout.

[0089] It should be understood that the footer may be composed in different ways in other embodiments. In these examples, the footer record is more complex than the header. Fields are added backward to maintain backward compatibility of the footer in case minor changes are added in future revisions. Similar to the header record, the footer contains a magic version for identifying it, as well as an offset to the start position of ROSA.

[0090] In these examples of footers, it may be possible to add extensions to a plain ROSA file. For example, regarding the security of access control, a dedicated data structure can be added to index the RDF data. Flags can be used in the footer, and the flags indicate whether such optional extensions exist. According to these flags, the footer may have additional records corresponding to them, such as a section representing the additional data structure and a security entry record with a count of the entries of this data structure.

[0091] In these examples of footers, the footer may have the SHA-256 of the ROSA payload. The payload of the read-only data structure is composed of the section (and thus the data) from the end of the header to the start of the footer. SHA-256 is known per se. SHA-256 takes the initials of Secure Hash Algorithm 256-bit and is a cryptographic hash function belonging to the SHA-2 family of hash functions. The output of SHA-256 is a fixed-size 256-bit (32-byte) hash value, usually represented in hexadecimal. This SHA-256 can be used to check the integrity of the ROSA file. Additionally, or alternatively, this SHA-256 can be used to provide a reliable way to detect the difference between an old SHA-256 output of the summary of the RDF data set and an invalid summary of the RDF data set. The summary of the RDF data set can be constructed as discussed in European Patent EP22306798.4 filed on December 6, 2022.

[0092] In an example, each entry of the read-only data structure can have a maximum size ranging from 32 bits to 256 bits. The expression "entry of the read-only data structure" means that the individual elements (or items) of the read-only data structure cannot have a size exceeding 256 bits and the minimum size is 32 bits. If the minimum size is not reached, padding can be used to conform to the minimum size. As an example, the minimum size is 32 bits and the maximum size is 64 bits. In another example, the individual elements (or items) are encoded in 32 bits. In a further example, the individual elements (or items) are encoded in 64 bits. Since the size is fixed, the assembly of the read-only data structure is easy.

[0093] As an example, the read-only data structure (ROSA archive) may be limited to 2 32 entries. This improves the compression rate of the read-only data structure. In fact, instead of having one huge ROSA file, multiple ROSA files can be queried. Furthermore, the access bandwidth to cloud storage is disadvantageous for downloading huge flat files. Therefore, a read-only data structure limited to 2 32 entries conforms to the limitations of general cloud storage.

[0094] Without changing the present invention, the read-only data structure can be limited to 2 64 entries, or 2 128 entries, or 225 6 entries. As already explained, the read-only data structure may be memory-mapped, and thus the overall size of the read-only data structure may be limited by the memory available in the system.

[0095] The read-only data structure of FIG. 1 shows an example in which all the sections described above in this specification are represented. It should be understood that the read-only data structure according to the present invention is not limited to its specific examples. In particular, the minimum requirements for the read-only data structure are the index list 130 of one or more graphs of the RDF data set, for each graph of the RDF data set, the index list 140 of the graph's predicates, the dictionary 150 that maps each RDF term of the data set to its respective index, and the existence of the inverse dictionary 160. And for each predicate indexed by the index list of the predicates of each graph, an adjacency matrix 180 representing a group of tuples of the RDF data set is generated.

[0096] Next, a comparison between the read-only data structure of the present invention and Apache Parquet will be described. Although Apache Parquet belongs to the field of relational databases, the present invention belongs to the field of graph databases. Even if the relational world and the graph world represent two different paradigms that have no relation to each other except for the purpose of data storage, this comparison shows the improvement of the read-only data structure of the present invention. Parquet files are composed in the form of an enumeration of column slabs. Each slab contains its own metadata and some statistical information. The slab has its own dictionary and can be operated on its own. This is intended to be operated without joins and is very efficient when there are many columns in the table and the query uses only a few columns. Although Parquet files are not indexed, if the query condition can use statistical information to exclude many slats, it is possible to skip a lot of data (e.g., in a time range query where the slabs are sorted temporally). In the case of ROSA files, 80% of the disk usage of the RDF graph is in the storage of the dictionary, and since SPARQL queries execute joins without the query engine accessing the dictionary using the dictionary index, many joins are executed. Incidentally, most of the expensive time is spent during access to the adjacency matrix, which only uses 20% of the memory usage. If the query uses only 10% (size) of the predicates, the query of the ROSA archive will use at most 2% of the corresponding size.

[0097] Referring to FIG. 2, a computer-implemented method for saving the modifications applied to the read-only data structure, that is, for incrementally updating the ROSA file, will be described. As already described, the present invention aims to provide a high-speed read access, compression, and updatable data structure for the RDF dataset. In other words, it is an RDF archive that can be updated almost in real time regardless of the size of the RDF knowledge graph.

[0098] The update of ROSA files is performed by the W3C standard SPARQL query language. In SPARQL 1.1 Update, changes can be represented as intentional or extensional definitions. An intentional definition gives meaning to a term by specifying the necessary and sufficient conditions for when the term should be used. This is an approach diametrically opposed to an extensional definition that defines by listing all those that fall under the definition. All triples of the RDF dataset to be updated are listed. The update of an RDF dataset includes writing, modifying, and deleting at least one triple of the RDF dataset.

[0099] In the example, the update of read-only data storage is performed using CDC technology. As is known per se, CDC (Change Data Capture) is a technology used to identify and capture changes added to the data of a database. The main purpose of CDC is to recognize and track changes so that other systems or processes can respond accordingly. This is particularly useful in scenarios where it is essential to synchronize multiple data sources or propagate changes to downstream systems.

[0100] Figure 5 is an example that, as a limitation of SPARQL 1.1 Update, defines CDC for SPARQL update and enables changes to be expressed only as extensional definitions. It should be understood that this is an example of how triple modification is expressed.

[0101] As long as information is given about which triples in which graph have been changed, it is also possible to use other representations without changing the present invention. This will be discussed here.

[0102] Next, referring to FIG. 2, a first list of added and / or deleted tuples of the RDF dataset is obtained (step S200). "Obtaining the first list of added and / or deleted tuples of the RDF dataset" means providing the first list of added and / or deleted tuples of the RDF dataset to this method.

[0103] Next, a first read-only data structure is calculated (step S210). The calculated first read-only data structure is defined as described above, which means that the (first) ROSA file is calculated from the first list of added and / or deleted tuples applied to the RDF dataset. The (first) ROSA file of the first list of added and / or deleted tuples of the RDF dataset is constructed such that for each predicate of the added and / or deleted tuples of the RDF dataset in the first list, the first adjacency matrix represents the added tuples of the RDF database containing the same predicate. Further, the second adjacency matrix represents the deleted tuples of the RDF database consisting of the same predicate.

[0104] Then, the calculated read-only data structure is stored in the first file (step S220).

[0105] The initially saved file is also called "Delta ROSA". Delta ROSA is a representation of CDC information and may be, for example, a binary representation of the CDC representation. Structurally, Delta ROSA is based on the ROSA file format. As a result, Delta ROSA files inherit all the characteristics of ROSA files, mainly the zero-copy reading characteristic and compactness. The generation of Delta ROSA depends only on the size of the changes, not on the size of the ROSA file (i.e., the snapshot file or archive of the RDF dataset). Delta ROSA is independent of the ROSA file. When constructing Delta ROSA, there is no need to access the ROSA file. This brings an obvious advantage in performance for generating Delta ROSA files because the ROSA file may become huge regardless of the ROSA file compression method.

[0106] Delta ROSA files are a new file format similar to ROSA. Delta ROSA, like ROSA, contains a read-only adjacency matrix and a list of read-only dictionaries. While the adjacency matrix of a ROSA file describes a set of triples (ROSA contains one adjacency matrix per predicate and per graph), Delta ROSA contains two adjacency matrices per predicate and per graph. One describes the set of triples created by the operations described in the CDC file, and the other describes the set of triples deleted by the same operations. For example, if a certain triple is added and then deleted by an operation in the CDC file, that triple does not appear in the Delta ROSA file at all.

[0107] After step S220 of FIG. 2, a second list of added and / or deleted tuples of the RDF dataset may be obtained. A second read-only data structure is calculated, and this second read-only data structure is defined as described above in this specification, that is, a (second) ROSA file is calculated. This (second) ROSA file, when saved, is called a second delta ROSA file. The (second) ROSA file is calculated from a first list of added and / or deleted tuples applied to the RDF dataset and a second list of added and / or deleted tuples applied to the RDF dataset. Similar to the first ROSA file calculated from the first list of added and / or deleted tuples, this second ROSA file has, for each predicate of the added and / or deleted tuples of the RDF dataset in the first and second lists, a first adjacency matrix constructed to represent the added tuples of the graph database (of the first and second lists) consisting of the same predicate. Further, the second adjacency matrix represents the deleted tuples of the graph database (of the first and second lists) consisting of the same predicate.

[0108] Next, the calculated read-only data structure is saved to a second file, forming a second delta ROSA file.

[0109] Each time a new list of added and / or deleted tuples of the RDF dataset is obtained, for example, the third, fourth... nth list of added and / or deleted tuples of the RDF dataset, the same process may be repeated. For example, the calculated nth delta ROSA file is calculated from the first, second to nth lists of added and / or deleted tuples, thereby obtaining a chain of modifications. That is, to obtain the entire history up to the nth list, only the nth delta ROSA file is required, which is important for performance.

[0110] There is no need to retain all the history of the RDF knowledge graph. Only one snapshot of the RDF knowledge graph is retained as a ROSA file, and delta ROSA files representing the modifications applied to the received ROSA file are chained.

[0111] Thanks to the adjacency matrix of the delta ROSA file, or the delta ROSA files of the chain of delta ROSA files, the triples added or removed by a list of operations can be easily listed. Subsequently, the delta ROSA file(s) is / are used to change the result of the BGP operation on the ROSA file (archive), and triples are added or removed from the result according to the content of the delta ROSA file. The use of BGP and delta ROSA files will be described later.

[0112] In the example, the ROSA file (archive) may be composed of two or more graphs as already described. In this case, the step of obtaining a list of added and / or removed tuples of the RDF dataset may include the step of obtaining a list of added and / or removed tuples of the RDF dataset for each graph of the RDF set, and the step of calculating a read-only data structure for the obtained list of added and / or removed tuples may further include the step of calculating a first adjacency matrix and a second adjacency matrix for each graph of the RDF set.

[0113] In the example, when the chain exceeds a threshold, a new complete snapshot may be regenerated as a ROSA file (archive). Regeneration means calculating a new archive consisting of updates to the delta ROSA chain. After regeneration, the previous archive and the new archive are swapped. Since the updates are input into the "original" RDF knowledge dataset, the delta ROSA chain may be suppressed. This threshold is a trade-off between the CPU cost for regeneration and the performance cost during read queries (the memory cost of the delta ROSA and the CPU cost for using them), and can be parameterized according to the needs of the application. The threshold may also be based on, for example, the size of the chain of ROSA files, the lifespan of the ROSA chain, etc.

[0114] In the example, a mapping may be calculated between the index of the dictionary of the RDF dataset and the index of each dictionary of each list of added and / or deleted tuples of the RDF dataset. The calculated mapping is saved to a file. This mapping avoids looking up values in the ROSA dictionary and allows looking up the index in the dictionary of the delta ROSA of each triple from the results of future BGP operations. In fact, the BGP operation on the ROSA file returns the index of the RDF terms defined in the ROSA dictionary instead of the value of the triple. Even if the delta ROSA file has a local dictionary, there is no reason for the local dictionary to contain the same index as the dictionary of the ROSA file. Therefore, a new mapping file format is defined. It contains a data structure (e.g., adjacency matrix) that can efficiently search for the ROSA index from the delta ROSA index. Originally, this mapping can only be calculated from the ROSA file and the delta ROSA file. This file is called DROSAM (short for delta ROSA mapping). DROSAM can be calculated every time a new delta ROSA file is calculated. Alternatively, DROSAM may be calculated every time a new delta ROSA file is used.

[0115] Therefore, the modifications applied to the archive of the RDF knowledge dataset are stored in one or more delta ROSA files, and a set of delta ROSA files can form a chain of modifications applied to the archive. The mapping may be calculated and saved in a file called DROSAM for the purpose of improving the identification of the ROSA index from the delta ROSA index and, conversely, for the purpose of improving the efficiency of future BGP operations thereby.

[0116] Next, with reference to FIG. 3, a method for SPARQL query in the read-only data structure of the present invention will be described. This query is executed against the ROSA file (archive) and the latest delta ROSA file (calculated).

[0117] In step S300, a SPARQL query is obtained. SPARQL is a W3C recommendation for querying RDF data and is a graph matching language built on patterns of RDF triples. The "pattern of RDF triples" means a pattern / template formed by an RDF graph. In other words, the pattern of RDF triples is an RDF graph (i.e., a set of RDF triples), and the subject, predicate, and object, or label of the graph can be replaced with variables (for the query). SPARQL is a query language for RDF data and can express queries across various data sources regardless of whether the data is stored natively as RDF or presented as RDF via middleware. SPARQL is mainly based on graph isomorphism. Graph isomorphism is a mapping that respects the structure of two graphs. More specifically, it is a function between the vertex sets of two graphs that maps adjacent vertices to adjacent vertices.

[0118] SPARQL has the ability to query for mandatory and optional graph patterns and their conjunctions and disjunctions. SPARQL also supports aggregation, subqueries, negation, value creation by expressions, test of extensible values, and query restriction by the source RDF graph. That is, a SPARQL query needs to answer eight different triple patterns possible in SPARQL. Such eight triple patterns are (S,P,O), (S,?P,O), (S,P,?O), (S,?P,?O), (?S,P,O), (?S,?P,O), (?S,P,?O), and (?S,?P,?O), where the symbol? is placed before the variable in these patterns. A variable is the output of a triple pattern and may also be the output of a SPARQL query. In some examples, a variable may be the output of a SELECT query. The output of a SPARQL query can be constructed using variables (e.g., aggregators such as sum). Variables in a query may be used to construct graph isomorphism (intermediate nodes necessary to obtain the result of the query). In some examples, variables in a query are not used for either output or intermediate results.

[0119] As is well known, BGP (the initials of Basic Graph Pattern) refers to a simple RDF query pattern used in RDF query languages such as SPARQL (SPARQL Protocol and RDF Query Language). Since BGP is basically a set of triple patterns, it can represent the pattern of triples that match in RDF data.

[0120] A SPARQL query may be obtained by a SPARQL query engine. As is known per se, a SPARQL query engine is a system designed to process a SPARQL query and obtain information from an RDF data store. The SPARQL query engine takes a SPARQL query as input and performs processing on RDF data. It interprets the query, optimizes it for efficient execution, and extracts relevant information from the RDF graph.

[0121] After the SPARQL query engine obtains the SPARQL query in step S300, the ROSA file and the delta ROSA file (i.e., the chain of delta ROSA files) are obtained by the present system in step S310. "Obtaining the ROSA file and the delta ROSA file" means providing the database to this method. In the example, such obtaining or providing means either downloading the ROSA file and the delta ROSA file (e.g., from an online database or an online cloud), or obtaining the database from a memory (e.g., a persistent memory), and / or memory mapping the ROSA file and the delta ROSA file, or including them.

[0122] Next, in step S320, for each triple pattern of the query, a tuple that answers the triple pattern of the query in the read-only data structure (ROSA file) is found. As a result of the query, the first result set is obtained.

[0123] Next, in step S330, for each triple pattern of the query in the chain of delta ROSA files, the following procedure is executed. First, it is determined whether the tuple matching the triple pattern is to be deleted or added.

[0124] In the negative case, the result of the query is what was found in the first result set.

[0125] In the positive case, two situations are considered: i) If the tuple is deleted as inferred from the chain of delta ROSA files, the tuple is deleted from the first result set. And / or ii) If the tuple is added as inferred from the chain of delta ROSA files, the tuple is added to the first result set.

[0126] When step S330 is achieved, the initial result consists of the added tuples and there are no more tuples to be deleted. Thus, the SPARQL query is executed against the modified archive. The mapping stored in the DROSAM file can also be used to improve the efficiency of BGP operations (in terms of resource consumption and execution speed).

[0127] Figure 3 shows an example of the method executed by the so-called ROSACE system (ROSACE is short for ROSA Combination ENGINE). The ROSACE engine system defines an RDF dataset as the combination of the state of the RDF knowledge dataset materialized by the ROSA file and the set of changes applied to this state materialized by the (latest) delta ROSA file, the DROSAM file (which can be calculated from the ROSA file and the delta ROSA file when the delta ROSA file is first used), and read statements can be applied to this dataset. The working area of ROSACE can eliminate the need to regenerate the ROSA file when a series of changes are small compared to the state of the RDF data and when the changes are small compared to the size of the ROSA file.

[0128] Referring to Figures 6 to 10, an example of the usage scenario of the ROSACE system will be described. In this scenario, an error occurred in the RDF dataset called ESCO. ESCO is a multilingual classification of European skills, competencies, and occupations. ESCO is part of the Europe 2020 strategy. (ESCO homepage (europa.eu)). As can be seen in Figure 6, it is an error that "Datalogue" is written instead of "Datalog" in English. The purpose of this scenario is to correct the ESCO dataset without regenerating the ROSA file representing it. Thus, the starting point is the ROSA file containing version 1.0.8 of ESCO and the error, and two SPARQL requests. Counts the number of instances of skosxl:label. And -Indicates the value of the skosliteralForm of -uri..<...e8>.

[0129] As shown in Figure 7, a CDC-formatted correction is obtained. Next, as shown in Figure 8, a delta ROSA file is calculated as the import result of the CDC. Next, in Figure 9, a delta ROSAM file is calculated to associate the index of the delta ROSA file with the index of the ROSA file. Finally, in Figure 10, a hybrid dataset is obtained, and the hybrid dataset is composed of a ROSA file, a delta ROSA file, and a delta ROSAM file. Therefore, it is possible to execute a query on the ROSA file considering the changes stored in the delta ROSA file.

[0130] Figure 11 shows the experimental results according to an embodiment of the present invention. It should be understood that the response time of the query may vary depending on the query executed. The important thing is to notice from these results that the queries (Datalog queries and counts of SKOSXL labels) are of the same order.

[0131] One or more examples of read-only data structures may be combined. Examples of read-only data structures can be applied indifferently to RDF triples or RDF quads. For the sake of convenience in the description, the concept of quads will be described here. As an example, the RDF tuples of an RDF dataset are RDF quadruples. An RDF quad is obtained by adding a graph label to an RDF triple. In such an example, the RDF tuple includes an RDF graph. A standard specification defining RDF quads (also called N-Quads) has been published by the W3C. See, for example, "RDF 1.1 N-Quads, A line-based syntax for RDF datasets", W3C Recommendation 25 February 2014. An RDF quad is obtained by adding a graph name to an RDF triple. The graph name can be specified as empty (i.e., the default or unnamed graph) or an IRI (i.e., a graph IRI). In the example, the predicate of the graph may have the same IRI as the graph IRI. The graph name of each quad is the graph to which that quad belongs in each RDF dataset. An RDF dataset represents a collection of graphs, as is known per se (e.g., https: / / www.w3.org / TR / rdf-SPARQL-query / #rdfDataset). In all the examples described so far, the term RDF tuple (or tuple) refers indifferently to RDF triples or RDF quads unless the use of one or the other is explicitly mentioned. Thus, a knowledge RDF dataset is composed of RDF triples or RDF quads. If the knowledge RDF dataset is composed of RDF quads, the BGP becomes a quad pattern that additionally has the label of the graph as a query variable. In a specific example where the method obtains one or more adjacency matrices as representations of groups of tuples respectively, the subject and the object can be queried in one adjacency matrix.

Claims

1. A computer-implemented read-only data structure for archiving a snapshot of an RDF dataset into a file, wherein the RDF dataset has one or more graphs, and the read-only data structure is directly queryable using all possible triple patterns of a SPARQL query engine, and the data structure includes a list of indexes of one or more graphs of the RDF dataset, a list of indexes of predicates for each graph of the RDF dataset, a dictionary that associates each RDF term of the dataset with its respective index and vice versa, and for each predicate indexed in the index list of predicates of each graph, an adjacency matrix representing a group of tuples of the RDF dataset, whereby a complete representation of the RDF dataset directly queryable using all possible triple patterns of a SPARQL query engine is obtained a computer-implemented read-only data structure.

2. When each value of the dictionary has a length equal to or less than a predetermined threshold, the dictionary has a fixed length, preferably, the predetermined threshold is 2, 4, or 8 bytes The computer-implemented read-only data structure according to claim 1.

3. When at least one value of the dictionary has a length greater than the predetermined threshold, the dictionary has a variable length, and each value having a length greater than the predetermined threshold is indexed and stored in an overflow data structure, and the overflow data structure is part of the dictionary The computer-implemented read-only data structure according to claim 2.

4. At least a part of the dictionary is encoded, preferably, the values are encoded using the IEEE floating-point arithmetic standard (IEEE754-2019 or a previous version), or the prefix of the values is encoded using a hexadecimal key The computer-implemented read-only data structure according to any one of claims 1 to 3.

5. The dictionary consists of two or more dictionaries, preferably, each dictionary associates each RDF term of an RDF data type of the dataset with its respective index and vice versa, or each dictionary associates each RDF term of an IRI prefix of the dataset with its respective index and vice versa A computer-implemented read-only data structure according to any one of claims 1 to 4. **Claim 6** The adjacency matrix for each indexed predicate is obtained by performing vertical partitioning for each predicate of the RDF database, and preferably, the vertical partitioning is performed by a technique selected from the following Standard K2-triples, preferably those made from dynamic K2, Two static B+ trees, the first static B+ tree representing the subject-to-object (S,O) correspondence, and the second static B+ tree representing the object-to-subject (O,S) correspondence, A static XAM tree engraved as a read-only XAM tree A computer-implemented read-only data structure according to any one of claims 1 to 5. **Claim 7** Each entry of the data structure has a maximum size configured between 32 bits and 256 bits, preferably between 32 bits and 64 bits, and more preferably, the size of each entry is encoded in 32 bits or 64 bits A computer-implemented read-only data structure according to any one of claims 1 to 6. **Claim 8** The read-only data structure is a binary file including the following A header composed of an offset to the end of the recording position of the footer, A section configured between the header and the footer, The footer composed of an offset to the recording start position of the header, The section The dictionary For each graph, a list of indices of the graph's predicates and an adjacency matrix for each predicate indexed by the list of indices of the graph's predicates To store A computer-implemented read-only data structure according to any one of claims 1 to 7. **Claim 9** A computer-implemented method for storing modifications applied to a read-only data structure according to any one of claims 1 to 8, comprising Obtaining a first list of added and / or deleted tuples of the RDF dataset, Calculating a first read-only data structure defined according to any one of claims 1 to 8, calculating for each predicate of the added and / or deleted tuples of the RDF dataset of the first list, The first adjacency matrix represents the added tuples of the RDF dataset consisting of the same predicate, and / or The second adjacency matrix represents the deleted tuples of the RDF dataset consisting of the same predicates, the step of calculating, the step of storing the calculated read-only data structure in a first file, and a computer-implemented method having the above.

10. the step of obtaining a second list of added and / or deleted tuples of the RDF dataset, the step of calculating a second read-only data structure defined according to any one of Claims 1 to 8, which is calculated for each predicate of the added and / or deleted tuples of the RDF dataset of the first and second lists, The first adjacency matrix represents the added tuples of the RDF database consisting of the same predicates, and / or The second adjacency matrix represents the deleted tuples of the RDF database consisting of the same predicates, the step of calculating, the step of storing the calculated read-only data structure in a second file, thereby forming a modification applied to the read-only data structure, and The computer-implemented method according to Claim 9, further comprising the above.

11. The step of obtaining further includes the step of obtaining a list of added and / or deleted tuples of the RDF dataset for each graph of the RDF set, The step of calculating includes the step of calculating a first adjacency matrix and a second adjacency matrix for each graph of the RDF set. The computer-implemented method according to Claim 9 or 10.

12. the step of calculating a mapping between the index of the dictionary of the RDF dataset and the index of each dictionary of each list of added and / or deleted tuples of the RDF dataset, the step of storing the calculated mapping in a file, and The computer-implemented method according to any one of Claims 9 to 11, further comprising the above.

13. A computer-implemented method for SPARQL queries on a read-only data structure stored in a file according to any one of Claims 1 to 8 and a calculated read-only data structure stored in a first file according to any one of Claims 9 to 12, the step of obtaining a SPARQL query by a SPARQL query engine, wherein the SPARQL query includes at least one triple pattern, the step of obtaining, Obtaining a read-only data structure according to any one of claims 1 to 8 and storing the read-only data structure calculated according to any one of claims 9 to 12 in a file; For each triple pattern of the query, finding a tuple in the data structure that responds to the triple pattern of the query according to any one of claims 1 to 8, thereby obtaining a first set of results; For each triple pattern of the query, converting the calculated read-only data structure into a file according to the method described in any one of claims 9 to 12; Determining whether a tuple answering the triple pattern is deleted or added; If the tuple is deleted, deleting the tuple from the first result set; If the tuple is added, adding the tuple to the first result set; A computer-implemented method having the above steps.

14. A computer program comprising instructions for executing the method according to any one of claims 9 to 12 and / or the method according to claim 13.

15. A system including a processor coupled to a memory, wherein the memory stores the computer program according to claim 14.