Encoding data in a hierarchical data structure using hash trees for integrity protection
Patent Information
- Application Number
- JP2023580769
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-11
- Filing Date
- 2022-07-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-01
AI Technical Summary
Existing data structures for genomic data lack efficient methods for integrity control, particularly in ensuring the integrity of large files over long periods, and current solutions do not effectively group data structures to provide evidence of file integrity, allow for updates while maintaining integrity, and lack tracking of changes and accountability.
A hierarchical data structure using hash trees, specifically Merkle trees, is employed to encode and verify genomic data, where a selected subset of nodes and leaves form a portion of the hash tree, allowing for quick verification and updating while reducing storage size, and incorporating digital signatures for enhanced integrity protection.
The proposed solution enables efficient integrity protection of genomic data over time, allowing for quick verification and tracking of changes, while reducing storage requirements and preventing unauthorized modifications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The disclosed subject matter relates to an encoding system for encoding genomic data in a digital data structure, a verification system for verifying selected genomic data in a digital data structure, an encoding method for encoding genomic data in a digital data structure, a verification method for verifying selected genomic data in a digital data structure, and a computer readable medium. [Background technology]
[0002] As the amount of genomic data is constantly increasing, it is important that such information is stored in appropriate data structures. ISO / IEC 23092, which is incorporated herein by reference, defines standards for encoding, compressing, and protecting genomic data. In particular, ISO / IEC DIS 23092-1, "Information technology-Genomic information representation-Part 1: Transport and storage of genomic information," which is also incorporated herein by reference, defines data structures for storing and / or streaming genomic information.
[0003] Known data structures disclose hierarchical data structures that can store genomic data, e.g., sequence data, and associate it with other information related to the genomic data. For example, Table 4 of ISO / IEC DIS23092-1 discloses a format structure and hierarchical encapsulation levels. Table 4 shows boxes for various types of data and their possible storage.
[0004] Files for genomic data can be very large - hundreds of gigabytes or even terabytes in size - and traditional completeness measures take a long time to compute across that much data. Summary of the Invention [Problem to be solved by the invention]
[0005] It would be advantageous to have improved data structures for genomic data that allow better integrity control. For example, the above-mentioned ISO / IEC 23092 does not describe how to group together all data structures that efficiently give proof of the integrity of the file, especially for the whole file. Also, it is not feasible to add or remove data structures from a file or update a genomic file while taking integrity into account. There is no tracking of how the file is updated or who is responsible for those changes. Since genomic data pertains not only to the user's health care data but also to the health care data of his / her descendants, it is advantageous to protect the integrity of genomic data over time. Using individual digital signatures on selective data components is not sufficient for this purpose. For example, it does not protect the relationships between data structures. An attacker may delete components or change their order. Any of these issues deserve to be addressed individually. Other issues are identified and addressed herein. [Means for solving the problem]
[0006] Some embodiments are directed to a digital data structure. The data structure includes a plurality of genome blocks and a portion of a first hash tree. The hash tree is computed from a plurality of hash values of the plurality of genome blocks. The included portion of the first hash tree includes a selected subset of nodes, which may be a combination of the highest level or levels of nodes and a selected number of leaves of the first hash tree. It should be understood that the first hash tree need not be the first hash tree that appears in the data structure.
[0007] Genomic data is a particularly advantageous application because it is typically large and also hierarchical. However, embodiments may be applied to any type of data, particularly hierarchically organized data. Although many embodiments are described in the context of genomic data, the invention is not limited to genomic data.
[0008] In general, a hash tree or Merkle tree is a tree structure in which each leaf node contains a hash of a data block of data or the root of a hash tree of subordinate containers, and non-leaf nodes contain a hash over the nodes at the next lower level, the latter nodes may be leaf nodes or non-leaf nodes. A special type of hash tree is a Verkle tree, in which the hash function used to compute an internal node (non-leaf node) from its child nodes is a vector commitment rather than a regular hash. An embodiment may use a regular hash, in particular a Merkle-Damgard type hash function (MD-type), such as the SHA family, e.g., SHA-3. The leaves of a Verkle type tree use, for example, an MD-type regular hash function. Further description of Verkle trees can be found in the document "Verkle Trees" by Kuszmaul, Technical report; Massachusetts Institute of Technology: Cambridge, MA, USA, 2018. Multiple genomic data blocks may have been received, for example as partitions of genomic data. For example, in one embodiment, genomic data is received, such as a genome sequence and / or other genomic data. The genomic data may already be partitioned into blocks or may be divided into blocks by an encoding system.
[0009] By including the top-level part, fast verification or update is maintained, but by not including the lower level parts, the storage size is reduced, which could be some of the leaves, all the leaves, or even some lower levels.
[0010] In one embodiment, the data structure is a hierarchical data structure. A block at a higher level may point to multiple blocks at a lower level. A block at a higher level may contain a portion of a hash tree calculated over a lower level. This may happen more than once. For example, a first level may contain a portion of a hash tree calculated over blocks at a second lower level. A second level may contain a portion of a hash tree calculated over blocks at a third lower level, and so on. The first level may contain a portion of another hash tree calculated over other blocks at a second lower level.
[0011] The data structure or portions thereof may be stored, retrieved, streamed, received, encoded, and verified. When streaming the data structure, portions that were not included may be recomputed and included during streaming. When streaming the data structure, the streaming may include only selected genomic data blocks and portions of the hash tree that are needed to recompute the root of the hash tree without accessing non-streamed data.
[0012] In one embodiment, the hash tree is not stored in its entirety. This can also happen across multiple levels. Interestingly, a first level includes a partial hash tree for a second lower level, which also includes a partial hash tree for a third lower level. A hierarchy of partially included hash trees is constructed. In one embodiment, a tree parameter is received. The size of the portion of the hash tree that is included is determined from the tree parameter. For example, the tree parameter is the number of levels to include. In one embodiment, the tree parameter is greater than or equal to 2.
[0013] Other aspects of the hash tree, such as the number of children per node, i.e., the k-aryness, are also set by tree parameters.
[0014] One aspect is a verification system for verifying selected data in a data structure, for example, encoded by an embodiment of the encoding system. The verification is performed by verifying the root of the hash tree. This is typically done by recalculating the root. However, for more advanced types of hash trees, such as Verkle trees, asymmetric algorithms are used. To recalculate or otherwise verify the hash, a portion of the hash tree is either taken out of the data structure or recalculated. For example, leaf values may not be present and may be recalculated. Some of the lower levels (or portions thereof) may not be present in the data structure and may be recalculated. However, some hash values may be present in the data structure and may be retrieved. To determine which values are needed, a path is identified starting from the data block selected for verification to the root of the corresponding hash tree. The hash values for the hash tree along the path that are needed for pre-computation of the root and / or verification not only include values of the nodes on the path, but also include child nodes of the nodes on the path. Interestingly, the root may be the root of a hash tree directly associated with the data block being verified, e.g., a hash tree that contains the hash values of the data block in its leaves, e.g., a first hash tree, but it may alternatively (or as well) be the root of a hierarchically higher hash tree, e.g., a second hash tree. Once the root is recomputed, it can be compared to the root in the data structure. Other types of validation include verifying a signature over the root, or performing vector commitment validation, as in the case of a Verkle tree.
[0015] The encoding system may optionally be an electronic system implemented in one or more electronic devices. The verification system may optionally be an electronic system implemented in one or more electronic devices.
[0016] Further aspects are encoding and verifying methods. An embodiment of the method is implemented on a computer as a computer-implemented method, or in dedicated hardware, or in a combination of both. Executable code for an embodiment of the method is stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product includes non-transitory program code stored on a computer readable medium for performing an embodiment of the method when said program product is executed on a computer.
[0017] In one embodiment, the computer program comprises computer program code adapted to perform all or part of the steps of an embodiment of the method when the computer program is run on a computer. Preferably, the computer program is embodied on a computer readable medium.
[0018] Further details, aspects and embodiments will now be described, by way of example only, with reference to the drawings in which elements are illustrated for simplicity and clarity and are not necessarily drawn to scale. In the drawings, elements corresponding to elements already described may have the same reference numerals. [Brief description of the drawings]
[0019] [Figure 1a] FIG. 1 is a schematic diagram illustrating an example of an embodiment of an encoding system. [Figure 1b] FIG. 1 is a schematic diagram illustrating an example of an embodiment of a verification system. [Figure 2a] FIG. 2 is a schematic diagram illustrating an example of an embodiment of an encoded data format for storage. [Figure 2b] FIG. 2 is a schematic diagram illustrating an example of an embodiment of an encoded data format for storage. [Figure 2c]FIG. 2 is a schematic diagram illustrating an example of an embodiment of an encoded data format for storage. [Diagram 3] FIG. 2 is a schematic diagram illustrating an example of an embodiment of a hierarchical hash tree. [Figure 4] FIG. 2 is a schematic diagram illustrating an example of an embodiment of a hash tree. [Figure 5a] FIG. 2 is a schematic diagram illustrating an example of an embodiment of a hash tree. [Figure 5b] FIG. 2 is a schematic diagram illustrating an example of an embodiment of a hash tree. [Figure 6] FIG. 2 illustrates a schematic diagram of an example of an embodiment of an encoded message for streaming or transport. [Figure 7a] FIG. 2 is a schematic diagram illustrating an example of an embodiment of an encoding method; [Figure 7b] FIG. 1 is a schematic diagram illustrating an example of an embodiment of a verification method. [Figure 8a] 1 is a schematic diagram illustrating a computer readable medium having a writeable portion containing a computer program according to one embodiment; [Figure 8b] FIG. 2 is a schematic diagram illustrating a representation of a processor system according to one embodiment.
[0020] Reference List The following list of reference numbers and abbreviations corresponds to FIGS. 1a through 3 and is provided to facilitate interpretation of the drawings and should not be construed as limiting the scope of the claims. 110 Coding Systems 130 Processor System 140 Storage 150 Communication Interface 160 Verification System 170 Processor System 180 Storage 190 Communication Interface 200 Coding Systems 210 Hash Tree Units 211~213 Hash Tree Coding A, A1, …, A34 hierarchical data 220 Storage Unit 230 Streaming Unit 240 Verification Units 300 Hash Tree 310 Highest Level 311 Second highest level 313 3rd highest level 320 Reef Level 330 Selected Part 1000 Computer Readable Medium 1010 Writable area 1020 Computer Programs 1110 Integrated circuits 1120 Processing Unit 1122 Memory 1124 Special Purpose Integrated Circuit 1126 Communication Elements 1130 Interconnect 1140 Processor System DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] While the disclosed subject matter is susceptible to embodiment in many different forms, one or more specific embodiments have been shown in the drawings and will be described in detail herein, with the understanding that the disclosure is to be considered as an exemplification of the principles of the disclosed subject matter and is not intended to be limited to the specific embodiments shown and described.
[0022] In the following, for the sake of understanding, elements of an operational embodiment are described, however, it will be apparent that each element is configured to perform the described functions performed by it.
[0023] Moreover, the disclosed subject matter is not limited to the embodiments but also encompasses any other combination of features described herein or recited in mutually different dependent claims.
[0024] Figure 1a illustrates a schematic diagram of an exemplary embodiment of an encoding system 110. Figure 1b illustrates a schematic diagram of an exemplary embodiment of a verification system 160. The encoding system 110 includes a processor system 130, a storage 140, and a communication interface 150. The verification system 160 includes a processor system 170, a storage 180, and a communication interface 190.
[0025] For example, the communication interface 150 comprises an input interface configured to receive genomic data, e.g., a genome sequence and / or other genomic data. The processor system 130 is configured, for example, by software stored in the storage 140 to generate a data structure. The data structure includes a plurality of genome blocks and a portion of the first hash tree. It should be noted that instead of genomic data, other types of data may be used, in particular other hierarchical data.
[0026] For example, the communication interface 190 comprises an input interface configured to receive the data structure. The processor system 170 is configured to verify the recomputed root of the first hash tree using the highest level of the first hash tree in the obtained data structure.
[0027] Storage 140 and / or 180 may be included in electronic memory. Storage may include non-volatile storage. Storage 140 and / or 180 may include non-local storage, such as cloud storage. In the latter case, storage may be implemented as a storage interface to the non-local storage.
[0028] The systems 110 and / or 160 communicate with each other, external storage, input devices, output devices, and / or one or more sensors via a computer network. The computer network can be the Internet, an intranet, a LAN, a WLAN, etc. The computer network can be the Internet. The system comprises a connection interface configured to communicate within the system or, if necessary, outside the system. For example, the connection interface comprises a connector, such as a wired connector, such as an Ethernet connector, an optical connector, etc., or a wireless connector, such as an antenna, such as a Wi-Fi, 4G, or 5G antenna. The sensor can be a sequencing device for obtaining genome sequencing data from a sample.
[0029] The system is configured for digital communication, including, for example, receiving genomic data, storing a data structure, streaming a data structure, obtaining a data structure, for example for validation, for example, receiving or retrieving a data structure.
[0030] The implementation of the systems 110 and 160 is carried out in a processor system, for example one or more processor circuits, for example microprocessors, examples of which are shown herein. The figures and description describe functional units that may be functional units of a processor system. For example, FIG. 2a is used as a blueprint for a possible functional organization of a processor system. The processor circuits are not shown separately from the units in these figures. For example, the illustrated functional units are implemented in the systems 110 and 160 in whole or in part in computer instructions stored, for example, in electronic memories of the systems 110 and 160, and executable by the microprocessors of the systems 110 and 160. In hybrid embodiments, the functional units are implemented partly in hardware, for example as coprocessors, for example cryptographic coprocessors, and partly in software stored and executed on the systems 110 and 160.
[0031] In various embodiments of systems 110 and 160, the communication interface is selected from a variety of options. For example, the interface may be a network interface to a local or wide area network, such as the Internet, a storage interface to internal or external data storage, a keyboard, an application program interface (API), etc.
[0032] Systems 110 and 160 have a user interface, which may include familiar elements such as one or more buttons, a keyboard, a display, a touch screen, etc. The user interface is configured to provide for user interaction for configuring the system, retrieving genomic data, displaying genomic data, verifying genomic data, etc.
[0033] The storage may be implemented as electronic memory, such as flash memory, or magnetic memory, such as a hard disk. The storage may include multiple discrete memories that together constitute the storage 140, 180. The storage may include temporary memory, such as RAM. The storage may be cloud storage.
[0034] The system 110 may be implemented in a single device. The system 160 may be implemented in a single device. In general, the systems 110 and 160 each comprise a microprocessor that executes appropriate software stored in the system, for example downloaded and / or stored in a corresponding memory, for example a volatile memory such as RAM, or a non-volatile memory such as Flash. Alternatively, the system is implemented wholly or partly in programmable logic, for example as a Field Programmable Gate Array (FPGA). The system may be implemented wholly or partly as a so-called Application Specific Integrated Circuit (ASIC), for example an integrated circuit (IC) customized for their specific application. For example, the circuit is implemented in CMOS using a hardware description language, for example Verilog, VHDL, etc. In particular, the systems 110 and 160 comprise circuits for cryptographic functions, such as hash functions or signatures.
[0035] The processor circuit may be implemented in a distributed fashion, for example as multiple sub-processor circuits. The storage may be distributed across multiple distributed sub-storages. Some or all of the memory may be electronic memory, magnetic memory, etc. For example, the storage may have volatile and non-volatile portions. Some of the storage may be read-only.
[0036] The encoding system 110 may be implemented as a single device or as multiple devices. The verification system 160 may be implemented as a single device or as multiple devices.
[0037] Fig. 2a shows a schematic diagram of an example of an embodiment of an encoding system 200. Shown in Fig. 2a is a hierarchical data structure with multiple levels. At the highest level, a single data block labeled A is shown. At the second highest level, multiple blocks are shown, in this example three data blocks labeled A1, A2 and A3 are shown. Two of the data blocks at the second highest level also have further hierarchical levels below them. At the third highest level, two data block sequences are shown, with data blocks A11, A12 to A14 subordinate to data block A1, and data blocks A31, A32 to A34 subordinate to data block A3. There may be more than three levels in the hierarchical data structure. Under any data block, there may be five or more sub-blocks, for example eight or more, twenty or more, etc.
[0038] For multiple blocks in a sequence, a hash tree is constructed. Figure 2 shows a hash tree unit 210. The hash tree unit may be a functional unit of the encoding system 200. The hash tree unit 210 is configured to construct a hash tree for a sequence of data blocks. A hash tree is also known as a Merkle tree. The hash tree may be a binary hash tree. Examples of hash trees are further illustrated herein, for example with reference to Figures 3, 4, and 5.
[0039] Data, such as genomic data, is divided into a sequence of genomic data blocks 214. Although the sequence 214 is shown with three blocks, it often has more than three blocks, e.g., many more blocks, e.g., more than 10 blocks, more than 100 blocks, etc. In addition to the genomic data, the blocks may contain additional data, e.g., additional data blocks with information related to the genomic data, e.g., its origin, meaning, purpose, etc. Additional blocks A11, A12, and genomic data blocks A13 to A14 are shown in FIG. 2. Additional blocks A31, A32, and genomic data blocks A33 to A34 are shown in FIG. 2.
[0040] The input interface is configured to receive genomic data, for example in the form of a data block, but may also be configured to receive further data items. For example, the further data items and / or the genomic data are taken from multiple files combined in the data structure. The further data items may be assigned to further leaves of the first hash tree, for example their hash values are included in the leaves of the hash tree.
[0041] A hash tree, also known as a Merkle tree, is a tree in which leaf nodes are labeled with a cryptographic hash of a data block, and every non-leaf node is labeled with a cryptographic hash of the labels of its child nodes. In this case, the hash tree constructed in 211 has hashes of blocks A11 through A14 for its leaves. Most nodes typically have two child nodes, but one or more nodes may have only one child node, for example, to account for a number of blocks that is not a power of two. As a variation, a hash tree has nodes with four or more child nodes.
[0042] The hierarchical portion of the data structure above the sequence A11-A14 includes a portion of a hash tree. For example, the hash tree unit 210 calculates a number of hash values for the number of genome blocks A13 to A14 by applying a hash function to the genome data in at least the number of genome data blocks. The hash function is preferably a cryptographically strong hash function, such as SHA-3. Optionally, additional hashes are calculated for additional data blocks, such as blocks A11 and A12. The hash tree unit 210 calculates a first hash tree for the number of hash values, thereby assigning the number of hash values to leaves of the first hash tree. In one embodiment, the first hash tree has at least three levels. In one embodiment, the first hash tree has at least four leaves.
[0043] The hash function is in particular a regular hash function of the Merkle-Dangard construction, including a one-way compression function. The nodes in the hash tree or Merkle tree are not necessarily calculated by applying such a hash function. Other types of fingerprint algorithms are used, such as vector commitment for Verkle trees or other authentication tags. There may be multiple sequences of data blocks, including multiple sequences of genomic data blocks. Figure 2 shows a second example of such a sequence, blocks A31-A34. Similar to the case of sequences A11-A13, hash values are calculated and hash trees are constructed for these blocks as well.
[0044] Interestingly, the hash tree information obtained for the hash tree computed for a sequence of data blocks is included in the blocks hierarchically above it. For example, block A1 is hierarchically above blocks A11-A14. The hash tree information is included in one or more blocks at a higher level. In that example, block A1 includes hash tree information HTA1. Similarly, in this example, block A3 includes hash tree information HTA3 obtained from the hash tree computed by hash tree unit 210 for sequence A31-A34.
[0045] In one embodiment, the hash tree information includes at least the root of the hash tree, and preferably also includes the levels immediately below it. For example, the hash tree information includes the two highest levels of the hash tree. Interestingly, the complete hash tree is not included in the hash tree information. Larger or smaller portions of the hash tree are excluded. For example, in one embodiment, one or more of the lowest levels starting from the leaf are excluded from the hash tree information. By excluding portions of the hash tree, the storage or streaming requirements for the data structure are reduced. The excluded information can be recalculated later if they are needed, for example for verification, but because these levels are closer to the leaf, they are calculated over relatively less information. For example, to recalculate the hash value in the leaf, only the hash over a single data block needs to be calculated. However, to recalculate the nodes closer to the root of the hash tree, for example, much more blocks need to be included in the calculation.
[0046] In one embodiment, the lower level hash trees are included in their entirety in the data structure, but the roots of the lower hash trees are used in the leaves of the further hash trees. The further hash trees may be included partially, in particular the leaves including the roots of the lower hash trees may be omitted. Also, other parts of the further hash trees may be omitted, in particular one or more of the lower or lowest levels.
[0047] Including the levels near the top, e.g., at least the first two levels (although more levels are possible), will reduce computation time the most, and excluding the levels near the bottom will reduce storage requirements the most.
[0048] Building a hash tree can be done more than once. For example, Fig. 2a shows another example of a sequence of blocks A31-A34. The hash tree unit 210 is configured to calculate hash values for the blocks, calculate a hash tree, and include a portion of the tree, e.g., the first few levels, e.g., the first two levels, but exclude at least a portion of the tree, e.g., the hash tree information for hierarchically higher blocks, in this case the lowest or lower levels in HTA3.
[0049] The blocks containing the hash tree information, in this case blocks A1 and A3, may be dedicated to storing hash tree information, which may be included along with other information, for example information about data hierarchically lower and / or in hierarchical levels.
[0050] Building a hash tree can be done at multiple levels. For example, FIG. 2a shows another example sequence of blocks A1-A3. Blocks A1-A3 are at a hierarchically higher level than blocks A11-A14 and A31-A34. One or more of these blocks, or all of these blocks, contain hash tree information for a hash tree at a lower hierarchical level, in this case blocks A1 and A3. The hash tree unit 213 calculates hash values for the blocks in this sequence, builds a hash tree, and includes a part of the hash tree, e.g., the top part, in a hierarchically higher block, in this case block A. The hash tree information is labeled HTA. At this level 2, there can be multiple sequences, e.g., sequences B1, B2, etc. (not shown in FIG. 2a). At the level of block A, there can be multiple blocks, e.g., blocks B, C (not shown in FIG. 2a).
[0051] The hash tree information may be used in the verification to verify the integrity of the information used to calculate the hash tree information. In general, it is advantageous to include more data in one or more hash trees so that more data can be verified, and preferably all genome information is used to calculate one or more hash trees. However, it may happen that data that is less sensitive than other data is included, such as reference data, instruction data, etc. Some parts of the genome are much more sensitive than other parts, such as so-called junk DNA. Since genome data can be very large, excluding data from integrity protection can significantly speed up the verification of the data structure. In one embodiment, a portion of a plurality of genome blocks is marked as with integrity protection, and a portion of a plurality of genome blocks is marked as without integrity protection, and only the portion marked as with integrity protection is included in the first hash tree.
[0052] A hash tree allows detection of changes to the data over which the hash tree was computed. Interestingly, a hash tree allows selective validation of data. For example, in the case of partial validation, the hash tree is recalculated insofar as it depends on the blocks selected for recalculation and insofar as hash values in the hash tree are removed from storage but use stored hash values in the hash tree that depend only on data blocks that were not selected.
[0053] The hash tree alone, however, does not prevent malicious modification, e.g., modifying the data and recalculating and replacing all the hash trees. To avoid this, the encoding system is configured to compute a digital signature over the root of the hash tree and include the digital signature in the data structure. Figure 2a shows an example of this. Block A contains hash tree information for blocks A1-A3, and also contains a signature over the hash tree information. To reduce computation, the signature is computed only over the root of the hash tree, not over the entire hash tree information.
[0054] In one embodiment, one or more or all of the genome blocks are stored in compressed form. For example, one or more of blocks A13-A14 or blocks A33-A34 or all of blocks A13-A14 or blocks A33-A34 are compressed. To aid in integrity protection, a hash tree is calculated over additional hash values, i.e., hash values calculated over the uncompressed blocks as well as hash values calculated over the compressed blocks.
[0055] In one embodiment, one or more or all genome blocks are stored in compressed and encrypted form. Typically, blocks are compressed before being encrypted. For example, one or more of blocks A13-A14 or blocks A33-A34 or all of blocks A13-A14 or blocks A33-A34 are compressed and then encrypted. To aid in integrity protection, the hash tree is calculated over additional hash values, i.e., hash values calculated over uncompressed and unencrypted blocks, hash values calculated over compressed and unencrypted blocks, and hash values calculated over compressed and encrypted blocks. These are optional extensions, and an embodiment may, for example, include only hash values over compressed and encrypted blocks, or may, for example, include only hash values over uncompressed and unencrypted blocks.
[0056] Another advantage of the hash tree is that if only selected blocks have changed, e.g., been modified, the hash tree can be quickly recalculated by using the hash values in the hash tree that were calculated and stored for the blocks that were not modified. If one or more signatures are used, they are similarly recalculated.
[0057] In one embodiment, the device tracks such modifications. For example, modifications to genomic data, such as modifications including one or more of addition, deletion, and / or modification, are received. One or more hash trees are recalculated in whole or in part, and the trees are updated, for example, in storage. Interestingly, the modifications can be stored as additional blocks. The modifications can be applied when using the genomic data. The modifications can be included in additional blocks that are used in new hash trees or included in existing hash trees. This has the advantage that modifications to genomic data can be tracked.
[0058] Modifications to genomic data include, for example, modifications to metadata. Modifications to genomic data include, for example, modifications to genome sequence data. Modifications to genomic data include, for example, new data, including new metadata. Modifications to the data structure are applied and any parts of the hash tree that rely on them, e.g., calculated from the modified parts, are recalculated. Additionally or alternatively, the modifications themselves are recorded in the data structure, e.g., to aid in accountability. The modifications may be stored in new data blocks. New data blocks at the leaf level also avoid large parts of the hash tree being recalculated. For example, if there were blocks 1-100 before the modification, the modification is placed in block 101. Advantageously, the modification is placed in the tail blocks so that large parts of the hash tree can be reused without recomputation. In this case, the hash over blocks 1 and 2 can be reused, the hash over blocks 3 and 4, etc. This also works at higher levels, where the hashes that rely on blocks 1-4 at the next level can also be reused, such as in this example. When reading the data, the modifications in block 101 are applied if needed in the data.
[0059] The data structures constructed in system 200 may be used in a variety of ways.
[0060] For example, system 200 includes a storage unit 220. The storage unit 220 is configured to write data structures to a computer-readable medium, such as a non-transitory computer-readable medium.
[0061] For example, the system 200 comprises a streaming unit 230. The streaming unit 230 is configured to stream the data structure or a part thereof on a computer-readable medium, for example a transitory computer-readable medium, for example, the streaming being performed over a computer network.
[0062] 2b and 2c show an example of an embodiment of stored encoding. These examples can also be used for streaming, but in streaming, one can choose to exclude parts of the data structure. For example, in streaming, the encoding device receives a parameter indicating a selected part of the data structure. The encoding device then streams only the selected part, and optionally also blocks hierarchically above the selected part. For example, in streaming, the encoding device calculates that part of the hash tree that is not stored and includes them in the streaming. For example, if the sending device is computationally more powerful, this speeds up the verification at the receiving device.
[0063] There are many ways in which hierarchical data structures can be linearized for storage and / or streaming. For example, hierarchical relationships may be indicated in the blocks using pointers, or labels, etc. The blocks may be written in various orders. Figure 2b shows an example where blocks are written out depth-first. Figure 2c shows an example where blocks are written out breadth-first.
[0064] In one embodiment, the encoding system is configured for streaming the data structure or a portion thereof. When streaming the integrity protected data blocks, a transport message may be included in the stream, for example at the beginning of the stream. The transport message may include: an entire signature of one of the hash trees or a hierarchy of the hash trees; All hash nodes required to calculate nodes in one or more paths starting from the transported node or nodes to the entire route, so that integrity verification can be performed immediately after receiving the transported data block on the receiving end; Includes.
[0065] Returning to FIG. 2a, for example, the system 200 comprises a verification unit 240. The verification unit 240 is configured to verify the data structure or a part thereof. For example, the verification unit 240 verifies the data structure, such as the stored or streamed one in FIG. 2b or FIG. 2c. The verification may involve the verification of all the integrity data. For example, all parts of the stored hash tree are recalculated and the recalculated parts are compared with the stored parts. The integrity verification may also include the verification of the signature. Interestingly, the verification may also be applied to a selected part of the data structure. For example, the verification unit recalculates a number of hash values for the selected part of the plurality of genome blocks in the data structure by applying a hash function to the genome data in at least the selected genome data block, and recalculates the root of the first hash tree from the recalculated number of hash values of the first hash tree for the selected part of the plurality of genome blocks and the included level of the first hash tree for the unselected part of the plurality of genome blocks.
[0066] The hash tree is recomputed to the extent that it depends on the block selected for validation and to the extent that hash values in the hash tree have been removed from storage. For portions of the block that were not selected but for which hash values are available, hash values from the stored / streamed hash tree may be used.
[0067] The verification is performed in a client-server verification system with the entire data structure stored on the server and the data block to be verified available on the client where the verification is performed. For example, in one embodiment, the following verification process is used. Note that the order may be different. 1. The client informs the server of the ID of the data block in the data structure that is to be verified. 2. The server determines a path starting from the node associated with the data block to be validated to the root of one of the hash trees or the entire hierarchy of the hash trees. 3. The server identifies the data blocks or containers needed to compute the nodes in the path identified in section 2. 4. The server retrieves or computes the hash node associated with the data block or container identified in section 3, if available. 5. The server sends the extracted / computed hash nodes calculated in section 4 to the client. 6. The client calculates multiple hash values for the data blocks in the selected data structure for data integrity verification by applying a hash function to the selected data blocks. 7. The client recomputes the entire root of the hash tree hierarchy from the computed hash values of the data block to be validated (section 6) and the computed / retrieved hash nodes received from the server (section 5) that are needed to compute the nodes along the path leading to the entire root. 8. The client validates one of the hash trees in the obtained data structure or the entire recomputed root at the top of the hierarchy of hash trees.
[0068] FIG. 3 illustrates generally an example of an embodiment of a hash tree 300. Hash trees, such as the one in FIG. 3, are used in hash tree calculations 211, 212, 213, etc. A hash tree has multiple levels. The highest level 310 has a single node, the root of the hash tree. The second highest level 311 includes at least two nodes. In one embodiment, two nodes of level 311 are included in storage. Also illustrated in FIG. 3 is a third highest level 312 and a leaf level 320, e.g., the lowest level. A selected portion of the tree is stored in a data structure. FIG. 3 illustrates a selected portion 330 for storage. The selected portion 330 may include the first two levels. The selected portion 330 may include the three or more highest levels. The selected portion 330 may include a portion of a level, e.g., a portion of the third highest level.
[0069] In one embodiment, a tree parameter, e.g., a tree size parameter, is used during encoding, e.g., received at an input interface. The tree size parameter indicates how much of the hash tree should be stored in the data structure. The tree size parameter may indicate that a larger portion of the hash tree should be stored. The tree size parameter may indicate that a smaller portion of the hash tree should be stored. In this way, a user can indicate whether file size or verification speed for a selected portion should be optimized.
[0070] A corresponding validation system, e.g. validation system 160, may work similarly to encoding system 110 and / or encoding system 200. The validation system may work not only at the level of the end user, but also at an intermediate level, e.g. at the level of a server between the source of the data structure and the end user of the data structure. Validation may be performed at the server as well as at the end user's device.
[0071] For example, the verification system comprises an input interface configured to receive at least a portion of a digital data structure, the data structure including a plurality of genome blocks and a portion of the first hash tree including the first two highest levels of the first hash tree, but not including one or more lower levels of the first hash tree. It should be noted that the verification system does not need to receive the entire data structure generated by the encoding device. For example, only a portion of the genome data blocks in the data structure is received, e.g., only data blocks that are currently significant. The verification device also does not need to receive all of the hash tree. For example, a portion of the hash tree that relies only on data blocks that have not been transmitted or received can be summarized by sending only the hash value of the highest level of the hash tree that relies only on data that has not been received. For example, if blocks 1-4 have not been received, then the hashes of blocks 1-4 are needed, not the hashes that rely on blocks 1 and 2 or blocks 3 and 4. Assuming that block 5 has been received, the hashes that rely on blocks 1-4 are sufficient information about blocks 1-4 to verify the signature over the hash tree root. In other words, if a particular hash tree node depends only on blocks that have not been received, then no hash tree nodes below the particular hash tree node need to be sent or received, i.e., no hash tree nodes that the particular hash tree node depends on. If a higher hash tree node that the particular hash tree node depends on is sent, and the higher node depends only on data blocks that have not been sent, then even the particular hash tree node may not be needed. Although this is optional, the complete hash tree may be sent as long as the complete hash tree is available in the data structure. For example, the received hash tree may include all of some higher levels and omit all of some lower levels.
[0072] Using the obtained genomic data block, e.g., the received genomic data block, and the portion of the hash tree received, the root of the hash tree is recalculated. For the portions of the hash tree that depend on the received blocks, hash values can be recalculated, and for hash values that depend on blocks that have not been received, the received hash value can be used. The recalculated hash tree nodes, in particular the hash tree root, can be compared with the received hash tree nodes. Furthermore, if there is a signature on the hash tree root, it can be verified as well. If a mismatch is found in the hash values and / or signature, appropriate error handling is performed, e.g., an error is reported to the user, the file is rejected, etc.
[0073] Instead of validating all received genomic data blocks, the same approach is used to validate portions of received genomic data blocks.
[0074] Below are some further optional refinements, details and embodiments. The following embodiments can be applied to enable long-term integrity protection of genomic files or big data files in general. The embodiments are described in the context of ISO / IEC 23092, which defines a standard for encoding, compressing and protecting genomic data. Although the embodiments are advantageous in that context, they can also be applied outside of that context.
[0075] The current MPEG-G security solution ISO / IEC FDIS 23092-3:2019(E) has several security issues. There is no signature that covers the entire MPEG-G file. There are no signatures to check the overall integrity of the dataset group. In 7.4.2, the rfmd, dgmd, and dtpr boxes are protected. However, other boxes, such as dtcn, dghd, rfgn, labl, and lbll, are not protected. The signature protects a collection of access units containing genomic data, or other lower level data structures containing functional data, which can be resolved into bytes of the value field (as defined in ISO / IEC 23092-1:2019, 6.3). However, if it is a collection of access units, this can be a set of up to 256 blocks, each of which contains up to 256 blocks. 32 access units, which makes it difficult to check the validity of a selected access unit for random access without having to make calculations with respect to other access units protected in the same set. With data structures that are selectively and individually protected by digital signatures, current solutions generally cannot prevent unauthorized additions, deletions, or reordering of the data structures by an intruder. The process for verifying that a data structure belongs to a given file is not specified, nor is it specified how to verify that two data structures belong together as part of the same file.
[0076] Embodiments address one or more of these problems, including one or more of the following: Problem 1: Current MPEG-G solutions apply digital signatures to individual data structures (or containers). However, there is no way to efficiently group all the elements together to provide proof of the integrity of the entire file. Preferably, this is accomplished with low overhead while preserving random access. Problem 2: Genomic data is compressed and decompressed, and encrypted and decrypted. Compression algorithms, and new software implementations, can be prone to software bugs and / or design errors. Therefore, it is necessary to be able to verify the integrity and correctness of the Encrypted Compressed Data (ECD), Plaintext Compressed Data (PCD), and Plaintext Decompressed Data (PDD). Problem 3: The genome file is updated. A mechanism for data file integrity verification preferably enables this need. Problem 4: A genome file is updated. A structure for data file integrity verification preferably allows for tracking of changes and accountability for those changes. Problem 5: Dataset group IDs are very short (8 bits). This makes it infeasible to have unique identifiers that enable some use cases, such as merging or retrieval.
[0077] Embodiments propose to enhance genomic files, such as MPEG-G files, with a hierarchy of Merkle trees, with each Merkle tree bound to a data structure in the file, thereby enabling: Long-term integrity protection of all or selected data structures of a file as a complete whole, with low storage and computational overhead without impeding random access, allowing efficient verification of data structures at any level, including in dataset groups, and stitching them all together. Protection of the integrity of genomic data while it is being encrypted and compressed, while it is being compressed, and while it is being decompressed. Traceability and accountability. Improved performance compared to current solutions.
[0078] The reason for the emphasis on long term in the first bullet above is that the genomic information is important for the health of the user and his / her relatives. This means that the genomic information preferably remains private and private not only during the user's lifetime, but also during the lifetime of his / her children and grandchildren. Current MPEG-G solutions rely on traditional digital signatures that are not quantum resistant and can be cracked in a foreseeable time. One option to deal with this problem is to replace the existing digital signatures with quantum resistant digital signatures, but quantum resistant signatures are larger and slower than ECDSA. This therefore leads to a less efficient solution that requires individual signatures for each data structure under integrity protection. The proposed solution relies on Merkle trees (based on hash functions) and is therefore a natural solution for ensuring integrity in the long term, with the added benefit that only the root of the tree needs to be signed.
[0079] The reason that the proposed approach does not prevent random access is that it is possible to access the container and retrieve the needed MT nodes to verify that the data in the container has not been modified and that the container is part of the entire file, without having to access the data of the entire file.
[0080] Although the embodiments are described in the context of the MPEG-G standard with reference to its particular hierarchy of data structures, most of the features and functionality are generally applicable to any data format that organizes data into individual components. The embodiments are particularly useful for processing large amounts of data that are distributed and stored in a hierarchy of smaller data units.
[0081] Several embodiments are described below, each of which builds on the previous embodiment, for example, embodiment 2 builds on embodiment 1, embodiment 3 builds on embodiment 2, etc. Embodiments 1, 2, 3, 4, and 5 address the above problems 1, 2, 3, 4, and 5, respectively. There are a total of 14 embodiments.
[0082] Furthermore, after presenting the embodiments, we will explain how the embodiments compare to alternative solutions and how the embodiments can be applied to the GA4GH security solution to protect file integrity.
[0083] EMBODIMENT 1 This embodiment addresses problem 1. It associates hierarchical data structures (or containers) in an MPEG-G file with a Merkle Tree (MT). In general, higher level containers encapsulate lower level data structures. Because data structures in an MPEG-G file are organized in a hierarchical fashion, the roots of lower level MTs serve as leaves of higher level MTs.
[0084] In this embodiment, five levels of a Merkle tree are described {1, 2, 3, 4, 5}. The lowest level is 1 and the highest level is 5. At every level, the root of the Merkle tree is obtained by hashing selected leaf hash codes of the tree concatenated in a predefined order, such as the order in which the leaf data was stored. Hash() denotes a function that generates a hash code for the data structure specified in the brackets.
[0085] Level 1 (MT_AU; MT_DS) relates to MT at the Access Unit (AU) level and Descriptor Stream (DS) level in an MPEG-G file. Each access unit forms a Merkle tree with leaves that may include: Hash (AU Header), Hash (AU Information), Hash (AU Protection), and Hash(block) for each data block in the access unit. In general, a leaf contains all elements or a subset of elements in an AU container defined in ISO / IEC DIS23092-1, Table 25. The root of each access unit MT is denoted as MTR_AU. Each Descriptor Stream forms a Merkle tree with leaves that may contain: Hash (DS Header), Hash (DS Protection), and Hash(block) for each block in the descriptor stream. In general, a leaf contains all elements or a subset of elements in a DS container as defined in ISO / IEC DIS23092-1, Table 32. The root of each descriptor stream MT is denoted as MTR_DS.
[0086] Level 2-a relates to MT at the Attribute Group (AG) level in MPEG-G files. Each attribute group forms a Merkle tree with leaves that may include: Hash (AG Header) One or more MTR_AU In general, a leaf can be all elements or a subset of elements in an attribute group container as defined in ISO / IEC DIS23092-6. The root of each attribute group MT is denoted as MTR_AG.
[0087] Level 2-b relates to MT at the Annotation Table (AT) level in MPEG-G files. Each annotation table forms a Merkle tree with leaves that may contain: Hash(AT Header) Hash(AT Metadata) Hash (AT Protection) One or more MTR_AG In general, a leaf can be all elements or a subset of elements in an annotation table container as defined in ISO / IEC DIS23092-6. The root of each annotation table MT is denoted MTR_AT.
[0088] Level 3 (MT_DT) relates to MT at the Dataset (DT) level in an MPEG-G file. Each dataset forms a Merkle tree with leaves that may include: Hash(DatasetHeader) Hash(Dataset Metadata) Hash (Dataset Protection) Hash(Dataset Parameter Set) Hash (Master Index Table) One or more MTR_AU / MTR_DS / MTR_AT In general, a leaf may be all elements or a subset of elements in a dataset container as defined in Table 19 of ISO / IEC DIS23092-1. The root of each dataset MT is denoted as MTR_DT.
[0089] Level 4 (MT_DSG) relates to MT at the Dataset Group (DG) level in an MPEG-G file. Each dataset group forms an MT with leaves that may contain: Hash (DG Header) Hash(Reference) Hash(Reference Metadata) Hash (Label List) Hash (DG Metadata) Hash (DG Protection) One or more MTR_DTs In general, a leaf may be all elements or a subset of elements in a dataset group container as defined in Table 9 of ISO / IEC DIS23092-1. The root of each dataset group MT is denoted as MTR_DG.
[0090] Level 5 (MT_MPEGG) relates to the MT at the highest level of an MPEG-G file that groups together multiple data sets. An MT at the file level contains leaves that contain: Hash(File Header) One or more MTR_DGs The root of MT at the file level is indicated as MTR_MPEGG. At the file level, a new Merkle Tree Integrity Data container (mtid) is introduced to store MT data used for file integrity verification. The mtid box can be placed immediately after the File Header (flhd) or at the end of the file for easy updates. In its minimal form, the mtid box contains a signature on the entire root MTR_MPEGG and the ID of the signer, who may be the file owner or administrator. The signature is generated using the function Sign(PrivK, MTR_MPEGG), where PrivK is the private key of the signer with an associated public key certificate signed by a certification authority (CA) as proof of identity.
[0091] In this embodiment, five levels of MT are used, but the number of levels is not limited to five. The number of levels can be increased to accommodate additional levels of containers introduced in the MPEG-G standard, or decreased for file formats with fewer levels of containers.
[0092] Assuming a binary Merkle tree is used to compute the root from the leaves, a technique for verifying the leaves is described in US Pat. No. 4,309,569, Method of providing digital signatures.
[0093] This process is illustrated in Figure 4 with an exemplary Merkle tree with four leaves L0, L1, L2 and L3 generated by hashing on four data elements D0, D1, D2 and D3, respectively, that are being protected for global integrity. The tree has a highest level 410, a second highest level 420, a leaf level 430 and data 440.
[0094] For example, D0 is an AU with ID, D1 is metadata, and D2-D3 are genome data blocks. Note that in transport mode, to send only stippled blocks, other stippled nodes may also be sent to allow recomputation of the root node. The root node, or signatures N0-3 on the root node, may also be sent to allow verification of the recomputed root value.
[0095] The tree includes two internal nodes N0-1 and N2-3 and a root N0-3. For example, to verify data element D1, the leaf L0 and the intermediate node N2-3 are disclosed. Given data element D1, it is possible to compute L1 as hash(D1). Using L0 and L1, it is possible to compute N0-1 as hash(L0|L1), and using N0-1 and N2-3, it is possible to compute the public root N0-3 as hash(N0-1|N2-3).
[0096] Note that there is a distinction between the levels of individual Merkle trees, which correspond to data structures in MPEG-G, and the levels of the internal nodes in the binary Merkle tree. The external Merkle tree levels range from the lowest level, associated with an access unit or a descriptor stream, to the highest level, associated with a top-level data container, such as a data set group. 2 n For a binary Merkle tree with data elements, the lowest level has 2 n 2 generated by individual hashing on data elements n leaves, then there are n levels of intermediate nodes, l=1,...,n. At each level l, the number of nodes is 2 (n-l) The root node of a binary Merkle tree is at level n.
[0097] When building a binary Merkle tree, if there are an odd number of leaves or nodes at a particular level, the last leaf or node can be concatenated with itself. Consider an example of a 5-node with leaves A, B, C, D, and E. The process is as follows: Leaf level: There are 5 leaves, A and B connect to node AB, C and D connect to node CD, E is alone and therefore connects to node EE. Level 1: There are three nodes, AB and CD connect to node ABCD, EE is a single node, hence EEEE. Level 2: Finally, there are two nodes ABCD and EEEE giving the root node ABCDEEEE.
[0098] Embodiment 2: This embodiment addresses problem 2. It involves introducing level 0 in embodiment 1, where level 0 is lower than level 1. This is done as follows. Hash(Block), which is used as a leaf in the Merkle tree for each access unit or descriptor stream, contains (up to) three values for that block: Hash(ECD), hash of compressed and encrypted blocks Hash(PCD), hash of the compressed block Hash(PDD), the hash of the block after decompression is replaced by the MT root of block MTR_block calculated from
[0099] All three values are not always required: depending on which values are included and verified, it is possible to check with more or less precision the reasons for lack of integrity.
[0100] Each of these hash values is assigned to a leaf. This allows pinpointing errors, e.g. as compression errors. Note that the integrity can be checked at different places: the server cannot check the decrypted data if it does not have access to the decryption key, but the server can check the encrypted data.
[0101] Furthermore, in one example, the root of the Merkle tree is MTR_Block=Hash(Hash(Hash(ECD)|Hash(PCD))|Hash(PDD)). Note that in this case a different calculation for the root, e.g. following a ternary tree structure, is also valid, e.g. Hash(Hash(ECD)|Hash(PCD)|Hash(PDD)). This makes it possible to check the integrity of the data at different stages of the decoding and decompression process, and at different locations (e.g., on the client and in the cloud) in a distributed implementation of the standard.
[0102] For example, assume that an MPEG-G client requests data from an MPEG-G file stored in the cloud. The cloud may check the user's credentials and use the information in the Merkle tree, i.e., Hash(ECD), to check that the encrypted and compressed data has not been altered. The cloud then sends the encrypted and compressed data to the MPEG-G client, which is (i) compressed to save bandwidth and (ii) encrypted since the cloud itself does not have access to the user data. When the MPEG-G client receives the data, it may first use Hash(ECD) to check the integrity of the data. It may then use Hash(PCD) to decrypt the data and check the integrity of the decoded and compressed data. Finally, it may decompress the data and use Hash(PDD) to check the integrity of the decoded and decompressed data.
[0103] Embodiment 3: This embodiment addresses problem 3. To this end, when data is added and / or modified in a data structure at levels 0, 1, 2, 3 or 4: A new / updated MT for the new data structure at level i is calculated. MTR_MPEGG, and therefore the following process is performed until the signature on MTR_MPEGG is updated at level 5. A new (or updated) leaf will be added to the next higher level MT. The value of the new (or updated) leaf is the value of the MTR of the MT in the current level i. Since the MT at the next higher level has new / updated leaves, its root is also updated.
[0104] Embodiment 4: This embodiment addresses problem 4. To this end, one option is to detect changes / updates ( **The first step is to include a tracking table in the mtid box described in embodiment 1 to keep track of the signatures generated by the file. Since the creation of the file, the i-th entry in the table, for i=0, 1, 2, ..., corresponding to the i-th generated signature, contains: All or selected root values MTR_DG, MTR_DT, MTR_AT, MTR_AG, MTR_AU, MTR_DS and Merkle tree nodes to enable efficient validation and storage. For further details, see embodiments 6-8. A signature on the current MTR_MPEGG, denoted SIGNATURE_MTR_MPEGG(i), computed as Sign(PrivK, i|MTR_MPEGG|SIGNATURE_MTR_MPEGG(i-1)), where SIGNATURE_MTR_MPEGG(-1)="" PrivK is the signer's private key that has an associated public key certificate signed by a certification authority (CA) as proof of identity. For long-term integrity protection purposes, hash-based signature algorithms such as LMS, i.e., e.g., RFC8554, or one of the quantum-resistant algorithms, e.g., Falcon, SPHINCS (currently being standardized by NIST), are preferred. In particular, if a pre-quantum signature algorithm (e.g., ECDSA) is used in some entries, then it may be impossible to verify those entries as soon as ECDSA is cracked, e.g., by a quantum computer. Signer's ID When a file is modified at time i, i=1, 2, ... MTR_MPEGG is updated as described in embodiment 3. A new entry is included in the trace table in mtid. To facilitate tracing, the trace data may include: SIGNATURE_MTR_MPEGG(i) The new added data (if data is added) or the difference between the changed data compared to the previous data. Leaf that has been modified or added. moreover, When data at a leaf is accessed, file integrity and all relevant changes for that piece of data can be checked by traversing a path up to the current Merkle tree root. When searching for a particular data item in the tracking table in the mtid, all the information about the addition / change of that piece of data is found.
[0105] Embodiment 5: This embodiment addresses problem 5. It does this by: Make the datasetGroupID long enough (e.g., 256 bits) and generate it randomly. In particular, replace datasetGroupID with MTR_DG. To ensure that datasetGroupID is unique, an additional input in the calculation of MTR_DG is included, namely a sufficiently long random nonce N_MTR_DG. Similarly, datasetID can be replaced with MTR_DS. To ensure that the datasetGroupID is unique, an additional input in the calculation of MTR_DG is included, namely a sufficiently long random nonce N_MTR_DG.
[0106] The benefit of this change is that each dataset or dataset group is assigned a unique identifier. This is an advantage since the current MPEG-G standard includes use cases for merging files and APIs to retrieve entries based on these identifiers. If the identifier range was only 8 bits, merging two dataset groups may require the identifiers to be renamed. If the identifiers were not renamed, calls to the API may return incorrect results. Using MTR_DG and MTR_DS as datasetGroupID and datasetID solves these two problems.
[0107] Embodiment 6: MT storage overhead and computational performance trade-offs and optimizations: In general, for data structures with a small number of instances, the entire Merkle Tree (MT) can be stored to improve the efficiency of integrity verification without incurring large storage overhead. On the other hand, for data structures with a large number of instances, a trade-off between storage overhead and computational resources is usually preferable. 2 n For a binary MT with leaves, 2 n If all the leaves are available, the computation of the MT node and the root takes (2 n -1) hash operations are required. In binary MT, all intermediate nodes and the root are included, for a total of 2 n -There is 1 node.
[0108] This embodiment proposes an optimization to store the top m levels of nodes, where 1 ≤ m ≤ n, counting from the root of the binary MT. If this set of nodes were stored together with all the leaves, the computational burden of verifying a leaf would be [2 (n-m+1) −1+(m−1)] hash operations. Storing the leaves avoids the expensive operation of hashing on the data structure, especially if the data structure is large.
[0109] This optimization is illustrated in Figure 5a, where the stippled triangles represent a partial Merkle tree. In Figure 5a, the root node 511 at the highest level 510, say level n, is shown. Also shown in 520 are levels n-(m-1), 530 are levels nm, 540 are levels 1, and 550 are the leaf level (level 0). 522 shows m levels, and 521 shows n levels. Level 520 is 2 n-(m-1) Contains nodes. Level 530 contains 2 n-m The dotted triangular part contains 2 nodes. m - Contains 1 node. Level 1 is 2 n-1 Level 0 has 2 nodes. nIt has nodes.
[0110] Triangle 560 indicates nodes that are generated on the fly when they are needed. n-m+1 -1 node. The non-stippled parts of the tree 500, e.g., below level 530, are not remembered. The top levels, e.g., from level 520 onwards, are remembered.
[0111] The entire Merkle tree, in this example, is 2 n leaves. Both leaves and nodes at the top m levels, including the root, are stored. Triangles 560 at the lowest (n-m+1) levels are recomputed on the fly when a particular leaf needs to be verified.
[0112] Many leaves, e.g. 2 32 There can be as many leaves as there are leaves. There can be more leaves if needed. The block size can also vary. A smaller block size allows for finer grained integrity checking with less overhead, but may increase the size of the hash tree. For example, a tree size parameter is used to determine the block size.
[0113] The stored portion of MT can be part of a file, it can be stored as a separate file, it can be stored as part of a data structure.
[0114] For example, if n=32 and m=24, assuming a hash code size of 32 bytes, that is (2 24 -1) hashing (approximately 2 24* 32 bytes) and 2 32 The number of leaves and the number of MTs are stored in memory. Validating a leaf requires (2 9 -1) hash operations, plus 23 hashes to reach the root. (2 9+22) operations are very fast. The memory overhead depends heavily on the number of leaves in the data structure.
[0115] Alternatively, 2 n If m leaves are not stored, then by just storing the top m levels of nodes, the number of leaves that need to be calculated for integrity verification can still be reduced by 2 (n-m+1) In this case, the storage overhead is reduced to (2 m -1) hashes, the time to recompute the Merkle tree root is T binary_MT =[2 (n-m+1) -1+(m-1)] * t node +2 (n-m+1)* t leaf is given by Here, t node is the time for hashing on two nodes, and t leaf is the time to generate a leaf by hashing on the data structure. Without leaf or node memory, the computation time is T binary_MT =[2 n -1] * t node +2 n* t leaf become.
[0116] The operation of binary MT in a client-server setting, where the files reside on the server and the validation is performed on the client, is further explained using the example shown in Fig. 5b, where n = 4 and m = 2. Fig. 5b shows a hash tree with levels 4 to 1 and a leaf level. Level 4 has node R, level 3 has nodes N31 and N32, level 2 has nodes N21, N22, N23 and N24, level 1 has nodes N11 to N18, and the leaf level has nodes L1 to L16 corresponding to data D1 to D16.
[0117] Assume that all nodes at levels 3 and 4 and all leaves have been precomputed and stored, for example, in a data block. To verify D12 (highlighted), the following steps are taken: 1. The server generates the (highlighted) nodes {N15, N17, N18, N24}. In general, the number of hash operations on the server is {[2 (n-m+1) -1]-(n-m+1)}. 2. The server sends hashes {L11, N15, N24, N31} to the client along with the root signature. In general, the number of hashes that need to be sent is n. 3. The client: Leaf L12 is hashed on data component D12 Generate Using the hashes received from the server, generate the nodes (highlighted): {N16, N23, N32, R}. In general, the number of hash operations on a pair of hashes at the client is n. 4. The client decrypts the received signature using the associated public key and compares the decrypted hash code with the generated root R. If they match each other, the verification is successful.
[0118] Compare the cost of this approach to the current MPEG-G requirement of one digital signature per data structure being integrity protected. As soon as two data structures need to be verified, the computational overhead is reduced because the cost of digital signature verification is much more expensive than a hash calculation. The MT approach involves a single signature verification for the MT root of the file. Assuming the same data structures are integrity protected, the storage overhead is reduced since MPEG-G currently requires a digital signature for each data structure, and the signature size is larger than the hash value, especially if quantum-safe signatures are used.
[0119] The table below summarizes the characteristics of different levels of a Merkle tree, including the tree name, the maximum number of leaves, storage needs, the number of hash operations to validate a leaf, and the number of interior levels in a binary tree.
[0120] [Table 1]
[0121] Embodiment 7: Using Authentication Tags Instead of Hash Codes for Encrypted Data Containers It should be noted that MPEG-G allows protecting different containers with symmetric keys by using AES in GCM mode. This means that data containers can also be encrypted and authenticated by using corresponding symmetric keys. Authentication is performed by checking a stored authentication tag that depends on the entire data container. The authentication tag in the existing MPEG-G solution can be reused as a fingerprint of the data container to reduce CPU and memory requirements. In particular, the AES-CCM authentication tag can replace, for example, Hash(block) in level 1 of embodiment 1 or Hash(ECD) in embodiment 2. If the authentication tag is reused as a fingerprint, the hash of the data container (MT leaf) does not need to be recalculated or stored.
[0122] Embodiment 8: Selective Integrity Protection In one embodiment, all data structures or containers are protected. However, in some situations, a user may choose partial protection. This embodiment describes how selective protection of data structures can be achieved.
[0123] To this end, each container type has a field to indicate how the data structure of that container type is generally selected. For example, the value of the field could be: 0 - none selected (the Merkle tree excludes all data structures below this container level) 1-All selected 2-Select a Data Structure by ID 3-Selecting data structures by flags, e.g., using a sequence of bits to indicate the selection of data structures by their position in a higher-level container
[0124] Below are example settings for the general selection mode of dataset group containers, dataset containers, annotation table containers, access unit containers, and block containers for their inclusion during MT validation. mt_select_dg=1 (select all dataset groups) mt_select_dt=2 (select dataset by ID) mt_select_at=2 (select annotation table by ID) mt_select_au=2 (select access unit by ID) mt_select_bl=1 (select all payload blocks)
[0125] These general preference settings for each container type can be overridden for each individual container. Note that when a container is excluded, all its subordinate data structures are excluded regardless of their preference settings.
[0126] Embodiment 9: Selective storage of hash codes to improve the speed of integrity verification Storing hash codes can improve the speed of signature verification, but at the expense of storage overhead. As described in embodiment 6, the closer the container level is to the root (fewer hash codes), the lower the storage cost and the greater the benefit of improving computation speed (each hash code represents a larger chunk of data). In this embodiment, we describe how MT hash codes (or nodes) can be selectively stored.
[0127] For this purpose, each container type has a field to indicate how its corresponding MT hash code, including both leaves and other nodes in the binary tree (see embodiment 6), should generally be selected for storage. For example, the value of the field could be: 0 - no nodes 1- All nodes in the binary MT, including all leaves and intermediate nodes 2- All leaves only 3- Nodes at the top m levels of binary MT 4 - first n leaves 5- All leaves and nodes at the top m levels of the binary MT 6-The first n leaves and nodes at the top m levels of the binary MT
[0128] Below are example settings for the general storage modes of MT at the file (MPEGG) level, dataset group (DG) level, dataset (DT) level, annotation table (AT) level and access unit (AU) level. mt_store_mpegg=1 (store all MT nodes at the entire file level, with the leaves being MTR_DG) mt_store_dg=1 (store all MT nodes at the dataset group level with their leaves being MTR_DT) mt_store_dt=1 (store all MT nodes at the dataset level with their leaves being MTR_AT) mt_store_at=1 (store all MT nodes at annotation table level with leaves that are MTR_AU) mt_store_au=0 (MT nodes at the access unit level are not stored)
[0129] For storage modes 3-6, additional fields are required to specify the number of top levels (m) and the number of leaves (n) in the binary MT to be stored. These general MT storage settings for each container type can be overridden for each individual container.
[0130] Embodiment 10: Full or targeted integrity verification If some or all of the hash codes on a Merkle tree are stored to verify the integrity of a particular component, then only the hash code for the component in question needs to be regenerated, then all hash codes from the component in question back up to the root of the tree are verified, and finally the signature is checked.
[0131] Embodiment 11: Flexible Merkle Tree Organization - Binary Trees vs. K-ary Trees, Single Tree Hierarchy vs. Multiple Tree Hierarchies The description in the above embodiments focuses on a hierarchy of Merkle trees organized similarly to the MPEG-G file data structure. Each protected data structure is associated with an independent Merkle tree. The embodiments generally assume a binary Merkle tree, e.g., each node in the tree has at most two child nodes. Thus, a data file, such as an MPEF-G file, has a hierarchy of binary Merkle trees, each corresponding to a data container. As an alternative to the binary tree configuration, a K-ary Merkle tree contains nodes each computed by hashing simultaneously on k subnodes. In the extreme case, k can be equal to the total number of leaves, thereby resulting in a tree depth of 1. In other words, the root is computed by direct hashing on the concatenation of all N leaves. It is assumed that the hash function has a linear time complexity of O(N), and for ease of comparison with the binary model, N=2 n Then, the time to recalculate the root of an N-ary Merkle tree is T N-ary_MT =2 (n-1)* t node +2 n* t leaf where t node is the time for hashing on two nodes, and t leaf is the time to generate a leaf by hashing on the data structure.
[0132] As described in embodiment 6, the time to recalculate the root of a binary Merkle tree, where the top m levels of nodes are stored, is T binary_MT =[2 (n-m+1) -1+(m-1)] * t node +2 (n-m+1)* t leaf is given by:
[0133] 1≦m≦n, (2 mComparing the two methods with the storage overhead of (1)t leaf >>t node (2) the number of leaves is relatively small, for example, for MT_AU, where the maximum number of leaves is 256; and (3) assuming there is no mixed storage of nodes and leaves for the binary Merkle tree approach, T N-ary_MT ≒(2 n -2 m +1) * t leaf T binary_MT ≒2 (n-m+1)* t leaf When m=n, T N-ary_MT ≒t leaf However, T binary_MT ≒2 * t leaf It is. When m=nl and 1≦l≦(n-1), T binary_MT ≒2 (n-m+1)* t leaf =2 (l+1)* t leaf T N-ary_MT ≒[2 n -2 (n-l) +1] * t leaf >2 (l+1)* t leaf * {2 [n-(l+1)] -2 [n-2(l+1)]}>2 (l+1)* t leaf ≒T binary_MT It is.
[0134] Based on the above calculations of the speed of integrity verification, under the assumption that the computation time of hashing on a node is negligible compared to generating a leaf, we can conclude the following: The memory overhead is (2 n The binary MT approach is faster when m=(n-1) or the memory overhead is [2 (n-1)-1], T binary_MT ≒4t leaf This computation time continues until the memory overhead becomes large enough to cover the next level, e.g., m=n or the memory overhead becomes (2 n -1), whereas the N-ary approach has a storage overhead of (2 n -4) hashes or less, T N-ary_MT ≧4t leaf It is. The N-ary tree method has a storage overhead of (2 n -3) hash and 2 n It is faster when there is a gap between hashes. The binary MT method has a storage overhead of 2 n When the hash count exceeds 2 n leaves may be stored first, and the remaining overhead may be used to store the internal MT nodes, thereby leading to further improvement in verification speed.
[0135] Considering only the computation time of hashing on a node, T binary_MT_nodes_only =[2 (n-m+1) -1+(m-1)] * t node <2 (n-1)* t node <T N-ary_MT_nodes only Therefore, the binary approach always outperforms the N-ary approach when m>1.
[0136] Overall, therefore, the binary MT approach corresponds to the case of storing only the top m ≥ (n-1) levels of nodes, or storing all leaves + root (m = 1). n -3) hashes from (2 n +2) hashes. Under such a particular storage overhead constraint, there is no clear winner in verification speed.
[0137] These analysis results can provide guidance for choosing between the binary and N-ary Merkle tree approaches based on the number of leaves selected (embodiment 8) and the storage mode setting (embodiment 9). In a client-server setting where the files reside on the server and the validation is done at the client, the binary approach sends only n nodes to the receiver, while the N-ary approach sends only n nodes (2 n -1) leaf nodes need to be sent.
[0138] There may also be advantages in allowing more flexible Merkle tree organization. One flexibility is to allow having multiple Merkle trees per file, e.g., MT_1, MT_2, ..., where each Merkle tree contains data belonging to the same cohort with integrity to be protected as a complete whole independently, or where each Merkle tree has a different set of parameters for specific integrity protection requirements. It is also possible to support nested Merkle trees, e.g., by merging k Merkle trees into one big tree with the MT-overall root computed as the hash of the roots of the k Merkle trees one level deep, MTR_Overall=Hash(MTR_1|MTR_2|...|MTR_k).
[0139] Embodiment 12: Partitioning into Merkle trees for function components and data components An alternative embodiment related to embodiment 10 is to organize the function components and data components into separate Merkle tree hierarchies and then combine those Merkle tree hierarchies to form a new root whole. The function components include header, metadata, and protection structures at different container levels, and the data components correspond to structures containing payload block data. Since the function components are generally critical and small in size, it seems reasonable to allow all the function components to be protected by the Merkle tree without selection, while allowing selective protection of the data components. One potential advantage of such an arrangement is faster verification of the function components. For example, to verify the metadata of a data set, the hash codes of the individual access units in the data set are not involved in the verification. Only the hash code at the root of the Merkle tree for the data components is used for verification. The following example illustrates the idea of splitting the function components and the data components into two separate Merkle trees and then combining them into a single root whole. MTR_MPEGG MTR_Functional MTR_Functional_DG_1=Hash(Header_DG_1|Metadata_DG_1|Protection_DG_1|MT_Functional_DT_11|MT_Functional_DT_12|…) MTR_Functional_DG_2=Hash(Header_DG_2|Metadata_DG_2|Protection_DG_2|MT_Functional_DT_21|MT_Functional_DT_22|…) … MTR_Data MTR_Data_DG_1 MTR_Data_DG_2 …
[0140] In this example, the whole root MTR_MPEGG depends on the roots of the Function Merkle Tree and the Data Merkle Tree as Hash(MTR_Functional|MTR_Data). MTR_Functional is obtained by hashing on the concatenation of Function MT roots at the Dataset Group level. Such a root, prefixed by MTR_Functional_DG, is generated by hashing on the concatenation of the Dataset Group Header, Metadata, Protection, and the lower Function MT roots at the Dataset level. This resembles a K-ary Merkle Tree, since in generating the root value, the hash function takes multiple inputs and data components at the same level.
[0141] Example 13: Time stamp A timestamp can be included in each Merkle tree data structure to indicate the freshness of the Merkle tree integrity protection. The timestamp is stored and signed along with the top-level Merkle tree root MTR_MPEGG. Additionally, a freshness period can be imposed so that the signature on the entire MT root is automatically regenerated with updated timestamps before the MT data expires. An outdated timestamp indicates that the file may not be up-to-date and may be an older version used by an intruder to masquerade as the current version and undo recent changes.
[0142] Alternatively, when some data structures are characterized by a specific date, that date is used as an input (e.g., as an additional leaf) in those data structures. For example, assume that a user has performed several analyses on a patient's genomic data over the course of a few days. If the analyses are reflected in multiple data structures that were modified at different moments in time (different times, different days), each of the data structures may have a different timestamp when building the entire new hierarchy of the Merkle tree. The user may include a timestamp when signing the MTR_MPEGG. The signature timestamp is preferably the most recent timestamp.
[0143] Embodiment 14: Use for transport mode MPEG-G defines options for data storage and transport. In one embodiment, when data is transported, messages can also be integrity protected.
[0144] When a data structure, for example a block of data, is sent, the message may begin with: The nodes in the Merkle tree that are needed to verify the integrity of a data structure / block of data, and File signature.
[0145] This makes it possible to verify the integrity of blocks of data that are part of a file, without first having to receive the entire file.
[0146] The message may include, for example, a signature on MTR_MPEGG if this is the first message. Subsequent messages do not need to include this signature again as it would be redundant. This signature also only needs to be verified for the first message. This is shown in FIG. 6, where n blocks of data are retrieved from the server by the client. These n blocks of data are exchanged in n messages. The first message includes the file signature, the signer's public key (if not already available at the receiver), the nodes involved in the Merkle tree path needed to verify the first block of data, and the block of data itself. Once the client receives this information, it computes a hash on the received data (block 1), calculates the root of the tree using the generated hash and the required nodes in the path, and verifies the signature using the generated root and the signer's public key.
[0147] 6 shows a client 610 and a server 620. Message 1 (630) includes a signature (631), a node (631) involved in the Merkle tree path needed to verify block 1, and block 1 (633). Message 2 (640) includes a node (641) involved in the Merkle tree path needed to verify block 2, and block 2 (642). Message 3 (650) includes a node (651) involved in the Merkle tree path needed to verify block 3, and block 3 (652).
[0148] Processing of subsequently received messages is similar, except that since the signature is unique for the entire file, there is no need to include / check the signature any more.
[0149] Comparison with alternative solution bases The above embodiments use Merkle trees in their data structures to achieve long-term integrity protection of MPEG-G files. Alternative solutions include modifying / extending the current MPEG-G solution in 23092-3:2019-3, 7.4.
[0150] For example, an alternative solution may involve two changes: the use of a long-term signature algorithm and defining a procedure that allows binding signed containers.
[0151] The first point can be addressed by using a quantum-resistant algorithm, e.g. a hash-based signature algorithm such as LMS, XMSS or SPHINCS, instead of ECDSA.
[0152] Regarding the second point, different approaches can be considered. Below, three options are described: 1. The first approach involves including a unique file identifier, for example a randomly generated 256-bit long identifier, in the file header. This file header is signed by a system administrator. To link all data containers in a file, the MPEG-G specification is updated in such a way that when a container is signed, this signed container includes a unique file identifier. If this is done, when a user retrieves a data container from a file, the user must: The signature main file header is retrieved and verified, and the file identifier ID_file is extracted. It takes the data container, verifies its signature, and checks that the ID in the data container is equal to ID_file. 2. The previous approach is that data in different MPEG G files can be merged. When two files are merged, a new file header is created that contains the new file identifier. The file header can contain a merging history that contains the identifier of the previous file. If this is done, when a user extracts a data container from the file, the user must: The signature main file header is retrieved and verified, and the file identifier ID_file is extracted, as well as the previous file identifier from the previous unmerged file. It takes the data container, verifies its signature, and checks that the ID in the data container is equal to the ID_file or one of the ID_files from the previous unmerged file. 3. Another alternative approach involves having a unique long identifier in each container (e.g., a randomly generated 256-bit long identifier). When a container is signed, the signature may include the identifier of the parent container. If the signature of a container is to be verified, the following process is followed: Verify the container signature and extract the parent container ID Repeat the following until the parent container ID corresponds to an MPEG-G header file: Verify the signature Get the parent container's ID 4. Another option is to use protection boxes as specified in MPEG-G Part 3 in section 7.4 and define a process to combine them. For example, when a block of data in an access unit is to be verified, one or more or all of the following steps may be performed: 1. A digital signature must be present in the access unit protection for a given block. 2. This first signature must verify successfully. 3. There must be an access unit protection signature for a given access unit at the dataset level in the Data Set Protection box. 4. This second signature must verify successfully. 5. There must be a Data Set Protection Box signature for a given dataset at the Data Set Group level in the Data Set Group Protection Box. 6. This third signature must verify successfully.
[0153] If any of the steps fail, the integrity verification process fails. If a step is successful, the next step is evaluated.
[0154] Even if all the above steps are performed individually, the current MPEG-G standard lacks this verification process that combines them. Furthermore, there is no definition of a signature on the whole file that combines all the data set groups.
[0155] Comparing the above method with the embodiment using Merkle trees, the embodiment using Merkle trees is more efficient and comprehensive for several reasons. a) An MT-based solution makes it possible to verify not only the integrity of individual components, but also that they belong together as part of an entire MPEG-G file. It prevents unauthorized addition, deletion or reordering of data structures in the file. b) The first reason for this is that the MPEG-G approach requires storing a signature for each container, and hash-based signatures (e.g., SPHINCS) are large (see, e.g., the paper "SPHINCS: practical stateless hash-based signatures" by Bernstein et al.). c) The computational overhead is much larger, since the hash-based signature for each container needs to be verified independently.
[0156] Application example to ensure integrity protection in GA4GH files The Global Alliance for Genomics and Health (GA4GH) describes in its paper ("GA4GH File Encryption Standard", October 21, 2019) how to encrypt and integrity protect the individual blocks of 64 kilobytes that are subsequently exchanged. Each of the blocks is encrypted and a Message Authentication Code (MAC) is appended. However, that solution does not prevent an attacker from inserting, deleting, or reordering entire blocks during communication (this is described in Crypt4GH, section 1.1).
[0157] An approach to dealing with this problem involves defining a Merkle tree where each of these blocks of 64 kilobytes is a leaf. This is equivalent to level 1 in embodiment 1. The blocks of data can also be individual smaller Merkle trees, as in embodiment 2.
[0158] The root of the Merkle tree can be signed, or alternatively, a MAC can be computed using the same key and MAC algorithm, as in Crypt4GH.
[0159] Assume there are n blocks of 64 kilobytes. When a block of data is to be transmitted, the following information should be sent: a) a signature (or MAC) on the Merkle tree root, and b) the corresponding log(n) nodes in the binary Merkle tree that allow for examining the block of data (this includes the block index and the Merkle tree node ID); c) An encrypted, MAC-protected block of data (up to 64 kilobytes in size, according to Crypt4gh).
[0160] Please note that due to potential fragmentation problems when transporting data, if the block size is chosen to be B=64 kilobytes, the additional integrity data in points a) and b) above may be taken into account, which means that the total size of the blocks may be up to B=64 kilobytes, and the data blocks may be somewhat smaller due to the overhead caused by a) and b).
[0161] Upon receiving a message, the receiving party recalculates the root of the Merkle tree using the received log(n) nodes in the Merkle tree to locate the block of data. The receiving party then checks the signature on the Merkle tree. Finally, the receiving party checks the received block of data. Note that the root signature only needs to be verified for the first message. This is similar to the case in embodiment 14.
[0162] Notwithstanding the statements in Crypt4GH, section 1.1, it should be noted that the current standard for Crypt4GH may also provide some integrity protection to prevent unauthorized insertion, deletion or reordering of data blocks, for example, if an index is included in the block of data. This feature is independently important. This index identifies the relative position of the block. If this index is included in each exchange block in Crypt4GH, it will be impossible to reorder the blocks any further. The receiver may also check for the absence of duplicates based on this index. If the receiver receives a block with index k, the receiver may also check if it has received all blocks with block index 1, 2, ..., (k-1).
[0163] This last "index-based" approach differs from the MT-based approach. One advantage is that while the "index-based" approach is about how the data is sent, the MT approach also gives information about how the data is stored. This means that if an attacker can influence the process of sending packets, he can put the indexes in the right order but swap the blocks. The receiving party may then assemble the file in the wrong order. This is not feasible with the MT-based approach, since MT gives information about how the blocks are organized and stored in the file.
[0164] Figure 7a illustrates a schematic of an example of an embodiment of an encoding method 700 for encoding data in a digital data structure. The encoding method comprises: Obtaining (705) the data as a plurality of data blocks; Calculating (710) a plurality of hash values for the plurality of data blocks by applying a hash function to the plurality of data blocks; Computing (715) a first hash tree for the plurality of hash values, where the plurality of hash values are assigned to leaves of the first hash tree and one or more higher levels of the first hash tree are generated; including (720) a plurality of data blocks and a portion of a first hash tree in a data structure, the portion of the first hash tree not including at least a portion of the leaves of the first hash tree; has. Obtaining the plurality of genomic data blocks may include directly receiving the genomic data blocks.
[0165] 7b illustrates generally an example of an embodiment of a validation method 750 for validating selected genomic data in a digital data structure. The validation method 750 includes: receiving at least a portion of a data structure (755), the data structure including a plurality of data blocks and a portion of a hash tree, the data structure not including at least a portion of a leaf of the first hash tree; calculating (760) a plurality of hash values for the selected data blocks by applying a hash function to the data blocks in the data structure selected for data integrity verification; determining (765) a path starting from the data block selected for verification to the root of the corresponding hash tree; retrieving (770) a hash value for the hash tree along the path, if available in the data structure, or computing such a hash value; a step (775) of verifying the root of the hash tree from at least the calculated hash values; has.
[0166] As will be apparent to one skilled in the art, many different ways of performing the method are possible. For example, although the order of steps may be performed in the order shown, the order of steps may be changed or some steps may be performed in parallel. Moreover, steps of other methods may be inserted between steps. The inserted steps may represent improvements to the method as described herein and may not be related to the method. For example, some steps may be performed at least partially in parallel. Moreover, a given step may not be fully completed before the next step is started.
[0167] The method embodiments may be implemented using software, including instructions for causing a processor system to perform the method 700 and / or 750. The software may include only steps taken by a particular subentity of the system. The software may be stored on a suitable storage medium, such as a hard disk, floppy, memory, optical disk, etc. The software may be sent as a signal, over a wire, or wirelessly, or using a data network, such as the Internet. The software may be made available for download and / or for remote use on a server. The method embodiments may be implemented using a bitstream configured to configure programmable logic, such as a field programmable gate array (FPGA), to perform the method.
[0168] It will be appreciated that the presently disclosed subject matter also extends to computer programs, particularly computer programs on or in a carrier adapted to practice the disclosed subject matter. The programs may be in the form of object code, such as source code, object code, code intermediate source, and partially compiled form, or in any other form suitable for use in implementing an embodiment of the present method. An embodiment of a computer program product includes computer executable instructions corresponding to each of the processing steps of at least one of the described methods. These instructions may be subdivided into subroutines and / or stored in one or more files that are statically or dynamically linked. Another embodiment of a computer program product includes computer executable instructions corresponding to each of the devices, units and / or parts of at least one of the described systems and / or products.
[0169] Fig. 8a shows a computer readable medium 1000 having a writable portion 1010 and a computer readable medium 1001 also having a writable portion. The computer readable medium 1000 is shown in the form of an optically readable medium. The computer readable medium 1001 is shown in the form of an electronic memory, in this case a memory card. The computer readable medium 1000 and 1001 store data 1020 which, when executed by a processor system, represent instructions that cause the processor system to perform an embodiment of a method for encoding or verifying genomic data according to an embodiment. The computer program 1020 is embodied on the computer readable medium 1000 as a physical mark on the computer readable medium 1000 or by magnetization of the computer readable medium 1000. However, any other suitable embodiment is conceivable. Further, although the computer readable medium 1000 is illustrated herein as an optical disk, it will be appreciated that the computer readable medium 1000 may be any suitable computer readable medium, such as a hard disk, solid state memory, flash memory, etc., and may be non-recordable or recordable. The computer program 1020 includes instructions for causing a processor system to perform the method of encoding and / or verifying data.
[0170] Fig. 8b shows a schematic representation of a processor system 1140 according to an embodiment of the encoding and / or verification system. The processor system comprises one or more integrated circuits 1110. In Fig. 8b, the architecture of the one or more integrated circuits 1110 is shown in a schematic manner. The circuit 1110 comprises a processing unit 1120, e.g. a CPU, for running computer program components for performing the method according to an embodiment and / or for implementing modules or units thereof. The circuit 1110 comprises a memory 1122 for storing programming code, data, etc. Parts of the memory 1122 may be read-only. The circuit 1110 comprises a communication element 1126, e.g. an antenna, a connector or both. The circuit 1110 comprises a dedicated integrated circuit 1124 for performing some or all of the processing defined in the method. The processor 1120, the memory 1122, the dedicated IC 1124 and the communication element 1126 are connected to each other via an interconnect 1130, e.g. a bus. The processor system 1110 may be configured for contact and / or contactless communication using antennas and / or connectors, respectively.
[0171] For example, in one embodiment, the processor system 1140, e.g., an encoding or verification system, comprises a processor circuit and a memory circuit, the processor configured to execute software stored in the memory circuit. The memory circuit may be a ROM circuit, or a non-volatile memory, e.g., a flash memory. The memory circuit may be a volatile memory, e.g., an SRAM memory. In the latter case, the device may comprise a non-volatile software interface, e.g., a hard drive, a network interface, etc., configured to provide the software.
[0172] Although device 1140 is shown as including one of each described component, various components may overlap in various embodiments. For example, the processor may include multiple microprocessors configured to execute the methods described herein alone, or may include multiple microprocessors configured to execute steps or subroutines of the methods described herein such that the multiple processors cooperate to achieve the functionality described herein. Furthermore, when device 1140 is implemented in a cloud computing system, various hardware components may belong to separate physical systems. For example, the processor may include a first processor in a first server and a second processor in a second server.
[0173] The present invention includes the following further embodiments.
[0174] Embodiment 1. An encoding system for encoding data in a data structure, the encoding system comprising: an input interface for receiving data; Processor system a processor system comprising: Obtaining the data as a plurality of data blocks; calculating a plurality of hash values for the plurality of data blocks by applying a hash function to the plurality of data blocks; Computing a first hash tree for a plurality of hash values, where the plurality of hash values are assigned to leaves of the first hash tree to generate one or more higher levels of the first hash tree; Including a plurality of data blocks and all or a portion of a first hash tree in a data structure, the portion of the first hash tree not including at least a portion of the leaves of the first hash tree; A coding system that performs the above.
[0175] Embodiment 2. The encoding system of embodiment 1, wherein the processor system is configured to compute a second hash tree, wherein a root of the first hash tree is assigned to a leaf of the second hash tree, and a number of further hash values are assigned to a number of further leaves of the second hash tree, and the further hash values are generated as roots of the further hash tree and / or by hashing on further data blocks, and include at least the root of the second hash tree in the data structure.
[0176] Embodiment 3. An encoding system as described in embodiment 1 or 2, wherein the data structure includes data blocks organized in a hierarchy of containers, and the processor system computes a hash tree for each container in the hierarchy, and each leaf of the hash tree is either a hash value of a data block in the container or the root of a hash tree corresponding to a lower container.
[0177] Embodiment 4. An encoding system as described in any one of embodiments 1 to 3, wherein the processor system calculates a digital signature over the root of the hash tree, in particular the root of the hash tree having the hash tree root among the leaves of the hash tree, and includes the digital signature in the data structure.
[0178] Embodiment 5. An encoding system as described in any one of embodiments 1 to 4, wherein the processor system stores a data structure and / or streams the data structure or a portion of the data structure, the portion of the data structure including at least a portion of a data block and at least a portion of a hash tree corresponding to the data block.
[0179] Embodiment 6. An encoding system as described in any one of embodiments 1 to 5, wherein a subset of the plurality of data blocks and / or data containers is marked as having integrity protection, and the remainder of the plurality of data blocks and / or data containers is marked as not having integrity protection, and only the portion marked as having integrity protection is included in the hash tree.
[0180] Embodiment 7. An encoding system described in any one of embodiments 1 to 6, wherein an input interface receives one or more tree parameters, and a processor system includes in a data structure a set of nodes selected from a hash tree or a hierarchy of hash trees in accordance with the one or more tree parameters.
[0181] Embodiment 8. An encoding system as described in any one of embodiments 1 to 7, wherein an input interface receives modifications to the data, the modifications including one or more of additions, deletions, and / or modifications, and a processor system applies the modifications and selectively recalculates and updates portions of the hash tree that correspond to the modified portions of the data.
[0182] Embodiment 9. The leaves of the hash tree further include hashes of the data blocks in uncompressed form and hashes of the data blocks in compressed form, and the data blocks are included in the data structure in compressed form, or An encoding system as described in any one of embodiments 1 to 8, wherein leaves of the hash tree further include hashes of data blocks in uncompressed and unencrypted form, hashes of data blocks in compressed and unencrypted form, and hashes of data blocks in compressed and encrypted form, and wherein data blocks are included in the data structure in compressed and encrypted form.
[0183] Embodiment 10. A validation system for validating selected data in a data structure, the validation system comprising: an input interface for receiving at least a portion of a data structure, the data structure including a plurality of data blocks and a portion of a hash tree, the input interface excluding at least a portion of a leaf of the first hash tree; Processor system a processor system comprising: calculating a plurality of hash values for selected data blocks by applying a hash function to the data blocks in the data structure selected for data integrity verification; determining a path starting from the data block selected for validation to the root of a corresponding hash tree; Retrieving or computing a hash value for the hash tree along the path, if available in the data structure; verifying the root of the hash tree from at least the calculated hash values; A verification system that performs the above.
[0184] Embodiment 11. A verification system as described in embodiment 10, wherein the data structure includes a hierarchy of hash trees and the processor system identifies a path starting from a leaf to the entire root of the hierarchy of the hash trees.
[0185] Embodiment 12. A coding and / or verification system according to any one of embodiments 1 to 11, wherein the coding and / or verification system is a device.
[0186] Embodiment 13. An encoding and / or verification system according to any one of embodiments 1 to 12, wherein the plurality of data blocks comprises genomic data.
[0187] Embodiment 14. An encoding method for encoding data in a data structure, the encoding method comprising: obtaining data as a plurality of data blocks; calculating a plurality of hash values for the plurality of data blocks by applying a hash function to the plurality of data blocks; Computing a first hash tree for a plurality of hash values, where the plurality of hash values are assigned to leaves of the first hash tree and one or more higher levels of the first hash tree are generated; including a plurality of data blocks and a portion of a first hash tree in a data structure, the portion of the first hash tree not including at least a portion of the leaves of the first hash tree; The encoding method according to claim 1,
[0188] Embodiment 15. A method for validating selected data in a data structure, the method comprising: receiving at least a portion of a data structure, the data structure including a plurality of data blocks and a portion of a hash tree, the data structure not including at least a portion of a leaf of the first hash tree; calculating a plurality of hash values for selected data blocks by applying a hash function to the data blocks in the data structure selected for data integrity verification; determining a path starting from the data block selected for verification to the root of the corresponding hash tree; retrieving a hash value for the hash tree along the path, if available in the data structure, or computing such a hash value; verifying the root of the hash tree from at least the calculated hash values; A verification method comprising:
[0189] Embodiment 16. An encoding system for encoding genomic data in a digital data structure, the encoding system comprising: an input interface for receiving genomic data; Processor system a processor system comprising: acquiring genomic data as a plurality of data blocks; calculating a plurality of hash values for the plurality of genomic blocks by applying a hash function to at least the genomic data in the plurality of genomic data blocks; Computing a first hash tree for a plurality of hash values, the plurality of hash values being assigned to leaves of the first hash tree, the first hash tree having at least three levels; including a plurality of genome blocks and a portion of a first hash tree in a data structure, the portion of the first hash tree including the first two highest levels of the first hash tree but not including one or more lower levels of the first hash tree; A coding system that performs the above.
[0190] Embodiment 17. An encoding system as described in any one of embodiments 1 to 16, wherein the processor system computes a second hash tree, wherein a root of the first hash tree is assigned to a leaf of the second hash tree and a plurality of further data items are assigned to a plurality of further leaves of the second hash tree, and includes at least the root of the second hash tree in the data structure.
[0191] Embodiment 18. An encoding system as described in any one of embodiments 1 to 17, wherein the data structure is a hierarchical data structure having multiple levels, and the processor system calculates a hash tree for each level of the hierarchical data structure and includes the root of the hash tree calculated for a lower level of the hierarchical data structure in the hash tree for the level above the lowest level.
[0192] Embodiment 19. An encoding system according to any one of embodiments 1 to 18, wherein an input interface receives a tree parameter and the processor system includes a larger or smaller portion of the first hash tree in the data structure depending on the tree parameter.
[0193] It should be noted that the above-described embodiments are illustrative rather than limiting of the disclosed subject matter, and that those skilled in the art will be able to design many alternative embodiments.
[0194] In the claims, reference signs placed in parentheses shall not be construed as limiting the claims. Use of the verbs "comprise" and "have" and their conjugations does not exclude the presence of elements or steps other than those stated in the claims. The use of a singular element does not exclude the presence of a plurality of such elements. Expressions such as "at least one of", when followed by a list of elements, denote the selection of all elements or any subset of elements from the list. For example, the expression "at least one of A, B, and C" is to be understood as including A only, B only, C only, both A and B, both A and C, both B and C, or all of A, B, and C. The disclosed subject matter can be implemented by means of hardware comprising several distinct elements and by a suitably programmed computer. In device claims enumerating several parts, several of these parts may be implemented by one and the same hardware. The mere fact that several measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0195] In the claims, reference signs in parentheses refer to the reference signs in the drawings of exemplary embodiments or to formulas of the embodiments, which facilitates the understanding of the claims, and shall not be construed as limiting the claims.
Claims
1. An encoding system for encoding data in a data structure, the encoding system comprising: an input interface for receiving the data; and a processor system, wherein the processor system obtains the data as a plurality of data blocks; calculates a plurality of hash values for the plurality of data blocks by applying a hash function to the plurality of data blocks; calculates a first hash tree for the plurality of hash values, wherein the plurality of hash values are assigned to leaves of the first hash tree and one or more higher levels of the first hash tree are generated; includes the plurality of data blocks and a portion of the first hash tree in the data structure, wherein the portion of the first hash tree includes the first two highest levels of the first hash tree but does not include one or more lower levels of the first hash tree. An encoding system that performs the above.
2. The processor system calculates a second hash tree, wherein the root of the first hash tree is assigned to a leaf of the second hash tree, a plurality of additional hash values are assigned to a plurality of additional leaves of the second hash tree, and an additional hash value is generated as the root of an additional hash tree and / or generated by hashing on an additional data block. The encoding system according to claim 1, wherein the processor system includes at least the root of the second hash tree in the data structure.
3. The data structure includes data blocks organized in a hierarchy of containers, and the processor system calculates a hash tree for each container in the hierarchy, and each leaf of the hash tree is either the hash value of the data block within the container or the root of the hash tree corresponding to a lower-level container. The encoding system according to claim 1 or 2.
4. The processor system calculates a digital signature over the root of a hash tree, in particular over the root of a hash tree having a hash tree root between the leaves of the hash tree, and includes the digital signature in the data structure. The encoding system according to any one of claims 1 to 3.
5. The processor system stores the data structure and / or streams the data structure or a part of the data structure, and the part of the data structure includes at least a part of the data block and at least a part of the hash tree corresponding to the data block. The encoding system according to any one of claims 1 to 4.
6. A subset of the plurality of data blocks and / or data containers is marked as having integrity protection, and the rest of the plurality of data blocks and / or data containers is marked as having no integrity protection, and only the part marked as having integrity protection is included in the hash tree. The encoding system according to any one of claims 1 to 5.
7. The input interface receives one or more tree parameters, and the processor system includes in the data structure a set of nodes selected from one hash tree or a hierarchy of hash trees according to the one or more tree parameters. The encoding system according to any one of claims 1 to 6.
8. The input interface receives a modification to the data, the modification includes one or more of addition, deletion, and / or change, and the processor system applies the modification and selectively recomputes and updates a part of the hash tree corresponding to the modified part of the data. The encoding system according to any one of claims 1 to 7.
9. The leaf of the hash tree further includes a hash of the data block in uncompressed form and a hash of the data block in compressed form, and the data block is included in the data structure in compressed form, or The leaf of the hash tree further includes the hash of the data block in uncompressed and unencrypted form, the hash of the data block in compressed and unencrypted form, and the hash of the data block in compressed and encrypted form, and the data block is included in the data structure in compressed and encrypted form. The encoding system according to any one of claims 1 to 8.
10. A verification system for verifying selected data in a data structure, the verification system comprising: An input interface for receiving at least a part of the data structure, the data structure including a plurality of data blocks and a part of a hash tree, the part including the first two highest levels of the hash tree, but not including one or more lower levels of the hash tree; Input interface, A processor system, wherein the processor system Calculating a plurality of hash values for the selected data block by applying a hash function to the data block in the data structure selected for data integrity verification; Identifying a path from the selected data block selected for verification to the root of the corresponding hash tree; Fetching the hash value for the hash tree along the path, or calculating such a hash value, if available in the data structure; Verifying the root of the hash tree from at least the calculated plurality of hash values A verification system that performs the above.
11. The data structure includes a hierarchy of hash trees, and the processor system The verification system according to claim 10, which identifies a path from a leaf to the entire root of the hierarchy of the hash tree.
12. The encoding and / or verification system according to any one of claims 1 to 11, wherein the encoding and / or verification system is a device.
13. The encoding and / or verification system according to any one of claims 1 to 12, wherein the plurality of data blocks includes genomic data.
14. An encoding method for encoding data in a data structure, the encoding method comprising: Obtaining the data as a plurality of data blocks; Calculating a plurality of hash values for the plurality of data blocks by applying a hash function to the plurality of data blocks; Calculating a first hash tree for the plurality of hash values, wherein the plurality of hash values are assigned to leaves of the first hash tree and one or more higher levels of the first hash tree are generated; Including the plurality of data blocks and a part of the first hash tree in the data structure, wherein the part includes the first two highest levels of the first hash tree but does not include one or more lower levels of the first hash tree; An encoding method having the above steps.
15. A verification method for verifying selected data in a data structure, the verification method comprising: Receiving at least a part of the data structure, wherein the data structure includes a plurality of data blocks and a part of a hash tree, and the part includes the first two highest levels of a first hash tree but does not include one or more lower levels of the first hash tree; Calculating a plurality of hash values for the selected data blocks by applying a hash function to the data blocks in the data structure selected for data integrity verification; Identifying a path from the selected data blocks for verification to the root of the corresponding hash tree; When available in the data structure, retrieving the hash values for the hash tree along the path or calculating such hash values; Verifying the root of the hash tree from at least the calculated plurality of hash values; A verification method having the above steps.
16. The verification method according to claim 15, wherein the data structure includes a hierarchy of hash trees, and the verification method includes identifying a path from a leaf to the entire root of the hierarchy of the hash tree.
17. A non-transitory or transitory computer-readable medium including data representing instructions that, when executed by a processor system, cause the processor system to execute the verification method according to claim 14, 15, and / or 16.