Data management method and device

The use of a distributed ledger to manage data integrity by hashing and verifying data items and programs addresses inconsistencies and security issues in existing systems, ensuring reliable and secure data management.

WO2025141161A1PCT designated stage expired Publication Date: 2025-07-03DIGICORP LABS HOLDING BV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/088568
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing data management systems face challenges with inconsistent data across multiple sources, leading to redundancy and security vulnerabilities, necessitating a more fundamental solution for uniform data management and integrity verification.

Method used

A method using a distributed ledger to store hashes of data items and programs, retrieving and verifying these hashes before executing programs to ensure data and program integrity, with the new hashes stored on the ledger.

Benefits of technology

This approach ensures data consistency and security by providing a tamper-proof record of data integrity and authenticity, reducing the need for costly infrastructure and enhancing trust in data management systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024088568_03072025_PF_FP_ABST
    Figure EP2024088568_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Some embodiments are directed to a method for data management using a distributed ledger. The method may include retrieving a first hash (241) for a data item (221) and a second hash (242) for a program (222) from the distributed ledger, and verifying the first and second hashes against a retrieved data item and a retrieved program. A new hash (243) derived from a new data item, obtained from executing the program using at least the data item as input, may be stored in the distributed ledger.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DATA MANAGEMENT METHOD AND DEVICE

[0002] TECHNICAL FIELD

[0003] The presently disclosed subject matter relates to a data management method, a data management device, and a computer readable medium.

[0004] BACKGROUND

[0005] In the existing technological landscape, the same Master Data is often stored at multiple data sources. In practice this often leads to inconsistent sets of Master Data. Attempts to address this issue include the deployment of costly infrastructure solutions specifically designed for data integration, such as the Asset Information Factory (AIF) and the Enterprise Data Warehouse (EDW). These solutions manage cross-references among various source systems. Compounding the problem is the isolated development of applications, resulting in an array of redundant tools that aim to address similar data integration challenges for Assets.

[0006] The current approach to these challenges is suboptimal, necessitating a more fundamental solution. Uniform management of data at the source level would alleviate various practical problems, as further elaborated herein.

[0007] In addition to the issue of inconsistent data across different sources, existing solutions also face significant security challenges. For example, the current infrastructure is susceptible to unauthorized alterations of both data and programs. This vulnerability not only compromises the integrity of the data but also poses risks to the overall reliability and trustworthiness of the system. Therefore, any comprehensive solution must also address these security gaps to ensure both data consistency and system integrity.

[0008] SUMMARY

[0009] There is a desire for an improved method for data management. A method for data management, and a data management device are defined in the claims.

[0010] In an embodiment, a method for data management uses a distributed ledger to store hashes of both data items and of one or more programs. The data items and programs are typically not stored on the distributed ledger. For example, before committing to the result of a program computation, the method may retrieve hashes of one or more data items and of the program from the distributed ledger, and the hashes to verify the integrity of the data item(s) and the program. A hash of a new data item that is obtained by executing the program may itself be stored on the distributed ledger.

[0011] The data management method described herein may be applied in a wide range of practical applications, e.g., in anti-counterfeiting applications, product tracking, and the like. Various practical applications are discussed herein.

[0012] An embodiment of the method may be implemented on a computer as a computer implemented method, or in dedicated hardware, or in a combination of both. Executable code for an embodiment of the method may be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product comprises non-transitory program code stored on a computer readable medium for performing an embodiment of the method when said program product is executed on a computer.

[0013] In an embodiment, the computer program comprises computer program code adapted to perform all or part of the steps of an embodiment of the method when the computer program is run on a computer. Preferably, the computer program is embodied on a computer readable medium.

[0014] Another aspect of the presently disclosed subject matter is a method of making the computer program available for downloading. This aspect is used when the computer program is uploaded into a server, and when the computer program is available for downloading from such a server.

[0015] BRIEF DESCRIPTION OF DRAWINGS

[0016] Further details, aspects, and embodiments will be described, by way of example only, with reference to the drawings. Elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. In the figures, elements which correspond to elements already described may have the same reference numerals. In the drawings,

[0017] Figure la schematically shows an example of an embodiment of a data management system,

[0018] Figure lb schematically shows an example of an embodiment of a data management system,

[0019] Figure 2a schematically shows an example of an embodiment of a data management system, Figure 2b.1 schematically shows an example of an embodiment of a data management device,

[0020] Figure 2b.2 schematically shows an example of an embodiment of a data management device,

[0021] Figure 2c schematically shows an example of part of an embodiment of a data management system,

[0022] Figure 2d schematically shows an example of an embodiment of a data management device,

[0023] Figure 3 schematically shows an example of an embodiment of a distributed ledger,

[0024] Figure 4 schematically shows an example of an embodiment of a data management system,

[0025] Figure 5 schematically shows an example of an embodiment of a data management system,

[0026] Figure 6 schematically shows an example of an embodiment of a data management system,

[0027] Figure 7a schematically shows an example of an embodiment of a data logging system,

[0028] Figure 7b schematically shows an example of an embodiment of a data logging system,

[0029] Figure 8 schematically shows an example of an embodiment of a data management system,

[0030] Figure 9a schematically shows an example of an embodiment of a data anonymizing system,

[0031] Figure 9b schematically shows an example of an embodiment of a data anonymizing system,

[0032] Figure 10 schematically shows an example of an embodiment of a method for data management using a distributed ledger,

[0033] Figure I la schematically shows a computer readable medium having a writable part comprising a computer program according to an embodiment,

[0034] Figure 11b schematically shows a representation of a processor system according to an embodiment.

[0035] Reference signs list The following list of references and abbreviations corresponds to figures la- 5, and 1 la-1 lb, and is provided for facilitating the interpretation of the drawings and shall not be construed as limiting the claims.

[0036] 100, 102 a data management system

[0037] 110 a data source

[0038] 110.1, 110.2 a data source

[0039] 120 a data management device

[0040] 130 a distributed ledger interface

[0041] 130.1, 130.2 a distributed ledger interface

[0042] 111 a processor system

[0043] 112 storage

[0044] 113 a communication interface

[0045] 121 a processor system

[0046] 122 a storage

[0047] 123 a communication interface

[0048] 131 a processor system

[0049] 132 a storage

[0050] 133 a communication interface

[0051] 172 a computer network

[0052] 200 a data management system

[0053] 210 a data management device

[0054] 211, 212 a data source

[0055] 211’, 212’ a data source

[0056] 213 a data source

[0057] 221 a data item

[0058] 222 a program

[0059] 223 a new data item

[0060] 224 a first transaction identifier

[0061] 225 a second transaction identifier

[0062] 240 a distributed ledger interface

[0063] 241 a first hash

[0064] 242 a second hash

[0065] 243 a new hash 244 a transaction identifier 251, 252 a private key 300 a distributed ledger 310, 320 distributed ledger block

[0066] 3 H, 321 a transaction identifier 312, 322 a hash 410 a data management system 411 a blockchain synchronization unit 421 a distributed ledger node 412 a core 422 external storage 413 a database 414 a web interface 510 Setup, e.g., Private key generation and configuration 520 MDM application 521 Bootstrap 522 Generate fund addresses 523 Wait for funds 524 Generate contract address 525 Root owner 526 Issuance Transaction for owner address 601 information stored on chain 602 function for providing an external URL

[0067] 800, 801 a logging system 811 a first application 812 a second application 813 a logging interface 810 a data management device 821, 822 a logging item

[0068] 841 a first hash 842 a second hash 843 a new hash 823 updated logging collection

[0069] 830 a third application

[0070] 844 a transaction identifier

[0071] 802 a data management system

[0072] 851 a first storage

[0073] 852 a second storage 861 a data collection

[0074] 853 a data provenance interface

[0075] 850 a data management device

[0076] 881 a first hash 882 a second hash

[0077] 883 a new hash

[0078] 863 a new data collection(optional)

[0079] 870 an application

[0080] 900, 901 a data anonymizing system

[0081] 910 a data source

[0082] 911 a data collection 920 a data anonymizer

[0083] 921 an anonymized data collection

[0084] 922 a key file

[0085] 930 a cryptographic proof system

[0086] 931 a cryptographic proof

[0087] 1000, 1001 a computer readable medium

[0088] 1010 a writable part

[0089] 1020 a computer program 1110 integrated circuit(s) 1120 a processing unit

[0090] 1122 a memory

[0091] 1124 a dedicated integrated circuit

[0092] 1126 a communication element

[0093] 1130 an interconnect

[0094] 1140 a processor system

[0095] DESCRIPTION OF EMBODIMENTS

[0096] While the presently disclosed subject matter is susceptible of embodiment in many different forms, there are shown in the drawings and will herein be described in detail one or more specific embodiments, with the understanding that the present disclosure is to be considered as exemplary of the principles of the presently disclosed subject matter and not intended to limit it to the specific embodiments shown and described.

[0097] In the following, for the sake of understanding, elements of embodiments are described in operation. However, it will be apparent that the respective elements are arranged to perform the functions being described as performed by them.

[0098] Further, the subject matter that is presently disclosed is not limited to the embodiments only, but also includes every other combination of features described herein or recited in mutually different dependent claims.

[0099] Figure la schematically shows an example of an embodiment of a data management system 100.

[0100] Shown is a data source 110, data management device 120, and distributed ledger interface 130, which may be part of a system 100.

[0101] Data source 110 may be configured for one or more of: storing and providing a data item, storing, and providing a program. Data management device 120 is configured to verify the data item and / or the program against hashes stored in a distributed ledger. Data management device 120 may further be configured to execute the program. Distributed ledger interface 130 is configured to provide read and / or write access to a distributed ledger. In particular, distributed ledger interface 130 may be configured to retrieve a hash stored in the distributed ledger, and / or to store a new hash in the distributed ledger. For example, system 100 may be used to provide a high assurance of the integrity of the data and / or program, yet at the same time require little storage on a distributed ledger. Distributed ledger interface 130 may be a distributed ledger interface device.

[0102] Data source 110 may comprise a processor system 111, a storage 112, and a communication interface 113. Data management device 120 may comprise a processor system 121, a storage 122, and a communication interface 123. Distributed ledger interface 130 may comprise a processor system 131, a storage 132, and a communication interface 133.

[0103] A data source may, for example, comprise one or more of: a local storage, e.g., local to data management device 120, e.g., a hard drive; a cloud storage; and a distributed storage network, e.g., IPFS.

[0104] In the various embodiments of communication interfaces 113, 123 and / or 133, the communication interfaces may be selected from various alternatives. For example, the interface may be a network interface to a local or wide area network, e.g., the Internet, a storage interface to an internal or external data storage, an application interface (API), etc.

[0105] Storage 112, 122 and 132 may be, e.g., electronic storage, magnetic storage, etc. The storage may comprise local storage, e.g., a local hard drive or electronic memory. Storage 112, 122 and 132 may comprise non-local storage, e.g., cloud storage. In the latter case, storage 112, 122 and 132 may comprise a storage interface to the non-local storage. Storage may comprise multiple discrete sub-storages together making up storage 112, 122, 132.

[0106] Storage 112, 122 and 132 may be non-transitory storage. For example, storage 112, 122 and 132 may store data in the presence of power such as a volatile memory device, e.g., a Random Access Memory (RAM). For example, storage 112, 122 and 132 may store data in the presence of power as well as outside the presence of power such as a non-volatile memory device, e.g., Flash memory. Storage may comprise a volatile writable part, say a RAM, a non-volatile writable part, e.g., Flash. Storage may comprise a non-volatile non-writable part, e.g., ROM.

[0107] The devices 110, 120 and 130 may communicate internally, with each other, with other devices, external storage, input devices, output devices, and / or one or more sensors over a computer network. The computer network may be an internet, an intranet, a LAN, a WLAN, a WAN, etc. The computer network may be the Internet. The devices 110, 120 and 130 may comprise a connection interface which is arranged to communicate within system 100 or outside of system 100 as needed. For example, the connection interface may comprise a connector, e.g., a wired connector, e.g., an Ethernet connector, an optical connector, etc., or a wireless connector, e.g., an antenna, e.g., a Wi-Fi, 4G or 5G antenna.

[0108] The communication interface 113 may be used to send or receive digital data, e.g., a data item, transaction identifier, program. The communication interface 123 may be used to send or receive digital data, e.g., a data item, program, hashes, new data, or hashes, and so on. The communication interface 133 may be used to send or receive digital data, e.g., data retrieved from a distributed ledger, and / or data for storing on the distributed ledger.

[0109] The execution of devices 110, 120 and 130 may be implemented in a processor system. The devices 110, 120 and 130 may comprise functional units to implement aspects of embodiments. The functional units may be part of the processor system. For example, functional units shown herein may be wholly or partially implemented in computer instructions that are stored in a storage of the device and executable by the processor system.

[0110] The processor system may comprise one or more processor circuits, e.g., microprocessors, CPUs, GPUs, etc. Devices 110, 120 and 130 may comprise multiple processors. A processor circuit may be implemented in a distributed fashion, e.g., as multiple sub-processor circuits. For example, devices 110, 120 and 130 may use cloud computing.

[0111] Typically, the data source 110, data management device 120, and distributed ledger interface 130 each comprise a microprocessor which executes appropriate software stored at the device; for example, that software may have been downloaded and / or stored in a corresponding memory, e.g., a volatile memory such as RAM or a non-volatile memory such as Flash.

[0112] Instead of using software to implement a function, the devices 110, 120 and / or 130 may, in whole or in part, be implemented in programmable logic, e.g., as field- programmable gate array (FPGA). The devices may be implemented, in whole or in part, as a so-called application-specific integrated circuit (ASIC), e.g., an integrated circuit (IC) customized for their particular use. For example, the circuits may be implemented in CMOS, e.g., using a hardware description language such as Verilog, VHDL, etc. In particular, data source 110, data management device 120 and distributed ledger interface 130 may comprise circuits, e.g., for cryptographic processing, and / or arithmetic processing. In hybrid embodiments, functional units are implemented partially in hardware, e.g., as coprocessors, e.g., cryptographic coprocessors, and partially in software stored and executed on the device.

[0113] Figure lb schematically shows an example of an embodiment of a data management system 102. System 102 may comprise multiple data sources; shown are data sources 110.1 and 110.2. For example, data sources 110.1 and 110.2 may be used to store different types of data, e.g., one or more of data items, programs, first and second hashes. For example, data sources 110.1 and 110.2 may store, at least in part, the same data items, and may be synchronized by querying a distributed ledger for hashes of data items they store, and verifying that the store the correct version, e.g., the latest version. For example, the last hash stored on the distributed ledger may correspond to the most recent update of a data item. The same holds for programs.

[0114] System 100 may comprise multiple distributed ledger interfaces; shown are distributed ledger interfaces 130.1 and 130.2. The devices are connected through a computer network 172, e.g., the Internet. The data source 110 and data management device 120 may be according to an embodiment.

[0115] Figure 2a schematically shows an example of an embodiment of a data management system 200. System 200 comprises a data management device 210. Data management device 210 is configured to communicate with one or more data sources. In this example, three data sources are shown: data sources 211, 212, and 213. There may be more or fewer data sources.

[0116] For example some of the shown data sources may be the same. For example, data source 211 and data source 212 may be the same data source. For example, data source 211 and data source 213 may be the same data source. For example, data source 211, data source 212, and data source 213 may be the same data source.

[0117] On the other hand, for example, data source 211 and data source 212 may be different. For example, data source 211, data source 212, and data source 213 may all be different.

[0118] There are many choices possible for these data sources. For example, they may be selected from a local storage, e.g., local to data management device 210, e.g., a hard drive; a cloud storage; and a distributed storage network, e.g., IPFS.

[0119] Data management device 210 is configured to work with computer programs and data, and to use a distributed ledger to increase the reliability of said computer programs and data. On the other hand, data management device 210 avoids storing large amounts of data in the distributed ledger as this is resource intensive. Storing data on a distributed ledger is typically costly and may be slow, typically much more expensive, and slower than other types of storage. Furthermore, distributed ledgers may not even support storing of data over a defined file size. For example, data management device 210 may be configured to work with a distributed ledger 300, further described below.

[0120] In order to, e.g., perform its data processing tasks, data management device 210 is configured to retrieve a data item 221 and a computer program 222. Program 222 may be a smart contract. For example, data management device 210 may be configured for retrieving a data item 221 from a data source 211 and for retrieving a program 222 from a data source 212. Program 222 is configured to perform a computation on data item 221, and possibly additional inputs which are not shown separately in figure 2a, and compute a new data item 223.

[0121] In an embodiment, data source 212 may be data management device 210 itself. For example, the program may be stored on local storage of data management device 210, e.g., an SSD. For example, data management device 210 may be configured for retrieving program 222 from local storage. Alternatively, the program is stored elsewhere. A hash of the program may be stored on the distributed ledger.

[0122] Note that typically, program 222 is not executed on data item 221 until both have been verified as explained below.

[0123] A particularly advantageous application is one in which computer program 222 updates the data item 221, e.g., computing an updated data item 223. For example, in an embodiment, data item 221 comprises a list of attributes. Program 222 may be configured to change at least one of said attributes. In particular, data item 221 may be associated with a physical object so that an embodiment may be used to securely update or inspect the attributes.

[0124] In various applications, the secure and authentic updating of attributes tied to physical objects enhances the reliability and integrity of the product in question.

[0125] For example, an embodiment may be used in anti -counterfeiting, product tracking, and maintenance control. An embodiment could help ensure the authenticity of products, e.g., from the manufacturer to the end consumer. For example, in the pharmaceutical industry, the product may comprise drugs, vaccines, and / or other medical products. Attributes such as batch number, expiration date, and chemical composition are linked to a physical object like a medicine bottle. As the object moves through the distribution network, these attributes are updated using an embodiment, thus improving data integrity, and making it difficult to tamper with the information. Any attempt to alter the attributes can be detected.

[0126] An embodiment may help to ensure that components are genuine and / or meet safety standards. For example, in the automotive sector, the components may be vehicle components, such as airbags or brake systems. Each component may be associated with attributes, such as a serial number, manufacturing date, or test results. These attributes may be updated when the component is installed, inspected, or repaired. Discrepancy in attributes could signal that the component is counterfeit, thus allowing for corrective action. Similarly, in the context of industrial equipment maintenance, an embodiment can maintain the quality and safety of machinery. Each piece of equipment is given attributes such as last maintenance date, parts replaced, or test results. These may be updated when the equipment undergoes maintenance or inspection.

[0127] For example, one of the attributes in the list may indicate an owner of the physical object. Updating this attribute may thus indicate a new owner of the physical object associated with data item 221.

[0128] In an embodiment, program 222 may be configured to perform various data operations on data item 221, e.g., delete in part, insert data, replace data, reorder data items, update, attach data, e.g., concatenate. The additional input may comprise which operation is to be performed. The additional input may comprise which data is to be added, altered, or the like. Additional input is not needed. For example a useful application is a counter, in which data item 221 comprises a counter, which is incremented by 1 by program 222. Such a counter could be used as a nonce.

[0129] A particularly advantageous application is one in which computer program 222 updates the data item 221, e.g., computing an updated data item 223.

[0130] In order to verify that data item 221 and program 222 are integer, e.g., has not been changed, data management device 210 retrieves hashes for them. In particular, a data management device 210 retrieves a first hash 241 for data item 221 and a second hash 242 for computer program 222. The hashes have previously been computed from data item 221 and program 222 respectively; that the hashes did not change since then is proof that the data item 221 and program 222 did not change either. For example, the hashes may have been computed with a hash function. Examples of hash functions include: SHA-1, SHA-256, SHA-512, Whirlpool, SHA-3.

[0131] The hashes are retrieved from a distributed ledger. Storing information on a distributed ledger relies on a consensus algorithm that verifies the integrity of each new entry across multiple nodes. Once data is written and confirmed, altering it would require changing the record on a majority of nodes, making unauthorized modifications to the distributed ledger virtually impossible.

[0132] For example, for a Bitcoin Blockchain the distributed ledger interface may comprise a full node server, which may be contacted to retrieve data, e.g., based on its Transaction ID. Alternatively, the data management device 210 may run a Bitcoin full node locally, in which case the blockchain may be queried locally.

[0133] For example, for an Ethereum Blockchain, the distributed ledger interface may include an Ethereum gateway server, which can be contacted to retrieve data, e.g., using a Transaction Hash or Contract Address along with a Function Signature. Alternatively, the data management device may run an Ethereum full node locally, allowing for direct JSON-RPC calls to query the blockchain.

[0134] For example, for a Hyperledger Fabric network, the distributed ledger interface may comprise a Peer node server. This server can be queried through a chain code or smart contract to retrieve data based on a Key or Composite Key. In cases where a data management device is part of the Hyperledger Fabric network, the blockchain can be queried directly, e.g., through an SDK or REST API provided by the specific implementation of Hyperledger Fabric in use.

[0135] Once data management device 210 has both data item 221 and first hash 241, data management device 210 verifies them against each other. Once data management device 210 has both program 222 and second hash 242, data management device 210 verifies them against each other. To verify a hash against data, data management device 210 may compute the hash itself over the data using the same hash function as was used to compute the hash, and compare the computed hash against the retrieved hash.

[0136] If the first and / or second hash does not verify correctly, this is evidence that the corresponding data item and / or program was changed unexpectedly, which may be evidence of tampering, or of a synchronization problem. Data management device 210 is configured to perform a contingency algorithm in that case. For example, the contingency algorithm may comprise raising an alarm, e.g., sending a warning message, e.g., to an operator of device 210. Typically, data management device 210 will refuse to perform the computation on data item 221 to produce new data item 223, in case of verification error. In an embodiment, however, the computation is done regardless of verification results, or at least started, but computation result 223 is only committed to once the verification is complete. This allows verification and computation to be done in parallel. In an embodiment, a stated computation is aborted once verification fails. In an embodiment, the hash function generates a string, e.g., a binary string which is compared against a retrieved hash. In an embodiment however, the hash is computed according to a Merkle tree. In a Merkle tree, each leaf node contains the hash of a data item, and each non-leaf node contains the hash of its child nodes. To compare a hash value against the data, you would traverse the tree from the leaf node containing the hash of the data item up to the root, recalculating and verifying hashes at each level. If the root hash matches the expected value, the data has not been altered.

[0137] In an embodiment, data management device 210 retrieves multiple data items from one or more data sources, all different from the distributed ledger. Furthermore, multiple corresponding first hashes are retrieved therefor and each data item is verified therewith. The program is executed on at least the multiple data items.

[0138] Interestingly, retrieving the multiple first hashes may be in parallel with executing the program. The first hashes may be part of a Merkle tree. The latter allows verifying the integrity of the data items, but also that the data items belong together, e.g., in a same set. Use of a Merkle tree may be convenient but is not necessary. For example, to manage multiple data items, it suffices to retrieve multiple corresponding first hashes. The use of a Merkle tree is optional.

[0139] In an embodiment, data management device 210 is configured to retrieve a first transaction identifier 224 identifying a first transaction on the distributed ledger, and a second transaction identifier 225 identifying a second transaction on the distributed ledger. The transaction identifiers are used to locate and retrieve the stored hashes from the distributed ledger. For example, the first hash is retrieved from the first transaction on the distributed ledger, and the second hash is retrieved from the second transaction. For example, the transaction identifier is communicated to a distributed ledger interface 240 which uses it to locate the hash stored thereon.

[0140] For example, for a bitcoin blockchain, the transaction identifier may comprise a Transaction ID (TXID) to locate and retrieve the stored hash. The TXID is a 64- character hexadecimal string that uniquely identifies a transaction. By searching this TXID in a Bitcoin block explorer, the specific transaction may be located that includes the data, e.g., in an OP RETURN field.

[0141] For example, for an Ethereum blockchain, the transaction identifier may comprise a Transaction Hash or a combination of Contract Address and Function Signature. For standard transactions, the Transaction Hash, a 66-character string starting with "Ox," can be used to find the hash. If the hash is stored within a smart contract, the Contract Address and potentially a specific function signature may be used to locate the hash.

[0142] For example, for a Hyperledger fabric, the transaction identifier may comprise a Key or Composite Key. Hyperledger Fabric typically stores key-value pairs for storing data. To retrieve the hash, the key, or a composite key, which may be obtained at the time of storing the data, may be used. For example, chain code, e.g., a smart contract, access may be used to query the data.

[0143] There are various ways to store and retrieve the transaction identifiers.

[0144] Figure 2b.l schematically shows an example of an embodiment of a data management device. In this example, which may be combined with figure 2a, data management device 210 may be configured to retrieve both data item 221 and first hash 224 from data source 211. Likewise, program 222 and second hash 225 may both be retrieved from data source 212. In other words, data and transaction identifier may be stored together, and may be transferred together.

[0145] Figure 2b.2 schematically shows an example of an embodiment of a data management device. In this example, which may be combined with figure 2a, data management device 210 may be configured to retrieve data item 221 from data source 211 and first hash 224 from different data source 211’. Likewise, program 222 may be retrieved from data source 212 and second hash 225 may be retrieved from different data source 212’. In other words, data and transaction identifier may be stored on different data sources, and may be transferred separately.

[0146] These embodiments may be combined. For example, hash and data item may come from the same data source, but program and hash from different data sources, or vice versa. If multiple data items are used, then for each a different choice may be made in this respect.

[0147] In an embodiment, the alternative data sources, e.g., data sources 211’ and 212’ may be integrated with data management device 210.

[0148] Returning to figure 2a.

[0149] Data management device 210 is configured to execute the program 222 using at least the data item 221 as input. In this way a new data item 223 is obtained, which results from said execution. The new data item 223 may be stored, e.g., in an external data source 213. Interestingly, in some embodiments, it may not actually be necessary to store data item 223 itself. For example, this may be done if data item 223 may be recomputed when necessary. In an embodiment, further input from the outside may be used as a further input for program. The further input may not be associated with a hash.

[0150] In an embodiment, the data management device 210 is configured for scheduled operation causing modification to the program’s output.

[0151] Data management device 210 is configured to derive a new hash 243 from at least the new data item 223. For example, data management device 210 may apply the hash function to at least data item 223.

[0152] The hash may be computed over additional data. For example, in an embodiment, program 222 may authenticate and / or authorize a user of program 222 before new data item 223 results from the execution. The new hash may be derived from the identity as well as from data item 223, e.g., by applying the hash function to both. In this case, the identity of the user may be stored to allow later verification of the hash; This is not necessary if the user is known, e.g., fixed, or is one of a limited number of users.

[0153] In an embodiment, program 222 may be configured to receive and verify one or more authorization codes before new data 223 results from the execution. New hash 243 is further derived from the one or more authorization codes. For example, an authorization code may comprise one of: OAuth Tokens, API keys, signature from a private key, etc.

[0154] New hash 243 is stored on the distributed ledger. For example, data management device 210 may contact distributed ledger interface 240 to store new hash 243.

[0155] In an embodiment, program 222 may be executed multiple times on the same input data. In that case, no modification to the program’s output may occur. Accordingly, no new hash needs to be written either. In an embodiment, data management device 210 is configured to compare data item 221 with new data item 223, e.g., by a direct comparison, or by comparing hash value for each, and to store new hash 243 only if data item 221 and new data item 223 differ.

[0156] For example, in the case of a Bitcoin Blockchain, data management device 210 may contact a Bitcoin full node to broadcast a new transaction containing the data. This transaction would then be added to the blockchain once it is verified and mined. For example, in the case of an Ethereum Blockchain, data management device 210 may contact an Ethereum gateway server or full node to send a new transaction or execute a smart contract function that stores the data.

[0157] For example, in the case of a Hyperledger Fabric network, data management device 210 may contact a Peer node server to invoke a chain code that stores the new data as a key -value pair in the ledger.

[0158] A new transaction identifier 244 may be obtained in conjunction with storing new hash 243. The new transaction identifier allows identifying a transaction on the distributed ledger that stores the new hash. The transaction identifier in a data source, which may be the same data source in which data item 223 is stored (if any), or the same data source from which a transaction identifier was obtained, e.g., data source 211’, or the new transaction identifier may be stored locally at data management device 210, etc. Figure 2c schematically shows an example of part of an embodiment of a data management system. Distributed ledger interface 240 receives new hash 243 for storing and returns new transaction identifier 244 to data management device 210 for identifying the transaction on the distributed ledger that stores new hash 243.

[0159] Returning to figure 2a.

[0160] In an embodiment, a database may be implemented. For example, the database may comprise multiple entries which each comprise entry data and an entry identifier. For each entry in the database, a first hash of entry data is stored in the distributed ledger. A transaction identifier suitable for locating the first hash in the ledger is stored in the entry identifier.

[0161] The program receives as input one or more entries for which the corresponding first hashes are retrieved and verified against the corresponding entry data. Additional data may comprise amendments to one or more of the entries. Accordingly, the program may store the new entry in the database, compute a first hash and store the hash on the distributed ledger. A new transaction identifier may be stored in the entry transaction. The storage on the ledger may comprise additional information, e.g., an identifier of the entry. The latter allows a verification of the database against the chain, e.g., to detect and / or prevent rollback attacks.

[0162] In an embodiment, the program may be a smart contract. For example, the execution of the program may be triggered, e.g., by a transaction that is sent to the smart contract's address. The trigger may come from a user, say, or from another smart contract sending the trigger. For example, data management device 210 and / or distributed ledger interface 240 may monitor the distributed ledger for a triggering transaction. A triggering transaction may be placed on the distributed ledger by a user of another smart contract. For example, a triggering transaction may be recognized from a destination address, e.g., a To field, transaction data, and other elements, e.g., even value transferred or transaction status. Which transactions trigger may be predetermined.

[0163] For example, in an embodiment, a full node may be run locally, e.g., as part of the data management device 210. The full node may be configured to look for triggering transactions.

[0164] Once a triggering transaction has been found the local program may be executed according to an embodiment, e.g., including verification of data items and / or of the program. The program may be retrieved and executed.

[0165] For example, the program may consume tokens, e.g., utility tokens, to regulate the use of limited resources. Although the program may be executed off chain, the program may for example, send transactions to the distributed ledger interface to cause a transfer to tokens.

[0166] For example, data items and / or programs may have an identifying name, preferably a unique name. The name may be a natural name or may be generated, e.g., randomly. A program may refer to a data item by its identifier to indicate to data management device 210 that the item needs to be retrieved and used. For example, this may comprise a so-called primary key, e.g., for a database, or a column and row number for an excel sheet, or the like. A program may refer to another program by its identifier to indicate to data management device 210 that the program needs to be retrieved and executed.

[0167] Figure 2d schematically shows an example of an embodiment of a data management device 210. Data management device 210 is configured to sign a transaction comprising new hash 243 with a private key. The private key that is used depends on the type of data item 223. For example, data management device 210 may be configured to determine a data type from multiple data types for the new data item. A private key is associated with each data type. Figure 2d shows two private keys: private keys 251 and 252 associated with different data types. Data management device 210 retrieves the private key associated with the determined data type, and uses it to sign the transaction with the retrieved private key. Different data types might have different handling requirements, protocols, or compliance standards. By using a unique private key for different data types, the data management device can compartmentalize the appropriate handling for each type of data together with the signing. A different data type may be handled by different software, or even a different computer. This improves both security, tracking and auditing of data handling. Furthermore, using different private keys adds additional security. If one private key is compromised, only the data associated with that specific key is at risk, rather than all data managed by the device.

[0168] Furthermore hashes of data items may use a different key than computer programs. This has the advantage that a device that is allowed to store a new hash on the ledger may not necessarily be allowed to store a hash for a computer program, and vice versa.

[0169] For example, in an embodiment, first hash 241 is retrieved from a first transaction on the distributed ledger. A corresponding signature of the first transaction is verified using a first public key. Second hash 242 is retrieved from a second transaction on the distributed ledger. A corresponding signature of the second transaction is verified using a different second public key.

[0170] Using multiple private keys, or using different private keys for different data types is optional. In an embodiment, the same private key is used for all data types.

[0171] Figure 3 schematically shows an example of an embodiment of a distributed ledger 300.

[0172] In the schematic illustration, distributed ledger blocks 310 and 320 represent individual blocks within a distributed ledger, e.g., of blockchain type. Each block contains a transaction identifier and stored data, labeled as 311 and 312 in block 310, and 321 and 322 in block 320. In an embodiment, the stored data comprises a hash, e.g., the first hash, second hash or new hash.

[0173] The transaction identifier serves as a unique label for locating the specific transaction within the larger blockchain network, and the stored data represents the actual information contained in that transaction. These blocks are connected, e.g., sequentially, forming a distributed ledger maintained across multiple nodes. Editing the blocks is hard because each block contains a hash of one or more previous blocks, creating an interdependent chain or graph. Any modification would invalidate this hash, requiring the re-computation of all subsequent blocks, which is computationally infeasible. In an embodiment, the distributed ledger is DigiByte blockchain, e.g., DigiByte Layer 1. Other possible choices include the Bitcoin blockchain, Ethereum Blockchain, Hyperledger Fabric, and so on.

[0174] Below several further optional refinements, details, and embodiments are illustrated.

[0175] In an embodiment, a system, referred to as a layer 2 system, is provided on top of a distributed ledger, e.g., a layer 1 system. For example, the layer 1 system may be the DigiByte Public Open-Source Layer 1 system. The layer 2 system enhances the functionality, scalability, and efficiency of the network. By leveraging the security of the underlying distributed ledger, e.g., the DigiByte UTXO blockchain network, the layer 2 system provides improved ways for users to transact and interact with blockchain-based applications, e.g., using a smart contract protocol. Other UTXO blockchains could also be used, instead of DigiByte.

[0176] The layer 2 system can be applied in one or more of: Master Data Management, Creation of audit trails, Verification and Validation of Authenticity, etc. In this example, the application of the layer 2 system in Master Data Management will be described. The embodiment may be modified to fit other applications as needed.

[0177] The decentralized nature of Blockchain technology, and the immutable storage of transactions improves Master Data Management. For example, it may be used to enforce a common, single source of truth data set. Blockchain helps with data reconciliation among multiple parties, and maintaining data integrity. Digital signatures may be used with transactions to improve authenticity. Cryptographic hash functions, and consensus algorithms make it hard to tamper with the data.

[0178] The transparency of data collection is an important requirement of any MDM solution. Silos and multiple movements of data from one system to another in enterprise environments make it hard to maintain an audit trail. The layer 2 system may be configured to provide audit trails for every MDM operation. This will improve the confidence of stakeholders in the system. Blockchain technology makes it easier to share master data among stakeholders. The efficiency gains of the decentralized network may reduce server management costs.

[0179] A problem when using blockchain technology is the fact that as the network grows so does the volume of data in every node, over time this results in blockchain bloat. In an embodiment of the system, e.g., the layer 2 system, the smart contract is not stored on chain. For example, a program, such as a smart contract, may be stored in a data source and not in the distributed ledger. This is different from blockchains such as Ethereum in which smart contracts are stored on chain. Accordingly, the layer 2 system has reduced bloat. In fact, in an embodiment, the only data stored on chain are hashes, e.g., the first hashes and second hashes. This makes the layer 2 system space efficient.

[0180] For example, in an embodiment of the system the distributed ledger nodes, e.g., full nodes, are included in the system. For example, a private blockchain may be used. The full nodes may be provided with access to the data sources storing data items and / or programs so that computation on the chain can be corrected. If data items and / or programs are not available to a full node, as, e.g., may be the case if a public layer 1 chain is used, the integrity of the chain may still be ensured due to normal chain verification mechanisms, for example, trust may be created because multiple nodes are running.

[0181] A system in which peer-to-peer transactions are desired, especially without middlemen, benefits from a decentralized blockchain. While one may assume that participants in a proposed business can be trusted, relying solely on trust is not advisable. Blockchain offers a transaction processing system that eliminates the need for explicit trust among participants by decentralizing transaction validation.

[0182] Access to the system, e.g., layer 2 system, may be regulated using utility tokens, which grant access to various platform features. For example, the utility tokens may be purchased or otherwise distributed among participants to regulate access to limited resources. For example, utility tokens may be locked in a smart contract when a new user is admitted into the system. For example, a new user may be provided with a license and one or more utility tokens.

[0183] For example, utility tokens may be required to perform transactions on the distributed ledger. For example, a program, e.g., a smart contract configured to change data, e.g., data attributes, may ensure that this can only be done after a certain condition is met, e.g., in the case of a particular update, a data value owner (DVO) may have to approve the change. Verification of utility tokens may be managed by a separate smart contract.

[0184] In an embodiment, a so-called digital twin may be created on chain that represents a physical object. This object may deliver a ground truth asset master data entity. This is sometimes called a golden table. The Digital Twin may be a Non-Fungible Data Entry (NFDe), which is like a NFT and carries the property of absolute uniqueness as an attribute as it exists on the Blockchain. For example, in an embodiment, a database may be represented by encoding each database entry or database transaction as an object on the chain, e.g., in an NFT like object. For example, the object may comprise one or more data tables that can be accessed by multiple people. In an embodiment, access rights may be added so that only devices with proper access rights can access the table(s). For example, a table may be encrypted and access to the decryption key may be regulated.

[0185] An attribute change may be requested for any object. For example, according to a set Smart Contract Governance Criteria, the change will be recorded, and an immutable audit trail will be created. For example, ‘fingerprints’ (hashes) may be recorded on the blockchain. The smart contract may include a validation process before recording the change.

[0186] Transfer of Ownership of the object, e.g., of the golden data table, is possible, where after a validation process, according to a set Smart Contract Governance Criteria, the change will be recorded and an immutable audit trail will be created, using hashes on the blockchain.

[0187] Provisioning of data to the smart contract, possibly temporary, is possible, where after a validation process, according to a set Smart Contract Governance Criteria, the data will be made available and an immutable audit trail will be created, using hashes on the blockchain. Note that data on the distributed blockchain is always permanent, but data in the layer 2 system may be temporary before being recorded. Changes may be recorded by hashing on the blockchain so that an audit trail is created. For example, data may be provided to the contract, which may decide to discard the data, in case some criterion is met, e.g., a failed authorization.

[0188] For example, in an embodiment, attribute changes are only possible with the use of a Smart Contract in which governance criteria are defined and to ensure that defined data value owners approve the change.

[0189] With this, data stewardship can ensure that data is accurate, secure, and available to those who need it, while also protecting the privacy and confidentiality of individuals. Validation of issuance, updates, transfer, and provisioning by authorized individuals defined in smart contract, can preferably only be done after authorization. A particularly effective authorization system is a passwordless QR-based authentication. For example, an authentication system is described in Dutch patent application N2035471, included herein by reference, which may be used to authorize data changes; for example, any of the claims as filed therewith may be used for passwordless authentication. In an embodiment, the program retrieved from the data source is configured to authenticate and / or authorize a user of the program, e.g., using a smart card, password, or other authentication mechanism. This could be done before, after, or during the executing of the program; however, if a new data item results from the execution, then the new hash if further derived from the identity. This has the advantage that the identity of the user is linked to the new hash. For example, this process allows the program to claim ownership of a physical or digital asset by validating the personal or corporate identity

[0190] Preferably, the program also stores an identity of the user, e.g., to ease future precomputation of the hash; this is not necessary if other access control mechanism or the like are in place.

[0191] A data steward may, e.g., with the use of a smart contract, apply policies and procedures for data collection, data sharing, data provisioning, and data transfer, ensuring that data is properly classified and protected, and overseeing the use of data to ensure compliance with legal and ethical standards.

[0192] A Smart Contract application, e.g., a program, may store hashes on the distributed ledger, e.g., public DigiByte UTXO blockchain, optionally the data itself may also be included on the ledger.

[0193] The distributed ledger improves control, e.g., in terms of securing the process. For example, the following advantages may be implemented in a layer 2 system.

[0194] 1. Ownership-Based: Every relation or data-entry has at least an owner. Owners are in full control of the underlying data

[0195] 2. Decentralized: The protocol need not be managed by a single party. For example, a subset of data owners must approve changes. Protocol allows it to be run on decentralized blockchain. For example, governance criteria such as defined in the Ethereum NFT protocol, e.g., Smart Contract Protocol.Erc-721, may be used. For example, every approver may use his own wallet to sign a transaction.

[0196] 3. Change-Request Management: Changes to data must be approved by their owners (l...n)

[0197] 4. Immutable: Once approved changes are added to the blockchain, they cannot be altered due to the immutable nature of the technology.

[0198] 5. Multiple Data sources: Allows the use of local or decentralized KV-databases to store the data. For example, a protocol may support basic setup templates for them, e.g., using IPFS. 6. Type restrictions: Stored data can / must follow specific type constraints or can consist of raw data. Type constraints must be specified when creating a relation

[0199] 7. Redundant: Multiple instances can be spun up. Due to the deterministic properties of the protocol every instance will have the same view on the data

[0200] 8. Programmable: The system implementing the protocol may provide a basic set of API functions that allow programmability

[0201] 9. Secure: Data sources may be configured for data encryption . Note that data is typically not stored on the chain. For example, data may be stored locally, e.g., in local storage, e.g., a Hard Disk Drive, or IPFS. In case of IPFS, the data is preferably encrypted, as otherwise people could get access to the data. The encryption key can be a company key, or other encryption key. A key may be provided by the layer 2 system.

[0202] For example, the layer 2 system may be implemented in one or more data management devices 210. A user may, for example, perform one or more of the following operations.

[0203] 1. Set-up. A setup operation may be the very first operation that is performed when setting up a database. For example, this may be the root command that empowers an owner to be the administrator of the database and root context. Administrators may have full access to users, owners, data, etc.

[0204] 2. Create a context that references to a parent context, or root node

[0205] 3. Empower another user to become an administrator of a context. For example, this may only be callable by the administrator of the parent context

[0206] 4. Remove access / control of an administrator of a specific context

[0207] 5. Register a data type that can be used in a context, or globally

[0208] 6. Create a relation, e.g., with given context-id and type constraints, within a context. The caller must be an owner of the context. Type and access constraints will be specified within the protocol

[0209] 7. Create a dataset within a relation. Context-ID and KV-Store-ID must be specified when using this operation. Succeeds if the caller is the owner of the relation. Otherwise, it will be in pending state until the owner approves the creation

[0210] 8. Alter a dataset within a relation. Dataset-ID and KV-Store-ID must be specified when using this operation. Succeeds if the caller is the owner of the relation. Otherwise, it will be in pending state until the owner approves the creation 9. Remove a dataset from a specific relation by specifying the dataset identifier. Succeeds if the caller is the owner of the relation. Otherwise, it will be in pending state until the owner approves the creation

[0211] 10. Change the access constraints of a context. For instance: Allow anyone to alter data, or require at least two / three owners to approve the change. Only callable by the administrator of the context

[0212] 11. Approves a change request. Only callable by the administrator of the parent context

[0213] For example, the level 2 system will create a context address, e.g., a target address, to which operations will be sent on the distributed ledger. It may be required that operations are sent by respective owners. The system may track UTXOs of the owners, so it will be able to decode sender addresses. For example a detailed embodiment may use the following elements.

[0214] System Startup Parameters

[0215] 1. Seed phrase: determines database contract address and provides gas

[0216] 2. DigiByte RPC node connection

[0217] 3. Data source Settings. For example: IPFS, content encryption: seed phrase System State

[0218] The state of the chain sync as well as the IPFS data may be stored in a local database, e.g., a sqlite database

[0219] System UI

[0220] The system may be programmed in NextJS or the like, to provide a modem user interface allowing URL access

[0221] Identifiers

[0222] 1. Operation IDs may be calculated as follows: RIPEMD160(txid). Collisions are unlikely

[0223] 2. Content IDs may be stored within the OP RETURN data. IPFS SHA256 contentlds are by design 32 + ~1 bytes, which will fit onto the 80-byte limit of modem UTXO blockchains

[0224] Figure 4 schematically shows an example of system components of an embodiment of a data management system. Shown in figure 4 is a data management system 410. System 410 comprises a blockchain synchronization unit 411 configured to connect to a distributed ledger node 421, e.g., a DigiByte core node, e.g., to exchange ledger data, e.g., blockchain data. System 410 comprises a core 412 configured to connect to external storage 422, e.g., to an IPFS node. System 410 comprises a database 413. Core 412 and database 413 are configured to exchange state and data of the database. System 410 further comprises a web interface 414.

[0225] The system consists of several core components. In order to have access to the blockchain, it may synchronize the blockchain data by maintaining a connection to a node, e.g., a DigiByte Core Node. For example, the synchronization unit 411 may push messages to core component 412 which is configured to handle all the inputs and fetch the required data, store data in a database, etc. Web interface 414 may be configured to display the tables, owners, and forms.

[0226] Figure 5 schematically shows an example of an embodiment of a data management system, in particular an operations scratchboard and flow. Figure 5 comprises

[0227] 510: Setup, e.g., Private key generation and configuration

[0228] 520: MDM application

[0229] 521 : Bootstrap

[0230] 522: Generate fund addresses

[0231] 523: Wait for funds

[0232] 524: Generate contract address

[0233] 525: Root owner

[0234] 526: Issuance Transaction for owner address

[0235] Waiting for funds 523 may further obtain a public key generated from the private key. Root owner 525 may obtain a contract address and persist a root owner address on a contract address.

[0236] For example, we may have the following operations. a) System administrator installs the application. While installing, he configures a connection to the distributed ledger, e.g., a DigiByte RPC node connection. Moreover, he generates a private key, e.g., by using the CLI. The web interface may be made accessible within the company using the embodiment. b) Database administrator can now access the web interface with his browser and initiate the application bootstrap. The provided private key file will be the key to the database c) He will have to send utility tokens to an address, e.g., a funding address generated by the system. The funding address may be dependent on the private key file. d) He then has to sign an ID request to register himself as the root administrator e) The system may issue a root transaction that writes the owner address to the distributed ledger, e.g., a transaction from owner address to contract address. In an embodiment, only the very first transaction to the address is valid, e.g., a number smaller than 5. Every subsequent transaction will be ignored by the system. A full blockchain scan may be performed from a specific block height, e.g., block 15M.

[0237] Figure 6 schematically shows an example of an embodiment of a data management system. In figure 6, 601 denotes information that is stored on chain. The tokenUri function at 602 is a function that gives an external URL for a specified token number. The token number may be an integer value.

[0238] The layer 2 system may create a copy of the dataset, e.g., in a specific binary format, and that binary data may be hashed. This binary data may be stored using, possibly multiple times. Data could be stored on a hard disk, but it could also be IPFS, any database system, or the like. The full dataset may be imported in the layer 2 system, e.g., using an oracle or a specific data connector. This allows changes of these data to be recorded in the layer 2 system. Hashes of these data, time stamps and possible changes may be stored as hashes on the distributed ledger, e.g., a layer 1 UTXO blockchain. Preferably, only the hash of the data is stored on the distributed ledger. The hashes are derived from the original content. This uses ledger storage in an efficient way, since only hashes are stored.

[0239] The setup may use a distributed ledger interface, e.g., a distributed ledger application, e.g., the DigiByte Core application, running within one setup. This could be offered as a service, but could also be an on-premise installation on private servers.

[0240] The layer 2 system may additionally be configured with further features, e.g., further layers, e.g., an LDAP system. The LDAP system may be configured to grant certain users access to the web interface, send mails to the users, and make the management easier. Distributed ledger addresses, e.g., a DGB address, and / or public keys may be stored on the ledger, e.g., on chain.

[0241] The layer 2 system may be used when connected to other applications, e.g., to register activities, time-stamped on the blockchain to create an audit trail. It may use connectors to an application. For example, the layer 2 system may be combined with peer- to-peer video conferencing, wherein an audit trail registers the participants and times.

[0242] The layer 2 system may be used when connected to other applications to validate and verify authenticity of digital content. It may issue a validation of digital content, hashed on the blockchain, and the possibility to verify that certain validated content is original and has not been changed.

[0243] Validation of issuance, updates, transfer, and provisioning by authorized individuals defined in smart contract, may require authorization, e.g., using a passwordless QR-based authentication using.

[0244] In an embodiment, a REST protocol may be used for communication with a user front-end for access and management of data.

[0245] The following shows in stripped down, pseudocode form some of the functions that may be used. All functions results refer to the current state, e.g., the best block.

[0246] / / Return info about the current application state info

[0247] / / Return funding info (returns contract address and amount of DGB) funding

[0248] / / Get context with id ‘id’ context / : id

[0249] / / Get table with id ‘id’ relation / :id

[0250] / / Get history of operations(paginated) history

[0251] / / Get change requests (paginated), excludes preview of data change-requests

[0252] / / Get information about single change request with the id ‘changeRequestld’. Includes preview of data change-request-diff / : changeRequestld

[0253] / / will return a tree of all the data tables tree

[0254] / / will return all the data types used by the tables types

[0255] The following events may be defined. Events are not necessary. They may be used implicitly.

[0256] • CHANGE REQUEST ACCEPTED

[0257] • APPROVAL

[0258] • OWNER EMPOWERED

[0259] • TYPE REGISTERED

[0260] • RELATION CREATED

[0261] • CONTEXT CREATED

[0262] Figure 7a schematically shows an example of an embodiment of a data logging system 800. System 800 is configured to facilitate secure and verifiable logging using a distributed ledger.

[0263] Data logging system 800 comprises multiple applications; shown are a first application 811 and a second application 812. The applications produce logging data. For example, first application 811 and second application 812 may produce logging item 821 and logging item 822, respectively.

[0264] For example, a logging item may comprise events related to an application. For example, a logging item may relate to a storage. In this example, a logging item may comprise events related to the storage, objects and / or files that are stored. The objects and / or files may be represented as a computer address, e.g., a URL, possibly together with a hash or signature over the stored object and / or file.

[0265] As an illustrative, non-limiting example, logging item 821 may comprise

[0266] EntAppl : event 1

[0267] EntAppl : event2

[0268] EntApp2: event 1

[0269] EntApp2: event2

[0270] EntStol : event 1

[0271] EntStol : event2

[0272] For example, logging item 822 may comprise EntStol :Objectl

[0273] EntSto2:Object2

[0274] EntSto2:File 1

[0275] EntSto2:File 2

[0276] Here EntAppl and EntApp2 refer to two applications, here ‘enterprise applications’, and EntStol and EntSto2 refer to two storages, here ‘enterprise storage’. In this case, the applications collect multiple logging events in one logging item. This is efficient but not necessary; e.g., a single logging event may be comprised in a logging item.

[0277] System 800 may comprise an optional logging interface 813 that serves as an interface between the applications, e.g., applications 811, 812, and a data management device 810. For example, data management device 810 may be configured as data management device 210, and / or as data management device 120. However, data management device 810 may be configured to directly receive logging items from applications. Logging interface 813 may be configured for the retrieval and processing of logging items. Logging interface 813 may be configured for verifying logging items.

[0278] Data management device 810 is configured to retrieve one or more logging items from the one or more applications.

[0279] Optionally, data management device 810 may be configured to retrieve a previous logging collection (not separately shown in Figure 7a). Data management device 810 may be configured to retrieve a first hash 841 for the previous logging collection. Data management device 810 may be configured to verify first hash 841 against the previous logging collection. First hash 841 for the previous logging collection may have been stored in the distributed ledger earlier.

[0280] Data management device 810 may be configured to append the new logging items, e.g., logging items 821, 822, to the previous logging collection, forming an updated logging collection 823. A new hash 843 is computed for updated logging collection 823.

[0281] Note that access to the previous logging collection is not strictly necessary, in which case the updated hash may be formed by hashing first hash 841 and the new logging items together. The updated logging collection 823 may then be restricted to, e.g., first hash 841 and the new logging items, e.g., items 821 and 822.

[0282] The processing may be done under the control of a program (not separately shown in Figure 7a); e.g., forming updated logging collection 823 and / or first hash 841. The program and a second hash 842 may be retrieved, and the second hash may be verified for the program. The new hash 843 may be stored on a distributed ledger, e.g., using a distributed ledger interface 240. A transaction identifier 844 is generated and returned by distributed ledger interface 240. Transaction identifier 844 identifies the transaction on the distributed ledger that contains new hash 843.

[0283] Updated logging collection 823 may be made available in a third application 830, e.g., a log file browser or a log file storage system. Likewise, transaction identifier 844 may be made available in third application 830, e.g., associated with updated logging collection 823.

[0284] Through this configuration, data logging system 800 improves the tamperproof logging of logging items and their associated updates, leveraging the distributed ledger to provide verifiable proof of data integrity and provenance.

[0285] Figure 7b schematically shows an example of an embodiment of a data logging system 801. System 801 is similar to system 800 but has been simplified by omitting the retrieving and checking of second hash 842. Even the updated logging collection may be omitted. For example, the updated logging collection may not be made, or alternatively may not be made available in third application 830. Intermediate systems between system 800 and 801 are possible, e.g., omitting second hash 842 while still making the updated logging collection available.

[0286] Below, an additional embodiment is described that may be implemented in data logging system 800 and / or data logging system 801. The embodiment provides Blockchain-Based Data Logging for Storage Management Platforms. Instead of blockchain any other distributed ledger may be used.

[0287] An embodiment provides a data logging system integrated with blockchain technology to ensure tamper-proof and verifiable records of storage management activities. The system includes a logging layer that captures data points related to storage operations, such as administrative actions, user logins, data access events, and object storage operations, e.g., creation, modification, deletion. These operations are recorded in data blocks, each containing a unique cryptographic hash that ensures immutability and verifiability.

[0288] The logging layer is designed to generate a secure audit trail by bundling logging events into data blocks that are secured on a blockchain. The blocks are cryptographically linked, creating a chain of events that can be independently verified for data integrity. A distributed ledger interface enables the logging system to store cryptographic hashes on the blockchain and retrieve transaction identifiers that reference the logged events.

[0289] This system integrates with a storage management platform and is configured to maintain an immutable record of all data interactions. The cryptographic properties of the blockchain ensure that any modifications to the data are detectable, thereby providing a robust mechanism for verifying the provenance and integrity of stored information. Utilizing cryptographic hashing and digital signatures, the distributed ledger / blockchain ensures that alterations to the data are computationally impractical to make without detection, thereby safeguarding data integrity. In particular, modifications to the logging data will be detectable.

[0290] The system is further designed to support compliance with regulatory requirements by offering a readily verifiable, tamper-proof record of storage activities.

[0291] An embodiment may provide one or more of the following implementation features:

[0292] 1. Data Capture: The logging layer captures events, such as administrative actions, user interactions, and file storage modifications, in real time.

[0293] 2. Tamper-Proof Audit Trail: Logging data is cryptographically hashed and secured on the blockchain, ensuring that any unauthorized changes are detectable.

[0294] 3. Hash-Based Verification: A cryptographic hash for each data block ensures immutability and provides a means to verify the integrity of the logged data.

[0295] The system is particularly applicable in scenarios requiring rigorous data integrity guarantees, such as legal services, regulatory compliance environments, and industries with stringent data security requirements. For example, law firms can utilize the system to maintain an immutable record of sensitive documents and client interactions, ensuring that the integrity of their data remains uncompromised.

[0296] Figure 8 schematically shows an example of an embodiment of a data management system 802. Data management system 802 is configured to provide data provenance.

[0297] Data management system 802 comprises multiple storages; shown are a first storage 851 and a second storage 852. These storages may store objects and / or files. A data collection 861 is collated, containing entries that associate computer addresses, e.g., URLs, with their respective objects or files. For example, data collection 861 may include:

[0298] • URL 1 -> Object 1

[0299] • URL 2 -> Object 2 . URL 3 -> File 1

[0300] . URL 4 -> File 2

[0301] • URL 5 -> Object 3

[0302] For example, data collection 861 may indicate a relationship between the files and objects, e.g., a joint subject, e.g., relating to a same computer program, simulation, neural network training data, a joint owner, etc. A data provenance interface 853 facilitates communication between the storages, applications, and a data management device 850. For example, data provenance interface 853 may be configured to allow a user to construct and / or edit a data collection 861. For example, data management device 850 may be configured as data management device 210, and / or as data management device 120. Data management device 850 is configured to create proof of data collection 861 on a distributed ledger by storing a first hash 881 of collection 861 on the distributed ledger.

[0303] When modifications, e.g., additions or other alterations, occur in data collection 861, data management device 850 retrieves first hash 881 and verifies it against collection 861. Also, a new data collection 863 is created, e.g., the modified collection 861. Data management device 850 computes a new hash 883 for the data collection 863.

[0304] Note that access to the initial collection 861 is not strictly necessary, in which case verifying hash 881 may be skipped.

[0305] The processing may be done under the control of a program (not separately shown in Figure 8); e.g., forming updated collection 863 and / or first hash 883. The program and a second hash 882 may be retrieved, and the second hash may be verified for the program, e.g., by data management device 850.

[0306] The new hash 883 may be stored on a distributed ledger, e.g., using a distributed ledger interface 240. A transaction identifier may be generated and returned by distributed ledger interface 240 (not separately shown in Figure 8). The transaction identifier identifies the transaction on the distributed ledger that contains new hash 883.

[0307] Updated logging collection 863 may be made available in a third application 870, e.g., a file browser or a file storage system. Likewise, the transaction identifier may be made available in third application 870, e.g., associated with updated collection 863.

[0308] Through this configuration, data logging system 802 provides a tamper-proof provenance indication.

[0309] Below, an additional embodiment is described that may be implemented in data management system 802. An embodiment provides a data management system configured to ensure secure and tamper-proof data provenance through the use of blockchain technology. The system operates by creating unique data objects that encapsulate location information (e.g., URL) of stored objects or files and associated metadata (e.g., type, size, creation date). These data objects are cryptographically linked to the blockchain.

[0310] The system stores data provenance information in the form of blockchain- secured records. For example, a hash of the data collection, including associated URLs and metadata, is computed and stored on the blockchain. When modifications to the data collection occur, the system retrieves and verifies the stored hash before computing a new hash for the updated data collection. This process ensures that the integrity of the original and modified data collections is cryptographically verified.

[0311] The system comprises a data provenance interface that facilitates communication between storage devices and a data management device. The interface allows users to construct, modify, and verify data collections.

[0312] A distributed ledger interface allows the data management device to store cryptographic hashes and retrieve transaction identifiers associated with those hashes. Data objects are uniquely identified by cryptographic tokens that include a hash of their contents, ensuring immutability and verifiability. The blockchain's cryptographic properties ensure that the records cannot be tampered with.

[0313] The system can generate a complete and verifiable audit trail for all data objects and associated metadata, providing clear documentation of data origin, ownership, and modifications.

[0314] This system is particularly suited to applications requiring strict data provenance and integrity, such as government agencies managing decentralized storage pools. The system enables users to track access, changes, ownership, and metadata across diverse storage technologies while maintaining an immutable record of provenance.

[0315] The embodiment supports complex environments where related files and objects are bundled into collections spanning multiple storage locations. The provenance of each case, including all its components, is guaranteed through blockchain verification, ensuring a secure and comprehensive audit trail regardless of the underlying infrastructure.

[0316] Figure 9a schematically shows an example of an embodiment of a data anonymizing system 900. Figure 9a shows a data source 910. For example, data source 910 provides a data collection 911, e.g., machine learning training data. Data collection 911 is anonymized by a data anonymizer 920, producing an anonymized data collection 921 and a key file 922. Anonymized data collection 921 is the same as data collection 911 except for the removal of private data. Key file 922 allows reconstruction of data collection 911 from the combination of key file 922 and anonymized data collection 921. For example, key file 922 may contain a list of locations in data collection 911, e.g., a byte location, and the removed data.

[0317] In an embodiment, data anonymizer 920 uses pseudonymization, e.g., replacing sensitive data values with cryptographically generated tokens. Key file 922 may comprise cryptographic data, e.g., a cryptographic key, allowing restoring the tokens back to the original values.

[0318] For example, private data may be recognized in collection 911 using metadata associated with the fields in structured data, where the metadata provides descriptive labels indicating the nature of the content, such as "Name," "Address," or "Social Security Number." Additionally, data may be recognized based on predefined patterns, such as email addresses conforming to the structure "username@domain.com," phone numbers following established numeric formats, or identifiers such as social security numbers matching specific syntactic patterns. In some embodiments, dictionaries or databases containing known sensitive terms, such as names or geographic locations, may be used to identify corresponding data. Machine learning models may also be employed to analyze and classify data based on learned features that indicate sensitivity, such as semantic context, frequency of occurrence, or proximity to other sensitive data.

[0319] A data management device 210 may be configured to receive anonymized data collection 921 and a key file 922. A hash of data collection 921, and optionally also of data collection 911, may be stored on the distributed ledger, as in an embodiment. Optionally, a stored hash may be updated as in an embodiment.

[0320] In an embodiment, data management device 210 may be configured to reconstruct data collection 911 from anonymized data collection 921 and a key file 922 and compute a hash of data collection 911.

[0321] Figure 9b schematically shows an example of an embodiment of a data anonymizing system 901. System 901 is the same as system 900 except for the addition of a cryptographic proof system 930. Cryptographic proof system 930 generates a cryptographic proof 931 that proves cryptographically that a key file exists that allows collection 921 to be amended to collection 911 and to produce the correct hash. For example, the following information may be provided on the distributed ledger: hl = hash (anonymized data collection 921) h2 = hash (data collection 911) cp = cryptographic proof that a key file 922 a reconstruction of data collection 921 from anonymized data collection 921 will have hash hl.

[0322] Optionally, instead of cp, a hash of cp may be provided. Anonymized data collection 921 and proof cp, e.g., in unhashed form, may be stored external to the distributed ledger.

[0323] Optionally, a hash may be stored on the distributed ledger of the program that computes and / or updates the values above, e.g., for verification as in an embodiment.

[0324] To construct the cryptographic proof cp, a non-interactive zero-knowledge proof (ZKP) method may be utilized. This approach allows the cryptographic proof system 930 to generate a proof that demonstrates, without revealing the key file or private data, that the transformation of anonymized data collection 921 using key file 922 results in data collection 911 with the correct hash h2. Specifically, the ZKP can attest to the correctness of the transformation and the resulting hashes by verifying that the steps of the anonymization and reconstruction processes conform to predefined rules encoded in a cryptographic protocol.

[0325] Cryptographic proof system 930 may employ a non -interactive zeroknowledge proof (NIZKP) system, such as Succinct Non-Interactive Arguments of Knowledge (SNARKs). SNARKs allow the generation of a compact proof that verifies the correctness of the transformation from anonymized data collection 921 to data collection 911 using key file 922. The proof is both succinct and computationally efficient, enabling verification without requiring interaction between the prover and verifier. This means that the proof can be generated once and verified multiple times by any party, ensuring that the transformation and resulting hashes hl and h2 are correct without exposing the underlying data or key file. The proof or its hash can be stored on the distributed ledger.

[0326] Below, an additional embodiment is described that may be implemented in data anonymizing system 900 and / or 901. An embodiment provides a system and method for anonymizing large datasets while maintaining data provenance and integrity using blockchain technology. The system is configured to retrieve data from a storage management platform, process it through an anonymization engine, and create a verifiable record of the anonymization process.

[0327] The anonymization engine replaces sensitive information, such as names and social security numbers, with pseudo-codes, ensuring that the structure of the data remains intact for subsequent analysis or machine learning tasks. A key file is generated during this process, mapping the original data to the pseudo-anonymized data. This key file maintains traceability between the datasets and may be safeguarded by cryptographic techniques.

[0328] The key file is hashed using a cryptographic hashing algorithm, such as SHA- 256, to produce a cryptographic hash that uniquely identifies the mapping. The generated hash is stored on a blockchain, providing an immutable and tamper-proof record of the relationship between the original and anonymized datasets.

[0329] An embodiment may implement one or more of the following features:

[0330] 1. Data Retrieval: A dataset is retrieved from a storage management platform for processing.

[0331] 2. Anonymization: An anonymization engine processes the data, replacing sensitive information with pseudo-codes.

[0332] 3. Key File Generation: A key file maps the original sensitive data to the pseudo-anonymized data, enabling traceability.

[0333] 4. Hashing and Blockchain Storage: The key file is hashed using a cryptographic algorithm, and the resulting hash is stored on a blockchain.

[0334] This system may be used, e.g., for anonymization of datasets for machine learning and artificial intelligence applications while adhering to data privacy regulations.

[0335] Below, an additional embodiment is described. The mechanism described below may be used in any of the embodiments described herein, in particular those described with reference to Figures 7a-9b. The system may comprise one or more or all of the following parts

[0336] 1) A mobile identity application

[0337] 2) An Identity Server

[0338] 3) A protect application: Manages the locking and unlocking of storage management platform accounts adhering to strict security protocols while ensuring user- friendly operations through a streamlined process and protocol.

[0339] 4) A storage management platform: A separate application integral to the process. By default, all storage management platform accounts are initially locked. Unlocking occurs only when all required signatures, as defined by the enterprise administrator, are collected. The system supports multi -signature protocols, such as 2-of- 4, for unlocking procedures. Here is an outline of a possible process:

[0340] 1) Initiation: An individual cannot access the protected storage management platform accounts without undergoing a structured unlocking process.

[0341] 2) Signature Submission: To unlock the account, an individual signs a message using their private key through the mobile identity application.

[0342] Example payload:

[0343] { "userid": "<uuid-v4>", "action": "unlock- storage management platform - account", "Accountld": "<uuid-v4>" }

[0344] In an embodiment, the payload contains more detailed information.

[0345] 3) Signature Verification: The payload, along with the signature, is sent to the protect application, which verifies the signature's validity.

[0346] 4) Notification: Once the signature is verified, the protect application notifies relevant employees that a colleague is requesting access to the storage management platform account.

[0347] 5) Approval Window: There is a limited time window for these employees to approve the unlock request.

[0348] 6) Access Grant: If the required number of signatories approve, the initiating user is allowed 60 minutes to activate the unlocking mechanism, granting temporary access as predetermined by the administrator. The time limit of 60 minutes can be replaced by another predefined time-window.

[0349] 7) Monitoring: Throughout the access period, the protect application monitors any alterations the user makes to the account storage.

[0350] 8) Re-locking: The user can manually re-lock the account, or it will automatically re-lock after the designated access period expires.

[0351] 9) Review and Confirmation: Once re-locked, designated verification officers, e.g., a data value owner (DVOs), can review changes made during the access period, which are logged and displayed interactively in the user interface.

[0352] 10) Final Approval (optional for now): DVOs can endorse the alterations by creating a digital signature of approval, similar to the initial JSON payload used for unlocking.

[0353] Additionally, cryptographic proofs (NFD logs) are systematically logged, SHA-256 hashed, and stored on a blockchain to ensure data integrity. It lies in the nature of the hash computation that the data is now anonymized. The hashes put to the blockchain contain metadata such as a timestamp, encoded in a binary message. Instead of SHA-256 another cryptographic hash may be used. For example, a post-quantum safe hash function may be used.

[0354] The above process uses an epoch-based extraction algorithm that hashes designated regions of the logfiles.

[0355] These hashed sections are then recorded on the blockchain, providing a tamper-proof record.

[0356] A tool may reconcile these transactions. As blockchain transactions are immutable, they establish a verifiable record of all user interactions and requests, enhancing the security of the system.

[0357] The log file service employs a private key and nonce system, ensuring distinct transactions are associated with the appropriate application instance. This allows for multiple instances of the application to run concurrently without data overlap. During reconciliation the CLI knows which transactions belong to its scope.

[0358] In an embodiment, an Ethereum Virtual Machine (EVM) chain is used for the distributed ledger, however other chain types may be used instead, e.g., UTXO blockchains.

[0359] In an embodiment, one or more of the following are implemented:

[0360] Monitoring: User activities are continuously monitored and compiled into reports

[0361] Auditing: All actions are auditable, with immutable logs sent to a blockchain

[0362] Enforced Access Policies: Approval-based, time-limited access controls enforce strict adherence to access management policies

[0363] Authentication: Secure authentication using QR code and zero-knowledge technology

[0364] Authorization: Access to protected accounts is authorized only after obtaining required cryptographic approvals

[0365] Least Privilege Access Model: Users are only granted the minimal necessary access to protected storage management platform accounts

[0366] Segregation of Duties: Access requests require approvals from multiple entities Figure 10 schematically shows an example of an embodiment of a method 700 for data management using a distributed ledger. Method 700 may be computer implemented and comprises retrieving (710) a data item (221) from a data source (211) different from the distributed ledger, retrieving (720) a program (222) from a data source (212) different from the distributed ledger, retrieving (730) a first hash (241) for the data item and a second hash (242) for the program from the distributed ledger, and verifying the first and second hashes against the retrieved data item and the retrieved program, executing (750) the program (222) using at least the data item (221) as input, obtaining therefrom a new data item (223) resulting from said execution, storing (760) a new hash (243) derived from at least the new data item on the distributed ledger (300).

[0367] Many different ways of executing the method are possible, as will be apparent to a person skilled in the art. For example, the order of the steps can be performed in the shown order, but the order of the steps can be varied, or some steps may be executed in parallel. Moreover, in between steps other method steps may be inserted. The inserted steps may represent refinements of the method such as described herein, or may be unrelated to the method. For example, some steps may be executed, at least partially, in parallel. Moreover, a given step may not have finished completely before a next step is started.

[0368] Embodiments of the method may be executed using software, which comprises instructions for causing a processor system to perform an embodiment of method 700. Software may only include those steps taken by a particular sub-entity of the system. The software may be stored in a suitable storage medium, such as a hard disk, a floppy, a memory, an optical disc, etc. The software may be sent as a signal along a wire, or wireless, or using a data network, e.g., the Internet. The software may be made available for download and / or for remote usage on a server. Embodiments of the method may be executed using a bitstream arranged to configure programmable logic, e.g., a field-programmable gate array (FPGA), to perform an embodiment of the method.

[0369] It will be appreciated that the presently disclosed subject matter also extends to computer programs, particularly computer programs on or in a carrier, adapted for putting the presently disclosed subject matter into practice. The program may be in the form of source code, object code, a code intermediate source, and object code such as partially compiled form, or in any other form suitable for use in the implementation of an embodiment of the method. An embodiment relating to a computer program product comprises computer executable instructions corresponding to each of the processing steps of at least one of the methods set forth. These instructions may be subdivided into subroutines and / or be stored in one or more files that may be linked statically or dynamically. Another embodiment relating to a computer program product comprises computer executable instructions corresponding to each of the devices, units and / or parts of at least one of the systems and / or products set forth.

[0370] Figure Ila shows a computer readable medium 1000 having a writable part 1010, and a computer readable medium 1001 also having a writable part. Computer readable medium 1000 is shown in the form of an optically readable medium. Computer readable medium 1001 is shown in the form of an electronic memory, in this case a memory card. Computer readable medium 1000 and 1001 may store data 1020 wherein the data may indicate instructions, which when executed by a processor system, cause a processor system to perform an embodiment of a data management method, according to an embodiment. The computer program 1020 may be embodied on the computer readable medium 1000 as physical marks or by magnetization of the computer readable medium 1000. However, any other suitable embodiment is conceivable as well. Furthermore, it will be appreciated that, although the computer readable medium 1000 is shown here as an optical disc, the computer readable medium 1000 may be any suitable computer readable medium, such as a hard disk, solid state memory, flash memory, etc., and may be non-recordable or recordable. The computer program 1020 comprises instructions for causing a processor system to perform an embodiment of said data management method.

[0371] Figure 11b shows in a schematic representation of a processor system 1140 according to an embodiment. The processor system comprises one or more integrated circuits 1110. The architecture of the one or more integrated circuits 1110 is schematically shown in Figure 11b. Circuit 1110 comprises a processing unit 1120, e.g., a CPU, for running computer program components to execute a method according to an embodiment and / or implement its modules or units. Circuit 1110 comprises a memory 1122 for storing programming code, data, etc. Part of memory 1122 may be read-only. Circuit 1110 may comprise a communication element 1126, e.g., an antenna, connectors or both, and the like. Circuit 1110 may comprise a dedicated integrated circuit 1124 for performing part or all of the processing defined in the method. Processor 1120, memory 1122, dedicated IC 1124 and communication element 1126 may be connected to each other via an interconnect 1130, say a bus. The processor system 1140 may be arranged for contact and / or contact-less communication, using an antenna and / or connectors, respectively.

[0372] For example, in an embodiment, processor system 1140, e.g., the data management device or system may comprise a processor circuit and a memory circuit, the processor being arranged to execute software stored in the memory circuit. For example, the processor circuit may be an Intel Core i7 processor, ARM Cortex -R8, etc. The memory circuit may be an ROM circuit, or a non-volatile memory, e.g., a flash memory. The memory circuit may be a volatile memory, e.g., an SRAM memory. In the latter case, the device may comprise a non-volatile software interface, e.g., a hard drive, a network interface, etc., arranged for providing the software.

[0373] While system 1140 is shown as including one of each described component, the various components may be duplicated in various embodiments. For example, the processing unit 1120 may include multiple microprocessors that are configured to independently execute the methods described herein or are configured to perform elements or subroutines of the methods described herein such that the multiple processors cooperate to achieve the functionality described herein. Further, where the system 1140 is implemented in a cloud computing system, the various hardware components may belong to separate physical systems. For example, the processor 1120 may include a first processor in a first server and a second processor in a second server.

[0374] The following clauses represent advantageous embodiments.

[0375] Clause 1. A method (700) for data management using a distributed ledger (300) comprising

[0376] - retrieving (710) a data item (221) from a data source (211) different from the distributed ledger,

[0377] - retrieving (720) a program (222) from a data source (212) different from the distributed ledger,

[0378] - retrieving (730) a first hash (241) for the data item and a second hash (242) for the program from the distributed ledger, and verifying the first and second hashes against the retrieved data item and the retrieved program,

[0379] - executing (750) the program (222) using at least the data item (221) as input, obtaining therefrom a new data item (223) resulting from said execution,

[0380] - storing (760) a new hash (243) derived from at least the new data item on the distributed ledger (300).

[0381] Clause 2. A method as in Clause 1, comprising - retrieving a first transaction identifier (224) identifying a first transaction on the distributed ledger, the first hash being retrieved from the first transaction, and / or

[0382] - retrieving a second transaction identifier (225) identifying a second transaction on the distributed ledger, the second hash being retrieved from the second transaction.

[0383] Clause 3. A method as in any of the previous clauses, wherein storing the new hash comprises obtaining a transaction identifier (244) identifying a transaction on the distributed ledger that stores the new hash, the method comprising storing the transaction identifier in a data source.

[0384] Clause 4. A method as in any of the previous clauses, comprising storing the new data item (223) in an external data source (213).

[0385] Clause 5. A method as in any of the previous clauses, wherein further first hashes for the data item are retrieved and verified against the data item according to a Merkle tree or similar data structure.

[0386] Clause 6. A method as in any of the previous clauses, wherein the data item comprises a list of attributes, the program being configured to change at least one of said attributes.

[0387] Clause 7. A method as in Clause 6, wherein the data item relates to a physical object, one of the attributes indicating an owner of the physical object indicating an owner of the physical object and digital object.

[0388] Clause 8. A method as in any of the previous clauses, wherein the program authenticates and / or authorizes a user of the program before the new data item results from the execution, the program storing an identity of the user, the new hash being further derived from the identity.

[0389] Clause 9. A method as in any of the preceding clauses, wherein

[0390] - the first hash is retrieved from a first transaction on the distributed ledger, the method comprises verifying a signature of the first transaction, and / or

[0391] - the second hash is retrieved from a second transaction on the distributed ledger, the method comprises verifying a signature of the second transaction.

[0392] Clause 10. A method as in any of the preceding clauses, wherein the new hash is stored in a transaction on the distributed ledger, the method comprising

[0393] - determining a data type from multiple data types for the new data item, a private key (251; 252) being associated with each data type,

[0394] - retrieving the private key associated with the determined data type, - signing the transaction with the retrieved private key.

[0395] Clause 11. A method as in any of the previous clauses, wherein

[0396] - the program is configured to receive and verify one or more authorization codes before the new data results from the execution, the new hash being further derived from the one or more authorization codes.

[0397] Clause 12. A method as in any of the previous clauses, wherein the distributed ledger is a UTXO blockchain, e.g., the DigiByte blockchain, e.g., DigiByte Public Open- Source Layer 1.

[0398] Clause 13. A method as in any of the previous clauses, wherein the data source is one of a local storage, e.g., a hard drive, a cloud storage, and a distributed storage network, e.g., IPFS.

[0399] Clause 14. A device comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations according to any of the previous clauses.

[0400] Clause 15. A transitory or non-transitory computer readable medium (1000) comprising data (1020) representing instructions, which when executed by a processor system, cause the processor system to perform the method according to any of clauses 1- 13.

[0401] Clause 16. A method for verifiable logging of data using a distributed ledger, comprising: generating, by one or more applications, a plurality of logging items corresponding to events associated with the one or more applications, retrieving the plurality of logging items, retrieving a first hash corresponding to a previous logging collection stored on the distributed ledger, appending the plurality of logging items to the previous logging collection to form an updated logging collection, generating a new hash for the updated logging collection, storing the new hash on the distributed ledger, and providing a transaction identifier corresponding to the new hash for retrieval from the distributed ledger.

[0402] Clause 17. The method of Clause 16, further comprising: verifying the first hash retrieved from the distributed ledger against the previous logging collection, and retrieving a program and a second hash for the program from the distributed ledger, and verifying the second hash against the retrieved program, the program being configured for the appending the plurality of logging items, generating a new hash for the updated logging collection, and optionally storing the new hash on the distributed ledger.

[0403] Clause 18. A method for managing data provenance using a distributed ledger, comprising:

[0404] - retrieving a data collection from a plurality of storages, the data collection comprising entries associating computer addresses with corresponding objects or files;

[0405] - generating a first hash of the data collection and storing the first hash on the distributed ledger;

[0406] - detecting a modification to the data collection, including at least one addition, removal, or alteration of entries in the data collection;

[0407] - verifying the first hash against the data collection prior to modification;

[0408] - generating a new hash for the modified data collection; and

[0409] - storing the new hash on the distributed ledger.

[0410] Clause 19. The method of Clause 18, further comprising:

[0411] - providing a data provenance interface configured to enable construction, modification, and verification of the data collection;

[0412] - retrieving a transaction identifier associated with the stored hash from the distributed ledger;

[0413] - associating the transaction identifier with the modified data collection; and making the modified data collection and its associated transaction identifier available in an application for access.

[0414] Clause 20. A method for anonymizing data using a distributed ledger, possibly as in any one of the preceding clauses, comprising:

[0415] - retrieving a dataset from a data source;

[0416] - processing the dataset through an anonymization engine to replace sensitive information with pseudo-codes, thereby creating an anonymized dataset;

[0417] - generating a key file that maps the sensitive information in the original dataset to the corresponding pseudo-codes in the anonymized dataset; - hashing the key file using a cryptographic hashing algorithm to produce a cryptographic hash; and

[0418] - storing the cryptographic hash of the key file on a distributed ledger, thereby providing a record of the relationship between the original and anonymized datasets.

[0419] Clause 21. The method of Clause 20, further comprising:

[0420] - recognizing sensitive information in the dataset using metadata associated with structured data fields, predefined patterns, or dictionaries of known sensitive terms; and

[0421] - optionally employing a machine learning model to classify data as sensitive based on semantic context, frequency of occurrence, or proximity to other sensitive information.

[0422] Clause 22. The method of any one of clauses 20-21, wherein the anonymized dataset and key file are configured to enable reconstruction of the original dataset by combining the anonymized dataset with the key file, and further comprising:

[0423] - verifying the integrity of the anonymized dataset by storing a hash of the anonymized dataset on the distributed ledger.

[0424] Clause 23. A data anonymizing system, possibly as in any one of the preceding clauses, comprising:

[0425] - a data source configured to provide a dataset;

[0426] - an anonymization engine configured to replace sensitive information in the dataset with pseudo-codes, producing an anonymized dataset;

[0427] - a key file generator configured to map the sensitive information in the dataset to the pseudo-codes in the anonymized dataset;

[0428] - a hashing module configured to generate a cryptographic hash of the key file; and

[0429] - a distributed ledger interface configured to store the cryptographic hash of the key file on a distributed ledger, providing an immutable record of the anonymization process.

[0430] Clause 24. The system of any one of clauses 20-23, further comprising a data management device configured to:

[0431] - retrieve the cryptographic hash from the distributed ledger;

[0432] - reconstruct the original dataset from the anonymized dataset and the key file; and - verify the integrity of the reconstructed dataset by computing and comparing a hash of the reconstructed dataset with a previously stored hash.

[0433] Clause 25. The method of any one of clauses 20-24, further comprising:

[0434] - generating a cryptographic proof that demonstrates the existence of a key file enabling the reconstruction of the original dataset from the anonymized dataset, the proof being generated without revealing the key file or the private data;

[0435] - using the cryptographic proof to verify that the transformation from the anonymized dataset to the original dataset produces a hash matching the original dataset's hash; and

[0436] - storing the cryptographic proof or a hash of the proof on a distributed ledger for verifiable evidence of the anonymization process.

[0437] Clause 26. The method of any one of clauses 20-25, wherein the cryptographic proof is generated using a non-interactive zero-knowledge proof (NIZKP) system, the method further comprising:

[0438] - generating a compact proof using the NIZKP system that attests to the correctness of the transformation from the anonymized dataset to the original dataset based on predefined cryptographic rules; and

[0439] - enabling the proof to be verified multiple times by independent parties without requiring interaction between the prover and the verifier.

[0440] Clause 27. A system for secure anonymization and verification of data, possibly as in any one of the preceding clauses, comprising:

[0441] - an anonymization engine configured to replace sensitive information in a dataset with pseudo-codes to produce an anonymized dataset;

[0442] - a key file generator configured to map sensitive information to pseudocodes and generate a key file;

[0443] - a cryptographic proof system configured to: generate a cryptographic proof demonstrating that the key file allows the reconstruction of the original dataset from the anonymized dataset, and verify that the transformation produces the correct hash for the original dataset;

[0444] - a distributed ledger interface configured to store the cryptographic proof or a hash of the proof on a distributed ledger for immutable verification.

[0445] Clause 28. The system of Clause 27, wherein the cryptographic proof system employs a non-interactive zero-knowledge proof (NIZKP) system, such as Succinct NonInteractive Arguments of Knowledge (SNARKs), to: - generate a succinct and computationally efficient proof attesting to the correctness of the transformation from the anonymized dataset to the original dataset;

[0446] - ensure the proof can be independently verified without exposing the key file or sensitive information; and

[0447] - store the proof or its hash on the distributed ledger to enable multiple verifications by independent parties.

[0448] Clause 29. The method of any one of clauses 20-25, further comprising:

[0449] - storing a hash of the anonymized dataset, a hash of the original dataset, and the cryptographic proof on the distributed ledger to ensure the integrity and traceability of the anonymization and reconstruction processes.

[0450] It should be noted that the above-mentioned embodiments illustrate rather than limit the presently disclosed subject matter, and that those skilled in the art will be able to design many alternative embodiments.

[0451] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. Use of the verb ‘comprise’ and its conjugations does not exclude the presence of elements or steps other than those stated in a claim. The article ‘a’ or ‘an’ preceding an element does not exclude the presence of a plurality of such elements. Expressions such as “at least one of’ when preceding a list of elements represent a selection of all or of any subset of elements from the list. For example, the expression, “at least one of A, B, and C” should be understood as including only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The presently disclosed subject matter may be implemented by hardware comprising several distinct elements, and by a suitably programmed computer. In the device claim enumerating several parts, several of these parts may be embodied by one and the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0452] In the claims references in parentheses refer to reference signs in drawings of exemplifying embodiments or to formulas of embodiments, thus increasing the intelligibility of the claim. These references shall not be construed as limiting the claim.

Claims

CLAIMSClaim 1. A method (700) for data management using a distributed ledger (300) comprising retrieving (710) a data item (221) from a data source (211) different from the distributed ledger, retrieving (720) a program (222) from a data source (212) different from the distributed ledger, retrieving (730) a first hash (241) for the data item and a second hash (242) for the program from the distributed ledger, and verifying the first and second hashes against the retrieved data item and the retrieved program, executing (750) the program (222) using at least the data item (221) as input, obtaining therefrom a new data item (223) resulting from said execution, storing (760) a new hash (243) derived from at least the new data item on the distributed ledger (300).Claim 2. A method as in Claim 1, comprising retrieving a first transaction identifier (224) identifying a first transaction on the distributed ledger, the first hash being retrieved from the first transaction, and / or retrieving a second transaction identifier (225) identifying a second transaction on the distributed ledger, the second hash being retrieved from the second transaction.Claim 3. A method as in any of the preceding claims, wherein storing the new hash comprises obtaining a transaction identifier (244) identifying a transaction on the distributed ledger that stores the new hash, the method comprising storing the transaction identifier in a data source.Claim 4. A method as in any of the preceding claims, comprising storing the new data item (223) in an external data source (213).Claim 5. A method as in any of the preceding claims, wherein further first hashes for the data item are retrieved and verified against the data item according to a Merkle tree or similar data structure.Claim 6. A method as in any of the preceding claims, wherein the data item comprises a list of attributes, the program being configured to change at least one of said attributes.Claim 7. A method as in Claim 6, wherein the data item relates to a physical object, one of the attributes indicating an owner of the physical object indicating an owner of the physical object and digital object.Claim 8. A method as in any of the preceding claims, wherein the program authenticates and / or authorizes a user of the program before the new data item results from the execution, the program storing an identity of the user, the new hash being further derived from the identity.Claim 9. A method as in any of the preceding claims, wherein the first hash is retrieved from a first transaction on the distributed ledger, the method comprises verifying a signature of the first transaction, and / or the second hash is retrieved from a second transaction on the distributed ledger, the method comprises verifying a signature of the second transaction.Claim 10. A method as in any of the preceding claims, wherein the new hash is stored in a transaction on the distributed ledger, the method comprising determining a data type from multiple data types for the new data item, a private key (251; 252) being associated with each data type, retrieving the private key associated with the determined data type, signing the transaction with the retrieved private key.Claim 11. A method as in any of the preceding claims, wherein the program is configured to receive and verify one or more authorization codes before the new data results from the execution, the new hash being further derived from the one or more authorization codes.Claim 12. A method as in any of the preceding claims, wherein the distributed ledger is a UTXO blockchain, e.g., the DigiByte blockchain, e.g., DigiByte Public Open-Source Layer 1.Claim 13. A method as in any of the preceding claims, wherein the data source is one of a local storage, e.g., a hard drive, a cloud storage, and a distributed storage network, e.g., IPFS.Claim 14. A method as in any one of the preceding claims, comprising retrieving a plurality of logging items generated by one or more applications, the plurality of logging items corresponding to events associated with the one or more applications, wherein the first hash corresponds to a previous logging collection stored on the distributed ledger, - obtaining a new data item comprises appending the plurality of logging items to the previous logging collection to form an updated logging collection,Claim 15. The method of Claim 14, wherein the program is configured for the appending the plurality of logging items, generating a new hash for the updated logging collection, and optionally storing the new hash on the distributed ledger.Claim 16. A method as in any one of the preceding claims, comprising retrieving a data collection from a plurality of storages, the data collection comprising entries associating computer addresses with corresponding objects or files; generating a first hash of the data collection and storing the first hash on the distributed ledger; detecting a modification to the data collection, including at least one addition, removal, or alteration of entries in the data collection; wherein the new data item comprises said modified data collection.Claim 17. The method of Claim 16, further comprising: providing a data provenance interface configured to enable construction, modification, and verification of the data collection; retrieving a transaction identifier associated with the stored hash from the distributed ledger; associating the transaction identifier with the modified data collection; and making the modified data collection and its associated transaction identifier available in an application for access.Claim 18. A method as in any one of the preceding claims for anonymizing data, comprising: retrieving a dataset from a data source, e.g., a storage management platform; processing the dataset through an anonymization engine to replace sensitive information with pseudo-codes, e.g., anonymized identifiers, thereby creating an anonymized dataset; generating a key file that maps the sensitive information in the original dataset to the corresponding pseudo-codes in the anonymized dataset; hashing the key file using a cryptographic hashing algorithm to produce a cryptographic hash; and storing the cryptographic hash of the key file on a distributed ledger, thereby providing a record of the relationship between the original and anonymized datasets.Claim 19. A device comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations according to any of the previous claims.Claim 20. A transitory or non-transitory computer readable medium (1000) comprising data (1020) representing instructions, which when executed by a processor system, cause the processor system to perform the method according to any of claims 1-18.

Citation Information

Patent Citations

  • IMPROVED SYSTEM FOR SECURE TRANSMISSION OF AUTHENTICATION DATA

    NL2035471A