Decentralized and reliable logging method and system for data exchange traceability
A decentralized data sharing system using DLT and smart contracts addresses the limitations of centralized management by validating data transfers and preventing fraudulent logging, enhancing flexibility and reliability in cross-domain data exchange.
Patent Information
- Application Number
- JP2024117557
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2024-07-23
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2044-07-23
AI Technical Summary
Conventional data sharing platforms rely on centralized management, which incurs administrative costs and lacks flexibility in cross-domain data exchange, and existing DLT-based audit methods fail to provide a trusted record of data flow.
A decentralized data sharing system using a distributed ledger technology (DLT) that stores reputation scores for computing nodes, validates data transfers through smart contracts, and employs a reputation-based mechanism to prevent false logging, enabling trust between data providers and users without a centralized administrator.
The system reduces administrative costs, enhances flexibility in cross-domain data exchange, and ensures trusted record-keeping and auditing by preventing fraudulent data transfer logging, thus promoting reliable data exchange across different domains.
Smart Images

Figure 2025108343000001_ABST
Abstract
Description
Technical Field
[0001]
[0001] The present invention generally relates to a method and system for validating data transfer, and more particularly to a method and system for validating data transfer using a distributed ledger.
Background Art
[0002]
[0002] Cloud-based data sharing platforms help organizations seamlessly share, purchase, and sell data. These highly virtualized high-performance data platforms can be built within a data-sharing-as-a-service model where subscribers can manage, curate, and tailor data for a fee.
[0003]
[0003] Data exchange is a component / service of a data sharing platform that is essential for facilitating the purchase and sale of data. It enables data providers to expose their data assets and allows data users to view, compare, and collect data. Data exchange platforms and service providers typically operate with a centralized system administrator / server to process exchange activities in a secure and trustworthy manner. The administrator acts as a trusted point, or intermediary, to establish a highly reliable data trading relationship between data providers (DPs) and data users (DUs). The administrator is responsible for providing essential data governance functions such as ID / access control, data exchange validation, and reliable record keeping for auditing and traceability. Audit information includes information about the actual data exchange process, from transaction information such as accounts and payment details, and what was agreed upon in the contract, to what is happening in the actual data transfer over the communication link. Such information may include what data and how much data is involved in the exchange, what data was transferred, the source and destination of the data flow, communication link quality, the time used to transmit the data, etc.
[0004]
[0004] The arrangements of the embodiments will be understood and recognized from the following forms for carrying out the invention, which are made by way of example only and are to be read in conjunction with the figures.
Brief Description of the Drawings
[0005]
Figure 1
[0005] FIG. 1 is a schematic diagram illustrating a distributed data sharing system according to an embodiment.
Figure 2
[0006] FIG. 2 is a block diagram illustrating a distributed auditing function according to an embodiment.
Figure 3
[0007] FIG. 3 is a block diagram illustrating a logging process according to an embodiment.
Figure 4
[0008] FIG. 4 is a flowchart of a process for validating data transfer using the distributed data sharing system of FIG. 1.
Figure 5
[0009] FIG. 5 is a block diagram illustrating a state machine.
Figure 6A
[0010] FIG. 6A is a flowchart of an exemplary method for updating a reputation score using a dummy data user.
Figure 6B
Figure 7
Figure 8
[0011] FIG. 8 is a block diagram of an exemplary implementation of the distributed data sharing system of FIG. 1.
DETAILED DESCRIPTION OF THE INVENTION
[0006]
[0012] According to an embodiment, a computer-implemented method is provided for validating a data transfer from a data provider to a data user via a first computing node of a network comprising a plurality of computing nodes. The plurality of computing nodes communicate with a distributed ledger storing, for each computing node, a respective reputation score derived based on a plurality of previous data transfers associated with that node. The method comprises obtaining, by the first computing node, a data transfer record specifying the data transfer, and submitting, by the first computing node, the data transfer record to a smart contract on the distributed ledger to validate the data transfer record by determining that the reputation score of the first computing node exceeds a predefined threshold.
[0007]
[0013] The distributed ledger may further store, for each computing node, respective data logs that specify a plurality of previous data transfers associated with that node. The method may further comprise sending, by a first computing node, a request to add a validated data transfer record to the data log associated with the first computing node to a smart contract on the distributed ledger.
[0008]
[0014] As an initial step, the method may comprise accessing the distributed ledger using a dummy user function to examine the data log of the first computing node, and determining, by the dummy user function and based on the examined data log, whether to update the reputation score of the first computing node stored on the distributed ledger, and calling, by the dummy user function, a smart contract based on determining to update the reputation score of the first computing node, and updating the reputation score of the first computing node based on the examined data log.
[0009]
[0015] The data log may further specify the status for data items involved in a plurality of previous data transfers. Determining whether to update the reputation score of the first computing node may comprise determining, by the dummy user function, whether a request by a data provider to revoke access to a data item is indicated by the status of the data item, and deciding to update the reputation score and decreasing the reputation score in response to determining that the status of the data item indicates the request. Additionally or alternatively, determining whether to update the reputation score of the first computing node may comprise determining, by the dummy user function, whether a plurality of data items meet a similarity criterion, and deciding to update the reputation score and decreasing the reputation score in response to determining that the plurality of data items meet the similarity criterion.
[0010]
[0016] Obtaining data transfer records may involve the first computing node monitoring network traffic between a data provider and a data user via a data plane associated with the first computing node.
[0011]
[0017] The data transfer record may specify one or more of contract information between the data provider and the data user, information describing a plurality of data items involved in the data transfer, information describing the relationship with a plurality of previous data transfers, information describing a data usage policy, a data transfer ID, a data transfer timestamp, source and destination IP addresses, and information describing data plane performance.
[0012]
[0018] According to an embodiment, a computer-implemented method is provided for validating a data transfer from a data provider to a data user via a first computing node of a network comprising a plurality of computing nodes using a distributed ledger. The distributed ledger stores respective reputation scores derived based on a plurality of previous data transfers associated with each node. The method comprises receiving, by the distributed ledger, a data transfer record specifying the data transfer, and validating, by the distributed ledger using a smart contract on the distributed ledger, the data transfer record when it is determined that the reputation score of the first computing node exceeds a predefined threshold.
[0013]
[0019] The distributed ledger may further store, for each computing node, a respective data log specifying a plurality of previous data transfers associated with each node. The method may further comprise receiving, by the distributed ledger, a request from the first computing node to add the validated data transfer record to the data log associated with the first computing node, and adding, by the distributed ledger using a smart contract on the distributed ledger, the validated data transfer record to the data log associated with the first computing node.
[0014]
[0020] As an initial step, this method may further comprise providing, by means of a distributed ledger, access to the data log of a first computing node for a dummy user function, and updating a reputation score of the first computing node based on the investigated data log in response to receiving a request from the dummy user function.
[0015]
[0021] The data log may further specify the status for data items associated with a plurality of previous data transfers, and the status of the data item may indicate a request by a data provider seeking to revoke access to the data item. Updating the reputation score of the first computing node based on the investigated data log may comprise updating the reputation score and decreasing the reputation score. Additionally or alternatively, the data log may further specify a plurality of data items associated with a plurality of previous data transfers. The plurality of data items may meet a similarity criterion. Updating the reputation score of the first computing node based on the investigated data log may comprise updating the reputation score and decreasing the reputation score.
[0016]
[0022] The data transfer record may be generated by a first computing node by monitoring network traffic between a data provider and a data user via a data plane associated with the first computing node.
[0017]
[0023] According to an embodiment, a computing node is provided that comprises a processor and a memory. The memory stores a plurality of instructions executable by the processor to implement the first or second embodiment.
[0018]
[0024] According to an embodiment, a computer-readable medium is provided that comprises a plurality of executable instructions that, when executed by a processor, cause the processor to implement the first or second embodiment.
[0019]
[0025] In summary, the present disclosure aims to overcome at least some of the drawbacks associated with conventional data sharing platforms that rely on centralized management. In particular, the present disclosure proposes a data sharing system that provides auditing and traceability capabilities without the need for a centralized administrator. Instead, data providers and data users can establish a trust relationship directly with each other through consensus. The elimination of the administrator reduces administrative costs and enables more flexible and cross-domain (e.g., not limited to a particular vendor of the data plane or data platform) data exchange.
[0020]
[0026] Broadly speaking, the present disclosure proposes embodiments that use a distributed ledger technology (DLT)-based method to track and validate data exchange activities between data users and data providers, thereby providing a system of record-keeping techniques for achieving trusted record-keeping and auditing. In particular, the present disclosure describes i) a method for generating an audit log (i.e., a log comprising details regarding, for example, a contract between a data user and a data provider, details regarding how data is transferred, and the like), ii) a method for validating a log creator / publisher, and iii) a method for publishing and storing the log (using smart contracts) on a distributed ledger. As will be described in detail below, embodiments may integrate a reputation-based mechanism with DLT to prevent false logging.
[0021]
[0027] FIG. 1 shows a distributed data sharing / exchange system 1 without a centralized administrator. The system 1 can be configured to enable a data user (DU) 3 to request data item(s) from a data provider (DP) 5 via a network 7 and obtain the data item(s) from the data provider (DP) 5. Generally, the (distributed) network 7 includes a plurality of computing nodes (also referred to as distributed data platform (DDP) nodes) and can be configured to implement a consensus-based alliance of data exchange ecosystem participants. As will be described in detail below, the network 7 can perform management functions, for example, creating data transfer records for an auditing process.
[0022]
[0028] The DDP node can be implemented as each computer program executed by one or more computers. The DDP node can typically interact with both the DP and the DU. In particular, each of the DDP nodes can be equipped with a data processing engine / core for processing data sharing / exchange requests from the DP and / or the DU. The DP / DU can be regarded as the "end user" of the DDP node. Further, the DDP node can be configured to collect, combine, and transmit data on a communication infrastructure (e.g., the Internet) to other DDPs. From this, in order to obtain access to the network (i.e., to be able to request data items from the DP connected to different DDP nodes and obtain data items from the DP), it is sufficient for the DP / DU to be connected to only one DDP node. Each DDP node can have a related data plane configured to implement the transfer of data items from the DP to the DU (e.g., to route / forward data packets received from the DP to the DU). Different DDP nodes can have different related data planes, that is, the network 7 can form a "cross-domain" data sharing platform / alliance. Providing different data planes for the DDP nodes enables cross-domain data platform (or cross-data space) auditing and traceability, and thus can promote an increase in data exchange in cross-domain scenarios where the level of trust between the data provider and the data user can be low.
[0023]
[0029] Generally, a DDP node can be configured to monitor and log any network traffic on its respective data plane. In particular, a DDP node can be configured to obtain (by monitoring the data plane) a data transfer record comprising information specifying data transfer between a DP and a DU in which the DDP node is involved. For this purpose, a DDP node can be equipped with a dedicated function for logging network traffic on its data plane and for fetching relevant logs when needed (e.g., when a data exchange agreement is agreed upon or when a data transfer is taking place). More specifically, a DDP node can be equipped with a data exchange control function configured to record information specifying a data exchange agreement (or contract) between a DP and a DU, such as a contract ID, a contract timestamp, actors or services involved in the contract (e.g., IP addresses, device IDs, etc.), a contract description, a detailed description of the data item(s) involved in the exchange, the relationship with previous exchanges / contracts, or the like. The detailed implementation of this function may vary between individual DDP vendors.
[0024]
[0030] Additionally or alternatively, the data exchange control function may record a data usage policy. The data usage policy may specify, for example, a policy for onward data processing and requirements for further reporting by the DU in order to maintain sufficient governance / control over the data item by the original data owner / provider.
[0025]
[0031] The DDP node may further comprise a transfer phase recording function for recording information during the data transfer phase. For example, this function may record data transfer IDs, corresponding data exchange contract IDs, data transfer timestamps, the source and destination of the data transfer (e.g., IP addresses, device IDs, etc.), a detailed description of the data being transferred (e.g., concatenation of a resource URL and a UUID for the DP, data plane queries, data fingerprints, watermarks, and data item IDs such as the like), data plane performance metrics (e.g., data rate), side information or custom tags as required in the usage policies of current and previous data exchange contracts.
[0026]
[0032] Information recorded by the DDP node (e.g., by the data exchange control function and / or the transfer phase recording function) may be accessed (e.g., fetched) through an application programming interface (API) by the log creator function (abbreviated as log creator) of the DDP node as a data transfer record.
[0027]
[0033] Referring to FIG. 1, DU3 communicates with a first DDP node 9, and DP5 communicates with a second DDP node 11. For example, DU3 may decide to purchase a data item provided by DP5, and thus may send a request for a specific data item to DDP node 9. DDP node 9 may then forward the request to DDP node 11 to facilitate the data transfer. In the embodiment of FIG. 1, DP5 and DU3 communicate with different DDPs, but it is understood that in other embodiments, DP5 and DU3 may communicate with the same DDP node.
[0028]
[0034] Network 7 can be further configured to have access to the distributed ledger, i.e., each of the DDP nodes (e.g., DDP nodes 9, 11) is configured to access the ledger (i.e., read from and / or write to the ledger). The distributed ledger can be configured to store a transfer history log (abbreviated as data log) that specifies previous data transfers executed via network 7. Generally, the purpose of the data log is to provide a traceable and auditable record of data transfers associated with network 7. The distributed ledger can store individual data logs for each DDS that specify previous data transfers associated with the respective DDS nodes.
[0029]
[0035] Generally, the distributed ledger can also host a smart contract (or multiple smart contracts) for implementing logic / functions. The term "smart contract" can refer to a self-executing application stored on and executed on the distributed ledger. For example, a smart contract application can be used to write data to and read data from the ledger. Thus, a smart contract application can provide the data stored on the distributed ledger to (authorized) clients / applications. In particular, the smart contract can implement the logic / functions necessary to implement the reputation-based validation method described below.
[0030]
[0036] The distributed ledger can be implemented as a permissioned blockchain network or in any other known suitable way. The distributed ledger can be hosted by the DDP nodes (i.e., the computing nodes of network 7 that implement the DDP nodes can also implement the nodes of the distributed ledger). In other embodiments, the distributed ledger can be hosted on a separate distributed computing network that communicates with network 7 to enable the DDP nodes to access the ledger.
[0031]
[0037] The log creator of the DDP node can fetch data transfer records (e.g., when data transfer occurs), and submit the records to the ledger so that the data transfer records can be added to the data log stored in the ledger (when successfully validated as described below). In this way, the data log forms an auditable log of transfer history accessible to all clients / users of the distributed ledger. This is in contrast to conventional DLT-based audit log recording methods in the field of data sharing that lack the ability to provide a trusted record of physical events. For example, conventional DLT-based methods for logging financial transactions in data trading cannot verify the data flow from the DP to the UP (e.g., the transfer of data items such as data files).
[0032]
[0038] Generally speaking, in order to build trust in the data log stored in the ledger, each data transfer record can be validated before being allowed to be added to the data log. This can be done by checking the level of reliability or reputation associated with the entity (i.e., a specific DDP node) submitting the data transfer record to the ledger. If the validation attempt fails (e.g., because previous events / disputes have caused a decrease in the reputation of the DDP node, as will be described in detail below), the data transfer may not be added to the ledger. Further, the corresponding DDP node can be prevented from participating in the data exchange system 1 (e.g., until the DDP node takes actions to restore its reputation level). Here, an embodiment of implementing such a reputation-based validation method is described.
[0033]
[0039] The distributed ledger can also store a reputation score for each DDP node. Generally, the reputation score can indicate the level of trust associated with a particular DDP node. For example, the reputation score for a DDP node can be {0, 1, 2,...}, where a higher numerical value indicates a better reputation for the DDP node. Since the reputation score is stored on the distributed ledger, all clients of the ledger can typically access the reputation scores of all DDP nodes. Smart contracts (on the distributed ledger) can be configured to create and modify reputation scores on the distributed ledger and implement the logic / functionality necessary to facilitate synchronization of the member nodes of the distributed ledger. As described below with reference to FIGS. 2-4, the smart contract can be configured to receive a data transfer record from the log creator of the DDP node for validation and add the validated transfer record to the distributed ledger.
[0034]
[0040] It is understood that the distributed ledger can store the reputation score for each DDP node in any suitable manner. For example, in an embodiment, the distributed ledger can store, in addition to or instead of, the reputation score for each DDP node as described above, an individual reputation score for each of the provided resources such that the reputation score of a particular DDP node can be calculated based on the reputation scores of the resources provided by the DDP node (e.g., by averaging the corresponding individual resource reputation scores). This can be advantageous in that the reputation can be validated for both the resource level and the overall DDP node level. For example, this can be useful in providing the DDP node with information useful for improving its overall reputation (e.g., by removing provided resources with a low reputation), and thus can be useful when a subset of the provided resources has a low reputation.
[0035]
[0041] The reputation score of a specific DDP node can be calculated based on the historical data exchange activities of the DDP node (e.g., by a reputation scoring function provided in a smart contract application). The reputation score of a specific DDP node can be (automatically) updated when the log recording events for each DDP are completed (e.g., when the data transfer records submitted by the DDP node are added to the data log).
[0036]
[0042] The reputation score can be further (periodically) updated based on "user-side" feedback provided by dummy data users (DDUs), as described in detail below. Thus, the reputation score can also be updated when feedback from a DDU is received.
[0037]
[0043] It is understood that the smart contract application can be configured to execute further logic / functions to implement the proper operation of the ledger. For example, the smart contract application can be further configured to perform an ID check to ensure that the smart contract can only be accessed (or called) by registered clients of the distributed ledger.
[0038]
[0044] The ledger may also store the product state associated with data items (or transfer IDs) within the data sharing / exchange process. For example, the DP may request that the DDP node store on the ledger the product state associated with a specific data item provided by the DP. The product state may be created (and updated) by using smart contracts on the ledger. For example, the product state of a data item may be one of {original, copy, split, shipped, received, investigated, labeled, sold, cancelled, etc...}. Each state value may be updated by smart contract logic when a new data transfer event is logged or when validation feedback from the DDU (as described below with reference to FIG. 5) is received. Generally, the logic / rules for state changes may be defined within a state machine in the smart contract.
[0039]
[0045] The DDP node may implement a dummy data user function (abbreviated as DDU) for querying the ledger to provide feedback for updating the reputation score. This means that based on the content of the ledger, the DDU may determine that the reputation score of a particular DDP node should be updated (e.g., decreased or increased), and accordingly may call a smart contract to update the reputation score. The DDU may be configured to update the reputation scores of DDP nodes other than the DDP node associated with the DDU. Smart contract rules may be used to prevent self-update (i.e., to prevent the DDU from updating the reputation score of the DDP node implementing the DDU).
[0040]
[0046] The DDU can be implemented as a separate client (i.e., the DDU can have a DLT-based client ID different from the DP / DU client ID) and can be activated to request a dummy transaction with the DDP node to investigate the ledger (thereby invoking the smart contract). During the dummy transaction, since the dummy transaction occurs within the same DDP node, there is no substantial data transfer / flow. As a result, the use of the DDU does not violate the common data usage / governance policy.
[0041]
[0047] The DDU can be activated according to a policy (set by an alliance of data exchange ecosystem participants or by the DDP node administrator), for example, according to a schedule or in response to a request (e.g., triggered by a predefined event). The details of the policy for activating the DDU can be kept confidential, but it is understood that the permission to use the DDU is given by the alliance. The record of the dummy transaction can also be recorded on the distributed ledger (and thus can be accessible by other clients of the ledger).
[0042]
[0048] Investigating the ledger by the DDU to provide feedback for updating the reputation score may involve investigating the data log and / or the status of data items. Two specific examples of how the DDU can derive an update to the reputation score are described below with reference to FIGS. 6 and 7. However, it is understood that the DDU can derive reputation score updates in several ways depending on the specific application. For example, the DDU can determine the (recent) change in the product status of a data item indicating a dispute, or the correlation between data items of separate DPs indicating replicated data items. In another example, the DDU can determine that the DDP node has manipulated the original record (e.g., to mask details such as the originator, ownership, etc. of the data item) before submitting the data transfer record to the smart contract.
[0043]
[0049] In an embodiment, appropriate DDU validation logic can be (in part) hard-coded functionality / logic within a smart contract for investigating the content of a submitted data transfer record when a smart contract application is called during a logging process (as described below with reference to FIG. 4). Generally, validation policies can be made public (i.e., known to all members of the distributed ledger) and can be automatically enforced such that they can be hard-coded into the chain code.
[0044]
[0050] Generally speaking, knowing that there are mechanisms for public records and post-event validation can prevent members of the distributed data exchange system 1 from fraudulently publishing / logging their data transfer events.
[0045]
[0051] An exemplary method of logging and validating data transfer between DP5 and DU3 is described with reference to FIGS. 2-4. FIG. 2 is a block diagram illustrating the data plane flow between various components of system 1. FIG. 3 is a block diagram illustrating the details of the validation process. FIG. 4 is a flowchart of an exemplary method.
[0046]
[0052] Generally speaking, each DDP node monitors its respective data exchange activities and reports and verifies the data exchange activities through the consensus process of the distributed ledger, such that all DDP nodes have access to synchronized and verified information regarding data exchange. Further, for example, the creation and update of data logs stored on the distributed ledger are sufficiently distributed among DDP nodes, except for minimal central management at the DDP node alliance level for defining usage policies and managing access to the distributed ledger.
[0047]
[0053] Referring to FIGS. 2 to 4, in the initial step S101, the DDP node 9 obtains a data transfer record that designates data transfer (e.g., by logging network traffic on its data plane as described above). More specifically, the log creator of the DDP node 9 can fetch (or receive) the corresponding data transfer record in response to, for example, the completion of data transfer from DP5 to DU3. The log creator component can then process each data transfer record by means of a hashing operation using the form of the Secure Hash Algorithm (SHA) before submitting the hashed record to the smart contract to validate the record (and add the record to the distributed ledger), as shown in FIG. 3. The hashing process can be implemented using any known suitable method and can be selected based on the data exchange requirements for protecting confidential information and privacy. Generally, the hashing process can provide sufficient granularity to separate individual parts of the data transfer record, for example, by creating a separate hash value for the "description of data" part. As illustrated in FIG. 2, a similar process can occur at the DDP node 11, i.e., the log creator of the DDP node 11 can fetch the corresponding data transfer record and execute the hashing operation.
[0048]
[0054] In step S102, the DDP node 9 submits the (hashed) data transfer record to the smart contract on the distributed ledger to validate the data record. More specifically, the log creator component of the DDP node 9 can be the user / client of the smart contract and can provide the data transfer record as input to the smart contract application.
[0049]
[0055] The smart contract can validate the received data transfer record by determining that the reputation score of DDP node 9 exceeds a predefined threshold. For this purpose, the smart contract can read each reputation score for DDP node 9 from the distributed ledger and compare the reputation score with the threshold (the threshold can also be stored on the ledger). Then, the smart contract can determine that the data transfer record submitted by DDP node 9 is valid if the reputation score of DDP node 9 exceeds the threshold. Generally speaking, the data transfer record can be considered valid because it is submitted by an entity with a (sufficiently) high reputation. Similarly, the smart contract can determine that the data transfer record submitted by DDP node 11 is valid if the reputation score of DDP node 11 exceeds the threshold.
[0050]
[0056] If the smart contract determines that the submitted data transfer record is valid, the smart contract enters the data transfer record into the ledger (i.e., adds the data transfer record to the data log). The smart contract can update each reputation score based on the successful validation of the data transfer record (i.e., the smart contract can increase each reputation score in response to successful validation). In an embodiment, the smart contract can also update the product status related to the relevant data item based on the successful validation.
[0051]
[0057] In an embodiment, when a (part of the) DDU logic is written into the smart contract, the data transfer record can also pass the corresponding function to be published onto the ledger. This can be integrated with the permission mechanism in the conventional DLT system.
[0052]
[0058] If the reputation score is below the threshold, the corresponding DDP node can be (temporarily) prohibited from interacting with the distributed ledger (i.e., future data transfer activities can be prohibited). The DDP node can be prohibited from data exchange activities until its reputation score recovers. The reputation score can recover, for example, by feedback provided by the DDU, by overriding / resetting the alliance permission, or by any other suitable criteria such as payment of a fine, provided external trustee / insurance, etc.
[0053]
[0059] From this, the distributed data sharing / exchange system 1 described above can reduce the various expense costs associated with conventional data sharing systems that rely on central management. Furthermore, as described above, system 1 enables the validation and investigation of what should be stored in the ledger without relying on a trusted external source / a trusted third party. Furthermore, before publishing a record on the ledger (e.g., to mask details such as the originator, owner, etc. of a data item), the risk that a DDP node will hide information from the data transfer record or modify the original record can be reduced because the DDU can detect such behavior and reduce the reputation score of the associated DDP node to prevent further data exchange activities of the DDP node.
[0054]
[0060] Here, detailed examples of specific implementations of some aspects of the above method are described. As described above, the ledger can store a data log and product status related to data items exchanged via the network 7. In one embodiment, this can be implemented by providing logic for creating (and updating) corresponding data objects within a smart contract. The provided logic can create a smart contract entity for each resource and can be capable of creating an association with the resource. For example, data items provided by the DP can be represented in the data log as "resource" data objects with attributes as described in Tables 1 and 2 below. Table 1: Resource Attributes
[0055]
Table 1
[0056] Table 2: Status Attributes
[0057]
Table 2
[0058]
[0061] From this, in this example, the resource object has an attribute "status" that specifies the state of the corresponding data item. The value of the status attribute can change according to the state machine model. FIG. 5 illustrates an exemplary state machine diagram for the status attribute of the resource object. Table 3 summarizes the corresponding permitted actions. As can be seen in Table 1, the resource object also has a "history" attribute in which the previous state values are logged. Table 3: Permitted Actions
[0059]
Table 3
[0060]
[0062] For example, a DDP node can identify new data items provided by the DP (by monitoring its data plane). After successful validation (as described above), the DDP node can call a smart contract to create a new resource data object (as part of the data log stored in the ledger). The "state" of the resource object can start in the published state. As seen in Table 1, a resource can be specified by a data plane URL and a resource ID. Policies regarding the name and conditions of use can also be passed to the state machine. A resource can have an associated reputation score attribute. The reputation score of the provided resource can start with the reputation score of the DDP node to which the DP is connected. This allows the reputation to be validated for both the resource level and the overall DDP level. The reputation score of a particular DDP node can depend on the individual reputation scores of the resources provided by that DDP node.
[0061]
[0063] Similarly, the smart contract can be equipped with logic to create a corresponding subscription object representing the subscription to a particular data item (by the DU) for each resource object representing the data item provided by the DP. The reputation of the resource subscription object can initially be inherited from the parent resource (i.e., the object representing the data item), i.e., the parent resource must exist in the audit chain. The ID of the resource can be appended with a URL or a data user data plane endpoint to create a unique subscription ID. The state machine can start in the subscribed state.
[0062]
[0064] A smart contract may have logic for updating a state machine, i.e., for transitioning from one state to another state on a resource or participating state machine. The permitted transition activation for a "cancel" action may be such that this action can only be executed by the DP and that the action applies to both the resource and participation. The permitted transitions for other actions may be such that they can only be executed by the DU. The DDU may be a special user who can execute a "validate" action.
[0063]
[0065] When a cancel action is invoked on a resource, the cancel action may also be invoked (by the corresponding logic in the smart contract) on all participation in the resource, any child resources associated with the resource, and any other derived products. In this way, all access to the associated resources will be blocked.
[0064]
[0066] The access action may control access to the corresponding data plane API, may only be executed in a validated state, and the data items may not be used until the corresponding resource participation object is validated. More specifically, the access action may be executed to obtain an access token required to access the data plane API. This may verify the state of the resource before issuing the required access token, such as a JSON web token. The token may only be issued if the participation in the resource is in a validated state.
[0065]
[0067] Referring to FIGS. 6A - B, an exemplary method for updating a reputation score using a DDU in response to the cancellation of access to a resource is described. FIG. 6A is a flowchart illustrating a process triggered by calling a cancellation action. In an initial step S601, the DP determines that access to data items provided by the DP needs to be cancelled. This can occur due to different reasons such as disputes or recalls resulting from GDPR requirements or other related reasons. In such cases, the DP communicates with the DDP node to request the cancellation of the data items.
[0066]
[0068] In step S602, the cancellation operation is executed by a smart contract and the state of the resource object is changed to "cancelled". This initiates the cancellation of all resources and subscriptions associated with its parent resource. Therefore, in step S603, it is determined whether there are child or subscription objects, and if so, step S602 is executed for the corresponding objects. When all subscriptions and child objects are cancelled, the process is completed (S104).
[0067]
[0069] Next, the DDU may perform an update of all corresponding reputations based on the number of related resources affected by the cancellation, as described below with reference to FIG. 6B. Thus, in step S605, the DDU examines the ledger and is activated to update the reputation score. As described above, this can occur based on a schedule or an event (e.g., the DDP node may activate the DDU after receiving a cancellation request from the DP).
[0068]
[0070] In step S606, the DDU obtains from the ledger all resources (lists) having the status attribute "cancelled". Since cancellation is typically triggered by a negative event such as a dispute, the DDU may decide to lower the associated reputation score. Thus, in step S607, the DDU calculates for each of the relevant resources an update to its respective reputation score. For example, the update may be determined based on the number of affected subscribers or the total number of child resources that utilize the parent resource.
[0069]
[0071] In an embodiment, the reputation score update may be calculated based on a reputation reduction factor given by: Reputation reduction = Number of subscribers related to the cancelled resource / Total number of subscribers related to the data provider, And the reputation reduction weight is given by: Reputation weight = Length of time until resolution / Acceptable resolution time.
[0070]
[0072] The overall score adjustment can then be calculated as the product of the reputation decrease and the reputation weight. This decrease can be applied for each "acceptable resolution time" cycle. Further, in step S607, the DDU calls the smart contract to update the reputation score, that is, to update each value of the "reputation" attribute of the relevant resources. More specifically, the dummy user can update the reputation of all resources associated with the DP in the resources that have been cancelled (via the audit chain code, that is, the smart contract) (that is, the DDU can not only lower the value of the reputation attribute of the cancelled resources, but also lower the other resources of the same DP). As a result, subsequent subscriptions to resources associated with the same data provider may not be permitted if the reputation threshold conditions are not met. In step S608, the DDU determines whether further resources (for example, resources of the "onward sale" DP, that is, the DP that provides data items incorporating data items originally provided by another DP that has cancelled the data item in question) need to be updated, and if so, executes step S607 for these resources. If there is no need to be updated, the process is completed (S609).
[0071]
[0073] Generally speaking, since a decrease in reputation may cause the DP to lose access to the market, the above process encourages the rapid resolution of disputes (or similar problems) that may affect many subscribers. These reputation updates are automatically performed by the DDU to ensure independence and avoid interference in the process.
[0072]
[0074] The above process also enables the (automatic) processing of a supply chain of data resources that can include an on-ward sales chain. For example, a DP that uses data provided by another DP in that on-ward product will also be affected by the reputation score update. In an embodiment, the on-ward sales DP may have a lower reputation decrease applied compared to the original DP of the resource being canceled (e.g., only 10% or 25% of the decrease). In this way, the loss of reputation resulting from using a seemingly reliable resource from a DP without knowledge can be controlled to ensure a level of fairness.
[0073]
[0075] A particular advantage of the process described with reference to FIGS. 6A - B is that the reputation score is updated directly (based on the revocation of access) to a data resource that can also be a further derived data resource. In this way, the reputation score can promote a reliable data supply chain and, at the same time, high availability (i.e., rapid resolution) in a completely decentralized manner across different domains.
[0074]
[0076] Referring to FIG. 7, a further exemplary method for updating the reputation score is described. Broadly speaking, this method can enable i) the detection of on-ward sales without the confirmation response of the original data provider, and ii) the reduction of the reputation score of the on-ward seller.
[0075]
[0077] Generally, a smart contract can be equipped with logic that enables the on-ward sale of data items such that the original data provider is confirmed. For example, a smart contract can be equipped with a "combine" operation that enables the creation of new resource objects (referred to as "child" and "parent" resources respectively) from an existing resource object. In this case, the reputation score of the child can be initially propagated from the parent resource.
[0076]
[0078] For example, a "highly reliable" DP may use a "combination" operation on a data resource that uses data from another resource. This may involve inserting the corresponding child resource ID into an existing parent resource when creating a new resource state machine within the audit chain.
[0077]
[0079] However, if the DP does not confirm and respond to this specialized resource as an input to the on - ward product, the DDU may detect such an omission and, as described below with reference to FIG. 7, may accordingly lower the reputation score of the DP. In step S701, the DDU is activated to investigate the ledger (as described above in S605). For example, the DDU may be activated periodically. In step S702, the DDU receives a plurality of resource objects (i.e., resources in the "published" state), for example, all resource objects on the ledger, or all resource objects associated with the DPs communicating with a specific DDP node.
[0078]
[0080] In step S703, the DDU calculates a first similarity score for each pair of objects within the plurality of resource objects. The similarity score may be determined based on criteria such as public data in the metadata and / or data resource descriptions (e.g., keywords describing the resource offering) (e.g., similar to conventional plagiarism detection based on matching keywords and word sequences).
[0079]
[0081] In step S704, the DDU determines whether any of the first similarity scores exceeds a threshold value, thereby determining whether a plurality of resource objects meet the first similarity criterion (to determine whether there is potential for on-ward sales). For example, the threshold value can be 40% (i.e., if the first similarity score exceeds 40%, the relevant resource can be further investigated by proceeding to step S705; if it does not exceed, the relevant resource is not further investigated and the process proceeds to S709). In an embodiment, to protect legitimate on-ward sales, meeting the first similarity criterion may further require that the potential parent resource does not include the data resource ID entry of the potential child resource (this indicates legitimate on-ward sales through the combined operation).
[0080]
[0082] If the first similarity score exceeds the threshold value (S705), the DDP joins the relevant resource to obtain further information. In step S706, the DDU calculates a second similarity score for the relevant pair of resources based on the ratio of data parameters that are the same in both resources (e.g., the name of the data element associated with the resource): Second similarity score = Number of matching data element names / Total number of data elements
[0081]
[0083] Advantageously, from this, the first and second similarity scores can be calculated from the information stored in the ledger without the DDU having to access the actual data.
[0082]
[0084] In step S707, the DDU determines whether the second similarity score exceeds the corresponding threshold. For example, the threshold for the second similarity score can be 80%. If the second similarity score exceeds the corresponding threshold, the data elements are considered to be similar enough to proceed with adjusting the associated reputation, and in step S708, the DDU calculates a reputation decrease for the last-created resource (since the previously created resource is the original resource). The decrease can be calculated as follows: Reputation decrease = Reputation * Similarity, where "Similarity" is either the first, the second, or a combination of the first and second similarity scores. The DDU can then call a smart contract to update the reputation score accordingly.
[0083]
[0085] From this, in this way, the reputation update can take into account the possibility that the later-created resource is an onward sale that has not actually been confirmed. By periodically updating the reputation, the DDU can also allow an acceptable resolution time before performing a reputation score update. Thus, when notified, the DP performs the combined actions necessary to confirm an onward sale and has one "acceptable resolution time" period to cancel the data resource offering.
[0084]
[0086] In step S709, the DDU determines whether the similarity of further resources needs to be determined. If it needs to be determined, step S703 is executed for these resources. If it does not need to be determined, the process is completed (S712).
[0085]
[0087] In an embodiment, the method further comprises steps S710 and S711, that is, when the DDU determines in step S707 that the second similarity score is not exceeded, the method may proceed to step S710. In this step, relevant data items are investigated to determine a third similarity score. For example, the third similarity score may be determined based on detecting matching data element values, or matching element sequences. For example, Third similarity score = number of matching data element values / total number of data elements, Or, Third similarity score = number of matching data element value sequences / total number of data element sequences.
[0086]
[0088] Alternatively, the third similarity score may be determined based on detecting the presence or absence of a specific known data watermark, such as a known sequence or data element value that would not normally be expected to occur. Such a watermark may be a synthetic or adjusted data element value that does not intentionally match the ground truth reality. For example, for a data element corresponding to a timestamp, the data and time can be adjusted by a predetermined and known amount from the actual timestamp of occurrence. In this way, the correlation score between resources can be reliably determined in a binary fashion (i.e., 0 or 1, where 1 indicates the presence of the watermark in the relevant resource) based on the discovery or non-discovery of the embedded watermark.
[0087]
[0089] In step S111, the DDU determines whether the third similarity score exceeds the corresponding threshold. For example, the threshold for the third similarity score may be 10%. From this, if the third similarity score exceeds the corresponding threshold, the process continues with step S708 as described above (i.e., since a portion of the original data is likely to be copied in the later offering, while adjusting the reputation).
[0088]
[0090] The method described above with reference to FIG. 7 can have the effect of reducing the reputation of very similar data provider resources, which is typically a desirable result since a market that includes many similar products will ultimately have a lower value for data users and become more like a commodity market. Thus, it is generally desirable to encourage more diverse data products (as described above).
[0089]
[0091] FIG. 8 is a block diagram illustrating a particular implementation of the distributed data sharing system of FIG. 1. It is understood that the embodiment of FIG. 8 is merely an example, and the system of FIG. 1 can be implemented in several different ways. In particular, any reference to specific commercially available software tools is made only to provide a better understanding of how some of the features described above can be advantageously implemented in practice, and is not intended to limit the scope of protection.
[0090]
[0092] The embodiment of FIG. 8 comprises a Hyperledger DLT approach (built on a blockchain managed by AWS). The system uses the International Data Spaces Association (IDSA) Data Space Connector reference implementation (IDSA provides a standardized mechanism for distributed data sharing, including support for different data planes and auditing). From this, this embodiment incorporates a conventional data plane that can be registered and used via conventional APIs.
[0091]
[0093] Within the blockchain, there are two members representing different organizations, associated with two separate AWS accounts. Member A publishes the chaincode used to validate and store the state data about the smart contract, and Member B approves the use of the chaincode for the channel. Subsequently, access to the chaincode is authorized based on an embedded policy that identifies the specific actions / roles and member types (DP and DU / DDU) required to execute the operations. Cognito IdP user groups are provided by Lambda functions and provide permissions associated with users who can access the fabric through the GraphQL API that appears through the fabric client API. In this way, the AuditIDS instance in the data IDS connector pod can access the fabric audit service.
[0092]
[0094] To incorporate the IDS connector approach into the DDP node-based distributed data market described above, control plane interactions are provided (IdP-based DDP registration, publishing and auditing the API by the data provider and the joining by the data user). Additionally, the IDSA data space connector may provide an optional contract negotiation step that can complement DDP API joining and monetization within the control plane. These can define the terms and conditions of data processing in the form of policies and assist in the selection of data providers / data planes. The additional auditing function provides logging support for data sharing between the DP and the DU and data usage according to the usage contract or data usage policy. This can be useful for providing dispute resolution, billing, and an immutable audit trail without room for debate. As described above, auditing is implemented using a DLT-based approach using smart contracts that use hashing and operation validation, allowing for decentralization to eliminate the need for a trusted central storage service or oracle.
[0093]
[0095] In the embodiment of FIG. 8, the IDS data space connector is a component that provides data plane functionality and supports integration with auditing, identity and broker, and app store functionality. These additional components are optional, for example, it is understood that these components are not required if the functionality is supported via an API manager (or considered unnecessary as a matter of design choice). For example, it may be possible to register a data connector and use an API manager to provide market integration. Thus, in this way, similar app store and auditing functions can also be provided partially or fully through the API manager integration function.
[0094]
[0096] The data space connector can be combined with AWS Hyperledger-based auditing and validation smart contract functionality. These, as described above with reference to FIGS. 5 and Tables 1-3, permit recording and validation of each data resource participation and use or access. Auditing is performed within the distributed ledger smart contract code, which validates and records resource disclosure, accesses requests to the smart contract, and performs the checks necessary to determine whether an action is permitted (e.g., for creating resource and participation objects), and can appear through the GraphQL API that provides the functionality described above.
[0095]
[0097] The AuditIDS implementation can use the GraphQL API exposed by the Fabric client lambda function. This permits invocation of chain code operations according to permitted actions, as described above with reference to Table 3. GraphQL operations can be directly mapped to the chain code by the lambda function and can also permit queries of the state machine database state. This enables AuditIDS and DDU to observe changes in the resource state information.
[0096]
[0098] The AuditIDS code enables asynchronous observation of changes in the resource table of the data plane connector database (connectordb) and can update the chain code accordingly. For example, each time a change occurs in the resource table of an IDS connector instance, the associated AuditIDS for that connector executes an update of the corresponding audit operation via the GraphQL API. For example, when a newly provided resource entry appears in the IDS connector, the corresponding GraphQL is used to update the audit chain using the "create resource object" operation.
[0097]
[0099] Cognito IdP can be used as an identity provider to manage access to the managed API. The IdP stores users within user pools / groups and their corresponding permissions to perform actions in the audit fabric. Thus, each domain (fabric member) can control what actions a user is permitted to perform within their member domain, and for this reason, for the AuditIDS to perform the required actions, it must use a role that has the permissions for those actions. DDU requires special permissions to enable reputation updates. The IdP issues a JWT access token that permits the corresponding GraphQL API actions for the fabric client.
[0098]
[0100] From this, the embodiment of FIG. 8 can be used to implement the methods described above with reference to FIGS. 1-7.
[0099]
[0101] Although certain arrangements have been described, they are presented by way of example only and are not intended to limit the scope of protection. The inventive concepts described herein can be implemented in a variety of other arrangements. Additionally, various additions, omissions, substitutions, and changes can be made to the arrangements described herein without departing from the scope of the invention as defined by the following claims.
Claims
1. A computer-implemented method for validating data transfer from a data provider to a data user via a first computing node of a network comprising a plurality of computing nodes, wherein the plurality of computing nodes communicate with a distributed ledger storing, for each computing node, respective reputation scores derived based on a plurality of previous data transfers associated with that node, the method comprising: obtaining, by the first computing node, a data transfer record specifying the data transfer; submitting, by the first computing node, the data transfer record to a smart contract on the distributed ledger to validate the data transfer record by determining that the reputation score of the first computing node exceeds a predefined threshold; A method comprising the steps of:
2. The distributed ledger further stores, for each computing node, respective data logs specifying the plurality of previous data transfers associated with that node, and the method further comprises: sending, by the first computing node, a request to the smart contract on the distributed ledger to append the validated data transfer record to the data log associated with the first computing node. The method according to claim 1.
3. As an initial step: accessing the distributed ledger using a dummy user function and examining the data log of the first computing node; determining, by the dummy user function and based on the examined data log, whether to update the reputation score of the first computing node stored on the distributed ledger; invoking, by the dummy user function, the smart contract based on a determination to update the reputation score of the first computing node and updating the reputation score of the first computing node based on the examined data log. The method according to claim 2. A method further comprising the steps of:
4. The data log further specifies the status of data items involved in the plurality of previous data transfers, and determining whether to update the reputation score of the first computing node involves judging, by the dummy user function, whether the status of the data item indicates a request by the data provider to cancel access to the data item; in response to determining that the status of the data item indicates the request, determining to update the reputation score and decreasing the reputation score; The method according to claim 3, comprising the above.
5. The data log further specifies a plurality of data items involved in the plurality of previous data transfers, and determining whether to update the reputation score of the first computing node involves judging, by the dummy user function, whether the plurality of data items meet a similarity criterion; in response to determining that the plurality of data items meet the similarity criterion, determining to update the reputation score and decreasing the reputation score; The method according to claim 3, comprising the above.
6. Obtaining the data transfer record comprises monitoring, by the first computing node, network traffic between the data provider and the data user via a data plane associated with the first computing node. The method according to claim 1.
7. The data transfer record contract information between the data provider and the data user; information describing a plurality of data items involved in the data transfer; information describing the relationship with a plurality of previous data transfers; information describing a data usage policy; a data transfer ID; a data transfer timestamp; source and destination IP addresses; information describing data plane performance The method according to claim 1, specifying one or more of the above.
8. A computing node comprising a processor and a memory, wherein the memory stores a plurality of instructions executable by the processor to implement the method according to claim 1.
9. A computer-readable medium comprising a plurality of executable instructions that, when executed by a processor, cause the processor to implement the method according to claim 1.
10. A computer-implemented method for validating data transfer from a data provider to a data user via a first computing node of a network comprising a plurality of computing nodes, using a distributed ledger, wherein the distributed ledger stores respective reputation scores derived based on a plurality of previous data transfers associated with each node the method comprises receiving, by the distributed ledger, a data transfer record specifying the data transfer validating, by the distributed ledger, the data transfer record using a smart contract on the distributed ledger by determining that the reputation score of the first computing node exceeds a predefined threshold A method comprising the above steps. **Claim 11** the distributed ledger further stores, for each computing node, a respective data log specifying the plurality of previous data transfers associated with each node, and the method comprises receiving, by the distributed ledger, a request from the first computing node to add the validated data transfer record to the data log associated with the first computing node adding, by using the smart contract on the distributed ledger, the validated data transfer record to the data log associated with the first computing node The method according to claim 10, further comprising the above steps. **Claim 12** As an initial step providing, by the distributed ledger, access to the data log of the first computing node to a dummy user function updating, in response to receiving a request from the dummy user function, the reputation score of the first computing node based on the investigated data log The method according to claim 11, further comprising the above steps. **Claim 13** The data log further specifies a state for data items associated with the plurality of previous data transfers, and the state of the data items indicates a request by the data provider to revoke access to the data items. Updating the reputation score of the first computing node based on the investigated data log comprises updating the reputation score to decrease the reputation score, the method according to claim 12.
14. The data log further specifies a plurality of data items associated with the plurality of previous data transfers, the plurality of data items meet a similarity criterion, and updating the reputation score of the first computing node based on the investigated data log comprises updating the reputation score to decrease the reputation score, the method according to claim 12.
15. The data transfer record is generated by the first computing node by monitoring network traffic between the data provider and the data user via a data plane associated with the first computing node, the method according to claim 10.
16. The data transfer record contract information between the data provider and the data user, information describing a plurality of data items involved in the data transfer, information describing the relationship with a plurality of previous data transfers, information describing a data usage policy, a data transfer ID, a data transfer timestamp, source and destination IP addresses, information describing data plane performance specifies one or more of, the method according to claim 10.
17. A computing node comprising a processor and a memory, the memory storing a plurality of instructions executable by the processor to implement the method according to claim 10.
18. A computer-readable medium comprising a plurality of executable instructions that, when executed by a processor, cause the processor to implement the method according to claim 10.
Citation Information
Patent Citations
Service data processing method and device, and service processing method and device
JP2020510331A
Blockchain-based service execution method and apparatus, and electronic device
JP2020516968A
Preventing Inappropriate Data Transfers Based on Reputation Scores
US20120291087A1