Method for processing data in a distributed network
The method addresses the lack of user control in data processing by implementing a four-level data model and access management within a distributed network, ensuring secure and standardized data access and storage while protecting data sovereignty.
Patent Information
- Application Number
- PCT/EP2024/084316
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-05
AI Technical Summary
Current data processing methods in distributed networks lack user control over data storage and access, leading to vulnerabilities such as unauthorized access and loss of data sovereignty, particularly highlighted by incidents like the Cambridge Analytica scandal.
A method for processing data in a distributed network utilizing a four-level data model, device-independent address reference standard, mandatory contract information linked to data records, and a control instance for managing access, ensuring standardized data storage and access across devices while enforcing subscriber-dependent access rights and restrictive conditions.
This approach enables users to maintain control over their data, ensures secure and standardized data access and storage, and prevents unauthorized access, thereby protecting data sovereignty and reducing the risk of data misuse.
Smart Images

Figure EP2024084316_05062025_PF_FP_ABST
Abstract
Description
[0001] Method for processing data in a distributed network
[0002] A method for processing data in a distributed network, comprising: a four-level data model; a device-independent address reference standard;
[0003] Contract information which is mandatorily linked to a data record; and a control instance for checking and releasing access to the data records, wherein at least the data records are standardised in such a way that they can be stored and read regardless of the device, and wherein the contract information contains subscriber-dependent access rights, wherein the following steps are carried out in response to an access request by the control instance: a. comparing the identity with the associated contract information and determining the subscriber's access authorisation; b. determining at least one restrictive access condition and checking whether the at least one determined restrictive access condition can be fulfilled; only if access is authorized and the access condition can be fulfilled: c. releasing access to the specific data record.
[0004] While the initial development and spread of the internet was dominated by the idea that individual (private or public) computers could share and send the data stored on them, regardless of location, the current approach is almost exclusively for large server centers to host the data, while private users have little or no data stored on an external server without a backup copy. Such server centers are generally privately operated. Often, the storage space offered is not only a paid service, but also requires private data. Thus, the individual user relinquishes control over their data and must trust the server service provider.
[0005] However, the Cambridge Analytica scandal has demonstrated that the agglomeration of personal data can have an impact on society as a whole, and that personal data is a resource worthy of protection, providing holistic insights into our political and sexual orientation, medically relevant information, and our personalities. Individuals and companies should be able to decide independently who receives this information.
[0006] A conventional solution with a dedicated server is only a pseudo-solution for private individuals and companies in the face of massive hacker attacks, and also causes significant maintenance costs. An approach (as described, for example, in the application)
[0007] US 2023 / 0 005 391 A1) is that the stored data is stored in fragments and thus a hacker attack can only capture a small amount of useless data.
[0008] Based on this, the present invention is based on the object of at least partially overcoming the disadvantages known from the prior art. The features of the invention are derived from the independent claims, for which advantageous embodiments are presented in the dependent claims. The features of the claims can be combined in any technically reasonable manner, whereby the explanations from the following description as well as features from the figures, which comprise additional embodiments of the invention, can also be consulted for this purpose.
[0009] The invention relates to a method for processing data in a distributed network, comprising at least the following components: a data model with four levels, of which the most basic level is the data level, and the higher-level levels define the data structures in the respective subordinate levels in an increasingly abstract manner; a device-independent address reference standard, with which a specific data record in the data level is addressed with an associated address reference;
[0010] Contract information which is mandatorily linked to a specific data record; and a control authority for checking and approving access to the data records, wherein at least the data records are standardised using the three subordinate levels in such a way that they can be stored and read regardless of the device, and wherein the contract information contains subscriber-dependent access rights to the associated specific data record, wherein the method comprises at least the following steps carried out by the control authority: upon an access request from a subscriber with a computer-readable identity to a specific data record via an associated address reference: a. comparing the identity of the subscriber with the associated contract information and determining an access authorisation of the requesting subscriber; b.Determining at least one restrictive access condition according to the associated contract information for the requesting subscriber and checking whether the at least one determined restrictive access condition can be fulfilled; and only if the requesting subscriber has access authorization and if the at least one access condition can be fulfilled: c. Releasing access within the scope of the at least one access condition to the requested specific data set, whereby without such releasing of access, a subscriber is prevented from accessing the data set in question by the control authority.
[0011] Unless explicitly stated otherwise, ordinal numbers used in the preceding and following descriptions serve only to clearly distinguish them and do not reflect the order or ranking of the designated components. An ordinal number greater than one does not necessarily imply that another such component must be present.
[0012] The procedure proposed here gives participants (natural and legal persons, as well as machines) the opportunity to retain control over their data within manageable time and cost-effective means. The key aspect here is the creation of a standardization that enables reading and saving (and all other types of access) of data records regardless of the device. At the same time, the standardization is not a rigid framework into which the data must fit. Rather, when creating or editing data records, either an existing standard is applied or, if no suitable one exists, a new one is created. However, the new standard fulfills the minimum requirements within the data model, allowing third parties to read and write (unless blocked).
[0013] The process also ensures that no unauthorized access to the participant's data records occurs. The participant thus has full control over their data records. If a third party gains access to the participant's data records without the participant's knowledge and consent, this third party can be proven to have obtained these records illegally and is therefore legally liable. This may not be sufficient protection for a hacker who is difficult to identify, but the stolen data records are thus unusable for potential buyers. The key requirement for processing data is that it be in a structured form. In one embodiment, it is possible to initially store data in an unstructured form and then subsequently structure it, for example, only when needed, in order to process it.Currently, many systems, and especially websites, contain a large amount of unstructured data, making machine processing difficult. The concept of the Semantic Web was developed based on a proposal by Tim Berners-Lee, the founder of the modern internet. The concept brings together various standards that enable self-description of data. The standards concern: a data model; data identifiers; the definition of vocabularies for describing data models (or ontologies), and serialization of the data model. In principle, the concept of the Semantic Web is compatible with various standards and is already being used in multiple contexts.
[0014] Standards based on the Semantic Web or comparable concepts (such as Web3) have in common that they are based on a meta-model-based data model. The meta-metamodel-based approach discussed below is advantageous for the present method because all abstracted meta-models used within it follow the same set of rules. The use of a meta-metamodel is not mandatory, but such a set of rules would have to be established for each other data model. The Resource Description Framework [RDF] is most commonly used. Other formats used in Semantic Web applications include JSON-LD, JSON Schema, and Activity Streams.
[0015] The existing standards have not resulted in the implementation of the Semantic Web. Furthermore, the idea of the Semantic Web does not aim to contribute to the attainment of data sovereignty by natural and legal persons. Therefore, this procedure is based on a data model that incorporates and extends the idea of the Semantic Web. The following section describes the implementation of the procedure proposed here based on the Essential Meta Object Family [EMOF] standard of the Object Management Group. However, the procedure proposed here can be implemented with other data models such as JSON-LD, Resource Description Framework, or relational databases. The use of the Essential Meta Object Facility [EMOF] of the Object Management Group is particularly recommended. EMOF is one of two possible implementation variants of the Meta Object Facility in version 2.5.1 [MOF 2].
[0016] Fundamentally, data models contain a syntax. However, the statements made in a data model require interpretation. A data model achieves interpretability through the use of a common vocabulary. If the vocabulary also contains rules that govern the correct use of resources defined by the vocabulary, the vocabulary is referred to as an ontology. Ontologies typically have the following components: classes, types, instances, relations, inheritance, and axioms.
[0017] Classes describe the description of common properties. Classes can be arranged hierarchically in the form of superclasses and subclasses. Types describe the representation of object types, which are represented in classes. Instances are representations of objects. Relations describe the relationships between instances and are also called properties. Relations of classes can also be inherited, so that the relations are passed on to the inheriting element. Multiple inheritance refers to the inheritance of more than one class. Axioms describe ubiquitously true statements. Axioms are usually used to express knowledge that cannot be derived from classes. Concrete serializations of data models are also heavily dependent on the latter.Serialization is the process of translating a data structure or object state into a storable or transferable format, allowing for later reconstruction of the stored or transferred data. These definitions can also be found in other data models, such as the one mentioned above, in addition to EMOF.
[0018] Data and data types are created, or can be created, in a standardized data model. All data is structured based on this data model, allowing automated storage and processing without the need for reprocessing. This makes it possible to reference the same data in different applications and then reuse it immediately, because its structure is understandable for both applications.
[0019] At the most basic level are concrete data, i.e., objects of reality. At the next higher (abstraction) level are participant models (domain models), which define the data of the most basic (i.e., concrete) level. At the next higher (second abstraction) level are meta-models. A meta-model is a model that defines a model. It therefore describes the structure of the model of concrete data. The higher (third and final abstraction) level contains a meta-metamodel. This models how meta-models are structured. It is important that the higher level represents the end of the meta-ization, because otherwise it could continue indefinitely. This is achieved by the meta-metamodel also defining itself.
[0020] The advantages of EMOF over other data model standards are that it uses particularly easy-to-understand rules for modeling metadata, which are fundamentally based on UML class modeling. The Unified Modeling Language [UML] is considered the industry standard for the specification, construction, documentation, and visualization of software, its components, and other systems and is standardized according to ISO 19505. Furthermore, EMOF is platform-independent and allows various technological mappings based on UML models and UML profiles. It also supports a wide variety of tools that allow modeling at the meta-level. All tools used for UML modeling can be used here. The advantage of relying on the UML standard in the aforementioned contexts is based on the widespread use of UML, because this modeling language corresponds to an industry standard.This allows a high degree of connectivity to previous metamodels of software infrastructure.
[0021] An address reference standard defines how a specific data record can be found using an associated address reference. For example, in a Semantic Web implementation, the address reference is widely used, integrated as a URI (Uniform Resource Identifier). URIs are based on the Internet Protocol standard Internationalized Resource Identifier (IRI). Each URI serves as a unique identifier for one and the same resource. The reverse is not necessarily true: A resource can have multiple URIs. The URI can also be used as an address for retrieving further data about the referenced resource. URIs can be used in various forms, for example as the most widely used URL (Uniform Resource Locator). Depending on the data model used, different URI schemes are available, in accordance with the schemes officially registered with the IANA (Internet Assigned Numbers Authority).
[0022] Contract information is information that regulates access to a specific data set and access rights. Multiple contracts with different contract information can exist for each data set, defining access to the data set by one or more participants. This enables various forms of collaboration. The contract information includes, for example, which (types of) participants or companies are permitted access, what rights a participant with access rights has, and how long or how often a participant with access rights can be granted access.For example, an insurer should only be able to read a data set containing the personal data of a person applying for insurance, only within a specific time window and / or only once, or alternatively only until the regular end of an insurance term.
[0023] It should be noted that the contract information contains access rights, which can be divided into access authorizations and access conditions. Access authorizations include the type of access, such as reading, writing, copying, moving, deleting, and other actions. Access conditions concern the accessing participant, the duration, the number of repetitions, a time window, and other conditions.
[0024] The control authority is either independent of the participant or provided by a service provider or a public institution. This control authority can be understood as a doorman. It should be noted that the control authority itself (preferably) does not, or at least not necessarily, have access to the data records. In one embodiment, the control authority's access authorization is regulated in individual contractual information; in one embodiment, access is generally denied to a control authority; in another embodiment, the control authority is granted access (e.g., for automated execution).
[0025] It should be noted that it may sometimes be possible to gain access to a data set indirectly or by chance without the control authority's review and approval. However, even then, it is ensured that the participant in question cannot provide proof of authorization to access such a data set. For this purpose, the contract information is preferably stored separately in such a way that only the owner of the data set and the control authority (if necessary, subject to the need for approval from an owner or an authorized participant who is not available to the control authority) can be brought together. Alternatively or additionally, write permissions for contract information generally do not exist for anyone other than the owner, or the owner's active consent to any changes to the contract information is required.Alternatively or additionally, the control authority is configured to record access to a protected data set and / or immediately trigger a notification to the owner. In one embodiment, access is possible for a relevant sovereign authority to protect the public, whereby the owner is necessarily informed and / or an authorizing order (e.g., a judicial search warrant) must be presented to the control authority. In one embodiment, this is necessarily and easily understandable, automatically included in the contract information and cannot be changed by the owner. This creates legal certainty, similar to what we might already expect in the physical world.
[0026] It should be noted that not every data set necessarily contains such data or has to have a type of content that must be protected from being read or copied. For example, goods and / or services offered online are completely freely readable, but they should not be modifiable (by everyone) and / or disseminated without conditions. For example, access to an article in a digital newspaper magazine is conditionally tied to the payment of a fee. For example, access is tied to the disclosure of information, such as the device used, gender, address, or profession. On the other hand, there is a group of participants who should have free access, such as the owner's employees. This may be person-specific, location-specific, and / or dependent on an access token.For example, access rights are limited to the duration of a project or a temporary employment contract.
[0027] It is further proposed in an advantageous embodiment of the method that the data records of the data level have a classification and the data content contains at least one attribute, wherein in the case of a plurality of attributes at least one of the attributes, preferably of an attribute-value pair, limits, preferably defines, the data form and / or data quantity of the at least one other attribute.
[0028] Within the meta-metamodel, an internal classification of data records in the data layer is useful because this makes the data machine-processable (see the above comments on the Semantic Web). Preferably, two variants of classification exist to ensure machine-processing of the data. Within this preferred embodiment, each instance has a type. In the context of classification, this refers to instances that are either of the class type or of the attribute type (see the comments on the data model) and are combined with the primitive types (see, for example, primitive types in EMOF). Instances of the class type contain information that describe the attributes of the attribute-value pairs of instances of these classes. Instances of the attribute type describe the valid values of each attribute-value pair in the instances of the class in which they are defined.Using this information, the Class type can describe itself together with the Attribute type.
[0029] The functionality of personal data is described as an example:
[0030] If a person record is to be stored in an instance, its type could belong to a class "Person." The record could also contain an attribute-value pair<Nachname: "Müller"> include.
[0031] If a data set has multiple attributes, the data content, within a preferred embodiment, contains exactly one attribute-value pair, which restricts or, in an advantageous embodiment, defines the data form and / or the data volume. The attribute-value pair thus describes the type of the data set classified in this way.
[0032] In a preferred embodiment, those data records which were not created within the model and do not correspond to the model-internal classification are assigned to a fallback classification and then transferred to the model-internal classification.
[0033] In a further advantageous embodiment of the method, it is proposed that the data structures have a classification and the data content contains at least one attribute, wherein an existing suitable classification is used when creating a data record and / or a data structure or, if a suitable classification does not exist, a suitable classification is generated. This embodiment achieves immediate machine readability when new data is created. At the same time, it allows the necessary flexibility to map human reality at the data level (and higher levels). Experience shows that even with the most careful consideration, not all variations and possible combinations can be foreseen.However, when new data is created, it is directly transferred into an (at least partially) existing structure and, as a rule, new data only contains some new properties.
[0034] In a preferred embodiment, a newly created classification will be used in the future for other applications, preferably for other, particularly preferably for all, participants.
[0035] It is further proposed in an advantageous embodiment of the method that the data records and / or the address references are individually encrypted, wherein preferably the control instance does not have a suitable key for this.
[0036] Any attributes and values discussed in the previous sections can optionally be encrypted at the type level, attribute level, and value level (as well as any subset thereof). This means that the type URIs used by an instance, the attribute URIs, and / or the values of an attribute-value pair are unreadable if the corresponding decryption key is unknown.
[0037] To maximize the data sovereignty of its participants, a type of encryption is provided; only with encryption is a data set hosted on a server no longer automatically available for full access from the server side. It should be noted that locally hosted data sets can also be encrypted, or that this data sovereignty is already ensured in unencrypted form due to an appropriate query structure with an internal or external control authority. Encryption can also be used to prevent third parties who gain unauthorized access to the server from also gaining immediate full access to the data stored there. The implementation of this aspect of the process is described below as an example based on the OpenPGP standard according to RFC4880.This standard specifies, among other things, formats and concrete procedures for hybrid encryption, cryptographic identities, and digital signatures (including the mapping of trust relationships between identities and / or key material). For message encryption and group encryption, among other things, the Messaging Layer Security [MLS] standard can be used as an alternative to OpenPGP.
[0038] Hybrid encryption combines techniques from symmetric and asymmetric cryptography. Unlike a digital signature method, each recipient has a key pair, the public key of which is made available to the sender. In the simplest case, the sender uses this public key to encrypt a newly generated symmetric key. The latter is used to encrypt the actual payload (for example, using AES [Advanced Encryption Standard]). The recipient decrypts the symmetric key using their private key and can then decrypt the payload. The main reasons for using hybrid encryption methods are the need for a simple and efficient key agreement for the symmetric method. Furthermore, the use of purely asymmetric methods for encrypting large amounts of data is impractical due to their inefficiency.
[0039] When using asymmetric cryptography techniques, the authenticity of the transmitted public keys must be protected to prevent man-in-the-middle attacks. An attacker could generate their own key pair, interpose themselves between the honest parties, and exchange public keys with each of them. Intercepted messages are then decrypted, intercepted, and, if necessary, modified, forwarded to the originally intended recipient. Possible solutions include the direct exchange of keys and / or certificates between the parties, a public key infrastructure including certificate authorities, or web-of-trust-based approaches. In one embodiment, the desired level of trust in the participants active within a web of trust should be known, i.e., machine-readable.On the other hand, this information should not be freely and / or always readable, i.e., interpretable, by embedding the participant's identity in the missing key. Thus, the identity can only be read together with the existing key, i.e., only in the situation where a verification must be performed, and furthermore, only by the controlling authority, and not (necessarily or by default) by the requesting participant.
[0040] A digital signature is a concept from asymmetric cryptography used to protect the authenticity of exchanged data. A sender generates a key pair consisting of a private and a public key. The private key must be kept secret. In contrast, the public key is made accessible to every recipient in a way that preserves authenticity. A digital signature created using the private key on a message can then be verified using the public key alone.
[0041] It is further proposed in an advantageous embodiment of the method that the associated contract information is encrypted, wherein preferably: the control authority has at least one less of a plurality of keys by means of which a key group can be formed together, wherein the relevant contract information can be made readable for the control authority exclusively by means of the formed key group, and with the access request of a subscriber the at least one missing key is made usable for forming the appropriate key group for the control authority, wherein further preferably the at least one missing key contains the computer-readable identity of the subscriber or is contained in this identity in a standardized manner that can be read by the control authority.
[0042] In an advantageous embodiment, the key group is formed by exactly two keys, and the control entity owns one of the two keys. The other key is provided upon an access request, either directly by the requesting participant or via a request initiated by the requesting participant from the owner of the relevant contract information.
[0043] In one embodiment, the trust desired for a Web of Trust in the participants who are active there should be known, i.e. machine-readable.
[0044] On the other hand, this information should not be freely and / or always readable, i.e., interpretable, by embedding the participant's identity in the missing key. Thus, the identity can only be read together with the existing key, i.e., only in the situation where a verification must be performed, and furthermore, only by the controlling authority, and not (necessarily or by default) by the requesting participant.
[0045] It is further proposed in an advantageous embodiment of the method that at least one of the following components is signed by means of the identity of a respective participant or owner of the data set: the contract information; the access request; the address references; and the data contents of a data set.
[0046] The goal of this implementation is to ensure the authenticity of data records and to restrict access to them through contracts. Protecting the authenticity of data records can be achieved through a digital signature of the participant writing the data. While a control authority can be entrusted with the enforcement of a contract, the authenticity of the contract information itself must also be maintained. The latter can be ensured through a digital signature of an authorized participant on the contract information.
[0047] A participant's trust in the authenticity of data records and contract information depends directly on the trust the participant places in the key used for the signature or in a participant's identity cryptographically associated with this key. This trust can be established using well-known techniques such as the direct exchange of keys / certificates between the parties, a public key infrastructure including certificate authorities, or web-of-trust-based approaches. The approach outlined in the last sections can, for example, be implemented using the means described in the OpenPGP standard.
[0048] In one embodiment, only identities certified by an external certification authority are permitted.
[0049] It is further proposed in an advantageous embodiment of the method that the contract information contains at least one of the following access conditions: current location of the subscriber or origin of the access request;
[0050] Number of accesses, preferably per requesting identity;
[0051] Period for access authorization, preferably dependent on the requesting identity; and
[0052] Scope of use of the data set.
[0053] The scope of use refers to the rights and obligations of the participant and the data storage provider. A distinction must be made between legal and technical requirements. An example of a legal requirement is a specific retention period. An example of a technical requirement is that the data is accessed via a TLS (Transport Layer Security) connection.
[0054] It is further proposed in an advantageous embodiment of the method that the owner of the data set or another authorized participant must grant approval as the first factor or second factor for step c.
[0055] In this embodiment, the owner of the requested data set receives a notification (e.g., a push message with an authentication request). Access is only possible once the owner or another authorized participant has authorized the access, provided that any other existing access conditions are met (e.g., location of the request). In one embodiment, the owner or the other authorized participant in question is a company that, in addition to an external control authority, wishes to at least register access. This is advantageous, for example, in the event of an error or data leak in order to potentially gain insights into the registered accesses. It should be noted that in one embodiment, the owner himself operates the control authority, and only preferably the control authority is operated by an independent entity.
[0056] It is further proposed in an advantageous embodiment of the method that when checking the identity, the access request and / or the data set by means of the control authority, a trustworthiness level is also checked, wherein preferably an external confirmation is obtained and / or a certificate created by an external certification authority is used.
[0057] For example, a trustworthiness level for a participant is an (explicit or implicit) assessment by (at least one) other participant(s), or for a natural person a so-called certificate of good conduct or comparable certificates from the real world.
[0058] In an advantageous embodiment, a digital certificate is obtained, which is not necessarily used in the real world. Rather, an external certification authority, preferably one independent of the controlling authority, is used, for example, by collecting data and / or through a one-time (or repeated) registration of a participant. This achieves a very high degree of web of trust. It should be noted that such a query is preferably an access condition and does not necessarily automatically apply to all types of data sets, for example, not to data sets for which the most barrier-free distribution possible is desired, such as advertising campaigns.
[0059] In a further advantageous embodiment of the method, it is proposed that the address reference used in the present access request be unreadable for a participant and readable for the control authority. The approach is comparable to a QR code, in which a (basically arbitrary) number of pieces of information can be read by humans only with considerable effort, but can be read by a machine in a fraction of a second. In contrast, an address reference is (preferably cryptographically) encrypted and cannot be read by anyone, not even with the help of a machine (for example, in the case of a QR code, using the camera of a smartphone). Rather, the address reference is, for example, incomplete and / or altered in a (seemingly) random manner.Only with the key or a fixed (preferably complicated) algorithm of the control authority is the address reference made readable, and preferably exclusively for the control authority itself.
[0060] This makes it possible for a representative of an address reference to be openly distributed, but the actual address reference can only be processed with the assistance of the control authority. This makes it possible, for example, to forgo further encryption of data records or to choose a relatively simple level of encryption, which would otherwise no longer be considered sufficiently secure on its own, but which is sufficient when combined with the unreadable address reference. This makes it possible to reduce the amount of data and / or increase processing speed.
[0061] The invention described above is explained in detail below against the relevant technical background with reference to the accompanying drawings, which show preferred embodiments. The invention is in no way limited by the purely schematic drawings, whereby it should be noted that the drawings are not to scale and are not suitable for defining proportions. It is shown in
[0062] Fig. 1 : a pyramid of a meta-metamodel;
[0063] Fig. 2: a sequence diagram of an encryption;
[0064] Fig. 3: A diagram of data processing in a multi-provider environment; Fig. 4: A sequence diagram of a participant registration process; Fig. 5: A sequence diagram of creating a resource; Fig. 6: A sequence diagram of adding an instance to a resource; Fig. 7: An overview of a sequence diagram of granting access to data based on a contract;
[0065] Fig. 7a: Part 1 of the sequence diagram according to Fig. 7;
[0066] Fig. 7b: Part 2 of the sequence diagram according to Fig. 7; and
[0067] Fig. 8: a diagram with the third and fourth levels of a practical implementation of a metamodel according to Fig. 1.
[0068] Figure 1 shows a pyramid of a meta-metamodel. The meta-metamodel has four different levels (or layers). At the first data level (most basic level), the meta-metamodel includes so-called real-world instances of objects. These objects are relevant within a data set but lie outside the application itself. These instances can be anything, for example, a real person, an object, or a virtual asset. Figure 1 shows three different people as examples.
[0069] On a second level above, concrete domain-specific models can be created. Within these models, data content, which is either abstracted from the first level and / or created independently, is entered into the system. Domain-specific models describe the form of data sets. Data sets always have a specific data structure defined by a domain-specific model, and data sets contain data content.
[0070] Each domain-specific model is an instantiation of a metamodel of the next, third, level. Domain-specific models can be created using any valid language (valid means that the language can be used to create domain-specific models). As an example, Fig.
[0071] 1 a domain model (consisting of the attributes FirstName, LastName and Birthday as well as the class Customer) and several records, some of which have relationships to the persons of the first level.
[0072] At the third level, metamodels create structures for defining concrete domain-specific models and their data. Each metamodel is a direct model of the language used at the second level. At the same time, each metamodel is an indirect model of a second-level domain-specific model. Each metamodel is an instantiation of metamodels at the fourth (and final) level.
[0073] Metamodels can be created using any valid language (valid means that the language can be translated into and from the meta-metamodel and used to define a language for creating domain-specific models). As an example, Figure 1 contains structures that can be used to define a domain model (e.g., the definition of an attribute containing a class name).
[0074] At the fourth level, we reach the highest level of abstraction: the meta-metamodel. The meta-metamodel defines generic building blocks, which in turn can be used to create metamodels. Each generic building block defines itself. The meta-metamodel is a direct model of the language used to create metamodels at the third level. At the same time, the meta-metamodel is an indirect model for each metamodel created with the metamodeling language. A meta-metamodel can be created with any valid language (valid means that the language is capable of defining its components and can be used to define a language for creating metamodels). As an example, the figure contains three generic building blocks that are used in most data models: data type, attribute, and class.
[0075] Fig. 2 shows a sequence diagram of an encryption algorithm. This example does not cover every possible use case. The example also omits common process steps to make the diagram readable. The server can be hosted on the client itself. The steps shown may also occur in a different order or in a combined format, for example, to improve performance. This also applies to the sequence diagrams and diagrams shown and explained below in Fig. 1 to Fig. 8.
[0076] The figure illustrates the addition of new data to an existing record for which a contract already exists. To enter the data entered by a participant, it must be encrypted, signed, and transferred to the server.
[0077] First, the appropriate encryption keys must be made available to the participant. This is necessary because different parts of a data set may need to be encrypted with different keys, and these keys can also change during the lifetime of a data set.
[0078] Information about the encryption keys is contained in the contract information associated with the data set. For example, a contract could contain the encryption keys in encrypted form. Alternatively, the contract could also contain a reference to a key that the participant has obtained by other means. The appropriate keys and, optionally, the contract information are downloaded from the server after the control authority has verified the participant's access rights to the information. Similarly, the participant verifies the authenticity and consistency of the contract information received from the server.
[0079] In addition to encrypting the data, the participant also signs it. The private parts of the signature keys used always remain with the participant. The signed data structure sent to the server contains not only the encrypted data but also server-readable metadata. This metadata is used by the server to verify the legitimacy of the participant's changes to the data set and, if successful, is entered into the database.
[0080] Figure 3 shows a diagram of data processing, specifically data storage, in a multi-provider environment. The first part of the figure explains a domain-specific model of classes, attributes, and instances of classes and attributes. The elements in the middle of Figure 3 explain the metadata types of the given types and are derived from a metamodel. In this example, we have only two metadata types: class and attribute. There are four domain-specific classes, named Person, Address, Student, and Institute. Each solid arrow means that the class or attribute of the domain-specific model is an instance of the metadata type. Each class has at least one attribute.Arrows (dashed with large gaps), marked "extends" in the legend, indicate that the extending class inherits the attributes or extends the form of the extended class. The attributes in Fig. 3 can be represented in the following variants:
[0081] In the first variant, attributes are primitive (they define a single or multiple, technically speaking, atomic value(s)), and thus the values in the attribute-value pair are implicitly contained (for example, Name is a string, a postal code, an integer, or similar). This applies to the elements referred to as attributes (here: Name (twice), Birthday, City, Postcode, Street, and StudentNo).
[0082] Attributes can also reference instances of classes, meaning that the value of the attribute-value pair defined by the attribute is not primitive. For example, a Person can have an attribute named Address. The class referenced by the attribute consists of two or more attribute-value pairs. The address in the figure has three attribute-value pairs (here: City, Postcode, Street). The arrows labeled "references" describe the relationships between the classes and are also considered attributes. The arrows labeled "contains" describe the relationships between instances and classes and are also considered attributes. However, instances of the target class are hard-linked to instances of the source class, or rather, linked instances live within them. In our example, each institute has an Address. Each institute has several Students, while each Student has an Institute.A person can have multiple friends.
[0083] The second part of Figure 3 explains the distribution of data within a distributed network based on the domain-specific model mentioned above. In this example, there are exactly two providers.
[0084] Provider A hosts data from the Institute TU Dortmund. As we can see, the type and name are contained in the same instance. For the address, we see an id reference to another instance. In this case, the address is stored with the same provider. Provider B hosts the data for student Horst. Horst belongs to an institute, which is represented by the Institute attribute. In this attribute, you can find the reference to the ID of the institute of the TU Dortmund. As we already know, the data for the Institute TU Dortmund is stored with a different provider. The Horst instance also contains an Address attribute, which refers to an instance of the Address type. This instance is also stored with Provider A, and even in the Student Horst instance because the corresponding attribute definition defines a "containment" reference. It can be seen that data can be located on multiple providers simultaneously.This even applies to different data within a dataset or even within a single instance.
[0085] Fig. 4 shows a sequence diagram of a participant's registration process. This shows a sequence diagram of the registration process within a system designed according to the method proposed here. In this example, the participant registers for the first time. They use a participant ID and a password. This data is entered on a client. The client then generates a public key / private key pair for the participant. The private key is then encrypted with the password entered by the participant. This now encrypted private key and the public key are now both ready for use. The client makes both available to the participant and publishes the public key on a server. The registration process could now be complete, depending on the implementation.Providers are free to introduce and / or use additional verification procedures.
[0086] Fig. 5 shows a sequence diagram of creating a resource R, designed according to the method proposed here. In this example, a participant has already opened a client and logged into their account. They then create a resource. The client then retrieves the participant's private key. The client then reserves an arbitrary, unassigned resource ID on the server for the participant's private key. The client then creates and signs a contract with this private key. The client then sends a request to the server to create a resource with the signed contract.
[0087] The server identifies the public key for the user ID, which is stored on a server (this server can be a second server). The signature is then verified using the public key. In our example, the verification process is successful, and the server proceeds to create the resource and the signed contract. The server informs the client of the successful completion of the process, which in turn informs the participant. The resource creation process is thus complete.
[0088] Fig. 6 shows a sequence diagram of adding an instance to a resource, which contains the following precondition as indicated at the beginning of the sequential flow shown: For this example, a resource already exists and is stored on the server.
[0089] The participant creates a new instance T, for example of a class. In this figure, the instance is not specified because it can be any instance already defined in the data set. The participant selects a classifier C for the newly created instance T. The participant enters this information into the client. The client is then able to create a new empty instance T and classifies the new instance T. The participant can continue editing their instance T. They can now set values. Because they have selected a specific type, the participant must set the values according to the type definition. When the participant is finished, the participant tells their client to save the data they have just created.
[0090] To store the data, the client serializes the instance T. In our preferred embodiment, the client also signs the instance with the participant's private key. At this point, the client has a signed instance T. The client attempts to store the signed instance T in the already created resource R. The server extracts the fingerprint from the signature of the signed instance T. The server then retrieves the public key PUB associated with the fingerprint. After the server has successfully verified the integrity and authority of the signed instance T using the public key PUB (as in our example), it persists the signed instance T in the resource.The server informs the client of the successful save, which in turn informs the participant. The process of adding an instance to a resource is thus complete.
[0091] Fig. 7 (or Fig. 7a and Fig. 7b) shows a sequence diagram of granting access to data based on a contract. Here, the following case is discussed: A contract may contain a condition that must be granted on demand by a privileged entity, such as Owner 0. For this example, there are several preconditions:
[0092] First, the participant P has a uniform resource identifier for data.
[0093] Second, the privileged entity (in our example, usually the owner 0 of a record) has the necessary permissions to grant the participant access, at least upon request.
[0094] In a first step, participant P attempts to access data D. To do so, they open the data D using an address reference (here: data-uri D). Participant P's client then attempts to load the data D from the server (derived or resolved from the address reference D) using a signed request (SigREQ). The server retrieves the contract (CTR) that refers to the existing address reference D. It checks the authorization of participant P based on this contract (CTR).
[0095] In our example, the CTR contract states that participant P can access the data set if owner 0 grants them access to data D. The server will request this permission from owner 0. It also retrieves participant P's public key and adds it to the SigREQ. The request to owner 0 now contains the signed original SigREQ request to the server, the address reference D, the present CTR contract, and the identity of the requesting participant P. The request is now ready for delivery, so the server makes it available to owner 0.
[0096] In our example, the server now waits for the decision of owner 0. A practical use case could look similar to our example: The server executes a loop in which it waits for the answer from owner 0 while the request is not answered.
[0097] Once the server receives a response, it has two options:
[0098] Either it retrieves the actual data DD and prepares a response RES to the client of participant P, which contains the actual data DD and a response key ANS. key, or it sends a response RES stating that owner 0 has denied participant P access to the data D in question, i.e., of the type Access Denied.
[0099] In both cases, the server informs the client of participant P, which in turn informs participant P.
[0100] On the side of the privileged entity (for example, Owner 0), the request process might look similar to our example. First, Owner 0's client receives new requests by executing an infinite loop. Once the server has received requests, it transmits the List of Inquiries (LI) to Owner 0's client. The client, in turn, notifies Owner 0 of the receipt of new requests, which then opens the new request on its client. Owner 0 has two options: It can either grant or deny access to the requested data.
[0101] If access is granted, the process might look like this: The owner's client first checks the integrity of the request. Then, it retrieves the public key of the requesting party P. The owner 0 himself grants the requested permissions to the address reference for the signed request in the request. The server receives the response from owner 0's client and proceeds as described above. This includes the option that the data may be encrypted. If this is the case, the server re-encrypts the data for the requesting party P and includes the encryption key (EK) encrypted for the requesting party P in its response.
[0102] If the owner denies access, the client sends the response to the server, which proceeds as described above. This completes the access granting or denial process, and the participant can continue or not continue with their process.
[0103] Fig. 8 shows and explains an example of a layer diagram of a practical implementation of a meta-metamodel according to Fig. 1. As can be seen, this diagram also contains four levels (or layers).
[0104] At a first layer (Layer 1), we have a real-world object. In this case, it's a person named John Peter. The objects of the first layer are represented by the contents of the second layer (Layer 2). In this second layer, there is an instance of John Peter. This instance contains several attribute-value pairs, such as the person's full name (FullName) and their birthday (Birthday). The John Peter instance is an instance of PersonClass. PersonClass defines the attributes that this domain-specific model considers important. PersonClass also contains references to the definitions of PersonClass's attributes: FullnameAttribute and BirthdayAttribute.
[0105] All instances of the domain-specific model are instances of a Meat metamodel created at Layer 3. At the third level, there are several abstract instances, such as Class (which cannot be instantiated), Attributes, NameAttribute, TypeAttribute, and AttributesAttribute. Using these instances, it is possible to create domain-specific models, such as the model at the second level. There are also several relationships between the instances of the metamodel. For example, AttributesAttribute defines the content of the Attributes attribute in the Class instance.
[0106] The fourth and highest level defines generic building blocks. These are the generic building blocks Instance, Attribute, and Class. As can be seen, every entity in the second and third levels is (either directly or indirectly) an instance of these generic building blocks.
[0107] The procedure proposed here enables private data sovereignty with easy implementation and at the same time device independence for the digital
[0108] Exchange of data sets created.
Claims
Patent claims 1 . A method for processing data in a distributed network, comprising at least the following components: - a data model with four levels, of which the most basic level is the data level and the higher levels define the data structures in the respective subordinate levels in an increasingly abstract manner; - a device-independent address reference standard with which a specific data record in the data layer is addressed with an associated address reference; - Contract information which is mandatorily linked to a specific data set; and - a control authority for checking and approving access to the data records, wherein at least the data records are standardised using the three subordinate levels in such a way that they can be stored and read regardless of the device, and wherein the contract information contains subscriber-dependent access rights to the associated concrete data record, wherein the method comprises at least the following steps carried out by the control authority: in response to an access request from a subscriber with a computer-readable identity to a concrete data record via an associated address reference: a. comparing the identity of the subscriber with the associated contract information and determining an access authorisation of the requesting subscriber; b.Determining at least one restrictive access condition according to the associated contract information for the requesting subscriber and checking whether the at least one determined restrictive access condition can be fulfilled; and only if the requesting subscriber has access authorization and if the at least one access condition can be fulfilled: c. Releasing access within the scope of the at least one access condition to the requested specific data set, whereby without such releasing. of access, a participant is prevented from accessing the data record in question by the control authority.
2. The method according to claim 1, wherein the data records of the data level have a classification and the data content includes at least one attribute, wherein in the case of a plurality of attributes at least one of the attributes, preferably of an attribute-value pair, limits, preferably defines, the data form and / or data quantity of the at least one other attribute.
3. The method according to claim 1 or claim 2, wherein the data structures have a classification and the data content includes at least one attribute, wherein when creating a data record and / or a data structure an existing suitable classification is used or, if a suitable classification does not exist, a suitable classification is generated.
4. Method according to one of the preceding claims, wherein the data records and / or the address references are individually encrypted, wherein preferably the control authority does not have a suitable key therefor.
5. Method according to one of the preceding claims, wherein the associated contract information is encrypted, preferably: - the control authority has at least one of several keys by means of which a key group can be formed together, whereby the relevant contract information can be made readable for the control authority exclusively by means of the formed key group, and - with the access request of a subscriber, the at least one missing key is made available to form the appropriate key group for the control authority, wherein more preferably the at least one missing key contains the computer-readable identity of the subscriber or is contained in this identity in a standardized manner that is readable by the control authority.
6. Method according to one of the preceding claims, wherein at least one of the following components is signed by means of the identity of a respective participant or owner of the data set: - the contract information; the access request; the address references; and the data contents of a data record.
7. Method according to one of the preceding claims, wherein the contract information contains at least one of the following access conditions: current location of the subscriber or origin of the access request; Number of accesses, preferably per requesting identity; - Period for access authorization, preferably dependent on the requesting identity; and - Scope of use of the data set.
8. Method according to one of the preceding claims, wherein the owner of the data set or another authorized participant must grant approval as the first factor or second factor for step c.
9. Method according to one of the preceding claims, wherein when checking the identity, the access request and / or the data set by means of the control authority, a trustworthiness level is also checked, wherein preferably an external confirmation is obtained and / or a certificate created by an external certification authority is used.
10. Method according to one of the preceding claims, wherein the address reference used in the present access request is unreadable for a participant and readable for the control authority.
Citation Information
Patent Citations
Polymorphic encryption for security of a data vault
US20230005391A1
Central computer supported encrypted medical data storage HyperCrypt service uses individual patient data symmetric key and centrally protected private asymmetric key
DE102004035424A1
Aggregation of ancillary data associated with source data in a system of networked collaborative datasets
US20190034491A1