Trusted data space blood relationship tracking management architecture and method

The trusted data space lineage tracing management architecture solves the security and integrity issues in data flow, realizes the authenticity and immutability of data flow, ensures data sovereignty and privacy security, and provides a trusted data circulation model.

CN121567375APending Publication Date: 2026-02-24SHANGHAI YUNBAO DATA WING DATA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511636257.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

In a trusted data space, data transfer presents security, susceptibility to tampering, and integrity issues. Existing technologies struggle to establish effective data circulation models to safeguard data sovereignty and privacy.

Method used

The trusted data space lineage tracing management architecture is adopted, which includes a data layer, a technology layer, and a contract layer. Through metadata preprocessing, trusted execution environment encryption processing, dynamic and static dual-mode flow characterization, and token authentication mechanism, a technical closed loop is formed to ensure the authenticity and immutability of data flow.

Benefits of technology

It realizes a complete link for metadata flow, ensures the authenticity and immutability of the information in the link, safeguards data sovereignty and privacy security, and provides a reliable data circulation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567375A_ABST
    Figure CN121567375A_ABST
Patent Text Reader

Abstract

The invention discloses a trusted data space blood relationship tracking management architecture and method. The data layer is used for preprocessing metadata of multiple parties; the technical layer is used for constructing a trusted execution environment, performing encryption processing on metadata from multiple parties in the trusted execution environment, and performing trusted flow evidence storage at the same time; the contract layer is used for carrying out full-link description on circulation of the encrypted metadata in a dynamic mode and a static mode; and the application layer is used for authenticating a user and granting a corresponding token according to an authentication result, the token corresponds to the trusted flow certificate, and the method is implemented by adopting the management architecture. According to the credible data space blood relationship tracking management architecture and method, the credible flow storage evidence and blood relationship tracking form a technical closed loop, a complete link for metadata flow is provided, the authenticity and non-tampering property of link information are ensured, and the data sovereignty and privacy security are ensured through a data flow mode combining the credible flow storage evidence and the blood relationship tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of trusted data space technology, and in particular to a trusted data space lineage tracing management architecture and method. Background Technology

[0002] Trusted data lineage tracing management technology is a technology used to trace the flow of data. When trusted data flows between multiple platforms, there are security issues in the flow. At the same time, the current characterization of the flow is also susceptible to tampering and integrity problems. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a trusted data space lineage tracing management architecture and method, which forms a technical closed loop with trusted circulation and evidence storage and lineage tracing, provides a complete link for metadata flow, ensures the authenticity and immutability of the link information, and the combined data circulation mode protects data sovereignty and privacy security.

[0004] This invention provides a trusted data space lineage tracing management architecture, including:

[0005] The data layer is used for the preprocessing of metadata from multiple parties.

[0006] The technical layer is used to build a trusted execution environment and encrypt metadata from multiple parties within the trusted execution environment, while also ensuring trusted circulation and evidence storage.

[0007] The contract layer is used to depict the entire flow of encrypted metadata through both dynamic and static modes.

[0008] The application layer is used to authenticate users and grant corresponding tokens based on the authentication results. The tokens correspond to trusted circulation and storage evidence.

[0009] In one embodiment, the contract layer includes:

[0010] Static modules are used to record the dependencies for retrieving metadata;

[0011] The dynamic module is used to capture the actual data flow trajectory through runtime logs.

[0012] In one embodiment, the static module includes:

[0013] The lexical analysis submodule is used for ANTLR4-based SQL syntax rule files, decomposing SQL statements into token sequences;

[0014] The syntax analysis submodule is used to construct an abstract syntax tree based on token sequences and using recursive descent parsing.

[0015] The semantic analysis submodule is used to infer the dependencies of metadata based on the abstract syntax tree using the TANE algorithm.

[0016] In one embodiment, the data layer includes:

[0017] The partitioning submodule is used to divide a Hive table into 128 partitions using a composite partition key consisting of time and business domain.

[0018] The node label index module is used to create a unique index for the qualifiedName attribute of the Table label;

[0019] The relation type index module is used to create a timestamp attribute range index for the DERIVED_FROM relation;

[0020] Query plan optimization module: Used to analyze the execution plan through PROFILE MATCH and return the complete path.

[0021] In one embodiment, the data layer further includes:

[0022] The caching strategy module is used to cache hotspot lineage paths in Redis and set the TTL to 30 minutes.

[0023] In one embodiment, the data layer further includes:

[0024] The parallel query module is used to configure Neo4j with dbms.threads.worker_count=32.

[0025] In one embodiment, the data layer further includes:

[0026] The data enhancement module is used to process metadata to obtain standard metadata.

[0027] This invention also provides a trusted data space lineage tracing management method, implemented using the aforementioned trusted data space lineage tracing management architecture, comprising the following steps:

[0028] Users are authenticated, and corresponding tokens are granted based on the authentication results. The tokens correspond to trusted circulation and storage evidence.

[0029] Preprocessing of metadata from multiple parties;

[0030] Construct a trusted execution environment and encrypt metadata from multiple parties within the trusted execution environment, while also performing trusted circulation and evidence storage.

[0031] The entire flow of encrypted metadata is depicted through both dynamic and static modes.

[0032] The trusted data space lineage tracing management architecture and method provided by this invention form a technical closed loop of trusted circulation and evidence storage and lineage tracing, providing a complete link for metadata flow, ensuring the authenticity and immutability of the information in the link, and the combined data circulation mode protects data sovereignty and privacy security. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 A schematic diagram of the trusted data space lineage tracing management architecture provided by the present invention.

[0035] Figure 2 This is a flowchart illustrating the trusted data space lineage tracing management method provided by the present invention. Detailed Implementation

[0036] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. Based on the description of the present invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present invention.

[0037] In the description of this invention, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.

[0038] The terms “upper,” “lower,” “left,” “right,” “front,” “back,” “top,” “bottom,” “inner,” and “outer,” etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of the invention is in use. They are only for the convenience of description and simplification, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention.

[0039] The terms “first,” “second,” “third,” etc., are used merely to distinguish elements with similar properties, not to indicate or imply relative importance or a specific order.

[0040] The terms “include,” “comprising,” or any other variation thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.

[0041] Example 1

[0042] Please see Figure 1 The trusted data space lineage tracing management architecture provided in this embodiment includes:

[0043] Data layer 1 is used for preprocessing metadata from multiple parties;

[0044] Technical layer 2 is used to build a trusted execution environment and encrypt metadata from multiple parties within the trusted execution environment, while also performing trusted circulation and evidence storage.

[0045] Contract layer 3 is used to depict the entire flow of encrypted metadata through both dynamic and static modes.

[0046] Application layer 4 is used to authenticate users and grant corresponding tokens based on the authentication results. The tokens correspond to trusted circulation evidence.

[0047] Understandably, the management architecture is first explained in its entirety. Data Layer 1 preprocesses metadata from multiple sources, such as adding tags or performing other processing operations. Application Layer 4 can be platforms like Hyperledger Fabric and Ethereum. Each platform can use OAuth 2.0 for enterprise SSO authentication. After authentication, Application Layer 4 assigns a corresponding JWT token based on the authentication information. The JWT token may embed a DID identifier, and the validity of the public key is verified through the technical layer. Contract Layer 3 verifies the JWT signature, DID permissions, and Solidity code. After verification, based on the request from Application Layer 4 and the permission management of DID permissions, the corresponding metadata is indexed from Data Layer 1. The metadata is encrypted by Technical Layer 2, and key links in the metadata flow are recorded by smart contracts based on Contract Layer 3. When a key event is triggered, the smart contract automatically writes information such as hash value, timestamp, and operation subject into the block, forming an immutable audit trail. This achieves multi-chain collaborative notarization of batch information and status, and its cross-chain interoperability also supports interconnection with platforms such as Hyperledger Fabric. Finally, the encrypted metadata is transmitted to Application Layer 4.

[0048] In some embodiments, contract layer 3 includes:

[0049] Static modules are used to record the dependencies for retrieving metadata;

[0050] The dynamic module is used to capture the actual data flow trajectory through runtime logs.

[0051] Understandably, in modern data architectures, static modules, which do not involve runtime metadata, are the foundation and starting point for building enterprise-level data graphs. They analyze the relationships between metadata to infer metadata dependencies. Dynamic modules, on the other hand, can capture the actual data flow trajectory through runtime logs. They can record the actual flow of metadata and be used to verify the dependencies between metadata inferred by static modules. Combining the two can effectively improve the metadata graph, as shown in the table below. Through the above operations, different query perspectives and application capabilities can be provided for different application scenarios.

[0052] Bloodline type Definition Example Application scenarios Positive bloodline Data link from Kafka to Flink to Hive to Tableau Data flow tracking Reverse bloodline The source path from Tableau to Hive to Flink to Kafka Problem Data Location Field-level lineage The Hive table's userid field originates from the Kafka userid field. Fine-grained data quality control Superior bloodline The Hive table dwuserclick originates from the Kafka Topic user_click Data asset genealogy management

[0053] In some embodiments, the static module includes:

[0054] The lexical analysis submodule is used to decompose SQL statements into token sequences from ANTLR4-based SQL syntax rule files (such as MySQL.g4).

[0055] It is known that ANTLR4 is an industry standard. It works through a predefined syntax rule file, and its processing flow can be to read an SQL string, ignore whitespace and comments according to the rules, and finally output a token sequence.

[0056] The syntax analysis submodule is used to construct an abstract syntax tree based on token sequences and using recursive descent parsing.

[0057] As we can see, the goal of this module is to understand the true meaning of SQL and extract the dependencies between fields. Key nodes include SelectClause, FromClause, and JoinNode. Let's take a complex SQL query as an example:

[0058] CREATE TABLE stu_tj AS

[0059] SELECT b.id, b.name

[0060] FROM (SELECT oldId id, name FROM stu WHERE id = '2') b.

[0061] The semantic analysis submodule is used to infer the dependencies of metadata based on the abstract syntax tree using the TANE algorithm.

[0062] Understandably, dependencies can be based on abstract syntax tree traversal. The TANE algorithm, for example, is a partitioning refinement-based algorithm used to automatically discover functional dependencies from metadata. In static analysis, we can apply its concepts to semantic constraints in SQL.

[0063] In some embodiments, data layer 1 includes:

[0064] The partitioning submodule is used to divide the Hive table into 128 partitions using a composite partition key of time and business domain. It can be understood that data layer 1 can use a graph database, the metadata table structure can include entity ID, type, name, and attributes, the synchronization process covers three stages: reading, transformation, and writing, and the query process quickly returns lineage relationships by parsing metadata.

[0065] The node label index module is used to create a unique index for the qualifiedName attribute of the Table label.

[0066] What we know is that this actually creates an index with uniqueness constraints. After creating the index, Neo4j maintains a structure similar to a B+ tree. During a query, it can directly locate the target node through the index, with a time complexity close to O(log n). Its role in lineage queries is crucial: almost all lineage queries start by finding a specific table or field node by name. This index ensures optimal performance for locating the starting point, making it the first and most important accelerator in the entire query chain, reducing query response time from 800ms for a full table scan to 12ms.

[0067] The relation type index module is used to create a timestamp attribute range index for the DERIVED_FROM relation.

[0068] It is known that it supports filtering lineage links by time range, and the query latency is <50ms with a scale of 10 million relationships.

[0069] Query plan optimization module: Used to analyze the execution plan through PROFILE MATCH and return the complete path.

[0070] As we know, PROFILE MATCH displays the execution plan of the query, the time spent at each step, and the number of records generated. PROFILE MATCH p=(t:Table)-[*1..3]->() can be understood as limiting the path depth. RETURN p analyzes the execution plan to avoid the relationship explosion problem.

[0071] In some embodiments, data layer 1 further includes:

[0072] The caching strategy module is used to cache hotspot lineage paths in Redis and set the TTL to 30 minutes.

[0073] Understandably, for some frequently queried metadata, the cached content can include the index conditions, path depth, lineage type and other filtering conditions mentioned above. Setting the TTL to 30 minutes can be understood as the cache expiration time being 30 minutes.

[0074] In some embodiments, data layer 1 further includes:

[0075] The parallel query module is used to configure Neo4j with dbms.threads.worker_count=32.

[0076] Understandably, this parameter sets the size of the worker thread pool used by Neo4j to handle query requests. The default value is usually equal to the number of CPU cores; setting it to 32 means that a maximum of 32 queries can be executed simultaneously, supporting 200+ complex lineage queries per second.

[0077] In some embodiments, data layer 1 further includes:

[0078] The data enhancement module is used to process metadata to obtain standard metadata.

[0079] Understandably, the data augmentation module can use AI technology and machine learning algorithms to deduplicate, correct, and classify massive amounts of raw data, transforming unstructured data into standardized data assets and providing a high-quality data foundation for kinship tracing.

[0080] Example 2

[0081] Please see Figure 2 This embodiment provides a trusted data space lineage tracing management method, implemented using the aforementioned trusted data space lineage tracing management architecture, and includes the following steps:

[0082] Users are authenticated, and corresponding tokens are granted based on the authentication results. The tokens correspond to trusted circulation and storage evidence.

[0083] Preprocessing of metadata from multiple parties;

[0084] Construct a trusted execution environment and encrypt metadata from multiple parties within the trusted execution environment, while also performing trusted circulation and evidence storage.

[0085] The entire flow of encrypted metadata is depicted through both dynamic and static modes.

[0086] As described above, the trusted data space lineage tracing management architecture and method provided by this invention form a technical closed loop of trusted circulation and evidence storage and lineage tracing, providing a complete link for metadata flow, ensuring the authenticity and immutability of the information in the link, and the combined data circulation mode protects data sovereignty and privacy security.

[0087] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A trusted data space lineage tracing management architecture, characterized in that, include; The data layer is used for the preprocessing of metadata from multiple parties. The technical layer is used to build a trusted execution environment and encrypt metadata from multiple parties within the trusted execution environment, while also ensuring trusted circulation and evidence storage. The contract layer is used to depict the entire flow of encrypted metadata through both dynamic and static modes. The application layer is used to authenticate users and grant corresponding tokens based on the authentication results. The tokens correspond to trusted circulation and storage evidence.

2. The trusted data space lineage tracing management architecture as described in claim 1, characterized in that, The contract layer includes: Static modules are used to record the dependencies for retrieving metadata; The dynamic module is used to capture the actual data flow trajectory through runtime logs.

3. The trusted data space lineage tracing management architecture as described in claim 2, characterized in that, Static modules include: The lexical analysis submodule is used for ANTLR4-based SQL syntax rule files, decomposing SQL statements into token sequences; The syntax analysis submodule is used to construct an abstract syntax tree based on token sequences and using recursive descent parsing. The semantic analysis submodule is used to infer the dependencies of metadata based on the abstract syntax tree using the TANE algorithm.

4. The trusted data space lineage tracing management architecture as described in claim 1, characterized in that, The data layer includes: The partitioning submodule is used to divide a Hive table into 128 partitions using a composite partition key consisting of time and business domain. The node label index module is used to create a unique index for the qualifiedName attribute of the Table label; The relation type index module is used to create a timestamp attribute range index for the DERIVED_FROM relation; Query plan optimization module: Used to analyze the execution plan through PROFILE MATCH and return the complete path.

5. The trusted data space lineage tracing management architecture as described in claim 4, characterized in that, The data layer also includes: The caching strategy module is used to cache hotspot lineage paths in Redis and set the TTL to 30 minutes.

6. The trusted data space lineage tracing management architecture as described in claim 4, characterized in that, The data layer also includes: The parallel query module is used to configure Neo4j with dbms.threads.worker_count=32.

7. The trusted data space lineage tracing management architecture as described in claim 4, characterized in that, The data layer also includes: The data enhancement module is used to process metadata to obtain standard metadata.

8. A trusted data space lineage tracing management method, characterized in that, The implementation using the trusted data space lineage tracing management architecture of any one of claims 1 to 7 includes the following steps: Users are authenticated, and corresponding tokens are granted based on the authentication results. The tokens correspond to trusted circulation and storage evidence. Preprocessing of metadata from multiple parties; Construct a trusted execution environment and encrypt metadata from multiple parties within the trusted execution environment, while also performing trusted circulation and evidence storage. The entire flow of encrypted metadata is depicted through both dynamic and static modes.