Blockchain-based data asset distributed encryption storage device and right confirmation method
Patent Information
- Application Number
- CN202611266501.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]当前主流可信数据空间与数据资产管控方案,业界通常采用中心化云平台进行数据托管,并结合基础的区块链技术进行数据摘要存证,这种现有技术在一定程度上解决了海量数据的基础物理存储问题,并利用区块链的不可篡改特性为数据提供了防伪打卡记录,促进了数据要素的初步流通与共享,然而,随着数据合规监管的趋严以及跨域流通场景的复杂化,现有技术逐渐暴露出明显的不足,无法满足数据流通与治理一体化的建设要求:其一,存储与治理存在割裂,且缺乏差异化安全机制
本发明向各存储节点统一同步密级分类标准及访问权限基线,并由前置治理模块对待入库数据执行敏感字段识别、密级标注、治理标签分类以及空值、违规涉密内容和低质量数据拦截,使数据在进入存储节点之前即完成统一治理,进一步根据数据密级确定分片策略和加密策略,将待入库数据拆分形成多个加密分片并分散存储至不同存储节点,从而降低完整数据集中存储于单一节点所带来的泄露及单点故障风险。
Smart Images

Figure CN122778451A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data ownership confirmation and evidence preservation technology, specifically to a blockchain-based distributed encrypted storage device for data assets and a method for ownership confirmation and evidence preservation. Background Technology
[0002] Trusted data space is the core infrastructure for realizing controllable sharing of data among multiple entities and market-oriented circulation of factors of production. Secure storage of data assets, confirmation of ownership, and compliant governance are the three core requirements for the construction of data space.
[0003] Current mainstream trusted data space and data asset management solutions typically employ centralized cloud platforms for data hosting, combined with basic blockchain technology for data digest notarization. This existing technology, to some extent, solves the fundamental physical storage problem of massive amounts of data and leverages the immutability of blockchain to provide anti-counterfeiting and record-keeping, promoting the initial circulation and sharing of data elements. However, with increasingly stringent data compliance regulations and the growing complexity of cross-domain circulation scenarios, existing technologies are gradually revealing significant shortcomings, failing to meet the requirements of integrated data circulation and governance: Firstly, storage and governance are disconnected, and differentiated security mechanisms are lacking. Existing solutions often treat data governance as an external add-on module, lagging behind the data entry process and unable to intercept low-quality or illegally classified data at the source. Simultaneously, the underlying storage often adopts a unified centralized architecture, failing to implement differentiated sharding and encryption strategies based on different data security levels. Once unauthorized access or a single point of failure occurs, it can easily lead to overall data leakage, making it difficult to balance the timeliness of compliant governance with the security of underlying storage. Secondly, the ownership boundaries are vague and there is a lack of effective on-chain and off-chain verification mechanisms. Traditional blockchain ownership confirmation schemes usually only authorize the original data as a whole, which does not adapt to the characteristics of multi-tenant and cross-domain circulation of data space. They cannot finely distinguish the ownership boundaries of the data owner, controller, and user. On-chain evidence storage is often separated from the actual physical storage and operation behavior off-chain. There is a lack of two-way cross-verification means between on-chain evidence storage rules and off-chain physical logs, which can easily lead to unauthorized data copying and ownership disputes, and it is difficult to accurately trace the data afterward.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a blockchain-based distributed encrypted storage device for data assets and a method for confirming and storing ownership, in order to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: Blockchain-based distributed encrypted storage devices for data assets specifically include: The global management module is used to complete tenant real-name authentication during the tenant onboarding phase, generate unique identity accounts and configure access permissions, uniformly distribute security classification standards and access permission baselines to each independent storage node that supports a hierarchical deployment architecture, establish the correspondence between the local entity and each independent storage node, so that each independent storage node undertakes the storage management of the corresponding local data assets according to the correspondence, and forms mutually isolated data asset ownership management scopes according to the local entity. The global governance layer receives the data asset storage status returned by each storage node, and the global governance layer is deployed independently of each storage node. The pre-processing module is used to identify sensitive fields, mark the security level, and classify the data to be entered into the database with governance tags. It also intercepts null values, illegal classified content, and low-quality data, and outputs the marked data to be entered into the database, the corresponding data security level, and governance tags. The sharded storage module is used to determine the sharding strategy and encryption strategy based on the labeled data to be stored and the data security level, split the data into multiple encrypted shards and store them in different storage nodes, record the storage location of each shard and generate storage logs. The rights confirmation and evidence storage module is used to receive tenant identity accounts, governance tags, data security levels, and shard storage locations. It generates asset digital fingerprints based on the labeled data to be stored. After the data to be stored completes governance, sharding, and rights confirmation, it forms corresponding data assets. Based on the data element three-rights separation rule, it divides the ownership boundaries of data ownership, data management rights, and data usage rights, and maps different ownership boundaries to three main roles: data owner, data controller, and data user. It configures corresponding subject permissions for different subject roles, encapsulates asset digital fingerprints, governance tags, data security levels, subject permissions, and shard storage locations into on-chain evidence storage data, and writes the full lifecycle operation logs with the on-chain evidence storage data into the blockchain. It also deploys ownership smart contracts on the blockchain. The contract audit module is used to call the ownership smart contracts deployed on the chain, retrieve on-chain evidence data and storage logs, perform compliance audits, and conduct two-way cross-verification between on-chain and off-chain.
[0007] Furthermore, the global management module is configured with a global governance layer independent of each storage node. The global governance layer is connected to the global data governance system and is used to issue a unified network-wide security classification standard and cross-subject access permission baseline to each storage node, thereby constraining the data processing, storage management and access permission execution of each storage node.
[0008] Furthermore, the data to be imported is subjected to sensitive field identification, security level labeling, and governance tag classification. The specific steps are as follows: Read the data to be imported into the database, iterate through each field of the data, perform sensitive field identification, and identify the sensitive fields contained within the data; Based on the identified sensitive fields and combined with the preset security classification standards, the security classification of the data to be entered into the database is marked to obtain the data security level of the data to be entered into the database. Based on the results of sensitive field identification and data security level, the governance labels are classified, and corresponding governance labels are configured for the data to be entered into the database. The governance label classification includes risk type labels, business attribute labels, and disposal and control labels.
[0009] Furthermore, based on the labeled data to be imported and its security level, the method for determining the sharding and encryption strategies is as follows: Erasure coding algorithm is used to encode the labeled data to be put into the database into multiple data fragments. Different encryption methods are used for data with different security levels. Among them, public security level data is encrypted with a single layer of AES, while sensitive and core security level data is encrypted with a combination of AES-256 and RSA with a double layer. The encryption process is bound to the tenant identity information.
[0010] Furthermore, the method for generating asset digital fingerprints based on the labeled data to be entered into the warehouse is as follows: Obtain the complete business content of the labeled data to be put into the database, and use a hash algorithm to perform hash operation on the complete business content to generate a unique digital fingerprint of the asset. The asset digital fingerprint is a fixed-length hash digest used for integrity verification and tamper identification of data assets; The complete business content refers to the labeled data to be stored, which has been processed by the pre-processing module to identify sensitive fields, mark confidentiality levels, and classify governance tags, but has not yet been fragmented, split, or encrypted.
[0011] Furthermore, the method for linking the entire lifecycle operation log with on-chain evidence storage data and writing it into the blockchain, and then deploying the ownership smart contract on the blockchain is as follows: Establish a correspondence between tenant identity, subject role, and permission set. Following the separation of data ownership, management rights, and usage rights, tenants associated with the target data are categorized into three subject roles: data owner, controller, and user. A matching permission set is configured for each subject role to obtain subject permissions. The asset digital fingerprint, governance tag, data security level, subject permissions, and shard storage location are encapsulated as on-chain evidence data and written to the blockchain. Upon receiving a data operation request, the subject role of the requesting tenant and the request operation type are matched and verified against the on-chain evidence-based permission set. If the matching verification passes, the corresponding data operation is allowed. If the matching verification fails, the data operation is intercepted and an operation log is generated. The generated operation log is then written into the blockchain as a full lifecycle operation log and associated with the corresponding on-chain evidence storage data. The rights confirmation and evidence storage module also deploys a rights confirmation smart contract with solidified rights confirmation rules, permission verification logic, audit trigger conditions, and cross-entity circulation traceability logic to the blockchain. The cross-entity circulation traceability logic supports the full-process traceability of data sharing, transactions, and authorization between entities in different locations. When data ownership is transferred or the scope of authorization changes, an on-chain ownership change record is automatically generated, the shard access permissions of each storage node are updated synchronously, and the entire circulation link is completely retained to the full lifecycle data.
[0012] Furthermore, the methods for conducting compliance audits and on-chain / off-chain two-way cross-verification are as follows: The system retrieves on-chain evidence data from within the blockchain and reads storage logs generated by each storage node. It then invokes the on-chain ownership smart contract to obtain the fixed ownership rules, permission verification logic, and audit trigger conditions within the contract. The system compares the digital fingerprint of the on-chain evidenced asset with the hash digest recalculated after off-chain sharding recovery to complete data integrity verification. Combining the permission verification logic within the ownership smart contract, it compares the operational behavior in the storage logs with the subject permissions of the on-chain evidenced entity. If the audit trigger conditions are met, a risky behavior is identified, and a risk audit alert is output. If the audit trigger conditions are not met, the data asset status is determined to be compliant, and an audit pass result is output. Simultaneously, it supports special audits of data assets for inclusion in the financial statements, automatically extracting asset ownership certificates, cost measurement basis, and full lifecycle operation records to generate data asset audit reports and ownership confirmation documents that comply with accounting standards.
[0013] To achieve the above objectives, the present invention also provides the following technical solution: A blockchain-based distributed encrypted ownership confirmation and storage method for data assets, wherein the method is used to execute the blockchain-based distributed encrypted storage device for data assets as described above, and the specific steps include: S1. During the tenant onboarding phase, complete the tenant's real-name authentication, generate a unique identity account and configure access permissions, uniformly distribute the security classification standard and access permission baseline to each independent storage node that supports the hierarchical deployment architecture, establish the correspondence between the local entity and each independent storage node, so that each independent storage node undertakes the storage management of the corresponding local data assets according to the correspondence, and forms mutually isolated data asset ownership management scopes according to the local entity. The global governance layer receives the data asset storage status returned by each storage node, and the global governance layer is deployed independently of each storage node. S2. Identify sensitive fields, classify security levels, and classify governance tags for the data to be imported. Intercept null values, illegal classified content, and low-quality data. Output the labeled data to be imported, the corresponding data security level, and governance tags. S3. Based on the labeled data to be stored and the data security level, determine the sharding strategy and encryption strategy, split the data into multiple encrypted shards and store them in different storage nodes, record the storage location of each shard and generate storage logs. S4. Receive tenant identity account, governance tag, data security level, and shard storage location. Generate asset digital fingerprints based on the labeled data to be added to the database. After the data to be added to the database completes governance, sharding, and rights confirmation, corresponding data assets are formed. Based on the data element three-rights separation rule, the ownership boundaries of data ownership, data management rights, and data usage rights are divided. Different ownership boundaries are mapped to three main roles: data owner, data controller, and data user. Corresponding subject permissions are configured for different subject roles. The asset digital fingerprint, governance tag, data security level, subject permissions, and shard storage location are encapsulated as on-chain evidence data. The operation logs of the entire lifecycle are associated with the on-chain evidence data and written to the blockchain. Deploy ownership smart contracts on the blockchain. S5 calls the ownership smart contract deployed on the chain, retrieves on-chain evidence data and storage logs, performs compliance audits and on-chain-off-chain two-way cross-verification.
[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention unifies and synchronizes security classification standards and access permission baselines across all storage nodes. A pre-processing governance module performs sensitive field identification, security classification, governance tagging, and interception of null values, illegal confidential content, and low-quality data on the data to be stored. This ensures that the data undergoes unified governance before entering the storage nodes. Furthermore, based on the data security level, a sharding strategy and encryption strategy are determined, splitting the data to be stored into multiple encrypted shards and distributing them across different storage nodes. This reduces the risk of leakage and single point of failure caused by storing complete data on a single node.
[0015] This invention also generates digital fingerprints of assets based on the labeled data to be stored, delineates the boundaries of subject permissions based on tenant identity accounts, and encapsulates the asset digital fingerprint, governance tag, data security level, subject permissions, and sharded storage location to form on-chain evidence data. At the same time, it associates the full lifecycle operation log with the on-chain evidence data and writes it into the blockchain. By calling the ownership smart contract, and combining the on-chain evidence data and storage logs to perform compliance audits and on-chain-off-chain two-way cross-verification, the content status, subject permissions, on-chain evidence information, and off-chain actual storage and operation behavior of data assets can be mutually verified, providing a basis for the ownership tracing, operation tracking, and abnormal behavior verification of data assets. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall storage device module composition of the present invention; Figure 2 This is a schematic diagram of the distributed encrypted ownership confirmation and evidence storage method for data assets based on blockchain according to the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0019] Example: Please see Figure 1 The present invention provides a technical solution: Blockchain-based distributed encrypted storage devices for data assets specifically include: The global management module is used to complete tenant real-name authentication during the tenant onboarding phase, generate unique identity accounts and configure access permissions, uniformly distribute security classification standards and access permission baselines to each independent storage node that supports a hierarchical deployment architecture, establish the correspondence between the local entity and each independent storage node, so that each independent storage node undertakes the storage management of the corresponding local data assets according to the correspondence, and forms mutually isolated data asset ownership management scopes according to the local entity. The global governance layer receives the data asset storage status returned by each storage node, and the global governance layer is deployed independently of each storage node. The global management module is configured with a global governance layer independent of each storage node. It serves as the unified governance entry point for the entire distributed encrypted storage device, deployed independently of each storage node, operating as an independent central hub for overall management, providing the highest standard of rules for subsequent network management. It does not directly store the complete business content of each data asset, but rather uses it for unified security classification standards and cross-subject access permission baselines. Specifically, the global governance layer obtains applicable data classification and grading rules and subject access control rules from the overall data governance system, and transforms them into security classification standards and access permission baselines that can be uniformly executed by each storage node. These are then synchronously distributed to all storage nodes in real-time or cyclically, achieving the purpose of constraining data processing, storage management, and access permission execution on each storage node. The "complete business content" refers to the data to be stored after the pre-governance module has completed sensitive field identification, security level labeling, and governance tag classification, and is not yet fragmented or encrypted.
[0020] For example, in the scenario of a trusted government data space, a dedicated consortium blockchain is built, and multiple encrypted storage devices of this invention are deployed as underlying storage nodes, connecting to a government-wide data governance platform. The governance platform uniformly issues four-level classification standards for government data, rules for identifying sensitive information, and cross-departmental permission baselines. At the same time, each government department accesses the data space as an independent tenant and completes identity registration.
[0021] During the tenant onboarding phase, the global management module receives the identity registration information submitted by the tenant and performs real-name authentication on the tenant's identity. The tenant's identity includes the enterprise, institution, or other authorized entity corresponding to the data owner, controller, and user. During the real-name authentication process, the consistency of the tenant's submitted entity name, unified identity information, entity qualification information, and authorized operator identity information is verified. When the identity information is verified, a unique identity account is generated for the tenant within the entire storage device system, and a binding relationship is established between the unique identity account and the tenant entity. When the identity information is not verified, the onboarding process of the corresponding tenant is terminated, and no data storage and access permissions are allocated to it.
[0022] After tenant real-name authentication is completed, the rules for determining the corresponding tenant's geographical location are further determined based on the administrative region, organizational affiliation, or business management scope recorded in the tenant's registration information. The geographical location information pre-configured on each independent storage node is then read to establish a correspondence between the geographical location and the independent storage nodes with corresponding geographical locations. A geographical location can correspond to one or more independent storage nodes, which together form the data storage node range for that geographical location. When a tenant submits data to be stored, the corresponding geographical location is determined based on the unique identity account associated with the data, and the data storage node range corresponding to that geographical location serves as a constraint for the subsequent selection of actual storage nodes by the sharding storage module.
[0023] After generating a unique identity account, the global management module configures access permissions for the unique identity account based on the tenant's subject category, business scope, and pre-configured access permission baseline. This limits the tenant's basic operation scope after entering the system, such as one or more of the following: data upload, data query, data retrieval, data modification, permission application, and audit information viewing. For the subsequent permissions of data owners, controllers, and users corresponding to specific data assets, the permissions are further divided according to the ownership relationship of the specific assets to form a hierarchical control relationship between system-level access permissions and subject permissions.
[0024] Meanwhile, the global control module uses different local entities as isolated units for data ownership management, establishing independent data asset ownership management scopes for each. This is used to limit the data ownership and subject permissions of data subsequently submitted by the corresponding local entity after completing pre-governance, fragmented storage, and confirmation of rights and evidence to form data assets. The data asset ownership records and subject permissions corresponding to different local entities are isolated from each other. The subject permissions formed under one local entity do not directly affect the data assets corresponding to other local entities. The global control module only divides the mutually isolated data asset ownership management scopes according to local entities. For specific data assets subsequently included in each ownership management scope, the data owner, data controller, data user, and corresponding permission set are further determined by the confirmation of rights and evidence module according to the data element three-rights separation rule.
[0025] The global governance layer obtains unified security classification rules from the global data governance system and forms corresponding security classification standards based on the security classification rules. The security classification standards include at least the sensitivity determination conditions corresponding to different data security levels and the data processing and storage requirements that should be adopted for different security levels. For example, data is divided into public, internal, sensitive and core security levels, and different security levels correspond to different fragmentation storage requirements and encryption strength requirements.
[0026] Simultaneously, the global governance layer generates cross-subject access permission baselines based on the subject access control requirements of the overall data governance system. These access permission baselines are pre-defined access permission configuration benchmarks for different subject categories and different data security levels. Specifically, they specify the basic operations allowed and prohibited for different subject categories under different data security levels. For example, authorized subjects are allowed to query and access publicly classified data; downloading, forwarding, or modifying sensitive data is further restricted; and only subjects meeting preset authorization conditions are allowed to perform limited operations on core classified data. The access permission baselines do not directly replace the subject permissions of specific data assets, but rather serve as a higher-level constraint on the configuration of specific subject permissions. That is, the permissions of data owners, controllers, and users configured for specific assets must not exceed the access permission baselines issued by the global governance layer.
[0027] After configuring the security classification standards and access permission baselines, the global governance layer synchronizes the latest versions of the security classification standards and access permission baselines to each storage node included in the current distributed storage system, along with the corresponding rule version identifiers. Upon receiving the security classification standards and access permission baselines, each storage node writes them into its local node governance configuration and returns synchronization confirmation information to the global governance layer. The global governance layer determines whether synchronization is complete based on the rule version identifiers returned by each storage node. If the rule version used by a storage node is inconsistent with the current rule version of the global governance layer, that storage node is marked as a node awaiting synchronization, and the corresponding governance rules are reissued until the storage node completes the rule update. Through this method, even if data from different tenants is distributed across different storage nodes, each node still performs data processing according to a unified data security classification method and access control standards.
[0028] Meanwhile, during the subsequent data storage and lifecycle management process, each storage node forms and updates the asset storage status based on the actual storage execution results of the corresponding data assets. The asset storage status is used to characterize the storage results and changes of the corresponding data assets in the current storage node. Each storage node returns the corresponding asset storage status to the global governance layer according to a preset status return cycle, or when the asset storage status is updated, so that the global governance layer can continuously obtain the current storage status of data assets distributed across different storage nodes, which serves as the data foundation for subsequent global data asset management and compliance auditing.
[0029] The pre-processing module is used to identify sensitive fields, mark the security level, and classify the data to be entered into the database with governance tags. It also intercepts null values, illegal classified content, and low-quality data, and outputs the marked data to be entered into the database, the corresponding data security level, and governance tags. The pre-governance module is used to complete unified data identification, classification and quality governance before the data actually enters each storage node. It receives the data to be stored submitted by the tenant and the tenant's unique identity account associated with the data to be stored, and reads the sensitive information identification rules, security classification standards and access permission baselines currently issued by the global governance layer. These serve as the unified rule basis for performing pre-governance on the data to be stored. By setting the governance process before sharding and encrypted storage, data with data quality problems or that does not meet the current governance requirements will not directly enter the subsequent sharding storage stage.
[0030] After reading the data to be imported, each field of the data is captured sequentially. Sensitive field identification rules issued by the global governance layer are invoked to perform sensitive field identification. For example, a semantic analysis engine is used for identification, and sensitive fields contained within the data are checked and identified one by one. For example, for field names containing features such as name, ID number, contact information, address, and bank account, they are determined to be candidate sensitive fields based on the field name identification results. For field names that cannot directly represent the meaning of the data, further matching is performed based on the character composition, length, data format, and preset keywords of the field content to determine whether the corresponding field is a sensitive field.
[0031] While performing sensitive field identification, a data quality pre-check is performed on the data to be imported. For required fields in the data to be imported, the field values are checked to see if they are empty. When a required field is empty or the field content is missing, the corresponding data record is marked as an anomaly. For fields with preset data formats or value range requirements, the current field value is matched with the corresponding field's data format rules and reasonable value range. When the field type, data length, data format, or field value exceeds the preset range, the corresponding data record is marked as low-quality data, so as to pre-screen out empty values, format errors, and low-quality data in the subsequent data to be imported.
[0032] For data to be imported into the database after sensitive fields have been identified, the data to be imported is labeled with a security level based on the identified sensitive fields and the security classification standards uniformly issued by the global governance layer. This maps the security level to which the data belongs and determines the data security level that the current data to be imported should correspond to.
[0033] For example, regarding the classification of data security levels, when the data to be entered into the database contains only publicly available business information, it can be marked as public security level; when it contains non-public business fields such as internal business numbers and internal circulation status, it can be marked as internal security level; when it contains personal identification information, contact information, or other restricted access fields, it can be marked as sensitive security level; and when it contains data content that requires the highest level of access restriction according to the full-domain data governance system, it is marked as core security level.
[0034] After completing the data security classification, the pre-governance module further performs governance label classification based on the sensitive field identification results and data security level, generating corresponding governance labels for the data to be entered into the database. The governance labels include risk type labels, business attribute labels, and disposal control labels. Different types of governance labels are used to describe the risk attributes, business affiliation, and subsequent data processing requirements of the data to be entered into the database. Multiple different types of governance labels can be configured for the same data to be entered into the database at the same time.
[0035] Among them, the risk type label is used to characterize the type of sensitive or restricted data risk in the data to be entered into the database. For example, based on the sensitive field identification results, it can be labeled as personal information, sensitive personal information, important business data, internal restricted data, etc.
[0036] Business attribute tags are used to characterize the business affiliation and data purpose of data to be added to the trusted data space. Based on the tenant identity information, the source of the data to be added, the business name, and the content of the data fields, the corresponding business area, department, data source, or data purpose can be determined. For example, in the government trusted data space, it can be marked as a business attribute such as population management, social security, market supervision, urban governance, or public services, so that the specific business scope to which the data asset belongs can be identified in the subsequent storage, authorization, and auditing process.
[0037] Disposal and control labels are used to characterize the governance and control methods that need to be implemented during the subsequent storage, access, and use of data to be included in the database, and are determined based on data security level, risk type labels, and access permission baselines. For example, a regular encryption label is configured for data with low sensitivity, a strengthened encryption, fragmented storage, or restricted access label is configured for data containing sensitive fields, an authorized access label is configured for data that can only be accessed by specific subjects, and a de-identification label is configured for data that needs to hide some fields during subsequent use.
[0038] For data with empty required fields or incorrect field formats, it will be classified as low-quality data and transmission to subsequent shard storage modules will be stopped. For data whose security level exceeds the current tenant's allowed security level, it will be classified as unauthorized confidential content and blocked.
[0039] The sharded storage module is used to determine the sharding strategy and encryption strategy based on the labeled data to be stored and the data security level, split the data into multiple encrypted shards and store them in different storage nodes, record the storage location of each shard and generate storage logs. The system receives the labeled data to be stored, the corresponding data security level, and the governance tag output by the pre-governance module. Based on the data security level, it determines the corresponding sharding strategy and encryption strategy, encodes the labeled data to be stored into multiple data shards, performs encryption processing on each data shard that matches the data security level, stores the encrypted shards to different storage nodes, records the actual storage nodes written to each encrypted shard and the storage execution results, and generates the corresponding storage log. The sharded storage module pre-establishes a correspondence between different data security levels and erasure coding parameters. The erasure coding parameters include at least the number of original data blocks, the number of redundant check blocks, and the minimum number of shards required for data recovery. The sharding strategy corresponding to different data security levels sets different numbers of original data blocks, redundant check blocks, and distributed storage requirements.
[0040] Erasure coding is a data encoding method that improves the fault tolerance and recovery capabilities of data in a distributed storage environment by adding redundant check data to the original data. Its principle is to divide the data to be stored into several original data blocks, and calculate and generate several redundant check blocks based on the original data blocks, so that the original data blocks and redundant check blocks together form multiple data fragments. When some of the data fragments cannot be retrieved due to storage node failure, fragment loss, or inaccessibility, as long as the number of remaining data fragments reaches the preset data recovery quantity, the original data content can be recalculated and recovered using the remaining data fragments. For any labeled data to be stored, the data security level determined by the pre-governance module is read, the erasure coding parameters used for the current data are determined according to the corresponding relationship, and the labeled data to be stored is divided into several original data blocks according to the preset data block size. Then, the erasure coding algorithm is used to generate several redundant check blocks based on the original data blocks. The original data blocks and redundant check blocks are used together as data fragments to be stored. Thus, when some storage nodes or some data fragments are abnormal, as long as the data fragments that meet the corresponding recovery requirements can be obtained, the original data can be recovered according to the erasure coding algorithm.
[0041] For example, using Reed-Solomon erasure coding, the data to be added to the database is encoded into several original data fragments and redundancy check fragments. If the encoding parameters corresponding to a certain data are six original data fragments and three redundancy check fragments, then a total of nine data fragments are formed. If any three of these fragments become unavailable, the remaining fragments that meet the recovery requirements can still be used to complete the data recovery.
[0042] After erasure coding is completed, the sharded storage module calls the corresponding encryption strategy according to the data security level of the data to be stored: for publicly classified data, and for both publicly classified and internally classified data, each data shard is encrypted using the AES symmetric encryption algorithm at a single layer, with the internally classified data using a key update cycle; for sensitive and core classified data, a combination of AES-256 and RSA encryption is used to improve the key security and access isolation capabilities of high-security data in the distributed storage process. During subsequent data recovery, only after the identity and permissions of the requesting subject are verified, the key information corresponding to the tenant's unique identity account is allowed to be called to complete the AES-256 data encryption key decryption and data shard decryption.
[0043] After encrypting each data shard, the shard storage module reads the node status of the currently available storage nodes, determines candidate storage nodes from the nodes that meet the current data security level storage requirements, and allocates corresponding storage nodes to each encrypted shard according to the principle of distributed storage, so that multiple encrypted shards of the same data asset are written to different storage nodes, avoiding the complete data content being stored in a single node. After determining the target storage node corresponding to each encrypted fragment, a unique fragment identifier is generated for each encrypted fragment, and the encrypted fragment is sent to the corresponding target storage node. After completing the fragment writing, each target storage node returns storage confirmation information. The fragment storage module records the correspondence between the data fragment identifier, data security level, target storage node identifier, and actual storage location based on the returned results.
[0044] Initial storage logs are generated for the generation, encryption, node allocation, and actual writing processes of data shards. Corresponding storage logs are generated during the lifecycle management of read, write, migration, recovery, update, and deletion operations on data shards at each storage node. These storage logs record at least the data asset identifier, shard identifier, tenant identity account, operation type, operation time, storage node, and operation execution result. They are used to record the actual storage and operation process of the corresponding data asset in the off-chain storage node, as well as the actual execution status of the data in the off-chain distributed storage process. The operation type refers to the category of data processing operations initiated by the tenant for the data asset, such as data registration, data upload, data query, data retrieval, data modification, data download, data authorization, authorization revocation, permission change, shard migration, data recovery, and data deletion.
[0045] The rights confirmation and evidence storage module is used to receive tenant identity accounts, governance tags, data security levels, and shard storage locations. It generates asset digital fingerprints based on the labeled data to be stored. After the data to be stored completes governance, sharding, and rights confirmation, it forms corresponding data assets. Based on the data element three-rights separation rule, it divides the ownership boundaries of data ownership, data management rights, and data usage rights, and maps different ownership boundaries to three main roles: data owner, data controller, and data user. It configures corresponding subject permissions for different subject roles, encapsulates asset digital fingerprints, governance tags, data security levels, subject permissions, and shard storage locations into on-chain evidence storage data, and writes the full lifecycle operation logs with the on-chain evidence storage data into the blockchain. It also deploys ownership smart contracts on the blockchain. The system receives tenant identity accounts generated by the global management module, governance tags and data security levels output by the front-end governance module, and shard storage locations formed by the shard storage module. It associates the above information with the data to be stored according to the same annotation. At the same time, for the data to be stored after the front-end governance module has completed the sensitive field identification, security level annotation and governance tag classification, before the data has undergone erasure coding sharding and encryption processing, it generates an asset digital fingerprint for its complete business content, so that the data shards that are subsequently stored in different storage nodes can be uniformly mapped to the same data asset. A hash algorithm is a data processing algorithm that maps data of arbitrary length to a fixed-length hash digest according to preset operation rules. The same data content, using the same hash algorithm and the same data representation, will yield the same hash digest. However, when the original data content changes, the recalculated hash digest will also change. Therefore, hash digests can be used to verify the consistency of data content. Furthermore, hash operations are unidirectional; that is, a corresponding hash digest can be calculated from the original data, but the original data content cannot be directly recovered from the hash digest. The specific method for generating digital fingerprints of assets based on this is as follows: For any labeled data to be stored, the rights confirmation and evidence storage module obtains its complete business content. Before performing the hash operation, it standardizes the labeled data according to the field arrangement order, character encoding method, and data serialization format. The standardized data content is taken as the complete business content to ensure that the same data asset with unchanged content has a consistent data expression form. The data content that has been processed by the pre-governance module for sensitive field identification, security level labeling, and governance tag classification, but has not yet been fragmented and encrypted, is taken as the complete business content. Then, the hash algorithm is used to perform a hash operation on the complete business content to obtain a fixed-length hash digest, and this hash digest is used as the asset digital fingerprint of the corresponding data asset.
[0046] By establishing a unique association between the digital fingerprint of the asset and the data asset, the content state of the current data asset is characterized when the pre-governance is completed but before sharding and encryption. Subsequently, when data integrity verification is required, the complete business content recovered from each off-chain data shard can be re-hashed using the same method, and the re-obtained hash digest can be compared with the asset digital fingerprint recorded on the chain to determine whether the data asset has undergone content changes during storage or circulation.
[0047] After generating the digital fingerprint of the asset, the rights confirmation and evidence storage module takes the current data asset as the object, reads the tenant identity account associated with it, and establishes the correspondence between tenant identity, subject role and permission set according to the source relationship, management relationship and authorized use relationship of the data asset. Based on the data element three-rights separation rule, the ownership boundaries of data ownership, data management right and data use right are divided. The data element three-rights separation rule refers to the ownership configuration rule for data ownership, data use right and data management right in the process of holding, using and operating data. It is used to determine the ownership type and the scope of operations that different tenant subjects have for the same data. Data ownership is mapped to the data owner subject role, data management right is mapped to the data controller subject role, and data use right is mapped to the data user subject role. Tenant identity accounts with corresponding ownership relationships are configured as corresponding subject roles, and matching permission sets are configured for different subject roles to form subject permissions. After completing the delineation of ownership, management, and usage rights of data, as well as the configuration of corresponding subject roles and permission sets, the rights confirmation and evidence storage module determines the corresponding local subject based on the tenant identity account associated with the current data. The resulting ownership boundaries, subject roles, and subject permissions are then incorporated into the data asset ownership management scope of the local subject, so that the data asset ownership records and subject permissions corresponding to different local subjects can be confirmed, stored, and isolated for management.
[0048] The data owner is used to identify the entity that has the original ownership or legal source of control over the target data asset. The data controller is used to identify the entity that assumes the responsibility for data governance and operation management according to the authorization of the data owner or the requirements of the governance system. The data user is used to identify the entity that, after obtaining authorization from the data owner or the data controller, performs restricted-use access to the target data asset.
[0049] For example, a data owner's set of permissions may include one or more of the following: data asset registration, authorization, revocation of authorization, permission change, query, and audit viewing; a data controller's set of permissions may include one or more of the following: governance information maintenance, storage management, access approval, audit viewing, and anomaly handling; and a data user's set of permissions may include one or more of the following: query, access, computation, or authorized download. The permission sets corresponding to each subject role are all constrained by the access permission baseline issued by the aforementioned global governance layer.
[0050] The rights confirmation and evidence preservation module associates and encapsulates the digital fingerprint, governance tag, data security level, subject permissions, and shard storage location returned by the shard storage module of the current data asset to form on-chain evidence preservation data of the corresponding data asset. It saves evidence preservation information that can represent the identity, governance attributes, ownership relationship and off-chain storage location of the data asset, and further generates a unique data asset identifier. This data asset identifier is used as the association index between on-chain evidence preservation data, off-chain shard storage location and subsequent operation logs. The encapsulated on-chain evidence preservation data is submitted to the blockchain. After the blockchain completes the transaction confirmation, it records the corresponding on-chain transaction identifier and block position, so that the data asset forms an initial rights confirmation record that can be traced through asset digital fingerprint, tenant identity and subject permissions. At the same time, when the data asset undergoes authorized shard migration or storage location adjustment in the subsequent life cycle, it does not overwrite the original on-chain evidence preservation record, but generates a corresponding shard storage location change record and writes it to the blockchain in association with the original on-chain evidence preservation data. The current valid shard storage location is determined by combining the initial shard storage location with the most recent shard storage location change record. Upon receiving a data operation request, the system matches and verifies the requesting tenant's role and operation type against the permission set in the on-chain evidence data. If the match passes, the corresponding data operation is allowed; otherwise, it is blocked. An operation log is generated for each data operation request, recording the corresponding permission matching result and data operation execution result. This log is then associated with the corresponding on-chain evidence data and written to the blockchain, forming a full lifecycle operation log for the corresponding data asset. The rights confirmation and evidence storage module also deploys a smart contract on the blockchain containing rights confirmation rules, permission verification logic, audit triggering conditions, and cross-entity circulation traceability logic. This cross-entity circulation traceability logic supports full-process traceability of data sharing, transactions, and authorization between entities in different locations. When data ownership transfer or authorization scope change occurs... During a change, an on-chain ownership change record is automatically generated, and the shard access permissions of each storage node are updated synchronously. The entire circulation chain is fully preserved to the end-lifecycle data. The cross-entity circulation traceability logic is as follows: When a request for ownership change or authorization scope change for a target data asset is received, the data asset identifier, the identity account of the requesting tenant, the identity account of the target tenant, the ownership type to be changed, and the authorized operation scope carried in the request are obtained. When the ownership smart contract is called, the subject role and permission set of the requesting tenant are matched and verified with the currently valid ownership record on the chain. When the matching verification is successful, an ownership change record is generated according to the change request, a new permission version identifier is generated for the changed subject permissions, and the ownership change record is associated with the original on-chain evidence data and written to the blockchain. When the matching verification fails, the corresponding change request is intercepted.
[0051] After the on-chain ownership change record is confirmed, the ownership confirmation and evidence storage module determines the corresponding shard and storage node based on the data asset identifier. It then sends an access control instruction carrying the data asset identifier, shard identifier, and latest access control version identifier to each relevant storage node. Each storage node updates the access permissions of the corresponding shard according to the access control instruction and returns the local access control version identifier and update result. The access control version identifier returned by each storage node is compared with the latest access control version identifier on the chain. If they match, the access control synchronization is considered complete. If they do not match, the corresponding storage node is marked as a node to be updated, and the access control update instruction is resent. Simultaneously, the subject, ownership type, authorization scope, change time, and access control synchronization result before and after the ownership change are recorded in the full lifecycle operation log. The contract audit module is used to call the ownership smart contracts deployed on the chain, retrieve on-chain evidence data and storage logs, perform compliance audits and on-chain-off-chain two-way cross verification; When an audit request for a specific data asset is received, the contract audit module connects to the blockchain network and retrieves the on-chain evidence data corresponding to the target data asset. At the same time, the contract audit module reads the storage logs generated by each storage node during actual operation in parallel according to the storage location indication of the on-chain evidence. Simultaneously, the contract audit module calls the ownership smart contract previously deployed on the chain by the rights confirmation and evidence storage module to obtain the rights confirmation rules, permission verification logic, and audit trigger conditions, which serve as the judgment benchmark for the subsequent verification process.
[0052] The contract audit module first performs a hash digest-based integrity check on the data content to check for inconsistencies in the data state of the storage nodes. The specific method is as follows: Encrypted fragments are extracted from each node based on the storage logs, and fragment recovery is performed according to a preset erasure coding strategy and decryption algorithm to restore the complete business content. Using the same hash algorithm as when the data was stored, the restored complete business content is re-hashd to obtain an in-time hash digest. This recalculated hash digest is then compared with the asset digital fingerprint in the on-chain evidence storage data. If the two are completely consistent, the data integrity check is complete, proving that the off-chain data and the on-chain fingerprint are accurately mapped. If the two are inconsistent, it indicates that the actual off-chain stored content has changed from the data content state recorded during on-chain rights confirmation, which is identified as a data integrity anomaly, and the corresponding data integrity anomaly audit trigger condition is met.
[0053] After completing the data integrity verification, the contract audit module further performs true on-chain and off-chain two-way cross-verification of storage status and operational behavior: On one hand, reverse authenticity verification is performed from off-chain to on-chain: the contract audit module uses the storage logs of the storage nodes as a benchmark to compare the asset status and full lifecycle operation logs stored on-chain; specifically, it verifies whether the shard storage location recorded in the on-chain evidence data matches the physical storage status actually reported by the storage nodes; and it verifies whether each data operation request recorded on-chain has a corresponding actual read / write call and decryption execution record in the node storage log. If the on-chain record has operation logs, but the storage node lacks corresponding execution records, it is determined that there is a discrepancy between the on-chain operation records and the actual off-chain execution results.
[0054] On the other hand, the positive permission constraint verification from on-chain to off-chain is performed: the contract audit module combines the permission verification logic extracted from the ownership smart contract to analyze the retrieved node storage logs and extract the actual operation behavior that occurred for the data asset; these actual operation behaviors that occurred off-chain are compared one by one with the subject permissions clearly defined in the on-chain evidence data to verify whether all read and write flow operations of the storage node are subject to the constraints of the on-chain rights confirmation strategy and whether there are any violations or unauthorized behaviors.
[0055] Finally, combining the permission verification method within the ownership smart contract with the results of the aforementioned integrity verification and two-way cross-verification, it is determined whether the audit triggering conditions are triggered, as follows: When the hash digest recalculated from the complete off-chain business recovery is inconsistent with the digital fingerprint of the asset in the on-chain evidence storage data, a data integrity anomaly is determined to have occurred; when the operation subject and operation type recorded in the storage log exceed the permission set of the corresponding subject role in the on-chain evidence storage data, a subject permission anomaly is determined to have occurred. When an on-chain full lifecycle operation log records that a certain data operation has been matched with permissions and allowed to be executed, but there is no actual execution record matching the data operation in the storage log of the corresponding storage node, or the operation execution result recorded on the chain is inconsistent with the actual execution result recorded in the off-chain storage log, an operation execution consistency anomaly is determined to have occurred. When the actual shard storage status returned by the storage node is inconsistent with the current valid shard storage location recorded on the chain, and there is no corresponding authorized shard storage location change record on the chain, it is determined that an abnormal shard storage status has occurred. When any of the above abnormalities occur, it is determined that the audit triggering conditions are met; otherwise, it is determined that the audit triggering conditions are not met.
[0056] Based on the above, if the audit triggering conditions are met, a risky behavior is identified and a risk audit alarm is output; if the audit triggering conditions are not met, and both two-way verifications are fully matched, the data asset status is determined to be compliant and an audit pass result is output.
[0057] Automatic extraction and calculation of cost measurement basis: The contract audit module extracts the node space occupied, network bandwidth consumption, and computing resource usage of encrypted shard calls for the corresponding data assets during the distributed storage process based on resource consumption logs. At the same time, based on governance computing logs, it extracts the computing power consumption of the pre-governance module when performing sensitive field identification, security level labeling, and data cleaning. After extracting the above resource occupancy and computing power consumption parameters, the system performs quantitative conversion based on the system's preset computing power cost conversion factor and resource billing rate, and summarizes and calculates the estimated value of resource consumption cost in the data asset processing and storage process, which serves as the quantitative basis for measuring the data asset's accounting cost.
[0058] Regarding the generation of data asset audit reports: The system invokes a pre-defined asset entry audit template to structurally aggregate asset ownership certificates, cost measurement basis, and full lifecycle operation records that have undergone two-way cross-verification. The core fields generated include: unique data asset identifier, asset ownership entity, data security level, initial recorded cost valuation after conversion, and on-chain two-way cross-verification compliance conclusion. The compliance basis originates from the results of the two-way cross-verification in this module, namely, the objective fact that the on-chain evidence data and the off-chain storage logs are completely consistent and no abnormal alarms are triggered. This technical means ensures that the report content meets the accounting standards regarding the legal ownership or control of data assets and that the cost can provide quantitative basis for accounting recognition.
[0059] Regarding the generation and output of ownership confirmation documents: The ownership confirmation document is output as a standardized electronic certificate with a digital signature. During generation, the system extracts the current effective subject permission configuration result of the ownership smart contract and the corresponding on-chain ownership confirmation block height. The above content is digitally signed and encapsulated using an electronic signature authentication certificate and the corresponding signature key. The document also embeds the asset's digital fingerprint (hash digest), owner's public key identifier, ownership confirmation timestamp, and the corresponding blockchain transaction hash value. Its legal effect is reflected in the fact that the ownership confirmation document is generated based on the blockchain's multi-node consensus mechanism and the underlying hash anti-tampering characteristics, which can guarantee the data's non-forgeability and non-repudiation. It serves as ownership verification and electronic evidence technology material for data element market asset registration, cross-entity transaction flow, and judicial ownership dispute evidence.
[0060] Please see Figure 2 The present invention also provides a blockchain-based distributed encrypted data asset ownership verification and storage method, used to execute the blockchain-based distributed encrypted data asset storage device described in any of the above claims, the specific steps of which include: S1. During the tenant onboarding phase, complete the tenant's real-name authentication, generate a unique identity account and configure access permissions, uniformly distribute the security classification standard and access permission baseline to each independent storage node that supports the hierarchical deployment architecture, establish the correspondence between the local entity and each independent storage node, so that each independent storage node undertakes the storage management of the corresponding local data assets according to the correspondence, and forms mutually isolated data asset ownership management scopes according to the local entity. The global governance layer receives the data asset storage status returned by each storage node, and the global governance layer is deployed independently of each storage node. S2. Identify sensitive fields, classify security levels, and classify governance tags for the data to be imported. Intercept null values, illegal classified content, and low-quality data. Output the labeled data to be imported, the corresponding data security level, and governance tags. S3. Based on the labeled data to be stored and the data security level, determine the sharding strategy and encryption strategy, split the data into multiple encrypted shards and store them in different storage nodes, record the storage location of each shard and generate storage logs. S4. Receive tenant identity account, governance tag, data security level, and shard storage location. Generate asset digital fingerprints based on the labeled data to be added to the database. After the data to be added to the database completes governance, sharding, and rights confirmation, corresponding data assets are formed. Based on the data element three-rights separation rule, the ownership boundaries of data ownership, data management rights, and data usage rights are divided. Different ownership boundaries are mapped to three main roles: data owner, data controller, and data user. Corresponding subject permissions are configured for different subject roles. The asset digital fingerprint, governance tag, data security level, subject permissions, and shard storage location are encapsulated as on-chain evidence data. The operation logs of the entire lifecycle are associated with the on-chain evidence data and written to the blockchain. Deploy ownership smart contracts on the blockchain. S5. Invoke the on-chain deployed smart contract for ownership, retrieve on-chain evidence data and storage logs, and perform compliance audits and on-chain / off-chain two-way cross-verification. The above formulas are all dimensionless numerical calculations, derived from software simulations using collected large amounts of data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0061] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0062] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0063] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A blockchain-based distributed encrypted storage device for data assets, characterized in that, Specifically, it includes: The global management module is used to complete tenant real-name authentication during the tenant onboarding phase, generate unique identity accounts and configure access permissions, uniformly distribute security classification standards and access permission baselines to each independent storage node that supports a hierarchical deployment architecture, establish the correspondence between the local entity and each independent storage node, so that each independent storage node undertakes the storage management of the corresponding local data assets according to the correspondence, and forms mutually isolated data asset ownership management scopes according to the local entity. The global governance layer receives the data asset storage status returned by each storage node, and the global governance layer is deployed independently of each storage node. The pre-processing module is used to identify sensitive fields, mark the security level, and classify the data to be entered into the database with governance tags. It also intercepts null values, illegal classified content, and low-quality data, and outputs the marked data to be entered into the database, the corresponding data security level, and governance tags. The sharded storage module is used to determine the sharding strategy and encryption strategy based on the labeled data to be stored and the data security level, split the data into multiple encrypted shards and store them in different storage nodes, record the storage location of each shard and generate storage logs. The rights confirmation and evidence storage module is used to receive tenant identity accounts, governance tags, data security levels, and shard storage locations. It generates asset digital fingerprints based on the labeled data to be stored. After the data to be stored completes governance, sharding, and rights confirmation, it forms corresponding data assets. Based on the data element three-rights separation rule, it divides the ownership boundaries of data ownership, data management rights, and data usage rights, and maps different ownership boundaries to three main roles: data owner, data controller, and data user. It configures corresponding subject permissions for different subject roles, encapsulates asset digital fingerprints, governance tags, data security levels, subject permissions, and shard storage locations into on-chain evidence storage data, and writes the full lifecycle operation logs with the on-chain evidence storage data into the blockchain. It also deploys ownership smart contracts on the blockchain. The contract audit module is used to call the ownership smart contracts deployed on the chain, retrieve on-chain evidence data and storage logs, perform compliance audits, and conduct two-way cross-verification between on-chain and off-chain.
2. The blockchain-based distributed encrypted storage device for data assets according to claim 1, characterized in that, The global management module is configured with a global governance layer independent of each storage node. The global governance layer is connected to the global data governance system and is used to issue a unified network-wide security classification standard and cross-subject access permission baseline to each storage node, thereby constraining the data processing, storage management and access permission execution of each storage node.
3. The distributed encrypted storage device for data assets based on blockchain according to claim 1, characterized in that, The specific steps for identifying sensitive fields, marking security levels, and classifying governance tags for data to be imported are as follows: Read the data to be imported into the database, iterate through each field of the data, perform sensitive field identification, and identify the sensitive fields contained within the data; Based on the identified sensitive fields and combined with the preset security classification standards, the security classification of the data to be entered into the database is marked to obtain the data security level of the data to be entered into the database. Based on the results of sensitive field identification and data security level, the governance labels are classified, and corresponding governance labels are configured for the data to be entered into the database. The governance label classification includes risk type labels, business attribute labels, and disposal and control labels.
4. The distributed encrypted storage device for data assets based on blockchain according to claim 1, characterized in that, The method for determining the sharding and encryption strategy based on the labeled data to be imported and the data security level is as follows: Erasure coding algorithm is used to encode the labeled data to be put into the database into multiple data fragments. Different encryption methods are used for data with different security levels. Among them, public security level data is encrypted with a single layer of AES, while sensitive and core security level data is encrypted with a combination of AES-256 and RSA with a double layer. The encryption process is bound to the tenant identity information.
5. The blockchain-based distributed encrypted storage device for data assets according to claim 4, characterized in that, The method for generating asset digital fingerprints based on the labeled data to be added to the warehouse is as follows: Obtain the complete business content of the labeled data to be put into the database, and use a hash algorithm to perform hash operation on the complete business content to generate a unique digital fingerprint of the asset. The asset digital fingerprint is a fixed-length hash digest used for integrity verification and tamper identification of data assets; The complete business content refers to the labeled data to be stored, which has been processed by the pre-processing module to identify sensitive fields, mark confidentiality levels, and classify governance tags, but has not yet been fragmented, split, or encrypted.
6. The distributed encrypted storage device for data assets based on blockchain according to claim 1, characterized in that, The method for linking the full lifecycle operation log with on-chain evidence storage data and writing it to the blockchain, and then deploying ownership smart contracts on the blockchain is as follows: Establish the correspondence between tenant identity, subject role and permission set. According to the rule of separation of data ownership, management right and use right, the tenants associated with the target data are marked as three subject roles: data owner, controller and user. The permission set matching each subject role is configured to obtain subject permissions. The asset digital fingerprint, governance tag, data security level, subject permissions and shard storage location are encapsulated as on-chain evidence data and written into the blockchain. Upon receiving a data operation request, the system matches and verifies the subject role and operation type of the requesting tenant with the permission set stored on the blockchain. If the matching verification passes, the corresponding data operation is allowed to be executed. If the matching verification fails, the data operation is intercepted and an operation log is generated. The generated operation log is then written to the blockchain as a full-lifecycle operation log and associated with the corresponding on-chain evidence data. The rights confirmation and evidence storage module also deploys a rights confirmation smart contract with solidified rights confirmation rules, permission verification logic, audit triggering conditions, and cross-entity circulation traceability logic to the blockchain. The cross-entity circulation traceability logic supports the full-process traceability of data sharing, transactions, and authorization between entities in different locations. When data ownership is transferred or the scope of authorization is changed, an on-chain ownership change record is automatically generated, the shard access permissions of each storage node are updated synchronously, and the entire transfer link is fully preserved to the entire lifecycle of the data.
7. The blockchain-based distributed encrypted storage device for data assets according to claim 6, characterized in that, The methods for conducting compliance audits and on-chain / off-chain two-way cross-verification are as follows: The system retrieves on-chain evidence data from within the blockchain and reads storage logs generated by each storage node. It then calls the ownership smart contract deployed on the chain to obtain the ownership confirmation rules, permission verification logic, and audit trigger conditions embedded within the contract. The system compares the digital fingerprint of the asset stored on the chain with the hash digest recalculated after off-chain sharding recovery to complete data integrity verification. Combined with the permission verification logic within the ownership smart contract, the system compares the operational behavior in the storage logs with the subject permissions of the entity stored on the chain. If the audit trigger conditions are met, a risky behavior is identified and a risk audit alert is issued. If the audit trigger conditions are not met, the data asset status is determined to be compliant and an audit pass result is issued. At the same time, it supports special audits of data assets entering the table, automatically extracts asset ownership certificates, cost measurement basis, and full life cycle operation records, and generates data asset audit reports and ownership confirmation documents that comply with accounting standards.
8. A blockchain-based distributed encrypted data asset ownership confirmation and storage method, wherein the method is executed by the blockchain-based distributed encrypted data asset storage device according to any one of claims 1-7, characterized in that, The specific steps include: S1. During the tenant onboarding phase, complete the tenant's real-name authentication, generate a unique identity account and configure access permissions, uniformly distribute the security classification standard and access permission baseline to each independent storage node that supports the hierarchical deployment architecture, establish the correspondence between the local entity and each independent storage node, so that each independent storage node undertakes the storage management of the corresponding local data assets according to the correspondence, and forms mutually isolated data asset ownership management scopes according to the local entity. The global governance layer receives the data asset storage status returned by each storage node, and the global governance layer is deployed independently of each storage node. S2. Identify sensitive fields, classify security levels, and classify governance tags for the data to be imported. Intercept null values, illegal classified content, and low-quality data. Output the labeled data to be imported, the corresponding data security level, and governance tags. S3. Based on the labeled data to be stored and the data security level, determine the sharding strategy and encryption strategy, split the data into multiple encrypted shards and store them in different storage nodes, record the storage location of each shard and generate storage logs. S4. Receive tenant identity account, governance tag, data security level, and shard storage location. Generate asset digital fingerprints based on the labeled data to be added to the database. After the data to be added to the database completes governance, sharding, and rights confirmation, corresponding data assets are formed. Based on the data element three-rights separation rule, the ownership boundaries of data ownership, data management rights, and data usage rights are divided. Different ownership boundaries are mapped to three main roles: data owner, data controller, and data user. Corresponding subject permissions are configured for different subject roles. The asset digital fingerprint, governance tag, data security level, subject permissions, and shard storage location are encapsulated as on-chain evidence data. The operation logs of the entire lifecycle are associated with the on-chain evidence data and written to the blockchain. Deploy ownership smart contracts on the blockchain. S5 calls the ownership smart contract deployed on the chain, retrieves on-chain evidence data and storage logs, performs compliance audits and on-chain-off-chain two-way cross-verification.