Data Elementalization, Assetization and Decentralized Circulation Method, Device and Storage Medium

By mapping data resources into data elements and converting them into digital assets, and using blockchain technology for decentralized circulation, the management and circulation of data elements in the existing technology in a centralized environment is solved, the privacy protection and verifiability of data is realized, and the transparency and trust of circulation are enhanced.

CN118096155BActive Publication Date: 2025-06-24SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410154773.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-06-24
Estimated Expiration
2044-02-02

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively manage and circulate data elements in a centralized environment, and there are problems of information tampering, single point of failure, security risks and trust, and there is a lack of an effective trustworthy verification mechanism for data elements.

Method used

By mapping data resources into data elements and converting them into digital assets using evaluation models, and using blockchain technology for decentralized circulation, data privacy protection and verifiability are achieved.

Benefits of technology

The decentralized circulation of data elements is realized, the privacy protection and verifiability of data is ensured, single point of failure and security risks are avoided, and the transparency and trust of data management and circulation are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118096155B_ABST
    Figure CN118096155B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, apparatus, and storage medium for data elementization, assetization, and decentralized circulation. The present invention generates data elements from data resources in any field according to the guidance of the domain knowledge of the field of the data resources through corresponding mapping rules; converts the data elements into digital assets using corresponding evaluation models according to the market demand of the market where the data elements are located; the assetized asset data is described from three dimensions and four levels of modality, privacy level, and granularity; the described digital assets are processed according to a set processing method to achieve privacy protection and verifiable data for on-chain, and the circulation of digital assets is defined as a transaction including a transaction hash, transaction status, transaction block, timestamp, sender address, and recipient address. Transactions form blocks, and blocks are organized into a blockchain through hash pointers, and the logical associations between transactions form a transaction network; the decentralized circulation of digital assets is achieved through verifiable transaction operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of blockchain, and particularly to a method, device and storage medium for data elementization, assetization and decentralized circulation. Background Art

[0002] Data elements refer to data resources that participate in social production and business activities and bring economic benefits to owners or users.

[0003] The value of data elements needs to be reflected in the process of element circulation. However, current data elements generally circulate, are shared or traded in a centralized environment, and element information may be easily copied or tampered with, resulting in infringement of the rights and interests of element owners, which dampens the enthusiasm of owners or users. In addition, the circulation, sharing or trading of data elements in a centralized environment often faces problems such as single-point failures, security risks, and opacity and trust. Blockchain technology is a new distributed infrastructure and computing paradigm that uses a block-chain data structure to verify and store data, uses a distributed node consensus algorithm to generate and update data, uses cryptography to ensure the security of data transmission and access, and uses smart contracts composed of automated script codes to program and operate data. Converting data elements into digital assets stored on the blockchain can endow the element circulation process with decentralized characteristics and avoid problems such as tampering of element information as much as possible. However, in a large-scale data environment, data management becomes very complex, and different types and privacy-constrained data are mixed and stored together, making it difficult to manage. This is mainly reflected in the following aspects: First, data resources often lack value anchoring and are difficult to participate in transaction circulation as digital assets, that is, they lack elementization and assetization. Second, with the continuous upgrading of privacy regulations, data privacy protection has become particularly important, and users have different privacy requirements for different data. Third, different types of data are suitable for different analysis methods. When managing existing data elements, different types of data are mixed, which is not conducive to data analysis. Fourth, a more flexible data sharing and access control mechanism is needed so that users can choose to share data at a specific granularity while restricting access to other data. Fifth, there is a lack of an effective data element trusted verification mechanism. Currently, there is no effective prior art that can simultaneously meet the above requirements to achieve secure and efficient representation and assetization of data elements, privacy protection of data, and trusted verification of data. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method, device and storage medium for data elementization, assetization and decentralized circulation.

[0005] In a first aspect, the present invention provides a method for data elementization, assetization and decentralized circulation, including:

[0006] Guided by the domain knowledge of the data resource in this field, the data resource is mapped into data elements through corresponding mapping rules; the data elements are transformed into digital assets according to the market demand of the market where they are located, using the corresponding evaluation model.

[0007] The digital asset data after assetization is described from three dimensions: modality, privacy level, and granularity. Among them, from the multi-modal dimension, it is divided into structured data, picture and video data, text information data, and other types of data; from the privacy level dimension, it is divided into raw data, desensitized data, modeled data, and artificial intelligence data; from the granularity dimension, it is divided into data items, sets, files, and assets.

[0008] After the description, the digital assets are processed according to the set processing method to achieve privacy protection and verifiable data for on-chain. The circulation of digital assets is defined as a transaction that includes transaction hash, transaction status, the block to which the transaction belongs, timestamp, sender address, and receiver address. Transactions form blocks, and blocks are organized into a blockchain through hash pointers, and the logical associations between transactions form a transaction network; the decentralized circulation of digital assets is achieved through verifiable transaction operations.

[0009] Furthermore, guided by the domain knowledge of the data resource in this field, mapping the data resource into data elements through the corresponding mapping rules includes:

[0010] Create a mapping rule MR for converting the original data resource into meaningful data elements through set data processing steps. The set data processing steps include data cleaning, feature engineering, dimensionality reduction, classification, and clustering. Create domain knowledge DK based on the professional knowledge of a specific field to guide the mapping of the data resource to data elements. Then, the process of mapping the data resource into data elements is expressed as follows:

[0011] DE = MR(DR, DK).

[0012] Furthermore, transforming the data elements into digital assets according to the market demand of the market where they are located, using the corresponding evaluation model includes:

[0013] Build an evaluation model EM and a market demand model MD for evaluating the value of the digital assets formed by the data elements according to the specific attributes of the data elements through regression models and machine learning algorithms. Then, the process of mapping the data elements into digital assets is expressed as follows:

[0014] DA = EM(DE, MD).

[0015] Furthermore, the rules for classifying the digital asset data after assetization from the modal dimension are as follows: Data with a clear pattern and format organization is classified as structured data; Digital assets containing large file formats such as pictures, audio, and video are classified as picture and video data; Digital assets organized in text form and having a semantic structure are classified as text information data; For various data forms other than structured data, picture and video data, and text information data, they are classified as other types of data;

[0016] The rules for classifying the digital asset data after assetization from the privacy level dimension are as follows: Digital assets that the owner of the digital asset actively gives up processing are classified as raw data; Digital assets that the owner of the digital asset requests to be desensitized are classified as desensitized data; Digital assets that the owner of the digital asset requests to be modeled are classified as modeled data; Artificial intelligence data that the owner of the digital asset needs to upload is classified as artificial intelligence data;

[0017] The rules for classifying digital assets from the granularity dimension are as follows: Data is divided into four categories: data items, sets, files, and assets according to the granularity size. Among them, a set is composed of data items, a file includes picture and video data and text information data, and an asset is composed of a combination of various types of data items, sets, and files.

[0018] Furthermore, in the process of uploading the digital asset after being processed according to the set processing method to achieve privacy protection and data verifiability, the ways to provide data verifiability include: When uploading off-chain data to the chain, first use a cryptographic hashing algorithm to process the data to obtain a digest of the data, and then use a digital signature algorithm to sign the digest of the data, and upload the signature to the chain to achieve data verifiability.

[0019] Furthermore, the uploading of the digital asset after being processed according to the set processing method to achieve privacy protection and data verifiability includes different processing of the data after weighing security and performance according to the division of different privacy levels before uploading, including: For raw data and desensitized data, use the national cryptographic algorithm standard to encrypt the data, and desensitize the desensitized data according to the set desensitization rules before encrypting and uploading; For modeled data and artificial intelligence data, choose whether to use an encryption algorithm for encryption according to the requirements.

[0020] Furthermore, the ways to provide verifiable transaction operations include:

[0021] Calculate the hash of each transaction in the block to obtain the transaction hash, sign the transaction hash of each transaction in the block using the private key, and save the signature in the transaction data as the first layer of signature;

[0022] Construct a Merkle hash tree using the transaction hashes within a block to obtain the Merkle root. When a transaction is packaged into a block, store the signature of the Merkle root in the block as the second-level signature.

[0023] For a blockchain formed by organizing multiple blocks, sign all the blocks in the blockchain as a whole to obtain a blockchain signature, which serves as the third-level signature.

[0024] Verify transaction operations through the three-level signatures of transactions, blocks, and the blockchain.

[0025] Furthermore, configure the transaction hash value of the previous block in the block header of each block to form a hash pointer.

[0026] In a second aspect, the present application provides an apparatus for realizing data elementization, assetization, and decentralized circulation, including: at least one processing unit, where the processing unit is connected to a storage unit through a bus unit, and the storage unit stores a computer program. When the computer program is executed by the processing unit, it realizes the data elementization, assetization, and decentralized circulation method.

[0027] In a third aspect, the present application provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, it realizes the data elementization, assetization, and decentralized circulation method.

[0028] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:

[0029] This application designs a mapping mechanism between data elements and digital assets, explores the value of the original complex and disordered data resources, extracts them into data elements, and maps the data elements into digital assets, laying a preliminary foundation for the subsequent circulation of elements. The digital asset data after assetization in this application is described from three dimensions: modality, privacy level, and granularity. Data of different types, privacy levels, and granularities are organized into a three-dimensional and four-layer structure, classifying the originally diverse data in different dimensions and at different levels to better represent the data. This mapping and organization method helps to better manage, understand, and apply the data, and at the same time provides more flexibility to process and share the data as needed. In particular, it is possible to better process and encrypt the digital asset data according to the classification in different dimensions and at different levels to protect privacy. This application regards the circulation of digital assets obtained by assetizing data elements as transactions. Transactions form blocks, and blocks are composed of blockchain through hash pointers. The logical association between transactions constitutes a transaction network, and the decentralized circulation of digital assets is realized through verifiable transaction operations. The decentralized element circulation mechanism formed by this application enables the elements to be chained, that is, introducing blockchain technology, providing a security guarantee for the circulation of data elements, making each transaction verifiable, and being able to detect tampered transactions in a timely manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments in accordance with the present invention, and are used together with the specification to explain the principles of the present invention.

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a schematic diagram of the architecture for data elementization, assetization, and decentralized circulation provided by an embodiment of the present disclosure.

[0033] Figure 2 It is a schematic diagram of the flow of the mapping mechanism between data elements and digital assets provided by an embodiment of the present disclosure.

[0034] Figure 3 It is a flowchart for describing the digital asset data after assetization from three dimensions: modality, privacy level, and granularity provided by an embodiment of the present disclosure.

[0035] Figure 4 It is a schematic diagram of the transaction operation verification method provided by an embodiment of the present disclosure.

[0036] Figure 5Flow chart of obtaining multi - layer signatures by digitally signing transactions, blocks, and blockchains using the SM2 algorithm provided by an embodiment of the present invention;

[0037] Figure 6 Flow chart of verifying the multi - layer signatures obtained by digitally signing transactions, blocks, and blockchains using the SM2 algorithm provided by an embodiment of the present invention;

[0038] Figure 7 Schematic diagram of a device for realizing data elementization, assetization, and decentralized circulation provided by an embodiment of the present invention. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] It should be noted that in this article, the terms "include", "comprise", or any other variant thereof are intended to cover non - exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0041] As Figure 1 shown, the present invention provides a method for data elementization, assetization, and decentralized circulation, including:

[0042] Data resources in any field are used to generate data elements according to the domain knowledge of the data resources in this field and corresponding mapping rules; the data elements are converted into digital assets using corresponding evaluation models according to the market demands of the market where the data elements are located. The present invention designs a mapping mechanism for data resource elementization and a mapping mechanism for data element assetization. Raw data resources off - chain often lack value anchoring and are difficult to participate in transaction circulation as digital assets. As Figure 2 shown, the mapping mechanism designed by the present invention realizes the mapping process from data resources to digital assets through the following steps, including:

[0043] S11, Data resource value discovery. Anchor data resources to subject matters or value measurement attributes through manual or automated means. Analyze and process data resources to reveal potential value, insights, and opportunities in the data resources through subject matters or value measurement attributes.

[0044] S12, Map data resources into data elements. For data resources in any domain, generate data elements according to the guidance of the knowledge of the domain to which the data resources belong through corresponding mapping rules.

[0045] To enable the element-based and asset-based management of data resources in various domains, the present disclosure makes the following definitions:

[0046] DR = {Raw_Data1, Raw_Data2,..., Raw_Datan} represents data resources. Where Raw_Datax refers to the original data owned or collected by the data resources, and Raw_Datax can be structured data such as database records and tabular data, or unstructured data such as text, images, and audio.

[0047] MR(Param1, Param2,......Paramn) represents an element-based mapping rule for converting raw data resources into meaningful data elements. The element-based mapping rule is any one or several data processing steps among data cleaning, feature engineering, dimensionality reduction, classification, and clustering. Where Paramx represents the parameters required by the mapping rule.

[0048] DK = {Expertise1, Expertise2,..., Expertisen} represents domain knowledge. Where Expertisex refers to the professional knowledge of a specific domain, which is used to guide the mapping process from data resources to data elements and includes the experience and understanding of domain experts.

[0049] DE = {Processed_Data1, Processed_Data2,..., Processed_Datan} represents data elements. Where Processed_Datax represents the data with certain value after mapping of the data resources.

[0050] The process of mapping data resources into data elements is expressed as follows:

[0051] DE = MR(DR, DK);

[0052] That is, for any domain data resource DR, under the guidance of the domain knowledge DK to which the data resource DR belongs, the data element DE is generated through the corresponding mapping rule MR.

[0053] S13. Map data elements to digital assets using an expression model so that the data elements can be circulated. According to the definition of the expression model, map the data elements. The expression model includes converting text data into digital codes, mapping attribute values to numerical values, and mapping relationships to connections or weights.

[0054] Furthermore, in order to map data elements into digital assets, the present invention makes the following definitions:

[0055] EM(Param1, Param2,..., Paramn) represents an evaluation model for evaluating the value of digital assets based on the specific attributes of data elements. Where Param represents the parameters required by the evaluation model. The evaluation model includes regression models, machine learning algorithms, and risk models.

[0056] MD = {Evaluation_Parameter1, Evaluation_Parameter2,..., Evaluation_Parametern} represents a market demand model, and Evaluation_Parameter is a market demand evaluation parameter for evaluating the value of digital assets based on the specific attributes of data elements.

[0057] The process of mapping data elements into data assets is represented as follows:

[0058] DA = EM(DE, MD);

[0059] That is, the process of converting data element DE into digital asset DA using evaluation model EM according to the market demand MD in the market where the data element is located.

[0060] In the second aspect of the present disclosure, a three-dimensional four-layer stereoscopic representation model based on a blockchain-network structure is constructed to describe the digital asset data after assetization from three dimensions: modality, privacy level, and granularity. That is, as Figure 3 shown, map it to a layer in the modality dimension according to the modality type of the digital asset data, map it to a layer in the privacy level according to the privacy degree of the digital asset data, and map it to a layer in the granularity dimension according to the granularity size of the digital asset data.

[0061] Specifically, classify the digital asset data after being elementized and assetized from three dimensions: modality, privacy level, and granularity. The above three dimensions are divided into four categories. The multi-modal dimension is divided into four layers: structured data, picture and video data, text information data, and other types of data; the privacy level dimension is divided into four layers: raw data, desensitized data, modeled data, and artificial intelligence data; the granularity dimension is divided into four layers: data item, set, file, and asset.

[0062] For a specific piece of data, it belongs to three dimensions simultaneously: modality, privacy level, and granularity. Under a specific dimension, it only belongs to a certain layer of that dimension. For example, a certain piece of data belongs to structured data under multi-modal, de-sensitized data under privacy level, and asset granularity data under granularity division.

[0063] Interpret the model from each dimension and each layer under the dimension:

[0064] The data exhibits multi-modality. In this disclosure, the data is classified into structured data, picture and video data, text information data, and other types of data according to modality.

[0065] Specifically, first, perform multi-modal division on the digital assets uploaded to the chain according to the different modalities of the digital assets. The rules for classifying the data after assetization from the modality dimension include: classifying digital assets with clear patterns and format organizations as structured data, such as spreadsheets, database table data, and structured communication transmission data. All fields of the structured data are described in key-value pairs, and the values in the key-value pairs contain sub-key-value pairs with more than one level; classifying digital assets containing large file formats such as pictures, audio, and video as picture and video data; classifying data organized in text form and having a semantic structure as text information data, such as documents, emails, and chat record data; classifying various data forms other than structured data, picture and video data, and text information data as other types of data, such as binary data, sensor data, log files, and geographical location information data.

[0066] Data has different requirements for privacy levels. In this disclosure, data is classified into raw data, desensitized data, modeled data, and artificial intelligence data according to privacy levels. The rules for classifying digital asset data after assetization from the dimension of privacy levels are as follows: Data actively abandoned by the owner of the digital asset for processing is classified as raw data. This type of data has not been processed, edited, or transformed, that is, the raw data directly obtained from the off-chain data source is directly mapped to the digital asset and encrypted and uploaded to the chain through the national cryptographic algorithm. Data required by the owner of the digital asset to be desensitized is classified as desensitized data. According to the desensitization rules preset by the owner of the digital asset, data desensitization is performed before mapping the off-chain data to the digital asset for encryption and uploading to the chain. For example, data with personal identity-related information or specific identifiers removed. Desensitization processing helps protect the privacy of the data, making the data more secure during analysis, sharing, or storage, while still retaining the usefulness of the data. Data required by the owner of the digital asset to be modeled is classified as modeled data. According to the modeling requirements of the asset owner, machine learning or statistical models are used to process the raw data. Models that can be used as modeled data include prediction models, classification models, and clustering models. The data obtained after model processing is mapped to the digital asset for uploading to the chain. More levels of information are obtained through modeling without directly exposing the details of the raw data, reducing the dimension of the data and extracting the key features of the data while protecting the privacy of the raw data. Artificial intelligence data that the owner of the digital asset needs to upload is classified as artificial intelligence data. The artificial intelligence data is used for private deep learning or federated learning after uploading, generally including deep learning models and neural network models. The processed data is mapped to the digital asset for uploading to the chain. This processing enables the owner of the digital asset to utilize the prediction or analysis results of the model without exposing the detailed information of the raw data.

[0067] Since the blockchain is a distributed ledger and the ledger is shared and visible among all consensus nodes, with transparency, it is necessary to further enhance the privacy protection of data. Therefore, this application provides a division of privacy levels. According to the division of different privacy levels, different processing is performed on the data after weighing security and performance: For raw data and desensitized data, in order to prevent the exposure of plaintext, the national cryptographic algorithm specification such as SM4 is used to encrypt the data; for modeled data and artificial intelligence data, considering that they already belong to a relatively high privacy protection level, it is selected whether to use an encryption algorithm for encryption according to requirements.

[0068] The rules for classifying digital assets from the granularity dimension include: Data is divided according to granularity. In this disclosure, data is divided into four categories: data items, sets, files, and assets according to the size of the granularity. Further, for the convenience of understanding, the following definitions are made in this disclosure:

[0069] The data item type in the digital asset is represented by Di, where i represents the subscript of the data item, indicating that multiple data items can be included in the digital asset. Common data items include integer type, floating-point type, boolean type, string type, etc.

[0070] The set in the digital asset is represented by Cj = {D1, D2, D3, ……, Dn}. The elements in the set are composed of the data item Di, where j represents the subscript of the set, indicating that multiple sets can be included in the digital asset.

[0071] The file type in the digital asset is represented by Fk. The file type includes picture and video data and text information data, where k represents the subscript of the file type, indicating that multiple files can be included in the digital asset.

[0072] The digital asset type is represented by Asset. The digital asset is composed of a combination of multiple types such as data items, sets, and files, and is represented as {D1, D2, ……, Dn, C1, C2, ……, Cn, F1, F2, ……, Fn}. When using a digital asset to represent an ID card, it includes data items such as name, gender, ethnicity, and file data photo.

[0073] The present disclosure designs a decentralized element circulation mechanism to meet the decentralized circulation, sharing, or trading of data assets.

[0074] After the description, the digital asset is processed according to the set processing method to achieve privacy protection and data verifiability, and then uploaded to the chain. The circulation of the digital asset is positioned as a transaction including transaction hash, transaction status, transaction belonging block, timestamp, sender address, and receiver address. Transactions form blocks, and blocks are organized into a blockchain through hash pointers. The logical association between transactions constitutes a transaction network; the decentralized circulation of the digital asset is achieved through verifiable transaction operations.

[0075] During the process of achieving privacy protection and data verifiability and uploading the digital asset to the chain after processing the digital asset according to the set processing method, the methods for providing data verifiability include: when off-chain data is uploaded to the chain, first use a cryptographic hashing algorithm to process the data to obtain the digest of the data, and then use a digital signature algorithm to sign the digest of the data, and upload the signature to the chain to achieve data verifiability. An exemplary method is: first use a cryptographic hashing algorithm, such as SM3, to process the data to obtain the digest of the data, and then use a digital signature algorithm, such as SM2, to sign the digest of the data, and upload the signature to the chain to achieve data verifiability.

[0076] The on-chain implementation of privacy protection and data verifiability after processing digital assets according to the set processing method, such as the description in the process of classifying data assets from the privacy dimension, including different processing of data and then on-chain after weighing security and performance according to different privacy levels, including: for the original data and desensitized data, in order to prevent the exposure of plaintext, use the national cryptographic algorithm standard to encrypt the data, and desensitize the desensitized data according to the set desensitization rules before encrypting and on-chain; for modeled data and artificial intelligence data, given that they already belong to a relatively high privacy protection level, choose whether to use an encryption algorithm for encryption according to the requirements.

[0077] The decentralized element circulation mechanism positions the circulation of digital assets as a transaction that includes a transaction hash, transaction status, the block to which the transaction belongs, timestamp, sender address, and recipient address.

[0078] In the specific implementation process, the transaction manifestation is as follows:

[0079] Tx = {Transaction_Hash, Status, Block, Timestamp, From, To};

[0080] Transaction_Hash represents the transaction hash, which is the unique identifier of the transaction. The transaction hash value is generated by hashing operations on other fields in the transaction and description information related to the transaction. Status = {Success, Pending, Failure,...} represents the transaction status, including the Success (successful), Pending (pending), and Failure (failed) statuses. Block represents the block number to which the transaction belongs, that is, the block height. Timestamp represents the timestamp, that is, the time when the block where it is located is generated. From represents the transaction sender address, and To represents the transaction recipient address.

[0081] In the decentralized element circulation mechanism, transactions form blocks, and the blocks are organized into a blockchain through hash pointers, and the logical associations between transactions form a transaction network. A preferred method is to use multi-layer signatures at each level of transactions, blocks, and blockchains to ensure the verifiability of transaction operations; refer to Figure 4 As shown, the methods for verifiable transaction operations include:

[0082] Calculate the hash of each transaction within the block to obtain the transaction hash. Sign the transaction hash of each transaction within the block using the private key, and save the signature in the transaction data as the first-layer signature. Construct a Merkle hash tree using the transaction hashes within the block to obtain the Merkle root. When the transactions are packaged into a block, store the signature of the Merkle root in the block as the second-layer signature. For a blockchain formed by multiple blocks, sign all the blocks in the blockchain as a whole to obtain the blockchain signature, which serves as the third-layer signature. Verify the transaction operations through the three-layer signatures of transactions, blocks, and the blockchain.

[0083] Taking the construction of a Merkle hash tree from transactions to blocks and then signing the three layers of transactions, blocks, and the blockchain as an example, demonstrate the principle and process of transaction verification in this application:

[0084] Suppose there is a block containing n transactions Tx1, Tx2, Tx3, ……Txn. Use the hash algorithm to perform hash operations on these n transactions to obtain n hash values H1, H2, H3, ……, Hn. Sign the transaction hash of each transaction within the block using the private key, and save the signature in the transaction data as the first-layer signature. The sender uses its private key to sign the transaction hash. The signature is generated by applying a digital signature algorithm to the digest. Usually, the digital signature generated using the elliptic curve digital signature algorithm will be stored in the transaction data in the blockchain together with the transaction.

[0085] Construct a Merkle hash tree using the transaction hashes within the block to obtain the Merkle root. When the transactions are packaged into a block, store the signature of the Merkle root in the block as the second-layer signature. The n transaction hash values within the block correspond to the leaf nodes of the Merkle hash tree. Perform a concatenation operation on every two of the n leaf nodes, and perform a hash operation on the concatenated value to obtain the parent node of these two leaf nodes. For example, for H1 and H2, perform a concatenation operation to get H1|H2, then H(H1|H2) is their parent node. Obtain a new layer of parent nodes. Traverse the nodes of each layer to construct the corresponding new layer of parent nodes until there is only one root node, i.e., the Merkle root, in the new layer. When multiple transactions are packaged into a block, the Merkle root is stored in the block. Since the Merkle root value corresponds to each block one by one, the Merkle root is signed as the second-layer signature of the corresponding block.

[0086] For a blockchain formed by multiple blocks, sign all the blocks in the blockchain as a whole, and the obtained result serves as the signature of the blockchain and the third-layer signature of the corresponding blockchain.

[0087] During verification, verifying the transaction hash signature of a certain transaction can verify whether the transaction has been tampered with. Verifying the signature of the Merkle root of a certain block can verify whether there are any tampered transactions in the block. Verifying the signature of a certain blockchain can verify whether there are any tampered blocks or transactions in the blockchain. For example, if you want to verify whether a specific transaction in a block has been tampered with, you only need to provide the Merkle path corresponding to the transaction for verification. Taking a block with four transactions as an example, if you want to verify whether transaction Tx1 in the block has been tampered with, you only need to provide the Merkle path of transaction Tx1: {H(Tx2), H(H(Tx3)|H(Tx4))}. Construct the Merkle root according to the Merkle path and verify the constructed Merkle root with the signature of the Merkle root stored on the blockchain to determine whether the transaction has been tampered with.

[0088] See Figure 5 and Figure 6 As shown, an example of implementing multi-layer verification with digital signatures using the SM2 algorithm is as follows:

[0089] The parameter definitions involved in the process are as follows: d A represents the private key of user A. e represents the output value of the cryptographic hash algorithm applied to message M. e′ represents the output value of the cryptographic hash algorithm applied to message M′. G represents a base point of the elliptic curve, and its order is a prime number. H v () represents a cryptographic hash algorithm with a message digest length of v bits. M represents the message to be signed. M′ represents the message to be verified. N represents the order of the base point G. P A represents the public key of user A. (r, s) represents the signature sent. (r′, s′) represents the received signature. [k]P represents the k-fold point of point P on the elliptic curve. mod represents taking the remainder.

[0090] See Figure 5 and Figure 6 , the process of using the SM2 algorithm to perform digital signatures on transactions, blocks, and blockchains to obtain multi-layer signatures and using multi-layer signatures for verification includes:

[0091] Set Calculate The random number generator generates a random number k ∈ [1, n - 1]; calculate the elliptic curve point (x1, y1) = [k]G. Z A The hash value of the distinguishable identifier of user A, part of the elliptic curve system parameters, and the public key of user A.

[0092] Calculate r = (e + x1) mod n; calculate s = ((1 + d A ) -1 ·(k - r·d A )) mod n.

[0093] If r = 0 or r + k = n or s = 0, then return to recalculate the random number.

[0094] Otherwise, obtain the signature (r, s) of the message M.

[0095] When verifying transactions, blocks, and blockchains, the following process is executed:

[0096] Upon receiving the information M' to be verified and the corresponding (r', s'), check whether r' ∈ [1, n - 1] holds. If it does not hold, the verification fails; check whether s' ∈ [1, n - 1] holds. If it does not hold, the verification fails.

[0097] Set Calculate Calculate t = (r' + s') mod n.

[0098] If t = 0, the verification fails.

[0099] Calculate the elliptic curve point (x'1, y'1) = [s']G + [t]P A , calculate R = (e' + x1') mod n.

[0100] Check whether R = r' holds. If it holds, the verification passes; otherwise, the verification fails.

[0101] Through the above steps, the present disclosure enables the circulated data resources after being factorized and assetized, and performs verifiable processing on transactions to timely detect tampered transactions.

[0102] Embodiment 2

[0103] Refer to Figure 7 As shown, the present invention provides a device for realizing data factorization, assetization, and decentralized circulation, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the data factorization, assetization, and decentralized circulation method is realized, including:

[0104] Under the guidance of the domain knowledge of the data resource field, map the data resource into data elements through corresponding mapping rules; according to the market demand of the market where the data elements are located, use the corresponding evaluation model to convert the data elements into digital assets;

[0105] Describe the digital asset data after assetization from three dimensions of modality, privacy level, and granularity. Among them, from the multi-modal dimension, it is divided into structured data, picture and video data, text information data, and other types of data; from the privacy level dimension, it is divided into raw data, desensitized data, modeled data, and artificial intelligence data; from the granularity dimension, it is divided into data items, sets, files, and assets;

[0106] After the description, the digital assets are processed according to the set processing method to achieve privacy protection and verifiable data for on-chain. The circulation of digital assets is defined as a transaction that includes a transaction hash, transaction status, block to which the transaction belongs, timestamp, sender address, and recipient address. Transactions form blocks, and the blocks are organized into a blockchain through hash pointers. The logical associations between transactions form a transaction network. The decentralized circulation of digital assets is achieved through verifiable transaction operations.

[0107] Of course, the computer program stored in the storage unit of a product recommendation device based on blockchain and machine learning provided in an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in a method for data elementization, assetization, and decentralized circulation provided in any embodiment of the present invention.

[0108] Embodiment 3

[0109] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, which when executed by a processor, implements the method for data elementization, assetization, and decentralized circulation, including: under the guidance of domain knowledge in the field of data resources, mapping data resources into data elements through corresponding mapping rules; converting data elements into digital assets according to market demands in the market where the data elements are located using corresponding evaluation models.

[0110] Describe the digital asset data after assetization from three dimensions: modality, privacy level, and granularity. Among them, from the multi-modal dimension, it is divided into structured data, picture and video data, text information data, and other types of data; from the privacy level dimension, it is divided into raw data, desensitized data, modeled data, and artificial intelligence data; from the granularity dimension, it is divided into data items, sets, files, and assets.

[0111] After the description, the digital assets are processed according to the set processing method to achieve privacy protection and verifiable data for on-chain. The circulation of digital assets is defined as a transaction that includes a transaction hash, transaction status, block to which the transaction belongs, timestamp, sender address, and recipient address. Transactions form blocks, and the blocks are organized into a blockchain through hash pointers. The logical associations between transactions form a transaction network. The decentralized circulation of digital assets is achieved through verifiable transaction operations.

[0112] Of course, for a computer-readable storage medium provided in an embodiment of the present invention, the computer program stored therein is not limited to the method operations described above, and can also execute related operations in a method for data elementization, assetization, and decentralized circulation provided in any embodiment of the present invention.

[0113] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of structures or units can be in electrical, mechanical or other forms.

[0114] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for factorizing, assetizing and decentralizing data circulation, characterized in that: include: According to the guidance of the domain knowledge of the data resource field, the data resource is mapped into data elements through corresponding mapping rules; The data elements are converted into digital assets using corresponding evaluation models according to the market demand of the market. The process includes: S11, Data Resource Value Discovery: Anchoring data resources to subject matter or value measurement attributes through manual or automatic means; analyzing and processing data resources to reveal potential value, insights and opportunities in data resources through subject matter or value measurement attributes; S12, mapping data resources into data elements: generating data elements from data resources in any field according to the guidance of the domain knowledge to which the data resources belong through corresponding mapping rules, including: definition Represents a data resource, where Raw_ Datan Refers to the raw data owned or collected by the data resource; It represents the factorization mapping rules, which are used to convert the original data resources into meaningful data elements. The factorization mapping rules are any one or several data processing steps in data cleaning, feature engineering, dimensionality reduction, classification and clustering. Paramn Indicates the parameters required by the mapping rule; represents domain knowledge; Expertise Refers to domain-specific expertise that guides the process of mapping data resources to data elements, including the experience and understanding of domain experts; represents a data element; Processed_Data Indicates data with certain value after data resource mapping; The process of mapping data resources into data elements is as follows: ; That is, data resources DR in any field are used to generate data elements DE through corresponding mapping rules MR according to the guidance of the domain knowledge DK to which the data resources DR belong; S13, mapping data elements into digital assets using an expression model to make the data elements circulate; mapping data elements according to the definition of the expression model, the expression model includes converting text data into digital codes, mapping attribute values ​​into numerical values, and mapping relationships into connections or weights; including: defining Represents a valuation model, which is used to evaluate the value of digital assets based on specific attributes of data elements; valuation models include regression models, machine learning algorithms, and risk models; Represents a market demand model, which is used to evaluate the value of digital assets based on the specific attributes of data elements; the process of mapping data elements into data assets is represented as follows: ; Upcoming data elements DE The process of converting data elements into digital assets DA using valuation models EM according to the market demand MD of the market in which they are located; The assetized digital asset data is described from three dimensions: modality, privacy level, and granularity. From the multimodal dimension, it is divided into structured data, image data, text information data, and other types of data; from the privacy level dimension, it is divided into raw data, desensitized data, modeled data, and artificial intelligence data; from the granularity dimension, it is divided into data items, collections, files, and assets; any digital asset has the three dimensions of modality, privacy level, and granularity at the same time; under a specific dimension, it only has a certain layer of that dimension; After the description, the digital assets are processed according to the set processing method to achieve privacy protection and data verifiability on the chain, and the circulation of digital assets is positioned as a transaction containing transaction hash, transaction status, block to which the transaction belongs, timestamp, sender address and receiver address. Transactions form blocks, and blocks are organized into blockchains through hash pointers. The logical associations between transactions constitute a transaction network; the decentralized circulation of digital assets is achieved through verifiable transaction operations, where the methods of providing verifiable transaction operations include: performing hash calculations on each transaction in the block to obtain a transaction hash, using a private key to sign the transaction hash of each transaction in the block, and saving the signature in the transaction data as the first-layer signature; using Use the transaction hashes in the block to build a Merkle hash tree to get the Merkle root. When the transaction is packaged into a block, the signature of the Merkle root is stored in the block as the second-layer signature. For a blockchain formed by multiple blocks, all blocks in the blockchain are signed as a whole to get the blockchain signature as the third-layer signature. Transaction operations are verified through three-layer signatures of transactions, blocks, and blockchains. During verification, verifying the transaction hash signature of a transaction can verify whether the transaction has been tampered with, verifying the signature of the Merkle root of a block can verify whether any transaction in the block has been tampered with, and verifying the signature of a blockchain can verify whether any blocks in the blockchain or transactions in the blocks have been tampered with.

2. The data elementization, assetization and decentralized circulation method according to claim 1, characterized in that: The rules for classifying digital asset data after assetization from the modality dimension include: classifying data with clear patterns and formats as structured data; classifying digital assets with large file formats such as pictures, audio, and video as picture image data; classifying digital assets organized in text form and with semantic structure as text information data; and classifying various data forms other than structured data, picture image data, and text information data as other types of data; The rules for classifying digital asset data after assetization from the privacy level dimension include: classifying digital assets that the digital asset owner actively abandons processing as raw data; classifying digital assets that the digital asset owner requires to be desensitized as desensitized data; classifying digital assets that the digital asset owner requires to be modeled as modeled data; and classifying artificial intelligence data that the digital asset owner needs to upload as artificial intelligence data. The rules for classifying digital assets from the granularity dimension include: dividing data into four categories according to the size of the granularity: data items, collections, files, and assets. Among them, collections are composed of data items, files include image data and text information data, and assets are composed of a combination of multiple types of data items, collections, and files.

3. The data elementization, assetization and decentralized circulation method according to claim 1, characterized in that: In the process of achieving privacy protection and data verifiability on the chain after the digital assets are processed according to the set processing method after the description, the methods of providing data verifiability include: when the off-chain data is on the chain, the data is first processed using a cryptographic hash algorithm to obtain a summary of the data, and then a digital signature algorithm is used to sign the summary of the data, and the signature is uploaded to the chain to achieve data verifiability.

4. The data elementization, assetization and decentralized circulation method according to claim 1, characterized in that: After the description, the digital assets are processed according to the set processing method to achieve privacy protection and data verifiability on the chain, including the division according to different privacy levels, and different processing of the data after balancing security and performance before uploading to the chain, including: for original data and desensitized data, the data is encrypted using the national secret algorithm specification, and the desensitized data is desensitized according to the set desensitization rules before being encrypted and uploaded to the chain; for modeled data and artificial intelligence data, encryption without using encryption algorithms is selected according to needs.

5. The method for data elementization, assetization and decentralized circulation according to claim 1, characterized in that: The transaction hash value of the previous block is configured in the block header of each block to form a hash pointer.

6. A device for realizing data elementization, assetization and decentralized circulation, characterized in that: include: At least one processing unit, wherein the processing unit is connected to a storage unit via a bus unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the data elementization, assetization and decentralized circulation method as described in any one of claims 1 to 5 is implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the data elementization, assetization and decentralized circulation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Balanced win-win data asset pricing mechanism

    CN110706018A

  • Data capitalization operation platform authorization approval process method, system and equipment

    CN115330341A