A consensus method and system based on data assets

By generating an initial set of data blocks associated with feature dimensions, and utilizing an improved Bitcoin-NG consensus mechanism and distributed consensus algorithm, the performance bottleneck of blockchain technology in data asset transactions is solved, achieving efficient and secure multi-dimensional data asset consensus.

CN120631983BActive Publication Date: 2025-11-18ZHEJIANG COMM SERVICES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511121701.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing blockchain technology suffers from performance bottlenecks in data asset transactions, making it difficult to balance high throughput, low latency, and adaptability to multi-dimensional data characteristics. Furthermore, existing solutions typically employ a single optimization objective, which can easily lead to fork risks or resource waste.

Method used

By receiving data asset transaction requests, an initial set of data blocks associated with feature dimensions is generated, feature mapping and decomposition are performed, multi-objective optimization is carried out using an improved Bitcoin-NG consensus mechanism, security is ensured by combining a PoW chain, and a distributed consensus algorithm is applied for verification and confirmation, ultimately generating a consensus result for data asset transactions.

Benefits of technology

It improves the efficiency and security of data asset consensus, can dynamically adapt to multiple characteristic requirements, and reduces fork risks and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631983B_ABST
    Figure CN120631983B_ABST
Patent Text Reader

Abstract

The application discloses a consensus method and system based on data assets, and the method comprises the following steps: receiving a data asset transaction request in a distributed network, generating an initial data block set based on the transaction request; performing feature mapping on the initial data block set, constructing a block feature vector set, decomposing the feature vectors in the block feature vector set into a plurality of sub-feature vectors to capture different characteristics of the data assets; performing multi-objective optimization on the sub-feature vectors by using an improved Bitcoin-NG consensus mechanism to generate an optimized data block set; verifying and confirming the optimized data block set by using a distributed consistency algorithm to determine a plurality of final data blocks as the consensus result of the corresponding data asset transaction request. By using the embodiment of the application, the efficiency and security of data asset consensus can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data asset technology, and in particular to a consensus method and system based on data assets. Background Technology

[0002] With the rapid development of blockchain technology, efficient trading and secure consensus of data assets have become core challenges for distributed networks. Traditional consensus mechanisms (such as PoW and PoS) suffer from performance bottlenecks in data asset trading scenarios, struggling to balance high throughput, low latency, and adaptability to multi-dimensional data characteristics. For example, while the Bitcoin-NG protocol improves transaction processing speed through microblocks, its native design does not consider the complex characteristics of data assets (such as ownership, timeliness, and privacy levels), making it difficult to balance consensus efficiency and data security. Furthermore, existing solutions typically employ a single optimization objective, failing to dynamically adapt to the diverse needs of data assets and easily leading to fork risks or resource waste. Summary of the Invention

[0003] The purpose of this invention is to provide a consensus method and system based on data assets to address the shortcomings of existing technologies and improve the efficiency and security of data asset consensus.

[0004] One embodiment of this application provides a consensus method based on data assets, the method comprising:

[0005] Receive data asset transaction requests in a distributed network, and generate a set of initial data blocks based on the transaction requests, wherein each initial data block is associated with a feature dimension of the data asset;

[0006] Feature mapping is performed on the initial data block set to construct a block feature vector set. The feature vectors in the block feature vector set are decomposed into several sub-feature vectors to capture different characteristics of the data assets. Each feature vector represents a multi-dimensional feature attribute of the corresponding initial data block.

[0007] The sub-feature vector is optimized using an improved Bitcoin-NG consensus mechanism to generate an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader through the main chain block and processes transactions through the micro-block, and combines the PoW chain to ensure the security and immutability of data assets.

[0008] The optimized data block set is verified and confirmed using a distributed consensus algorithm to determine several final data blocks as the consensus results for the corresponding data asset transaction requests.

[0009] Optionally, the step of receiving data asset transaction requests in a distributed network and generating an initial set of data blocks based on the transaction requests, wherein each initial data block is associated with a feature dimension of the data asset, including:

[0010] The system receives data asset transaction requests through distributed network nodes, parses the asset type, digital signatures of both parties, and the identifier of the underlying asset in the request, verifies the validity of the signature using elliptic curve cryptography, filters out malicious requests with invalid signatures, and outputs a list of valid transaction requests.

[0011] For each request in the list of valid transaction requests, extract the feature dimensions of the corresponding data asset, including data size, generation time, ownership chain length, encryption level, and circulation frequency. Eliminate the difference in units through feature standardization and output a data asset feature dimension table.

[0012] Based on the data asset feature dimension table, a clustering algorithm is used to group transaction requests according to feature similarity. Each group corresponds to a feature dimension combination. A unique initial block identifier is assigned to each group, and a block-feature dimension mapping table is output.

[0013] Based on the block-feature dimension mapping table, the detailed information of each group of transaction requests is packaged into an initial data block, and the corresponding feature dimension label is embedded in the header of each block to generate an initial data block set.

[0014] Optionally, the initial data block set is feature-mapped to construct a block feature vector set. The feature vectors in the block feature vector set are then decomposed into several sub-feature vectors to capture different characteristics of the data assets. Each feature vector represents a multi-dimensional feature attribute corresponding to the initial data block, including:

[0015] Traverse the initial data block set, extract the specific values ​​corresponding to the feature dimension labels of each block, and combine the derived features of the block, such as the transaction throughput and verification time, to construct the original feature matrix;

[0016] The original feature matrix is ​​processed by minimax normalization to eliminate the influence of different feature dimensions. Multicollinearity features are detected and removed by variance inflation factor, and the redundant normalized feature matrix is ​​output.

[0017] Based on the redundancy-removing normalized feature matrix, a deep learning embedding model is used to map the multidimensional features of each block into a high-dimensional vector. Each dimension of the vector corresponds to the abstract expression of the feature, and the block feature vector set is output.

[0018] Apply the nonnegative matrix factorization algorithm to the block feature vector set, set the decomposition dimension according to the core characteristics of the data asset, decompose each high-dimensional feature vector into multiple sub-feature vectors, each corresponding to the quantitative expression of a single characteristic, and output the initial sub-feature vector set;

[0019] Calculate the mutual information value of the initial sub-feature vector set, remove redundant sub-vectors with mutual information higher than the threshold, verify the independence of the sub-vectors through principal component analysis, and output the optimized sub-feature vector set.

[0020] Optionally, the improved Bitcoin-NG consensus mechanism is used to perform multi-objective optimization on the sub-feature vector to generate an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader through main chain blocks and processes transactions through micro-blocks, and combines a PoW chain to ensure the security and immutability of data assets, including:

[0021] Security sub-vectors are extracted from the sub-feature vector set, and a comprehensive score is constructed by combining the node computing power value. The main chain block leader is elected through a computing power competition, and the leader node identifier and the hash record of the election process are output.

[0022] The leader node generates main chain blocks based on security sub-vectors. The block header contains the root hash of the sub-feature vector set, the leader's public key, and term information. It completes main chain block mining through the PoW mechanism and outputs main chain blocks with PoW proofs.

[0023] The leader node sorts transaction requests according to the liquidity sub-vector, and prioritizes packaging high-liquidity requests into micro-blocks. Each micro-block contains a transaction digest corresponding to the sub-feature vector and the leader's signature. The micro-block sequence is generated according to the transaction order, and the set of unverified micro-blocks is output.

[0024] Network nodes perform PoW chain verification on the unverified microblock set, calculate the hash value of each microblock and compare it with the root hash of the main chain block. Verified microblocks are included in the temporary transaction pool. At the same time, the nodes synchronously update the verification status of the local sub-feature vector and output the pool of verified microblocks.

[0025] Based on the value stability sub-vector, multi-objective optimization weights are set, and the main chain blocks and the verified micro-block pool are merged to select blocks that meet the weight thresholds and integrate them into an optimized data block set according to timestamps.

[0026] Optionally, the step of applying a distributed consensus algorithm to verify and confirm the optimized data block set, and determining several final data blocks as the consensus result of the corresponding data asset transaction request, includes:

[0027] The optimized data block set is broadcast to all consensus nodes in the distributed network. Each node calls the smart contract to execute the transaction logic in the block locally, checks the consistency between the block hash value and the sub-feature vector, and outputs the node's local verification result.

[0028] The practical Byzantine fault-tolerant algorithm is used to collect the local verification results of all nodes. When the percentage of nodes that agree to the verification exceeds 2 / 3, the consensus confirmation process is triggered. The list of agreeing nodes and the reasons for rejection by dissenting nodes for each block are recorded, and the consensus confirmation certificate is output.

[0029] Based on consensus confirmation credentials, the passed blocks undergo final hash chain verification to ensure the continuity and immutability between blocks, filter out abnormal blocks with broken hash chains, and output a sequence of candidate blocks that have passed verification.

[0030] The candidate block sequence is arranged in ascending order by transaction timestamp to generate a final data blockchain containing a complete ownership chain, transaction records, and sub-feature vector labels, which serves as the consensus result for the corresponding data asset transaction request.

[0031] Another embodiment of this application provides a consensus system based on data assets, the system comprising:

[0032] The receiving module is used to receive data asset transaction requests in a distributed network and generate a set of initial data blocks based on the transaction requests, wherein each initial data block is associated with a feature dimension of the data asset.

[0033] The construction module is used to perform feature mapping on the initial data block set, construct a block feature vector set, and decompose the feature vectors in the block feature vector set into several sub-feature vectors to capture different characteristics of the data assets. Each feature vector represents a multi-dimensional feature attribute of the corresponding initial data block.

[0034] The optimization module is used to perform multi-objective optimization on the sub-feature vector using the improved Bitcoin-NG consensus mechanism to generate an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader through the main chain block and processes transactions through the micro-block, and combines the PoW chain to ensure the security and immutability of data assets.

[0035] The verification module is used to apply a distributed consensus algorithm to verify and confirm the optimized data block set, and determine several final data blocks as the consensus results of the corresponding data asset transaction requests.

[0036] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0037] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0038] Compared with existing technologies, this invention provides a consensus method based on data assets, which involves receiving data asset transaction requests in a distributed network, generating an initial set of data blocks based on the transaction requests, performing feature mapping on the initial data block set to construct a block feature vector set, decomposing the feature vectors in the block feature vector set into several sub-feature vectors to capture different characteristics of the data assets, using an improved Bitcoin-NG consensus mechanism to perform multi-objective optimization on the sub-feature vectors to generate an optimized data block set, and applying a distributed consensus algorithm to the optimized data block set for verification and confirmation to determine several final data blocks as the consensus results of the corresponding data asset transaction requests, thereby improving the efficiency and security of data asset consensus. Attached Figure Description

[0039] Figure 1 A hardware structure block diagram of a computer terminal for a consensus method based on data assets, provided in an embodiment of the present invention;

[0040] Figure 2 A flowchart illustrating a consensus method based on data assets provided in an embodiment of the present invention;

[0041] Figure 3 This is a schematic diagram of the structure of a consensus system based on data assets provided in an embodiment of the present invention. Detailed Implementation

[0042] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0043] This invention first provides a consensus method based on data assets, which can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0044] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a consensus method based on data assets, provided as an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0045] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any consensus method based on data assets.

[0046] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0047] Internal memory provides an environment for the execution of computer programs in non-volatile storage media, which, when executed by a processor, enable the processor to execute any consensus method based on data assets.

[0048] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0049] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0050] See Figure 2 The embodiments of the present invention provide a consensus method based on data assets, which may include the following steps:

[0051] S201, Receive a data asset transaction request in a distributed network, and generate a set of initial data blocks based on the transaction request, wherein each initial data block is associated with a feature dimension of the data asset;

[0052] Specifically, data asset transaction requests can be received through distributed network nodes, the asset type, digital signatures of both parties and the identifier of the underlying asset in the request can be parsed, the validity of the signature can be verified using elliptic curve cryptography, malicious requests with invalid signatures can be filtered out, and a list of valid transaction requests can be output.

[0053] A distributed network consists of multiple nodes located in different physical locations. These nodes communicate via a peer-to-peer (P2P) protocol and collectively maintain records of data asset transactions. When a data asset transaction request is initiated, it is broadcast to all nodes in the network, and each node has the capability to receive and initially process the request. For example, a node might receive a data asset transaction request from a corporate user transferring a customer behavior dataset, or from individual users sharing intellectual property data.

[0054] When parsing a transaction request, three core elements need to be extracted: asset type, digital signatures of both parties, and the identifier of the underlying asset. The asset type distinguishes the category of data assets, such as "structured transaction data," "unstructured document," and "encryption algorithm model," etc. Different types of assets have different subsequent processing logics. The digital signatures of both parties are encrypted strings generated by the transaction initiator and receiver using their respective private keys, used to prove the authenticity and non-repudiation of the transaction. The identifier of the underlying asset is a unique identifier for the data asset, such as the hash value "a1b2c3d4...", which allows the specific asset to be located on the network.

[0055] Elliptic Curve Cryptography (ECC) is a key technology for verifying the validity of digital signatures. Based on the mathematical principles of elliptic curves, it offers shorter key lengths than RSA at the same security level (e.g., 256-bit ECC is comparable in security to 3072-bit RSA) and is more computationally efficient. The verification process consists of three steps: First, the public keys of both parties are extracted from the transaction request (the public key is the public key corresponding to the private key and is used to verify the signature); second, the core information in the transaction request (such as asset type, underlying asset identifier, and transaction timestamp) is hashed using SHA-256 to obtain a message digest (e.g., "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"); finally, the digital signature is decrypted using the sender's public key to obtain a decrypted hash value, which is compared with the message digest. If they match, the signature is valid; otherwise, it is invalid.

[0056] For example, user A initiates a data asset transaction with user B, signing the transaction information with their private key to generate a signature "SigA". Upon receiving the request, the node extracts user A's public key "PubA", calculates a hash of the transaction information to obtain "Hash1", and decrypts "SigA" using "PubA" to obtain "Hash2". If "Hash1" and "Hash2" are identical, the signature is valid, confirming that the transaction was initiated by user A and has not been tampered with; if they are inconsistent, it is determined to be a malicious request (possibly due to a forged signature or information tampering) and is filtered out.

[0057] After filtering out malicious requests, the remaining requests form a list of valid transaction requests. Each request in the list contains complete transaction information: asset type (e.g., "medical image dataset"), transaction party IDs (e.g., "UserID1001" and "UserID1002"), underlying asset identifier (e.g., "AssetHash789"), transaction timestamp (e.g., "2025-07-25 15:30:00"), and signature verification result ("valid"). This list provides the foundational data for generating the initial data block, ensuring that all transaction requests entering the system are genuine and legitimate.

[0058] For each request in the list of valid transaction requests, extract the feature dimensions of the corresponding data asset, including data size, generation time, ownership chain length, encryption level, and circulation frequency. Eliminate the difference in units through feature standardization and output a data asset feature dimension table.

[0059] The characteristic dimensions of data assets are key indicators describing their attributes, and these indicators directly affect the generation of subsequent blocks and the consensus process. For each request in the list of valid transaction requests, five core characteristic dimensions need to be extracted one by one: data size, generation time, ownership chain length, encryption level, and circulation frequency.

[0060] Data size refers to the size of a data asset, usually measured in bytes (B), reflecting the storage and processing costs of the asset. For example, a customer's data set might be 5GB (5 × 1024 bytes). 3 B); Generation time is the timestamp of the first creation of the data asset, such as "2024-03-15 09:20:00", used to determine the timeliness of the asset; Ownership chain length refers to the number of times the ownership of the asset has been transferred from its creation to the current transaction. For example, if an asset is transferred from creator A to B, and then to C, the ownership chain length is 2, reflecting the complexity of the asset's circulation history; Encryption level is a rating of the asset's encryption strength, divided into levels 1-5 (level 1 is the weakest, level 5 is the strongest). For example, an asset using AES-256 encryption is level 5; Circulation frequency refers to the number of times the asset has been traded in the past 30 days. For example, a popular dataset has a circulation frequency of 12 times, reflecting the activity level of the asset.

[0061] These feature dimensions have significantly different units (e.g., data size is measured in GB, while circulation frequency is measured in times). Directly using them for analysis would lead to large numerical features such as "data size" dominating the results, while smaller numerical features such as "ownership chain length" would be ignored. Therefore, feature standardization is necessary. The min-max standardization method is used to map all features to the [0,1] interval. The formula is: Standardized value = (Original value - Minimum value) / (Maximum value - Minimum value).

[0062] For example, in a batch of transaction requests, the minimum data size is 1GB and the maximum is 10GB. If the original size of an asset is 5GB, then the standardized value = (5-1) / (10-1) ≈ 0.44. The minimum ownership chain length is 0 and the maximum is 5. If the chain length of an asset is 2, then the standardized value = (2-0) / (5-0) = 0.4. The encryption level itself is 1-5, with a minimum of 1 and a maximum of 5 during standardization. If the asset's level is 3, then the standardized value = (3-1) / (5-1) = 0.5. Through standardization, all features are comparable on the same order of magnitude, avoiding interference from differences in dimensions in subsequent processing.

[0063] The data asset characteristic dimension table integrates the above information. Each record corresponds to a valid transaction request and includes the request ID, asset type, original characteristic values ​​(data size, generation time, etc.), and standardized characteristic values. For example, a record might be: "Request ID 001, Asset type 'Financial Transaction Data', Data size 2GB (standardized 0.11), Generation time '2024-05-20' (standardized 0.6), Ownership chain length 3 (0.6), Encryption level 4 (0.75), Circulation frequency 8 (0.67)". This table clearly presents the characteristic attributes of each data asset, providing a quantitative basis for subsequent clustering and grouping.

[0064] Based on the data asset feature dimension table, a clustering algorithm is used to group transaction requests according to feature similarity. Each group corresponds to a feature dimension combination. A unique initial block identifier is assigned to each group, and a block-feature dimension mapping table is output.

[0065] Clustering algorithms are used to group transaction requests with similar characteristics together, ensuring high consistency in data asset characteristics within each group, which facilitates the subsequent generation of targeted initial data blocks. A commonly used clustering algorithm is K-means, whose core principle is to divide the data into K clusters, ensuring high similarity within clusters and low similarity between clusters.

[0066] The determination of the K value (number of clusters) needs to consider the diversity of data asset characteristics and can be selected using the elbow method: calculate the sum of squared errors (SSE) within clusters for different K values. SSE decreases as K increases. When K increases to a certain value, the rate of decrease in SSE drops sharply; this K value is the optimal one. For example, when K=3, SSE=100; when K=4, SSE=80; when K=5, SSE=78; when K=6, SSE=77. The elbow is at K=4, so K=4 is chosen, dividing the transaction requests into 4 groups.

[0067] Feature similarity is calculated using Euclidean distance. For the standardized feature vectors of two transaction requests (e.g., vector A=[0.44,0.6,0.6,0.75,0.67], vector B=[0.55,0.5,0.7,0.8,0.7]), the Euclidean distance formula is: The smaller the distance, the higher the similarity. The K-means algorithm updates the cluster centers (the mean of the feature vectors of each cluster) iteratively until the cluster centers stabilize (the change in the cluster centers between two iterations is less than 0.01), and finally obtains 4 clusters, each cluster representing a combination of feature dimensions.

[0068] For example, Cluster 1 is characterized by "small data size (normalization 0.1-0.3), short ownership chain (0.1-0.4), and high encryption level (0.6-0.9)," and includes small-value encrypted asset transactions in the financial sector; Cluster 2 is characterized by "large data size (0.7-0.9) and high circulation frequency (0.7-1.0)," and includes large-scale dataset transactions with high frequency of circulation. The combination of feature dimensions for each group is represented by the feature vector at the cluster center. For example, the center vector of Cluster 1 is [0.2, 0.5, 0.3, 0.8, 0.4], representing the typical characteristics of this group.

[0069] Each group is assigned a unique initial block identifier in the format "BlockInit-XXXXX", where "XXXXX" is a 5-digit number (00001-99999) to ensure uniqueness. For example, cluster 1 is assigned "BlockInit-00001", cluster 2 is assigned "BlockInit-00002", and so on. Each initial block identifier corresponds one-to-one with a feature dimension combination within the group, and this is recorded in the block-feature dimension mapping table.

[0070] The block-feature dimension mapping table contains the initial block identifier, cluster center feature vector, feature dimension combination description, and a list of included request IDs. For example, a record might be: "BlockInit-00001, cluster center [0.2,0.5,0.3,0.8,0.4], feature dimension combination 'small scale, short chain, high encryption', request ID list [ID001,ID003,ID005]". This table establishes the association between the initial block and the feature dimensions of the data asset, clarifying the transaction request range and feature attributes corresponding to each block, and providing a grouping basis for subsequent packaging of initial data blocks.

[0071] Based on the block-feature dimension mapping table, the detailed information of each group of transaction requests is packaged into an initial data block, and the corresponding feature dimension label is embedded in the header of each block to generate an initial data block set.

[0072] The initial data block packaging must strictly adhere to the block-feature dimension mapping table, ensuring that each block contains only transaction requests within its corresponding group, and that its feature dimensions are explicitly marked in the block header. The block structure consists of two parts: a block header and a block body. The block header contains metadata, and the block body contains transaction details.

[0073] The fields in the block header include: initial block identifier (e.g., "BlockInit-00001"), feature dimension label (a string converted from the cluster center feature vector, e.g., "small scale_short chain_high encryption"), preceding block hash ("000000" if it is the first block), timestamp (block creation time, e.g., "2025-07-25 16:00:00"), number of transactions (number of transaction requests within the group, e.g., 5), and Merkle root (the hash value of all transaction hashes in the block body, used for quick verification of transaction integrity).

[0074] The block body contains detailed information on all transaction requests within the group. Each transaction record includes: transaction ID (consistent with the ID in the list of valid transaction requests), the IDs of both parties, the identifier of the underlying asset, the transaction amount (if any), digital signature, and transaction status ("Pending Confirmation"). For example, the block body of BlockInit-00001 contains three transaction records: ID001, ID003, and ID005, each detailing the corresponding asset transaction.

[0075] The embedding of feature dimension labels involves converting the "feature dimension combination description" in the block-feature dimension mapping table into machine-recognizable labels. These labels are in key-value pair format, such as "data size = small, ownership chain = short, encryption level = high," facilitating the rapid extraction of block feature attributes during subsequent feature mapping. The labels are written to specific fields in the block header and bound to the initial block identifier, ensuring that the feature dimensions of each block can be directly read.

[0076] Integrity verification is required during the packaging process, and the Merkle root of the block body is calculated: First, the SHA-256 hash of each transaction record is calculated to obtain a list of transaction hashes; then, the hash list is paired up in pairs, and the hash of each pair is calculated. This process is repeated until a root hash, i.e., the Merkle root, is obtained. For example, if the hashes of 3 transactions are H1, H2, and H3, first calculate H12 = hash(H1 + H2), then calculate the Merkle root = hash(H12 + H3), and write it into the block header to ensure that the transaction information in the block body has not been tampered with (any modification to a transaction will cause the Merkle root to change).

[0077] The initial data block set is a collection of all packaged initial blocks. Each block conforms to the structure described above and is associated with a feature dimension of the data asset. For example, the set contains four blocks, BlockInit-00001 to BlockInit-00004, corresponding to four sets of transaction requests. Each block header has a clear feature dimension label, and the block body contains complete transaction details. This set provides the basic data unit for subsequent feature mapping and consensus processing, ensuring that each block accurately reflects the characteristics of its corresponding data asset.

[0078] S202, perform feature mapping on the initial data block set, construct a block feature vector set, decompose the feature vectors in the block feature vector set into several sub-feature vectors to capture different characteristics of the data assets, wherein each feature vector represents the multi-dimensional feature attribute of the corresponding initial data block;

[0079] Specifically, the initial data block set can be traversed, and the specific values ​​corresponding to the feature dimension labels of each block can be extracted. Combined with the derived features of the block, such as the transaction throughput and verification time, the original feature matrix can be constructed.

[0080] The initial data block set contains multiple blocks associated with the feature dimensions of the data asset. The traversal process requires parsing the feature dimension labels and transaction details of each block to extract the basic and derived features needed to construct the original feature matrix. Each initial data block header embeds feature dimension labels, such as "data size = 0.3, generation time = 0.6, ownership chain length = 0.5, encryption level = 0.8, circulation frequency = 0.4". The specific values ​​corresponding to these labels are the standardized feature values ​​from the block-feature dimension mapping table, directly reflecting the core attributes of the data asset.

[0081] When extracting specific values ​​from feature dimension labels, they must precisely match the label fields in the block header. For example, if the feature dimension labels for BlockInit-0001 are "Data size = 0.2, Generation time = 0.5, Ownership chain length = 0.3, Encryption level = 0.9, Circulation frequency = 0.2", then the extracted values ​​will be [0.2, 0.5, 0.3, 0.9, 0.2], corresponding to the five basic features: data size, generation time, ownership chain length, encryption level, and circulation frequency, respectively. These values ​​are rounded to two decimal places to ensure precision while avoiding data redundancy.

[0082] The extraction of derived features requires calculation based on the transaction behavior data of the block. Core derived features include transaction throughput and verification time. Transaction throughput refers to the number of transactions processed by a block per unit of time, calculated as "Transaction throughput = Total number of transactions in the block / Block generation time interval," where the block generation time interval is the difference in timestamps between the current block and the previous block (in seconds). For example, BlockInit-0001 contains 8 transactions, with a generation timestamp of 1627230000, and the previous block timestamp is 1627229980, with an interval of 20 seconds. Therefore, the transaction throughput = 8 / 20 = 0.4 transactions / second. Verification time refers to the average time from when a block is received by a network node to when the initial verification is completed. It is calculated by averaging the verification times of 5 random nodes. For example, if node A takes 0.8 seconds, node B takes 1.2 seconds, node C takes 0.9 seconds, node D takes 1.1 seconds, and node E takes 1.0 seconds, then the verification time = (0.8 + 1.2 + 0.9 + 1.1 + 1.0) / 5 = 1.0 seconds.

[0083] The original feature matrix is ​​a two-dimensional matrix that integrates basic and derived features. The number of rows in the matrix equals the number of initial data blocks, and the number of columns equals the total number of features (5 basic features + 2 derived features = 7). For example, if the initial data block set contains 3 blocks, then the original feature matrix will be 3 rows and 7 columns, with each row corresponding to the feature vector of one block.

[0084] The first line (BlockInit-0001) contains [0.2, 0.5, 0.3, 0.9, 0.2, 0.4, 1.0].

[0085] The second line (BlockInit-0002) contains [0.7, 0.3, 0.8, 0.5, 0.9, 1.2, 0.5].

[0086] The third row (BlockInit-0003) contains [0.5, 0.7, 0.6, 0.7, 0.6, 0.8, 0.8]. Each element in the matrix is ​​a specific numerical value, clearly showing the performance of each block across the seven feature dimensions, providing the raw data foundation for subsequent feature processing.

[0087] When constructing the original feature matrix, a data integrity check is required to ensure that all seven features of each block are complete. If the "verification time" of a block is not collected due to node failure, it is filled with the average verification time of blocks in the same group (e.g., if the average verification time of the group containing BlockInit-0004 is 0.9 seconds, then it is filled with 0.9 seconds). Meanwhile, outliers are handled. For example, if a block's transaction throughput is 5.0 transactions / second (far exceeding the group average of 1.0 transactions / second), it is considered an outlier and replaced with the 95th percentile value of the group (e.g., 1.5 transactions / second) to avoid interference from outliers in subsequent feature processing.

[0088] The original feature matrix is ​​processed by minimax normalization to eliminate the influence of different feature dimensions. Multicollinearity features are detected and removed by variance inflation factor, and the redundant normalized feature matrix is ​​output.

[0089] Although the features in the original feature matrix have been initially standardized, the value ranges of different features still differ (e.g., transaction throughput ranges from 0.4 to 1.2, and verification time ranges from 0.5 to 1.0). These differences in scale may lead to an imbalance in feature weights (e.g., features with large numerical ranges have an excessively high weighting in the calculation). Max-min normalization maps all features to the [0,1] interval, completely eliminating the influence of scale. The formula is: Normalized value = (Original value - Minimum feature value) / (Maximum feature value - Minimum feature value), where the minimum and maximum feature values ​​are the global minimum and maximum values ​​of that feature in the original feature matrix.

[0090] For example, in the original feature matrix, the minimum value of "transaction throughput" is 0.4 and the maximum value is 1.2. The original value of a certain block is 0.8. Therefore, the normalized value = (0.8-0.4) / (1.2-0.4) = 0.4 / 0.8 = 0.5. Similarly, the minimum value of "verification time" is 0.5 and the maximum value is 1.0. The original value of a certain block is 0.8. Therefore, the normalized value = (0.8-0.5) / (1.0-0.5) = 0.3 / 0.5 = 0.6. Through this processing, the values ​​of all features are compressed to [0,1], ensuring that each feature has an equal weight in subsequent analysis.

[0091] After normalization, it is necessary to detect multicollinearity among features, i.e., multiple features are highly correlated (such as "circulation frequency" and "transaction throughput" which may be highly correlated because they both reflect asset activity). This leads to a decrease in the model's explanatory power for the features and unstable parameter estimation. Variance inflation factor (VIF) is a commonly used indicator to measure multicollinearity, and its calculation formula is: VIF_i = 1 / (1-R_i) 2 ), where R_i 2 R_i is the coefficient of determination obtained by performing linear regression of the i-th feature on all other features. 2 The closer to 1, the stronger the collinearity of this feature with other features, and the larger VIF_i is.

[0092] When calculating VIF, a linear regression needs to be performed on each feature: with "circulation frequency" as the dependent variable and the other 6 features as independent variables, the regression equation is fitted using the least squares method to obtain R. 2 =0.9, then VIF=1 / (1-0.9)=10; with "encryption level" as the dependent variable and other characteristics as independent variables, we get R. 2 =0.3, then VIF=1 / (1-0.3)≈1.43. The VIF threshold is usually set to 10. When the VIF of a feature > 10, it is considered to have severe multicollinearity and needs to be removed from the feature matrix.

[0093] For example, the VIF of "Circulation Frequency" was calculated to be 12 (>10), and its correlation coefficient with "Transaction Throughput" was 0.85 (highly correlated), so "Circulation Frequency" was removed; the VIFs of other features were all <10 (such as "Data Size" VIF=2.1, "Generation Time" VIF=3.5) and were retained. After removing redundant features, the original feature matrix was reduced from 7 columns to 6 columns, including "Data Size, Generation Time, Ownership Chain Length, Encryption Level, Transaction Throughput, and Verification Time".

[0094] The redundancy-reduced normalized feature matrix is ​​a matrix that has undergone normalization and decollinearity removal processing. For example, the original 3x7 matrix becomes 3x6: the first row [0.2,0.5,0.3,0.9,0.5,0.6], the second row [0.7,0.3,0.8,0.5,1.0,0.0], and the third row [0.5,0.7,0.6,0.7,0.75,0.6]. The features in this matrix eliminate both dimensional differences and multicollinearity interference, providing high-quality input data for subsequent feature mapping.

[0095] To ensure the rationality of redundancy removal, the variance explained by the feature before and after feature removal needs to be calculated (by principal component analysis). If the total variance explained decreases by no more than 5% after removal (e.g., from 90% to 88%), it indicates that the removed feature is a redundant feature and the processing is effective. If the decrease exceeds 10%, the VIF threshold needs to be re-evaluated or other features need to be removed to balance feature simplification and information retention.

[0096] Based on the redundancy-removing normalized feature matrix, a deep learning embedding model is used to map the multidimensional features of each block into a high-dimensional vector. Each dimension of the vector corresponds to the abstract expression of the feature, and the block feature vector set is output.

[0097] Although the features in the redundancy-removed normalized feature matrix have been optimized, they are still low-dimensional explicit features (6 dimensions), making it difficult to capture the complex nonlinear relationships between features (such as the implicit correlation between "encryption level" and "verification time"). Deep learning embedding models can map these low-dimensional features into high-dimensional vectors (such as 256 dimensions), where each dimension represents an abstract combination of features, thus more comprehensively characterizing the feature attributes of the block.

[0098] The commonly used deep learning embedding model is the autoencoder, whose structure includes an input layer, hidden layers (encoding layer), a bottleneck layer (embedding layer), hidden layers (decoding layer), and an output layer. The number of neurons in the input layer is equal to the dimension of the deduplicated features (6 neurons); the encoding layer compresses features step by step through two hidden layers (e.g., 128 neurons or 64 neurons); the bottleneck layer is the embedding layer, with the number of neurons equal to the dimension of a high-dimensional vector (256), and the output is the high-dimensional embedding vector of the block; the decoding layer reconstructs the input features through a structure symmetrical to the encoding layer (64 neurons or 128 neurons), and the output layer has the same dimension as the input layer (6 neurons).

[0099] The goal of model training is to minimize the reconstruction error between the decoding layer output and the input layer features. The mean squared error (MSE) is used as the loss function: Loss = Σ(output features - input features) 2 / Number of samples. The training data is a redundancy-removed normalized feature matrix. Stochastic gradient descent is used to optimize the parameters, with a learning rate of 0.001 and 1000 iterations. Training stops when the loss function changes less than 1e-5 for 100 consecutive iterations. For example, the input features of a certain block are [0.2,0.5,0.3,0.9,0.5,0.6]. After model encoding, the embedding layer outputs a 256-dimensional vector [0.12,0.35,...,0.28] (intermediate dimensions omitted). The decoding layer reconstructs the features as [0.21,0.49,0.32,0.89,0.51,0.59], with MSE=0.0001, indicating good reconstruction results.

[0100] Each dimension of a high-dimensional vector has no explicit physical meaning, but rather is an abstract expression of the complex relationships between features. For example, the 10th dimension may comprehensively reflect the feature combination of "high encryption level and long verification time", while the 50th dimension may correspond to the pattern of "large data scale and large transaction throughput". These abstract dimensions can capture implicit correlations that low-dimensional features cannot reflect (such as the non-linear relationship between "ownership chain length" and "transaction throughput").

[0101] The block feature vector set is a collection of high-dimensional embedding vectors from all blocks. For example, in a set containing 3 blocks, each block corresponds to a 256-dimensional vector, with each element value in the range [0,1] (since the input features have been normalized, the embedding layer uses the sigmoid activation function). The vectors in this set retain the key information of the original features, while enhancing the discriminative power of the features through mapping in high-dimensional space (blocks of different types are farther apart in high-dimensional space). For example, the cosine similarity of the embedding vectors of financial data blocks and medical data blocks in high-dimensional space is 0.2 (low similarity), while the similarity of blocks of the same type is 0.8 (high similarity).

[0102] To evaluate the embedding effect, the intra-class similarity and inter-class similarity of the high-dimensional vectors can be calculated: intra-class similarity is the average vector similarity of blocks in the same feature group (e.g., 0.75), and inter-class similarity is the average vector similarity of blocks in different groups (e.g., 0.3). When the intra-class similarity is significantly higher than the inter-class similarity (difference > 0.4), it indicates that the embedding model has effectively captured the feature differences of the blocks, and a set of block feature vectors can be output for subsequent processing.

[0103] Apply the nonnegative matrix factorization algorithm to the block feature vector set, set the decomposition dimension according to the core characteristics of the data asset, decompose each high-dimensional feature vector into multiple sub-feature vectors, each corresponding to the quantitative expression of a single characteristic, and output the initial sub-feature vector set;

[0104] While high-dimensional vectors (e.g., 256-dimensional) in the block feature vector set can comprehensively characterize block features, their excessive dimensionality and abstractness make it difficult to directly correlate them with the core characteristics of data assets (such as security, liquidity, and value stability). Non-negative matrix factorization (NMF) can decompose high-dimensional vectors into multiple low-dimensional sub-feature vectors, each corresponding to a quantified expression of a core characteristic. The principle is to decompose a high-dimensional matrix V (of shape n×m, where n is the number of blocks and m is the dimension of the high-dimensional vector) into two non-negative matrices W (n×k) and H (k×m), such that V≈W×H, where k is the decomposition dimension (the number of core characteristics), and each row of W represents the k sub-feature vectors of a block.

[0105] The core characteristics of data assets typically include security, liquidity, and value stability; therefore, the decomposition dimension k is set to 3. Security reflects the tamper-proof and encryption protection capabilities of data asset transactions, and is related to "encryption level" and "ownership chain length." Liquidity reflects the trading activity and circulation efficiency of the asset, and is related to "transaction throughput" and "verification time." Value stability reflects the degree of fluctuation in asset value, and is related to "generation time" and "data scale." The setting of the decomposition dimension k needs to be combined with the business scenario. If cross-border transactions are involved, the "compliance" characteristic can be added, setting k to 4.

[0106] The iterative process of the NMF algorithm is as follows: Initialize W and H as non-negative random matrices (element values ​​are in [0,1]); calculate the reconstruction error ||VW×H|| 2 (Frobenius norm); update W and H using the multiplicative update rule (the update formula for W is W_ij = W_ij × (V×H^T)_ij / (W×H×H^T)_ij, and the update formula for H is similar); repeat the iteration until the reconstruction error is less than the threshold (e.g., 1e-3) or the maximum number of iterations (e.g., 1000 times) is reached.

[0107] For example, if the matrix V of the block feature vector set is 3×256 and the decomposition dimension k=3, then W is 3×3 (each row represents 3 sub-feature vectors of a block), and H is 3×256 (each column represents the weight of the core feature on the higher-dimensional vector dimension). After decomposition, the first column of W corresponds to the security sub-vector (higher values ​​indicate stronger security), the second column corresponds to the liquidity sub-vector, and the third column corresponds to the value stability sub-vector. The sub-feature vector of a certain financial block is [0.85, 0.6, 0.7], indicating that it has high security (0.85), medium liquidity (0.6), and relatively high value stability (0.7).

[0108] Each sub-feature vector has a value range of [0,1] (because NMF requires the matrix to be non-negative), and the higher the value, the stronger the corresponding characteristic. For example, a security sub-vector value of 0.9 indicates that the asset transactions in this block have undergone high-strength encryption (high encryption level) and the ownership chain is clear (tamper-proof); a liquidity sub-vector value of 0.2 indicates low transaction throughput and long verification time (inactive circulation); and a value stability sub-vector value of 0.5 indicates moderate asset value fluctuation (relatively recent generation time and stable data scale).

[0109] The initial sub-feature vector set is a collection of sub-feature vectors from all blocks. For example, in a set of 3 blocks, each block contains 3 sub-vectors (security, liquidity, and value stability). The set format is one row per block, with each row containing 3 values ​​(e.g., [0.85, 0.6, 0.7], [0.7, 0.9, 0.5], [0.6, 0.7, 0.8]). This set transforms high-dimensional abstract features into quantitative indicators directly related to core characteristics, facilitating subsequent multi-objective optimization (e.g., leader election based on security sub-vectors, and transaction ranking based on liquidity sub-vectors).

[0110] The evaluation of the decomposition effect is achieved through reconstruction error and feature correlation: the reconstruction error must be less than a preset threshold (e.g., 1e-3) to ensure that the decomposition retains the key information of the high-dimensional vector; feature correlation refers to the correlation coefficient between the sub-feature vector and the corresponding core feature index (e.g., the correlation coefficient between the security sub-vector and the encryption level > 0.7). When the correlation coefficient between all sub-vectors and their corresponding features is > 0.6, it indicates that the decomposition is effective and the initial sub-feature vector set can be output.

[0111] Calculate the mutual information value of the initial sub-feature vector set, remove redundant sub-vectors with mutual information higher than the threshold, verify the independence of the sub-vectors through principal component analysis, and output the optimized sub-feature vector set.

[0112] The initial sub-feature vector set may contain redundant sub-vectors (security, liquidity, value stability), meaning there is a strong dependency between different sub-vectors (e.g., the sub-vectors of liquidity and value stability are highly correlated). This can lead to repeated calculation of feature weights in subsequent optimization processes. The mutual information value can measure the degree of dependency between two sub-vectors. The higher the value, the stronger the redundancy. The calculation formula is: I (X;Y)=ΣΣp (x,y) log (p (x,y) / (p (x) p (y))), where p(x,y) is the joint probability distribution of sub-vectors X and Y, and p (x) and p (y) are the marginal probability distributions. The mutual information value ranges from [0,log (min (|X|,|Y|))].

[0113] Calculate the mutual information value of each pair of sub-vectors in the initial sub-feature vector set. For example, the mutual information I(X;Y) between security (X) and liquidity (Y) is 0.3, the mutual information I(Y;Z) between liquidity (Y) and value stability (Z) is 0.6, and the mutual information I(X;Z) between security and value stability is 0.2. Set the mutual information threshold to 0.5 (the threshold is set according to the business's requirements for independence; the higher the requirements, the lower the threshold). Then, I(Y;Z) = 0.6 > 0.5, indicating that the liquidity and value stability sub-vectors are redundant.

[0114] The removal of redundant subvectors should be considered in conjunction with business priorities, retaining subvectors that are more important to the core objectives. For example, in data asset consensus, if liquidity has a higher priority than value stability, then the value stability subvector should be removed, while security and liquidity should be retained. If the priorities of the two are equal, the correlation between the subvector and the objective function (such as the correlation coefficient with consensus efficiency) can be calculated, and the subvector with higher correlation should be retained (e.g., if the correlation coefficient of the liquidity subvector is 0.7 > 0.5 of the value stability subvector, then the value stability subvector should be removed).

[0115] After removing the value stability sub-vector, the remaining safety and liquidity sub-vectors need to have their independence verified using Principal Component Analysis (PCA). PCA projects the sub-vectors onto a new coordinate system, making the covariance between the principal components zero (complete independence). The variance explained by the two projected principal components is calculated. If the first principal component explains 90% of the variance of the safety sub-vector, and the second principal component explains 85% of the variance of the liquidity sub-vector, and their covariance is 0.02 (close to zero), it indicates good independence of the sub-vectors and no significant redundancy.

[0116] The optimized sub-feature vector set is a set of sub-vectors after removing redundancy, such as retaining security and liquidity sub-vectors, with the set being [[0.85,0.6], [0.7,0.9], [0.6,0.7]]. The dimension of each sub-vector is determined according to the number of remaining core features (e.g., 2-dimensional), and the mutual information value between sub-vectors is < 0.5, and the covariance verified by PCA is < 0.1, ensuring that each sub-vector corresponds to an independent core feature, which can be used for subsequent multi-objective optimization of the improved Bitcoin-NG consensus mechanism.

[0117] To ensure the rationality of the optimization, it is necessary to re-evaluate the information retention rate of the sub-vectors to the original high-dimensional features (by comparing the reconstruction error through NMF, the increase in reconstruction error after removing redundancy should not exceed 10%), and verify the performance of the sub-vectors in downstream tasks (such as improving the efficiency of consensus based on the optimized sub-vectors by 15%). When both the information retention rate and task performance meet the standards, the optimized sub-feature vector set is output.

[0118] S203, the sub-feature vector is optimized using an improved Bitcoin-NG consensus mechanism to generate an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader through the main chain block and processes transactions through the micro-block, and combines the PoW chain to ensure the security and immutability of data assets.

[0119] Specifically, security sub-vectors can be extracted from the sub-feature vector set, combined with node computing power values ​​to construct a comprehensive score, elect the main chain block leader through computing power competition, and output the leader node identifier and hash record of the election process;

[0120] The sub-feature vector set contains multiple sub-vectors reflecting the core characteristics of data assets. Among them, the security sub-vector is a key indicator for measuring the tamper-proof capability and encryption protection level of data asset transactions. The extraction of the security sub-vector needs to be related to features such as the encryption level and ownership chain length of the data asset. For example, in a certain sub-feature vector set, the value range of the security sub-vector is [0,1], with higher values ​​indicating stronger security. Specifically, the security sub-vector can be calculated by "normalized encryption level value × 0.5 + normalized ownership chain length value × 0.3 + reverse value of historical tampering record × 0.2", where the reverse value of historical tampering record is "1 - number of tamperings / total number of transactions", ensuring that security is reflected from multiple dimensions. For example, if a data asset has an encryption level of 0.9, an ownership chain length of 0.8, and 0 historical tampering records (reverse value 1.0), then its security subvector value is 0.9×0.5 + 0.8×0.3 + 1.0×0.2 = 0.45 + 0.24 + 0.2 = 0.89, which is considered high security.

[0121] Node computing power is a metric that measures the computing ability of network nodes. It is usually expressed as hash rate (the number of hash calculations completed per second), with the unit being H / s (hashes per second). Commonly used units include TH / s (terahashes per second, 1 TH / s = 10^65). 12 H / s), PH / s (petashes per second, 1PH / s = 10 15 (H / s). For example, node A has a computing power of 5TH / s, node B has 3TH / s, and node C has 8TH / s. The higher the computing power, the stronger the node's ability to handle cryptographic computation and verification tasks, and the more suitable it is as the leader of the main chain block.

[0122] The construction of the comprehensive score needs to balance the security sub-vector and the node computing power value. The calculation formula is "Comprehensive Score = Security Sub-vector Value × 0.6 + Node Computing Power Standardized Value × 0.4", where the node computing power standardized value is "Node Computing Power Value / Total Network Computing Power Value", ensuring a reasonable computing power ratio. For example, if node C has a security sub-vector value of 0.89 and a total network computing power of 20TH / s, its computing power standardized value is 8 / 20 = 0.4, then the comprehensive score = 0.89 × 0.6 + 0.4 × 0.4 = 0.534 + 0.16 = 0.694; the comprehensive score of node A = 0.8 × 0.6 + (5 / 20) × 0.4 = 0.48 + 0.1 = 0.58. Node C has a higher score and is more competitive.

[0123] The computing power competition is the core of the leader election process. Essentially, nodes compete for the right to package blocks by solving cryptographic puzzles, following the rule of "the first to solve the puzzle wins." The puzzle requires nodes to calculate a random number (Nonce) such that the hash value of the block header information (including the security subvector, node public key, timestamp, etc.) and the Nonce satisfies a preset condition (e.g., a prefix of 18 consecutive zeros). For example, if the block header information is "security subvector = 0.89, public key = PubC, timestamp = 1627230000", node C, by exhaustively searching for the Nonce, calculates the SHA-256 hash when Nonce = 123456, obtaining "000000000000000000a1b2c3d4e5f6...", which satisfies the condition of a prefix of 18 zeros, thus successfully solving the puzzle.

[0124] The hash records of the election process are used to trace the fairness of the election. The records include the public keys of all participating nodes, the submitted nonces, the calculated hash values, and the timestamps of the solutions. Each record is concatenated into a Merkle chain using SHA-256 hashes, ultimately generating the root hash of the election process. For example, the root hash might be “f0e4c2f76c58916ec258f246851be44c55a94f567c0e78e5a0c6f44c8d4aa4”, stored in a distributed network. Any node can verify whether cheating (such as tampering with the solution time) occurred during the competition.

[0125] The leader node is identified by a unique identifier of the winning node (e.g., "NodeC-789"), which is bound to the election root hash and broadcast to the entire network. If no node solves the puzzle within a preset time (e.g., 10 minutes), the puzzle difficulty is increased (the number of prefix zeros is reduced to 17), and the competition is restarted to ensure that the leader election is conducted efficiently.

[0126] The leader node generates main chain blocks based on security sub-vectors. The block header contains the root hash of the sub-feature vector set, the leader's public key, and term information. It completes main chain block mining through the PoW mechanism and outputs main chain blocks with PoW proofs.

[0127] The process of leader nodes generating main chain blocks must be based on the security subvector to ensure that the block content meets the security requirements of data assets. The structure of the main chain block is divided into a block header and a block body. The block body contains the core transaction records that have been screened (transactions that are highly correlated with the security subvector, such as transactions with a cryptographic level ≥ 0.8), while the block header centrally stores metadata and is the key carrier of the PoW mechanism.

[0128] The root hash of the sub-feature vector set is a core field in the block header, obtained by calculating the Merkle root of all sub-feature vectors. For example, if the sub-feature vector set contains 3 vectors [S1=0.89, S2=0.75, S3=0.92], first calculate the hash H1=hash(S1), H2=hash(S2), and H3=hash(S3) for each vector, then calculate... Root hash = The final root hash is “e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855”, and the embedded block header is used to verify the integrity of the subvector set.

[0129] The leader's public key (e.g., "PubC") is the public key of the leader node, used to prove the legitimacy of blocks. Any node can verify whether a block was generated by the elected leader using the public key. The term information specifies the validity period of the leader's permissions, typically 30 minutes. During their term, the leader has the right to package micro-blocks. A new election is required after the term ends to prevent a single node from monopolizing the market for an extended period. For example, term information like "start = 1627230000, end = 1627231800" clearly defines the leader's responsibilities and time frame.

[0130] The core of the Proof-of-Work (PoW) mechanism for main chain block mining is ensuring the immutability of blocks through computational power investment. The mining process is similar to the computational power competition for leader election, but more difficult (e.g., the prefix contains 20 zeros). After generating the block body, the leader node combines the block header information (root hash, public key, term, previous block hash, etc.) with a random number (Nonce) to calculate the hash value until the target condition is met. For example, if the previous block hash is "a1b2c3d4...", the leader node performs trillions of hash calculations, and when Nonce = 987654, it obtains the target hash "00000000000000000000f1e2d3c4b5a6...", thus completing the PoW proof.

[0131] A Proof of Work (PoW) includes the target hash, the Nonce used, and the computation time (e.g., 2 minutes), which, together with the block header, forms a complete main chain block. For example, a main chain block with a PoW proof might have the following header: "Root hash = e3b0c442..., Public key = PubC, Term = 30 minutes, Precedence hash = a1b2c3d4..., PoW proof = (Nonce = 987654, Target hash = 000000000000000000000f1e2d3..., Time = 120s)", and the block body would contain 50 high-security transaction records.

[0132] During mining, the PoW difficulty is dynamically adjusted based on the network's total hashrate, changing every 2016 blocks to ensure the average mining time remains stable at around 10 minutes. For example, if the average mining time for the most recent 2016 blocks is 8 minutes (too high hashrate), the difficulty is increased (by adding prefix zeros to 21); if the average time is 12 minutes (too low hashrate), the difficulty is decreased (to 19) to maintain network stability.

[0133] The leader node sorts transaction requests according to the liquidity sub-vector, and prioritizes packaging high-liquidity requests into micro-blocks. Each micro-block contains a transaction digest corresponding to the sub-feature vector and the leader's signature. The micro-block sequence is generated according to the transaction order, and the set of unverified micro-blocks is output.

[0134] The liquidity subvector reflects the activity and processing efficiency of data asset transactions. Its core parameters include transaction throughput (number of transactions per unit time), verification time (transaction confirmation speed), and circulation frequency (number of historical transactions). After standardization, these parameters are combined into a value between 0 and 1, with higher values ​​indicating stronger liquidity. For example, a liquidity subvector value of 0.9 for a transaction request indicates high transaction throughput (1.2 transactions / second), short verification time (0.5 seconds), and high circulation frequency (10 times / month), classifying it as a high-liquidity transaction.

[0135] The leader node sorts transaction requests in descending order based on their liquidity subvector values, prioritizing requests with higher liquidity to improve overall network transaction efficiency. Fairness must be considered during sorting; if the difference in liquidity subvector values ​​between two transactions is less than 0.05 (e.g., 0.85 and 0.83), they are sorted in ascending order by their initiation timestamps to prevent "queue jumping." For example, if the transaction list contains requests A (liquidity 0.9, timestamp = 1627230001), B (0.85, timestamp = 1627230002), and C (0.85, timestamp = 1627230000), the sorting result is A, C, B, prioritizing higher liquidity while ensuring the time order of transactions of the same liquidity level.

[0136] A microblock is a lightweight data unit for processing high-frequency transactions. Each microblock is limited to 1MB in size (containing approximately 200 transactions) and its structure includes a microblock header and a microblock body. The microblock header contains the block sequence number (e.g., "Micro-001"), the parent block hash (main chain block hash), the number of transactions, and a timestamp. The microblock body contains a transaction digest and a leader signature. The transaction digest is a compressed representation of the key information of each transaction (transaction ID, public keys of both parties, underlying asset identifier, and hash value), such as "Transaction ID=Tx123, hash=h1; Tx124, hash=h2...". The leader signature is the leader's signature on the microblock header using their private key, ensuring the authenticity of the microblock.

[0137] For example, the micro-block header of Micro-001 is "Sequence Number = 001, Parent Hash = Main Chain Hash, Transaction Count = 150, Timestamp = 1627230060", the micro-block body contains a digest of 150 transactions, and the leader signature is "SigC = Private Key C Signature (Micro-block Header)". After each micro-block is generated, the leader node immediately broadcasts it to the network and records it in its local micro-block sequence.

[0138] The micro-block sequence is arranged in ascending order of its generation timestamp, forming a continuous transaction processing chain. The sequence number of each micro-block is incremented, and the parent block hash points to the main chain block, ensuring its correlation with the main chain. For example, the sequence is Micro-001 (1627230060) → Micro-002 (1627230120) → Micro-003 (1627230180), with each micro-block spaced approximately 60 seconds apart, adapting to the processing needs of high-liquidity transactions.

[0139] The unverified microblock set is a collection of all generated but unverified microblocks from network nodes. This set includes the complete data of each microblock, its generation time, leader signature, and corresponding liquidity sub-vector value. The leader node packages and broadcasts this unverified microblock set every 10 microblocks generated (or every 5 minutes) to facilitate batch verification by network nodes. For example, the unverified microblock set might contain Micro-001 to Micro-010, along with the root hash "d1e2f3a4b5c6...", for nodes to verify its integrity.

[0140] Traffic control is required when generating microblocks. If there is a surge in high-liquidity transactions in a short period of time (such as 1,000 transactions per second), the leader node will activate the dynamic microblock size adjustment mechanism to temporarily increase the microblock size to 2MB and increase the generation frequency (one every 30 seconds) to avoid transaction congestion. If the transaction traffic is below the threshold (such as 10 transactions per second), it will be reduced to 0.5MB to reduce network redundancy.

[0141] Network nodes perform PoW chain verification on the unverified microblock set, calculate the hash value of each microblock and compare it with the root hash of the main chain block. Verified microblocks are included in the temporary transaction pool. At the same time, the nodes synchronously update the verification status of the local sub-feature vector and output the pool of verified microblocks.

[0142] Network nodes are non-leader nodes participating in verification within a distributed network. Each node maintains a complete copy of the main chain blocks and a set of sub-feature vectors, possessing the ability to independently verify micro-blocks. The verification process employs a combination of parallel and serial methods. Nodes first verify the root hash of the unverified micro-block set (ensuring the set's integrity), and then perform PoW chain verification on each micro-block one by one.

[0143] The hash value of each microblock is calculated using the SHA-256 algorithm. The input is the complete content of the microblock (microblock header + microblock body). The output hash value must satisfy the correlation with the root hash of the main chain block—the last 8 bits of the microblock hash must match the last 8 bits of the main chain root hash to ensure that the microblock belongs to a branch of the current main chain, rather than an isolated block. For example, if the last 8 bits of the main chain root hash are "a1b2c3d4", and the hash of microblock Micro-001 is "...e5f6a1b2c3d4", the last 8 bits match, passing the correlation check.

[0144] The verification also includes: the validity of the leader's signature (decrypting the signature using the leader's public key and comparing it with the micro-block header hash), the completeness of the transaction digests (the hash of each transaction digest must match the transaction hash recorded locally by the node), and the continuity of the micro-block sequence numbers (e.g., Micro-001 should be followed by Micro-002, with no skipped numbers). For example, when a node verifies Micro-001, if it finds that the decrypted leader signature matches the micro-block header hash, the hashes of all transaction digests match the local records, and the sequence numbers are continuous, the verification is considered successful.

[0145] The temporary transaction pool is a temporary space for nodes to store verified micro-blocks. It is managed using a FIFO (First-In, First-Out) strategy and has a capacity of 100 micro-blocks. When the capacity is exceeded, the oldest micro-block is written to the local disk for archiving. Micro-blocks in the temporary transaction pool need to be marked with a verification timestamp and verification node ID, such as "Micro-001, verification time = 1627230065, verification nodes = NodeA, NodeB, NodeD", to record the verification results of multiple nodes and enhance credibility.

[0146] The verification status of local sub-feature vectors is used to track the credibility of each sub-vector. The status is divided into "pending verification," "partial verification," and "fully verified." When a micro-block is verified, the number of verification nodes for its corresponding sub-feature vector (such as the liquidity sub-vector) increases. When the number of verification nodes is ≥3, the status is updated from "pending verification" to "partial verification." When ≥10 nodes have verified it and no node raises any objections, it is updated to "fully verified." For example, the liquidity sub-vector 0.9 is updated to "partial verification" after being verified by 3 nodes, and eventually becomes "fully verified" as more nodes verify it.

[0147] The verified microblock pool is a collection of all microblocks in the "partially verified" state or higher, arranged in ascending order of microblock number. Each microblock is accompanied by a verification signature from multiple nodes (node ​​public key + verification time). For example, the microblock pool contains Micro-001 (verified by 5 nodes) and Micro-002 (verified by 8 nodes). The set of verification signatures for each microblock forms a "verification consensus," providing a basis for subsequent integration.

[0148] Microblocks that fail verification (such as hash mismatch or invalid signature) will be marked as "invalid" and the reason for failure will be recorded (such as "leader signature forgery" or "last 8 bits of hash mismatch"). This will be broadcast to the entire network to warn other nodes and trigger the leader node to regenerate the microblock to ensure that transactions are not lost.

[0149] Based on the value stability sub-vector, multi-objective optimization weights are set, and the main chain blocks and the verified micro-block pool are merged to select blocks that meet the weight thresholds and integrate them into an optimized data block set according to timestamps.

[0150] The value stability subvector measures the volatility of a data asset's value. Key parameters include generation time (newer is more stable), data size (larger size is more stable), and value volatility (historical rate of value change). These are standardized and merged into a value between 0 and 1, with higher values ​​indicating greater value stability. For example, a value stability subvector of 0.8 for a data asset indicates that it was recently generated (within one month), has a large data size (10GB), and low value volatility (monthly volatility ≤5%), classifying it as a highly stable asset.

[0151] Multi-objective optimization weights are used to balance the three major characteristics of security, liquidity, and value stability. Weight settings must be tailored to the specific business scenario. For example, in financial data asset trading, the weights are: security 0.4, liquidity 0.3, and value stability 0.3; in social data asset trading, the liquidity weight can be increased to 0.5. The weights must sum to 1, and each weight must be at least 0.1 to prevent any characteristic from being ignored. For example, the optimization weights in this case are [security 0.4, liquidity 0.3, stability 0.3].

[0152] The process of merging main chain blocks and micro-block pools involves calculating the overall score of each block, using the formula: "Overall Score = Security Sub-vector × Security Weight + Flow Sub-vector × Flow Weight + Stability Sub-vector × Stability Weight". The overall score of a main chain block is calculated based on its core transactions, while the overall score of a micro-block is the average score of all transactions within it. For example, if a main chain block has a security sub-vector of 0.89, a flow of 0.7, and a stability of 0.8, its overall score is: 0.89 × 0.4 + 0.7 × 0.3 + 0.8 × 0.3 = 0.356 + 0.21 + 0.24 = 0.806; the average score of a certain micro-block is 0.78, both exceeding the weight threshold of 0.7.

[0153] When filtering blocks that meet the weight threshold, a comprehensive score ≥ 0.7 is set as the passing condition. Blocks with a score lower than this (e.g., a comprehensive score of 0.65) are marked as "needing optimization" and their characteristics need to be re-evaluated to ensure they meet the requirements. The filtering process also checks the logical consistency between blocks, such as whether there are conflicts between transactions in main chain blocks and micro-blocks (e.g., the same asset being traded repeatedly). If conflicts exist, transactions in main chain blocks are prioritized (because the main chain is more secure).

[0154] Timestamp-based consolidation arranges the filtered main chain blocks and micro-blocks in ascending order of their generation timestamps, forming a coherent set of optimized data blocks. The main chain blocks serve as "anchors," with micro-blocks following sequentially. Each micro-block points to the hash of a main chain block, creating a hierarchical structure of "main chain + micro chain." For example, the consolidated set might be in the following order: Main chain block (timestamp = 1627230000) → Micro-001 (1627230060) → Micro-002 (1627230120), with continuous timestamps and a clear structure.

[0155] Each block in the optimized data block set comes with a comprehensive score and the number of verification nodes. For example, the main chain block has a score of 0.806 and 20 verification nodes, while Micro-001 has a score of 0.78 and 15 verification nodes, ensuring the quality of the set is traceable. Simultaneously, the root hash of the set (the Merkle root of all block hashes) is calculated and broadcast for all network nodes to verify the integrity of the integration. If a node's integration result root hash matches the broadcast value, the optimized set is accepted; otherwise, the data is resynchronized to ensure data consistency across the entire network.

[0156] During the integration process, if duplicate transactions are found (the same transaction ID appears in multiple microblocks), the transaction with the earliest timestamp will be retained, and the rest will be marked as "duplicate and invalid" to avoid double trading of assets. If there is an error in the calculation of the transaction amount (if any) (such as violating "payment amount = unit price × quantity"), the transaction's overall score in the microblock will be reduced by 0.1. If the score is still ≥0.7 after the reduction, it will be retained; otherwise, it will be removed to ensure the accuracy of the optimized set.

[0157] S204, the optimized data block set is verified and confirmed using a distributed consensus algorithm to determine several final data blocks as the consensus results of the corresponding data asset transaction requests.

[0158] Specifically, the optimized data block set can be broadcast to all consensus nodes in the distributed network. Each node calls the smart contract to execute the transaction logic in the block locally, checks the consistency between the block hash value and the sub-feature vector, and outputs the node's local verification result.

[0159] The optimized data block set includes main chain blocks and micro-blocks optimized through multi-objective processes. Its broadcast is implemented using the gossip protocol. Upon receiving the set, each consensus node immediately forwards it to 3-5 randomly selected neighboring nodes. These neighboring nodes repeat this process until all consensus nodes in the network (typically 5-20 nodes, dynamically adjusted based on network size) have received the complete set. During broadcasting, each block is accompanied by a digital signature (signed using the leader node's private key). Receiving nodes first verify the signature's validity (decrypted using the leader's public key) to ensure the block has not been tampered with. For example, after receiving the optimized set, node A verifies the signature using the leader's public key PubC, confirming that the signature matches the block hash before proceeding with further processing.

[0160] Consensus nodes are trusted nodes that participate in the final verification. They must meet preset qualification conditions (such as computing power ≥ 1 TH / s, node online time ≥ 90%), and register through distributed identity authentication (DID), for example, node identifiers such as "Consensus-Node1" and "Consensus-Node2". Each node maintains an independent local database, storing blockchain copies, smart contract code, and sub-feature vector sets, and has complete verification capabilities.

[0161] Smart contracts are predefined, automatically executed code that encapsulates the business logic of data asset transactions. Examples include "ownership transfer rules" (requiring digital signature confirmation from the original owner), "transaction limit verification" (single transaction amount ≤ 1 million RMB), and "compliance checks" (cross-border transactions require compliance identification). When a smart contract is invoked, the node takes the transaction details from the block (transaction party IDs, underlying asset identifier, transfer quantity, etc.) as input, executes the contract code, and outputs a "valid" or "invalid" logical verification result. For example, if the ownership chain of the underlying asset in a transaction shows the current owner as UserA, but the transaction initiator is UserB (without UserA's authorized signature), the smart contract will output "invalid" after execution, marking the transaction logic as illegal.

[0162] The consistency check between the block hash value and the sub-feature vector is the core of the verification process. The block hash value is calculated from the block header information (including sub-feature vectors, timestamps, and previous block hashes) using SHA-256. For example, if the block header information is "sub-feature vector = [0.8, 0.7, 0.9], timestamp = 1627230000, previous hash = H0", the hash calculated is H1 = SHA-256 (header information). The consistency check needs to verify two points: first, whether the hash value recorded in the block is consistent with the H1 recalculated by the node (to prevent tampering); second, whether the hash value contains the features of the sub-feature vector. For example, a specific algorithm (such as standardizing the sub-feature vector and using it as part of the hash input) is used to ensure the binding relationship between the hash and the sub-feature. If the recalculated hash deviates from the recorded value by more than 0.0001 (hexadecimal difference), it is considered inconsistent.

[0163] The local verification result of a node consists of three parts: transaction logic verification result (valid / invalid), hash consistency verification result (consistent / inconsistent), and comprehensive verification conclusion (pass / fail). The rule for determining the comprehensive conclusion is: only when the transaction logic is "valid" and the hash is "consistent" is the conclusion "pass"; otherwise, it is "fail". For example, the verification result of node 1 for block A is "transaction logic valid, hash consistent → pass"; the result of node 2 is "transaction logic valid, hash inconsistent → fail", and the specific differences in the inconsistency are recorded (e.g., the hash is recorded as H1', ​​H1 is calculated, and the difference is in the 8th position).

[0164] The verification result must include the node signature and timestamp to prove its authenticity and timeliness. For example, the signature of node 1's result is "Sig-Node1 = Private Key 1 Signature (Result + Timestamp)", with the timestamp accurate to milliseconds (e.g., 1627230001.123). The local verification results of all nodes will be temporarily stored in the local result pool, awaiting collection by the practical Byzantine fault-tolerant algorithm. The average time of the verification process must be controlled within 10 seconds to ensure consensus efficiency.

[0165] The practical Byzantine fault-tolerant algorithm is used to collect the local verification results of all nodes. When the percentage of nodes that agree to the verification exceeds 2 / 3, the consensus confirmation process is triggered. The list of agreeing nodes and the reasons for rejection by dissenting nodes for each block are recorded, and the consensus confirmation certificate is output.

[0166] Practical Byzantine Fault Tolerance (PBFT) is a consensus algorithm designed to address malicious nodes (sending incorrect information or failing to respond) in distributed networks. Its core principle is a three-phase protocol (pre-preparation, preparation, confirmation) that ensures consensus is reached even when malicious nodes comprise no more than one-third of the network. Assuming a distributed network with four consensus nodes (Node1-Node4), allowing a maximum of one malicious node, at least three nodes must agree to reach consensus (more than two-thirds).

[0167] The pre-prepare phase of PBFT is initiated by the master node (usually the leader node or rotating node from step one). The master node collects the local verification results from all nodes, groups them by block, and generates a pre-prepare message for each block, containing the block hash, master node signature, and verification result digest (e.g., "3 passes, 1 fails"). For example, the master node generates the pre-prepare message "Pre-Prepare = Block A hash + H0, Signature = PubC, Digest = 3 / 1" for block A and broadcasts it to all consensus nodes.

[0168] During the preparation phase, each node, upon receiving the pre-preparation message, verifies its validity (master node signature, block hash matching). If valid, it broadcasts the preparation message to the entire network, including an "agree / disagree" flag and evidence supporting that flag (such as the hash calculation process in the local verification result). For example, if Node1 agrees to block A, its preparation message would be "Prepare=Node1, Block A, Agree, Evidence = Hash Calculation Record"; if Node4 is a malicious node and sends "Disagree" without providing evidence, its message will be ignored.

[0169] The confirmation phase is the final consensus-building phase. Nodes collect preparation messages from other nodes. When a node receives "agree" preparation messages (containing valid evidence) from more than 2 / 3 of the nodes, it broadcasts a confirmation message to the entire network. If more than 1 / 3 of the nodes receive "disagree" preparation messages, it broadcasts a rejection confirmation message. For example, if Node1 receives agreement preparation messages from Node2 and Node3 (including itself, a total of 3 agree messages, representing 3 / 4 > 2 / 3), it broadcasts the confirmation message "Commit=Node1, Block A, Agree".

[0170] When the number of confirmation messages for a block exceeds 2 / 3 of the nodes, the consensus confirmation process is triggered. For example, if block A receives confirmation messages from Nodes 1-3 (3 / 4 > 2 / 3), it is considered passed; if block B only receives confirmation messages from Nodes 1-2 (2 / 4 < 2 / 3), it is considered failed. For failed blocks, the reasons for rejection from dissenting nodes need to be analyzed. Common reasons include "transaction logic violation (e.g., ownership transfer without the original owner's signature)," "hash inconsistency (difference between recorded value and calculated value)," and "sub-feature vector mismatch (hash does not contain feature information)," etc. For example, Node 4's reason for rejecting block A is "the calculated hash and the recorded hash differ in the 8th bit (0x3 vs 0x5)."

[0171] When recording the list of nodes that agree, it must include the node identifier, signature, and confirmation timestamp. For example, the list of nodes that agree for block A is “Node1 (Sig1, 1627230002.123), Node2 (Sig2, 1627230002.145), Node3 (Sig3, 1627230002.156)”. In addition to the node identifier, the list of nodes that disagree must also record the reasons for rejection and evidence (such as a hash calculation comparison table, screenshots of violations of transaction logic), for example, “Node4: Hash difference (record H1'=...3..., calculate H1=...5...), evidence = Attachment 1”.

[0172] A consensus confirmation certificate is a structured document containing the consensus results of all blocks. Its format is: Block Identifier + Consensus Result (Pass / Fail) + List of Agreeing Nodes + List of Disagreeing Nodes + Root Hash of the Certificate. The root hash of the certificate is calculated by concatenating the hashes of the consensus results of all blocks. For example, if the hash of block A is HA and the hash of block B is HB, the root hash = SHA-256(HA+HB), ensuring the certificate is immutable. After the certificate is generated, it is jointly signed by all agreeing nodes and stored in a distributed ledger, which can be queried and verified by any node.

[0173] For blocks that fail to pass, the consensus confirmation process triggers a secondary verification mechanism: the reasons for rejection from dissenting nodes are extracted, and the master node organizes a review group of 3 random nodes to re-verify the block. If the review result is "passed" (2 / 3 of the review nodes agree), the consensus result is updated to "passed". Otherwise, it is marked as an "invalid block" and needs to be returned to the data block set optimization stage for reprocessing to ensure that all transaction requests are handled properly.

[0174] Based on consensus confirmation credentials, the passed blocks undergo final hash chain verification to ensure the continuity and immutability between blocks, filter out abnormal blocks with broken hash chains, and output a sequence of candidate blocks that have passed verification.

[0175] Hash chain verification is a crucial step in ensuring the continuity and immutability of a blockchain. Its core principle is that "the hash value of each block contains the hash value of the previous block," forming a "chain structure." Any tampering with a historical block will invalidate the hash value of subsequent blocks, thus making it detectable. For example, in a blockchain sequence Block0→Block1→Block2, the hash H1 of Block1 contains the hash H0 of Block0, and the hash H2 of Block2 contains H1. That is, H1 = SHA-256(Block1 content + H0), and H2 = SHA-256(Block2 content + H1).

[0176] Based on the consensus confirmation credentials, the "passed" blocks are first selected. Assuming the passed blocks are Block0 (H0), Block1 (H1), Block2 (H2), and Block3 (H3), they need to be sorted by timestamp (to ensure verification according to the generation order). The sorted sequence is Block0 (t0) → Block1 (t1) → Block2 (t2) → Block3 (t3), where t0... <t1<t2<t3。

[0177] The specific process of hash chain verification is as follows: Starting from the first block (Block0), record its hash H0; verify whether the hash H1 of the second block (Block1) contains H0, that is, extract the "preceding hash" field from the block header of Block1, check whether it is equal to H0, and recalculate H1' = SHA-256 (Block1 content + H0). If H1' = H1 and the preceding hash = H0, then Block1 passes the verification; then use H1 to verify Block2, and repeat the above process until all blocks that have passed the verification are completed.

[0178] For example, the preceding hash field of Block 1 is recorded as H0. Recalculating H1' = SHA-256 (Block 1 content + H0) = H1, which matches the recorded H1, passes the verification. The preceding hash field of Block 2 is recorded as H1, but recalculating H2' = SHA-256 (Block 2 content + H1) = H2'' ≠ H2, and the recorded H2 does not contain H1 (the preceding hash field was found to have been tampered with to H1' through hash parsing). Therefore, it is determined that the hash chain of Block 2 is broken, making it an abnormal block.

[0179] Immutability verification requires the use of historical hash records. Each block's hash value is broadcast to all network nodes and written to a local hash database after generation. Verification compares the current block's hash with historical records in the hash database. If the difference exceeds a preset threshold (e.g., one different digit in hexadecimal), it is determined to have been tampered with. For example, if Block0's historical hash record is H0, and the current calculated hash is H0', a comparison reveals that H0' differs from H0 at the 10th bit (0x7 vs 0x3). Therefore, Block0 is determined to have been tampered with, and all subsequent blocks (dependent on H0) are marked as abnormal.

[0180] The filtering of abnormal blocks must follow the principle of "removal upon breakage": if a block's hash chain is broken (e.g., Block 2), then that block and all subsequent blocks (Block 3) are removed, retaining only the valid blocks before the break (Block 0, Block 1); if a block is tampered with (e.g., Block 0), then all blocks are removed, and the process returns to the initial stage to regenerate blocks. After filtering, the reason for the anomaly must be recorded, such as "Block 2: inconsistent hash calculation, previous hash field tampered with" or "Block 0: inconsistent with historical hash database records, suspected tampering," to provide a basis for subsequent tracing.

[0181] The verified candidate block sequence is a set of blocks whose hash chains are complete and tamper-proof, arranged in ascending order of timestamps. Each block includes a verification identifier (e.g., "verified_hash chain complete"), verification node signatures (at least 3 node signatures), and a verification timestamp. For example, a candidate block sequence of [Block0(t0, verified), Block1(t1, verified)] ensures the continuity from the initial block to the current block, and the hash of each block can be traced back to its source, laying the foundation for the final consensus result.

[0182] To improve verification efficiency, the hash chain verification adopts a parallel computing mode, with each node responsible for verifying a portion of the blocks (e.g., 4 nodes share the task of verifying 10 blocks, with each node verifying 2-3 blocks). The verification results are synchronized through a P2P network and finally aggregated by the master node to form a global verification result. The average verification time is controlled within 30 seconds to avoid affecting the overall efficiency of the consensus process.

[0183] The candidate block sequence is arranged in ascending order by transaction timestamp to generate a final data blockchain containing a complete ownership chain, transaction records, and sub-feature vector labels, which serves as the consensus result for the corresponding data asset transaction request.

[0184] Although the candidate block sequence has passed hash chain verification, there may be timestamp disorder (e.g., the timestamp of a later-generated block is earlier than that of a previous block). Therefore, they need to be sorted in ascending order of transaction timestamps to ensure the temporal consistency of the blockchain. A transaction timestamp is a record of the time each transaction was initiated (accurate to milliseconds). The block timestamp is the earliest transaction timestamp contained within it. Sorting is based on the block timestamps. If the difference between two block timestamps is less than 1 second (possibly due to network latency), they are sorted lexicographically by their hash values ​​(smaller hash value first).

[0185] For example, in the candidate block sequence, Block0's timestamp is 1627230000.123, and Block1's is 1627230000.098 (slightly earlier due to network latency). The difference between the two block timestamps is 0.025 seconds (less than 1 second). Arranged in lexicographical order, Block1's hash is "a1b2c3...", and Block0's is "b1c2d3...". Since "a" is earlier than "b" in lexicographical order, the order is [Block1, Block0], ensuring the relative rationality of the time sequence.

[0186] The integration of a complete ownership chain requires tracing the ownership transfer history of each data asset, from the initial owner to the current owner. Each transfer record must include the signature of the previous owner, the signature of the current owner, the transfer timestamp, and the identifier of the underlying asset, forming a complete chain of "Owner A → Owner B → Owner C". For example, if the ownership chain of a certain data asset records "Initial Owner = UserA" in Block 0 and "UserA → UserB (signature verification passed)" in Block 1, the integrated ownership chain is "UserA (t0) → UserB (t1)", ensuring that the transfer of asset ownership is traceable.

[0187] The integration of transaction records must include detailed information for each transaction: transaction ID, IDs of both parties, identifier of the underlying asset, transaction amount (if any), transaction status ("confirmed"), digital signature, and corresponding sub-feature vector label (e.g., "security = 0.9, liquidity = 0.8, value stability = 0.7"). For example, the integrated transaction records contained in Block1 would be "Tx001: UserA→UserB, Asset Identifier = Asset123, Amount = 500 yuan, Status = Confirmed, Signature = SigA+SigB, Sub-feature Label = Security 0.9_Liquidity 0.8_Stability 0.7".

[0188] The embedding of sub-feature vector labels must correspond one-to-one with blocks. Each block header contains a set of sub-feature vectors for the transactions processed in that block (e.g., [0.9, 0.8, 0.7]). The label format is a key-value pair of "feature type = value", which facilitates subsequent querying of the core characteristics of the transactions in that block. For example, the sub-feature labels for Block 0 are "security = 0.85, liquidity = 0.7, value stability = 0.8", and the labels for Block 1 are "security = 0.9, liquidity = 0.8, value stability = 0.7", clearly reflecting the changes in transaction characteristics at different stages.

[0189] The final data blockchain is a complete chain integrating the above information. Its structure includes: a block header (block identifier, timestamp, previous hash, sub-feature tag, number of transactions), a block body (list of transaction records, ownership chain fragment, consensus confirmation credential digest), and a block tail (verification node signature, hash value). For example, the first block of the final blockchain, Block 0, contains the initial transaction records and ownership starting point. Subsequent blocks are connected sequentially to form a complete closed loop of "transaction initiation → verification → consensus → confirmation".

[0190] This blockchain, serving as the consensus result for data asset transaction requests, will be written into a distributed ledger and broadcast to the entire network. Both parties to the transaction can query the blockchain using the target asset's identifier to obtain the final state and complete record of the transaction. Simultaneously, the blockchain's hash value will be synchronized to multiple backup nodes to ensure data persistence and resilience; any node failure or data loss will not affect the validity of the consensus result.

[0191] As can be seen, the process involves receiving data asset transaction requests from a distributed network, generating an initial set of data blocks based on these requests, performing feature mapping on the initial data block set to construct a block feature vector set, decomposing the feature vectors in the block feature vector set into several sub-feature vectors to capture different characteristics of the data assets, using an improved Bitcoin-NG consensus mechanism to perform multi-objective optimization on the sub-feature vectors to generate an optimized data block set, and applying a distributed consensus algorithm to the optimized data block set for verification and confirmation to determine several final data blocks as the consensus results for the corresponding data asset transaction requests. This process improves the efficiency and security of data asset consensus.

[0192] Another embodiment of the present invention provides a consensus system based on data assets, see [link to relevant documentation]. Figure 3 The system may include:

[0193] The receiving module 301 is used to receive data asset transaction requests in a distributed network and generate a set of initial data blocks based on the transaction requests, wherein each initial data block is associated with a feature dimension of the data asset;

[0194] The construction module 302 is used to perform feature mapping on the initial data block set, construct a block feature vector set, and decompose the feature vectors in the block feature vector set into several sub-feature vectors to capture different characteristics of the data assets. Each feature vector represents a multi-dimensional feature attribute of the corresponding initial data block.

[0195] The optimization module 303 is used to perform multi-objective optimization on the sub-feature vector using the improved Bitcoin-NG consensus mechanism to generate an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader through the main chain block and processes transactions through the micro-block, and combines the PoW chain to ensure the security and immutability of data assets.

[0196] The verification module 304 is used to apply a distributed consensus algorithm to verify and confirm the optimized data block set, and determine several final data blocks as the consensus results of the corresponding data asset transaction requests.

[0197] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0198] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0199] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.

[0200] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A data asset based consensus method, characterized in that, The method comprises: Receiving a data asset transaction request in a distributed network, generating a set of initial data block sets based on the transaction request, wherein each initial data block is associated with a characteristic dimension of a data asset; Feature mapping of the initial data block set is performed to construct a block feature vector set, and the feature vectors in the block feature vector set are decomposed into a plurality of sub-feature vectors to capture different characteristics of the data asset, wherein each feature vector represents the multi-dimensional feature attributes of the corresponding initial data block; A multi-objective optimization is performed on the sub-feature vectors using an improved Bitcoin-NG consensus mechanism to generate an optimized data block set, wherein the improved Bitcoin-NG consensus mechanism confirms the leader through the main chain block and processes the transaction through the micro block, and combines PoW chain to ensure the security and tamper resistance of the data asset; A distributed consistency algorithm is applied to the optimized data block set to verify and confirm a plurality of final data blocks as the consensus result of the corresponding data asset transaction request.

2. The method of claim 1, wherein, The receiving of the data asset transaction request in the distributed network and the generation of a set of initial data block sets based on the transaction request, wherein each initial data block is associated with a characteristic dimension of a data asset, comprises: Receiving a data asset transaction request through a distributed network node, parsing the asset type, transaction digital signature and subject asset identification in the request, verifying the signature validity using an elliptic curve encryption algorithm, filtering out malicious requests with invalid signatures, and outputting a list of valid transaction requests; For each request in the list of valid transaction requests, extract the characteristic dimensions of the corresponding data asset including data size, generation time, ownership chain length, encryption level, and circulation frequency, eliminate dimensional differences through feature standardization processing, and output a data asset characteristic dimension table; Based on the data asset characteristic dimension table, use clustering algorithm to group the transaction requests according to feature similarity, each group corresponds to a feature dimension combination, assign a unique initial block identifier to each group, and output a block-feature dimension mapping table; According to the block-feature dimension mapping table, pack the detailed information of each group of transaction requests into an initial data block, embed the corresponding feature dimension label in the header of each block, and generate an initial data block set.

3. The method of claim 2, wherein, The feature mapping of the initial data block set, the construction of the block feature vector set, and the decomposition of the feature vectors in the block feature vector set into a plurality of sub-feature vectors to capture different characteristics of the data asset, wherein each feature vector represents the multi-dimensional feature attributes of the corresponding initial data block, comprises: Iterate through the initial data block set, extract the specific numerical value corresponding to the feature dimension label for each block, construct an original feature matrix combining the derived features of the block including transaction throughput and verification time consumption; Perform maximum and minimum normalization on the original feature matrix to eliminate the dimensional influence of different feature dimensions, detect and remove multicollinearity features through variance inflation factor, and output a de-redundancy normalized feature matrix; Based on the de-redundant normalized feature matrix, a deep learning embedding model is used to map the multi-dimensional features of each block to a high-dimensional vector, and each dimension of the vector corresponds to an abstract expression of the feature. The output is a set of block feature vectors. A non-negative matrix factorization algorithm is applied to the block feature vector set, and the decomposition dimension is set according to the core characteristics of the data asset. Each high-dimensional feature vector is decomposed into multiple sub-feature vectors, each corresponding to a single characteristic quantization expression. The output is an initial set of sub-feature vectors. Calculate the mutual information value of the initial sub-feature vector set, remove redundant sub-vectors with mutual information higher than the threshold, and verify the independence of the sub-vectors through principal component analysis. The output is an optimized set of sub-feature vectors.

4. The method of claim 3, wherein, The improved Bitcoin-NG consensus mechanism is used to optimize the sub-feature vectors for multiple objectives, generating an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader of the main chain block and processes transactions through microblocks, ensuring the security and tamper resistance of the data asset through the PoW chain. This includes: Extract security sub-vectors from the sub-feature vector set and combine them with node computing power values to build a comprehensive score. The leader of the main chain block is elected through a computing power competition, and the leader node identifier and the hash record of the election process are output. The leader node generates a main chain block based on the security sub-vector. The block header contains the root hash of the sub-feature vector set, the leader public key, and the term information. The main chain block is mined through the PoW mechanism, and the main chain block with PoW proof is output. The leader node sorts transaction requests according to the liquidity sub-vector, and prioritizes high-liquidity requests for packaging into microblocks. Each microblock contains a transaction summary corresponding to the sub-feature vector and a leader signature. The microblock sequence is generated in transaction order, and the unverified microblock set is output. Network nodes perform PoW chain verification on the unverified microblock set. The hash value of each microblock is calculated and compared with the root hash of the main chain block. The verified microblocks are included in the temporary transaction pool, and the node synchronously updates the local sub-feature vector verification status. The verified microblock pool is output. Based on the value stability sub-vector, set the multi-objective optimization weight, fuse the main chain block and the verified microblock pool, and select the blocks that meet the weight threshold. The optimized data block set is integrated by timestamp.

5. The method of claim 4, wherein, A distributed consistency algorithm is applied to the optimized data block set for verification and confirmation, determining a number of final data blocks as the consensus result of the corresponding data asset transaction request. This includes: Broadcast the optimized data block set to all consensus nodes in the distributed network. Each node calls the smart contract to locally execute the transaction logic in the block, checks the consistency of the block hash value and the sub-feature vector, and outputs the local verification result of the node. Use the practical Byzantine fault tolerance algorithm to collect the local verification results of all nodes. When the proportion of nodes that agree to verify exceeds 2 / 3, trigger the consensus confirmation process, record the list of nodes that agree to verify and the reasons for the rejection of dissenting nodes, and output the consensus confirmation credentials. Based on consensus confirmation credentials, the passed blocks undergo final hash chain verification to ensure the continuity and immutability between blocks, filter out abnormal blocks with broken hash chains, and output a sequence of candidate blocks that have passed verification. The candidate block sequence is arranged in ascending order by transaction timestamp to generate a final data blockchain containing a complete ownership chain, transaction records, and sub-feature vector labels, which serves as the consensus result for the corresponding data asset transaction request. 6.A consensus system based on data assets, characterized in that, The system includes: The receiving module is used to receive data asset transaction requests in a distributed network and generate a set of initial data blocks based on the transaction requests, wherein each initial data block is associated with a feature dimension of the data asset. The construction module is used to perform feature mapping on the initial data block set, construct a block feature vector set, and decompose the feature vectors in the block feature vector set into several sub-feature vectors to capture different characteristics of the data assets. Each feature vector represents a multi-dimensional feature attribute of the corresponding initial data block. The optimization module is used to perform multi-objective optimization on the sub-feature vector using the improved Bitcoin-NG consensus mechanism to generate an optimized data block set. The improved Bitcoin-NG consensus mechanism confirms the leader through the main chain block and processes transactions through the micro-block, and combines the PoW chain to ensure the security and immutability of data assets. The verification module is used to apply a distributed consensus algorithm to verify and confirm the optimized data block set, and determine several final data blocks as the consensus results of the corresponding data asset transaction requests.

7. The system of claim 6, wherein, The receiving module is specifically used for: The system receives data asset transaction requests through distributed network nodes, parses the asset type, digital signatures of both parties, and the identifier of the underlying asset in the request, verifies the validity of the signature using elliptic curve cryptography, filters out malicious requests with invalid signatures, and outputs a list of valid transaction requests. For each request in the list of valid transaction requests, extract the feature dimensions of the corresponding data asset, including data size, generation time, ownership chain length, encryption level, and circulation frequency. Eliminate the difference in units through feature standardization and output a data asset feature dimension table. Based on the data asset feature dimension table, a clustering algorithm is used to group transaction requests according to feature similarity. Each group corresponds to a feature dimension combination. A unique initial block identifier is assigned to each group, and a block-feature dimension mapping table is output. Based on the block-feature dimension mapping table, the detailed information of each group of transaction requests is packaged into an initial data block, and the corresponding feature dimension label is embedded in the header of each block to generate an initial data block set.

8. The system of claim 7, wherein, The building module is specifically used for: Traverse the initial data block set, extract the specific values ​​corresponding to the feature dimension labels of each block, and combine the derived features of the block, such as the transaction throughput and verification time, to construct the original feature matrix; The original feature matrix is ​​processed by minimax normalization to eliminate the influence of different feature dimensions. Multicollinearity features are detected and removed by variance inflation factor, and the redundant normalized feature matrix is ​​output. Based on the redundancy-removing normalized feature matrix, a deep learning embedding model is used to map the multidimensional features of each block into a high-dimensional vector. Each dimension of the vector corresponds to the abstract expression of the feature, and the block feature vector set is output. Apply the nonnegative matrix factorization algorithm to the block feature vector set, set the decomposition dimension according to the core characteristics of the data asset, decompose each high-dimensional feature vector into multiple sub-feature vectors, each corresponding to the quantitative expression of a single characteristic, and output the initial sub-feature vector set; Calculate the mutual information value of the initial sub-feature vector set, remove redundant sub-vectors with mutual information higher than the threshold, verify the independence of the sub-vectors through principal component analysis, and output the optimized sub-feature vector set.

9. A storage medium, characterized by The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-5 when it is run.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Commodity transaction data management system and method

    CN116579775A

  • Digital asset security verification and information monitoring method and system

    CN119416179A