Blockchain-driven end-to-end monitoring system and method for data exchange

By generating environmental fingerprints and digitally signing them, verifying TEE signatures and device behavior baselines, and implementing hierarchical and sharded storage and dynamic selection of consensus algorithms, the issues of data credibility and storage security in data exchange monitoring are solved, achieving reliability and traceability throughout the entire data exchange process.

CN120825496BActive Publication Date: 2026-04-03GUANGZHOU YITUO SOFTWARE DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing data exchange monitoring technologies cannot guarantee the original credibility of data, lack trust assessment of data sources and device behavior, and lack sharding and layering storage mechanisms, resulting in insufficient rationality and security of data storage. This makes it difficult to meet the evidence preservation needs of data exchange in a multi-chain environment, affecting the traceability and credibility of data exchange.

Method used

By collecting business data to generate environmental fingerprints and digitally signing them, the validity of TEE signatures is verified. The environmental fingerprints are extracted and compared with the device behavior baseline to calculate real-time trust scores. Trusted data packets are fragmented and stored in the corresponding storage layer according to storage attributes. A consensus algorithm is dynamically selected to generate cross-chain verifiable evidence, thereby achieving effective supervision of the entire data process.

Benefits of technology

Ensure the reliability and authenticity of data sources, guarantee data quality from the source, improve storage efficiency and data manageability, ensure data consistency and verifiability in cross-chain environments, and enhance the security and reliability of the entire data exchange process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825496B_ABST
    Figure CN120825496B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data exchange technology, and specifically discloses a blockchain-driven end-to-end data exchange monitoring system and method. The system includes: a trusted data acquisition terminal that collects business data and synchronously generates an environmental fingerprint, and digitally signs the combination of business data and environmental fingerprint to generate a trusted data packet; an edge verification gateway that receives the trusted data packet and verifies the validity of the TEE signature; when the TEE signature validity verification passes, it extracts the environmental fingerprint, compares it with the corresponding device behavior baseline, and calculates a real-time trust score; a hierarchical sharding storage terminal that, when the trust score is higher than the trust score threshold, shards the trusted data packet according to storage attributes and stores it in the corresponding storage layer to obtain the storage result; and an elastic consensus selection terminal that determines the current network load based on the storage result, and dynamically selects a consensus algorithm based on the current network load to generate cross-chain verifiable evidence for the trusted data packet; ensuring the smooth operation of the entire data exchange process and the reliable flow of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data exchange technology, and in particular to a blockchain-driven data exchange end-to-end monitoring system and method. Background Technology

[0002] In the digital age, data has become a core asset for enterprises and organizations. Data exchange occurs frequently in business processes across various fields, encompassing scenarios such as financial transactions, supply chain management, and medical data sharing. Ensuring the security, reliability, and traceability of data exchange is crucial. With the explosive growth of data volume and the increasing complexity of data interaction, traditional data monitoring methods face numerous challenges. Data is susceptible to tampering and leakage during transmission and storage, potentially leading to severe economic losses, privacy violations, and business disruptions. Furthermore, data exchange involves multiple parties, making the establishment and maintenance of trust among them increasingly difficult, and ensuring the authenticity and integrity of data is challenging. Blockchain technology, with its decentralized, immutable, and traceable characteristics, offers a new solution for data exchange monitoring. As digital transformation accelerates, industries are increasingly demanding higher standards for data security and privacy protection. Blockchain-driven end-to-end data exchange monitoring systems have broad application prospects, enabling secure and efficient data exchange in numerous sectors such as finance, healthcare, and government, driving industries towards digitalization and intelligence.

[0003] However, existing data exchange monitoring technologies cannot guarantee the original credibility of data and lack trust assessment of data sources and device behavior. The lack of a sharded and layered storage mechanism results in insufficient rationality and security of data storage, and makes it difficult to meet the evidence preservation requirements of data exchange in a multi-chain environment, affecting the traceability and credibility of data exchange.

[0004] Therefore, this invention proposes a blockchain-driven data exchange end-to-end monitoring system and method. Summary of the Invention

[0005] This invention provides a blockchain-driven end-to-end data exchange monitoring system and method. By collecting business data and generating an environmental fingerprint, and then combining these two data points to digitally sign and generate a trusted data packet, the reliability of the data source and the authenticity of the data during the collection process are ensured, guaranteeing data quality from the source. Upon receiving the trusted data packet, the validity of the TEE signature is first verified. If verified, the environmental fingerprint is extracted and compared with the device behavior baseline, and a real-time trust score is calculated to further verify the data's credibility, providing a screening basis for subsequent storage and ensuring data storage security. When the trust score exceeds a certain threshold, the trusted data packet is fragmented and stored in the corresponding storage layer according to storage attributes, achieving reasonable data storage and management, improving storage efficiency and data manageability. Based on the storage results, the network load is determined, and a consensus algorithm is dynamically selected to generate cross-chain verifiable evidence. This adapts to different network conditions, ensuring data consistency and verifiability in a cross-chain environment, guaranteeing the smooth operation of the entire data exchange process and reliable data flow. By combining blockchain technology with the data monitoring process, the entire process from data collection to storage and evidence generation can be effectively supervised, improving data credibility and security.

[0006] This invention provides a blockchain-driven end-to-end data exchange monitoring system, comprising:

[0007] The trusted data acquisition terminal is used to collect business data and synchronously generate environmental fingerprints, and digitally sign the combination of business data and environmental fingerprints to generate trusted data packets;

[0008] The edge verification gateway is used to receive trusted data packets and verify the validity of TEE signatures. When the validity of the TEE signature is verified, the environmental fingerprint is extracted, compared with the corresponding device behavior baseline, and a real-time trust score is calculated.

[0009] The hierarchical and fragmented storage layer is used to fragment trusted data packets according to storage attributes and store them in the corresponding storage layer when the trust score is higher than the trust score threshold, thereby obtaining the storage result.

[0010] The elastic consensus selection end is used to determine the current network load based on the storage results, and dynamically select a consensus algorithm to generate cross-chain verifiable evidence for trusted data packets based on the current network load.

[0011] Optionally, the method for extracting environmental fingerprints at the edge verification gateway, comparing them with the corresponding device behavior baseline, and calculating a real-time trust score includes:

[0012] Extracting feature values ​​of multi-dimensional feature terms of environmental fingerprints from trusted data packets;

[0013] The normal fluctuation range of each characteristic item of the corresponding device is collected by the sliding window algorithm and used as the baseline of the corresponding device behavior.

[0014] The deviation of the feature values ​​of all dimensions of the environmental fingerprint from the device behavior baseline is calculated, and the real-time trust score is calculated by combining the weights of all dimensions of the feature.

[0015] Optionally, the tiered and fragmented storage end includes:

[0016] The feature parsing module is used to treat trusted data packets and all corresponding synchronous trusted data packets as all trusted data packets to be evaluated when the trust score is higher than the trust score threshold. It parses all static features and all dynamic feature dimensions of each trusted data packet to be evaluated, generates static feature vectors based on all static features of each trusted data packet to be evaluated, and performs sliding time window sampling on each dynamic feature dimension of each trusted data packet to be evaluated and statistically obtains the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated.

[0017] The feature operation module is used to calculate the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated and the storage attraction between every two trusted data packets to be evaluated, based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated and the numerical sequence of all sliding time window samples of each dynamic feature dimension of each trusted data packet to be evaluated.

[0018] The attribute determination module is used to determine the storage attributes of the current trusted data packet based on the structural entropy of all dynamic feature dimensions of each trusted data packet to be evaluated, the storage attraction between every two trusted data packets to be evaluated, and the static feature vector of each trusted data packet to be evaluated, and to determine the storage layer and storage algorithm of the current trusted data packet based on the storage attributes of the current trusted data packet.

[0019] The storage execution module is used to store the current trusted data packet based on the storage algorithm and storage layer, and obtain the storage result.

[0020] Optionally, the feature calculation module includes:

[0021] The information entropy calculation submodule is used to calculate the information entropy of each dynamic feature dimension of each trusted data packet to be evaluated based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated.

[0022] The storage gravity calculation submodule is used to calculate the storage gravity between two trusted data packets to be evaluated based on the information entropy of each pair of trusted data packets to be evaluated and the feature space distance between the current values ​​of all dynamic feature dimensions of the corresponding two trusted data packets to be evaluated.

[0023] The structural entropy calculation submodule is used to calculate the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated based on the k-nearest neighbor algorithm and the numerical sequence of all sliding time windows sampled for each dynamic feature dimension of each trusted data packet to be evaluated.

[0024] Optionally, the attribute determination module includes:

[0025] The Gravity Clustering Submodule is used to cluster all trusted data packets to be evaluated based on the structural entropy of all dynamic feature dimensions of each trusted data packet to be evaluated and the storage gravity between every two trusted data packets to be evaluated, thereby obtaining multiple data packet clusters.

[0026] The thermal sorting submodule is used to perform thermal sorting on all data packet clusters and obtain the thermal sorting value of all data packet clusters.

[0027] The storage attribute determination submodule is used to determine the storage attribute of the current trusted data packet based on the heat ranking value and static feature vector of the data packet cluster to which the current trusted data packet belongs, and to determine the storage layer and storage algorithm of the current trusted data packet based on the storage attribute of the current trusted data packet.

[0028] Optionally, the gravity clustering submodule includes:

[0029] The first clustering unit is used to merge two corresponding trusted data packets to be evaluated with a storage gravity greater than a preset gravity threshold as an initial cluster, continue to calculate the storage gravity between each pair of initial clusters, and continue to merge two corresponding initial clusters with a storage gravity greater than a preset gravity threshold as a merged cluster. At the same time, the average structural entropy of the structural entropy of all dynamic feature dimensions of all trusted data packets to be evaluated contained in each merged cluster is calculated.

[0030] The second clustering unit is used to determine whether there are any merged clusters whose average structural entropy exceeds the structural entropy threshold. If so, the corresponding merged cluster is split, and the split clusters obtained after splitting are merged with all the remaining merged clusters until there are no cluster pairs with a storage gravity greater than the preset gravity threshold. In this case, multiple data packet clusters are obtained. Otherwise, all merged clusters are merged until there are no cluster pairs with a storage gravity greater than the preset gravity threshold. In this case, multiple data packet clusters are obtained.

[0031] Optionally, the thermal sorting submodule includes:

[0032] The cluster feature extraction unit is used to determine the activity, timeliness, and importance of each data packet cluster;

[0033] The heat value determination unit is used to determine the heat value of each data packet cluster based on its activity, timeliness, and importance.

[0034] The thermal sorting unit is used to sort all data packet clusters thermally according to the principle of thermal value from largest to smallest, and obtain the thermal sorting value of all data packet clusters.

[0035] Optionally, the storage attribute determination submodule includes:

[0036] Based on the feature aggregation unit, the storage attribute of the current trusted data packet is generated based on the heat ranking value of the data packet cluster to which the current trusted data packet belongs and the static feature vector.

[0037] The mapping table retrieval unit is used to retrieve the storage attribute based on the feature-attribute feature mapping table of the current trusted data packet, determine the storage attribute of the current trusted data packet, and determine the storage layer and storage algorithm of the current trusted data packet based on the storage attribute of the current trusted data packet.

[0038] Optionally, the flexible consensus selection endpoint includes:

[0039] The network load determination module is used to determine the current network load based on the stored results, where the current network load includes the number of nodes, average transmission latency, and transaction conflict probability.

[0040] The consensus algorithm decision module is used to input the current network load into the consensus decision tree to select a consensus algorithm;

[0041] The consensus switching execution module is used to switch the consensus algorithm at the next block height, while freezing the queue of unconfirmed transactions until the switch is completed and cross-chain verifiable evidence is generated.

[0042] This invention provides a blockchain-driven method for monitoring the entire data exchange process, including:

[0043] Collect business data and generate environmental fingerprints simultaneously, and digitally sign the combination of business data and environmental fingerprints to generate trusted data packets;

[0044] Receive trusted data packets and verify the validity of the TEE signature. When the validity of the TEE signature is verified, extract the environmental fingerprint, compare it with the corresponding device behavior baseline, and calculate the real-time trust score.

[0045] When the trust score is higher than the trust score threshold, the trusted data packet is fragmented and stored in the corresponding storage layer according to the storage attribute to obtain the storage result;

[0046] Based on the storage results, the current network load is determined, and a consensus algorithm is dynamically selected to generate cross-chain verifiable evidence for trusted data packets.

[0047] The beneficial effects of this invention compared to existing technologies are as follows: By collecting business data and generating environmental fingerprints, and then combining the two to digitally sign and generate trusted data packets, the reliability of the data source and the authenticity of the data during the collection process are ensured, guaranteeing data quality from the source. After receiving trusted data packets, the validity of the TEE signature is first verified. If it passes, the environmental fingerprint is extracted and compared with the device behavior baseline, and a real-time trust score is calculated to further verify the credibility of the data, providing a screening basis for subsequent storage and ensuring the security of data storage. When the trust score is higher than the trust score threshold, the trusted data packets are fragmented and stored in the corresponding storage layer according to storage attributes, realizing reasonable data storage and management, improving storage efficiency and data manageability. Based on the storage results, the network load is determined, and a consensus algorithm is dynamically selected to generate cross-chain verifiable evidence, which can adapt to different network conditions, ensuring the consistency and verifiability of data in the cross-chain environment, and ensuring the smooth operation of the entire data exchange process and the reliable flow of data. By combining blockchain technology with data monitoring processes, the entire process of data collection, storage, and evidence generation can be effectively supervised, improving the credibility and security of the data.

[0048] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.

[0049] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0050] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0051] Figure 1 This is a schematic diagram of a blockchain-driven data exchange end-to-end monitoring system according to an embodiment of the present invention;

[0052] Figure 2 This is a schematic diagram of the method for calculating real-time trust scores in an embodiment of the present invention. Detailed Implementation

[0053] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0054] like Figure 1 As shown, this invention provides an embodiment of a blockchain-driven end-to-end data exchange monitoring system, comprising:

[0055] The trusted data acquisition terminal is used to collect business data and synchronously generate environmental fingerprints, and digitally sign the combination of business data and environmental fingerprints to generate trusted data packets;

[0056] The edge verification gateway is used to receive trusted data packets and verify the validity of TEE signatures. When the validity of the TEE signature is verified, the environmental fingerprint is extracted, compared with the corresponding device behavior baseline, and a real-time trust score is calculated.

[0057] The hierarchical and fragmented storage layer is used to fragment trusted data packets according to storage attributes and store them in the corresponding storage layer when the trust score is higher than the trust score threshold, thereby obtaining the storage result.

[0058] The elastic consensus selection end is used to determine the current network load based on the storage results, and dynamically select a consensus algorithm to generate cross-chain verifiable evidence for trusted data packets based on the current network load.

[0059] In this embodiment, business data refers to the data generated and involved in various business processes under the data exchange scenario, such as transaction amount and information of both parties in financial transactions. It is the core object of monitoring and the basis for subsequent operations.

[0060] In this embodiment, environmental fingerprint is a feature identifier related to the business data collection environment. It is composed of multi-dimensional environmental information features such as device hardware, software and network, and is used to identify the collection environment to verify the reliability of the data source.

[0061] In this embodiment, business data is collected and environmental fingerprints are generated simultaneously: this means that the trusted data collection terminal collects multi-dimensional information about the collection environment while acquiring business data to generate an environmental fingerprint, so that the business data is closely associated with the collection environment and the reliability of the data source is guaranteed.

[0062] In this embodiment, a digital signature is performed on the combination of business data and environmental fingerprint to generate a trusted data packet: the trusted data acquisition end uses the private key of the asymmetric encryption algorithm to perform encryption operation on the combination of business data and environmental fingerprint to generate a digital signature. This signature, together with the business data and environmental fingerprint, constitutes a trusted data packet, ensuring data integrity and authenticity.

[0063] In this embodiment, the trusted data packet is an overall data packet formed by combining business data, environmental fingerprints, and digitally signing the combination of the two. It includes the data itself, its source environment information, and authenticity verification identifiers, and is used for transmission and verification at various stages of the system.

[0064] In this embodiment, the validity of the TEE signature is verified: the edge verification gateway checks the TEE (Trusted Execution Environment) signature in the trusted data packet to determine whether it was generated by a trusted party through a legitimate and secure means, thereby confirming the authenticity and integrity of the data packet and deciding whether to proceed with further processing. This is generally done in the following steps:

[0065] Extraction from trusted data packets: The edge verification gateway first extracts the TEE signature and signature-related metadata, such as the public key and certificate, from the received trusted data packets. This information is fundamental to verifying the validity of the signature. The public key is used to subsequently verify whether the signature was generated by the corresponding private key, and the certificate is used to confirm the legitimacy of the public key and its owner.

[0066] Certificate Authority (CA) Verification: This checks whether the certificate was issued by a trusted Certificate Authority (CA). The edge verification gateway maintains a list of trusted CAs and compares the certificate authority of the retrieved certificate with this list. If the certificate authority is not on the trusted list, signature verification fails, the data packet may be untrusted, and further processing is stopped.

[0067] Verify that the certificate is valid. Certificates have specific start and end dates, and the edge verification gateway checks if the current time is within the certificate's validity period. If the certificate has expired, the signature is also considered invalid, and the data packet is untrustworthy.

[0068] Check if a certificate has been revoked by using a Certificate Revocation List (CRL) or Online Certificate Status Protocol (OCSP). Even if a certificate was issued by a trusted CA and is still valid, if it has been revoked, it means that the private key associated with the certificate may have been compromised or that there are other security issues, the signature is invalid, and the data packets are untrustworthy.

[0069] Subject Information Comparison: Extract the signing subject information (such as device identifier, organization name, etc.) from the certificate and compare it with the source subject information claimed by the trusted data packet. Ensure that the subject corresponding to the public key matches the source subject of the data packet. If they do not match, the signature verification fails, and the data packet may have a source fraud issue.

[0070] Obtaining the raw data hash value: The edge verification gateway uses the same hash algorithm as when generating the signature to perform a hash calculation on the combination of business data and environmental fingerprint in the trusted data packet, resulting in a hash value. This hash value represents a feature summary of the data packet content.

[0071] Decrypt the signature: Decrypt the TEE signature using the public key extracted from the certificate. If the signature was correctly generated using the corresponding private key, decryption will yield a hash value, which should match the hash value calculated in the previous step.

[0072] Hash value comparison: The hash value obtained from decrypting the signature is compared with the recalculated hash value. If the two are exactly the same, it indicates that the TEE signature is valid, that is, the data packet has not been tampered with during transmission and was indeed generated by a trusted party holding the corresponding private key; if the comparison is inconsistent, the signature verification fails, and the authenticity or integrity of the data packet is compromised.

[0073] Integrity and Authenticity Verification: Only when the certificate's legitimacy, the consistency between the public key and the signing subject, and the signature itself all pass verification can the edge verification gateway confirm the authenticity and integrity of the trusted data packet. At this point, the data packet is considered to have been generated by a trusted party through legitimate and secure means.

[0074] Decision on further processing: If signature verification passes, the edge verification gateway will continue to process the trusted data packet, such as extracting environmental fingerprints, comparing them with the device behavior baseline, and calculating real-time trust scores; if signature verification fails, the edge verification gateway will discard the data packet or take other security measures, such as recording abnormal logs and alerting relevant systems, to ensure the security and reliability of the entire data exchange process.

[0075] In this embodiment, real-time trust scoring involves the edge verification gateway comparing the environmental fingerprint extracted from the trusted data packet with the device behavior baseline, calculating the deviation of all dimensional feature values ​​from the baseline, and combining the scores with weights to reflect the current trustworthiness of the data packet, providing a basis for subsequent storage.

[0076] In this embodiment, the trust score threshold is a value used to compare with the real-time trust score. Only when the real-time trust score is higher than this trust score threshold will the trusted data packet enter the subsequent fragmented storage process based on storage attributes.

[0077] In this embodiment, storage attributes are attributes used to determine the storage method and location of data packets based on the static characteristics (such as data type, source, etc.) and dynamic characteristics (such as activity level, timeliness, etc.) of trusted data packets, through a series of operations such as feature parsing, calculation, clustering and sorting. These attributes include determining the storage layer and storage algorithm.

[0078] In this embodiment, trusted data packets are fragmented and stored in corresponding storage layers according to their storage attributes to obtain storage results: The hierarchical fragmented storage end divides trusted data packets into different parts according to the determined storage attributes and stores them in the matching storage layer (such as high-speed storage layer, large-capacity storage layer, etc.). The corresponding storage algorithm is used for storage, and the storage results are obtained after completion, so as to realize reasonable data storage and management.

[0079] In this embodiment, the consensus algorithm is a dynamically selected algorithm by the elastic consensus selector based on the current network load. This algorithm enables nodes in the blockchain network to reach consensus on the state and order of trusted data packets, ensuring data consistency and verifiability in a cross-chain environment. Assume a cross-chain data exchange scenario involving financial institutions, supply chain companies, and regulatory authorities, with multiple nodes jointly maintaining the blockchain network. The elastic consensus selector will dynamically select a suitable consensus algorithm based on the current network load.

[0080] When the number of nodes in the network is relatively small (e.g., only 10 nodes), the average transmission latency is very low (e.g., each data transmission takes only 0.1 seconds on average), and the probability of transaction conflicts is also very low (e.g., only 1 conflict in every 1000 transactions), the elastic consensus mechanism may choose the Proof-of-Stake (PoS) algorithm. The PoS algorithm is relatively simple and efficient, and in this network environment, nodes can quickly reach consensus.

[0081] As business grows, the number of nodes in the network increases significantly to 100, the average transmission latency rises to 1 second, and the probability of transaction conflicts also increases to 5 conflicts per 100 transactions. At this point, the elastic consensus mechanism may switch to the PBFT (Practical Byzantine Fault Tolerance) algorithm. The PBFT algorithm can tolerate a certain number of malicious nodes and can still efficiently reach consensus even with a large number of nodes and high network latency.

[0082] For example, in a supply chain, multiple companies simultaneously conduct goods transactions and fund transfers, generating a large number of trusted data packets. The PBFT algorithm, through multiple rounds of message passing and verification between nodes, ensures that even if some nodes malfunction or engage in malicious behavior, normal nodes can still reach a consensus on the status (such as goods have been shipped, funds have been received) and order of these trusted data packets. It first achieves consensus among nodes through three stages: pre-preparation, preparation, and submission, ensuring the consistency and verifiability of data in complex cross-chain environments.

[0083] In this embodiment, cross-chain verifiable notarization refers to the generation of a verifiable notarization record across different blockchains based on a dynamically selected consensus algorithm for trusted data packets. This ensures the traceability and credibility of data exchange in a multi-chain environment, guaranteeing that the authenticity and integrity of the data can be verified across chains. When the elastic consensus selection endpoint selects a suitable consensus algorithm based on network load (e.g., PBFT algorithm in complex network scenarios), it processes the trusted data packets. For example, a trusted data packet representing a financial institution paying a supply chain company, after reaching consensus using the PBFT algorithm, will generate cross-chain verifiable notarization based on this.

[0084] This notarized record not only documents basic payment information (amount, time, and information of both parties), but also includes crucial information from the consensus process, such as the signatures of participating nodes and the timestamp of consensus achievement. This information is encrypted and linked using cryptographic techniques, forming an immutable record. In subsequent business processes, whether it's a supply chain company confirming receipt of payment or a regulatory audit by regulatory authorities, the authenticity and integrity of this transaction data must be verified. Because it generates a cross-chain verifiable notarized record, different blockchains can verify it through specific cross-chain protocols and interfaces.

[0085] For example, regulatory authorities can use cross-chain technology on their consortium blockchain to access evidence of payment transactions by financial institutions on the Ethereum blockchain. By verifying the signatures, timestamps, and information related to the consensus algorithm in the evidence records, they can confirm whether the payment transaction actually occurred and whether the data is complete and tamper-proof. This ensures the traceability and credibility of data exchange in a multi-chain environment, guaranteeing that the authenticity and integrity of the data can be verified across chains.

[0086] like Figure 2 As shown, in order to achieve the comparison between environmental fingerprints and device behavior baselines and calculate real-time trust scores by extracting environmental fingerprint feature values, collecting device behavior baselines, and calculating deviations and weights, a method for extracting environmental fingerprints and comparing them with corresponding device behavior baselines at the edge verification gateway and calculating real-time trust scores is further proposed, including:

[0087] Extracting feature values ​​of multi-dimensional feature terms of environmental fingerprints from trusted data packets;

[0088] The normal fluctuation range of each characteristic item of the corresponding device is collected by the sliding window algorithm and used as the baseline of the corresponding device behavior.

[0089] The deviation of the feature values ​​of all dimensions of the environmental fingerprint from the device behavior baseline is calculated, and the real-time trust score is calculated by combining the weights of all dimensions of the feature.

[0090] In this embodiment, feature values ​​of multi-dimensional feature items of the environmental fingerprint are extracted from the trusted data packet. The trusted data packet contains business data, environmental fingerprint, and digital signature. The environmental fingerprint consists of multi-dimensional feature items, such as device hardware information (model, CPU specifications, etc.), software environment information (operating system version, application version, etc.), and network environment information (IP address, network bandwidth, etc.). After parsing the environmental fingerprint from the trusted data packet, the specific feature value is obtained for each dimension of the feature item. For example, the feature value corresponding to the device model feature item might be "a certain brand and model of mobile phone", and the feature value of the operating system version feature item might be "Android 12", etc. These feature values ​​are used for subsequent comparison with the device behavior baseline.

[0091] In this embodiment, the corresponding device for the trusted data packet—that is, the device used to collect business data and generate the trusted data packet—is the carrier of the environment described by the environmental fingerprint. For example, in a medical data collection scenario, it might be a specific medical examination device; in financial transaction data collection, it might be a terminal computer used by bank staff or a user's mobile payment device, etc. The various characteristic information of this device constitutes part of the environmental fingerprint.

[0092] In this embodiment, the device's various characteristics refer to the different aspects of the device's features corresponding to trusted data packets. These include hardware characteristics (such as the device's processor model, memory capacity, and storage capacity), software characteristics (such as the installed operating system, various applications and their versions), and network-related characteristics (such as the device's MAC address and the type of network it is on). These characteristics collectively characterize the device's state and attributes, and are the fundamental elements for generating environmental fingerprints and determining the device's behavioral baseline.

[0093] In this embodiment, a sliding window algorithm is used to collect the normal fluctuation range of various characteristic items of the corresponding device for trusted data packets. In this scenario, a fixed-size time window is set based on a time series, and data on various characteristic items of the corresponding device for trusted data packets is collected within this window. As time progresses, the window moves continuously like a sliding mechanism, continuously collecting data. By analyzing this data, the fluctuation range of each characteristic item under normal conditions is determined. For example, the CPU utilization of the device over a period of time, as shown by data collected through the sliding window, has a normal fluctuation range between 30% and 60%. This range serves as part of the device behavior baseline for evaluating the reliability of subsequently collected data.

[0094] In this embodiment, the deviation of the feature values ​​of all dimensional feature items of the environmental fingerprint from the device behavior baseline is calculated, and the real-time trust score is calculated by combining the weights of all dimensional feature items: the feature values ​​of each dimensional feature item extracted from the environmental fingerprint are compared with the corresponding device behavior baseline (normal fluctuation range of each feature item) obtained by the sliding window algorithm, and the degree to which the feature value of each feature item deviates from the normal fluctuation range is calculated, i.e., the deviation, which specifically includes:

[0095] First, calculate the difference between the characteristic value and the midpoint of the normal fluctuation range. Divide the difference by half of the normal fluctuation range and convert it into a percentage to obtain the deviation.

[0096] Simultaneously, based on the importance of different dimensional features to data credibility, each feature is assigned a corresponding weight. The deviation of all dimensional features is combined with their respective weights to calculate a numerical value as a real-time trust score. For example, if the deviations of the three features—device model, operating system version, and network bandwidth—are 0.1, 0.2, and 0.15, respectively, and their weights are 0.3, 0.4, and 0.3, then the real-time trust score might be calculated using a formula similar to 0.1 × 0.3 + 0.2 × 0.4 + 0.15 × 0.3.

[0097] To achieve hierarchical and fragmented storage of trusted data packets with trust scores exceeding a trust score threshold based on storage attributes by parsing trusted data packet characteristics, calculating characteristic data, determining storage attributes, and executing storage, a hierarchical and fragmented storage endpoint is further proposed, including:

[0098] The feature parsing module is used to treat trusted data packets and all corresponding synchronous trusted data packets as all trusted data packets to be evaluated when the trust score is higher than the trust score threshold. It parses all static features and all dynamic feature dimensions of each trusted data packet to be evaluated, generates static feature vectors based on all static features of each trusted data packet to be evaluated, and performs sliding time window sampling on each dynamic feature dimension of each trusted data packet to be evaluated and statistically obtains the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated.

[0099] The feature operation module is used to calculate the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated and the storage attraction between every two trusted data packets to be evaluated, based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated and the numerical sequence of all sliding time window samples of each dynamic feature dimension of each trusted data packet to be evaluated.

[0100] The attribute determination module is used to determine the storage attributes of the current trusted data packet based on the structural entropy of all dynamic feature dimensions of each trusted data packet to be evaluated, the storage attraction between every two trusted data packets to be evaluated, and the static feature vector of each trusted data packet to be evaluated, and to determine the storage layer and storage algorithm of the current trusted data packet based on the storage attributes of the current trusted data packet.

[0101] The storage execution module is used to store the current trusted data packet based on the storage algorithm and storage layer, and obtain the storage result.

[0102] In this embodiment, trusted data packets and all corresponding synchronous trusted data packets refer to the following: a trusted data packet is a data unit composed of business data, an environmental fingerprint, and a digital signature of the combination of the two. All corresponding synchronous trusted data packets refer to other trusted data packets generated from similar environments or related devices within the same or similar time period. For example, in a distributed data acquisition system, multiple devices simultaneously collect different parts of business data and generate their own trusted data packets. These data packets belong to the same batch of synchronous trusted data packets, and they may be correlated and analyzed together in subsequent processing to determine storage strategies, etc.

[0103] In this embodiment, all static features and all dynamic feature dimensions of each trusted data packet to be evaluated are analyzed: the trusted data packet to be evaluated refers to the aforementioned trusted data packet and the synchronous trusted data packet. Static features refer to characteristics that do not change frequently with time or other factors, such as data source identifier, data type (text, image, or numerical value, etc.), and business category. Dynamic feature dimensions reflect aspects of data changes over time or under other conditions, such as data update frequency and data volume trends.

[0104] In this embodiment, a sliding time window sampling method is used for each dynamic feature dimension of each trusted data packet to be evaluated, and the probability distribution of values ​​for all dynamic feature dimensions of each trusted data packet to be evaluated is statistically obtained: For each dynamic feature dimension of each trusted data packet to be evaluated, a sliding time window sampling method is adopted. That is, a fixed-duration time window is set, and the window moves continuously as time progresses, recording the value of the dynamic feature dimension within the window at each move. For example, for the dynamic feature dimension of data update frequency, an hour is used as a sliding time window, and the number of data updates per hour is recorded. By statistically analyzing the sampled data of multiple sliding time windows, the probability of different values ​​of the dynamic feature dimension occurring is calculated, thereby obtaining the probability distribution of values. For example, statistical analysis shows that for the data update frequency dimension, the probability of updating 1-2 times per hour is 60%, and the probability of updating 3-4 times per hour is 30%, etc.

[0105] In this embodiment, the storage layer and storage algorithm of the current trusted data packet are determined based on its storage attributes. These storage attributes are characteristics determined by comprehensively considering the static and dynamic features of the trusted data packet, along with related calculation results (such as structural entropy and storage attraction). A suitable storage layer is selected based on these attributes. Storage layers may include different types such as high-speed storage layers (for storing data requiring fast access) and large-capacity storage layers (suitable for storing large amounts of data with relatively low access frequency). Simultaneously, corresponding storage algorithms are determined, such as selecting specific compression algorithms, encryption algorithms, and other storage operation methods for different data structures and security requirements.

[0106] In this embodiment, the current trusted data packet is stored based on the storage algorithm and storage layer to obtain the storage result: The storage operation is performed on the current trusted data packet according to the previously determined storage algorithm and storage layer. For example, if the storage layer is a large-capacity storage layer and the storage algorithm uses a specific compression and encryption algorithm, the data packet is first compressed and encrypted, and then stored in a designated location on the large-capacity storage device. After the storage operation is completed, the storage result is obtained.

[0107] To calculate information entropy, storage gravity, and structural entropy to perform dynamic feature calculations on trusted data packets and provide a basis for determining storage attributes, a feature calculation module is further proposed, including:

[0108] The information entropy calculation submodule is used to calculate the information entropy of each dynamic feature dimension of each trusted data packet to be evaluated based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated.

[0109] The storage gravity calculation submodule is used to calculate the storage gravity between two trusted data packets to be evaluated based on the information entropy of each pair of trusted data packets to be evaluated and the feature space distance between the current values ​​of all dynamic feature dimensions of the corresponding two trusted data packets to be evaluated.

[0110] The structural entropy calculation submodule is used to calculate the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated based on the k-nearest neighbor algorithm and the numerical sequence of all sliding time windows sampled for each dynamic feature dimension of each trusted data packet to be evaluated.

[0111] In this embodiment, the information entropy of each dynamic feature dimension of each trusted data packet to be evaluated is calculated based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated:

[0112] Information entropy is a metric in information theory used to measure the uncertainty or randomness of information. For a given dynamic feature dimension of each trustworthy data packet to be evaluated, its probability distribution reflects the likelihood of different values ​​occurring in that dimension. Information entropy is calculated using the formula for calculating information entropy (emphasizing the common Shannon entropy formula). For example, for the dynamic feature dimension of data update frequency, if the probability of "updating once per hour" is 0.3, "updating twice per hour" is 0.5, and "updating three times per hour" is 0.2, the information entropy of that dynamic feature dimension can be calculated by substituting these values ​​into the formula.

[0113] In this embodiment, the storage attraction between two trusted data packets to be evaluated is calculated based on the information entropy of each pair of trusted data packets to be evaluated and the feature space distance between the current values ​​of all dynamic feature dimensions of the corresponding two trusted data packets to be evaluated:

[0114] Euclidean distance can be used to calculate the distance in the feature space;

[0115] The quotient of the product of the mean of the information entropy of all dynamic feature dimensions of each pair of trusted data packets to be evaluated and the square of the feature space distance between the two pairs of trusted data packets to be evaluated is taken as the storage attraction between the two pairs of trusted data packets to be evaluated.

[0116] In this embodiment, the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated is calculated based on the k-nearest neighbor algorithm and the numerical sequence of all sliding time window samples for each dynamic feature dimension of each trusted data packet to be evaluated:

[0117] The k-nearest neighbor algorithm is used to find the k most similar points in a dataset to a given data point. For each dynamic feature dimension of a trusted data packet to be evaluated, its sliding time window sampled numerical sequence contains a series of data points that change over time for that dimension. Using the k-nearest neighbor algorithm, the relationship between each data point in these numerical sequences and its k neighbors is analyzed, thereby calculating the dynamic feature dimension. structural entropy :

[0118] ;

[0119] In the formula, This represents the total number of dynamic feature dimensions. Indicates the first Number of samples in each neighborhood This represents the total number of samples (which is the total number of samples in all neighborhoods of all dynamic feature dimensions). Represents the logarithmic function with base 2;

[0120] Each data point is converted into a two-dimensional vector using its numerical value and acquisition time, and the distance between each pair of data points is calculated based on the two-dimensional vector and the Euclidean distance formula.

[0121] Centered on the current data point, divide the data into three neighborhoods with a radius of m (e.g., three) nearest neighbors, and count the number of data points covered in each neighborhood as the sample number in the corresponding neighborhood.

[0122] Structural entropy reflects the complexity or regularity of the structure of data in a dynamic feature dimension. For example, if a data sequence in a dynamic feature dimension exhibits relatively regular changes, the structural entropy calculated using the k-nearest neighbor algorithm may be low, indicating that its structure is relatively simple and highly regular; conversely, if the data sequence changes in a complex manner without obvious regularity, the structural entropy will be high. This is of great significance for understanding the internal structure of data, determining the storage attributes and storage methods of data packets, as data packets with similar structural entropy may have similar storage requirements.

[0123] To achieve cluster partitioning and thermal sorting based on structural entropy and storage gravity, and subsequently determine the storage attributes, storage layer, and storage algorithm of trusted data packets, an attribute determination module is further proposed, including:

[0124] The Gravity Clustering Submodule is used to cluster all trusted data packets to be evaluated based on the structural entropy of all dynamic feature dimensions of each trusted data packet to be evaluated and the storage gravity between every two trusted data packets to be evaluated, thereby obtaining multiple data packet clusters.

[0125] The thermal sorting submodule is used to perform thermal sorting on all data packet clusters and obtain the thermal sorting value of all data packet clusters.

[0126] The storage attribute determination submodule is used to determine the storage attribute of the current trusted data packet based on the heat ranking value and static feature vector of the data packet cluster to which the current trusted data packet belongs, and to determine the storage layer and storage algorithm of the current trusted data packet based on the storage attribute of the current trusted data packet.

[0127] To obtain a reasonable data packet cluster partitioning based on storage gravity merging and splitting clusters, thus preparing for subsequent thermal sorting and storage attribute determination, a gravity clustering submodule is further proposed, including:

[0128] The first clustering unit is used to merge two corresponding trusted data packets to be evaluated with a storage gravity greater than a preset gravity threshold as an initial cluster, continue to calculate the storage gravity between each pair of initial clusters, and continue to merge two corresponding initial clusters with a storage gravity greater than a preset gravity threshold as a merged cluster. At the same time, the average structural entropy of the structural entropy of all dynamic feature dimensions of all trusted data packets to be evaluated contained in each merged cluster is calculated.

[0129] The second clustering unit is used to determine whether there are any merged clusters whose average structural entropy exceeds the structural entropy threshold. If so, the corresponding merged cluster is split, and the split clusters obtained after splitting are merged with all the remaining merged clusters until there are no cluster pairs with a storage gravity greater than the preset gravity threshold. In this case, multiple data packet clusters are obtained. Otherwise, all merged clusters are merged until there are no cluster pairs with a storage gravity greater than the preset gravity threshold. In this case, multiple data packet clusters are obtained.

[0130] In this embodiment, a preset gravity threshold is used: this is a pre-set fixed value used to determine whether the storage gravity between two trusted data packets to be evaluated or two initialization clusters is large enough to decide whether to merge them. For example, if the preset gravity threshold is set to 0.6, and the calculated storage gravity between two trusted data packets to be evaluated is 0.7, which is greater than the threshold, then the two data packets will be merged into one initialization cluster.

[0131] In this embodiment, the storage attraction between every two initialization clusters is calculated: an initialization cluster is formed by merging two corresponding trusted data packets to be evaluated whose storage attraction is greater than a preset attraction threshold. For these initialization clusters, the storage attraction between each pair of clusters is also calculated. The calculation method is similar to that for calculating the storage attraction between two trusted data packets to be evaluated: the mean of the information entropy of all dynamic feature dimensions of all data packets in each initialization cluster is taken as the information entropy of each initialization cluster, and the ratio of the product of the information entropy of every two initialization clusters to the feature space distance between the current values ​​of all dynamic feature dimensions is taken as the storage attraction between the corresponding two initialization clusters.

[0132] In this embodiment, the average structural entropy of the structural entropy of all dynamic feature dimensions of all trusted data packets to be evaluated contained in each merged cluster is calculated. After merging initial clusters with storage gravity greater than a preset gravity threshold to form merged clusters, each merged cluster needs to be analyzed. For all trusted data packets to be evaluated contained in the merged cluster, the structural entropy of all dynamic feature dimensions of each data packet is first calculated, and then these structural entropy values ​​are averaged to obtain the average structural entropy. This average structural entropy reflects the overall level of structural complexity of the data packets in the dynamic feature dimensions within the merged cluster. For example, if there are 5 trusted data packets to be evaluated in a merged cluster, the structural entropy of each dynamic feature dimension of each data packet is calculated separately, and then these structural entropy values ​​are added together and divided by 5 to obtain the average structural entropy, which is used for subsequent evaluation and processing of the merged cluster.

[0133] In this embodiment, the structural entropy threshold is a pre-set value used to determine whether the average structural entropy of a merged cluster exceeds a reasonable range. When the average structural entropy of a merged cluster exceeds the structural entropy threshold, it indicates that the dynamic feature dimension structure of the data packets within that cluster may be too complex or inconsistent, requiring splitting to obtain a more reasonable data packet cluster division. For example, if the structural entropy threshold is set to 0.8, and the calculated average structural entropy of a merged cluster is 0.9, which is greater than the threshold, then this merged cluster needs to be split.

[0134] In this embodiment, split clusters are formed when the average structural entropy of a merged cluster exceeds a structural entropy threshold. To make the data packet clustering more reasonable, the merged cluster is split into multiple smaller clusters, which are called split clusters. The splitting method may be based on factors such as differences in the dynamic feature dimension of data packets and the distribution of structural entropy. For example, it is found that some data packets in the merged cluster exhibit significantly different patterns of change in the dynamic feature dimension of "data update frequency" (taking the frequency of user-published content as an example). Some data packets correspond to users who publish content very frequently, almost multiple times a day, while other data packets correspond to users who publish content relatively infrequently, perhaps only once or twice a week. Based on this obvious difference, the merged cluster can be split into two split clusters. The first split cluster contains data packets corresponding to users who publish content frequently, and the second split cluster contains data packets corresponding to users who publish content less frequently. In this way, the data packets in each split cluster are more similar in terms of the dynamic feature of "data update frequency". For example, the data packets in split cluster 1 correspond to users whose content publication frequency is mostly concentrated between 2-5 times per day, while the data packets in split cluster 2 correspond to users whose content publication frequency is mostly concentrated between 1-2 times per week.

[0135] In addition to data update frequency, we also focused on two dynamic feature dimensions: "like frequency" and "comment frequency." First, we calculated the structural entropy of each data packet along these two dimensions. Structural entropy reflects the structural complexity or regularity of the data in these dimensions. Calculations revealed that the distribution of structural entropy across these two dimensions in the merged cluster exhibited two different patterns. For one group of data packets, the like and comment frequencies showed a strong correlation—high like frequency was accompanied by high comment frequency—and the structural entropy was relatively low, indicating a relatively regular relationship. For another group of data packets, there was no significant correlation between like and comment frequencies, and the structural entropy was relatively high.

[0136] Based on this difference in structural entropy distribution, the merged clusters are split according to the structural entropy of like and comment frequencies. Data packets with a strong correlation between like and comment frequencies (low structural entropy) are assigned to one split cluster, while data packets with a weak correlation (high structural entropy) are assigned to another. For example, data packets in split cluster A have a correlation coefficient of 0.8 between like and comment frequencies and a structural entropy between 0.3 and 0.5; while data packets in split cluster B have a correlation coefficient of only 0.2 and a structural entropy between 0.7 and 0.9. This splitting makes the data packets within each split cluster more similar in their relationship across these two dynamic feature dimensions, which helps in subsequent more accurate data analysis and storage management.

[0137] In this embodiment, the split clusters obtained after splitting and all remaining merged clusters are merged again until there are no cluster pairs with a storage gravity greater than a preset gravity threshold. Multiple data packet clusters are then obtained. After splitting the merged clusters that exceed the structural entropy threshold, these split clusters are combined with other merged clusters that have not yet been further processed, and the storage gravity between each pair is recalculated. If there are cluster pairs with a storage gravity greater than the preset gravity threshold (which could be between split clusters, between split clusters and merged clusters, or between merged clusters), they are merged. This process is repeated continuously, merging cluster pairs, until the storage gravity between all pairs of clusters is no greater than the preset gravity threshold. The resulting multiple clusters are the final data packet clusters.

[0138] To achieve heat ranking of data packet clusters by determining their activity, timeliness, and importance, and calculating their heat values, a heat ranking submodule is further proposed, including:

[0139] The cluster feature extraction unit is used to determine the activity, timeliness, and importance of each data packet cluster;

[0140] The heat value determination unit is used to determine the heat value of each data packet cluster based on its activity, timeliness, and importance.

[0141] The thermal sorting unit is used to sort all data packet clusters thermally according to the principle of thermal value from largest to smallest, and obtain the thermal sorting value of all data packet clusters.

[0142] In this embodiment, the activity, timeliness, and importance of each data packet cluster are determined:

[0143] Activity level: Reflects the activity level of data within a data packet cluster, and can be measured in various ways. For example, the proportion of accesses to this data packet cluster within a day to the total accesses to the platform can be used as the activity level of the data packet cluster.

[0144] Timeliness: Reflects the timeliness of data within a data packet cluster. It primarily considers how the value of data to the business changes over time. The ratio of data volume published within 24 hours to the total data volume of the platform is used as the timeliness of the data packet cluster.

[0145] Importance: This assesses the degree of importance of data within a data packet cluster within the overall business process. This may depend on factors such as the business area the data relates to and its impact on business decisions. For example, data packets containing core financial data or key business process control instructions have high importance; while auxiliary or reference data have relatively lower importance. Taking an Enterprise Resource Planning (ERP) system as an example, data on production orders and raw material procurement are crucial to business operations, and their data packet clusters have higher importance than those containing employee attendance records. The average importance of the data packet cluster is calculated by taking the original importance values ​​of all data packets within it.

[0146] In this embodiment, the heat value of each data packet cluster is determined based on its activity, timeliness, and importance. Different weights are assigned to activity, timeliness, and importance (set according to business needs and data characteristics), and then their weighted sum is calculated; the result is the heat value. For example, assuming the activity weight is 0.3, the timeliness weight is 0.3, and the importance weight is 0.4, for a certain data packet cluster, its activity score is 8 (out of 10), its timeliness score is 7, and its importance score is 9. Then, the heat value of this data packet cluster is 0.3×8 + 0.3×7 + 0.4×9 = 7.5. A higher heat value indicates a higher overall importance and level of attention for the data packet cluster in the entire system. Subsequent operations such as sorting data packet clusters based on the heat value can then be performed to better manage and store data.

[0147] To determine the storage attributes, storage layer, and storage algorithm of trusted data packets by generating storage attribute-based features and retrieving a mapping table, a storage attribute determination submodule is further proposed, including:

[0148] Based on the feature aggregation unit, the storage attribute of the current trusted data packet is generated based on the heat ranking value of the data packet cluster to which the current trusted data packet belongs and the static feature vector.

[0149] The mapping table retrieval unit is used to retrieve the storage attribute based on the feature-attribute feature mapping table of the current trusted data packet, determine the storage attribute of the current trusted data packet, and determine the storage layer and storage algorithm of the current trusted data packet based on the storage attribute of the current trusted data packet.

[0150] In this embodiment, a storage attribute-based feature-attribute feature mapping table is used. This is a pre-built mapping table that records the correspondence between different storage attribute-based features and their corresponding attribute features. The attribute features are storage-related, including information such as the suitable storage layer (e.g., high-speed storage layer, conventional storage layer, large-capacity storage layer, etc.) and the storage algorithm to be used (e.g., specific encryption algorithm, compression algorithm, etc.). For example, a storage attribute-based feature of "high heat index and data type is financial transaction record" might correspond to the attribute feature "high-speed storage layer and high-strength encryption algorithm" in the mapping table. By establishing this mapping relationship, the system can quickly find the appropriate storage solution based on the storage attribute-based features.

[0151] In this embodiment, the storage attributes of the current trusted data packet are determined by retrieving the storage attribute based on the feature-attribute feature mapping table.

[0152] After obtaining the storage attribute criteria of the current trusted data packet, the system uses this as an index to search in the storage attribute criteria-attribute feature mapping table. For example, if the storage attribute criteria indicate that the data packet belongs to a cluster of high-frequency and sensitive data types, the system will retrieve the corresponding attribute feature in the mapping table as "suitable for storage in a high-security storage layer, and using encryption and compression algorithms".

[0153] To determine network load based on storage results, a consensus algorithm is selected and switched at the next block height using a consensus decision tree. Simultaneously, the queue of unconfirmed transactions is frozen to generate cross-chain verifiable evidence. Furthermore, a flexible consensus selection mechanism is proposed, including:

[0154] The network load determination module is used to determine the current network load based on the stored results, where the current network load includes the number of nodes, average transmission latency, and transaction conflict probability.

[0155] The consensus algorithm decision module is used to input the current network load into the consensus decision tree to select a consensus algorithm;

[0156] The consensus switching execution module is used to switch the consensus algorithm at the next block height, while freezing the queue of unconfirmed transactions until the switch is completed and cross-chain verifiable evidence is generated.

[0157] In this embodiment, the number of nodes, average transmission latency, and transaction conflict probability are as follows:

[0158] Number of nodes: This refers to the total number of devices or computers participating in data processing, storage, and verification operations within a blockchain network. The number of nodes has a significant impact on network performance and security.

[0159] Average transmission latency refers to the average time it takes for data (such as transaction information, block data, etc.) to be transmitted from one node to another in a blockchain network.

[0160] Transaction conflict probability: In a blockchain network, a conflict may occur when multiple transactions attempt to modify the same data or resource simultaneously. The transaction conflict probability is a metric that measures the likelihood of such a conflict occurring.

[0161] In this embodiment, the current network load is input into the consensus decision tree to select a consensus algorithm: network load-related data is provided as input to the consensus decision tree, which analyzes and judges this data according to preset rules and algorithms. For example, if the number of nodes is large and the average transmission latency is low, but the probability of transaction conflicts is high, the decision tree may determine based on these conditions that the current network is more suitable for a consensus algorithm that can effectively handle high-concurrency transaction conflicts.

[0162] In this embodiment, a consensus decision tree is used to select a consensus algorithm. It is organized in a tree structure. The nodes of the tree contain judgments related to network load conditions, such as whether the number of nodes exceeds a certain threshold, whether the average transmission delay is within a specific range, and the probability of transaction conflicts. Starting from the root node, the conditions are judged based on the current network load data, and the tree extends downwards along branches that meet the conditions until it reaches the leaf nodes. The leaf nodes correspond to specific consensus algorithms. For example, the root node judges the number of nodes; if the number of nodes is greater than 100, it enters a branch, and then judges the average transmission delay under that branch. If the delay is less than a certain value, it enters another branch, finally reaching the leaf node to determine the specific consensus algorithm to be used, such as the Practical Byzantine Fault Tolerance (PBFT) algorithm.

[0163] In this embodiment, the consensus algorithm is switched at the next block height, while the queue of unconfirmed transactions is frozen until the switch is complete and cross-chain verifiable evidence is generated. In a blockchain, block height refers to the sequential numbering of blocks in the blockchain according to their creation order. When the consensus decision tree determines that a consensus algorithm switch is necessary, the switch is performed at the next block height. This ensures that the algorithm transition is completed at a relatively clear point in time, while maintaining the consistency and continuity of blockchain data.

[0164] During the consensus algorithm switch, to avoid confusion caused by unconfirmed transactions under different consensus rules, the system freezes the unconfirmed transaction queue and suspends processing of these transactions. The queue is unfrozen and processing resumes only after the consensus algorithm switch is complete and the new algorithm is running stably. After the consensus algorithm switch is completed and transactions are successfully processed, cross-chain verifiable notarization is generated. Cross-chain verifiable notarization enables the verification of the authenticity and integrity of transaction data across different blockchains, ensuring the credibility and traceability of data in a cross-chain environment, thereby guaranteeing the smooth operation of the entire blockchain-driven data exchange process.

[0165] This invention provides an implementation method for a blockchain-driven end-to-end data exchange monitoring method, comprising:

[0166] Collect business data and generate environmental fingerprints simultaneously, and digitally sign the combination of business data and environmental fingerprints to generate trusted data packets;

[0167] Receive trusted data packets and verify the validity of the TEE signature. When the validity of the TEE signature is verified, extract the environmental fingerprint, compare it with the corresponding device behavior baseline, and calculate the real-time trust score.

[0168] When the trust score is higher than the trust score threshold, the trusted data packet is fragmented and stored in the corresponding storage layer according to the storage attribute to obtain the storage result;

[0169] Based on the storage results, the current network load is determined, and a consensus algorithm is dynamically selected to generate cross-chain verifiable evidence for trusted data packets.

[0170] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A blockchain-driven end-to-end data exchange monitoring system, characterized in that: include: The trusted data acquisition terminal is used to collect business data and synchronously generate environmental fingerprints, and digitally sign the combination of business data and environmental fingerprints to generate trusted data packets; The edge verification gateway is used to receive trusted data packets and verify the validity of TEE signatures. When the validity of the TEE signature is verified, the environmental fingerprint is extracted, compared with the corresponding device behavior baseline, and a real-time trust score is calculated. The hierarchical and fragmented storage layer is used to fragment trusted data packets according to storage attributes and store them in the corresponding storage layer when the trust score is higher than the trust score threshold, thereby obtaining the storage result. The elastic consensus selection end is used to determine the current network load based on the storage results, and dynamically select a consensus algorithm to generate cross-chain verifiable evidence for trusted data packets based on the current network load. The tiered and fragmented storage end includes: The feature parsing module is used to treat trusted data packets and all corresponding synchronous trusted data packets as all trusted data packets to be evaluated when the trust score is higher than the trust score threshold. It parses all static features and all dynamic feature dimensions of each trusted data packet to be evaluated, generates static feature vectors based on all static features of each trusted data packet to be evaluated, and performs sliding time window sampling on each dynamic feature dimension of each trusted data packet to be evaluated and statistically obtains the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated. The feature operation module is used to calculate the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated and the storage attraction between every two trusted data packets to be evaluated, based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated and the numerical sequence of all sliding time window samples of each dynamic feature dimension of each trusted data packet to be evaluated. The attribute determination module is used to determine the storage attributes of the current trusted data packet based on the structural entropy of all dynamic feature dimensions of each trusted data packet to be evaluated, the storage attraction between every two trusted data packets to be evaluated, and the static feature vector of each trusted data packet to be evaluated, and to determine the storage layer and storage algorithm of the current trusted data packet based on the storage attributes of the current trusted data packet. The storage execution module is used to store the current trusted data packet based on the storage algorithm and storage layer, and obtain the storage result; The feature operation module includes: The storage gravity calculation submodule is used to calculate the storage gravity between two trusted data packets to be evaluated based on the information entropy of each pair of trusted data packets to be evaluated and the feature space distance between the current values ​​of all dynamic feature dimensions of the corresponding two trusted data packets to be evaluated. The structural entropy calculation submodule is used to calculate the structural entropy of each dynamic feature dimension of each trusted data packet to be evaluated based on the k-nearest neighbor algorithm and the numerical sequence of all sliding time windows sampled for each dynamic feature dimension of each trusted data packet to be evaluated.

2. The blockchain-driven data exchange end-to-end monitoring system according to claim 1, characterized in that, A method for extracting environmental fingerprints at the edge verification gateway, comparing them with the corresponding device behavior baseline, and calculating a real-time trust score includes: Extracting feature values ​​of multi-dimensional feature terms of environmental fingerprints from trusted data packets; The normal fluctuation range of each characteristic item of the corresponding device is collected by the sliding window algorithm and used as the baseline of the corresponding device behavior. The deviation of the feature values ​​of all dimensions of the environmental fingerprint from the device behavior baseline is calculated, and the real-time trust score is calculated by combining the weights of all dimensions of the feature.

3. The blockchain-driven data exchange end-to-end monitoring system according to claim 1, characterized in that, The feature calculation module also includes: The information entropy calculation submodule is used to calculate the information entropy of each dynamic feature dimension of each trusted data packet to be evaluated based on the probability distribution of the values ​​of all dynamic feature dimensions of each trusted data packet to be evaluated.

4. The blockchain-driven data exchange end-to-end monitoring system according to claim 1, characterized in that, The attribute determination module includes: The Gravity Clustering Submodule is used to cluster all trusted data packets to be evaluated based on the structural entropy of all dynamic feature dimensions of each trusted data packet to be evaluated and the storage gravity between every two trusted data packets to be evaluated, thereby obtaining multiple data packet clusters. The thermal sorting submodule is used to perform thermal sorting on all data packet clusters and obtain the thermal sorting value of all data packet clusters. The storage attribute determination submodule is used to determine the storage attribute of the current trusted data packet based on the heat ranking value and static feature vector of the data packet cluster to which the current trusted data packet belongs, and to determine the storage layer and storage algorithm of the current trusted data packet based on the storage attribute of the current trusted data packet.

5. The blockchain-driven data exchange end-to-end monitoring system according to claim 4, characterized in that, The gravity clustering submodule includes: The first clustering unit is used to merge two corresponding trusted data packets to be evaluated with a storage gravity greater than a preset gravity threshold as an initial cluster, continue to calculate the storage gravity between each pair of initial clusters, and continue to merge two corresponding initial clusters with a storage gravity greater than a preset gravity threshold as a merged cluster. At the same time, the average structural entropy of the structural entropy of all dynamic feature dimensions of all trusted data packets to be evaluated contained in each merged cluster is calculated. The second clustering unit is used to determine whether there are any merged clusters whose average structural entropy exceeds the structural entropy threshold. If so, the corresponding merged cluster is split, and the split clusters obtained after splitting are merged with all the remaining merged clusters until there are no cluster pairs with a storage gravity greater than the preset gravity threshold. In this case, multiple data packet clusters are obtained. Otherwise, all merged clusters are merged until there are no cluster pairs with a storage gravity greater than the preset gravity threshold. In this case, multiple data packet clusters are obtained.

6. The blockchain-driven data exchange end-to-end monitoring system according to claim 4, characterized in that, The thermal sorting submodule includes: The cluster feature extraction unit is used to determine the activity, timeliness, and importance of each data packet cluster; The heat value determination unit is used to determine the heat value of each data packet cluster based on its activity, timeliness, and importance. The thermal sorting unit is used to sort all data packet clusters thermally according to the principle of thermal value from largest to smallest, and obtain the thermal sorting value of all data packet clusters.

7. The blockchain-driven data exchange end-to-end monitoring system according to claim 4, characterized in that, The storage attribute determination submodule includes: Based on the feature aggregation unit, the storage attribute of the current trusted data packet is generated based on the heat ranking value of the data packet cluster to which the current trusted data packet belongs and the static feature vector. The mapping table retrieval unit is used to retrieve the storage attribute based on the feature-attribute feature mapping table of the current trusted data packet, determine the storage attribute of the current trusted data packet, and determine the storage layer and storage algorithm of the current trusted data packet based on the storage attribute of the current trusted data packet.

8. The blockchain-driven data exchange end-to-end monitoring system according to claim 1, characterized in that, Flexible consensus selection endpoints include: The network load determination module is used to determine the current network load based on the stored results, where the current network load includes the number of nodes, average transmission latency, and transaction conflict probability. The consensus algorithm decision module is used to input the current network load into the consensus decision tree to select a consensus algorithm; The consensus switching execution module is used to switch the consensus algorithm at the next block height, while freezing the queue of unconfirmed transactions until the switch is completed and cross-chain verifiable evidence is generated.

9. A blockchain-driven method for monitoring the entire data exchange process, characterized in that: A blockchain-driven data exchange end-to-end monitoring system applied to any one of claims 1 to 8, comprising: Collect business data and generate environmental fingerprints simultaneously, and digitally sign the combination of business data and environmental fingerprints to generate trusted data packets; Receive trusted data packets and verify the validity of the TEE signature. When the validity of the TEE signature is verified, extract the environmental fingerprint, compare it with the corresponding device behavior baseline, and calculate the real-time trust score. When the trust score is higher than the trust score threshold, the trusted data packet is fragmented and stored in the corresponding storage layer according to the storage attribute to obtain the storage result; Based on the storage results, the current network load is determined, and a consensus algorithm is dynamically selected to generate cross-chain verifiable evidence for trusted data packets.

Citation Information

Patent Citations

  • Electronic seal verification method and system based on block chain

    CN118277489A

  • Edge computing gateway security authentication system and method based on trusted computing and asymmetric encryption

    CN119743251A

  • Cache data security protection system based on cloud computing

    CN120217413A