A lightweight integrated data link traceability management method and system
By obtaining the implicit anomaly coefficients of the target subject's data, dynamically configuring the data integration rate and the number of verifications, and using sparse Merkle tree verification, the error risk and computing power waste in data chain traceability verification are solved, achieving efficient and reliable data chain traceability management.
Patent Information
- Application Number
- CN202511116623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Existing technologies for data chain traceability verification suffer from high verification error risk, significant waste of computing power, and difficulty in meeting timeliness requirements.
By acquiring the implicit anomaly coefficients of the target subject's data, dynamically configuring the data integration rate and the number of verifications, and employing sparse Merkle tree verification, a lightweight data chain traceability management method and system are constructed to achieve differentiated data processing and verification.
It significantly improves the efficiency of computing power utilization and the accuracy of verification in large-scale data traceability scenarios, reduces the waste of computing power in full-node verification, and achieves lightweight operation while ensuring the reliability of traceability.
Smart Images

Figure CN120896754B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data security technology, specifically to a lightweight integrated data chain traceability management method and system. Background Technology
[0002] In the current field of data traceability management, blockchain-based distributed technology is widely used for data chain traceability in scenarios such as supply chains and financial transactions due to its immutability. Existing technologies typically employ a full verification mechanism for data chain traceability verification, but this mechanism has significant drawbacks. For example, a large number of verifications may increase the probability of verification errors and potentially waste computing power. Furthermore, the massive computational load of verification may be insufficient to meet the timeliness requirements of data verification. Summary of the Invention
[0003] This application provides a lightweight integrated data chain traceability management method and system to address the technical problems of insufficient data chain verification reliability and low integrated verification efficiency in the prior art.
[0004] In view of the above problems, this application provides a lightweight and integrated data chain traceability management method and system.
[0005] Firstly, this application provides a lightweight, integrated data chain traceability management method, the method comprising:
[0006] Obtain the subject data of multiple target subjects, and obtain the data implicit anomaly coefficients of the multiple target subjects;
[0007] Based on multiple hidden anomaly coefficients, configure the data integration rate of multiple target subjects, and randomly select data from multiple subjects according to the multiple data integration rates to construct an integrated data chain sequence, wherein random selection is based on traversal constraints;
[0008] Based on the multiple data hidden anomaly coefficients, configure the number of traceability verifications, and randomly select the verification data chain that obtains the number of traceability verifications.
[0009] Based on the data chain sequence, multiple data selection rates for multiple target subjects are calculated. Combining the multiple data hidden anomaly coefficients, subject data within each verification data chain is randomly selected for sparse Merkle tree verification to obtain data verification coefficients, which serve as traceability management results.
[0010] Secondly, this application provides a lightweight integrated data chain traceability management system, including:
[0011] The data acquisition module is used to acquire the subject data of multiple target subjects and to acquire the data hidden anomaly coefficients of the multiple target subjects;
[0012] The random selection module is used to configure the data integration rate of multiple target subjects based on multiple data hidden anomaly coefficients, and randomly select data from multiple subjects according to multiple data integration rates to construct an integrated data chain sequence, wherein random selection is based on traversal constraints;
[0013] The verification quantity acquisition module is used to configure the traceability verification quantity based on the multiple data hidden anomaly coefficients, and randomly select the verification data chain to obtain the traceability verification quantity.
[0014] The traceability management module is used to calculate multiple data selection rates for multiple target subjects based on the data chain sequence, combine the multiple data hidden anomaly coefficients, randomly select subject data in each verification data chain, perform sparse Merkle tree verification, and obtain data verification coefficients as traceability management results.
[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0016] This application proposes a lightweight integrated data chain traceability management method and system. By dynamically sensing data anomaly risks and adaptively adjusting data integration and verification strategies, it significantly improves the computing power utilization efficiency and verification accuracy in large-scale data traceability scenarios. Compared with traditional methods, the technical solution provided in this application can significantly reduce the computing power waste caused by full-node verification while improving verification accuracy. It achieves lightweight operation while ensuring traceability reliability, and achieves the technical effect of dynamically balancing verification accuracy and verification efficiency with limited computing resources. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a lightweight integrated data chain traceability management method provided in an embodiment of this application.
[0019] Figure 2 This is a schematic diagram of a lightweight integrated data chain traceability management system provided in an embodiment of this application.
[0020] The components represented by each number in the attached diagram are explained below:
[0021] The system includes a data acquisition module (100), a random selection module (200), a verification quantity acquisition module (300), and a traceability management module (400). Detailed Implementation
[0022] This application provides a lightweight and integrated data chain traceability management method and system to address the technical problems of insufficient reliability of data chain verification and low efficiency of integrated verification in the prior art.
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0024] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0025] Example 1, as Figure 1 As shown, this application provides a lightweight integrated data chain traceability management method, wherein the method includes:
[0026] S10: Obtain the subject data of multiple target subjects, and obtain the data implicit anomaly coefficients of the multiple target subjects.
[0027] Traditional data traceability methods treat all entities uniformly, ignoring differences in individual historical abnormal behaviors. This results in high-risk and low-risk data receiving the same weight in subsequent processes. The lack of quantitative assessment of the credibility of entity data makes it difficult to predict potential risk distribution before data integration, leading to an imbalance in resource allocation. Low-risk data excessively consumes verification resources, while high-risk data that truly requires close monitoring may evade detection, ultimately reducing the overall verification accuracy of the traceability system.
[0028] Step S10 in the method provided in this application embodiment includes:
[0029] Acquire subject data for multiple target entities to be processed;
[0030] Based on the percentage of times that the target subjects' data anomalies occurred within a historical period, the implicit anomaly coefficients of the data for the multiple target subjects are calculated.
[0031] In this embodiment, multiple target entities' subject data to be processed are obtained. A target entity is the subject data of the target entity to be processed, such as multiple users or multiple enterprises. The subject data to be processed is the data that the target entity needs to store for subsequent verification, such as transaction data between multiple users or enterprises.
[0032] Based on the percentage of abnormal data tampering behavior verified over a past period for each target entity, multiple implicit anomaly coefficients are calculated. The implicit anomaly coefficient is calculated as: Implicit Anomaly Coefficient = Number of Abnormal Behaviors ÷ Total Number of Verifications. For example, if there are 30 abnormal behavior instances and 100 total verifications, the Implicit Anomaly Coefficient = 30 ÷ 100 = 0.3. The Implicit Anomaly Coefficient characterizes the data tampering activity of the target entity over a past period. The more times data has been tampered with in the past, the larger the Implicit Anomaly Coefficient, indicating a greater implicit risk of data tampering for the target entity, and requiring more traceability verification to ensure security.
[0033] By dynamically quantifying the implicit risk coefficient of subject data by the proportion of historical anomalies, an objective basis is provided for subsequent differentiated processing. The method provided in this application can pre-identify high-risk subjects and mark their anomaly probabilities, avoiding the resource misallocation problem caused by traditional homogenization processing, improving risk sensitivity from the source, and laying the foundation for data screening for subsequent lightweight integration and accurate verification.
[0034] S20: Based on multiple hidden anomaly coefficients, configure the data integration rate of multiple target subjects, randomly select data from multiple subjects according to the multiple data integration rates, and construct an integrated data chain sequence, wherein random selection is based on traversal constraints.
[0035] Current blockchain data integration uses a fixed-ratio strategy, which cannot dynamically adjust the data volume according to data risk. The excessive selection of low-risk data leads to redundant writing, while high-risk data may be missed, resulting in critical risk data not being included in the traceability chain and creating verification vulnerabilities.
[0036] Step S20 in the method provided in this application embodiment includes:
[0037] Based on multiple data hidden anomaly coefficients, the data integration rate of multiple target subjects is calculated, where each data integration rate includes the probability that the subject data of each target subject is selected;
[0038] Based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data, and a first data chain is constructed based on blockchain, wherein the preset number is less than the number of subjects of multiple target subjects;
[0039] Specifically, based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data sets, and a first data chain is constructed using blockchain, including:
[0040] Based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data.
[0041] Based on blockchain, a predetermined number of subject data will be constructed into data blocks to obtain the first data chain;
[0042] Multiple data chains are constructed to obtain an integrated data chain sequence, wherein the selection and construction are carried out according to traversal constraints, the traversal constraints including that each subject data appears at least once in the integrated data chain sequence.
[0043] In this embodiment, a data integration rate is calculated and allocated to multiple target subjects based on multiple data hidden anomaly coefficients. Each data integration rate includes the probability that the subject data of each target subject is selected. The larger the historical hidden anomaly coefficient, the greater the hidden risk of data tampering of the target subject. Therefore, a larger data integration rate is allocated to the target subject, increasing the probability that the corresponding target subject data is selected. Data integration rate = data hidden anomaly coefficient.
[0044] Based on multiple data integration rates, a preset number of subject data sets are randomly selected from multiple subject data sets. For example, when the data integration rate is 0.8, a preset number of subject data sets representing 80% of the total subject data volume is selected.
[0045] Based on blockchain, a predetermined number of subject data will be selected and constructed into data blocks to obtain the first data chain.
[0046] Multiple data chains are then constructed to obtain an integrated data chain sequence. The selection and construction are based on a traversal constraint: each principal data item must appear at least once within the integrated data chain sequence. This constraint ensures that every principal data item participates in the construction of the integrated data chain, preventing omissions.
[0047] The method provided in this application adaptively allocates the data integration rate based on the anomaly coefficient, realizing a dynamic sampling mechanism that integrates high-risk entities with a high probability and low-risk entities with a low probability. Combined with traversal constraints, it ensures that all entity data is integrated at least once, which significantly reduces redundant data and ensures that key risk data is prioritized for inclusion in the traceability chain, maintaining global coverage of the data chain while achieving lightweight storage.
[0048] S30: Configure the number of traceability verifications based on the multiple data hidden anomaly coefficients, and randomly select the verification data chain that obtains the number of traceability verifications.
[0049] Traditional verification methods employ a fixed number or proportion of data chains for verification, failing to consider the overall anomaly level of the current data chain. When the overall risk is low, over-verification can lead to wasted computing power; when the overall risk is high, insufficient verification will affect accuracy.
[0050] Step S30 in the method provided in this application embodiment includes:
[0051] Calculate the mean of the multiple data hidden anomaly coefficients to obtain the average data hidden anomaly coefficient;
[0052] Obtain the historical average data of all other subjects within a historical period, including the implied anomaly coefficient.
[0053] Get the preset number of traceability verifications;
[0054] Based on the ratio of the average data hidden anomaly coefficient to the historical average data hidden anomaly coefficient, the preset traceability verification quantity is adaptively adjusted to obtain the traceability verification quantity.
[0055] Within the integrated data chain sequence, a verification data chain is randomly selected to obtain the number of traceability verifications.
[0056] In this embodiment of the application, the arithmetic mean of multiple data hidden anomaly coefficients is calculated to obtain the average data hidden anomaly coefficient.
[0057] Based on past verification information, obtain the historical average data implied anomaly coefficients of all other subjects within a historical period.
[0058] Obtain the preset traceability verification quantity. The preset traceability verification quantity refers to the number of traceability verifications set in advance. For example, the preset traceability verification quantity is set to 40% of the data volume in the integrated data chain sequence. For example, if the data volume in the integrated data chain sequence is 1000, then the preset traceability verification quantity = 1000 × 40% = 400.
[0059] Based on the ratio of the average implied anomaly coefficient to the historical average implied anomaly coefficient, the preset traceability verification quantity is adaptively adjusted to obtain the traceability verification quantity. Traceability verification quantity = (Average implied anomaly coefficient ÷ Historical average implied anomaly coefficient) × Preset traceability verification quantity. For example, if the average implied anomaly coefficient is 0.1, the historical average implied anomaly coefficient is 0.2, and the preset traceability verification quantity is 400, then the traceability verification quantity = (0.1 ÷ 0.2) × 400 = 800.
[0060] Within the integrated data chain sequence, a verification data chain is randomly selected to obtain the number of traceable verifications.
[0061] The number of verification data chains is dynamically adjusted based on a comparison between the current average anomaly coefficient and historical benchmarks. When the overall risk increases, the verification quantity is automatically increased to enhance verification intensity; when the risk decreases, the verification quantity is reduced to save resources. This achieves an adaptive match between the verification scale and the system's real-time risk level, optimizing anomaly detection capabilities while controlling computational overhead.
[0062] S40: Based on the data chain sequence, calculate the selection rate of multiple data for multiple target subjects, combine the multiple data hidden anomaly coefficients, randomly select subject data in each verification data chain, perform sparse Merkle tree verification, and obtain the data verification coefficient as the traceability management result.
[0063] Traditional verification requires all data to participate in hash calculations, resulting in high computing power demands in large-scale traceability scenarios. Existing technologies use uniform verification of on-chain data, which cannot incorporate information such as data risk characteristics. This causes high-risk or high-frequency integrated data to escape verification due to random omissions, while low-risk data is repeatedly verified, reducing verification efficiency and accuracy.
[0064] Step S40 in the method provided in this application embodiment includes:
[0065] Based on the data chain sequence, the occurrence ratio of the subject data of each target subject is calculated as multiple data selection rates;
[0066] Multiple data validation rates are calculated based on multiple data hidden anomaly coefficients and multiple data selection rates;
[0067] Within the first verification data chain among multiple verification data chains, according to the multiple data verification rates, a preset number of subject data is randomly selected, and subject data corresponding to the target subject is randomly selected in other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient.
[0068] Within the first verification data chain of multiple verification data chains, a preset number of subject data are randomly selected according to the multiple data verification rates, and subject data corresponding to the target subject are randomly selected from other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient, including:
[0069] Within the first verification data chain among multiple verification data chains, a preset number of main data is randomly selected according to the multiple data verification rates, serving as the first benchmark main data for the preset number of verifications.
[0070] Within other data chains, randomly select the subject data corresponding to the target subject of the selected subject data to obtain the first verification subject data of the preset verification quantity;
[0071] Based on the first baseline main data with a preset number of verifications, a first baseline hash tree is constructed using a Merkle tree.
[0072] Based on the first verification subject data with preset verification data, a first verification hash tree is constructed using a Merkle tree;
[0073] Calculate the similarity between the first baseline hash tree and the first verification hash tree to obtain the first single-chain data verification coefficient;
[0074] Continue to perform sparse Merkle tree verification on multiple other verification data chains to obtain multiple single-chain data verification coefficients. Calculate the average to obtain the data verification coefficient, which serves as the result for traceability management.
[0075] In this embodiment, the occurrence ratio of the subject data for each target subject is calculated based on the data chain sequence, serving as multiple data selection rates. Data selection rate = Number of times the subject data appears in the data chain sequence ÷ Total number of data entries in the data chain sequence. For example, if the subject data appears 400 times in the data chain sequence and the total number of data entries is 1000, then the data selection rate is 400 ÷ 1000 = 0.4. A higher data selection rate indicates a greater probability of the subject data appearing in the data volume.
[0076] Multiple data validation rates are calculated based on multiple implied anomaly coefficients and multiple data selection rates. The data validation rate = (implied anomaly coefficient + selection rate) ÷ 2. For example, if a data point has an implied anomaly coefficient of 0.3 and a selection rate of 0.4, then the data validation rate = (0.3 + 0.4) ÷ 2 = 0.35. The data validation rate is a comprehensive analysis of the probability that each data point is selected during the validation process. A higher data validation rate indicates a greater probability that the data point will be selected and validated.
[0077] Within the first verification data chain of multiple verification data chains, subject data of a preset verification quantity from multiple subjects is randomly selected according to multiple data verification rates, serving as the first baseline subject data of the preset verification quantity. The data verification rate reflects the probability of each subject data being selected. The data verification rate is normalized: the normalized data verification rate = data verification rate ÷ the sum of all data verification rates. For example, if a subject data set is the transaction data between Company A and Company B in 2024, and its data verification rate is 0.35, and the sum of multiple data verification rates is 2.5, then the normalized data verification rate = 0.35 ÷ 2.5 = 0.14. The preset verification quantity is the number of data entries to be selected for verification. For example, the preset verification data can be set to 100 entries. Therefore, the number of transaction data entries between Company A and Company B in 2024 selected as the first baseline subject data = the normalized data verification rate result × the preset verification quantity = 0.14 × 100 = 14 entries. These 14 entries are randomly selected from the transaction data between Company A and Company B in 2024 and added to the first baseline subject data. Using the same processing method, within the first verification data chain of multiple verification data chains, multiple data points are selected and added to the first baseline main data according to multiple data verification rates.
[0078] Within other data chains, the subject data corresponding to the target subject of the selected subject data is randomly selected to obtain a preset number of first verification subject data. For example, if one of the selected subject data is the transaction data of company A in 2024, and the corresponding target subject is the transaction data of company A, then subject data from other data chains whose subject is the transaction data of company A is selected. For example, the transaction data of company A in 2023 from other data chains can be selected and added to the first verification subject data. The same selection quantity and method as the first baseline subject data are used for selection. For example, if the first baseline subject data randomly selects 14 items from the transaction data of target subject company A, then the first verification subject data also randomly selects 14 items from the transaction data of target subject company A and adds them to the first verification subject data. Multiple data items are selected and added to the first verification subject data according to the number of target subject data selected from multiple first baseline subject data sets and the target subject.
[0079] Based on a preset number of baseline main data entries, a first baseline hash tree is constructed using a Merkle tree. Merkle trees can be used to verify any data stored, processed, and transmitted within and between computers. They ensure that data transmission speed is not affected and that there is no corruption or alteration in peer-to-peer networks. A Merkle tree with a binary tree structure is constructed using the standard SHA-256 hash function, where the leaf nodes are the hash values of individual data entries.
[0080] Based on the first verification subject data with preset verification data, a first verification hash tree is constructed using a Merkle tree. The first verification hash tree is constructed using the same structure as the first baseline hash tree.
[0081] Comparing the hash values of the first benchmark hash tree and the first verification hash tree, when the root node hash values are the same, the verification coefficient of the first single-chain verification data is 1. When the root node hash values are different, the verification coefficient of the first single-chain verification data = tree similarity = number of leaf nodes with the same hash value ÷ total number of leaf nodes. For example, if both the first benchmark hash tree and the first verification hash tree have a total of 256 leaf nodes, and the number of leaf nodes with the same hash value is 183, then the verification coefficient of the first single-chain verification data = 183 ÷ 256 = 0.71. The larger the value of the first single-chain verification data, the higher the similarity between the first benchmark main data and the first verification main data, the higher the verification pass rate, and the lower the possibility of data tampering.
[0082] Continue performing sparse Merkle tree verification on multiple other verification data chains to obtain multiple single-chain data verification coefficients. Calculate the arithmetic mean of these single-chain data verification coefficients and use it as the data verification coefficient. Output this data verification coefficient as the traceability management result. Sparse Merkle tree verification does not use all data for verification, but instead selects data with high representativeness and high verification requirements using the methods provided above. This makes the data traceability management method more lightweight and efficient while maintaining high verification accuracy.
[0083] By fusing data selection rate and anomaly coefficient to generate a differentiated verification rate, non-uniform sampling of data within the verification chain is guided. High-risk or high-frequency data obtains a higher verification probability. Combined with Merkle tree architecture to generate a hash tree, the computational load required for a single verification is significantly reduced while maintaining verification reliability. This improves the efficiency of traceability output while ensuring the reliability of the output traceability results.
[0084] Example 2, as Figure 2 As shown, based on the same inventive concept as the lightweight integrated data chain traceability management method provided in Embodiment 1, this embodiment of the invention also provides a lightweight integrated data chain traceability management system, including:
[0085] The data acquisition module 100 is used to acquire the subject data of multiple target subjects and acquire the data implicit anomaly coefficients of the multiple target subjects;
[0086] The random selection module 200 is used to configure the data integration rate of multiple target subjects based on multiple data hidden anomaly coefficients, and to randomly select data from multiple subjects according to multiple data integration rates to construct an integrated data chain sequence, wherein random selection is based on traversal constraints;
[0087] The verification quantity acquisition module 300 is used to configure the traceability verification quantity based on the multiple data hidden anomaly coefficients and randomly select the verification data chain to obtain the traceability verification quantity.
[0088] The traceability management module 400 is used to calculate multiple data selection rates of multiple target subjects based on the data chain sequence, combine the multiple data hidden anomaly coefficients, randomly select subject data in each verification data chain, perform sparse Merkle tree verification, and obtain data verification coefficients as traceability management results.
[0089] In one embodiment, the data acquisition module 100 is further configured to:
[0090] Acquire subject data for multiple target entities to be processed;
[0091] Based on the percentage of times that the target subjects' data anomalies occurred within a historical period, the implicit anomaly coefficients of the data for the multiple target subjects are calculated.
[0092] In one embodiment, the random selection module 200 is further configured to:
[0093] Based on multiple data hidden anomaly coefficients, the data integration rate of multiple target subjects is calculated, where each data integration rate includes the probability that the subject data of each target subject is selected;
[0094] Based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data, and a first data chain is constructed based on blockchain, wherein the preset number is less than the number of subjects of multiple target subjects;
[0095] Specifically, based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data sets, and a first data chain is constructed using blockchain, including:
[0096] Based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data.
[0097] Based on blockchain, a predetermined number of subject data will be constructed into data blocks to obtain the first data chain;
[0098] Multiple data chains are constructed to obtain an integrated data chain sequence, wherein the selection and construction are carried out according to traversal constraints, the traversal constraints including that each subject data appears at least once in the integrated data chain sequence.
[0099] In one embodiment, the verification quantity acquisition module 300 is further configured to:
[0100] Calculate the mean of the multiple data hidden anomaly coefficients to obtain the average data hidden anomaly coefficient;
[0101] Obtain the historical average data of all other subjects within a historical period, including the implied anomaly coefficient.
[0102] Get the preset number of traceability verifications;
[0103] Based on the ratio of the average data hidden anomaly coefficient to the historical average data hidden anomaly coefficient, the preset traceability verification quantity is adaptively adjusted to obtain the traceability verification quantity.
[0104] Within the integrated data chain sequence, a verification data chain is randomly selected to obtain the number of traceability verifications.
[0105] In one embodiment, the traceability management module 400 is further configured to:
[0106] Based on the data chain sequence, the occurrence ratio of the subject data of each target subject is calculated as multiple data selection rates;
[0107] Multiple data validation rates are calculated based on multiple data hidden anomaly coefficients and multiple data selection rates;
[0108] Within the first verification data chain among multiple verification data chains, according to the multiple data verification rates, a preset number of subject data is randomly selected, and subject data corresponding to the target subject is randomly selected in other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient.
[0109] Within the first verification data chain of multiple verification data chains, a preset number of subject data are randomly selected according to the multiple data verification rates, and subject data corresponding to the target subject are randomly selected from other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient, including:
[0110] Within the first verification data chain among multiple verification data chains, a preset number of main data is randomly selected according to the multiple data verification rates, serving as the first benchmark main data for the preset number of verifications.
[0111] Within other data chains, randomly select the subject data corresponding to the target subject of the selected subject data to obtain the first verification subject data of the preset verification quantity;
[0112] Based on the first baseline main data with a preset number of verifications, a first baseline hash tree is constructed using a Merkle tree.
[0113] Based on the first verification subject data with preset verification data, a first verification hash tree is constructed using a Merkle tree;
[0114] Calculate the similarity between the first baseline hash tree and the first verification hash tree to obtain the first single-chain data verification coefficient;
[0115] Continue to perform sparse Merkle tree verification on multiple other verification data chains to obtain multiple single-chain data verification coefficients. Calculate the average to obtain the data verification coefficient, which serves as the result for traceability management.
[0116] In summary, the embodiments of this application have at least the following technical effects:
[0117] This application proposes a lightweight integrated data chain traceability management method and system. By dynamically sensing data anomaly risks and adaptively adjusting data integration and verification strategies, it significantly improves the computing power utilization efficiency and verification accuracy in large-scale data traceability scenarios. Specifically, this application optimizes the sampling ratio of data integration based on a dynamic probability model of data anomaly characteristics, avoiding the repeated on-chaining of low-risk data and reducing the amount of inefficient data on the blockchain. Simultaneously, through an anomaly risk-driven verification resource allocation mechanism, it increases the traceability coverage density of high-risk subject data, effectively suppressing the risk of missing key anomaly data, while reducing unnecessary verification rounds for low-risk data, thus reducing redundant computation overall. Furthermore, by combining a Merkle tree architecture, it minimizes the data processing volume required for a single verification while maintaining the integrity of the Merkle tree, accelerating the verification response speed. Compared to traditional methods, the technical solution provided in this application can significantly reduce the computing power waste caused by full-node verification while improving verification accuracy, achieving lightweight operation while ensuring traceability reliability, and achieving the technical effect of dynamically balancing verification accuracy and verification efficiency with limited computing resources.
[0118] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0119] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0120] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A lightweight, integrated data chain traceability management method, characterized in that, The method includes: Obtain subject data of multiple target subjects, and obtain the data implicit anomaly coefficients of the multiple target subjects; Based on multiple hidden anomaly coefficients, configure the data integration rate of multiple target subjects, and randomly select data from multiple subjects according to the multiple data integration rates to construct an integrated data chain sequence, wherein random selection is based on traversal constraints; Based on the multiple data hidden anomaly coefficients, configure the number of traceability verifications, and randomly select the verification data chain to obtain the number of traceability verifications. Based on the data chain sequence, multiple data selection rates for multiple target subjects are calculated. Combined with the multiple data hidden anomaly coefficients, subject data within each verification data chain is randomly selected for sparse Merkle tree verification to obtain data verification coefficients, which serve as the traceability management results, including: Based on the data chain sequence, the occurrence ratio of the subject data of each target subject is calculated as multiple data selection rates; Multiple data validation rates are calculated based on multiple data hidden anomaly coefficients and multiple data selection rates; Within the first verification data chain among multiple verification data chains, according to the multiple data verification rates, a preset number of subject data is randomly selected, and subject data corresponding to the target subject is randomly selected in other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient. Continue to perform sparse Merkle tree verification on multiple other verification data chains to obtain multiple single-chain data verification coefficients. Calculate the average to obtain the data verification coefficient, which serves as the result for traceability management.
2. The lightweight integrated data chain traceability management method according to claim 1, characterized in that, Acquire subject data for multiple target subjects, and obtain the data latent anomaly coefficients of the multiple target subjects, including: Acquire subject data for multiple target entities to be processed; Based on the percentage of times that the target subjects' data anomalies occurred within a historical period, the implicit anomaly coefficients of the data for the multiple target subjects are calculated.
3. The lightweight integrated data chain traceability management method according to claim 1, characterized in that, Based on multiple implicit anomaly coefficients, configure data integration rates for multiple target subjects. Then, randomly select data from multiple subjects according to these integration rates to construct an integrated data chain sequence, including: Based on multiple data hidden anomaly coefficients, the data integration rate of multiple target subjects is calculated, where each data integration rate includes the probability that the subject data of each target subject is selected; Based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data, and a first data chain is constructed based on blockchain, wherein the preset number is less than the number of subjects of multiple target subjects; Multiple data chains are constructed to obtain an integrated data chain sequence, wherein the selection and construction are carried out according to traversal constraints, the traversal constraints including that each subject data appears at least once in the integrated data chain sequence.
4. The lightweight integrated data chain traceability management method according to claim 3, characterized in that, Based on multiple data integration rates, a predetermined number of subject data are randomly selected from multiple subject data sets. A first data chain is constructed using blockchain technology, including: Based on multiple data integration rates, a preset number of subject data are randomly selected from multiple subject data. Based on blockchain, a predetermined number of subject data will be selected and constructed into data blocks to obtain the first data chain.
5. The lightweight integrated data chain traceability management method according to claim 1, characterized in that, Based on the multiple implicit anomaly coefficients of the data, configure the number of traceability verifications, and randomly select verification data chains to obtain the number of traceability verifications, including: Calculate the mean of the multiple data hidden anomaly coefficients to obtain the average data hidden anomaly coefficient; Obtain the historical average data of all other subjects within a historical period, including the implied anomaly coefficient. Get the preset number of traceability verifications; Based on the ratio of the average data hidden anomaly coefficient to the historical average data hidden anomaly coefficient, the preset traceability verification quantity is adaptively adjusted to obtain the traceability verification quantity. Within the integrated data chain sequence, a verification data chain is randomly selected to obtain the number of traceability verifications.
6. The lightweight integrated data chain traceability management method according to claim 1, characterized in that, Within the first verification data chain among multiple verification data chains, a preset number of subject data are randomly selected according to the multiple data verification rates, and subject data corresponding to the target subject are randomly selected from other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient, including: Within the first verification data chain among multiple verification data chains, a preset number of main data is randomly selected according to the multiple data verification rates, serving as the first benchmark main data for the preset number of verifications. Within other data chains, randomly select the subject data corresponding to the target subject of the selected subject data to obtain the first verification subject data of the preset verification quantity; Based on the first baseline main data with a preset number of verifications, a first baseline hash tree is constructed using a Merkle tree. Based on the first verification subject data with preset verification data, a first verification hash tree is constructed using a Merkle tree; Calculate the similarity between the first baseline hash tree and the first verification hash tree to obtain the first single-chain data verification coefficient.
7. A lightweight, integrated data chain traceability management system, characterized in that, The system for implementing the lightweight integrated data chain traceability management method according to any one of claims 1-6, the system comprising: The data acquisition module is used to acquire the subject data of multiple target subjects and to acquire the data hidden anomaly coefficients of the multiple target subjects; The random selection module is used to configure the data integration rate of multiple target subjects based on multiple data hidden anomaly coefficients, and randomly select data from multiple subjects according to multiple data integration rates to construct an integrated data chain sequence, wherein random selection is based on traversal constraints; The verification quantity acquisition module is used to configure the traceability verification quantity based on the multiple data hidden anomaly coefficients, and randomly select the verification data chain to obtain the traceability verification quantity. The traceability management module is used to calculate multiple data selection rates for multiple target subjects based on the data chain sequence, combine the multiple data hidden anomaly coefficients, randomly select subject data within each verification data chain, perform sparse Merkle tree verification, and obtain data verification coefficients as traceability management results, including: Based on the data chain sequence, the occurrence ratio of the subject data of each target subject is calculated as multiple data selection rates; Multiple data validation rates are calculated based on multiple data hidden anomaly coefficients and multiple data selection rates; Within the first verification data chain among multiple verification data chains, according to the multiple data verification rates, a preset number of subject data is randomly selected, and subject data corresponding to the target subject is randomly selected in other data chains to obtain a preset number of subject data combinations. Sparse Merkle tree verification is then performed to obtain the first single-chain data verification coefficient. Continue to perform sparse Merkle tree verification on multiple other verification data chains to obtain multiple single-chain data verification coefficients. Calculate the average to obtain the data verification coefficient, which serves as the result for traceability management.
Citation Information
Patent Citations
Additive self-service inquiring and tracing system
CN119624477A
Workpiece tracing error-proofing system
CN119863155A