Traceability verification method and device for trusted data space
By embedding dynamic watermark identification in the data operation flow and capturing multi-dimensional traceable data in real time, combining cross-domain feature fusion and credibility evaluation matrix, the problem of difficulty in monitoring and traceability of data operation is solved, and efficient and secure management of data and recovery of trusted data space is achieved.
Patent Information
- Application Number
- CN202510349447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-24
AI Technical Summary
In the prior art, data operation process is difficult to monitor throughout the process, data is easily tampered with, and timely discovery, and incomplete traceability information leads to difficulty in traceability.
Embed dynamic watermarks in the data operation flow to capture multi-dimensional traceability data in real time, and through cross-domain feature fusion, data evolution topology map and trustworthiness evaluation matrix, data tampering events are located, traceability chain break risk level is evaluated, and verification instruction sets are generated to restore the trustworthy data space of the data operation flow.
It realizes monitoring and management of the entire process of data operation, and improves data security, integrity and traceability.
Smart Images

Figure CN120217451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security, and particularly to a traceability verification method and device for a trusted data space. Background Art
[0002] The data circulation mode faces problems such as difficult data security guarantee, lack of trust mechanism, and data island phenomenon. Based on this, the trusted data space meets the urgent needs of digital development for secure, trusted, and efficient data circulation. However, in the field of data traceability verification, conventional data traceability methods often have problems such as incomplete traceability information and single credibility evaluation. For example, data traceability is usually centralized storage, which has disadvantages such as poor data security, and there is currently no unified standard for data traceability information, which increases the difficulty of information integration and analysis between data platforms. In addition, most existing traceability technologies are implemented based on a centralized architecture, and there is information asymmetry between the data owners and service providers during the traceability process, resulting in trust problems.
[0003] The prior art has technical problems such as difficult full-process monitoring of the data operation process, easy data tampering and difficult timely discovery, and incomplete traceability information leading to difficult traceability. Summary of the Invention
[0004] This application provides a traceability verification method and device for a trusted data space, which are used to solve the technical problems in the prior art that the data operation process is difficult to monitor throughout the process, the data is easily tampered with and difficult to discover in a timely manner, and the incomplete traceability information leads to difficult traceability.
[0005] In view of the above problems, this application provides a traceability verification method and device for a trusted data space.
[0006] In the first aspect of this application, a traceability verification method for a trusted data space is provided, and the method includes:
[0007] Embedding a dynamic watermark identifier in the data operation flow to capture multi-dimensional traceability data including data version fingerprints, access trajectory maps, and permission change sequences in real time; performing cross-domain feature fusion on the multi-dimensional traceability data to construct a data evolution topology graph, and generating a credibility evaluation matrix by combining a data integrity hash tree and operation behavior entropy value analysis; positioning data tampering events according to the data evolution topology graph and the credibility evaluation matrix, and evaluating the risk level of traceability chain breakage; inputting the feature vector and risk level of the data tampering event into a verification mapping array to obtain a verification instruction set including a data snapshot repair path, a permission traceability backtracking scheme, and a watermark reconstruction strategy; monitoring the execution result according to the verification instruction set, generating a credibility calibration report, and reconstructing the trusted data space of the data operation flow.
[0008] In a second aspect of the present application, a traceability verification device for a trusted data space is provided. The device includes:
[0009] A multi-dimensional traceability data acquisition module for embedding a dynamic watermark identifier in a data operation flow and capturing in real time multi-dimensional traceability data including a data version fingerprint, an access trajectory map, and a permission change sequence; a credibility evaluation matrix generation module for performing cross-domain feature fusion on the multi-dimensional traceability data, constructing a data evolution topology graph, and generating a credibility evaluation matrix by combining a data integrity hash tree and an operation behavior entropy value analysis; a risk level evaluation module for locating a data tampering event and evaluating the risk level of a broken traceability chain according to the data evolution topology graph and the credibility evaluation matrix; a verification instruction set acquisition module for inputting the feature vector and risk level of the data tampering event into a test mapping array to obtain a verification instruction set including a data snapshot repair path, a permission traceability backtracking scheme, and a watermark reconstruction strategy; and a trusted data space reconstruction module for monitoring the execution result according to the verification instruction set, generating a credibility calibration report, and reconstructing the trusted data space of the data operation flow.
[0010] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0011] Embedding a dynamic watermark identifier in a data operation flow and capturing multi-dimensional traceability data in real time; performing cross-domain feature fusion on the multi-dimensional traceability data, constructing a data evolution topology graph, and generating a credibility evaluation matrix; locating a data tampering event and evaluating the risk level of a broken traceability chain according to the data evolution topology graph and the credibility evaluation matrix; inputting the feature vector and risk level of the data tampering event into a test mapping array to obtain a verification instruction set; and monitoring the execution result according to the verification instruction set, generating a credibility calibration report, and reconstructing the trusted data space of the data operation flow. The technical effect of monitoring and managing the entire process of data operation is achieved, and the data security, integrity, and traceability are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0013] Figure 1 It is a schematic flow chart of a traceability verification method for a trusted data space provided by an embodiment of the present application;
[0014] Figure 2Schematic structural diagram of a traceability verification device for a trusted data space provided by an embodiment of the present application.
[0015] Explanation of reference numerals: Multi-dimensional traceability data acquisition module 10, credibility evaluation matrix generation module 20, risk level evaluation module 30, verification instruction set acquisition module 40, trusted data space reconstruction module 50. Detailed implementation manners
[0016] The present application provides a traceability verification method and device for a trusted data space, aiming to solve the technical problems in the prior art that it is difficult to monitor the entire process of data operation, data is easily tampered with and difficult to detect in a timely manner, and the traceability information is incomplete, resulting in difficulties in tracing.
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0018] Embodiment 1, as Figure 1 shown, the present application provides a traceability verification method for a trusted data space, and the method includes:
[0019] Step S100: Embed a dynamic watermark identifier in the data operation flow, and capture multi-dimensional traceability data including data version fingerprints, access track maps, and permission change sequences in real time.
[0020] Specifically, at the metadata layer of the data operation flow, with the help of a reversible encryption algorithm, a time-sensitive key and an operator identity identifier are encoded as a dynamic watermark and injected into it. By setting an event-driven watermark dynamic update logic, using the event hooks of data version fingerprint changes and permission changes to trigger watermark reconstruction and nested encryption operations respectively to maintain the effectiveness and security of the watermark. Adopt event listening and data collection technologies to capture multi-dimensional traceability data in real time. For data version fingerprints, use a hash algorithm to calculate the data content to generate fingerprint information that uniquely identifies the data version; record and analyze the calls to the system access interface to construct an access track map, and detail information such as the access source, access time, and access method of the data; with the help of the log recording function of the permission management system, collect and sort out the permission change sequence, including operation records such as permission granting, revocation, and change. These multi-dimensional traceability data provide comprehensive and accurate basic information for subsequent data processing and security verification.
[0021] Step S200: Perform cross-domain feature fusion on the multi-dimensional traceability data, construct a data evolution topology map, and combine the data integrity hash tree with the operation behavior entropy value analysis to generate a credibility assessment matrix.
[0022] Specifically, for the captured multi-dimensional traceability data, cross-domain feature fusion is first carried out. By configuring the data operation trajectory, it is converted into a multi-dimensional vector containing the operation type weight and context relevance on the evolution coordinate axis, and then the cross-domain operation nodes are aggregated with the help of graph neural network to realize the cross-domain feature fusion of multi-dimensional traceability data, so as to construct a data evolution topology diagram, and graphically present the evolution relationship of data under different operations and time nodes. At the same time, the data integrity hash tree is used to verify the data integrity, and the change of hash value is used to determine whether the data has been tampered with. By analyzing the entropy value of the operation behavior, the uncertainty and complexity of the operation behavior are measured to determine the baseline of the normal behavior pattern. Finally, the normal behavior pattern baseline is used as a constraint condition, combined with the analysis results of the data integrity hash tree and the operation behavior entropy value, to generate a credibility assessment matrix, which can quantify the credibility of the evaluation data in each operation link and state, and provide an important basis for the subsequent positioning of data tampering events and the assessment of the risk level of the traceability chain break.
[0023] Step S300: Locate the data tampering event and evaluate the risk level of the traceability chain break based on the data evolution topology map and the credibility assessment matrix.
[0024] Specifically, with the help of the data operation relationship presented in the data evolution topology diagram and the quantitative credibility information provided by the credibility assessment matrix, we begin to locate data tampering events and evaluate the risk level of traceability chain rupture. We first connect to the tampering event feature template library containing a variety of typical abnormal modes such as permission leapfrogging and watermark rupture modes, and compare the real-time operation mode in the data evolution topology diagram with it. At the same time, based on the abnormal fluctuation of credibility in the credibility assessment matrix, we use the KL divergence between the calculation and the standard template to detect potential abnormal operations. When the KL divergence exceeds the preset divergence threshold, it can be determined that a data tampering event has occurred, and this preset divergence threshold will be adjusted in a hierarchical and flexible manner according to the anti-attack effectiveness matrix to adapt to different security environments. After locating the data tampering incident, comprehensively consider factors such as the position of the tampering incident in the data evolution topology diagram, the amount of data involved, the degree of impact of the tampering on the data credibility, and combine the extent of the credibility decline in the credibility assessment matrix and the credibility changes of related nodes in the traceability chain to evaluate the risk level of traceability chain breakage, determine the possibility of the traceability chain being broken or unreliable due to data tampering, and provide key basis for subsequent targeted measures.
[0025] Step S400: Input the feature vector of the data tampering event and the risk level into the inspection mapping array to obtain a verification instruction set including the data snapshot repair path, the permission traceability backtracking scheme, and the watermark reconstruction strategy.
[0026] Specifically, a correlation mapping algorithm based on machine learning is used to process the data tampering event. First, one-hot encoding is performed on the feature vector of the data tampering event to convert each feature dimension into a binary vector to highlight feature differences; for the risk level, a standardization method is used to map it to the interval [0, 1] to unify the dimension. Then, the processed feature vector and risk level data are input into a pre-trained convolutional neural network (CNN) model. The CNN model extracts local features of the data through the convolutional layer, uses the pooling layer to reduce the data dimension and retain key information, and after multiple layers of processing, the data is input into the fully connected layer. In the fully connected layer, the model performs complex linear and non-linear transformations according to the weights and biases obtained from training, and finally outputs a multi-dimensional vector. This multi-dimensional vector is decoded and parsed into a data snapshot repair path, a permission traceability backtracking scheme, and a watermark reconstruction strategy according to specific rules, thereby forming a verification instruction set.
[0027] Step S500: Monitor the execution result according to the verification instruction set, generate a credibility calibration report, and reconstruct the trusted data space of the data operation flow.
[0028] Specifically, the monitoring and execution work is carried out according to the verification instruction set. For the data snapshot repair path in the instruction set, accurately trace back to the reliable state before the data was tampered with according to the steps and version information indicated by the path to ensure the accuracy and integrity of the data content; for the permission traceability backtracking scheme, trace back to the source of the permission change along the track of the permission change in detail, correct the abnormal permission situation, and restore the normal permission management order; according to the watermark reconstruction strategy, regenerate and embed the watermark to ensure the identification and tracking function of the watermark for data operations. During the entire execution process, the execution status, data changes, and repair effects of each operation are monitored in real time. After the execution is completed, a credibility calibration report is generated according to the monitoring records and verification results. The report details the credibility changes before and after data repair, the execution of each operation link, and the impact assessment on the security of the data space. Finally, based on these operation and report results, the trusted data space of the data operation flow is reconstructed to restore the data to a safe, reliable, and traceable state, providing a solid guarantee for subsequent data operations and usage.
[0029] In a possible implementation manner, step S100 further includes:
[0030] Step S110: Inject a reversible encrypted watermark into the metadata layer of the data operation flow, and the reversible encrypted watermark includes a time-sensitive key and an operator identity identifier.
[0031] Step S120: Set a watermark dynamic update rule based on the reversible encrypted watermark. The watermark dynamic update rule is used to trigger watermark reconstruction when the data version fingerprint changes, or to trigger nested encryption of the watermark when the permission changes.
[0032] Specifically, to achieve data traceability and security guarantee, a reversible encrypted watermark is injected into the metadata layer of the data operation flow. This watermark contains two important elements, namely, a time-sensitive key and an operator identity identifier. The time-sensitive key makes the watermark time-limited, capable of reflecting the state of the data according to different time stages, enhancing the security of the watermark and the traceability of the data; the operator identity identifier clarifies the subject that operates on the data, which helps to quickly locate the responsible party in case of problems in the future. At the beginning of the data operation, this reversible encrypted watermark is automatically embedded into the metadata layer of the data operation flow without affecting the normal use and transfer of the data. It can not only restore the time and operator information contained therein through decryption operations when needed later, but also resist external illegal tampering and attacks to a certain extent, providing a solid foundation for subsequent data traceability and security verification.
[0033] After the reversible encrypted watermark injection is completed, focus on setting the watermark dynamic update rule to ensure that the watermark can continuously and accurately reflect the latest state of the data. This rule mainly focuses on two key triggering conditions. When the data version fingerprint changes, it means that the data content has been modified. At this time, the watermark reconstruction mechanism will be automatically triggered. This mechanism will recalculate the information containing the time-sensitive key and the operator identity identifier, and combine the new data version characteristics to generate a new watermark and replace the original watermark, so as to ensure that the watermark exactly matches the current version of the data. When the permission of the data changes, in order to further strengthen the security of the data and the anti-counterfeiting ability of the watermark, nested encryption of the watermark will be triggered. On the basis of the original reversible encrypted watermark, an additional encryption algorithm and key are used to encrypt it again to form a nested structure. Even if an external attacker tries to crack the watermark, it will be difficult to succeed due to the protection of multiple-layer encryption. At the same time, it can clearly record this important event of permission change, providing strong support for data security management and traceability.
[0034] In a possible implementation manner, step S200 further includes:
[0035] Step S210: Configure the data operation track through the multi-dimensional traceability data.
[0036] Step S220: Map the data operation track to a multi-dimensional vector on the evolution coordinate axis, and each vector contains an operation type weight and a context correlation degree.
[0037] Step S230: Based on the multi-dimensional vectors on the evolution coordinate axis, perform cross-domain feature fusion on the multi-dimensional traceability data.
[0038] Specifically, extract data operation-related fields from the multi-dimensional traceability data, such as operation timestamps, operation subject identifiers, operation content summaries, and data flow path information, etc., and construct a data operation trajectory in chronological order to form an ordered sequence of operation records.
[0039] To further perform quantitative analysis on the data operation trajectory, it needs to be mapped to multi-dimensional vectors on the evolution coordinate axis. First, determine the dimensions of the evolution coordinate axis, which correspond to different attributes of data operations, such as operation time, operation object, operation source, etc. For each data operation trajectory, according to factors such as the importance, influence range, and frequency of the operation, assign corresponding operation type weights to different operation types. At the same time, use data mining and association analysis techniques to calculate the context correlation degree between each operation and the previous and subsequent operations as well as the overall data environment. This correlation degree reflects the logical dependencies and data flow relationships between operations. Finally, use the operation type weights and context correlation degree as different components of the vector to construct multi-dimensional vectors, completing the mapping of the data operation trajectory to multi-dimensional vectors on the evolution coordinate axis, providing a standardized input for subsequent cross-domain feature fusion and data analysis.
[0040] Using graph neural network (GNN) technology, with the multi-dimensional vectors on the evolution coordinate axis as node features, construct a graph structure, where nodes represent data operations and edges represent the association relationships between operations. Through the aggregation function of GNN, such as the aggregation function based on the attention mechanism, aggregate the features of adjacent nodes to achieve cross-domain feature fusion of multi-dimensional traceability data, so that the fused data contains more comprehensive and relevant information.
[0041] In a possible implementation manner, step S230 further includes:
[0042] Step S231: Based on the multi-dimensional vectors on the evolution coordinate axis, use a graph neural network to aggregate cross-domain operation nodes.
[0043] Step S232: Based on the cross-domain operation nodes, configure a weighted directional vector, and extract the propagation path features and abnormal diffusion patterns of the data version fingerprint in the data evolution topology graph.
[0044] Specifically, based on the multi-dimensional vectors on the evolution coordinate axis, graph neural network technology is used to aggregate cross-domain operation nodes. First, each multi-dimensional vector is used as the feature representation of the nodes in the graph neural network, and these nodes represent different cross-domain operations. The graph neural network will update and aggregate the node features by passing information between nodes through the message passing mechanism according to the connection relationships between nodes. In this process, the network will learn the complex associations between nodes, discover the hidden internal connections between cross-domain operations, and make the originally scattered cross-domain operation nodes form an organic whole.
[0045] Based on the aggregated cross-domain operation nodes, the configuration of the weighted directional vector is carried out. Corresponding weights are assigned to each edge according to factors such as the connection strength between nodes and the importance of operations to form a weighted directional vector, so as to accurately describe the interaction and influence direction between cross-domain operations. With the help of these weighted directional vectors, the propagation path characteristics of the data version fingerprint in the data evolution topology graph are further extracted. Analyze how the data version fingerprint flows and changes between different operation nodes, and at the same time identify possible abnormal diffusion patterns. For example, check whether there is a situation where the data version fingerprint spreads against the normal permissions or operation processes, so as to timely discover signs that the data may be tampered with or abnormally operated, providing a key basis for subsequent credibility evaluation and data security maintenance.
[0046] In a possible implementation manner, step S200 further includes:
[0047] Step S240: Obtain a double-chain anchored index for credibility evaluation of the data block hash value and the operation behavior entropy value within the current traceability verification time period.
[0048] Step S250: Determine the operation deposit chain with the double-chain anchored index, and establish a normal behavior pattern baseline through the fluctuation characteristics of the operation behavior entropy value.
[0049] Step S260: Use the normal behavior pattern baseline as a constraint condition to generate the credibility evaluation matrix.
[0050] Specifically, first, for the data blocks within the current traceability verification time period, a specific hashing algorithm (such as SHA-256) is used for processing to convert the content of the data blocks into hash values of a fixed length. These hash values become the unique identifiers of the data blocks. Even if there are minor changes in the data, the hash values will be significantly different, thus ensuring the precise identification of data integrity. At the same time, the operation behaviors within this period are carefully sorted out, various operations are classified and counted, the probability of each operation occurrence is calculated, and then the operation behavior entropy value is calculated according to the information entropy formula. This entropy value reflects the complexity and uncertainty of the operation behaviors. Subsequently, in order to closely associate the data block hash values and the operation behavior entropy values for subsequent credibility evaluation, a double-chain anchored index is constructed. The specific approach is to create a special data structure (such as in the form of nested dictionaries and lists in Python). During the process of traversing all data blocks and operation behavior records, each data block hash value corresponds to the relevant operation behavior entropy value and operation sequence number, and each operation behavior entropy value also inversely corresponds to the corresponding data block hash value and operation sequence number, thus completing the construction of the double-chain anchored index and providing an efficient index path for subsequent data credibility evaluation based on these two key indicators.
[0051] After determining the operation evidence chain with the help of the double-chain anchored index, the Isolation Forest algorithm is used to establish a baseline of normal behavior patterns based on the fluctuation characteristics of the operation behavior entropy value. Using the correspondence between the operation behavior entropy value and the data block hash value in the double-chain anchored index, the operation behavior entropy values are arranged in chronological order to construct the operation evidence chain, which completely records the operation trajectory of the data within the current traceability verification time period. Then, a series of operation behavior entropy values in the operation evidence chain are collected as a data set, and this data set is input into the Isolation Forest algorithm. This algorithm is an unsupervised learning algorithm and is very suitable for anomaly detection. The algorithm randomly selects features and randomly selects split points within the value range of these features, recursively divides the data, and constructs multiple isolation trees. Each data point starts from the root node in the tree and is divided until it becomes a leaf node. The path length that the data point passes through in the tree is an important basis for its anomaly score. For data points of normal behavior, due to their similar characteristic patterns, the paths in the tree are relatively long; while abnormal data points are often isolated more quickly and have shorter paths. The algorithm repeats this process multiple times to construct a forest, and then combines the results of all trees to calculate a final anomaly score for each data point. By learning the normal operation behavior data, the algorithm will identify the normal fluctuation range and pattern of the operation behavior entropy value. Based on the characteristics of these normal patterns, a threshold for the anomaly score is determined. The operation behavior entropy value pattern corresponding to the data points below this threshold is regarded as the normal behavior pattern, and thus the normal behavior pattern baseline is established. Subsequently, when evaluating the data credibility, if the anomaly score corresponding to the new operation behavior entropy value exceeds this threshold, it can be determined that the operation may be abnormal.
[0052] Comprehensively sort out all data operations and statuses within the current traceability verification time period. For each data block hash value and the corresponding operation behavior entropy value, carefully compare them with the baseline of the normal behavior pattern. For data operations that conform to the normal behavior pattern baseline, assign a relatively high credibility score. For example, if the operation behavior entropy value is within the range of plus or minus one standard deviation from the mean of the baseline, and the change pattern of the data block hash value is consistent with that under the normal process, a credibility score close to full marks, such as 90 to 100 points, can be given. Conversely, for data operations that deviate from the normal behavior pattern baseline, reduce the credibility score accordingly based on the degree of deviation. If the operation behavior entropy value exceeds two standard deviations above the baseline mean, or the data block hash value shows abnormal changes, such as being significantly different from the previous normal data block hash value, a credibility score of only 0 to 30 points may be given. Arrange the credibility scores of each data operation and status in a certain matrix structure, where the rows represent different data operation types and the columns represent different data statuses or time nodes, thus generating a credibility evaluation matrix. This matrix can intuitively and quantitatively display the credibility of data in each operation link and status, providing an important quantitative basis for subsequent data security analysis and decision-making.
[0053] In a possible implementation manner, step S300 further includes:
[0054] Step S310: Connect to the tampering event feature template library, which includes privilege escalation diffusion and watermark breakage patterns.
[0055] Step S320: Based on the tampering event feature template library, detect abnormal operation patterns and calculate the KL divergence from the standard template.
[0056] Step S330: When the KL divergence exceeds the preset divergence threshold, locate the data tampering event, and the preset divergence threshold is adjusted elastically in grades according to the anti-attack effectiveness matrix.
[0057] Specifically, first connect to the tampering event feature template library, which pre-stores various typical tampering event feature templates, such as privilege escalation diffusion features and watermark breakage pattern features. The privilege escalation diffusion feature records the behavior pattern of abnormal elevation or diffusion of privileges during data operations; the watermark breakage pattern feature details the abnormal breakage, missing, or damage of the watermark due to data tampering. These feature templates provide important reference bases for subsequent detection of data tampering.
[0058] Based on the connected tampering event feature template library, real-time monitoring of the data operation flow is started to detect whether there is an abnormal operation mode. During the monitoring process, the actually observed operation mode is compared and analyzed with the standard templates in the template library. To more precisely measure the degree of difference between the two, the method of calculating the KL divergence is adopted. The KL divergence is an index used to measure the difference between two probability distributions. Here, it can quantify the deviation degree between the actual operation mode and the normal mode represented by the standard template. Through feature extraction and probability distribution construction of the operation data, for each possible tampering event feature template, the KL divergence value corresponding to the actual operation mode is calculated.
[0059] A divergence threshold is preset in advance. When the calculated KL divergence exceeds this preset divergence threshold, it can be determined that a data tampering event has occurred, and the specific location of the data tampering event in the data operation flow can be quickly located. It should be noted that the preset divergence threshold is not fixed, but is adjusted elastically in grades according to the anti-attack effectiveness matrix. The anti-attack effectiveness matrix comprehensively considers various factors such as the security threat level faced by the system, the importance of the data, and the anti-attack ability of the system itself. For example, in the case of a high security threat or extremely important data, the anti-attack effectiveness matrix will instruct the system to lower the preset divergence threshold to make the detection of data tampering events more sensitive, so as to timely discover potential security risks; while in a relatively stable security environment and relatively low data importance, the preset divergence threshold will be appropriately increased to avoid excessive false alarms due to over-sensitivity, so as to ensure the efficient and accurate location of data tampering events in different security scenarios.
[0060] In a possible implementation manner, step S330 further includes:
[0061] Step S331: Count the types and frequencies of tampering attacks suffered within multiple traceability verification time periods, and set a heat map of the attack mode distribution.
[0062] Step S332: Introduce the success rate index and response delay index of the watermark reconstruction strategy, and calculate the success rate fuzzy parameter and response delay fuzzy parameter according to the heat map of the attack mode distribution.
[0063] Step S333: Configure the anti-attack effectiveness matrix based on the success rate fuzzy parameter and response delay fuzzy parameter.
[0064] Specifically, in-depth data statistics are carried out for multiple traceability verification time periods. First, the data operation records, security audit logs and information captured by relevant monitoring systems in each period are comprehensively scanned. In the process of sorting, each tampering attack event is accurately identified, and its attack type is carefully distinguished, such as clearly judging whether it is an attack involving illegal changes in permissions such as permission diffusion, or an attack that causes watermark breakage and destroys the integrity of watermarks. At the same time, each type of tampering attack is counted and its frequency of occurrence in different traceability verification time periods is recorded. After the statistics are completed, these statistical data are converted into attack mode distribution heat maps using visualization tools. In this heat map, the time period is used as the horizontal axis, and the key links of data operations or the data areas involved are used as the vertical axis. The distribution of various types of tampering attacks is intuitively presented through different degrees of color depth. Areas with high attack frequency will appear darker, indicating that the time period or data area is more likely to be tampered; while areas with low attack frequency have lighter colors, which means they are relatively safe. This enables clear and intuitive insight into the tampering attacks faced by data at different stages and links, providing strong data visualization support for subsequent targeted analysis and response.
[0065] Two key indicators of the watermark reconstruction strategy are introduced, namely the success rate indicator and the response delay indicator. The success rate indicator reflects the proportion of the watermark reconstruction strategy that successfully recovers the watermark after being tampered with, reflecting the effectiveness of the strategy; the response delay indicator measures the time interval from attack detection to watermark reconstruction completion, reflecting the timeliness of the strategy. Combined with the previously generated attack mode distribution heat map, analysis is performed for different attack types and occurrence areas. For areas with high attack frequency and darker colors, the success rate and response delay data of the watermark reconstruction strategy when multiple attacks are carried out in the area are collected. Using fuzzy mathematics methods, the membership function is set according to the distribution characteristics of the data and experience, the success rate and response delay data are converted into fuzzy sets, and the success rate fuzzy parameters and response delay fuzzy parameters are calculated. These fuzzy parameters can more comprehensively and flexibly describe the performance of the watermark reconstruction strategy under different attack modes, and provide a more accurate basis for the configuration of the subsequent anti-attack effectiveness matrix.
[0066] Based on the calculated success rate fuzzy parameter and response delay fuzzy parameter, start to configure the anti-attack effectiveness matrix. Each element of the anti-attack effectiveness matrix corresponds to the evaluation of the anti-attack ability of the watermark reconstruction strategy under different attack modes. The element values in the matrix are obtained through specific operations and combinations of the success rate fuzzy parameter and the response delay fuzzy parameter. For example, a high value of the success rate fuzzy parameter and a low value of the response delay fuzzy parameter are given a higher anti-attack effectiveness score, and vice versa. In this way, the anti-attack effectiveness matrix comprehensively and quantitatively reflects the comprehensive performance of the watermark reconstruction strategy in the face of various tampering attacks, providing an important basis for flexibly adjusting the preset divergence threshold and optimizing the watermark reconstruction strategy according to different security scenarios.
[0067] In a possible implementation manner, step S332 further includes:
[0068] Define the success rate fuzzy parameter where n is the total number of attack types, is the frequency weight of the i-th type of attack in the attack mode distribution heat map, f i represents the cumulative frequency of the i-th type of attack within the traceability verification time period, N success,i is the number of successful watermark reconstructions in the i-th type of attack, N total,i is the total number of triggers of the i-th type of attack.
[0069] Specifically, the calculation formula for the success rate fuzzy parameter is n is the total number of attack types. In the actual data security monitoring scenario, there are various types of attacks, such as unauthorized privilege escalation, malicious tampering resulting in watermark breakage, etc. n is the sum of these attack types. w i is the frequency weight of the i-th type of attack in the attack mode distribution heat map, and its calculation method is where f i represents the cumulative frequency of the i-th type of attack within the traceability verification time period. This means that the more frequently an attack occurs, the greater the weight of this type of attack in calculating the success rate fuzzy parameter. For example, if a certain type of attack appears frequently in multiple time periods, then its impact on the overall success rate fuzzy parameter is more significant. N success,i is the number of successful watermark reconstructions in the i-th type of attack, N total,i is the total number of triggers of the i-th type of attack. By the ratio of N success,i to N total,i the success rate of watermark reconstruction under the i-th type of attack can be obtained. Multiply this success rate by the corresponding frequency weight w i, and accumulate all attack types, then divide by the total frequency weights of all attack types to finally obtain the success rate fuzzy parameter s. This parameter comprehensively considers the occurrence frequencies of different attack types and the success of watermark reconstruction under various attacks, and can more comprehensively and accurately reflect the comprehensive success rate performance of the watermark reconstruction strategy in a variety of attack environments, providing an important basis for subsequent evaluation of the system's anti-attack ability.
[0070] In a possible implementation manner, step S332 further includes:
[0071] Define the response delay fuzzy parameter where T is the total number of statistical time windows within the traceability verification time period, is the average delay of watermark reconstruction within the t-th window, is the maximum delay within the t-th window, and α, β are weight coefficients (α + β = 1), used to balance the influence of the average and peak delays, and λ is the time decay factor.
[0072] Output the anti-attack effectiveness matrix with the success rate fuzzy parameter and the success rate segmentation boundary under the confidence level, and the response delay fuzzy parameter and the reference response delay tolerance value.
[0073] Specifically, first define the response delay fuzzy parameter, and the formula is where T represents the total number of statistical time windows within the traceability verification time period, which is the number of time segments for dividing the entire traceability verification period. is the average delay of watermark reconstruction within the t-th window, reflecting the average level of the watermark reconstruction time within this window; is the maximum delay within the t-th window, reflecting the longest situation of the watermark reconstruction time within this window. α, β are weight coefficients and satisfy (α + β = 1), used to adjust the influence degrees of the average delay and the peak delay in the calculation. For example, when more attention is paid to the average delay, the value of α can be appropriately increased. λ is the time decay factor, used to reflect the change of the influence degree of the delays in different time windows on the overall response delay fuzzy parameter over time. The closer the window is to the current time, the greater the influence of its delay on the overall parameter.
[0074] Based on the success rate fuzzy parameter and the success rate segmentation boundary under the confidence level, as well as the response delay fuzzy parameter and the reference response delay tolerance value, an anti-attack effectiveness matrix is output. The success rate segmentation boundary under the confidence level is a pre-set threshold for dividing different levels of the success rate fuzzy parameter. By comparing the success rate fuzzy parameter with these boundary values, the level of the success rate is judged. The reference response delay tolerance value is also a pre-set standard for measuring whether the response delay fuzzy parameter is within an acceptable range. By comparing these two fuzzy parameters with the corresponding boundary values and tolerance values respectively, comprehensively considering the success rate of watermark reconstruction and the response delay situation, the element values of the anti-attack effectiveness matrix are finally determined, thereby constructing a matrix that can reflect the anti-attack ability under different attack situations, providing strong support for security assessment and strategy adjustment.
[0075] Embodiment 2, based on the same inventive concept as the traceability verification method for a trusted data space in the foregoing embodiment, as Figure 2 shown, the present application provides a traceability verification device for a trusted data space. The device in the embodiments of the present application and the method embodiments are based on the same inventive concept. Among them, the device includes:
[0076] A multi-dimensional traceability data acquisition module 10, configured to embed a dynamic watermark identifier in a data operation stream and capture multi-dimensional traceability data including a data version fingerprint, an access track map, and a permission change sequence in real time.
[0077] A credibility evaluation matrix generation module 20, configured to perform cross-domain feature fusion on the multi-dimensional traceability data, construct a data evolution topology graph, and generate a credibility evaluation matrix in combination with data integrity hash tree and operation behavior entropy value analysis.
[0078] A risk level evaluation module 30, configured to locate a data tampering event according to the data evolution topology graph and the credibility evaluation matrix, and evaluate the risk level of traceability chain breakage.
[0079] A verification instruction set acquisition module 40, configured to input the feature vector and risk level of the data tampering event into a verification mapping array to obtain a verification instruction set including a data snapshot repair path, a permission traceability backtracking scheme, and a watermark reconstruction strategy.
[0080] A trusted data space reconstruction module 50, configured to monitor the execution result according to the verification instruction set, generate a credibility calibration report, and reconstruct the trusted data space of the data operation stream.
[0081] Furthermore, the device is also used to implement the following functions:
[0082] A reversible encrypted watermark is injected into the metadata layer of the data operation flow, and the reversible encrypted watermark includes a time-sensitive key and an operator identity. Based on the reversible encrypted watermark, a watermark dynamic update rule is set, and the watermark dynamic update rule is used to trigger watermark reconstruction when the data version fingerprint changes, or trigger watermark nested encryption when the permission changes.
[0083] Furthermore, the device is also used to achieve the following functions:
[0084] The data operation trajectory is configured through the multi-dimensional traceability data; the data operation trajectory is mapped into a multi-dimensional vector on the evolution coordinate axis, each vector containing an operation type weight and a context association degree; based on the multi-dimensional vector on the evolution coordinate axis, the multi-dimensional traceability data is subjected to cross-domain feature fusion.
[0085] Furthermore, the device is also used to achieve the following functions:
[0086] Based on the multidimensional vectors on the evolution coordinate axis, a graph neural network is used to aggregate cross-domain operation nodes; based on the cross-domain operation nodes, a weight directional vector is configured to extract the propagation path characteristics and abnormal diffusion pattern of the data version fingerprint in the data evolution topology graph.
[0087] Furthermore, the device is also used to achieve the following functions:
[0088] Obtain a double-chain anchor index for credibility assessment of the data block hash value and the operation behavior entropy value within the current traceability verification time period; determine the operation evidence chain with the double-chain anchor index, and establish a normal behavior pattern baseline through the fluctuation characteristics of the operation behavior entropy value; use the normal behavior pattern baseline as a constraint condition to generate the credibility assessment matrix.
[0089] Furthermore, the device is also used to achieve the following functions:
[0090] Connecting to a tampering event feature template library, the tampering event feature template library includes permission leapfrogging and watermark breakage modes; based on the tampering event feature template library, detecting abnormal operation modes, and calculating the KL divergence between the mode and the standard template; locating the data tampering event when the KL divergence exceeds a preset divergence threshold, and the preset divergence threshold is adjusted hierarchically and flexibly with an anti-attack effectiveness matrix.
[0091] Furthermore, the device is also used to achieve the following functions:
[0092] Statistically analyze the types and frequencies of tampering attacks suffered within multiple traceability verification time periods, and set up a heat map of the attack pattern distribution; introduce the success rate index and response delay index of the watermark reconstruction strategy, and calculate the success rate fuzzy parameter and response delay fuzzy parameter based on the heat map of the attack pattern distribution; configure an anti-attack effectiveness matrix based on the success rate fuzzy parameter and response delay fuzzy parameter.
[0093] Furthermore, the device is also used to implement the following functions:
[0094] Define the success rate fuzzy parameter where n is the total number of attack types, is the frequency weight of the i-th type of attack in the heat map of the attack pattern distribution, and f i represents the cumulative frequency of the i-th type of attack within the traceability verification time period, and N success,i is the number of successful watermark reconstructions in the i-th type of attack, and N total,i is the total number of trigger times of the i-th type of attack.
[0095] Furthermore, the device is also used to implement the following functions:
[0096] Define the response delay fuzzy parameter where T is the total number of statistical time windows within the traceability verification time period, is the average delay of watermark reconstruction within the t-th window, is the maximum delay within the t-th window, α and β are weight coefficients (α + β = 1) used to balance the influence of the average and peak delays, and λ is a time decay factor; output the anti-attack effectiveness matrix with the success rate fuzzy parameter and the success rate segmentation boundary under the confidence level, and the response delay fuzzy parameter and the reference response delay tolerance value.
[0097] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above description of specific embodiments of this specification is provided. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0098] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0099] This specification and the drawings are merely illustrative of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications therein.
Claims
1. A provenance verification method for a trusted data space, characterized in that: The method comprises: Embed dynamic watermarks in the data operation flow to capture multi-dimensional traceability data including data version fingerprints, access trajectory maps, and permission change sequences in real time; Perform cross-domain feature fusion on the multi-dimensional traceability data, construct a data evolution topology map, and combine the data integrity hash tree with the operation behavior entropy value analysis to generate a credibility assessment matrix; According to the data evolution topology diagram and credibility assessment matrix, locate the data tampering event and assess the risk level of the traceability chain break; Input the feature vector and risk level of the data tampering event into the verification mapping array to obtain a verification instruction set including a data snapshot repair path, a permission tracing back scheme, and a watermark reconstruction strategy; The execution result is monitored according to the verification instruction set, a credibility calibration report is generated, and the trusted data space of the data operation flow is reconstructed.
2. A traceability verification method for a trusted data space as claimed in claim 1, characterized in that: Embedding a dynamic watermark identifier in a data operation stream, the method comprising: Injecting a reversible encrypted watermark at the metadata layer of the data operation flow, wherein the reversible encrypted watermark includes a time-sensitive key and an operator identity; Based on the reversibly encrypted watermark, a watermark dynamic update rule is set, and the watermark dynamic update rule is used to trigger watermark reconstruction when the data version fingerprint changes, or trigger watermark nested encryption when the authority changes.
3. A traceability verification method for a trusted data space as claimed in claim 1, characterized in that: Perform cross-domain feature fusion on the multi-dimensional traceability data to construct a data evolution topology diagram, including: Through the multi-dimensional traceability data, configure the data operation track; Mapping the data operation trajectory into a multidimensional vector on an evolution coordinate axis, each vector including an operation type weight and a context association degree; Based on the multidimensional vectors on the evolution coordinate axis, cross-domain feature fusion is performed on the multidimensional traceability data.
4. A traceability verification method for a trusted data space as claimed in claim 3, characterized in that: Based on the multidimensional vector on the evolution coordinate axis, cross-domain feature fusion is performed on the multidimensional traceability data, including: Based on the multidimensional vectors on the evolution coordinate axis, a graph neural network is used to aggregate cross-domain operation nodes; Based on the cross-domain operation node, a weight directional vector is configured to extract the propagation path characteristics and abnormal diffusion mode of the data version fingerprint in the data evolution topology diagram.
5. A provenance verification method for a trusted data space as claimed in claim 1, characterized in that: Combining the data integrity hash tree with the entropy analysis of the operation behavior, a credibility assessment matrix is generated, including: Obtain the double-chain anchor index for credibility assessment of data block hash value and operation behavior entropy value within the current traceability verification time period; Determine the operation evidence chain with the double-chain anchor index, and establish a normal behavior pattern baseline through the fluctuation characteristics of the entropy value of the operation behavior; The credibility evaluation matrix is generated by taking the normal behavior pattern baseline as a constraint condition.
6. A traceability verification method for a trusted data space as claimed in claim 1, characterized in that: According to the data evolution topology diagram and the credibility evaluation matrix, the data tampering event is located, including: Connecting to a tampering event feature template library, wherein the tampering event feature template library includes permission leapfrogging and watermark breaking modes; Based on the tampering event feature template library, detect abnormal operation mode and calculate the KL divergence between it and the standard template; When the KL divergence exceeds a preset divergence threshold, the data tampering event is located, and the preset divergence threshold is adjusted hierarchically and flexibly according to an anti-attack performance matrix.
7. A traceability verification method for a trusted data space as claimed in claim 6, characterized in that: The preset divergence threshold is adjusted elastically in a hierarchical manner using an anti-attack effectiveness matrix, including: Count the types and frequencies of tampering attacks suffered during multiple traceability verification time periods, and set up a heat map of attack mode distribution; The success rate index and response delay index of the watermark reconstruction strategy are introduced, and the success rate fuzzy parameter and the response delay fuzzy parameter are calculated according to the attack mode distribution heat map; Based on the success rate fuzzy parameter and the response delay fuzzy parameter, an anti-attack effectiveness matrix is configured.
8. A traceability verification method for a trusted data space as claimed in claim 7, characterized in that: The fuzzy parameter of success rate is calculated, and the method comprises: Defining the fuzzy parameter of success rate Where n is the total number of attack types, is the frequency weight of the i-th attack in the attack mode distribution heat map, f i N represents the cumulative frequency of the i-th type of attack during the traceability verification period. success,i is the number of successful watermark reconstructions in the i-th attack, N total,i is the total number of triggering of the i-th type of attack.
9. A traceability verification method for a trusted data space as claimed in claim 8, characterized in that: Based on the success rate fuzzy parameter and the response delay fuzzy parameter, an anti-attack effectiveness matrix is configured, and the method includes: Define the response delay fuzzy parameter Among them, T is the total number of statistical time windows within the traceability verification time period, is the average delay of watermark reconstruction in the t-th window, is the maximum delay in the tth window, α and β are weight coefficients (α+β=1), which are used to balance the impact of average and peak delays, and λ is the time decay factor; The anti-attack performance matrix is outputted based on the success rate fuzzy parameter and the success rate segmentation boundary under the confidence level, the response delay fuzzy parameter and the benchmark response delay tolerance value.
10. A traceability verification device for a trusted data space, characterized in that: The device is used to implement a traceability verification method for a trusted data space according to any one of claims 1 to 9, and the device includes: The multi-dimensional traceability data acquisition module is used to embed dynamic watermarks in the data operation flow and capture multi-dimensional traceability data including data version fingerprints, access trajectory maps, and permission change sequences in real time; A credibility evaluation matrix generation module is used to perform cross-domain feature fusion on the multi-dimensional traceability data, construct a data evolution topology map, and combine the data integrity hash tree with the operation behavior entropy value analysis to generate a credibility evaluation matrix; A risk level assessment module, used to locate data tampering events and assess the risk level of traceability chain breakage based on the data evolution topology map and the credibility assessment matrix; A verification instruction set acquisition module is used to input the feature vector and risk level of the data tampering event into the inspection mapping array to obtain a verification instruction set including a data snapshot repair path, a permission tracing back scheme and a watermark reconstruction strategy; A trusted data space reconstruction module is used to monitor the execution result according to the verification instruction set, generate a credibility calibration report and reconstruct the trusted data space of the data operation flow.
Citation Information
Patent Citations
Multi-dimensional data authority management and privacy protection method for electric power information network
CN119538276A
Performance-lossless watermark credible traceability method and system based on block chain
CN119622672A
Cross-modal image-watermark joint generation and detection device and method thereof
US12125119B1
Cited By
Database verification method and electronic equipment
CN120687647A
Secure computing system and method based on data hierarchical storage and key distribution
CN121486217A
A secure computing system and method based on data tiered storage and key distribution
CN121486217B
Data management method and system for processing traceability of traditional Chinese medicine decoction pieces
CN122087506A