A method and device for provenance verification of trusted data spaces

By embedding dynamic watermark identifiers and cross-domain feature fusion into the data operation flow, a data evolution topology map and a credibility assessment matrix are constructed, solving the monitoring problem in the data operation process and realizing data security and traceability.

CN120217451BActive Publication Date: 2025-12-12LINGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510349447.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-12-12
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to monitor the entire data operation process, data is easily tampered with and difficult to detect in a timely manner, and incomplete traceability information makes tracing difficult.

Method used

By embedding dynamic watermarks into the data operation flow, multi-dimensional traceability data is captured in real time. A data evolution topology map and a credibility assessment matrix are constructed through cross-domain feature fusion to locate data tampering events and generate a verification instruction set to reconstruct a trustworthy data space.

Benefits of technology

It enables monitoring and management of the entire data operation process, improving data security, integrity, and traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217451B_ABST
    Figure CN120217451B_ABST
Patent Text Reader

Abstract

The application discloses a kind of traceability verification method and device for trusted data space, it is related to data security technical field, the method includes: embedding dynamic watermark identification in data operation flow, real-time multidimensional traceability data;Cross-domain feature fusion is carried out, data evolution topology graph is constructed, and trusted evaluation matrix is generated;According to the data evolution topology graph and trusted evaluation matrix, the risk level of traceability chain fracture is evaluated;The feature vector of the data tampering event is input into the test mapping array with risk level, to obtain verification instruction set;According to the verification instruction set, generate trusted calibration report and reconstruct trusted data space.The application solves the technical problems that the data operation process of prior art is difficult to monitor throughout, data is easy to be tampered with and difficult to be found in time, traceability information is incomplete, leading to difficult to trace, achieves monitoring and management to data operation whole process, improves the technical effects of data security, integrity and traceability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data security, in particular to a traceability verification method and device for a trusted data space. BACKGROUND

[0002] The data flow mode is faced with the problems of difficult guarantee of data security, lack of trust mechanism and data island phenomenon. Therefore, the trusted data space meets the urgent needs of safe, reliable and efficient data flow for digital development. However, in the field of data traceability verification, the conventional data traceability method often has problems such as incomplete traceability information and single reliability evaluation. For example, the data traceability is usually stored in a centralized manner, which has the disadvantage of poor data security. In addition, there is no uniform standard for the traceability information of data at present, which increases the difficulty of information integration and analysis between data platforms. Moreover, the existing traceability technology is mostly implemented based on a centralized architecture, and the information between the data owner and the service provider in the traceability process is not equal, which causes trust problems.

[0003] The existing technology has the technical problems of difficult whole-process monitoring of data operation process, easy tampering of data and difficult timely discovery, and difficult tracing caused by incomplete traceability information. SUMMARY

[0004] The present application provides a traceability verification method and device for a trusted data space, which is used to solve the technical problems of difficult whole-process monitoring of data operation process, easy tampering of data and difficult timely discovery, and difficult tracing caused by incomplete traceability information in the prior art.

[0005] In view of the above problems, the present application provides a traceability verification method and device for a trusted data space.

[0006] In a first aspect, the present application provides a traceability verification method for a trusted data space, which comprises:

[0007] A dynamic watermark identifier is embedded in a data operation flow, and multi-dimensional traceability data including data version fingerprints, access trajectory maps and permission change sequences are captured in real time. Cross-domain feature fusion is performed on the multi-dimensional traceability data to construct a data evolution topology graph, and a trustworthiness evaluation matrix is generated by combining a data integrity hash tree and an operation behavior entropy analysis. According to the data evolution topology graph and the trustworthiness evaluation matrix, a data tampering event is located and a traceability chain breakage risk level is evaluated. The feature vector and risk level of the data tampering event are input into a test mapping array to obtain a verification instruction set including a data snapshot repair path, a permission traceability backtracking scheme and a watermark reconstruction strategy. According to the verification instruction set, the execution result is monitored, a trustworthiness calibration report is generated, and the trusted data space of the data operation flow is reconstructed.

[0008] In a second aspect of the present application, a traceability verification device for a trusted data space is provided, the device comprising:

[0009] A multi-dimensional traceability data acquisition module is configured to embed dynamic watermark identification in a data operation flow and capture multi-dimensional traceability data including data version fingerprints, access trajectory maps and permission change sequences in real time. A trustworthiness evaluation matrix generation module is configured to perform cross-domain feature fusion on the multi-dimensional traceability data, construct a data evolution topology graph, and generate a trustworthiness evaluation matrix in combination with a data integrity hash tree and an operation behavior entropy analysis. A risk level evaluation module is configured to locate a data tampering event and evaluate a traceability chain breakage risk level according to the data evolution topology graph and the trustworthiness evaluation matrix. A verification instruction set acquisition module is configured to input a feature vector and a risk level of the data tampering event into a verification mapping array to obtain a verification instruction set including a data snapshot repair path, a permission traceability rollback scheme and a watermark reconstruction strategy. A trusted data space reconstruction module is configured to monitor execution results according to the verification instruction set, generate a trustworthiness calibration report and reconstruct a trusted data space of the data operation flow.

[0010] The one or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0011] The multi-dimensional traceability data is captured in real time by embedding dynamic watermark identification in a data operation flow. The cross-domain feature fusion is performed on the multi-dimensional traceability data to construct a data evolution topology graph and generate a trustworthiness evaluation matrix. The data tampering event is located and the traceability chain breakage risk level is evaluated according to the data evolution topology graph and the trustworthiness evaluation matrix. The feature vector and the risk level of the data tampering event are input into a verification mapping array to obtain a verification instruction set. The execution results are monitored according to the verification instruction set, the trustworthiness calibration report is generated and the trusted data space of the data operation flow is reconstructed. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] Figure 1 A traceability verification method flowchart for a trusted data space is provided for the embodiments of the present application.

[0014] Figure 2A structure schematic diagram of a traceability verification device for a trusted data space is provided for the embodiments of the present application.

[0015] Label explanation: multidimensional traceability data acquisition module 10, trustworthiness evaluation matrix generation module 20, risk level evaluation module 30, verification instruction set acquisition module 40, trusted data space reconstruction module 50. DETAILED DESCRIPTION

[0016] The present application provides a traceability verification method and device for a trusted data space, which is used to solve the technical problems that the data operation process is difficult to monitor throughout, the data is easy to tamper and difficult to find in time, and the incomplete traceability information leads to difficult tracing in the prior art.

[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the scope of protection of the present application.

[0018] Embodiment one, as shown in the present application, a traceability verification method for a trusted data space is provided, which comprises: Figure 1

[0019] Step S100: embedding dynamic watermark identification in the data operation flow, and capturing multidimensional traceability data including data version fingerprint, access trajectory map and permission change sequence in real time.

[0020] Specifically, in the metadata layer of the data operation flow, a time-sensitive key and an operator identity are encoded into a dynamic watermark by means of a reversible encryption algorithm, and are injected into the metadata layer. By setting event-driven watermark dynamic update logic, the events of data version fingerprint change and permission change are used to trigger watermark reconstruction and nested encryption operation respectively, so as to maintain the effectiveness and security of the watermark. Event listening and data acquisition technology is used to capture multidimensional traceability data in real time. For the data version fingerprint, a hash algorithm is used to calculate the data content to generate fingerprint information uniquely identifying the data version; by recording and analyzing the calling of the system access interface, an access trajectory map is constructed to record the access source, access time, access mode and other information of the data in detail; by means of the log recording function of the permission management system, the permission change sequence is collected and sorted, including the operation records of permission granting, withdrawing and changing. These multidimensional traceability data provide comprehensive and accurate basic information for subsequent data processing and security verification.

[0021] ​Step S200: Cross-domain feature fusion is performed on the multi-dimensional traceability data, a data evolution topology graph is constructed, and a credibility evaluation matrix is generated by combining data integrity hash tree and operation behavior entropy analysis.

[0022] Specifically, for the captured multi-dimensional traceability data, first, cross-domain feature fusion is carried out. By configuring the data operation track, it is converted into a multi-dimensional vector containing operation type weight and context correlation degree on the evolution coordinate axis, and then the cross-domain operation nodes are aggregated by means of graph neural network, realizing the cross-domain feature fusion of multi-dimensional traceability data, and thus constructing a data evolution topology graph to intuitively present the evolution relationship of data under different operations and time nodes in a graphical manner. At the same time, the data integrity is checked by using the data integrity hash tree, and the data integrity is judged according to the change of the hash value. By analyzing the operation behavior entropy, the uncertainty and complexity of the operation behavior are measured to determine the normal behavior mode baseline. Finally, the normal behavior mode baseline is taken as the constraint condition, and the data integrity hash tree and the operation behavior entropy analysis result are combined to generate a credibility evaluation matrix, which can quantitatively evaluate the credibility of data at each operation link and state, providing an important basis for subsequent positioning of data tampering events and evaluation of traceability chain rupture risk level.

[0023] Step S300: According to the data evolution topology graph and the credibility evaluation matrix, the data tampering event is located, and the traceability chain rupture risk level is evaluated.

[0024] Specifically, by means of the data operation relationship presented by the data evolution topology graph and the quantitative credibility information provided by the credibility evaluation matrix, the data tampering event is located and the traceability chain rupture risk level is evaluated. First, connect the tampering event feature template library containing multiple typical abnormal modes such as permission escalation diffusion and watermark breaking mode, compare the real-time operation mode in the data evolution topology graph with the template library, and at the same time, according to the abnormal fluctuation of credibility in the credibility evaluation matrix, use the method of calculating the KL divergence between the standard template to detect potential abnormal operations. When the KL divergence exceeds the preset divergence threshold, it can be determined that a data tampering event has occurred, and this preset divergence threshold will be flexibly adjusted according to the anti-attack efficiency matrix to adapt to different security environments. After locating the data tampering event, considering the position of the tampering event in the data evolution topology graph, the amount of data involved, the influence degree of tampering on data credibility and other factors, combining the magnitude of credibility decrease in the credibility evaluation matrix and the credibility change of related nodes on the traceability chain, the traceability chain rupture risk level is evaluated, and the possibility of traceability chain rupture or unreliability due to data tampering is determined, providing a key basis for subsequent targeted measures.

[0025] Step S400: input the feature vector of the data tampering event and the risk level into a verification mapping array to obtain a verification instruction set containing a data snapshot repair path, an authority trace back scheme and a watermark reconstruction strategy.

[0026] Specifically, the correlation mapping algorithm based on machine learning is used to process the data tampering event. First, the feature vector of the data tampering event is one-hot encoded, and each feature dimension is converted into a binary vector to highlight the feature difference. For the risk level, a standardization method is used to map it to the [0, 1] interval to unify the dimension. Then, the processed feature vector and risk level data are input into a pre-trained convolutional neural network (CNN) model. The CNN model extracts local features of the data through convolutional layers, reduces the data dimension and retains key information through pooling layers, and after multiple layers of processing, the data is input into a fully connected layer. In the fully connected layer, the model performs complex linear and nonlinear transformations according to the weights and biases obtained through training, and finally outputs a multi-dimensional vector. This multi-dimensional vector is decoded according to specific rules to parse the data snapshot repair path, the authority trace back scheme and the watermark reconstruction strategy, and then form a verification instruction set.

[0027] Step S500: monitor the execution results according to the verification instruction set, generate a credibility calibration report and reconstruct the trusted data space of the data operation flow.

[0028] Specifically, according to the verification instruction set, the monitoring execution work is carried out. For the data snapshot repair path in the instruction set, according to the steps and version information indicated by the path, the reliable state before the data is tampered with is accurately traced back to ensure the accuracy and integrity of the data content; for the authority trace back scheme, along the track of the authority change, the source of the authority change is traced in detail, the abnormal authority situation is corrected, and the normal authority management order is restored; according to the watermark reconstruction strategy, the watermark is regenerated and embedded to ensure the identification and tracking function of the watermark on the data operation. During the entire execution process, the execution status of each operation, the data change and the repair effect are monitored in real time. After the execution is completed, a credibility calibration report is generated according to the monitoring records and verification results. The report records the credibility changes before and after the data repair, the execution of each operation link and the influence evaluation on the data space security, etc. Finally, based on these operations and report results, the trusted data space of the data operation flow is reconstructed, so that the data is restored to a safe, reliable and traceable state, providing a solid guarantee for subsequent data operation and use.

[0029] In one possible implementation, step S100 further includes:

[0030] Step S110: inject a reversible encryption watermark into the metadata layer of the data operation flow, the reversible encryption watermark containing a time-sensitive key and an operator identity.

[0031] Step S120: setting a watermark dynamic update rule based on the reversible encryption watermark, the watermark dynamic update rule being used to trigger watermark reconstruction when a data version fingerprint changes or trigger watermark nested encryption when a permission changes.

[0032] Specifically, in order to realize the traceability and security guarantee of data, a reversible encryption watermark is injected into the metadata layer of the data operation flow. This watermark contains two important elements, namely time-sensitive key and operator identity. The time-sensitive key makes the watermark time-limited, which can reflect the state of the data according to different time stages, enhancing the security of the watermark and the traceability of the data; the operator identity explicitly identifies the subject who operates the data, which helps to quickly locate the responsible party when problems occur later. At the beginning of data operation, this reversible encryption watermark is automatically embedded into the metadata layer of the data operation flow without affecting the normal use and circulation of the data. It can restore the time and operator information contained therein through decryption operation when needed later, and can resist external illegal tampering and attacks to some extent, providing a solid foundation for subsequent data tracing and security verification.

[0033] After completing the reversible encryption watermark injection, the watermark dynamic update rule is set to ensure that the watermark can accurately reflect the latest state of the data. This rule mainly revolves around two key trigger conditions. When the data version fingerprint changes, it means that the data content has been modified, at which time the watermark reconstruction mechanism is automatically triggered. This mechanism will recalculate the information containing the time-sensitive key and the operator identity, combine the new data version characteristics, generate a new watermark and replace the original watermark, so as to ensure that the watermark is accurately matched with the current version of the data. When the permission of the data changes, in order to further strengthen the security of the data and the anti-fake ability of the watermark, the watermark nested encryption is triggered. On the basis of the original reversible encryption watermark, it is encrypted again using additional encryption algorithm and key, forming a nested structure. Even if external attackers try to crack the watermark, it will be difficult to succeed due to the protection of multi-layer encryption, and at the same time, the important event of permission change can be clearly recorded, providing strong support for the security management and tracing of data.

[0034] In one possible implementation manner, step S200 further includes:

[0035] Step S210: configuring a data operation track through the multi-dimensional traceability data.

[0036] Step S220: mapping the data operation track into a multi-dimensional vector on an evolution coordinate axis, each vector containing an operation type weight and a context correlation degree.

[0037] Step S230: Based on the multi-dimensional vector on the evolution coordinate axis, cross-domain feature fusion is performed on the multi-dimensional provenance data.

[0038] Specifically, data operation related fields such as operation timestamp, operation subject identifier, operation content summary, and data flow path information are extracted from multi-dimensional provenance data, and data operation trajectories are constructed in chronological order to form an ordered operation record sequence.

[0039] To further quantitatively analyze the data operation trajectory, it needs to be mapped to a multi-dimensional vector on the evolution coordinate axis. First, the dimensions of the evolution coordinate axis are determined, which correspond to different attributes of data operations such as operation time, operation object, and operation source. For each data operation trajectory, different operation types are assigned corresponding operation type weights according to factors such as operation importance, influence range, and frequency. At the same time, using data mining and correlation analysis techniques, the contextual correlation degree of each operation with the previous and subsequent operations and the overall data environment is calculated, which reflects the logical dependence and data flow relationship between operations. Finally, the operation type weight and the contextual correlation degree are used as different components of the vector to construct a multi-dimensional vector, completing the mapping of the data operation trajectory to the multi-dimensional vector on the evolution coordinate axis, and providing standardized input for subsequent cross-domain feature fusion and data analysis.

[0040] Using graph neural network (GNN) technology, the multi-dimensional vector on the evolution coordinate axis is used as the node feature to construct a graph structure, where the node represents the data operation and the edge represents the association relationship between the operations. Through the aggregation function of GNN, such as the aggregation function based on attention mechanism, the features of adjacent nodes are aggregated to realize cross-domain feature fusion of multi-dimensional provenance data, so that the fused data contains more comprehensive and more correlated information.

[0041] In one possible implementation, step S230 further includes:

[0042] Step S231: Based on the multi-dimensional vector on the evolution coordinate axis, graph neural network aggregation is performed on the cross-domain operation nodes.

[0043] Step S232: Based on the cross-domain operation nodes, a weight directional vector is configured to extract the propagation path features and abnormal diffusion patterns of the data version fingerprint in the data evolution topology graph.

[0044] Specifically, based on multi-dimensional vectors on the evolution coordinate axis, graph neural network technology is used to aggregate cross-domain operation nodes. First, each multi-dimensional vector is taken as a feature representation of a node in the graph neural network, and these nodes represent different cross-domain operations. The graph neural network transmits information between nodes through a message passing mechanism according to the connection relationship between nodes, updates and aggregates node features. In this process, the network learns the complex associations between nodes and uncovers the hidden internal relationships between cross-domain operations, so that the originally scattered cross-domain operation nodes form an organic whole.

[0045] Based on the aggregated cross-domain operation nodes, the weight directional vector is configured. According to the connection strength between nodes, the importance of operations and other factors, each edge is assigned a corresponding weight to form a weight directional vector to accurately describe the interaction and influence direction between cross-domain operations. With the help of these weight directional vectors, the propagation path characteristics of the data version fingerprint in the data evolution topology graph are further extracted. Analyze how the data version fingerprint flows and changes between different operation nodes, and identify possible abnormal diffusion patterns. For example, check if there is a data version fingerprint propagation that violates normal permissions or operation procedures, so as to timely discover signs that the data may be tampered with or abnormally operated, and provide key basis for subsequent credibility assessment and data security maintenance.

[0046] In one possible implementation manner, step S200 further includes:

[0047] Step S240: Obtain the double-chain anchored index for credibility assessment of the data block hash value and the operation behavior entropy value in the current traceability verification time period.

[0048] Step S250: Determine the operation evidence chain with the double-chain anchored index, and establish a normal behavior mode baseline through the fluctuation characteristics of the operation behavior entropy value.

[0049] Step S260: Take the normal behavior mode baseline as a constraint condition to generate the credibility assessment matrix.

[0050] Specifically, first, for the data blocks in the current traceability verification time period, a specific hash algorithm (such as SHA-256) is used to process the data blocks, and the contents of the data blocks are converted into fixed-length hash values. These hash values become the unique identifier of the data block. Even if the data changes slightly, the hash value will be significantly different, thus ensuring accurate identification of data integrity. At the same time, the operation behavior in this period is carefully analyzed, and various operations are classified and counted to calculate the probability of each operation. Then, according to the information entropy formula, the operation behavior entropy value is calculated. This entropy value reflects the complexity and uncertainty of the operation behavior. Subsequently, in order to closely associate the data block hash value and the operation behavior entropy value for subsequent credibility evaluation, a double-chain anchoring index is constructed. Specifically, a special data structure (such as a dictionary nested list in Python) is created. During the traversal of all data blocks and operation behavior records, each data block hash value is associated with the relevant operation behavior entropy value and operation sequence number. Each operation behavior entropy value is also reversely associated with the corresponding data block hash value and operation sequence number, thus completing the construction of the double-chain anchoring index, providing an efficient indexing path for subsequent data credibility evaluation based on these two key indicators.

[0051] After determining the operation evidence chain by means of the double-chain anchoring index, the Isolation Forest algorithm is used to establish a normal behavior pattern baseline according to the fluctuation characteristics of the operation behavior entropy value. The corresponding relationship between the operation behavior entropy value and the data block hash value in the double-chain anchoring index is used to arrange the operation behavior entropy value in chronological order, thereby constructing the operation evidence chain, which completely records the operation track of the data within the current traceability verification time period. Then, a series of operation behavior entropy values in the operation evidence chain are collected as a data set, and the data set is input into the Isolation Forest algorithm. The Isolation Forest algorithm is an unsupervised learning algorithm and is very suitable for anomaly detection. The algorithm randomly selects features and randomly selects split points within the value range of the features, recursively divides the data, and constructs multiple isolated trees. Each data point is divided from the root node in the tree until it becomes a leaf node. The path length of the data point in the tree is an important basis for its anomaly score. For normal behavior data points, because they have similar feature patterns, the path in the tree is relatively long; and abnormal data points tend to be isolated more quickly, with a shorter path. The algorithm repeats this process multiple times to construct a forest, and then integrates the results of all trees to calculate a final anomaly score for each data point. Through learning of normal operation behavior data, the algorithm identifies the normal fluctuation range and pattern of the operation behavior entropy value, and determines a threshold value of the anomaly score according to the characteristics of these normal patterns. The operation behavior entropy value pattern corresponding to the data point below the threshold value is regarded as a normal behavior pattern, and a normal behavior pattern baseline is established in this way. When performing data credibility evaluation subsequently, if the anomaly score corresponding to the new operation behavior entropy value exceeds the threshold value, it can be determined that the operation may be abnormal.

[0052] All data operations and states in the current traceability verification time period are comprehensively sorted out, and for each data block hash value and corresponding operation behavior entropy value, it is compared in detail with the normal behavior mode baseline. For data operations within the normal behavior mode baseline range, a higher credibility score is given. For example, if the operation behavior entropy value is within one standard deviation of the mean value of the baseline, and the data block hash value is consistent with the data block hash value change rule under the normal process, a credibility score close to full score, such as 90 to 100 points, can be given. Conversely, for data operations deviating from the normal behavior mode baseline, the credibility score is reduced according to the degree of deviation. If the operation behavior entropy value exceeds the baseline mean value by more than two standard deviations, or the data block hash value appears abnormal fluctuation, such as too large difference with the previous normal data block hash value, a credibility score of 0 to 30 points can be given. The credibility scores of each data operation and state are arranged according to a certain matrix structure, with rows representing different data operation types and columns representing different data states or time nodes, to generate a credibility evaluation matrix. This matrix can intuitively and quantitatively show the credibility of data in each operation link and state, providing an important quantitative basis for subsequent data security analysis and decision-making.

[0053] In one possible implementation manner, step S300 further includes:

[0054] Step S310: connecting a tampering event feature template library, the tampering event feature template library including privilege escalation diffusion and watermark fracture mode.

[0055] Step S320: detecting an abnormal operation mode based on the tampering event feature template library, and calculating a KL divergence between the abnormal operation mode and a standard template.

[0056] Step S330: locating the data tampering event when the KL divergence exceeds a preset divergence threshold, and the preset divergence threshold is graded and elastically adjusted based on an anti-attack performance matrix.

[0057] Specifically, first, a tampering event feature template library is connected, which pre-stores a plurality of typical tampering event feature templates, such as privilege escalation diffusion features and watermark fracture mode features. The privilege escalation diffusion features record the behavior mode of abnormal promotion or diffusion of the privilege in the data operation process; the watermark fracture mode features describe in detail the situation of abnormal fracture, loss or damage of the watermark due to data tampering. These feature templates provide an important reference basis for subsequent detection of data tampering.

[0058] Based on the connected tampering event feature template library, real-time monitoring of the data operation flow is started to detect whether there is an abnormal operation mode. In the monitoring process, the actual observed operation mode is compared and analyzed with the standard template in the template library. In order to more accurately measure the degree of difference between the two, the method of calculating the KL divergence is adopted. The KL divergence is an index for measuring the difference between two probability distributions, and here it can quantify the degree of deviation between the actual operation mode and the normal mode represented by the standard template. Through feature extraction and probability distribution construction of operation data, for each possible tampering event feature template, the KL divergence value of the actual operation mode corresponding to it is calculated.

[0059] The divergence threshold is preset, and when the calculated KL divergence exceeds the preset divergence threshold, it is determined that a data tampering event has occurred, and the specific location of the data tampering event in the data operation flow is quickly located. It is worth noting that the preset divergence threshold is not fixed, but is hierarchically and elastically adjusted according to the anti-attack performance matrix. The anti-attack performance matrix comprehensively considers various factors such as the security threat level faced by the system, the importance of the data, and the anti-attack ability of the system itself. For example, in the case of high security threat or extremely important data, the anti-attack performance matrix will instruct the system to lower the preset divergence threshold, making the detection of data tampering events more sensitive, so as to timely discover potential security risks; when the security environment is relatively stable and the importance of the data is relatively low, the preset divergence threshold will be appropriately increased, avoiding excessive sensitivity and generating too many false positives, so as to ensure that the data tampering event can be efficiently and accurately located in different security scenarios.

[0060] In one possible implementation, step S330 further includes:

[0061] Step S331: Statistically count the tampering attack types and frequencies suffered in a plurality of traceability verification time periods, and set an attack mode distribution heat map.

[0062] Step S332: Introduce the success rate indicator and response delay indicator of the watermark reconstruction strategy, and calculate the success rate fuzzy parameter and response delay fuzzy parameter according to the attack mode distribution heat map.

[0063] Step S333: Configure the anti-attack performance matrix based on the success rate fuzzy parameter and the response delay fuzzy parameter.

[0064] Specifically, in-depth data statistical work is carried out for multiple traceability verification time periods. First, the data operation records, security audit logs and information captured by related monitoring systems in each period are comprehensively scanned. In the process of combing, each tampering attack event is accurately identified, and its attack type is carefully distinguished, such as determining whether it is an attack involving permission violation change, or an attack leading to watermark breakage and destroying the integrity of the watermark. At the same time, each type of tampering attack is counted, and its frequency in different traceability verification time periods is recorded. After completing the statistics, the statistical data is converted into an attack pattern distribution heat map using a visualization tool. In this heat map, the time period is taken as the horizontal axis, and the key link of data operation or the data area involved is taken as the vertical axis. The distribution of various tampering attacks is intuitively presented through different degrees of color depth. The area with high attack frequency will present a darker color, indicating that this time period or data area is more likely to be attacked; while the area with low attack frequency is lighter in color, meaning that it is relatively safe. In this way, the tampering attack situation of data in different stages and links can be clearly and intuitively understood, providing strong data visualization support for subsequent targeted analysis and response.

[0065] Two key indicators of the watermark reconstruction strategy are introduced, namely the success rate indicator and the response delay indicator. The success rate indicator reflects the proportion of the watermark reconstruction strategy that successfully recovers the watermark after being attacked, reflecting the effectiveness of the strategy; the response delay indicator measures the time interval from attack detection to watermark reconstruction completion, reflecting the timeliness of the strategy. Combined with the attack pattern distribution heat map generated earlier, different attack types and occurrence areas are analyzed. For areas with high attack frequency and dark color, the success rate and response delay data of the watermark reconstruction strategy in multiple attacks in this area are collected. Using fuzzy mathematics method, according to the distribution characteristics of the data and the experience setting of the membership function, the success rate and response delay data are converted into fuzzy sets, and the success rate fuzzy variable and the response delay fuzzy variable are calculated. These fuzzy variables can more comprehensively and flexibly describe the performance of the watermark reconstruction strategy under different attack patterns, providing more accurate basis for the configuration of the subsequent anti-attack efficiency matrix.

[0066] Based on the calculated success rate fuzzy variable and response delay fuzzy variable, the anti-attack performance matrix is configured. Each element of the anti-attack performance matrix corresponds to the anti-attack ability evaluation of the watermark reconstruction strategy under different attack modes. The element value in the matrix is obtained by performing specific operations and combinations on the success rate fuzzy variable and the response delay fuzzy variable. For example, a high value of the success rate fuzzy variable and a low value of the response delay fuzzy variable are assigned a higher anti-attack performance score, and vice versa. In this way, the anti-attack performance matrix comprehensively and quantitatively reflects the comprehensive performance of the watermark reconstruction strategy when facing various tampering attacks, providing an important basis for flexibly adjusting the preset divergence threshold and optimizing the watermark reconstruction strategy according to different security scenarios.

[0067] In one possible implementation, step S332 further includes:

[0068] Defining a success rate fuzzy variable wherein n is the total number of attack types, w is the frequency weight of the i-th attack in the attack mode distribution heat map, f i N represents the cumulative frequency of the i-th attack within the traceability verification time period, N success,i N is the number of successful watermark reconstruction in the i-th attack, N total,i N is the total number of triggers of the i-th attack.

[0069] Specifically, the success rate fuzzy variable calculation formula is n is the total number of attack types, and in actual data security monitoring scenarios, there are various types of attacks, such as privilege escalation diffusion and malicious tampering leading to watermark rupture, and n is the sum of these attack types. i w is the frequency weight of the i-th attack in the attack mode distribution heat map, and its calculation method is wherein f i N represents the cumulative frequency of the i-th attack within the traceability verification time period. This means that the more frequent the attack occurs, the greater the weight of this type of attack in calculating the success rate fuzzy variable. For example, if a certain type of attack frequently occurs in multiple time periods, its influence on the overall success rate fuzzy variable will be more significant. success,i N is the number of successful watermark reconstruction in the i-th attack, N total,i N is the total number of triggers of the i-th attack. By the ratio of N success,i to N total,i , the success rate of watermark reconstruction under the i-th attack can be obtained. Multiply this success rate by the corresponding frequency weight w iThe success rate fuzzy variable s is obtained by accumulating all attack types and dividing by the total frequency weight of all attack types. The variable comprehensively considers the frequency of different attack types and the success of watermark reconstruction under various attacks, and can more comprehensively and accurately reflect the comprehensive success rate performance of the watermark reconstruction strategy under various attack environments, thereby providing an important basis for subsequent evaluation of system attack resistance.

[0070] In one possible implementation manner, step S332 further includes:

[0071] Defining a response delay fuzzy variable wherein T represents the total number of statistical time windows in the traceability verification time period, T is the total number of statistical time windows in the traceability verification time period, and T is the total number of statistical time windows in the traceability verification time period. is the average delay of watermark reconstruction in the tth window, and is the average delay of watermark reconstruction in the tth window. is the maximum delay in the tth window, and is the maximum delay in the tth window. α and β are weight coefficients (α+β=1), and are used to balance the influence of average and peak delay. λ is a time decay factor.

[0072] The anti-attack performance matrix is output by using the success rate fuzzy variable, the success rate segmentation boundary under the confidence level, the response delay fuzzy variable, and the reference response delay tolerance value.

[0073] Specifically, a response delay fuzzy variable is first defined, and the formula is wherein T represents the total number of statistical time windows in the traceability verification time period, and T is the total number of statistical time windows in the traceability verification time period. is the average delay of watermark reconstruction in the tth window, and is the average delay of watermark reconstruction in the tth window. is the maximum delay in the tth window, and is the maximum delay in the tth window. α and β are weight coefficients (α+β=1), and are used to balance the influence of average and peak delay. λ is a time decay factor.

[0074] The success rate fuzzy variable and the confidence level success rate split boundary, the response delay fuzzy variable and the reference response delay tolerance value are used as the basis to output the attack resistance performance matrix. The confidence level success rate split boundary is a threshold value preset for dividing different levels of the success rate fuzzy variable. The level of the success rate fuzzy variable is determined by comparing the success rate fuzzy variable with the boundary values. The reference response delay tolerance value is also a standard preset for measuring whether the response delay fuzzy variable is within an acceptable range. The two fuzzy variables are compared and analyzed with the corresponding boundary values and tolerance values, and the success rate of watermark reconstruction and the response delay are comprehensively considered to determine the element values of the attack resistance performance matrix, so as to construct a matrix that can reflect the attack resistance ability under different attack conditions, thereby providing strong support for security evaluation and strategy adjustment.

[0075] In the second embodiment, based on the same inventive concept as the traceability verification method for a trusted data space in the foregoing embodiments, as shown in the drawings, the application provides a traceability verification device for a trusted data space. The device and method embodiments in the application are based on the same inventive concept. The device comprises: Figure 2

[0076] The multi-dimensional traceability data acquisition module 10 is configured to embed dynamic watermark identifiers in the data operation flow and capture multi-dimensional traceability data including data version fingerprints, access trajectory maps and permission change sequences in real time.

[0077] The trustworthiness evaluation matrix generation module 20 is configured to perform cross-domain feature fusion on the multi-dimensional traceability data, construct a data evolution topology graph, and generate a trustworthiness evaluation matrix in combination with data integrity hash trees and operation behavior entropy analysis.

[0078] The risk level evaluation module 30 is configured to locate data tampering events and evaluate traceability chain fracture risk levels according to the data evolution topology graph and the trustworthiness evaluation matrix.

[0079] The verification instruction set acquisition module 40 is configured to input the feature vectors and risk levels of the data tampering events into a verification mapping array to obtain a verification instruction set including data snapshot repair paths, permission traceability backtracking schemes and watermark reconstruction strategies.

[0080] The trusted data space reconstruction module 50 is configured to monitor execution results according to the verification instruction set, generate a trustworthiness calibration report and reconstruct the trusted data space of the data operation flow.

[0081] Further, the device is also configured to implement the following functions:

[0082] ​Inject reversible encryption watermark in metadata layer of data operation flow, the reversible encryption watermark contains time-sensitive key and operator identity; based on the reversible encryption watermark, set watermark dynamic update rule, the watermark dynamic update rule is used to trigger watermark reconstruction when data version fingerprint changes, or trigger watermark nested encryption when permission changes.

[0083] Further, the device is also used to realize the following functions:

[0084] By the multi-dimensional traceability data, configure data operation track; map the data operation track to a multi-dimensional vector on the evolution coordinate axis, each vector contains operation type weight and context correlation degree; based on the multi-dimensional vector on the evolution coordinate axis, perform cross-domain feature fusion on the multi-dimensional traceability data.

[0085] Further, the device is also used to realize the following functions:

[0086] Based on the multi-dimensional vector on the evolution coordinate axis, aggregate cross-domain operation nodes using graph neural network; based on the cross-domain operation nodes, configure weight directional vector, extract data version fingerprint propagation path features and abnormal diffusion patterns in the data evolution topology graph.

[0087] Further, the device is also used to realize the following functions:

[0088] Obtain double-chain anchoring index for trustworthiness evaluation of data block hash value and operation behavior entropy value within the current traceability verification time period; determine operation evidence chain with the double-chain anchoring index, and establish normal behavior mode baseline through the fluctuation characteristics of operation behavior entropy value; take the normal behavior mode baseline as a constraint condition to generate the trustworthiness evaluation matrix.

[0089] Further, the device is also used to realize the following functions:

[0090] Connect tampering event feature template library, the tampering event feature template library contains permission overstep diffusion and watermark breaking mode; based on the tampering event feature template library, detect abnormal operation mode and calculate KL divergence with the standard template; when the KL divergence exceeds the preset divergence threshold, locate the data tampering event, and the preset divergence threshold is hierarchically and elastically adjusted with the attack resistance performance matrix.

[0091] Further, the device is also used to realize the following functions:

[0092] The attack mode distribution heat map is set by counting the types and frequencies of tampering attacks suffered in a plurality of traceability verification time periods; the success rate indicator and the response delay indicator of the watermark reconstruction strategy are introduced, and the success rate fuzzy variable and the response delay fuzzy variable are calculated according to the attack mode distribution heat map; and the attack resistance performance matrix is configured based on the success rate fuzzy variable and the response delay fuzzy variable.

[0093] Further, the apparatus is further configured to implement the following functions:

[0094] Define the success rate fuzzy variable wherein n is the total number of attack types, is the frequency weight of the i th attack in the attack mode distribution heat map, f i represents the cumulative frequency of the i th attack in the traceability verification time period, N success,i is the number of successful watermark reconstruction in the i th attack, N total,i is the total number of triggers of the i th attack.

[0095] Further, the apparatus is further configured to implement the following functions:

[0096] Define the response delay fuzzy variable wherein T is the total number of statistical time windows in the traceability verification time period, is the average delay of watermark reconstruction in the t th window, is the maximum delay in the t th window, and α and β are weight coefficients (α+β=1) for balancing the influence of average and peak delay, and λ is a time decay factor; the attack resistance performance matrix is outputted with the success rate fuzzy variable, the success rate segmentation boundary under the confidence level, the response delay fuzzy variable, and the reference response delay tolerance value.

[0097] It should be noted that the above sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above describes a specific embodiment of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

[0098] The above only describes the preferred embodiments of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0099] The specification and drawings are, of course, to be regarded in an illustrative rather than a restrictive sense. It is to be understood that any such modifications, variations, combinations or equivalents that fall within the scope of the application are intended to be embraced herein.

Claims

1. A method for provenance verification of a trusted data space, characterized in that, The method comprises: Embedding dynamic watermark identification in the data operation flow, capturing multi-dimensional traceability data containing data version fingerprints, access trajectory maps and permission change sequences in real time; Cross-domain feature fusion is performed on the multi-dimensional traceability data to construct a data evolution topology graph, and a credibility evaluation matrix is generated by combining data integrity hash tree and operation behavior entropy analysis; According to the data evolution topology graph and the credibility evaluation matrix, the data tampering event is located, and the traceability chain breakage risk level is evaluated; The feature vector and risk level of the data tampering event are input into the test mapping array to obtain a verification instruction set containing data snapshot repair path, permission traceability backtracking scheme and watermark reconstruction strategy; According to the verification instruction set, the monitoring execution result is generated, the credibility calibration report is generated, and the credible data space of the data operation flow is reconstructed; Wherein, embedding dynamic watermark identification in the data operation flow comprises: Injecting reversible encryption watermark in the metadata layer of the data operation flow, the reversible encryption watermark containing time-sensitive key and operator identity; Based on the reversible encryption watermark, set the watermark dynamic update rule, the watermark dynamic update rule is used to trigger watermark reconstruction when the data version fingerprint changes, or trigger watermark nested encryption when the permission changes; Wherein, the cross-domain feature fusion of the multi-dimensional traceability data to construct the data evolution topology graph comprises: Configure the data operation trajectory through the multi-dimensional traceability data; Map the data operation trajectory to a multi-dimensional vector on the evolution coordinate axis, each vector containing operation type weight and context correlation degree; Based on the multi-dimensional vector on the evolution coordinate axis, the multi-dimensional traceability data is cross-domain feature fused; Wherein, based on the multi-dimensional vector on the evolution coordinate axis, the multi-dimensional traceability data is cross-domain feature fused, comprising: Based on the multi-dimensional vector on the evolution coordinate axis, the cross-domain operation node is aggregated by using graph neural network; Based on the cross-domain operation node, configure the weight directional vector, extract the propagation path feature and abnormal diffusion mode of the data version fingerprint in the data evolution topology graph.

2. The provenance verification method for trusted data space of claim 1, wherein, Combining data integrity hash tree and operation behavior entropy analysis to generate credibility evaluation matrix, comprising: Obtain the double-chain anchoring index of the credibility evaluation of the data block hash value and the operation behavior entropy in the current traceability verification time period; Determine the operation evidence chain by the double-chain anchoring index, and establish the normal behavior mode baseline through the fluctuation characteristics of the operation behavior entropy; The normal behavior mode baseline is used as a constraint condition to generate the credibility evaluation matrix.

3. The provenance verification method for trusted data spaces of claim 1, wherein, According to the data evolution topology graph and the credibility evaluation matrix, the data tampering event is located, comprising: Connecting the tampering event feature template library, the tampering event feature template library contains permission overstep diffusion and watermark breakage mode; Based on the tampering event feature template library, detect abnormal operation mode and calculate the KL divergence with the standard template; When the KL divergence exceeds the preset divergence threshold, the data tampering event is located, and the preset divergence threshold is hierarchically and elastically adjusted by the attack resistance performance matrix.

4. A provenance verification method for a trusted data space as claimed in claim 3, wherein, The preset divergence threshold is hierarchically and elastically adjusted by the attack resistance performance matrix, comprising: Set an attack mode distribution heat map by counting the types and frequencies of tampering attacks suffered in multiple provenance verification time periods; Introduce a success rate indicator and a response delay indicator of the watermark reconstruction strategy, and calculate a success rate fuzzy variable and a response delay fuzzy variable according to the attack mode distribution heat map; Configure an anti-attack performance matrix based on the success rate fuzzy variable and the response delay fuzzy variable.

5. A provenance verification method for a trusted data space as claimed in claim 4, wherein, The method for calculating the success rate fuzzy variable comprises: Defining success rate fuzzy variables where n is the total number of attack types, is the frequency weight of the i-th attack in the attack pattern distribution heat map, represents the cumulative frequency of the i-th attack within the traceability verification time period, is the number of successful watermark reconstruction in the i-th attack, is the total number of triggers of the i-th attack.

6. A provenance verification method for a trusted data space as claimed in claim 5, wherein, The method for configuring the anti-attack performance matrix based on the success rate fuzzy variable and the response delay fuzzy variable comprises: Defining response latency ambiguity variables where T is the total number of statistical time windows within the provenance verification time period, is the average latency for watermark reconstruction within the tth window, is the maximum latency within the tth window, , is a weight coefficient (0 < a < 1) to balance the influence of average and peak latency, + = 1), to balance the influence of average and peak latency, is a time decay factor; Output the anti-attack performance matrix with the success rate fuzzy variable and a success rate segmentation boundary at a confidence level, and the response delay fuzzy variable and a reference response delay tolerance value.

7. An apparatus for provenance verification of a trusted data space, characterized in that, The device is used to implement the provenance verification method for a trusted data space according to any one of claims 1-6, and the device comprises: A multi-dimensional provenance data acquisition module is configured to embed dynamic watermark identification in a data operation flow, and to capture multi-dimensional provenance data including data version fingerprints, access trajectory maps and permission change sequences in real time; A trustworthiness evaluation matrix generation module is configured to perform cross-domain feature fusion on the multi-dimensional provenance data, construct a data evolution topology graph, and generate a trustworthiness evaluation matrix in combination with a data integrity hash tree and operation behavior entropy analysis; A risk level evaluation module is configured to locate a data tampering event and evaluate a provenance chain breakage risk level according to the data evolution topology graph and the trustworthiness evaluation matrix; A verification instruction set acquisition module is configured to input a feature vector and a risk level of the data tampering event into a verification mapping array to obtain a verification instruction set including a data snapshot repair path, a permission provenance backtracking scheme and a watermark reconstruction strategy; A trusted data space reconstruction module is configured to monitor execution results according to the verification instruction set, generate a trustworthiness calibration report and reconstruct a trusted data space of the data operation flow; The method for embedding dynamic watermark identification in a data operation flow comprises: Inject reversible encryption watermark in a metadata layer of the data operation flow, wherein the reversible encryption watermark contains a time-sensitive key and an operator identity; Set a watermark dynamic update rule based on the reversible encryption watermark, wherein the watermark dynamic update rule is used to trigger watermark reconstruction when a data version fingerprint is changed, or trigger watermark nested encryption when a permission is changed; The method for performing cross-domain feature fusion on the multi-dimensional provenance data to construct a data evolution topology graph comprises: Configure a data operation trajectory through the multi-dimensional provenance data; Map the data operation trajectory into a multi-dimensional vector on an evolution coordinate axis, wherein each vector contains an operation type weight and a context correlation degree; Perform cross-domain feature fusion on the multi-dimensional provenance data based on the multi-dimensional vector on the evolution coordinate axis; The method for performing cross-domain feature fusion on the multi-dimensional provenance data based on the multi-dimensional vector on the evolution coordinate axis comprises: Aggregate cross-domain operation nodes by using a graph neural network based on the multi-dimensional vector on the evolution coordinate axis; Based on the cross-domain operation node, a weight value directional vector is configured, and a propagation path feature and an abnormal diffusion mode of a data version fingerprint in the data evolution topology graph are extracted.

Citation Information

Patent Citations

  • Performance-lossless watermark credible traceability method and system based on block chain

    CN119622672A

  • Cross-modal image-watermark joint generation and detection device and method thereof

    US12125119B1