A data trust verification method and device in the scenario of data elements

By combining the trustworthy measurement engine and distributed ledger at the data supplier, trustworthy verification of the data access and computing stages is achieved, the problems of data integrity and confidentiality in data circulation are solved, and the trustworthiness guarantee of data circulation is improved.

CN117113299BActive Publication Date: 2025-07-29BEIJING INTERNETWARE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311028357.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-15
Publication Date
2025-07-29
Estimated Expiration
2043-08-15

AI Technical Summary

Technical Problem

In data circulation, it is difficult for the prior art to effectively ensure the integrity of the data itself provided by the data supplier and the confidentiality of the execution results of the data algorithm, especially in the data transmission and storage links, there is a risk of data tampering and privacy calculation results leaking.

Method used

Through the trusted metric engine of the data provider, run-time measurement of the target data interface execution process, obtain multi-dimensional legal measurement values, and combine it with a distributed ledger for integrity verification, and data leakage detection is carried out on the data algorithm to ensure security when processing data in a private computing environment.

Benefits of technology

It enhances the integrity guarantee in the data access stage and the confidentiality detection in the data calculation stage, ensures the credibility of data circulation, shifts from the intermediate side to the supply and demand side, and improves the credibility of data circulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117113299B_ABST
    Figure CN117113299B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for data trust verification in the scenario of data elements. The method includes: during the process of obtaining target data through a target data interface, performing runtime measurement on the execution process of the target data interface through the trust measurement engine of the data provider to obtain a measurement result; obtaining multi-dimensional legal measurement values corresponding to the target data interface; comparing the measurement result with the multi-dimensional legal measurement values to obtain a target data access integrity verification result; performing data leakage detection on the received data algorithm to obtain a data leakage detection result; in the case where both the target data access integrity verification result and the data leakage detection result are qualified, processing the target data through the data algorithm to obtain a processing result; and sending the processing result to the data demander. The aim is to simultaneously perform trust verification on the integrity of the data itself provided by the data provider and the potential execution result leakage defects of the data algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data circulation trusted verification, and particularly to a data trusted verification method and device in the scenario of data elements. Background Art

[0002] Today, data has been widely used to guide the role of various resources and promote the development of productivity, profoundly changing the social production and lifestyle. As an important production factor, data needs to circulate among different entities to create value. Data circulation refers to the behavior of exchanging data between data suppliers and demanders according to certain circulation rules, which can promote the in-depth mining and integration of data, improve the decision-making ability driven by data, and thus realize the improvement of the performance and value of products and services.

[0003] However, while bringing huge value, data circulation also has many trust issues. For example, data suppliers may provide forged malicious data, resulting in errors in the relevant operations of data demanders; data demanders may use or steal data illegally, infringing on the data suppliers' ownership of the data. Such trust issues that hinder or even prevent data circulation have become the key challenges that data circulation needs to address. Although relevant research and practices have made remarkable progress and achievements in the trusted guarantee of data circulation, there are still two deficiencies: (1) The guarantee of data integrity is insufficient. Data circulation is a multi-link complex process. Once data tampering occurs in a certain link, all subsequent links will be affected. Existing work pays more attention to how to guarantee the integrity of data in intermediate links such as transmission and storage, and the guarantee of the integrity of the data provided by data suppliers themselves is not sufficient, making it difficult to confirm whether the accessed data is true and effective. (2) The guarantee of data confidentiality is insufficient. When there are defects in the data algorithms executed in the privacy computing environment, the results of privacy computing may leak confidential data that should be protected. Although existing work on data confidentiality guarantee has enhanced data confidentiality from multiple aspects such as original data and computing environment, the evaluation of whether the execution results of data algorithms may damage data confidentiality is still not sufficient, making it difficult to discover potential data leakage defects in the algorithms. Summary of the Invention

[0004] In view of this, the present invention provides a data trusted verification method in the scenario of data elements. It aims to simultaneously perform trusted verification on the integrity of the data itself provided by data suppliers and potential execution result leakage defects of data algorithms, so as to shift the perspective of trusted guarantee of data circulation from being oriented towards intermediate parties to equally emphasizing both suppliers and demanders.

[0005] In the first aspect of the embodiments of the present invention, a data trusted verification method in the scenario of data elements is provided. The method includes:

[0006] During the process of the data provider obtaining target data through the target data interface, the running process of the target data interface is measured at runtime by the trusted measurement engine of the data provider to obtain a measurement result;

[0007] Obtain the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger;

[0008] Compare the measurement result with the multi-dimensional legal measurement values to obtain the target data access integrity verification result;

[0009] Perform data leakage detection on the received data algorithm to obtain a data leakage detection result;

[0010] When both the target data access integrity verification result and the data leakage detection result are qualified, process the target data through the data algorithm in the privacy computing environment to obtain a processing result;

[0011] Send the processing result to the data demander.

[0012] Optionally, before obtaining the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger, the method further includes:

[0013] In the offline state, obtain a first target program by temporarily instrumenting a first original target program, and obtain a static control flow graph by performing static analysis on the first target program;

[0014] During the execution of the target data interface in the first target program, monitor the execution process of the target data interface through the trusted measurement engine of the data platform in the trusted execution environment to construct a corresponding runtime behavior model, and obtain the execution result of the target data interface;

[0015] Match the execution result with the data flow information in the runtime behavior model to determine the entry basic block and the exit basic block of the target data interface;

[0016] In the static control flow graph, starting from the entry basic block, determine all basic blocks that can be reached from the entry basic block through depth-first search to obtain a first search result, and starting from the exit basic block, determine all basic blocks that can reach the exit basic block through depth-first search to obtain a second search result;

[0017] Determine the code segment corresponding to the target data interface according to the first search result and the second search result;

[0018] Construct an equivalent code segment corresponding to the code segment by slicing the code segment;

[0019] Perform permanent instrumentation on the equivalent code snippet, and perform runtime measurement on the static code and execution process of the permanently instrumented equivalent code snippet to obtain multi-dimensional legal measurement values;

[0020] Store the multi-dimensional legal measurement values in a distributed ledger.

[0021] Optionally, constructing an equivalent code snippet corresponding to the code snippet includes:

[0022] Construct an initial equivalent code snippet decoupled from other business functions according to the code snippet of the target data interface;

[0023] Obtain the equivalent code snippet by restoring the context environment of the initial equivalent code snippet.

[0024] Optionally, during the process of the data provider obtaining target data through the target data interface, perform runtime measurement on the execution process of the target data interface through the trusted measurement engine of the data provider to obtain a measurement result, including:

[0025] Call the permanently instrumented equivalent code snippet to execute to obtain target data;

[0026] During the execution process of the permanently instrumented equivalent code snippet, measure and sign the static code, execution process, and execution result through the trusted measurement engine of the data provider to obtain a measurement result.

[0027] Optionally, performing data leakage detection on the received data algorithm to obtain a data leakage detection result includes:

[0028] Instrument the second original target program corresponding to the data algorithm to obtain a second target program;

[0029] Use the initial seed file as input and execute the second target program through fuzz testing to obtain an execution log;

[0030] Obtain an execution log analysis result by analyzing the execution log;

[0031] Determine whether there is a new execution path during the current execution process according to the execution log analysis result;

[0032] In the case where there is a new execution path during the current execution process, add the seed files involved in the current execution process to the input queue, and perform data leakage detection on the seed files involved in the current execution process to update the data leakage detection result;

[0033] Obtain the latest seed file in the input queue for mutation, and use the mutated seed file as input to execute the second target program through fuzz testing to obtain an execution log. Return to the step: Analyze the execution log to obtain an execution log analysis result.

[0034] Optionally, perform data leakage detection on the seed files involved in the current execution process to update the data leakage detection result, including:

[0035] Through a path-sensitive byte-granularity test sampling strategy, perform byte-granularity input-output sampling on the seed files involved in the current execution process to obtain multiple sampling results;

[0036] Perform entropy analysis on each of the multiple sampling results to obtain an entropy analysis result corresponding to each sampling result;

[0037] Update the data leakage detection result according to the entropy analysis result corresponding to each sampling result.

[0038] Optionally, after analyzing the execution log to obtain an execution log analysis result, the method further includes:

[0039] Determine whether there are unexecuted branches in the current execution process according to the execution log analysis result;

[0040] In the case where there are unexecuted branches in the current execution process, expand the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label;

[0041] Guide the new taint label information to perform targeted input mutations one by one, and use the mutated seed file as input to execute the second target program through fuzz testing to obtain an execution log. Return to the step: Analyze the execution log to obtain an execution log analysis result.

[0042] Optionally, expanding the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label includes:

[0043] Record the execution path of the second target program in the current execution process to obtain relevant taint propagation information;

[0044] During the process of executing the second target program again through the seed files in the current execution process, pass the identifier of the unexecuted branch to the second target program to determine the unexecuted branch for dynamic information flow analysis;

[0045] Change each byte of the seed file in the current execution process one by one and re-execute the second target program to obtain the value of the branch variable of the unexecuted branch when running under the changed seed file each time;

[0046] When the obtained running value of the branch variable of the unexecuted branch is valid and different from the initial value, mark the position corresponding to the currently changed byte as a tag to be extended;

[0047] Determine the taint label of the branch variable corresponding to the unexecuted branch according to the relevant taint propagation information;

[0048] Use the tag to be extended to extend the taint label of the branch variable corresponding to the unexecuted branch to obtain a new taint label.

[0049] The second aspect of the present invention provides a data trust verification device in a data element scenario, and the device includes:

[0050] A trusted measurement engine for performing runtime measurement on the execution process of the target data interface during the process of the data provider obtaining target data through the target data interface to obtain a measurement result;

[0051] A data acquisition module for acquiring the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger;

[0052] A data integrity validator for comparing the measurement result with the multi-dimensional legal measurement values to obtain a target data access integrity verification result;

[0053] A data leakage measurer for performing data leakage detection on the received data algorithm to obtain a data leakage detection result;

[0054] An algorithm executor for processing the target data through the data algorithm in a privacy computing environment to obtain a processing result when both the target data access integrity verification result and the data leakage detection result are qualified;

[0055] A processing result return module for sending the processing result to the data demander.

[0056] Regarding the prior art, the present invention has the following advantages:

[0057] A data trustworthy verification method in a data element scenario provided by an embodiment of the present invention is applied to a data platform. During the process that a data provider obtains target data through a target data interface, a runtime measurement is performed on the execution process of the target data interface by the trustworthy measurement engine of the data provider to obtain a measurement result; a multi-dimensional legal measurement value corresponding to the target data interface stored in a distributed ledger is obtained; the measurement result is compared with the multi-dimensional legal measurement value to obtain a target data access integrity verification result; a data leakage detection is performed on the received data algorithm to obtain a data leakage detection result; in the case where both the target data access integrity verification result and the data leakage detection result are qualified, the target data is processed by the data algorithm in a privacy computing environment to obtain a processing result; and the processing result is sent to a data demander. Thus, the present invention adds a link to perform trustworthy verification on data respectively in the data access stage and the data usage stage. In the data access stage, the collection, evidence storage, and verification of runtime information for verifying the integrity of the data access process in the data provider's information system are added, so as to enhance the ability to guarantee data integrity; in the data calculation stage, an automated detection of potential data leakage defects in the data algorithm that accesses and analyzes data is added, so as to enhance the ability to guarantee data confidentiality. Thereby, the trustworthy guarantee perspective of data circulation shifts from being oriented to the intermediate party to equally emphasizing both the provider and the demander.

[0058] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically gives the specific implementation manners of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0060] Figure 1 It is a flowchart of a data trustworthy verification method in a data element scenario provided by an embodiment of the present invention;

[0061] Figure 2 It is a schematic diagram of an example of quantifying the leakage degree based on information entropy in a data trustworthy verification method in a data element scenario provided by an embodiment of the present invention;

[0062] Figure 3 It is a schematic diagram of another example of quantifying the leakage degree based on information entropy in a data trustworthy verification method in a data element scenario provided by an embodiment of the present invention;

[0063] Figure 4Schematic diagram of a data trust verification device in a data element scenario provided by an embodiment of the present invention. Detailed implementation manners

[0064] The exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings.

[0065] Figure 1 Flowchart of a data trust verification method in a data element scenario provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0066] Step S11: During the process that the data provider obtains target data through the target data interface, perform runtime measurement on the execution process of the target data interface through the trust measurement engine of the data provider to obtain a measurement result;

[0067] Step S12: Obtain the multi-dimensional legal measurement values stored in the distributed ledger corresponding to the target data interface;

[0068] Step S13: Compare the measurement result with the multi-dimensional legal measurement values to obtain a target data access integrity verification result;

[0069] Step S14: Perform data leakage detection on the received data algorithm to obtain a data leakage detection result;

[0070] Step S15: When both the target data access integrity verification result and the data leakage detection result are qualified, process the target data through the data algorithm in a privacy computing environment to obtain a processing result;

[0071] Step S16: Send the processing result to the data demander.

[0072] In the embodiment of the present invention, in the scenario of data circulation, data usually comes from the access results of specific data interfaces (API, Application Programming Interface) in the information system of the data provider. The data platform refers to a platform that provides multiple data interfaces and common services for the data provider and the data demander. All subsequent descriptions of data circulation in the present invention are for platform-based data circulation. The circulation process of platform-based data circulation is that the data platform obtains the data of the data provider, processes the obtained data through a data algorithm to obtain a processing result, and sends the obtained processing result to the data demander.

[0073] With the transformation of the trust model in the data circulation scenario from the previous untrusted trust model of the data platform to the trust model dominated by official management of the data platform, in order to adapt to the transformation of the trust model, the present invention changes the trust guarantee from the previous orientation towards the data platform to an equal emphasis on both the provider and the demander.

[0074] In this circulation process, the data platform will send a target data acquisition request to the data provider. When the data provider receives the target data acquisition request, it obtains the target data by executing the target data interface corresponding to the target data in the target program. In order to increase the credibility guarantee of the data, in the process of the data provider sending the target data to the data platform, the trusted measurement engine in the data provider performs runtime measurement on the execution process of the target data interface to obtain the corresponding measurement result. The distributed ledger stores multi-dimensional legal measurement values for verifying the integrity of the measurement result of the execution process of the target data interface.

[0075] After obtaining the measurement result of the execution process of the target data interface, the data platform obtains the multi-dimensional legal measurement values stored in the distributed ledger for verifying the integrity of the measurement result of the execution process of the target data interface and the measurement result obtained from the trusted measurement engine in the data provider. The data platform compares the measurement result with the obtained multi-dimensional legal measurement values to obtain the target data access integrity verification result, and this target data access integrity verification result will indicate whether the target data obtained by executing the target data interface is trustworthy.

[0076] The data platform also receives a data algorithm for processing the target data. After receiving the data algorithm for processing the target data, it performs data leakage detection on the data algorithm to obtain the corresponding data leakage detection result. The data algorithm can be provided by the data demander or other algorithm providers.

[0077] When the target data access integrity verification result indicates that the target data obtained by executing the target data interface is trustworthy, and the data leakage detection result indicates that the data algorithm is trustworthy, it shows that both the target data access integrity verification result and the data leakage detection result are qualified. At this time, the target data will be processed by the data algorithm in the privacy computing environment of the data platform to obtain a processing result. Then the obtained processing result is sent to the data demander, thus completing the trusted verification of platform-based data circulation in the data element scenario.

[0078] Among them, the data provider refers to the data provider information system for providing data externally. The target data represents the data that the current data platform needs to request. The trusted measurement engine refers to the software running in the trusted execution environment for performing measurement and constructing a proof report (i.e., data integrity evidence).

[0079] In an embodiment of the present invention, in the data circulation scenario, for the data provider, what it is trusted to do is to provide real and effective access data as promised, but it may provide forged malicious data, resulting in errors in the relevant operations of the data demander; for the data demander, what it is trusted to do is to use the data as promised, but it may use or steal the data illegally, thus infringing on the data provider's ownership of the data. Trusted guarantee refers to avoiding the occurrence of the above-mentioned trust problems.

[0080] A data trusted verification method in a data element scenario provided by an embodiment of the present invention is applied to a data platform. During the process of the data provider obtaining target data through a target data interface, the trusted measurement engine of the data provider performs runtime measurement on the execution process of the target data interface to obtain a measurement result; obtains the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger; compares the measurement result with the multi-dimensional legal measurement values to obtain a target data access integrity verification result; performs data leakage detection on the received data algorithm to obtain a data leakage detection result; in the case where both the target data access integrity verification result and the data leakage detection result are qualified, processes the target data through the data algorithm in a privacy computing environment to obtain a processing result; and sends the processing result to the data demander. Thus, the present invention adds a link to perform trusted verification on the data in the data access stage and the data usage stage respectively. In the data access stage, it adds the collection, storage, and verification of runtime information for verifying the integrity of the data access process in the data provider's information system, so as to enhance the ability to guarantee data integrity; in the data calculation stage, it adds the automated detection of potential data leakage defects in the data algorithm that accesses and analyzes the data, so as to enhance the ability to guarantee data confidentiality. Thereby, the trusted guarantee perspective of data circulation shifts from being oriented towards the intermediate party to equally emphasizing both the provider and the demander.

[0081] In the present invention, before obtaining the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger, the method further includes: in an offline state, obtaining a first target program by temporarily instrumenting a first original target program, and obtaining a static control flow graph by performing static analysis on the first target program; during the execution of the target data interface in the first target program, monitoring the execution process of the target data interface by the trusted measurement engine of the data platform in the trusted execution environment to construct a corresponding runtime behavior model, and obtaining the execution result of the target data interface; matching the execution result with the data flow information in the runtime behavior model to determine the entry basic block and the exit basic block of the target data interface; in the static control flow graph, starting from the entry basic block, determining all basic blocks that can be reached from the entry basic block through depth-first search to obtain a first search result, and starting from the exit basic block, determining all basic blocks that can reach the exit basic block through depth-first search to obtain a second search result; determining the code segment corresponding to the target data interface according to the first search result and the second search result; slicing the code segment to construct an equivalent code segment corresponding to the code segment; performing permanent instrumentation on the equivalent code segment, and performing runtime measurement on the static code and the execution process of the permanently instrumented equivalent code segment to obtain multi-dimensional legal measurement values; storing the multi-dimensional legal measurement values in the distributed ledger.

[0082] In an embodiment of the present invention, in order to reduce the time delay of obtaining data when the data platform performs online integrity verification on the target data access process, multi-dimensional legal measurement values for integrity verification of the measurement results of the execution process of the target data interface are obtained in advance through offline static program analysis and stored in the distributed ledger to ensure the credibility of the multi-dimensional legal measurement values.

[0083] Thus, when the data platform performs online integrity verification on the target data access process, it can directly obtain the multi-dimensional legal measurement values stored in the distributed ledger for integrity verification of the measurement results of the execution process of the target data interface, without generating the multi-dimensional legal measurement values for integrity verification of the measurement results of the execution process of the target data interface during online integrity verification, thereby reducing the time overhead of online integrity verification.

[0084] In an embodiment of the present invention, a distributed ledger stores multi-dimensional legal measurement values corresponding to various data interfaces. For example, for a data interface a for obtaining data A, a multi-dimensional legal measurement value a' for performing integrity verification on the measurement result of the execution process of the data interface a is stored in the distributed ledger; for a data interface b for obtaining data B, a multi-dimensional legal measurement value b' for performing integrity verification on the measurement result of the execution process of the data interface b is stored in the distributed ledger; for a data interface c for obtaining data C, a multi-dimensional legal measurement value c' for performing integrity verification on the measurement result of the execution process of the data interface c is stored in the distributed ledger, and so on. At the same time, each multi-dimensional legal measurement value for performing integrity verification on the measurement result of the execution process of a data interface can be reused repeatedly. Just obtain the stored multi-dimensional legal measurement value for performing integrity verification on the measurement result of the execution process of the data interface from the distributed ledger according to the data interface for online integrity verification, without the need for multiple calculations. In this way, the time overhead of online verification can be greatly reduced.

[0085] However, this offline static program analysis method requires instrumenting specific program locations of the target data interface. This instrumentation is used to record runtime control flow information and interact with a trusted measurement engine to determine the runtime processing logic for the target data interface. In current technologies, such instrumentation is usually performed on the complete program of the data provider information system. This approach will bring a large code bloat rate and additional overhead. On the one hand, the control flow graph of complex business processing programs is usually extremely complex. Performing offline static program analysis on the complete program to obtain legal measurement values is likely to lead to state explosion, bringing a huge verification overhead. On the other hand, since modular reuse is a common practice in the design of such data provider information systems, when performing offline static program analysis and instrumenting the measurement logic for the complete program, the instrumented program module may be called when other business functions unrelated to the target data interface are executed. This may result in unnecessary additional overhead and interfere with the measurement result of the execution process of the target data interface being built. To solve this problem, the present invention performs program slicing on the complete program where the target data interface is located and performs integrity verification only on the sliced code segment corresponding to the target data interface rather than the complete program. The specific implementation manner will be described later.

[0086] In an embodiment of the present invention, an implementation manner of obtaining a multi-dimensional legal measurement value for integrity verification of the measurement result of the execution process of a target data interface through offline static program analysis is as follows: In the offline state, first, temporary instrumentation is performed on a first original target program to obtain a first target program. Then, static analysis is performed on the obtained first target program to obtain a static control flow graph. During the process of offline executing the target data interface in the first target program, the execution process of the target data interface is monitored by the trusted measurement engine of the data platform in the trusted execution environment to construct a corresponding runtime behavior model and obtain the execution result of executing the target data interface in the first target program. The runtime behavior model will record the data flow information during the execution process of the target data interface. Among them, the temporary instrumentation is used to locate the code segment corresponding to the target data interface in the complete program corresponding to the target data interface and only takes effect during the process of locating the code segment corresponding to the target data interface; the first original target program refers to the complete program corresponding to the target data interface executed by the data provider information system.

[0087] Compare and match the obtained execution result of executing the target data interface in the first target program with the recorded data flow information during the execution process of the target data interface, and the entry basic block and the exit basic block of the code segment corresponding to the target data interface can be determined.

[0088] In the obtained static control flow graph, starting from the obtained entry basic block, perform a depth-first search to determine all basic blocks that can be reached from this entry basic block. All paths composed of all basic blocks that can be reached from this entry basic block constitute the first search result. Starting from the obtained exit basic block, perform a depth-first search to determine all basic blocks that can reach this exit basic block. All paths composed of all basic blocks that can reach this exit basic block constitute the second search result. Compare the obtained first search result and the second search result, and the intersection path of the two search results is the code segment corresponding to the target data interface.

[0089] By slicing the code snippet corresponding to the target data interface in the determined first target program, an equivalent code snippet corresponding to the target data interface is constructed. Then, permanent instrumentation of the measurement logic is performed on the equivalent code snippet, and runtime measurement is carried out on the static code of the permanently instrumented equivalent code snippet and the execution process of the permanently instrumented equivalent code snippet, to obtain a multi-dimensional legal measurement value for integrity verification of the measurement result of the execution process of the target data interface. The multi-dimensional legal measurement value includes: the legal measurement value corresponding to the static code of the permanently instrumented equivalent code snippet, and the set of legal measurement values corresponding to the execution process of the permanently instrumented equivalent code snippet. Among them, the permanent instrumentation represents the permanent instrumentation of the measurement logic on the equivalent code, which will take effect every time it is executed and will not fail, aiming to interact with the trusted measurement engine at runtime to achieve runtime measurement.

[0090] After obtaining the multi-dimensional legal measurement value for integrity verification of the measurement result of the execution process of the target data interface, the multi-dimensional legal measurement value is stored in the distributed ledger for subsequent reuse in integrity verification of the target data access process. That is, as long as the distributed ledger stores the multi-dimensional legal measurement value for integrity verification of the measurement result of the execution process of the target data interface, when subsequently performing integrity verification on the target data access process, directly obtain the multi-dimensional legal measurement value stored in the distributed ledger for integrity verification of the measurement result of the execution process of the target data interface, without the need to generate the multi-dimensional legal measurement value for integrity verification of the measurement result of the execution process of the target data interface every time integrity verification is performed on the target data access process.

[0091] In an embodiment of the present invention, for the set of legal measurement values corresponding to the execution process of the equivalent code segment after permanent instrumentation, first, a static control flow graph of the target data interface is obtained by performing static analysis on the equivalent code segment after permanent instrumentation. It should be noted that since a unique ID has been assigned to each basic block during permanent instrumentation, and the sub - process markers representing the start and end of loops or recursive logic can also be obtained, the IDs of each basic block and all the start and end nodes of all sub - processes can be obtained when performing static analysis on the equivalent code segment after permanent instrumentation. Then, further static analysis is performed on the static control flow graph of the target data interface. The entire target data interface is regarded as the initial sub - process, and all sub - processes in the initial sub - process are accessed in a depth - first strategy to calculate the legal cumulative hash values for each sub - process. Whenever the cumulative hash value of a sub - process is calculated, all nodes of the sub - process except the start node, and the edges connected to all these sub - nodes, are deleted from the static control flow graph of the target data interface until only one node remains in the static control flow graph of the target data interface. In this way, a set of cumulative hash values representing the legal execution paths of each sub - process can be calculated. The set composed of the respective cumulative hash value sets of all sub - processes is the set of legal measurement values corresponding to the execution process of the equivalent code segment after permanent instrumentation. When performing online integrity verification on the measurement results of the execution process of the target data interface, the execution process measurement value of each sub - process in the measurement results of the execution process of the target data interface is only compared with the set of cumulative hash values corresponding to that sub - process.

[0092] In an embodiment of the present invention, a runtime behavior model describes the program behavior of an application over a period of time, including the execution behavior of the application and the data required for its execution, and is composed of a set of basic blocks, a set of data variable states, a relationship representing the execution order between two basic blocks, and a relationship representing the association between a basic block and the state of a data object.

[0093] RBM=(B, D, ε, σ)

[0094] Wherein, B={bb i} is the set of basic blocks, where bb i represents a basic block composed of a set of sequentially executed statements; D={(id, address, data) i} is the set of runtime states of data variables, where (id, address, data) i represents the number of variable i, the address of the variable in memory, and the recorded value of the variable; represents the execution order relationship between two basic blocks; represents the corresponding relationship between a basic block and the state of relevant data variables.

[0095] In the present invention, during the process that the data provider obtains target data through the target data interface, the runtime measurement of the execution process of the target data interface is performed by the trusted measurement engine of the data provider to obtain a measurement result, including: calling and executing an equivalent code segment with permanent instrumentation to obtain target data; during the execution process of the equivalent code segment with permanent instrumentation, the trusted measurement engine of the data provider measures and signs the static code, the execution process, and the execution result to obtain a measurement result.

[0096] In an embodiment of the present invention, during the process that the data provider obtains target data through the target data interface, an implementation manner of performing runtime measurement on the execution process of the target data interface by the trusted measurement engine of the data provider to obtain a measurement result is as follows: before the data provider calls and executes an equivalent code segment with permanent instrumentation, the data platform first generates a random number Nonce that has not been used to avoid replay attacks, and then sends the random number Nonce and the request for calling the target data interface to the data provider information system together by the data platform. Then, the data provider calls and executes the equivalent code segment with permanent instrumentation based on the request including the random number Nonce to obtain target data, and the equivalent code segment with permanent instrumentation corresponds to the code segment corresponding to the target data interface. During the execution process of the equivalent code segment with permanent instrumentation, the trusted measurement engine of the data provider performs runtime measurement on the static code, the execution process, and the execution result of the equivalent code segment with permanent instrumentation to respectively obtain the measurement values corresponding to the three, signs these measurement values and Nonce to construct a measurement result. Then, the measurement result and the obtained target data are returned to the data platform.

[0097] In an embodiment of the present invention, an implementation manner of performing runtime measurement and signing on the static code, the execution process, and the execution result of the equivalent code segment with permanent instrumentation by the trusted measurement engine of the data provider during the execution process of the equivalent code segment with permanent instrumentation to obtain a measurement result is as follows: when the target data interface starts to execute, the equivalent code segment with permanent instrumentation is executed to send a "Hello" message to the trusted measurement engine; after receiving the "Hello" message, the trusted measurement engine initializes the measurement status and measures the static code of the equivalent code segment with permanent instrumentation to form a static code measurement value.

[0098] Whenever the target data interface enters a new basic block, execute the equivalent code snippet with permanent instrumentation. Send the unique basic block ID corresponding to the current basic block to the trusted measurement engine through the measurement trigger, and then continue to execute the subsequent instructions; after receiving the basic block ID, the trusted measurement engine performs cumulative hashing calculations.

[0099] Before the execution of the target data interface ends, execute the equivalent code snippet with permanent instrumentation, send a "Goodbye" message, the execution result to be returned, and a one-time use random number Nonce for replay attack prevention to the trusted measurement engine; after receiving the "Goodbye" message, the trusted measurement engine ends the cumulative hashing calculation and forms a measurement value for the execution process, then performs a hashing calculation on the execution result to form an execution result measurement value; finally, the trusted measurement engine performs a hashing calculation on the static code measurement value, the execution process measurement value, the execution result measurement value, and the Nonce, and signs the calculation result to construct the measurement result.

[0100] In an embodiment of the present invention, a way to compare the measurement result with the multi-dimensional legitimate measurement value to obtain the target data access integrity verification result is as follows: use the public key of the trusted measurement engine of the data platform to decrypt the signature of the measurement result to obtain a hash value, and then use the Nonce of this request saved by the data platform, the static code measurement value, the execution process measurement value, and the execution result measurement value in the measurement result to perform a hashing calculation to obtain another hash value; if the one hash value is the same as the other hash value, it means that the measurement result itself is true and valid, otherwise it means that there is an error in the measurement result itself and subsequent verification is not necessary. In this way, it is possible to prevent attackers from using old measurement results for replay attacks or impersonating the measurement engine to provide forged measurement results and other situations.

[0101] Compare the static code measurement value in the measurement result with the corresponding pre-calculated legitimate static code measurement value to verify the integrity of the static code of the target data interface; if the two are the same, it means that the static code of the target data interface has not been tampered with, otherwise it means that it has been tampered with. In this way, it is possible to prevent attackers from changing its subsequent behavior and then tampering with data by tampering with its executable file before the target data interface is executed.

[0102] Compare the execution process measurement value in the measurement result with the corresponding set of legal measurement values pre-calculated for the execution process to verify the integrity of the data interface execution process; if the execution process measurement value is the same as the legal measurement value in the corresponding set of legal measurement values pre-calculated for the execution process, it means that the execution process of the target data interface has not been tampered with, otherwise it means that it has been tampered with. In this way, it is possible to avoid the situation where an attacker modifies control data or critical constraint data using software vulnerabilities to change its runtime behavior to achieve data tampering when the target data interface is executed.

[0103] Compare the target data received by the data platform with the execution result measurement value to verify the integrity of the data interface execution result; if the two are the same, it means that the execution result has not been tampered with during the process of the data provider information system sending the execution result to the data platform and receiving the execution result, otherwise it means that it has been tampered with. In this way, it is possible to avoid the situation where an attacker modifies its return result through methods such as HTTP hijacking to achieve data tampering after the target data interface is executed.

[0104] If all of the above verifications pass, then the integrity of the target data access process passes the verification; otherwise, the integrity verification of the data access process fails and the data may have been tampered with.

[0105] In the present invention, constructing an equivalent code snippet corresponding to the code snippet includes: constructing an initial equivalent code snippet decoupled from other business functions according to the code snippet of the target data interface; obtaining the equivalent code snippet by restoring the context environment of the initial equivalent code snippet.

[0106] In an embodiment of the present invention, one implementation manner of constructing an equivalent code snippet corresponding to the code snippet of the target data interface is: constructing an equivalent code snippet corresponding to the code snippet of the target data interface includes two steps. The first step is to construct an initial equivalent code snippet decoupled from other business functions based on the code snippet of the target data interface; the second step is to restore the context environment of the constructed initial equivalent code snippet to obtain the equivalent code snippet.

[0107] For the first step, first copy all the basic blocks in the code snippet corresponding to the target data interface and instrument them into the first target program. Then determine the set of functions where the code snippet corresponding to the target data interface is located, copy the set of functions as new functions prefixed with _cfa_, delete the redundant basic blocks that are not in the code snippet corresponding to the target data interface from the new functions, and insert the new functions into the class where the original functions are located. Since the copying of functions is only a copy of the functions themselves, the call relationships of the new functions still point to the functions corresponding to the original function call relationships, that is, the original functions and the copied functions still point to the same functions. Therefore, it is necessary to reconstruct the call relationships between the copied functions to avoid control flow errors when executing the copied code snippets and affecting the measurement results. To this end, according to the inter-function call relationships in the static function call graph of the code snippet corresponding to the target data interface, traverse all the instructions used to define function call relationships in the copied code snippet (that is, the code snippet obtained by copying all the basic blocks in the code snippet corresponding to the target data interface), find the instructions representing the original functions used to define function calls between the original functions, and replace the original functions in the instructions with the new functions obtained by copying the corresponding original functions. In addition, reconstruct the jump relationships between the basic blocks in the new functions. The specific reconstruction process is similar to the reconstruction of the inter-function call relationships, that is, replace the target addresses of the jump statements in the basic blocks in the new functions with reference to the original jump relationships. This will not be elaborated here. After completing the reconstruction of the function call relationships and jump relationships of the copied code snippet, the initial equivalent code snippet can be obtained. By instrumenting the initial equivalent code snippet for constructing the measurement results instead of the original code snippet corresponding to the target data interface, the problem of mutual interference between integrity measurement and other coupled business functions can be avoided.

[0108] For the second step, to ensure the executability of the initial equivalent code snippet, it is also necessary to correctly construct the context environment on which the execution of the code snippet depends.

[0109] A Data Dependence Graph (DDG) is a directed graph representing the data dependence relationships between program statements (or basic blocks). Data dependence represents the dependence of a statement (or basic block) that references a certain variable in a program on the statement that defines the variable, that is, a "definition-use" dependence relationship. If a and b are two nodes in the program control flow graph respectively, and v is a variable in the program, when nodes a and b meet the following conditions, it is said that node b directly depends on node a with respect to variable v. The following conditions include: node a defines variable v; variable v is referenced in node b; there is an executable path between node a and node b, and there is no other statement that defines variable v on this path.

[0110] By performing static analysis on the first target program, a data dependence graph DDG=(V, E) can be constructed, where V represents the set of nodes corresponding to all statements in the program, and E represents the set of data dependence relationships between statements.

[0111] Variables with data dependencies can be divided into three types: local variables within a function, non-static member variables of a class, and global variables (or static member variables of a class) (global variables do not exist in some programming languages, but the effect of global variables can be achieved in the form of static member variables). Among them, for non-static member variables and global variables of a class, since each basic block is copied to its corresponding class when constructing the initial equivalent code snippet, the access to these types of variables can still be carried out normally. For local variables within a function, since the declarations and assignments to them may be lost when copying the code snippet, it is necessary to determine the relevant missing instructions according to the data dependence graph and restore the missing instructions at the entrance of the corresponding basic block. In this way, the obtained initial equivalent code snippet can have a correct context environment during execution, thereby ensuring its executability, that is, after restoring the context environment of the initial equivalent code snippet, an equivalent code snippet can be obtained.

[0112] By performing program instrumentation for measurement on the equivalent code snippet obtained by slicing the above program instead of the code snippet corresponding to the target data interface, the overhead of measurement and verification can be effectively reduced, and interference with the remaining business functions can be avoided.

[0113] In the present invention, the data leakage detection of the received data algorithm to obtain a data leakage detection result includes: instrumenting the second original target program corresponding to the data algorithm to obtain a second target program; using the initial seed file as input and performing fuzz testing on the second target program to obtain an execution log; analyzing the execution log to obtain an execution log analysis result; according to the execution log analysis result, determining whether there is a new execution path during the current execution process; in the case where there is a new execution path during the current execution process, adding the seed file involved in the current execution process to the input queue and performing data leakage detection on the seed file involved in the current execution process to update the data leakage detection result; obtaining the latest seed file in the input queue for mutation, using the mutated seed file as input, performing fuzz testing on the second target program to obtain an execution log, and returning to the step: analyzing the execution log to obtain an execution log analysis result.

[0114] In an embodiment of the present invention, after the data platform obtains the target data, it will process the obtained target data through a data algorithm to obtain a corresponding processing result. And processing the obtained target data through a data algorithm needs to be completed by executing a second original target program corresponding to the data algorithm. Therefore, detecting data leakage of the received data algorithm by the data platform represents detecting data leakage of the second original target program corresponding to the received data algorithm by the data platform. Specifically, instrument the second original target program corresponding to the data algorithm received by the data platform to obtain a second target program. Instrumenting the second original target program includes: instrumenting at the entrance of each basic block for recording the current execution path; instrumenting at each jump statement for recording relevant branch variables; and instrumenting instructions for dynamic taint tracking of input statements.

[0115] Use the initial seed file as input data and execute the second target program through fuzz testing to obtain corresponding execution logs. Wherein, the initial seed file represents a simple test case that will not trigger data leakage defects in the program. Analyze the obtained execution logs to obtain a corresponding execution log analysis result. According to the execution log analysis result, it can be determined whether there is a new execution path in the current execution process. Since there has been no execution of the second target program before the first execution with the initial seed file as the data input, there is a new execution path in the first execution log analysis result with the initial seed file as the data input.

[0116] At this time, add the seed files involved in the current execution process to the input queue, and perform data leakage detection on the seed files involved in the current execution process to update the data leakage detection result. Since there has been no execution of the second target program before the first execution with the initial seed file as the data input, there is no data leakage detection result. Therefore, the first data leakage detection with the initial seed file as the data input will obtain the initial data leakage detection result.

[0117] Obtain the latest seed file in the input queue for mutation to stimulate new execution paths of the second target program, thereby improving the coverage rate of data leakage detection for the second target program. Use the mutated seed file as input data and execute the second target program through fuzz testing again to obtain corresponding execution logs. Analyze the obtained execution logs to obtain a corresponding execution log analysis result. According to the execution log analysis result, it can be determined whether there is a new execution path in the current execution process.

[0118] When there is a new execution path in the current execution process compared to the previous execution process, data leakage detection is performed using the seed file in the current execution process to determine whether there is a risk of data leakage in the new execution path, thereby improving the coverage rate of data leakage detection for the second target program. Specifically: when there is a new execution path in the current execution process, the seed file involved in the current execution process is added to the input queue, and data leakage detection is performed on the seed file involved in the current execution process to update the data leakage detection result obtained in the previous execution process, obtaining a new data leakage detection result.

[0119] Then continue to obtain the latest seed file in the input queue and continue to mutate it to stimulate new execution paths of the second target program, thereby improving the coverage rate of data leakage detection for the second target program. Continue to use the currently mutated seed file as input data and execute the second target program again through fuzz testing to obtain the corresponding execution log. Analyze the obtained execution log to obtain the corresponding execution log analysis result. According to the execution log analysis result, it can be determined whether there is a new execution path in the current execution process. Loop like this until there are no new execution paths in the current execution process.

[0120] In the present invention, data leakage detection is performed on the seed file involved in the current execution process to update the data leakage detection result, including: through a path-sensitive byte-level test sampling strategy, byte-level input-output sampling is performed on the seed file involved in the current execution process to obtain multiple sampling results; by performing entropy analysis on each of the multiple sampling results respectively, the entropy analysis result corresponding to each sampling result is obtained; according to the entropy analysis result corresponding to each sampling result, the data leakage detection result is updated.

[0121] In an embodiment of the present invention, the present invention directly uses the concept of information entropy to describe the uncertainty of target data and quantifies the degree of data leakage in the process of processing target data based on this.

[0122] For a mapping relationship F (the calculation process of the data algorithm) from the target data input X to the calculation result Y

[0123] Y = F(X)

[0124] The information entropy of the target data input X is:

[0125]

[0126] The conditional entropy of the target data input X when the calculation result Y is known is:

[0127]

[0128] The information gain brought by the calculation result Y to the target data input X is as follows:

[0129] g(X,Y) = H(X) - H(X|Y)

[0130] Generally speaking, the greater the information gain g(X,Y) provided by the calculation result Y to the target data input X, the more uncertainty of the target data input X is eliminated, and the greater the data leakage degree of the calculation process Y = F(X).

[0131] For example Figure 2 as shown Figure 2 shows the quantization results of the leakage degrees of three calculation processes of Y = 0, Y = X % 2, and Y = X * 2 based on information entropy. Assume that the value range of X is all integers from 1 to 6. Then, when the value of Y is not observed, X has 6 possible values; after observing the value of Y, X in Y = 0 still has 6 possible values, X in Y = X % 2 remains 3 possible values, and X in Y = X * 2 has only one possible value. Calculate the information gain values for the above three processes respectively, and the results are 0, 0.30, and 0.78 respectively. It can be seen that for a calculation process with a larger information gain value, it is easier for an attacker to guess the correct target data input by observing the calculation result, and its data leakage degree is greater. In this way, by calculating the information gain value of the calculation process of the data algorithm and setting the corresponding leakage threshold, it is possible to judge whether there are data leakage defects in the calculation process.

[0132] However, when the data algorithm only produces data leakage when the target data input meets very few special cases, the above-mentioned method for quantifying the data leakage degree based on information entropy may produce false negatives. Take the calculation process of Y = (0:1? X == 0) as shown Figure 3 as an example. When quantifying the leakage degree based on information entropy, its information gain value is only 0.14, which represents a very low data leakage degree. However, once the attacker observes that the output result is 1, it can be directly inferred that the target data input is 0, which indicates that there are still data leakage defects in the calculation process. We call this type of data leakage situation local leakage.

[0133] The generation of such false negatives is because the description of the uncertainty of random variables by information entropy is holistic, and its quantization process considers all possible values of the random variable. Therefore, the quantification of the data leakage degree based on information entropy only describes the overall leakage of the target calculation process to the confidential input, and cannot accurately detect such local leakage defects.

[0134] For such local leaks, the present invention further introduces minimum entropy to quantify data leaks. Different from information entropy, minimum entropy only focuses on the maximum case of the value-taking possibility of a random variable and calculates entropy based on this. In minimum entropy, the probability of the maximum possible value of the above-mentioned random variable is called its vulnerability, that is:

[0135] V(X) = max x∈X p(x)

[0136] where p(x) is the probability that the discrete random variable X takes the value x; V(X) is the vulnerability of the random variable X. Vulnerability represents the maximum possible probability that an attacker can guess the correct value of the random variable in just one guess. The larger the vulnerability of a random variable, the easier it is for the attacker to accurately guess its correct value. Based on the vulnerability of the target data input X, its minimum entropy can be calculated as follows:

[0137] H ∞ (X) = -logV(X)

[0138] where H ∞ (X) is the minimum entropy of the random variable X.

[0139] When the calculation result Y is known, the conditional minimum entropy of the target data input X is

[0140] H ∞ (X|Y) = min y∈Y H ∞ (X|Y = y)

[0141] where H ∞ (X|Y = y) is the minimum entropy of the random variable X when the random variable Y takes the value y; H ∞ (X|Y) is the conditional minimum entropy of the random variable X. The minimum entropy and conditional minimum entropy describe the worst-case uncertainty of the target data input X before and after the calculation result Y is observed. Their difference describes the data leakage degree of the calculation process F in the worst case, and this value is called the minimum entropy leakage, as follows:

[0142] L(X,Y) = H ∞ (X) - H ∞ (X|Y)

[0143] where L(X,Y) is the minimum entropy leakage, that is, the leakage degree of the random variable X (data input) due to the observation of the random variable Y (calculation result).

[0144] For Figure 3For the calculation process of Y = (0:1? X == 0) shown, although its information gain value is only 0.14, its minimum entropy leakage value is 1.00, indicating that there is leakage in the worst case. In the present invention, we will use both the information gain based on Shannon entropy and the minimum entropy leakage based on minimum entropy to measure the data leakage degree of the data algorithm. When the information gain value is large but the minimum entropy leakage value is small, it indicates that there is no locality leakage but serious global leakage in the calculation process of the data algorithm, such as Y = X % 2; while when the information gain value is small but the minimum entropy leakage value is large, it indicates that the global leakage degree of the data algorithm is low but there is likely to be locality leakage, such as Y = (0:1? X == 0).

[0145] Considering the complexity and diversity of the data algorithm, it is almost impossible to infer the above theoretical values of information gain or minimum entropy leakage by analyzing the entire second target program corresponding to the data algorithm. Therefore, the present invention calculates the approximate values of information gain and minimum entropy leakage by means of test sampling. Since the input of the data algorithm is usually an indefinite-length data set and its value space is infinite, that is, the value space of the random variable X representing the target data input in the information gain g(X, Y) is infinite. For the leakage detection technology based on test sampling, this state explosion of the input space will lead to insufficient coverage of the input space by test cases, thereby affecting the accuracy of the detection results. To this end, the present invention adopts a path-sensitive byte-level test sampling strategy to reduce the dimension of the input space, thereby avoiding state explosion.

[0146] In the embodiment of the present invention, the path-sensitive byte-level test sampling strategy is: the detection algorithm adopts a strategy of controlling variables byte by byte, mutates and samples each byte of the given input a certain number of times (samplingTimes), and records the relevant mutation content and the corresponding calculation results for each mutation sampling.

[0147] Independently calculate the information gain and minimum entropy leakage brought by the calculation results for each input byte according to the sampling results. Compare the information gain brought by the calculation result corresponding to each input byte with the information gain threshold. Determine the input node whose corresponding information gain brought by the calculation result is greater than the information gain threshold as the leaked byte. And, compare the minimum entropy leakage brought by the calculation result corresponding to each input byte with the minimum entropy leakage threshold. Determine the input node whose corresponding minimum entropy leakage brought by the calculation result is greater than the minimum entropy leakage threshold as the leaked byte. When the information gain brought by the calculation result corresponding to an input byte is greater than the information gain threshold, and / or the minimum entropy leakage brought by the calculation result corresponding to the input byte is greater than the minimum entropy leakage threshold, the input byte will be determined as the leaked byte. Further record the position range of the calculation result that causes leakage to the input byte. Preferably, the number of byte mutations in the present invention is 10 times, and this value can be specified by the user. The larger the value, the more sampling times, and the detection effect is usually more accurate, but the time overhead of detection will also be greater; the information gain threshold is log(samplingTimes / 3), which represents the information gain value when each unique output corresponds to three equally probable unique input bytes; the minimum entropy leakage threshold is 1, which represents the minimum entropy leakage value when there is an output that can uniquely determine a certain input byte.

[0148] Specifically, an implementation manner of performing data leakage detection on the seed file involved in the current execution process to update the data leakage detection result is: through a path-sensitive byte-granularity test sampling strategy, perform byte-granularity input and output sampling on the seed file involved in the current execution process to obtain multiple sampling results. Perform entropy analysis on each sampling result, thereby obtaining entropy analysis results respectively corresponding to the sampling results. Each entropy analysis result includes information gain and minimum entropy leakage. Update the data leakage detection result with the entropy analysis results respectively corresponding to the obtained sampling results to obtain the current data leakage detection result.

[0149] In the present invention, after obtaining the execution log analysis result by analyzing the execution log, the method further includes: determining whether there is an unexecuted branch in the current execution process according to the execution log analysis result; in the case that there is an unexecuted branch in the current execution process, expand the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label; guide the new taint label information to perform targeted input mutations one by one, and use the mutated seed file as the input to execute the second target program through fuzz testing to obtain an execution log, and return to the step: obtaining the execution log analysis result by analyzing the execution log.

[0150] In an embodiment of the present invention, for a given seed file, that is, a test case, the above-mentioned data leakage detection algorithm based on byte-level entropy analysis can effectively detect whether there is a data leakage defect in the second target program execution path corresponding to the test case. Since the data leakage detection algorithm based on byte-level entropy analysis adopts a dynamic analysis method based on testing, in order to comprehensively detect data leakage during the calculation process of the data algorithm using this data leakage detection algorithm, it is necessary to execute the test cases that can trigger different execution paths of the second target program corresponding to the data algorithm one by one as the input of the second target program.

[0151] In an embodiment of the present invention, the present invention obtains test cases that can trigger different execution paths of the second target program corresponding to the data algorithm through fuzz testing, and uses this test case as the input for the aforementioned data leakage detection. However, in the case where the fuzz testing process cannot effectively trigger the execution of the code segment with a data leakage defect, then the corresponding test case cannot be obtained, resulting in difficulty in detecting the data leakage defect of the unexecuted code segment here, which poses a high requirement for the code coverage rate of the fuzz testing process. Therefore, in order to efficiently mutate and generate test cases that can cover more program paths, the present invention introduces dynamic taint tracking on the basis of fuzz testing to specifically mutate the unexecuted branches, thereby improving the code coverage rate of the fuzz testing process.

[0152] Specifically, when there are unexecuted branches in the seed file involved in the current execution process, the taint label of the branch variable corresponding to the unexecuted branch is extended through dynamic information flow analysis to obtain a new taint label. Guided by this new taint label, targeted input mutation is performed, and the mutated seed file is used as the input data to execute the second target program again through fuzz testing to obtain the corresponding execution log. Then return to the step: analyze the execution log to obtain the execution log analysis result. Thus, the seed file can be more specifically mutated into a new seed file that can execute the unexecuted branch, making it more likely that the mutated seed file can execute the unexecuted branch, thereby improving the coverage rate of data leakage detection for the second target program.

[0153] In the present invention, the taint label of the branch variable corresponding to the unexecuted branch is extended through dynamic information flow analysis to obtain a new taint label, including: recording the execution path of the second target program during the current execution process to obtain relevant taint propagation information; during the process of re-executing the second target program with the seed file during the current execution process, passing the identifier of the unexecuted branch to the second target program to determine the unexecuted branch for which dynamic information flow analysis is to be performed; changing each byte of the seed file during the current execution process one by one and re-executing the second target program to obtain the value of the branch variable of the unexecuted branch when running under the changed seed file each time; in the case where the obtained runtime value of the branch variable of the unexecuted branch is valid and different from the initial value, marking the position corresponding to the currently changed byte as a label to be extended; determining the taint label of the branch variable corresponding to the unexecuted branch according to the relevant taint propagation information; and extending the taint label of the branch variable corresponding to the unexecuted branch through the label to be extended to obtain a new taint label.

[0154] In an embodiment of the present invention, an implementation manner of expanding the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label is as follows: record the execution path of the second target program during the current execution to obtain relevant taint propagation information. Then, for each unexecuted jump direction (i.e., unexecuted branch) recorded in the path record, obtain the taint label of the branch variable of the unexecuted branch and expand it. When expanding the taint label of the unexecuted branch, first re-execute the second target program with the same test case, and an identifier of the unexecuted branch will be additionally passed to the second target program during this execution process to determine the unexecuted branch for which dynamic information flow analysis will be performed, so as to obtain the value of the branch variable when running under this test case during execution. Then, change each byte in the test case one by one and re-execute the second target program to obtain the values of the branch variable when running under each new test case after the change. If the newly obtained runtime value is valid and different from the original runtime value, mark the position of the currently changed byte as the label to be expanded. Determine the taint label of the branch variable corresponding to the unexecuted branch according to the relevant taint propagation information. Finally, use the label to be expanded to expand the taint label of the unexecuted branch to obtain a new taint label. Generally speaking, the above process shields the calculation process between the test case input and the target branch variable (that is, the branch variable of the unexecuted branch), directly performs dynamic analysis on the correlation between the two, so as to solve the problem that taint information is easily lost under complex data processing logic during the fuzz testing of the second target program, more effectively guide the mutation process of subsequent test cases, and improve the coverage ability of the fuzz testing for the execution path of the second target program corresponding to the data algorithm, that is, to make as many execution paths in the second target program be executed as possible, so as to obtain more accurate data leakage results.

[0155] The second aspect of the present invention provides a data trusted verification device in a data element scenario. The device 400 includes:

[0156] A trusted measurement engine 401, configured to perform runtime measurement on the execution process of the target data interface during the process of the data provider obtaining target data through the target data interface, and obtain a measurement result;

[0157] A data acquisition module 402, configured to acquire multi-dimensional legal measurement values stored in the distributed ledger and corresponding to the target data interface;

[0158] A data integrity validator 403, configured to compare the measurement result with the multi-dimensional legal measurement values to obtain a target data access integrity verification result;

[0159] A data leakage measurer 404 for detecting data leakage of the received data algorithm and obtaining a data leakage detection result;

[0160] An algorithm executor 405 for processing the target data through the data algorithm in a privacy computing environment to obtain a processing result when both the target data access integrity verification result and the data leakage detection result are qualified;

[0161] A processing result return module 406 for sending the processing result to the data requester.

[0162] Optionally, the device 400 further includes:

[0163] A temporary instrumentation module for obtaining a first target program by performing temporary instrumentation on a first original target program in an offline state and obtaining a static control flow graph by performing static analysis on the first target program;

[0164] A runtime behavior model construction module for monitoring the execution process of the target data interface through the trusted measurement engine of the data platform in a trusted execution environment during the execution of the target data interface in the first target program to construct a corresponding runtime behavior model and obtaining the execution result of the target data interface;

[0165] An entrance and exit basic block determination module for matching the execution result with the data flow information in the runtime behavior model to determine the entrance basic block and the exit basic block of the target data interface;

[0166] A depth-first search module for, in the static control flow graph, starting from the entrance basic block, determining all basic blocks that can be reached from the entrance basic block through depth-first search to obtain a first search result, and starting from the exit basic block, determining all basic blocks that can reach the exit basic block through depth-first search to obtain a second search result;

[0167] A code snippet determination module for determining the code snippet corresponding to the target data interface according to the first search result and the second search result;

[0168] An equivalent code snippet determination module for constructing an equivalent code snippet corresponding to the code snippet by slicing the code snippet;

[0169] A multi-dimensional legal measurement value determination module for performing permanent instrumentation on the equivalent code snippet and performing runtime measurement on the static code and execution process of the permanently instrumented equivalent code snippet to obtain multi-dimensional legal measurement values;

[0170] A multi-dimensional legal measurement value storage module for storing the multi-dimensional legal measurement value into a distributed ledger.

[0171] Optionally, the trusted measurement engine 401 includes:

[0172] An equivalent code snippet execution module for calling and executing a permanently instrumented equivalent code snippet to obtain target data;

[0173] A trusted measurement engine sub-module for measuring and signing static code, the execution process, and the execution result during the execution of the permanently instrumented equivalent code snippet to obtain a measurement result.

[0174] Optionally, the data leakage measurer 404 includes:

[0175] A program instrumenter for instrumenting a second original target program corresponding to the data algorithm to obtain a second target program;

[0176] A fuzz testing module for taking an initial seed file as input and executing the second target program through fuzz testing to obtain an execution log;

[0177] An execution log analysis module for analyzing the execution log to obtain an execution log analysis result;

[0178] A new execution path determination module for determining whether there is a new execution path during the current execution process according to the execution log analysis result;

[0179] A data leakage detection module for adding the seed file involved in the current execution process to an input queue and performing data leakage detection on the seed file involved in the current execution process to update the data leakage detection result when there is a new execution path during the current execution process;

[0180] The fuzz testing module is further configured to obtain the latest seed file in the input queue for mutation, and take the mutated seed file as input to execute the second target program through fuzz testing to obtain an execution log, and call the execution log analysis module for analyzing the execution log to obtain an execution log analysis result.

[0181] Optionally, the data leakage detection module includes:

[0182] A sampling module for performing byte-level input / output sampling on the seed file involved in the current execution process through a path-sensitive byte-level test sampling strategy to obtain a plurality of sampling results;

[0183] An entropy analysis module, configured to obtain entropy analysis results corresponding to respective sampling results by performing entropy analysis on the plurality of sampling results respectively;

[0184] A data leakage detection result update module, configured to update the data leakage detection result according to the entropy analysis results corresponding to the respective sampling results.

[0185] Optionally, the apparatus 400 further includes:

[0186] An unexecuted branch determination module, configured to determine whether there is an unexecuted branch in the current execution process according to the execution log analysis result;

[0187] A taint label extension module, configured to, when there is an unexecuted branch in the current execution process, extend the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label;

[0188] A mutation module, configured to guide the new taint label information to perform targeted input mutations one by one, and use the mutated seed file as an input to execute the second target program through fuzz testing to obtain an execution log, and call an execution log analysis module, configured to analyze the execution log to obtain an execution log analysis result.

[0189] Optionally, the taint label extension module includes:

[0190] A taint propagation information determination module, configured to record the execution path of the second target program in the current execution process to obtain relevant taint propagation information;

[0191] A dynamic information flow analysis object determination module, configured to, in the process of re-executing the second target program through the seed file in the current execution process, pass the identifier of the unexecuted branch to the second target program to determine the unexecuted branch for which dynamic information flow analysis is to be performed;

[0192] A value determination module for the seed file during runtime, configured to change each byte of the seed file in the current execution process one by one and re-execute the second target program to obtain the value of the branch variable of the unexecuted branch when running under the changed seed file each time;

[0193] A to-be-extended label determination module, configured to, when the obtained value of the branch variable of the unexecuted branch is valid and different from the initial value, mark the position corresponding to the currently changed byte as a to-be-extended label;

[0194] A taint label determination module, configured to determine the taint label of the branch variable corresponding to the unexecuted branch according to the relevant taint propagation information;

[0195] The stain label extension sub-module is used to extend the stain label of the branch variable corresponding to the unexecuted branch through the to-be-extended label to obtain a new stain label.

[0196] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0197] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0198] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0199] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.

Claims

1. A data credibility verification method in the scenario of data elements, characterized in that, Applied to a data platform, the method includes: During the process of the data provider obtaining target data through a target data interface, the trusted measurement engine of the data provider performs runtime measurement on the execution process of the target data interface to obtain a measurement result; Obtain the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger; Compare the measurement result with the multi-dimensional legal measurement values to obtain a target data access integrity verification result; Perform data leakage detection on the received data algorithm to obtain a data leakage detection result; When both the target data access integrity verification result and the data leakage detection result are qualified, process the target data through the data algorithm in a privacy computing environment to obtain a processing result; Send the processing result to the data requirer; Among them, during the process of the data provider obtaining target data through a target data interface, the trusted measurement engine of the data provider performs runtime measurement on the execution process of the target data interface to obtain a measurement result, including: calling and executing an equivalent code segment with permanent instrumentation to obtain target data; during the execution process of the equivalent code segment with permanent instrumentation, the trusted measurement engine of the data provider measures and signs the static code, execution process, and execution result to obtain a measurement result; Among them, performing data leakage detection on the received data algorithm to obtain a data leakage detection result includes: instrumenting the second original target program corresponding to the data algorithm to obtain a second target program; using the initial seed file as input, executing the second target program through fuzz testing to obtain an execution log; analyzing the execution log to obtain an execution log analysis result; according to the execution log analysis result, determine whether there is a new execution path during the current execution process; when there is a new execution path during the current execution process, add the seed file involved in the current execution process to the input queue, and perform data leakage detection on the seed file involved in the current execution process to update the data leakage detection result; obtain the latest seed file in the input queue for mutation, and use the mutated seed file as input to execute the second target program through fuzz testing to obtain an execution log, and return to the step: analyzing the execution log to obtain an execution log analysis result.

2. The data credibility verification method in the data element scenario according to claim 1, wherein, Before obtaining the multi-dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger, the method further includes: In an offline state, perform temporary instrumentation on the first original target program to obtain a first target program, and obtain a static control flow graph through static analysis of the first target program; During the execution process of the target data interface in the first target program, the trusted measurement engine of the data platform in the trusted execution environment monitors the execution process of the target data interface to construct a corresponding runtime behavior model, and obtains the execution result of the target data interface; Match the execution result with the data flow information in the runtime behavior model to determine the entry basic block and the exit basic block of the target data interface; In the static control flow graph, starting from the entry basic block, determine all basic blocks reachable from the entry basic block through depth-first search to obtain a first search result, and, starting from the exit basic block, determine all basic blocks reachable to the exit basic block through depth-first search to obtain a second search result; Determine the code snippet corresponding to the target data interface according to the first search result and the second search result; Construct an equivalent code snippet corresponding to the code snippet by slicing the code snippet; Perform permanent instrumentation on the equivalent code snippet, and perform runtime measurement on the static code and execution process of the permanently instrumented equivalent code snippet to obtain multi-dimensional legal measurement values; Store the multi-dimensional legal measurement values in a distributed ledger.

3. The data trust verification method in the data element scenario according to claim 2, characterized in that Constructing an equivalent code snippet corresponding to the code snippet includes: Construct an initial equivalent code snippet decoupled from other business functions according to the code snippet of the target data interface; Obtain the equivalent code snippet by restoring the context environment of the initial equivalent code snippet.

4. The data trust verification method in the data element scenario according to claim 1, characterized in that, Perform data leakage detection on the seed files involved in the current execution process to update the data leakage detection result, including: Perform byte-level input / output sampling on the seed files involved in the current execution process through a path-sensitive byte-level test sampling strategy to obtain multiple sampling results; Obtain the entropy analysis result corresponding to each sampling result by performing entropy analysis on each of the multiple sampling results; Update the data leakage detection result according to the entropy analysis result corresponding to each sampling result.

5. The data trust verification method in the data element scenario according to claim 1, characterized in that After obtaining the execution log analysis result by analyzing the execution log, the method further includes: Determine whether there are unexecuted branches in the current execution process according to the execution log analysis result; In the case where there are unexecuted branches in the current execution process, expand the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label; Guide the new taint label information to perform targeted input mutations one by one, and use the mutated seed files as inputs to execute the second target program through fuzz testing to obtain an execution log, and return to the step: obtain the execution log analysis result by analyzing the execution log.

6. The data trust verification method in the data element scenario according to claim 5, characterized in that Expanding the taint label of the branch variable corresponding to the unexecuted branch through dynamic information flow analysis to obtain a new taint label includes: Record the execution path of the second target program in the current execution process to obtain relevant taint propagation information; During the process of executing the second target program again through the seed files in the current execution process, pass the identifier of the unexecuted branch to the second target program to determine the unexecuted branch for dynamic information flow analysis; Change each byte of the seed file in the current execution process one by one and re - execute the second target program to obtain the value of the branch variable of the unexecuted branch when running under the changed seed file each time; When the obtained value of the branch variable of the unexecuted branch during runtime is valid and different from the initial value, mark the position corresponding to the currently changed byte as a tag to be extended; Determine the taint tag of the branch variable corresponding to the unexecuted branch according to the relevant taint propagation information; Use the tag to be extended to extend the taint tag of the branch variable corresponding to the unexecuted branch to obtain a new taint tag.

7. A data trust verification device in a data element scenario, characterized in that, The device includes: A trusted measurement engine, used to perform runtime measurement on the execution process of the target data interface during the process of the data provider obtaining target data through the target data interface to obtain a measurement result; A data acquisition module, used to acquire the multi - dimensional legal measurement values corresponding to the target data interface stored in the distributed ledger; A data integrity verifier, used to compare the measurement result with the multi - dimensional legal measurement values to obtain a target data access integrity verification result; A data leakage measurer, used to perform data leakage detection on the received data algorithm to obtain a data leakage detection result; An algorithm executor, used to process the target data through the data algorithm in a privacy - computing environment to obtain a processing result when both the target data access integrity verification result and the data leakage detection result are qualified; A processing result return module, used to send the processing result to the data demander; Among them, the trusted measurement engine includes: An equivalent code snippet execution module, used to call and execute the permanently instrumented equivalent code snippet to obtain target data; A trusted measurement engine sub - module, used to measure and sign the static code, execution process, and execution result during the execution process of the permanently instrumented equivalent code snippet to obtain a measurement result; Among them, the data leakage measurer includes: A program instrumenter, used to instrument the second original target program corresponding to the data algorithm to obtain a second target program; A fuzz testing module, used to use the initial seed file as input and execute the second target program through fuzz testing to obtain an execution log; An execution log analysis module, used to analyze the execution log to obtain an execution log analysis result; A new execution path determination module, used to determine whether there is a new execution path in the current execution process according to the execution log analysis result; A data leakage detection module, used to add the seed file involved in the current execution process to the input queue and perform data leakage detection on the seed file involved in the current execution process to update the data leakage detection result when there is a new execution path in the current execution process; The fuzz testing module is further configured to obtain the latest seed file in the input queue for mutation, and use the mutated seed file as input to execute the second target program through fuzz testing to obtain an execution log, and call an execution log analysis module to analyze the execution log to obtain an execution log analysis result.

Citation Information

Patent Citations

  • Dynamic evaluation method and system for information sensitivity of smart power grid

    CN114036570A

  • Remote proof-based interface type digital object authenticity verification method and device

    CN115051810A