Credible data space cross-domain collaborative analysis method based on privacy calculation
Through the collaborative data analysis across institutions and industries, a trusted data space based on privacy computing is built, and distributed edge computing and blockchain technology is adopted to solve the problems of data resource integration and utilization and privacy compliance, and achieve secure and efficient collaborative analysis of data.
Patent Information
- Application Number
- CN202510740348.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, cross-institutional and cross-industry data collaborative analysis has the problem that data resources are difficult to integrate and utilize and cannot meet privacy compliance requirements, especially in the process of data circulation, the risk of being vulnerable to man-in-the-middle attacks and internal personnel's illegal disclosure of data.
The cross-domain collaborative analysis method of trusted data space based on privacy computing is adopted. By deploying edge computing nodes on participants, building a distributed trusted data space, configuring a policy engine to define data permissions and privacy protection rules, selecting appropriate privacy computing protocols, performing data preprocessing and encryption processing, and realizing identity authentication and security monitoring through blockchain proof storage and smart contracts to ensure that data is synergistically analyzed in ciphertext form.
It realizes the full use of data resources to prevent malicious tampering while meeting privacy compliance requirements, and improves the security and efficiency of data collaborative analysis.
Smart Images

Figure CN120263387A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and specifically to a cross-domain collaborative analysis method for trusted data space based on privacy computing. Background Art
[0002] In the digital age, the demand for the circulation of data elements is continuing to rise, and cross-institutional and cross-industry data collaborative analysis has become the core way to fully explore and enhance the value of data. However, current technologies face many challenges in practical applications.
[0003] On the one hand, different institutions cannot directly exchange raw data due to privacy protection and compliance considerations. This makes a large number of valuable data resources like "information islands", which are difficult to effectively integrate and fully utilize, hindering the full release of data value.
[0004] On the other hand, traditional data collaboration methods have obvious drawbacks. For example, data transmission in plain text and centralized aggregation are extremely vulnerable to middleman attacks during data circulation. There is also the risk of internal personnel illegally leaking data, which cannot meet the current strict privacy compliance requirements.
[0005] Therefore, this application proposes a cross-domain collaborative analysis method in a trusted data space based on privacy computing. Summary of the invention
[0006] To this end, this application provides a cross-domain collaborative analysis method for a trusted data space based on privacy computing to solve the problem that data resources in the existing technology are difficult to integrate and utilize and cannot meet privacy compliance requirements.
[0007] In order to achieve the above objectives, this application provides the following technical solutions:
[0008] The cross-domain collaborative analysis method of trusted data space based on privacy computing includes the following steps:
[0009] Step 1: Building a trusted data space. Deploy edge computing nodes on multiple data owners participating in cross-domain collaborative analysis. Each edge computing node is interconnected through a blockchain network to form a distributed trusted data space. A policy engine is configured in the trusted data space to define data usage permissions, privacy protection levels, and computing task lifecycle management rules.
[0010] Step 2: Dynamically select a privacy computing protocol. Select at least one privacy computing protocol from a preconfigured protocol library based on the data type, privacy protection requirements, and computational complexity of the collaborative analysis task. The protocol library includes a combination of homomorphic encryption, federated learning, secure multi-party computing, and differential privacy algorithms.
[0011] Step 3, Cross-domain Identity Authentication and Authorization. The participating parties submit digital identity credentials through the blockchain network, and the smart contract verifies their legitimacy. The participating parties that pass the verification obtain dynamic authorization tokens, which contain the accessible data scope, operation type, and validity period.
[0012] Step 4, Data Preprocessing and Privacy Protection. The data owner performs desensitization, encryption, or sharding on the original data locally to generate input data that meets the requirements of the privacy computing protocol. The preprocessing process is recorded on the blockchain for evidence.
[0013] Step 5, Collaborative Analysis Task Execution. The edge nodes of each participating party perform distributed computing based on the selected privacy computing protocol. During the process, the intermediate results are transmitted through an encrypted channel, and the result aggregation is completed within the trusted data space.
[0014] Step 6, Dynamic Security Monitoring. During the task execution, the data stream, computing node status, and communication link are monitored in real time. If abnormal behaviors or privacy leakage risks are detected, protocol switching or task termination is triggered.
[0015] Step 7, Result Output and Traceability. The final analysis result is output after being jointly decrypted by the data owner. At the same time, an audit report containing the full process log of the task is generated, and the report is anchored to the blockchain through a hash value.
[0016] Preferably, the construction of the trusted data space in Step 1 includes
[0017] Configuring a hardware security module for each edge computing node, which is used for key management and secure storage.
[0018] Defining data classification labels through a policy engine, which include public data, restricted data, and confidential data.
[0019] Implementing distributed storage and fast retrieval of cross-domain data using the IPFS protocol.
[0020] Preferably, the rules for dynamically selecting the privacy computing protocol in Step 2 include:
[0021] Preferably select the homomorphic encryption protocol for structured data analysis tasks;
[0022] Select the federated learning protocol for unstructured data or tasks that require model training;
[0023] When the number of participating parties exceeds a preset threshold, automatically switch to the lightweight secure multi-party computing protocol.
[0024] Preferably, zero-knowledge proof technology is used for cross-domain identity authentication in Step 3, and the participating parties complete the verification without revealing identity details.
[0025] Preferably, the data preprocessing in step four includes:
[0026] Performing k-anonymization on sensitive fields;
[0027] Implementing differential privacy protection by adding Laplace noise;
[0028] Using threshold encryption technology to fragment and distribute data to multiple participating parties.
[0029] Preferably, the execution of the collaborative analysis task in step five includes:
[0030] Splitting the computing task into multiple subtasks for parallel execution;
[0031] Adopting a compression algorithm to reduce the amount of data transmitted for intermediate results;
[0032] Using a redundancy check mechanism in the aggregation stage to ensure the integrity of the results.
[0033] Preferably, the indicators of the dynamic security monitoring in step six include:
[0034] Monitoring abnormal fluctuations in the CPU / memory occupancy rate of the monitoring node;
[0035] Monitoring whether the communication link delay exceeds the threshold;
[0036] Monitoring the deviation degree of the intermediate result from the expected distribution.
[0037] Preferably, the auditable report in step seven includes the following data:
[0038] Data for evaluating the data contribution degree of each participating party;
[0039] Privacy protection algorithm parameters and data on execution effects;
[0040] Data on task execution time and resource consumption statistics.
[0041] Preferably, it further includes deploying a security gateway between the trusted data space and the external system, and the security gateway provides protocol conversion, traffic filtering and intrusion prevention functions.
[0042] Compared with the prior art, the present application has at least the following beneficial effects:
[0043] 1. The data participates in the calculation in ciphertext or desensitized form throughout the process, and the original data does not leave the domain. On the premise of meeting the privacy compliance requirements, the data can also be fully utilized;
[0044] 2. By using blockchain for evidence storage and smart contracts to automatically execute data usage rules, malicious node tampering is prevented. Description of the Drawings
[0045] To more intuitively illustrate the prior art and this application, exemplary drawings are given below. It should be understood that the specific shapes and structures shown in the drawings generally should not be regarded as limiting conditions when implementing this application; for example, those skilled in the art are capable of making routine adjustments or further optimizations to the addition / deletion / attribution division of certain units (components), specific shapes, positional relationships, connection methods, dimensional proportional relationships, etc. based on the technical concept disclosed in this application and the exemplary drawings.
[0046] Figure 1 This is a flowchart of the cross-domain collaborative analysis method for a trusted data space based on privacy computing in this application. Detailed implementation manners
[0047] The following further details this application through specific embodiments in conjunction with the drawings.
[0048] As Figure 1 shown, the cross-domain collaborative analysis method for a trusted data space based on privacy computing is characterized by including the following steps
[0049] Step 1, construction of a trusted data space: Deploy edge computing nodes at multiple data owners participating in cross-domain collaborative analysis, and each node is interconnected through a blockchain network to form a distributed trusted data space; a policy engine is configured in the trusted data space to define data usage permissions, privacy protection levels, and calculation task lifecycle management rules, and a decentralized data collaboration environment is constructed through the interconnection of distributed edge nodes and the blockchain network. Each node deploys a policy engine to define data usage rules (such as access permissions, privacy levels) to ensure that data only circulates within the authorized scope. The blockchain serves as a trust anchor to record the full process log of tasks to prevent data tampering or illegal operations;
[0050] Step 2, dynamic selection of privacy computing protocols: Select at least one privacy computing protocol from a pre-configured protocol library according to the data type, privacy protection requirements, and calculation complexity of the collaborative analysis task; the protocol library includes combinations of homomorphic encryption, federated learning, secure multi-party computation, and differential privacy algorithms, and the optimal privacy computing protocol is dynamically selected according to task requirements (such as data type, privacy level). For example, structured data uses homomorphic encryption for direct ciphertext calculation, and unstructured data uses federated learning for distributed model training to balance security and computing efficiency;
[0051] Step 3, Cross - domain Identity Authentication and Authorization: The participating parties submit digital identity credentials through the blockchain network, and the smart contract verifies their legitimacy; The participating parties that pass the verification obtain dynamic authorization tokens, which contain the accessible data scope, operation type, and validity period. The dynamic authentication and authorization of the participating parties' identities are realized through the blockchain smart contract. Combined with the Hardware Security Module (HSM) to manage keys, it ensures the credibility of the participating parties' identities and the traceability of data operations;
[0052] Step 4, Data Pre - processing and Privacy Protection: The data owner performs desensitization, encryption, or sharding on the original data locally to generate input data that meets the requirements of the privacy computing protocol; The pre - processing process is recorded on the blockchain for evidence preservation. Technologies such as desensitization and sharding are adopted in the pre - processing stage to ensure that the original data is not exposed;
[0053] Step 5, Collaborative Analysis Task Execution: The edge nodes of each participating party perform distributed computing based on the selected privacy computing protocol. During the process, the intermediate results are transmitted through an encrypted channel, and the result aggregation is completed within the trusted data space. When the task is executed, the calculation is split into subtasks for parallel processing, and the communication overhead is reduced by compressing the intermediate results;
[0054] Step 6, Dynamic Security Monitoring: During the task execution, the data stream, computing node status, and communication link are monitored in real - time. If abnormal behaviors or privacy leakage risks are detected, protocol switching or task termination is triggered. The node status and data stream are monitored in real - time, and the protocol is dynamically adjusted or abnormal tasks are terminated to ensure the security and controllability of the analysis process;
[0055] Step 7, Result Output and Traceability: The final analysis result is output after being jointly decrypted by the data owners. At the same time, an audit report containing the full - process log of the task is generated, and the report is anchored to the blockchain through a hash value.
[0056] The construction of the trusted data space described in Step 1 includes,
[0057] Configure a Hardware Security Module (HSM) for each edge computing node for key management and secure storage;
[0058] Define data classification labels through a policy engine, and the labels include public data, restricted data, and confidential data;
[0059] Adopt the IPFS protocol to realize the distributed storage and fast retrieval of cross - domain data.
[0060] The HSM manages keys through a physically isolated encryption chip to prevent key leakage or tampering; data classification tags (public / restricted / confidential) enable fine-grained permission control to ensure that data at different sensitivity levels is used in isolation according to rules; the IPFS protocol distributes and stores data hash values, solving the problem of cross-domain retrieval efficiency and avoiding single-point failures at the same time. This solution enhances the security of data storage and access and improves the feasibility of large-scale data collaboration.
[0061] The rules for dynamically selecting the privacy computing protocol described in step 2 include:
[0062] For structured data analysis tasks, the homomorphic encryption protocol is preferentially selected;
[0063] For unstructured data or tasks requiring model training, the federated learning protocol is selected;
[0064] When the number of participants exceeds a preset threshold, it is automatically switched to the lightweight secure multi-party computing protocol.
[0065] For structured data (such as database tables), homomorphic encryption is selected because it is suitable for numerical calculations and has high ciphertext operation efficiency; for unstructured data (such as images and texts), federated learning is adopted to avoid the exchange of raw data through distributed model training; when the number of participants exceeds the threshold (such as 10), it is switched to lightweight secure multi-party computing (SMPC) to reduce communication complexity. This rule realizes automated decision-making through a pre-configured policy library, optimizing computing efficiency and resource consumption.
[0066] The cross-domain identity authentication described in step 3 adopts zero-knowledge proof technology. The participants complete the verification without revealing the identity details. The participants submit the encrypted hash value of the identity credentials to the blockchain and prove its legality through the ZKP protocol (such as zk-SNARK) without disclosing the specific identity information. For example, a medical institution can prove that it has compliance qualifications without disclosing the business license number. This technology prevents the leakage of identity information and meets the needs of anonymized auditing at the same time, and is applicable to scenarios with extremely high privacy requirements.
[0067] The data preprocessing described in step 4 includes:
[0068] Perform k-anonymization processing on sensitive fields;
[0069] Implement differential privacy protection by adding Laplace noise;
[0070] Use threshold encryption technology to fragment and distribute data to multiple participants.
[0071] k-anonymization generalizes sensitive fields (such as ID numbers) into group attributes (such as age groups), making individual records unable to be uniquely identified; differential privacy confuses statistical results by adding noise (such as Laplace distribution) to prevent reverse inference; threshold encryption splits data into multiple shards and can only be decrypted when a sufficient number of participating parties cooperate. The combination of the three achieves multi-level privacy protection, covering both the static storage and dynamic calculation phases of data.
[0072] The collaborative analysis task execution described in step five further includes:
[0073] Splitting the computing task into multiple subtasks for parallel execution;
[0074] Using a compression algorithm to reduce the amount of data transmitted for intermediate results;
[0075] Using a redundancy check mechanism in the aggregation phase to ensure result integrity.
[0076] Split the task into subtasks for parallel processing (such as the MapReduce framework), leveraging the computing power of edge nodes to accelerate calculations; use a lossless compression algorithm (such as Huffman coding) to reduce the amount of intermediate result transmission; in the aggregation phase, verify data integrity through redundancy checks (such as CRC checksum) to prevent transmission errors or malicious tampering.
[0077] The metrics for the dynamic security monitoring described in step six include:
[0078] Abnormal fluctuations in the CPU / memory occupancy rate of nodes;
[0079] Communication link latency exceeding the threshold;
[0080] The deviation degree of intermediate results from the expected distribution.
[0081] Abnormal node resources (such as CPU occupancy continuously exceeding 90%) may indicate malicious mining or DDoS attacks; a sudden increase in communication latency (such as exceeding 500 ms) may indicate network eavesdropping or node failure; intermediate results deviating from the expected distribution (such as statistical feature differences exceeding 10%) may imply data leakage or algorithm tampering. The monitoring system automatically triggers alarms or intervention measures based on thresholds to enhance the system's robustness.
[0082] The auditable report in step seven includes the following data:
[0083] Data for evaluating the data contribution of each participating party;
[0084] Privacy protection algorithm parameters and data on execution effects;
[0085] Data on task execution time and resource consumption statistics.
[0086] The data contribution degree quantifies the marginal impact of each party's data on the result through the Shapley value algorithm to ensure collaboration fairness; the privacy protection algorithm parameters record the algorithm strength to provide third-party verification of compliance; the resource consumption statistics help optimize the task scheduling strategy. This report supports post-event traceability and responsibility definition, enhancing the transparency of cross-domain collaboration.
[0087] It also includes deploying a security gateway between the trusted data space and external systems. The security gateway provides protocol conversion, traffic filtering, and intrusion prevention functions, providing protocol conversion (such as HTTP to gRPC), traffic filtering (intercepting malicious requests based on a rule engine), and intrusion prevention (such as AI-based abnormal traffic detection). As a security boundary, the gateway prevents external attacks from penetrating into the internal data space and is compatible with the enterprise's existing IT infrastructure, reducing the implementation cost of this method.
[0088] The technical features of the above embodiments can be combined arbitrarily (as long as there is no contradiction in the combination of these technical features). For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described; these embodiments that are not explicitly written should also be considered as within the scope described in this specification.
Claims
1. A cross - domain collaborative analysis method for a trusted data space based on privacy computing, characterized in that, It includes the following steps: Step 1, Construction of a trusted data space: Deploy edge computing nodes at multiple data owners participating in cross-domain collaborative analysis. Each edge computing node is interconnected through a blockchain network to form a distributed trusted data space. A policy engine is configured within the trusted data space, and the policy engine is used to define data usage permissions, privacy protection levels, and calculation task lifecycle management rules; Step 2, Dynamic selection of privacy computing protocols: Select at least one privacy computing protocol from a pre-configured protocol library according to the data type, privacy protection requirements, and calculation complexity of the collaborative analysis task. The protocol library includes combinations of homomorphic encryption, federated learning, secure multi-party computing, and differential privacy algorithms; Step 3, Cross-domain identity authentication and authorization: The participating parties submit digital identity credentials through the blockchain network, and the smart contract verifies their legitimacy; The participating parties that pass the verification obtain a dynamic authorization token, and the token contains the accessible data range, operation type, and validity period; Step 4, Data preprocessing and privacy protection: The data owner locally performs desensitization, encryption, or sharding processing on the original data to generate input data that meets the requirements of the privacy computing protocol. The preprocessing process is recorded on the blockchain for evidence; Step 5, Execution of collaborative analysis tasks: The edge nodes of each participating party perform distributed computing based on the selected privacy computing protocol. During the process, intermediate results are transmitted through an encrypted channel, and result aggregation is completed within the trusted data space; Step 6, Dynamic security monitoring: During the task execution, the data stream, computing node status, and communication link are monitored in real time. If abnormal behavior or privacy leakage risks are detected, protocol switching or task termination is triggered; Step 7, Result output and traceability: The final analysis result is output after being jointly decrypted by the data owners. At the same time, an audit report containing the full process log of the task is generated, and the report is anchored to the blockchain through a hash value.
2. The cross-domain collaborative analysis method for a trusted data space based on privacy computing according to claim 1, wherein The construction of the trusted data space in Step 1 includes, Configure a hardware security module for each edge computing node, and the hardware security module is used for key management and secure storage; Define data classification labels through the policy engine, and the labels include public data, restricted data, and confidential data; Use the IPFS protocol to achieve distributed storage and fast retrieval of cross-domain data.
3. The cross-domain collaborative analysis method for a trusted data space based on privacy computing according to claim 1, wherein The rules for dynamic selection of privacy computing protocols in Step 2 include: Preferentially select the homomorphic encryption protocol for structured data analysis tasks; Select the federated learning protocol for unstructured data or tasks that require model training; When the number of participating parties exceeds a preset threshold, automatically switch to a lightweight secure multi-party computing protocol.
4. The cross - domain collaborative analysis method for a trusted data space based on privacy computing according to claim 1, wherein The cross-domain identity authentication in Step 3 adopts zero-knowledge proof technology, and the participating parties complete the verification without revealing identity details.
5. The cross - domain collaborative analysis method of the trusted data space based on privacy computing according to claim 1, wherein, The data preprocessing in Step 4 includes: Perform k-anonymization processing on sensitive fields; Achieve differential privacy protection by adding Laplace noise; Use threshold encryption technology to shard and distribute data to multiple participating parties.
6. The cross - domain collaborative analysis method for a trusted data space based on privacy computing according to claim 1, wherein, The execution of collaborative analysis tasks in Step 5 includes: Split the computing task into multiple subtasks for parallel execution; Adopt a compression algorithm to reduce the amount of transmitted data for intermediate results; Use a redundancy check mechanism during the aggregation phase to ensure result integrity.
7. The cross - domain collaborative analysis method for a trusted data space based on privacy computing according to claim 1, characterized in that, The indicators of the dynamic security monitoring described in Step 6 include: Abnormal fluctuations in the CPU / memory occupancy rate of the monitoring nodes; Monitoring whether the communication link delay exceeds the threshold; Monitoring the deviation degree of the intermediate result from the expected distribution.
8. The cross-domain collaborative analysis method of the trusted data space based on privacy computing according to claim 1, wherein The audit report in Step 7 includes the following data: Data contribution degree evaluation data of each participating party; Privacy protection algorithm parameters and execution effect data; Data on task execution time and resource consumption statistics.
9. The cross - domain collaborative analysis method for a trusted data space based on privacy computing according to claim 1, characterized in that, It also includes deploying a security gateway between the trusted data space and the external system, and the security gateway provides protocol conversion, traffic filtering and intrusion prevention functions.
Citation Information
Patent Citations
Secure multi-party computing collusion attack defending method based on alliance chain
CN116980117A
Privacy protection method for multi-mechanism joint data value sharing service
CN118869260A
Secure multi-agent system for privacy-preserving distributed computation
US12316753B1
Cited By
Structured data field use control method and device, medium and product
CN120561161A
Block chain-based traffic data security encryption and circulation system and method
CN120729629A
Cross-domain computing task processing method, program product, equipment and medium
CN120856354A
A cross-domain computing task processing method, program product, device and medium
CN120856354B
Emergency rescue integrated data security protection method and system based on block chain
CN121000416A