Cross-industry data security sharing method and system based on data desensitization and medium

Through the integration of dynamic hierarchical desensitization and a trusted execution environment, the problems of high privacy leakage risks, rigid static desensitization rules, and low multi-party computing efficiency in cross-industry data sharing are solved, and end-to-end protection of data and a secure and mutually trusted data circulation ecosystem are realized.

CN120223391APending Publication Date: 2025-06-27LINGSHU TECH CO LTD
View PDF 0 Cites 20 Cited by

Patent Information

Application Number
CN202510365682.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

Smart Images

  • Figure CN120223391A_ABST
    Figure CN120223391A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-industry data security sharing method and system based on data desensitization and a medium, and relates to the related field of privacy computing, and the method comprises the following steps: carrying out dynamic desensitization processing on original data, generating a shared data set for real-time recording, constructing an access behavior characteristic spectrum to configure a dynamic access strategy, and establishing a third-party security data sharing platform, storing the shared data set in a trusted execution environment in a fragmented manner, obtaining fragmented encrypted data, executing a dynamic access strategy, performing conjoint analysis on the fragmented encrypted data through a security multi-party computing protocol, and generating a data security sharing suggestion. The technical problems that in cross-industry data sharing, the privacy leakage risk is high, the static desensitization rule is rigid, and the multi-party calculation efficiency is low are solved, and the technical effects of breaking data islands and constructing safe and mutual-trust data circulation ecology through fusion of dynamic hierarchical desensitization and a trusted execution environment are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of privacy computing, and in particular to a cross-industry data security sharing method, system and medium based on data desensitization. Background Art

[0002] With the explosive growth of data volume and the increasing demand for cross-industry data sharing, traditional data desensitization technologies are facing severe challenges in terms of efficiency, dynamic adaptability, and security. Existing technologies mainly adopt static desensitization rules (such as fixed masks, hash replacements, etc.), which are difficult to meet the dynamic processing requirements of multi-source heterogeneous data. Especially in sensitive industries such as finance and healthcare, there are problems such as a decrease in data utility caused by over-desensitization or privacy leakage caused by insufficient desensitization. In addition, although existing solutions such as the full-flow desensitization patent of Huoyin Technology introduce machine learning to identify sensitive data, they still rely on a single desensitization strategy in cross-industry scenarios and do not combine trusted execution environment (TEE) or secure multi-party computation (MPC) technologies, making it difficult to achieve end-to-end protection during the data sharing process. At the same time, cross-industry data sharing faces the problem of fragmented policies. For example, static desensitization rules cannot adapt to the privacy policies of different industries (such as GDPR, HIPAA), resulting in difficulty in balancing data availability and compliance. Summary of the Invention

[0003] This application provides a cross-industry data security sharing method, system and medium based on data desensitization, which solves the technical problems of high privacy leakage risk, rigid static desensitization rules, and low multi-party computation efficiency in cross-industry data sharing, and realizes the technical effect of breaking data islands and building a secure and trustworthy data circulation ecosystem through the integration of dynamic hierarchical desensitization and trusted execution environment.

[0004] This application provides a cross-industry data security sharing method based on data desensitization. The method is applied to a cross-industry data security sharing system based on data desensitization and includes: performing dynamic desensitization processing on the original data to generate a shared data set; based on the shared data set, performing real-time recording to construct an access behavior feature map, and configuring a dynamic access policy according to the access behavior feature map; establishing a third-party secure data sharing platform, storing the shared data set in slices in a trusted execution environment to obtain sliced encrypted data; executing the dynamic access policy, and performing joint analysis on the sliced encrypted data through a secure multi-party computation protocol to generate data security sharing suggestions.

[0005] In a possible implementation, dynamic desensitization processing is performed on the original data to generate a shared data set, and the following processing is executed: parsing is performed based on the original data to obtain geographical location data and numerical data; masking rules are set according to the data sensitivity level, and a generalization algorithm is used for the geographical location data according to the masking rules to obtain first masking information; random noise perturbation is added to the numerical data according to the masking rules to obtain second masking information; a desensitized metadata label is generated according to the first masking information and the second masking information, and the original data is filtered according to the desensitized metadata label to generate the shared data set.

[0006] In a possible implementation, masking rules are set according to the data sensitivity level, and the following processing is executed: semantic analysis is performed based on the original data to obtain field semantic features, and management analysis is performed based on the original data to obtain context association relationships; the data sensitivity level is divided according to the field semantic features and the context association relationships; predefined masking intensity parameters are matched according to the data sensitivity level to generate the masking rules.

[0007] In a possible implementation, the original data is filtered according to the desensitized metadata label to generate the shared data set, and the following processing is executed: a cross-industry sharing scenario is retrieved, and the desensitized metadata label is matched with the cross-industry sharing scenario to generate multiple industry information, where the multiple industry information includes first industry information and second industry information; when the data recipient is the first industry information, a generalization strategy is formulated to filter the original data to generate first filtered data; when the data recipient is the second industry information, a perturbation strategy is formulated to filter the original data to generate second filtered data; the first filtered data and the second filtered data are jointly modeled to generate the shared data set.

[0008] In a possible implementation, real-time recording is performed based on the shared data set to construct an access behavior feature map, and the following processing is executed: data capture is performed on the shared data set to generate quadruple logs; time series analysis is performed based on the quadruple logs, and a multi-dimensional behavior vector is constructed according to the behavior time series; correlation analysis is performed on the quadruple logs through a graph neural network to construct a graph structure space; the multi-dimensional behavior vector is mapped to the graph structure space for update to obtain the access behavior feature map.

[0009] In a possible implementation, a third-party secure data sharing platform is established, and the shared data set is shard-stored in a trusted execution environment to obtain shard-encrypted data, and the following processing is performed: A sharding rule is generated according to the field correlation of the shared data set, and the shared data set is split according to the sharding rule into multiple shard data; The trusted execution environment is initialized and verified based on the third-party secure data sharing platform, and an encrypted shard storage module is loaded according to the verification result; The multiple shard data are encrypted through the encrypted shard storage module to generate multiple encrypted nodes, and the multiple encrypted nodes are dynamically aggregated to generate the shard-encrypted data.

[0010] In a possible implementation, the dynamic access policy is executed, and the shard-encrypted data is jointly analyzed through a secure multi-party computation protocol to generate a data security sharing recommendation, and the following processing is performed: When a data requester initiates a joint analysis, the multiple encrypted nodes are activated to perform privacy computation on the shard-encrypted data through the secure multi-party computation protocol to generate an intermediate result set; A risk assessment is performed based on the intermediate result set to generate risk assessment indicators, and the dynamic access policy is adjusted according to the risk assessment indicators to generate a multi-party access adjustment result; The multi-party access adjustment results are aggregated to generate a data security sharing recommendation.

[0011] In a possible implementation, when a data requester initiates a joint analysis, the multiple encrypted nodes are activated to perform privacy computation on the shard-encrypted data through the secure multi-party computation protocol to generate an intermediate result set, and the following processing is performed: In response to a joint analysis instruction initiated by the data requester, the multiple encrypted nodes are parsed to determine node requirement characteristics; The trusted node pool is traversed according to the node requirement characteristics for dynamic screening to construct hierarchical computation topology data; The secure multi-party computation protocol is triggered according to the hierarchical computation topology data, and verification is performed through zero-knowledge proof within the trusted execution environment to generate the intermediate result set.

[0012] This application also provides a cross-industry data security sharing system for data masking, including: A masking processing module for performing dynamic masking processing on original data to generate a shared data set; A policy configuration module for performing real-time recording based on the shared data set, constructing an access behavior feature map, and configuring a dynamic access policy according to the access behavior feature map; A data acquisition module for establishing a third-party secure data sharing platform, shard-storing the shared data set in a trusted execution environment to obtain shard-encrypted data; A joint analysis module for executing the dynamic access policy and jointly analyzing the shard-encrypted data through a secure multi-party computation protocol to generate a data security sharing recommendation.

[0013] The present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a cross-industry data security sharing method based on data desensitization are implemented.

[0014] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0015] The cross-industry data security sharing method, system and medium based on data desensitization provided by the present application relate to the technical field of privacy computing, solve the technical problems of high risk of privacy leakage, rigid static desensitization rules and low multi-party computing efficiency in cross-industry data sharing, and achieve the technical effect of breaking data islands and building a secure and trustworthy data circulation ecosystem through the integration of dynamic hierarchical desensitization and a trusted execution environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings of the embodiments of the present application will be briefly introduced below. Flowcharts are used in the present application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations described above or below do not necessarily need to be executed precisely in order. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0017] Figure 1 It is a schematic flowchart of the cross-industry data security sharing method based on data desensitization provided by the embodiment of the present application;

[0018] Figure 2 It is a schematic structural diagram of the cross-industry data security sharing system based on data desensitization provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the detailed description of the present application.

[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0021] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art belonging to the technical field of this application. The terms used herein are only for the purpose of describing the embodiments of this application.

[0022] Embodiment 1 provides a cross-industry data security sharing method for data masking, and the method is applied to a cross-industry data security sharing system for data masking, as Figure 1 shown, the method includes:

[0023] Step A100: Perform dynamic masking processing on the original data to generate a shared data set; in a possible implementation manner, step A100 further includes step A110: Parse the original data to obtain geographical location data and numerical data; execute step A120: Set a masking rule according to the data sensitivity level, and use a generalization algorithm for the geographical location data according to the masking rule to obtain first masking information; execute step A130: Add random noise perturbation to the numerical data according to the masking rule to obtain second masking information; execute step A140: Generate a masked metadata label according to the first masking information and the second masking information, and screen the original data according to the masked metadata label to generate the shared data set.

[0024] First, based on the characteristics of structured data fields, intelligent parsing of the original data is performed using regular expression matching and a semantic analysis model (such as BERT). For geographical location data, Geohash encoding is used to identify latitude and longitude coordinates or text addresses (such as "No. 1, Zhongguancun Street, Haidian District, Beijing"), and the field type is marked as "geographical location"; for numerical data (such as transaction amounts, medical test values), its numerical attributes are determined through statistical distribution analysis (skewness, kurtosis) and a business rule library (such as the financial field amount threshold). During the parsing process, parallel processing of TB-level data is achieved through a distributed computing framework (such as Spark), and the processing delay is controlled within milliseconds.

[0025] Furthermore, by constructing a three-dimensional scoring model, comprehensively considering the regulatory relevance (mapping of GDPR / HIPAA and other terms), field association risk (statistics of historical leakage events), and business value weight (reusability rate of data in cross-industry analysis), the field sensitivity score is calculated. For example, the "patient address" in medical data involves HIPAA privacy terms and has a score of 0.92 (high sensitivity); the "user consumption amount" in retail data has no direct identity association and has a score of 0.45 (low sensitivity). The scoring results are updated in real time to the metadata database to support dynamic policy adjustment. And according to the data sensitivity level, mask rules are set, and a generalization algorithm is adopted for the geographical location data according to the mask rules, which means adopting a hierarchical mapping strategy. For high-sensitivity data (score ≥ 0.8), the precise coordinates are converted into administrative division codes (such as "116.3°E, 39.9°N" → "110108"); for medium-sensitivity data (0.5 ≤ score < 0.8), it is retained to the street level (such as "Zhongguancun Avenue"); for low-sensitivity data (score < 0.5), only the city name is retained. The generalization granularity is dynamically controlled through the Geohash precision parameter (for example, 7-bit Geohash corresponds to an accuracy of ±19 meters, and when reduced to 4 bits, the accuracy is expanded to ±20 kilometers) to generate the first mask information.

[0026] Furthermore, according to the mask rules, random noise perturbations are added to the numerical data, which means that based on the differential privacy framework, Laplace noise is added to high-sensitivity numerical values (Δf / ε = 0.1, ε = 0.3) to ensure that individual data cannot be inferred; Gaussian noise (σ = 0.5 × data standard deviation) is used for medium-sensitivity data; light perturbations (±5% random floating) are performed on low-sensitivity data. The noise parameters are optimized through Monte Carlo simulation to ensure that the deviation of the statistical characteristics (mean, variance) of the perturbed data is ≤ 3% to generate the second mask information.

[0027] According to the first mask information and the second mask information, a desensitized metadata tag is generated. A desensitized metadata tag system can be constructed, including desensitization methods (such as generalization / noise), parameters (such as ε value, Geohash accuracy), and original-desensitized mapping tables (encrypted and stored in blockchain), and then the original data is screened according to the desensitized metadata tags, that is, the data usage scenarios are analyzed through a SQL-like strategy engine (such as Apache Calcite), such as automatically filtering highly sensitive fields in medical research scenarios, and only retaining generalized administrative divisions; financial risk control scenarios retain perturbed values ​​but shield precise geographic locations. When the shared data set is finally generated, column storage (Parquet format) is used and metadata tag indexes are attached for subsequent audit tracing, and the shared data set is generated to achieve the optimal balance between privacy protection and data utility. Actual measurements show that the re-identification rate of medical address data has dropped from 12% of traditional methods to 0.5%, the statistical error rate of financial transaction data is ≤2.1%, and the processing throughput reaches 100,000 records / second, supporting real-time and secure sharing in cross-industry scenarios.

[0028] In one possible implementation, step A120 further includes step A121, performing semantic analysis based on the original data to obtain field semantic features, performing management analysis based on the original data to obtain contextual association relationships; executing step A122, dividing the data sensitivity level according to the field semantic features and the contextual association relationships; executing step A123, matching predefined mask strength parameters according to the data sensitivity level to generate the mask rules.

[0029] First, use a distributed data parsing engine (such as Apache Arrow) to standardize the format of the raw data, identify field types (numeric, text, spatiotemporal, etc.), enable the natural language processing (NLP) pipeline for unstructured text fields (such as medical record descriptions), and use pre-trained models (such as BioBERT for medical text and FinBERT for financial text) to extract entity information (such as person names, organization names, disease names) and semantic roles (agent, patient, time).

[0030] Then, we build a field relationship graph and use the graph neural network (GNN) to mine implicit associations between fields. For example, in financial scenarios, when the co-occurrence frequency of "transaction amount" and "receiving account" exceeds the threshold, it is marked as a highly associated combination; in medical scenarios, the association path length between "patient ID" and "diagnosis result" is used as the basis for sensitivity assessment. At the same time, we use the LSTM model to analyze the temporal context (such as multiple consecutive visits by the same user) and identify abnormal association patterns (such as short-term high-frequency access to cross-industry data).

[0031] Furthermore, the data sensitivity level can be divided according to the semantic features of the fields and the context association relationship. By matching the fields with privacy terms (such as the definition of personal data in Article 4 of GDPR) based on a knowledge graph, calculating the matching strength (0-1), etc., the context risk value can be determined by calculating the node centrality through the association graph. At the same time, the business value weight can be determined by counting the field reuse rate according to cross-industry sharing requirements. On this basis, a fuzzy logic controller (FLC) is used to dynamically classify the scoring results, which can include high sensitivity (≥0.8): direct identifiers (such as mobile phone numbers), highly associated combinations (such as "name + address"), medium sensitivity (0.5-0.8): quasi-identifiers (such as age, gender), time-series sensitive data (such as continuous location tracks), low sensitivity (<0.5): aggregated data (such as regional statistical values), de-identified data. Real-time monitoring of changes in data usage scenarios (such as new cross-border sharing requirements), when it is detected that the field reuse rate increases by 10% or the associated risk value fluctuates by more than 15%, trigger a recalculation of the score and an adjustment of the level.

[0032] Match the predefined mask strength parameters according to the data sensitivity level, and establish a three-layer mask strength parameter matrix. The three-layer mask strength parameter matrix includes strong masks, which are applicable to high-sensitivity data, such as format-preserving hashing (FPH), full-field replacement; medium masks, which are applicable to medium-sensitivity data, such as partial masking (mobile phone number "138****1234"), generalization (date "2023-10-05 → 2023-10"); light masks, which are applicable to low-sensitivity data, such as noise injection (transaction amount ±5% perturbation), character truncation (email "ex***@domain.com"). At the same time, generate the optimal mask combination based on the sensitivity level and industry policies, determine the mask rules, and achieve fine-grained sensitivity dynamic evaluation. The classification accuracy rate reaches 94.7% (38% higher than the traditional rule engine), and the mask rule generation takes less than 50 ms, supporting real-time processing of tens of thousands of fields.

[0033] In a possible implementation, step A140 further includes step A141, retrieving cross-industry sharing scenarios, matching the desensitization metadata tags with the cross-industry sharing scenarios to generate multiple industry information, where the multiple industry information includes first industry information and second industry information; performing step A142, when the data recipient is the first industry information, formulating a generalization strategy to screen the original data to generate first screened data; performing step A143, when the data recipient is the second industry information, formulating a perturbation strategy to screen the original data to generate second screened data; performing step A144, jointly modeling the first screened data and the second screened data to generate the shared dataset.

[0034] Retrieving cross - industry sharing scenarios can extract key features based on a data usage statement (such as the industry type and usage description of the data recipient) through a natural language processing (NLP) model (such as BERT). At the same time, match de - sensitized metadata tags (such as sensitivity level, de - sensitization method, compliance terms) with the cross - industry sharing scenario to determine multiple industry information, and extract the first industry information and the second industry information based on the multiple industry information, where the first industry information is the medical industry and the second industry information is the financial industry.

[0035] Further, when the data recipient is the first industry information, formulating a generalization strategy to screen the original data means converting an exact address (such as "No. 1, Zhongguancun Street, Haidian District, Beijing") into a Geohash code (precision level 4, corresponding administrative division code "110108") to ensure k - anonymity (k≥5) for geographical location processing, then using format - preserving encryption (FPE) to replace the disease name (such as "diabetes") with a pseudonym (such as "DIA - 0032"), while retaining the disease classification code (ICD - 10) for research use to complete the desensitization of diagnostic information, and filtering out direct identifiers (patient ID, mobile phone number), retaining the generalized diagnosis time, age range (such as "30 - 40 years old"), and generalized address to generate the first screened data.

[0036] When the data recipient is the second industry information, formulating a perturbation strategy to screen the original data means adding Laplace noise that satisfies (ε = 0.1, δ = 1e - 5) - differential privacy to the transaction amount, with the statistical error of the perturbed data ≤2% to complete numerical perturbation, retaining the last four digits of the bank card number and replacing the rest with a hash value (SHA - 256) to ensure irreversibility, masking the exact bits of the IP address and generalizing it to the city level (such as "Shanghai"), and retaining the perturbed transaction time and amount range to generate the second screened data.

[0037] Finally, jointly modeling the first screened data and the second screened data means that within a trusted execution environment (TEE), aligning the common features (such as city, time range) of the two - industry data through a private set intersection (PSI) protocol to avoid directly exposing intersection information, and then adopting a horizontal federated learning framework. The medical data participants locally train model parameters (such as logistic regression weights), and the financial data participants upload encrypted gradients through a secure aggregation protocol. The central server aggregates and updates the global model. Throughout the process, the plaintext data does not leave the domain. Encrypt the desensitization rule version, data hash value, and joint model parameters and upload them to the blockchain to support multi - party auditing. At the same time, monitor the data access pattern in real - time. If a cross - industry high - frequency query is detected (such as a medical data recipient suddenly requesting financial features), immediately interrupt the session and initiate manual review, and generate the shared dataset.

[0038] Execute step A200 to perform real-time recording based on the shared dataset, construct an access behavior feature map, and configure a dynamic access policy according to the access behavior feature map; in a possible implementation, step A200 further includes step A210 of generating a quadruple log by performing data capture on the shared dataset; execute step A220 to perform time series analysis based on the quadruple log and construct a multi-dimensional behavior vector according to the behavior time series; execute step A230 to perform correlation analysis on the quadruple log through a graph neural network and construct a graph structure space; execute step A240 to map the multi-dimensional behavior vector to the graph structure space for updating to obtain the access behavior feature map.

[0039] Based on the shared dataset, real-time recording of data access behaviors and usage logs is performed by deploying lightweight probes at each node of the data sharing platform to capture user access behaviors in real time and extract core fields, which can include identity characteristics, operation characteristics, time characteristics, and object characteristics. The raw logs are windowed (sliding window = 1 second) through Apache Flink, and duplicate removal, field desensitization (such as replacing IP with geographical area encoding), and format standardization are performed to output a structured quadruple log. At the same time, performing time series analysis based on the quadruple log refers to extracting time series features from the quadruple log, that is, calculating the operation frequency, time interval variance, and peak period within a sliding window (5 minutes) for feature statistics. The operation period can also be detected through FFT (such as batch export every 15 minutes), and the periodic intensity factor is extracted for periodic analysis. The LSTM model is used to predict the behavior trend within the next 10 minutes, and the hidden state vector is output for trend prediction. The time series features (statistics + period + trend) and static features (user role, device type) are concatenated into a 512-dimensional behavior vector, which is reduced to 256 dimensions through PCA to reduce redundancy and retain 95% of the variance information, thereby constructing a multi-dimensional behavior vector.

[0040] Furthermore, the correlation analysis of the quadruple logs through a graph neural network means that a view graph can be constructed in batches based on historical logs and incrementally updated according to real-time stream data. That is, when a new user visits for the first time, a node is created and the "first access time" attribute is added. At the same time, the edge weights are dynamically adjusted according to the time decay factor. The constructed graph structure space contains nodes and edges. The node types include user nodes (attributes: role, authentication level), data nodes (attributes: sensitivity level, industry label), operation nodes (attributes: action type, risk score), and device nodes (attributes: geographical location, security status). The edge relationships can include user - execute - operation (weight = operation frequency), operation - act on - data (weight = data sensitivity), data - store - device (weight = access frequency). Finally, the multi-dimensional behavior vector is mapped to the graph structure space for update. The behavior vector can be mapped to graph nodes by using the graph attention mechanism (GAT) to calculate the attention coefficients between user nodes and their associated operation nodes. At the same time, the full-graph feature propagation can be triggered every 5 minutes to update the node embeddings (Node2Vec algorithm) to identify potential risk clusters (such as user groups that frequently access highly sensitive data). On this basis, the access behavior feature map is obtained, and the map update delay is controlled within 200 ms, supporting real-time processing of 500,000 logs per second. For example, when it is detected that a certain user frequently accesses patient address data (deviating from the normal pattern) at 3 am, the system blocks the access and issues an alarm within 10 seconds, and the traceability efficiency is improved by 90% compared with the traditional solution.

[0041] Execute step A300 to establish a third-party secure data sharing platform, store the shared data set in slices in a trusted execution environment, and obtain sliced encrypted data. In a possible implementation, step A300 further includes step A310 of generating a slicing rule according to the field relevance of the shared data set, splitting the shared data set according to the slicing rule into multiple sliced data; execute step A320 of initializing and validating the trusted execution environment based on the third-party secure data sharing platform, and loading an encrypted slice storage module according to the verification result; execute step A330 of encrypting the multiple sliced data through the encrypted slice storage module to generate multiple encrypted nodes, and dynamically aggregating the multiple encrypted nodes to generate the sliced encrypted data.

[0042] First, when generating a slicing rule based on the field relevance of the shared data set, it is necessary to analyze the business coupling degree and access pattern between fields, calculate the field correlation degree using a graph model, and aggregate highly correlated fields (such as user ID and transaction records) into logical slice units. Subsequently, the original data set is split into multiple slices using the consistent hashing algorithm to ensure balanced data distribution and minimize cross-slice query requirements.

[0043] During the initialization phase of the Trusted Execution Environment (TEE), the third-party platform verifies the integrity of the TEE startup chain through a remote authentication mechanism, including measuring the hash value of the startup code, checking the security configuration policy, and generating an environmental trust report. After successful verification, a cryptographic shard storage module based on SGX or TrustZone is dynamically loaded, which uses in-memory encryption technology to protect temporary data during the shard processing.

[0044] Encrypt the multiple shard data through the cryptographic shard storage module. The shard-level encrypted data generates a dynamic key based on the shard identifier using the AES-GCM algorithm; for field-level encryption, format-preserving encryption (FPE) technology is used for sensitive fields (such as ID numbers). The encrypted shard data is encapsulated as encrypted nodes and indexed and managed through a distributed hash table (DHT), supporting dynamic aggregation of encrypted nodes according to access permissions. During the aggregation process, homomorphic encryption technology is used to perform data association calculations in the ciphertext state to ensure that query operations do not expose plaintext information, thereby obtaining shard-encrypted data, which is concatenated through an automated orchestration engine to achieve full-life cycle management of shard rule generation, encryption key management, and TEE status monitoring.

[0045] Execute step A400, execute the dynamic access policy, and perform a joint analysis on the shard-encrypted data through a secure multi-party computation protocol to generate data security sharing recommendations.

[0046] In a possible implementation, step A400 further includes step A410. When the data requester initiates a joint analysis, activate the multiple encrypted nodes to perform privacy calculations on the shard-encrypted data through a secure multi-party computation protocol to generate an intermediate result set. In a possible implementation, step A410 further includes step A411. In response to the joint analysis instruction initiated by the data requester, parse the multiple encrypted nodes to determine the node demand characteristics; execute step A412, traverse the trusted node pool according to the node demand characteristics for dynamic screening, and construct hierarchical computing topology data; execute step A413, trigger the secure multi-party computation protocol according to the hierarchical computing topology data, and perform verification through zero-knowledge proof within the trusted execution environment to generate the intermediate result set.

[0047] When the data requester initiates a joint analysis, the system parses the node demand characteristics in the analysis instruction through a smart contract, including parameters such as data type, calculation dimension, and security level, and dynamically constructs a cross-institutional trusted node pool. Based on the node capability map (covering computing power, algorithm library, and security authentication level), three-layer screening is performed: first, filter candidate nodes that meet the basic security threshold; second, match dedicated algorithm nodes according to the calculation task type (such as federated learning nodes, homomorphic encryption computing nodes); finally, select the node combination with the optimal cooperation efficiency through a game theory mechanism.

[0048] Dynamically screen by traversing the trusted node pool according to the node demand characteristics. This is achieved by adopting a directed acyclic graph (DAG) structure, which decomposes the computing task into an input layer (data preprocessing), a computing layer (privacy computing), and a verification layer (zero-knowledge proof). Nodes in the input layer perform local differential privacy processing. Nodes in the computing layer implement multi-party data collaborative computing through the SMPC protocol (such as the ABY3 framework). Nodes in the verification layer use zk-SNARK to generate proofs of computing correctness, thereby obtaining hierarchical computing topology data.

[0049] Furthermore, when executed in the TEE environment, the confidentiality of the computing process is ensured through hardware-level memory isolation, and a remote verification mechanism is adopted to ensure the integrity of the execution environment. The zero-knowledge proof process adopts a circuitized verification scheme, converting the computing result into a verifiable circuit, generating a proof file through elliptic curve encryption. The verifier only needs to verify the validity of the proof to confirm the computing correctness, without exposing the original data and intermediate states throughout the process. Thus, the intermediate result set is determined, that is, dynamic load balancing is achieved through an automated topology optimization engine. When the computing load of a certain node exceeds the threshold, topology reconstruction is automatically triggered, and part of the computing tasks are migrated to standby nodes. The finally generated intermediate result set is protected by a homomorphic hash chain, and the decryption permission is only open to authorized nodes, realizing the minimum necessary exposure of the computing result.

[0050] Execute step A420, perform risk assessment based on the intermediate result set, generate risk assessment indicators, and adjust the dynamic access policy according to the risk assessment indicators to generate a multi-party access adjustment result; execute step A430, aggregate the multi-party access adjustment result to generate a data security sharing recommendation.

[0051] The risk assessment and access policy adjustment process based on the intermediate result set adopts a closed-loop feedback mechanism to achieve intelligent decision-making. First, by constructing a multi-modal risk assessment model, comprehensively calculate data sensitivity indicators (such as PHI value for medical data identification), access behavior anomaly degree (sequence pattern detection based on LSTM), and environmental risk coefficient (IP reputation library and device fingerprint verification). The model outputs a three-dimensional risk vector, which is mapped to a dynamic access policy matrix, where each policy entry contains parameters such as permission level, validity period, and encryption granularity.

[0052] The policy adjustment engine adopts a reinforcement learning framework, uses the risk vector as the state input, and selects the optimal policy action through the Q-learning algorithm. For example, when it is detected that a certain node continuously accesses highly sensitive fields, the engine automatically triggers policy contraction: upgrading column-level access control to row-level filtering, dynamically adjusting the execution intensity of the access policy and enabling dynamic desensitization algorithms. Policy changes are broadcast to the trusted node pool through digital signatures to ensure the atomicity and non-repudiation of policy updates.

[0053] Furthermore, the multi-party access adjustment results are aggregated. By adopting a federated learning framework, each participating party trains a policy evaluation model locally and only exchanges model parameter gradients instead of raw data. The gradients are aggregated through homomorphic encryption to generate a global policy optimization direction. The finally generated shared recommendations include a differential sharing scheme: attribute-based encryption sharing is used for low-risk requests, multi-party secure computation is implemented for medium risks, and a manual review process is initiated for high-risk operations, and an audit log containing data watermarks is generated. That is, based on the aggregated multi-party computation results, data security sharing recommendations are generated, and the data security sharing recommendations include data usage constraints and compliance operation guidelines.

[0054] The embodiments of the present application solve the technical problems of high privacy leakage risk, rigid static desensitization rules, and low multi-party computation efficiency in cross-industry data sharing, and achieve the technical effect of breaking data islands and building a secure and trustworthy data circulation ecosystem through the integration of dynamic hierarchical desensitization and a trusted execution environment.

[0055] In the above, reference is made to Figure 1 The cross-industry data security sharing method for data desensitization according to the embodiments of the present application is described in detail. Next, reference will be made to Figure 2 Describe the cross-industry data security sharing system for data desensitization according to the embodiments of the present application.

[0056] Embodiment 2, a cross-industry data security sharing system for data desensitization, which solves the technical problems of high privacy leakage risk, rigid static desensitization rules, and low multi-party computation efficiency in cross-industry data sharing, and achieves the technical effect of breaking data islands and building a secure and trustworthy data circulation ecosystem through the integration of dynamic hierarchical desensitization and a trusted execution environment. The cross-industry data security sharing system for data desensitization includes: a desensitization processing module 10, a policy configuration module 20, a data acquisition module 30, and a joint analysis module 40.

[0057] The desensitization processing module 10 is used to perform dynamic desensitization processing on the original data to generate a shared data set; the policy configuration module 20 is used to perform real-time recording based on the shared data set, construct an access behavior feature map, and configure a dynamic access policy according to the access behavior feature map; the data acquisition module 30 is used to establish a third-party secure data sharing platform, store the shared data set in slices in a trusted execution environment, and obtain sliced encrypted data; the joint analysis module 40 is used to execute the dynamic access policy, and perform joint analysis on the sliced encrypted data through a secure multi-party computation protocol to generate data security sharing recommendations.

[0058] Next, the specific configuration of the desensitization processing module 10 will be described in detail. As described above, for dynamic desensitization processing of the original data to generate a shared data set, the desensitization processing module 10 may further include: parsing the original data to obtain geographical location data and numerical data; setting a masking rule according to the data sensitivity level, and using a generalization algorithm for the geographical location data according to the masking rule to obtain first masking information; adding random noise perturbation to the numerical data according to the masking rule to obtain second masking information; generating a desensitized metadata label according to the first masking information and the second masking information, and screening the original data according to the desensitized metadata label to generate the shared data set.

[0059] Next, the specific configuration of the desensitization processing module 10 will be described in detail. As described above, for setting a masking rule according to the data sensitivity level, the desensitization processing module 10 may further include: performing semantic analysis on the original data to obtain field semantic features, and performing management analysis on the original data to obtain context association relationships; dividing the data sensitivity level according to the field semantic features and the context association relationships; matching predefined masking intensity parameters according to the data sensitivity level to generate the masking rule.

[0060] Next, the specific configuration of the desensitization processing module 10 will be described in detail. As described above, for screening the original data according to the desensitized metadata label to generate the shared data set, the desensitization processing module 10 may further include: invoking cross-industry sharing scenarios, matching the desensitized metadata label with the cross-industry sharing scenarios to generate multiple industry information, where the multiple industry information includes first industry information and second industry information; when the data recipient is the first industry information, formulating a generalization strategy to screen the original data to generate first screening data; when the data recipient is the second industry information, formulating a perturbation strategy to screen the original data to generate second screening data; jointly modeling the first screening data and the second screening data to generate the shared data set.

[0061] Next, the specific configuration of the policy configuration module 20 will be described in detail. As described above, for real-time recording based on the shared data set to construct an access behavior feature map, the policy configuration module 20 may further include: capturing data from the shared data set to generate a quadruple log; performing time series analysis on the quadruple log and constructing a multi-dimensional behavior vector according to the behavior time series; performing association analysis on the quadruple log through a graph neural network to construct a graph structure space; mapping the multi-dimensional behavior vector to the graph structure space for updating to obtain the access behavior feature map.

[0062] Next, the specific configuration of the data acquisition module 30 will be described in detail. As described above, a third-party secure data sharing platform is established, and the shared data set is fragmented and stored in the trusted execution environment to obtain fragmented encrypted data. The data acquisition module 30 may further include: generating a fragmentation rule according to the field correlation of the shared data set, splitting the shared data set according to the fragmentation rule into multiple fragmented data; initializing and verifying the trusted execution environment based on the third-party secure data sharing platform, and loading the encrypted fragmentation storage module according to the verification result; encrypting the multiple fragmented data through the encrypted fragmentation storage module to generate multiple encrypted nodes, and dynamically aggregating the multiple encrypted nodes to generate the fragmented encrypted data.

[0063] Next, the specific configuration of the joint analysis module 40 will be described in detail. As described above, the dynamic access policy is executed, and the fragmented encrypted data is jointly analyzed through a secure multi-party computation protocol to generate data security sharing suggestions. The joint analysis module 40 may further include: when a data requester initiates a joint analysis, activating the multiple encrypted nodes to perform privacy computation on the fragmented encrypted data through the secure multi-party computation protocol to generate an intermediate result set; performing a risk assessment based on the intermediate result set to generate risk assessment indicators, and adjusting the dynamic access policy according to the risk assessment indicators to generate a multi-party access adjustment result; aggregating the multi-party access adjustment result to generate data security sharing suggestions.

[0064] Next, the specific configuration of the joint analysis module 40 will be described in detail. As described above, when a data requester initiates a joint analysis, activating the multiple encrypted nodes to perform privacy computation on the fragmented encrypted data through the secure multi-party computation protocol to generate an intermediate result set. The joint analysis module 40 may further include: responding to a joint analysis instruction initiated by the data requester, parsing the multiple encrypted nodes to determine the node requirement characteristics; dynamically screening through the trusted node pool according to the node requirement characteristics to construct hierarchical computation topology data; triggering the secure multi-party computation protocol according to the hierarchical computation topology data, and verifying through zero-knowledge proof within the trusted execution environment to generate the intermediate result set.

[0065] The cross-industry data security sharing system for data desensitization provided by the embodiments of the present application can execute the cross-industry data security sharing method for data desensitization provided by any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the method.

[0066] Through the foregoing detailed description of the cross-industry data security sharing method for data desensitization, those skilled in the art can clearly understand the cross-industry data security sharing system for data desensitization in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For related parts, reference may be made to the description in the method part.

[0067] Embodiment 3 provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, any step of Embodiment 1 is implemented.

[0068] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0069] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A cross-industry data security sharing method based on data desensitization, characterized in that: The method comprises: Perform dynamic desensitization on the original data to generate a shared data set; Based on the shared data set, real-time recording is performed to build an access behavior feature map, and a dynamic access strategy is configured according to the access behavior feature map; Establish a third-party secure data sharing platform, store the shared data set shards in a trusted execution environment, and obtain shard encrypted data; The dynamic access strategy is executed, and the sharded encrypted data is jointly analyzed through a secure multi-party computing protocol to generate data security sharing recommendations.

2. The cross-industry data security sharing method based on data desensitization as claimed in claim 1 is characterized in that: The original data is dynamically desensitized to generate a shared data set. The method includes: Analyze the original data to obtain geographic location data and numerical data; Setting a mask rule according to the data sensitivity level, and applying a generalization algorithm to the geographic location data according to the mask rule to obtain first mask information; Adding random noise disturbance to the numerical data according to the masking rule to obtain second masking information; A desensitized metadata label is generated according to the first mask information and the second mask information, and the original data is screened according to the desensitized metadata label to generate the shared data set.

3. The cross-industry data security sharing method based on data desensitization as claimed in claim 2 is characterized in that: Set masking rules based on data sensitivity levels, including: Performing semantic analysis based on the original data to obtain field semantic features, and performing management analysis based on the original data to obtain contextual association relationships; Divide the data sensitivity level according to the field semantic features and the context association relationship; The mask rule is generated according to matching the predefined mask strength parameter with the data sensitivity level.

4. The cross-industry data security sharing method based on data desensitization as claimed in claim 2 is characterized in that: The original data is screened according to the desensitized metadata tag to generate the shared data set, the method comprising: Retrieving a cross-industry sharing scenario, matching the desensitized metadata tag with the cross-industry sharing scenario, and generating a plurality of industry information, wherein the plurality of industry information includes first industry information and second industry information; When the data recipient is the first industry information, a generalization strategy is formulated to filter the original data to generate first filtered data; When the data recipient is the second industry information, a disturbance strategy is formulated to filter the original data to generate second filtered data; The first screening data and the second screening data are jointly modeled to generate the shared data set.

5. The cross-industry data security sharing method based on data desensitization as claimed in claim 1 is characterized in that: Based on the shared data set, real-time recording is performed to construct an access behavior feature map, and the method includes: Generate a quadruple log by performing data capture on the shared data set; Performing time series analysis based on the four-tuple log and constructing a multi-dimensional behavior vector according to the behavior time series; Performing correlation analysis on the four-tuple logs through a graph neural network to construct a graph structure space; The multi-dimensional behavior vector is mapped to the graph structure space for updating to obtain the access behavior feature map.

6. The cross-industry data security sharing method based on data desensitization as claimed in claim 1 is characterized in that: Establish a third-party secure data sharing platform, store the shared data set in a trusted execution environment, and obtain encrypted data in the shared data set. The method includes: Generate a sharding rule according to the field association of the shared data set, and split the shared data set into multiple shard data according to the sharding rule; Initialize and verify the trusted execution environment based on a third-party secure data sharing platform, and load the encrypted shard storage module based on the verification results; The plurality of shard data are encrypted by the encrypted shard storage module to generate a plurality of encryption nodes, and the plurality of encryption nodes are dynamically aggregated to generate the shard encrypted data.

7. The cross-industry data security sharing method based on data desensitization as claimed in claim 6 is characterized in that: Executing the dynamic access strategy, jointly analyzing the sharded encrypted data through a secure multi-party computing protocol, and generating a data security sharing suggestion, the method includes: When the data requester initiates joint analysis, the multiple encryption nodes are activated to perform privacy calculations on the sharded encrypted data through a secure multi-party computing protocol to generate an intermediate result set; Perform risk assessment based on the intermediate result set to generate risk assessment indicators, adjust the dynamic access policy according to the risk assessment indicators, and generate a multi-party access adjustment result; The multi-party access adjustment results are aggregated to generate data security sharing recommendations.

8. The cross-industry data security sharing method based on data desensitization as claimed in claim 7 is characterized in that: When the data requester initiates joint analysis, the multiple encryption nodes are activated to perform privacy calculations on the sharded encrypted data through a secure multi-party computing protocol to generate an intermediate result set, and the method includes: In response to the joint analysis instruction initiated by the data requester, the plurality of encryption nodes are analyzed to determine the node requirement characteristics; Traversing the trusted node pool for dynamic screening according to the node demand characteristics, and constructing hierarchical computing topology data; The secure multi-party computing protocol is triggered according to the hierarchical computing topology data, and is verified through zero-knowledge proof in a trusted execution environment to generate the intermediate result set.

9. A cross-industry data security sharing system based on data desensitization, characterized by: The system is used to implement the cross-industry data security sharing method based on data desensitization according to any one of claims 1 to 8, and the system includes: The desensitization processing module is used to dynamically desensitize the original data and generate a shared data set; A policy configuration module, used to record in real time based on the shared data set, build an access behavior feature map, and configure a dynamic access policy according to the access behavior feature map; A data acquisition module is used to establish a third-party secure data sharing platform, store the shared data set slices in a trusted execution environment, and obtain the slice encrypted data; The joint analysis module is used to execute the dynamic access strategy, conduct joint analysis on the sharded encrypted data through a secure multi-party computing protocol, and generate data security sharing recommendations.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the cross-industry data security sharing method based on data desensitization described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Cross-industry big data sharing method and system

    CN120389868A

  • Data anonymization adjustment method and system based on dynamic association risk analysis

    CN120449213A

  • Data anonymization adjustment method and system based on dynamic correlation risk analysis

    CN120449213B

  • Structured data field use control method and device, medium and product

    CN120561161A

  • Government affair data circulation management method and system based on block chain

    CN120579219A