Market supervision data asset management method and system based on compliance and privacy protection
By modeling metadata labels, encapsulating privacy compliance policy and verifying knowledge graphs of market supervision data, combined with the privacy protection and processing mechanism, the dynamic management and value evaluation of market supervision data resources are solved, efficient compliance and privacy protection of data assets are achieved, and data management is improved.
Patent Information
- Application Number
- CN202510399619.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In the existing technology, the market supervision field lacks data governance means that dynamically adapt to complex regulatory scenarios, there is tension between privacy protection and data availability, and the data asset value evaluation mechanism is imperfect, which makes it difficult to effectively utilize data resources and difficult to guarantee compliance.
Through collection, cleaning and abstract modeling, data assets with metadata labels are formed, privacy compliance policies are generated and encapsulated into data capsules, privacy compliance knowledge graphs are built for semantic matching and compliance verification, differential privacy disturbance mechanisms, homomorphic encryption and secure multi-party computing and other processing mechanisms are called, and data asset value dynamic evaluation model is built to realize dynamic management and compliance guarantee of data assets.
It realizes efficient data management and value utilization under the requirements of privacy protection and compliance, improves the identification, manageability and traceability of data assets, reduces the risk of data leakage, improves compliance and security, and supports the differentiated management and value-driven operation of data assets.
Smart Images

Figure CN120296788A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data asset management, and particularly to a method and system for market supervision data asset management based on compliance and privacy protection. Background Art
[0002] With the acceleration of the digital governance process, a large amount of heterogeneous and multi-source data resources have gradually accumulated in the market supervision field. These data resources are of great value in serving regulatory decisions, assisting risk warnings, and promoting the improvement of governance efficiency. However, regulatory data usually involves enterprise operations, natural person behaviors, and other sensitive information. How to achieve the efficient utilization of data on the premise of ensuring data security and personal privacy protection has become an important issue that needs to be solved urgently.
[0003] In the prior art, the management of regulatory data mainly focuses on static classification and grading and access control, lacking data governance means that can dynamically adapt to complex regulatory scenarios. At the same time, privacy compliance requirements are becoming increasingly strict, and the differences in laws and regulations between countries and regions make the cross-domain use and sharing of data have relatively high compliance risks. In addition, there is a natural tension between privacy protection and data availability. The lack of a unified mechanism to coordinate and optimize the two often leads to the difficulty of effectively releasing the data value.
[0004] On the other hand, although the management of data assets has become a trend, the data asset value evaluation mechanism in the regulatory field is still imperfect, lacking the comprehensive judgment ability of data quality, risk level, and usage scenarios. This not only affects the scientific allocation of data resources but also restricts the development of intelligent regulatory means driven by data. Therefore, there is an urgent need for a data asset management method that can take into account privacy protection, compliance requirements, and data value realization to support the modernization transformation of the market supervision system. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a method and system for market supervision data asset management based on compliance and privacy protection, aiming to achieve the efficient management and value utilization of heterogeneous data resources in the market supervision field on the premise of meeting privacy protection and compliance requirements.
[0006] The present invention realizes the above object through the following technical solutions:
[0007] A method for market supervision data asset management based on compliance and privacy protection, the method comprising:
[0008] Collecting, cleaning, and abstracting and modeling the heterogeneous data resources of the regulatory object to form data assets with metadata tags;
[0009] Generating a privacy compliance policy according to the metadata tags, and encapsulating the privacy compliance policy and the data assets into a data capsule;
[0010] Construct a privacy compliance knowledge graph and use a large language model to achieve semantic matching and compliance verification between the privacy compliance policies in the data capsule and the privacy compliance knowledge graph;
[0011] Call at least one privacy protection processing mechanism to perform compliance processing on the data asset according to the compliance verification result, and the privacy protection processing mechanism includes differential privacy perturbation mechanism, homomorphic encryption, and secure multi-party computation;
[0012] Before the data is externally released, determine the anonymization processing parameters through optimized calculation and perform data distortion processing;
[0013] Construct a dynamic evaluation model for the value of data assets, and evaluate the comprehensive value and risk level of data assets in real time to guide the differential management of data assets.
[0014] As a preferred solution of the present invention, the metadata tags of the data asset at least include data attribution, sensitivity level, processing method, usage scope, integrity, update frequency, accuracy, and relevance.
[0015] As a preferred solution of the present invention, the encapsulation of the data capsule includes: binding and storing the data asset content, privacy compliance policies, and data processing context information in a unified structure, and adopting a machine-readable formal policy language in the data capsule to achieve consistent management and automatic execution of the privacy compliance policies and the data asset life cycle;
[0016] The privacy compliance policy is optimized through a feedback reinforcement learning algorithm, including:
[0017] Construct a mapping library of historical privacy compliance policies and compliance verification results;
[0018] Define a policy optimization reward function R(π), and the expression is:
[0019] R(π) = α·C(π,R) + β·P(π,D) - γ·L(π,V);
[0020] In the formula, C(π,R) represents the matching degree score of the policy π and the regulatory requirement R, P(π,D) represents the data protection effect index after the implementation of the policy, L(π,V) represents the data value loss caused by the policy; α, β, γ are weight coefficients, and satisfy α + β + γ = 1;
[0021] Based on the policy optimization reward function, use the policy gradient or Q-learning method to optimize the policy version and improve the regulatory adaptability and data usability of the privacy compliance policy.
[0022] As a preferred embodiment of the present invention, the method for constructing a privacy compliance knowledge graph includes:
[0023] Collect laws, regulations, regulatory policies, and industry standards related to privacy protection, and construct a privacy compliance knowledge graph with regulatory clauses, terms, and logical relationships as node-edge structures;
[0024] Use a graph embedding algorithm to generate low-dimensional semantic feature vectors for the nodes of the knowledge graph;
[0025] Regularly and automatically update the content of the knowledge graph to reflect the latest regulatory changes, and record the update history to provide audit traceability support.
[0026] As a preferred embodiment of the present invention, the implementation method of semantic matching and compliance verification includes:
[0027] Adopt a retrieval-enhanced generation mechanism, call a large language model to semantically analyze the privacy compliance policy clauses in the data capsule, and extract key semantic features;
[0028] Based on the key semantic features, retrieve relevant regulatory clause nodes in the privacy compliance knowledge graph;
[0029] Use a graph embedding algorithm to represent the retrieved regulatory clause nodes and privacy compliance policy clauses as low-dimensional semantic vectors, and quantify the semantic matching degree between the policy clauses and the regulatory nodes through a similarity calculation method;
[0030] According to the quantification result of the semantic matching degree, output the compliance matching score for each privacy compliance policy clause, and automatically identify the policy clauses that are inconsistent with or insufficiently covered by the current regulations;
[0031] For the non-compliant policy clauses, generate amendment suggestions based on the semantics of the regulatory clauses, and provide data input for policy optimization learning.
[0032] As a preferred embodiment of the present invention, the differential privacy perturbation mechanism is applicable to a multi-level data processing network architecture, and the data processing network architecture includes edge devices, intermediate nodes, and a central server;
[0033] The differential privacy perturbation mechanism constructs and executes a differential privacy injection optimization model based on the trust models of each node in the multi-layer network, and is used to dynamically determine the injection level and injection method of differential privacy noise, so as to optimize the model training performance while satisfying the privacy budget constraint. Specifically, it includes:
[0034] 1) Obtain network structure parameters, including the number of layers of the processing system, the set of nodes in each layer, and their connection relationships;
[0035] 2) Establish a node trust model, evaluate the credibility of intermediate nodes, and generate trust labels or trust probability values for each node to distinguish differential privacy processing paths;
[0036] 3) Set differential privacy budget parameters, including the overall privacy budget (∈, δ) and the upper limit of budget allocation for each training round;
[0037] 4) Construct a joint optimization objective function with the model training error, privacy risk, and resource cost as optimization objectives. The expression is:
[0038]
[0039] In the formula, l t is the noise injection level, and σ t is the noise intensity; L model represents the upper bound of the model training error, R privacy represents the cumulative privacy budget consumption, and C resource represents the communication and computing resource cost; λ1, λ2, and λ3 are weight coefficients;
[0040] 5) Set optimization variables and constraint conditions. The optimization variables include the noise injection level, noise intensity, and the proportion of participating devices; the constraint conditions include privacy budget limitations, system resource upper limits, and the effectiveness requirements of the differential privacy mechanism;
[0041] 6) Use reinforcement learning, dynamic programming, or gradient-based optimization algorithms to solve the above model, generate a differential privacy injection strategy, and determine the injection nodes and corresponding noise intensity parameters for each round of training;
[0042] 7) Execute the injection strategy, which specifically includes: injecting calibration noise into trusted intermediate nodes, pre-injecting protective noise by child nodes in untrusted paths, and dynamically adjusting to achieve the co-optimal of privacy protection and model performance;
[0043] The differential privacy perturbation parameter, the current privacy risk level, and the policy execution effect are dynamically adjusted using an adaptive algorithm. The formula is:
[0044]
[0045] In the formula, ∈ t is the differentially private budget adjusted in real time, ∈0 is the benchmark budget value, R t is the currently estimated privacy leakage risk value, θ is the set risk threshold, and η is the adjustment coefficient.
[0046] As a preferred solution of the present invention, the anonymization processing parameters are determined through optimization calculations. The method includes:
[0047] Construct a data distortion convex optimization model that quantifies the privacy leakage risk using mutual information, expressed as:
[0048] min I(S; Y), s.t. D(X, Y) ≤ δ;
[0049] In the formula, I(S; Y) represents the mutual information between the sensitive data S and the released data Y, which is used to quantify the privacy leakage risk during data release; D(X, Y) is the data distortion degree between the original data X and the released data Y, and δ is the preset data distortion constraint threshold;
[0050] Solve the data distortion convex optimization model through a convex optimization algorithm to automatically determine the optimal anonymization processing parameters for privacy protection under the condition of meeting the data availability requirements.
[0051] As a preferred solution of the present invention, the data asset value dynamic evaluation model is constructed based on a reinforcement learning algorithm, with the integrity, accuracy, update frequency, relevance, and risk level of the data asset as the state space input, to predict the current comprehensive value score and future value trend of the data asset;
[0052] The reinforcement learning algorithm uses the value prediction error and risk assessment as the joint reward function, and optimizes the value evaluation accuracy through continuous training to achieve dynamic management support for different data assets in different regulatory scenarios.
[0053] As a preferred solution of the present invention, the method further includes: when derivative data is generated after processing the data asset, based on the static analysis and causal inference mechanism, automatically generate a residual privacy policy applicable to the derivative data, and bind and store the residual privacy policy with the derivative data to achieve the continuation and tracking management of the privacy compliance policy during the data usage process, specifically including:
[0054] Establish a directed graph model of the data processing process to clarify the influence path of each processing link on the privacy attributes;
[0055] Use static analysis to determine the influence of the data processing steps on the original privacy compliance policy, and use causal inference analysis to clarify the policy constraints applicable to the derivative data;
[0056] Bind and store the generated residual privacy policy with the derivative data in the form of metadata;
[0057] According to the compliance verification and value evaluation results, use the reinforcement learning algorithm to dynamically optimize the privacy compliance policy;
[0058] Record the audit log for the entire process of data processing and policy optimization, use the non-tamperable storage technology to ensure the security of the audit data, and provide the functions of the whole process audit tracking and retrospective analysis.
[0059] A market supervision data asset management system based on compliance and privacy protection, which is applied to the market supervision data asset management method based on compliance and privacy protection as described above. The system includes:
[0060] A data asset modeling module, which is used to collect, clean, and abstractly model heterogeneous data resources of regulatory objects to form data assets with metadata tags;
[0061] A data capsule generation module, which is used to generate privacy compliance policies based on metadata tags and encapsulate the privacy compliance policies and data assets into data capsules;
[0062] A knowledge graph and semantic matching module, which is used to construct a privacy compliance knowledge graph and use a large language model to perform semantic matching and compliance verification between the privacy compliance policies in the data capsules and the privacy compliance knowledge graph;
[0063] A privacy protection processing module, which is used to call a privacy protection processing mechanism to perform compliance processing on data assets according to the compliance verification results;
[0064] An anonymization publishing processing module, which is used to determine anonymization processing parameters through optimized calculation and implement data distortion processing before the data is externally published;
[0065] A value evaluation module, which is used to construct a dynamic evaluation model for the value of data assets and real-time evaluate the comprehensive value and risk level of data assets;
[0066] A management decision-making module, which is used to implement a differentiated management strategy for data assets according to the evaluation results;
[0067] An audit tracking module, which is used to optimize and adjust privacy compliance policies, record audit logs, and provide full-process audit tracking and retrospective analysis.
[0068] The beneficial effects of the present invention are as follows: By collecting, cleaning, and abstracting and modeling heterogeneous data resources, data assets with metadata tags are formed, realizing the structured and standardized processing of heterogeneous data of regulatory objects, enhancing the identifiability, manageability, and traceability of data assets, and laying a foundation for subsequent compliance strategy generation and processing; Privacy compliance policies are generated based on metadata tags and encapsulated into data capsules to realize the binding of data and compliance policies, ensuring that privacy protection requirements run through the data life cycle and enhancing the compliance and controllability of data use; By constructing a privacy compliance knowledge graph and combining it with a large language model for semantic matching and compliance verification, the semantic consistency and matching degree between privacy compliance policies and regulatory regulations are improved, the intelligent and automated level of compliance auditing is enhanced, and manual intervention is reduced; According to the results of compliance verification, a privacy protection processing mechanism is called to ensure that the data processing process complies with the requirements of privacy protection regulations, dynamically apply appropriate data protection technologies, and reduce the risk of data leakage; Before the data is externally released, the anonymization processing parameters are optimized, and data distortion processing is implemented to effectively control the risk of privacy leakage on the premise of ensuring data availability, and realize the security and compliance of external data release; A dynamic evaluation model for the value of data assets is constructed to evaluate the comprehensive value and risk level of data, support the refined and differentiated management of data assets, and realize compliance operation and strategy optimization driven by value. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Wherein:
[0071] Figure 1 is the method flow chart of the present invention;
[0072] Figure 2 is the system modular structure schematic diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.
[0074] Such as Figure 1As shown in the figure, this is an embodiment of the present invention, which provides a method for managing market supervision data assets based on compliance and privacy protection, including the following:
[0075] S1: Collect, clean, and abstractly model heterogeneous data resources of regulatory objects to form data assets with metadata tags.
[0076] Data of regulatory objects collected from different sources needs to be cleaned to remove incorrect or duplicate information, then a unified model is established and labeled with tags (metadata tags) describing its attributes. The metadata tags of data assets should at least include the following information:
[0077] Data ownership: The owner or source of the data;
[0078] Sensitivity level: The sensitivity of the data, such as public, internal, confidential, etc.;
[0079] Processing method: How the data should be processed, such as encryption, anonymization, etc.;
[0080] Scope of use: For which purposes or fields the data can be used;
[0081] Integrity: Whether the data is complete or missing;
[0082] Update frequency: The update cycle of the data, such as daily, weekly, etc.;
[0083] Accuracy: The accuracy level of the data;
[0084] Relevance: The degree of association of the data with a specific task or field.
[0085] S2: Generate privacy compliance policies based on metadata tags and encapsulate the privacy compliance policies and data assets into data capsules.
[0086] During the encapsulation process of the data capsule, the data asset content, privacy compliance policy, and data processing context information are bound and stored in a unified structure, and a machine-readable formal policy language is used inside the data capsule to achieve consistent management and automatic execution of the privacy compliance policy and the data asset life cycle;
[0087] The privacy compliance policy is optimized through a feedback enhanced learning algorithm, including:
[0088] Build a mapping library of historical privacy compliance policies and compliance verification results, recording past privacy compliance policies and their corresponding compliance verification results;
[0089] Define a policy optimization reward function R(π), comprehensively considering the matching degree of the policy with regulatory requirements, the data protection effect after policy implementation, and the data value loss caused by the policy. The expression is:
[0090] R(π) = α·C(π,R) + β·P(π,D) - γ·L(π,V);
[0091] Wherein, C(π,R) represents the matching degree score between the policy π and the regulatory requirement R, P(π,D) represents the data protection effect index after the implementation of the policy, and L(π,V) represents the data value loss caused by the policy; α, β, and γ are weight coefficients, and α + β + γ = 1;
[0092] Based on the policy optimization reward function, use the policy gradient or Q-learning method to optimize the policy version and improve the regulatory adaptability and data usability of the privacy compliance policy.
[0093] S3: Construct a privacy compliance knowledge graph, and use a large language model to implement semantic matching and compliance verification between the privacy compliance policy in the data capsule and the privacy compliance knowledge graph.
[0094] In one embodiment, the method for constructing a privacy compliance knowledge graph includes:
[0095] Collect laws, regulations, regulatory policies, and industry standards related to privacy protection, and construct a privacy compliance knowledge graph with regulatory clauses, terms, and logical relationships as the node-edge structure;
[0096] Adopt a graph embedding algorithm to generate low-dimensional semantic feature vectors of the knowledge graph nodes for easy computer processing;
[0097] Regularly and automatically update the content of the knowledge graph to reflect the latest regulatory changes, and record the update history to provide audit traceability support.
[0098] The implementation method of semantic matching and compliance verification includes:
[0099] Adopt a retrieval-augmented generation mechanism, call a large language model to perform semantic parsing on the privacy compliance policy clauses in the data capsule, and extract key semantic features (key contents that can reflect the intent, scope of application, object, operation restrictions, and obligation requirements of the privacy compliance policy clauses);
[0100] For example, the clause content of the privacy compliance policy is: For the positioning information of users, after obtaining their explicit consent, it should only be used for navigation services and should not be saved for more than 24 hours.
[0101] Then the key semantic features extracted after semantic parsing may include:
[0102] Data type: positioning information;
[0103] Operation behaviors: collection, use, save;
[0104] Prerequisite: Express consent;
[0105] Usage restriction: Navigation service;
[0106] Time limit: Not exceeding 24 hours;
[0107] Restriction clause: Not to be used for other purposes.
[0108] Based on key semantic features, retrieve relevant regulatory clause nodes in the privacy compliance knowledge graph (such as "Processing of personal information requires authorization", "Location data belongs to sensitive information");
[0109] Use the graph embedding algorithm to represent the retrieved regulatory clause nodes and privacy compliance policy clauses as low-dimensional semantic vectors, and quantify the semantic matching degree between the policy clauses and the regulatory nodes through a similarity calculation method (such as cosine similarity);
[0110] According to the quantification results of the semantic matching degree, output the compliance matching score for each privacy compliance policy clause, and automatically identify the policy clauses that are inconsistent with or insufficiently covered by the current regulations;
[0111] For the non-compliant policy clauses, generate amendment suggestions based on the semantics of the regulatory clauses, and provide data input for policy optimization learning.
[0112] S4: According to the compliance verification results, call at least one privacy protection processing mechanism (differential privacy perturbation mechanism, homomorphic encryption, secure multi-party computation, etc.) to perform compliance processing on the data assets.
[0113] Differential privacy perturbation: Prevent individual information from being identified by adding random noise to the data. For example, when counting the age distribution of users, add a small random value to each age data to ensure that the specific age cannot be determined.
[0114] Among them, the differential privacy perturbation mechanism is applicable to multi-level data processing network architectures, and the architectures include edge devices, intermediate nodes, and central servers;
[0115] The differential privacy perturbation mechanism constructs and executes a differential privacy injection optimization model based on the trust model of each node in the multi-layer network, and is used to dynamically determine the injection level and injection method of the differential privacy noise, so as to optimize the model training performance while satisfying the privacy budget constraint. Specifically, it includes:
[0116] 1) Obtain network structure parameters, including the number of layers of the processing system, the set of nodes in each layer and their connection relationships;
[0117] 2) Establish a node trust model, evaluate the credibility of the intermediate nodes, and generate trust labels or trust probability values for each node, which are used to distinguish the differential privacy processing paths;
[0118] 3) Set differential privacy budget parameters, including the overall privacy budget (∈, δ) and the upper limit of budget allocation for each training round;
[0119] 4) Construct a joint optimization objective function, with the model training error, privacy risk, and resource cost as the optimization objectives. The expression is:
[0120]
[0121] In the formula, l t is the noise injection level, and σ t is the noise intensity; L model represents the upper bound of the model training error, R privacy represents the cumulative privacy budget consumption, and C resource represents the communication and computing resource cost; λ1, λ2, and λ3 are weight coefficients;
[0122] 5) Set the optimization variables and constraints. The optimization variables include the noise injection level, noise intensity, and the proportion of participating devices; the constraints include privacy budget limits, system resource upper limits, and the effectiveness requirements of the differential privacy mechanism;
[0123] 6) Use reinforcement learning, dynamic programming, or gradient-based optimization algorithms to solve the above model, generate a differential privacy injection strategy, and determine the injection nodes and corresponding noise intensity parameters for each round of training;
[0124] 7) Execute the injection strategy, which specifically includes: injecting calibration noise in trusted intermediate nodes, pre-injecting protective noise by child nodes in untrusted paths, and dynamically adjusting to achieve the co-optimal of privacy protection and model performance;
[0125] The differential privacy perturbation parameters, the current privacy risk level, and the policy execution effect are dynamically adjusted using an adaptive algorithm. The formula is:
[0126]
[0127] In the formula, ∈ t is the differentially private budget adjusted in real time, ∈0 is the benchmark budget value, R t is the currently estimated privacy leakage risk value, θ is the set risk threshold, and η is the adjustment coefficient;
[0128] This mechanism realizes minimizing the model performance loss and system resource consumption while ensuring privacy compliance through joint optimization of differential privacy configuration and training performance.
[0129] Homomorphic encryption allows for performing calculations on data in an encrypted state, ensuring that the data remains encrypted throughout the processing. Example: Summing encrypted financial data results in an encrypted outcome, and only authorized parties can decrypt and view it.
[0130] Secure multi-party computation: Multiple participating parties jointly complete a computational task without revealing their respective data. Example: Multiple hospitals jointly count the total number of patients with a certain disease without sharing their individual patient lists.
[0131] S5: Before the data is externally released, determine the anonymization processing parameters through optimization calculations and perform data distortion processing.
[0132] Specifically, the method for determining the anonymization processing parameters through optimization calculations includes:
[0133] Construct a data distortion convex optimization model that quantifies the privacy leakage risk using mutual information, expressed as:
[0134]
[0135] In the formula, I(S; Y) represents the mutual information between the sensitive data S and the released data Y, which is used to quantify the privacy leakage risk during data release; D(X, Y) is the data distortion degree between the original data X and the released data Y, and δ is a preset data distortion constraint threshold;
[0136] Solve the data distortion convex optimization model through a convex optimization algorithm to automatically determine the optimal anonymization processing parameters for privacy protection under the condition of meeting the data availability requirements.
[0137] S6: Construct a dynamic evaluation model for the value of data assets to real-time evaluate the comprehensive value and risk level of data assets, so as to guide the differential management of data assets.
[0138] Furthermore, the dynamic evaluation model for the value of data assets is constructed based on a reinforcement learning algorithm, with the integrity, accuracy, update frequency, relevance, and risk level of data assets as the state space input to predict the current comprehensive value score and future value trend of data assets;
[0139] The reinforcement learning algorithm uses the value prediction error and risk assessment as a joint reward function, and optimizes the value evaluation accuracy through continuous training to achieve dynamic management support for different data assets under different regulatory scenarios.
[0140] In a preferred embodiment of the method of the present invention, the following is also included:
[0141] When derivative data is generated from processed data assets, based on static analysis and causal inference mechanisms, a residual privacy policy applicable to the derivative data is automatically generated and stored in a bound manner with the derivative data to achieve the continuation and tracking management of privacy compliance policies during the data usage process, specifically including:
[0142] Establish a directed graph model of the data processing process to clarify the impact path of each processing link on privacy attributes;
[0143] Use static analysis to determine the impact of data processing steps on the original privacy compliance policy, and use causal inference analysis to clarify the policy constraints applicable to the derivative data;
[0144] Store the generated residual privacy policy in a bound manner with the derivative data in the form of metadata;
[0145] According to the results of compliance verification and value evaluation, use reinforcement learning algorithms to dynamically optimize privacy compliance policies;
[0146] Record audit logs for the entire process of data processing and policy optimization, use tamper-proof storage technology to ensure the security of audit data, and provide functions for the entire process of audit tracking and retrospective analysis.
[0147] Through the above methods, it is possible to effectively manage data assets in market supervision, ensure data compliance and privacy protection, and at the same time improve the value and utilization efficiency of data.
[0148] Such as Figure 2 shown, which is another embodiment of the present invention. This embodiment provides a market supervision data asset management system based on compliance and privacy protection, which is applied to the market supervision data asset management method based on compliance and privacy protection as described above, including:
[0149] A data asset modeling module for collecting, cleaning, and abstractly modeling heterogeneous data resources of regulatory objects to form data assets with metadata tags;
[0150] A data capsule generation module for generating privacy compliance policies based on metadata tags and encapsulating the privacy compliance policies and data assets into data capsules;
[0151] A knowledge graph and semantic matching module for constructing a privacy compliance knowledge graph and using large language models to perform semantic matching and compliance verification between the privacy compliance policies in the data capsules and the privacy compliance knowledge graph;
[0152] A privacy protection processing module for calling a privacy protection processing mechanism to perform compliance processing on data assets according to the results of compliance verification;
[0153] An anonymization publishing processing module, which is used to determine anonymization processing parameters through optimized calculation and implement data distortion processing before the data is externally published;
[0154] A value evaluation module, which is used to build a dynamic evaluation model for the value of data assets and real-time evaluate the comprehensive value and risk level of data assets;
[0155] A management decision-making module, which is used to implement differentiated management strategies for data assets according to the evaluation results;
[0156] An audit tracking module, which is used for optimizing and adjusting privacy compliance policies, recording audit logs, and providing full-process audit tracking and retrospective analysis.
[0157] Through the data acquisition interface, heterogeneous raw data (such as transaction data, device logs, declaration records, etc.) of regulatory objects is called and input into the data asset modeling module for cleaning (missing value filling, format standardization) and abstraction modeling (forming a structured data representation), generating metadata tags (such as attribution, sensitivity level, etc.), and outputting structured data assets + metadata tags as the input of the data capsule generation module.
[0158] The data capsule generation module calls the privacy compliance policy library / rule engine, generates initial privacy compliance policies according to the tags, binds the data, policies, and context to form a "data capsule" object (structured container), and outputs standardized data capsules (including data subjects + compliance policies + meta-information) for semantic verification and subsequent privacy protection processing modules to call.
[0159] The knowledge graph and semantic matching module constructs / updates the privacy regulation knowledge graph, calls the large language model, semantically analyzes the policies, matches the corresponding regulation nodes in the knowledge graph, calculates the compliance matching degree, and outputs: verification results (compliance / inconsistent / non-covered), correction suggestions or prompts, and feeds back the semantic verification logs to the audit tracking module, provides the policy verification results to the privacy protection processing module, and decides whether to execute the protection mechanism.
[0160] The privacy protection processing module receives the compliance verification results and the original data capsule (including the data ontology and policies), selects mechanisms such as differential privacy perturbation, homomorphic encryption, or SMPC according to the inconsistent clauses, and outputs the intermediate data assets that have been compliant processed for further processing by the anonymization publishing processing module.
[0161] The anonymization publishing processing module constructs a distortion optimization model (such as mutual information constraint), solves the optimal anonymization parameters (such as k value, ε, etc.), applies data distortion transformation, and outputs the data version to be published (which has satisfied the trade-off between anonymity and usability); this data flows to the publishing channel, and its value feedback is tracked by the value evaluation module.
[0162] The value assessment module evaluates the current value of data based on metadata and risk level assessment data, uses reinforcement learning algorithms to predict future value trends, outputs a comprehensive value score and a risk-value mapping, and feeds them back to the management decision-making module as the basis for decision-making.
[0163] For different data assets, the management decision-making module formulates differentiated opening, storage, scheduling, or sharing strategies, updates the data asset lifecycle management plan, outputs policy suggestions or instructions, which may act on the input configuration in the data asset modeling stage.
[0164] The audit tracking module runs through multiple stages, constructs audit logs and stores them in an immutable block structure; outputs tracking reports and backtracking analysis interfaces.
[0165] In summary, the present invention realizes the structured and standardized representation of data assets by collecting, cleaning, and abstractly modeling heterogeneous data resources of regulatory objects, and classifying and managing them in combination with metadata tags, enhancing the unified management ability and controllability of data assets. The present invention introduces a privacy compliance strategy and an encapsulation mechanism for data assets, namely the "data capsule" technology, uses large language models for semantic parsing and compliance verification, and combines a privacy compliance knowledge graph to achieve automatic compliance review of the data usage process, significantly improving the automation level of privacy compliance management. By constructing a policy optimization reward function and dynamically optimizing the privacy compliance policy in combination with reinforcement learning methods, not only the regulatory adaptability of the policy is improved, but also the usage value of the data is maximally retained, achieving a balance between privacy protection and data availability.
[0166] The present invention uses a convex optimization model based on mutual information to dynamically determine data anonymization processing parameters, effectively reducing the risk of privacy leakage, and ensuring the security of data during sharing and publishing through mechanisms such as differential privacy and adaptive encryption. By constructing a reinforcement learning-driven dynamic evaluation model for the value of data assets, the comprehensive value and risk level of data assets are evaluated in real time, providing a scientific basis for the differentiated management and usage decisions of data assets in different regulatory scenarios.
[0167] The present invention supports the generation and binding of residual privacy policies for derivative data, and combines an immutable audit log mechanism to provide full-process compliance tracking and backtracking capabilities, enhancing the transparency and compliance guarantee of the data usage process. The proposed system solution integrates multiple functions such as data modeling, compliance policy generation, privacy processing, value assessment, management decision-making, and audit tracking in the form of functional modules, forming a closed-loop data asset compliance management system with good scalability and adaptability.
[0168] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0169] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0170] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for market supervision data asset management based on compliance and privacy protection, characterized in that The method includes: Collecting, cleaning, and abstractly modeling the heterogeneous data resources of the supervised objects to form data assets with metadata tags; generating privacy compliance policies based on the metadata tags, and encapsulating the privacy compliance policies and data assets into data capsules; constructing a privacy compliance knowledge graph, and using a large language model to achieve semantic matching and compliance verification between the privacy compliance policies in the data capsules and the privacy compliance knowledge graph; Invoking at least one privacy protection processing mechanism to perform compliance processing on the data assets according to the compliance verification results, where the privacy protection processing mechanisms include differential privacy perturbation mechanism, homomorphic encryption, and secure multi-party computation; Before the data is externally published, determining anonymization processing parameters through optimized calculation and performing data distortion processing; Constructing a dynamic evaluation model for the value of data assets to evaluate the comprehensive value and risk level of data assets in real time to guide the differential management of data assets.
2. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The metadata tags of the data assets at least include data ownership, sensitivity level, processing method, usage scope, integrity, update frequency, accuracy, and relevance.
3. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The encapsulation of the data capsule includes: binding and storing the data asset content, privacy compliance policies, and data processing context information in a unified structure, and adopting a machine-readable formal policy language in the data capsule to achieve the consistency management and automatic execution of privacy compliance policies and the data asset life cycle; The privacy compliance policy is optimized through a feedback enhanced learning algorithm, including: Constructing a mapping library of historical privacy compliance policies and compliance verification results; Defining a policy optimization reward function R(π), and the expression is: R(π) = α·C(π,R) + β·P(π,D) - γ·L(π,V); In the formula, C(π,R) represents the matching degree score of the policy π and the regulatory requirement R, P(π,D) represents the data protection effect index after the implementation of the policy, L(π,V) represents the data value loss caused by the policy; α, β, γ are weight coefficients and satisfy α + β + γ = 1; based on the policy optimization reward function, use the policy gradient or Q-learning method to optimize the policy version and improve the regulatory adaptability and data usability of the privacy compliance policy.
4. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The method for constructing the privacy compliance knowledge graph includes: Collecting laws, regulations, regulatory policies, and industry standards related to privacy protection, and constructing a privacy compliance knowledge graph with regulatory clauses, terms, and logical relationships as node-edge structures; Adopting a graph embedding algorithm to generate low-dimensional semantic feature vectors of the knowledge graph nodes; Automatically updating the content of the knowledge graph regularly to reflect the latest regulatory changes, and recording the update history to provide audit traceability support.
5. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The implementation method of the semantic matching and compliance verification includes: Adopting a retrieval enhanced generation mechanism, invoking a large language model to semantically analyze the privacy compliance policy clauses in the data capsule, and extracting key semantic features; Based on the key semantic features, retrieving relevant regulatory clause nodes in the privacy compliance knowledge graph; Using the graph embedding algorithm, the retrieved regulatory clause nodes and privacy compliance policy clauses are represented as low-dimensional semantic vectors, and the semantic matching degree between the policy clauses and the regulatory nodes is quantified through a similarity calculation method; According to the quantification result of the semantic matching degree, the compliance matching score of each privacy compliance policy clause is output, and the policy clauses inconsistent with or insufficiently covered by the current regulations are automatically identified; For the non-compliant policy clauses, amendment suggestions based on the semantics of the regulatory clauses are generated, and data input for policy optimization learning is provided.
6. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein, The differential privacy perturbation mechanism is applicable to a multi-level data processing network architecture, and the data processing network architecture includes edge devices, intermediate nodes, and a central server; The differential privacy perturbation mechanism constructs and executes a differential privacy injection optimization model based on the trust models of each node in the multi-layer network, for dynamically determining the injection level and injection method of differential privacy noise, so as to optimize the model training performance while satisfying the privacy budget constraint, specifically including: 1) Obtain network structure parameters, including the number of layers of the processing system, the set of nodes in each layer, and their connection relationships; 2) Establish a node trust model, evaluate the credibility of the intermediate nodes, and generate trust labels or trust probability values for each node, for distinguishing differential privacy processing paths; 3) Set differential privacy budget parameters, including the overall privacy budget (∈, δ) and the budget allocation upper limit for each training round; 4) Construct a joint optimization objective function, with the model training error, privacy risk, and resource cost as the optimization objectives, and the expression is: where \(l\) t is the noise injection level, and \(\sigma\) t is the noise intensity; \(L\) model represents the upper bound of the model training error, and \(R\) privacy represents the cumulative privacy budget consumption, and \(C\) resource represents the communication and computing resource cost; \(\lambda_1\), \(\lambda_2\), and \(\lambda_3\) are weight coefficients; 5) Set optimization variables and constraint conditions. The optimization variables include the noise injection level, noise intensity, and the proportion of participating devices; the constraint conditions include privacy budget limitations, system resource upper limits, and the effectiveness requirements of the differential privacy mechanism; 6) Use reinforcement learning, dynamic programming, or gradient-based optimization algorithms to solve the above model, generate a differential privacy injection strategy, and determine the injection nodes and corresponding noise intensity parameters for each round of training; 7) Execute the injection strategy, specifically including: injecting calibration noise in the trusted intermediate nodes, pre-injecting protective noise by the child nodes in the untrusted paths, and dynamically adjusting to achieve the co-optimal of privacy protection and model performance; The differential privacy perturbation parameters, the current privacy risk level, and the policy execution effect are dynamically adjusted using an adaptive algorithm, and the formula is: where ∈ t is the differentially private budget adjusted in real time, ∈0 is the baseline budget value, R t is the currently estimated privacy leakage risk value, θ is the set risk threshold, and η is the adjustment coefficient.
7. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The anonymization processing parameters are determined through optimization calculations, and the method includes: Construct a data distortion convex optimization model that quantifies the privacy leakage risk with mutual information, and is expressed as: min I(S; Y), s.t. D(X, Y) ≤ δ; In the formula, I(S; Y) represents the mutual information between the sensitive data S and the published data Y, which is used to quantify the privacy leakage risk when the data is published; D(X, Y) is the data distortion degree between the original data X and the published data Y, and δ is the preset data distortion constraint threshold; Solve the data distortion convex optimization model through a convex optimization algorithm, and automatically determine the optimal anonymization processing parameters for privacy protection under the condition of meeting the data availability requirements.
8. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The dynamic evaluation model of data asset value is constructed based on the reinforcement learning algorithm, taking the integrity, accuracy, update frequency, relevance, and risk level of data assets as the state space input to predict the current comprehensive value score and future value trend of data assets; The reinforcement learning algorithm uses the value prediction error and risk assessment as the joint reward function, and optimizes the value evaluation accuracy through continuous training to achieve dynamic management support for different data assets in different regulatory scenarios.
9. The method for managing market supervision data assets based on compliance and privacy protection according to claim 1, wherein The method further includes: when derivative data is generated after data assets are processed, based on the static analysis and causal inference mechanism, automatically generating a residual privacy policy applicable to the derivative data, and binding and storing the residual privacy policy with the derivative data to achieve the continuation and tracking management of privacy compliance policies during the data usage process, specifically including: Establishing a directed graph model of the data processing process to clarify the influence path of each processing link on privacy attributes; using static analysis to determine the influence of data processing steps on the original privacy compliance policy, and using causal inference analysis to clarify the policy constraints applicable to derivative data; Binding and storing the generated residual privacy policy with the derivative data in the form of metadata; Dynamically optimizing the privacy compliance policy using the reinforcement learning algorithm according to the compliance verification and value evaluation results; Recording audit logs for the entire process of data processing and policy optimization, using tamper-proof storage technology to ensure the security of audit data, and providing the functions of full-process audit tracking and retrospective analysis.
10. A market supervision data asset management system based on compliance and privacy protection, which is applied to the market supervision data asset management method based on compliance and privacy protection according to any one of claims 1-9, and is characterized in that, The system includes: a data asset modeling module for collecting, cleaning, and abstractly modeling heterogeneous data resources of regulatory objects to form data assets with metadata tags; A data capsule generation module for generating privacy compliance policies based on metadata tags and encapsulating the privacy compliance policies and data assets into data capsules; A knowledge graph and semantic matching module for constructing a privacy compliance knowledge graph and using a large language model to perform semantic matching and compliance verification between the privacy compliance policies in the data capsules and the privacy compliance knowledge graph; A privacy protection processing module for calling the privacy protection processing mechanism to perform compliance processing on data assets according to the compliance verification results; An anonymization publishing processing module for determining anonymization processing parameters through optimized calculation and implementing data distortion processing before the data is externally published; A value evaluation module for constructing a dynamic evaluation model of data asset value to real-time evaluate the comprehensive value and risk level of data assets; A management decision-making module for implementing differential management strategies for data assets according to the evaluation results; An audit tracking module for optimizing and adjusting privacy compliance policies, recording audit logs, and providing full-process audit tracking and retrospective analysis.
Citation Information
Patent Citations
Private data full life cycle protection method and system based on data platform
CN117972779A
Archive management method and system based on big data
CN118551414A
Federal learning-based privacy protection type large-scale model training and deployment method
CN118734360A
Method for constructing compliance risk control AI knowledge base field by using LLM
CN118886489A
Security risk assessment method and system for digital assets
CN119417612A
Cited By
Cross-domain agent knowledge migration and privacy barrier system
CN120671194A
Generative privacy data protection method and system based on GRL learning and adaptive ADAGAN
CN120744985A
Metadata-based multi-source data processing method, device and equipment
CN120872904A
Metropolitan area network privacy protection computing migration task scheduling optimization method
CN121037119A
A method for optimizing metask scheduling of privacy protection computing in a metropolitan area network
CN121037119B