Data desensitization access management method and device and medium

CN122885960APending Publication Date: 2026-10-09CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610984312.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0002]现有技术中,主要关注于识别数据中的敏感信息并进行脱敏,然而,该方法存在一个根本性的局限:它进行的是“一刀切”式的数据处理,忽略了数据价值与隐私保护之间的平衡,以及不同应用场景下对数据颗粒度的差异化需求

Benefits of technology

[0134]1.动态数据安全分级方法,基于自然语言模型动态分析文本数据上下文,生成一个动态安全评分,该评分与预定义的静态分级标签共同构成数据的增强安全属性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122885960A_ABST
    Figure CN122885960A_ABST
Patent Text Reader

Abstract

The application provides a data desensitization access management method and device and medium, and relates to the technical field of data. The method comprises the following steps: acquiring first storage data according to a data access request, wherein the first storage data comprises first original data and first sensitive classification information; identifying a first access user, a first access environment and a first access purpose according to the data access request, acquiring a first data reconstruction strategy according to the first access user, the first access environment, the first access purpose and the first sensitive classification information; and performing data reconstruction on the first original data according to the first data reconstruction strategy, so as to acquire first reconstructed data and output the first reconstructed data as a response to the data access request. According to the application, the original data is provided with preset sensitive classification information, in the access process, the reconstruction strategy of sensitive information is acquired according to the access user, the environment and the purpose, the data is reconstructed according to the reconstruction strategy to respond to the access, the data meeting the use purpose of the user is provided, and the leakage of sensitive information is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates at least to the field of data technology, and in particular to a data anonymization access management method, apparatus and medium. Background Technology

[0002] Existing technologies mainly focus on identifying sensitive information in data and desensitizing it. However, this method has a fundamental limitation: it performs "one-size-fits-all" data processing, ignoring the balance between data value and privacy protection, as well as the differentiated needs for data granularity in different application scenarios. Summary of the Invention

[0003] To address the aforementioned shortcomings, this application provides a data anonymization access management method, apparatus, and medium to solve the following technical problems: how to achieve a balance between data value and privacy protection, and the differentiated needs for data granularity in different application scenarios.

[0004] Firstly, this application provides a data anonymization access management method, the method comprising:

[0005] The first stored data is obtained according to the data access request. The first stored data includes the first raw data and its first sensitivity classification information.

[0006] Identify the first access user, the first access environment, and the first access purpose based on the data access request, and obtain the first data reconstruction strategy based on the first access user, the first access environment, the first access purpose, and the first sensitivity classification information;

[0007] The first original data is reconstructed according to the first data reconstruction strategy to obtain the first reconstructed data and output as a response to the data access request.

[0008] Furthermore, among which:

[0009] The first sensitivity classification information includes the first static sensitivity classification, the first dynamic sensitivity score, the first sensitive words and their contextual features;

[0010] The first user to access the data may be an internal or external person, the first environment may be an internal or external environment, and the purpose of the first access may be statistical analysis, content understanding, or direct access to data.

[0011] Furthermore, a first data reconstruction strategy is obtained based on the first accessing user, the first accessing environment, the first accessing purpose, and the first sensitivity classification information, specifically including:

[0012] Based on the first access user and the first static sensitivity classification, determine whether the first access user has access rights to the first raw data;

[0013] If the first accessing user has access rights to the first raw data, the reconstruction density of the first sensitive word and its contextual features is determined based on at least a portion of the first accessing user, the first access environment, the first access purpose, and the first dynamic sensitivity score.

[0014] Furthermore, based on at least a portion of the first accessing user, the first accessing environment, the first accessing purpose, and the first dynamic sensitivity score, the reconstruction density of the first sensitive word and its contextual features is determined, specifically including:

[0015] If the first user is an internal staff member and the purpose of the first visit is statistical analysis, then aggregate and statistically analyze the first sensitive word and its contextual features.

[0016] If the first user is an internal staff member and the purpose of the first visit is content understanding, then determine to extract the summary of the first sensitive word and its contextual features;

[0017] If the first access purpose is to directly access data, the first access environment is an internal environment, or the first dynamic sensitivity score is lower than the first threshold, then the entity generalization label will be used to replace the first sensitive word.

[0018] If the first access purpose is to directly access data, the first access environment is an external environment, and the first dynamic sensitivity score is lower than the first threshold, then the first sensitive word and its context features are subject to format-preserving encryption or homomorphic encryption.

[0019] Furthermore, the method also includes:

[0020] Raw data is obtained from the original data source, wherein the raw data is unstructured natural language data;

[0021] According to the preset data security classification and grading permission table, obtain the static sensitivity classification of the original data;

[0022] Using a pre-trained natural language model, sensitive words and their contextual features in the original data are identified, and a dynamic sensitivity score of the original data is obtained based on the sensitive words and their contextual features.

[0023] The static sensitivity classification, dynamic sensitivity score, sensitive words and their contextual features are used as sensitivity classification information. The original data and sensitivity classification information are combined to form the stored data.

[0024] Furthermore, among which:

[0025] The data security classification and grading permission table includes: an internal data security classification and grading permission table and a general data security classification and grading permission table;

[0026] The internal data security classification and hierarchical permission table sets different access permissions for different internal personnel roles;

[0027] The general data security classification and hierarchical permission table sets different access permissions for different external personnel roles;

[0028] The internal data security classification and grading permission table and the general data security classification and grading permission table each have different static sensitivity levels set according to different data types.

[0029] Further, based on the first accessing user and the first static sensitivity classification, it is determined whether the first accessing user has access rights to the first raw data, specifically including:

[0030] Identify whether the first accessing user is an internal or external user. Based on the internal or external user role, query the internal data security classification and grading permission table or the general data security classification and grading permission table. Based on the internal or general data security classification and grading permission table, determine whether the first accessing user has access rights to the first static sensitivity level of the first original data.

[0031] Further, the first original data is reconstructed according to the first data reconstruction strategy to obtain the first reconstructed data and output as a response to the data access request, specifically including:

[0032] Based on the reconstruction density of the first sensitive word and its context features, obtain the first reconstructed data after data reconstruction of the first original data;

[0033] After outputting the first reconstructed data as a response to the data access request, the feedback from the first accessing user to the response is obtained, and the natural language model is trained again based on the feedback.

[0034] Secondly, this application provides a computer device, the computer device including a processor and a memory, the memory storing a computer program, and when the processor runs the computer program stored in the memory, the processor executes the data desensitization access management method as described above.

[0035] Thirdly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the data desensitization access management method described above.

[0036] This application provides a data anonymization access management method, device, and medium. By pre-setting sensitive classification information on the original data, during the access process, a reconstruction strategy is obtained for sensitive information based on the accessing user, environment, and purpose. The data is reconstructed according to the reconstruction strategy to respond to the access. This provides data that meets the user's purpose while avoiding the leakage of sensitive information, achieving a balance between data value and privacy protection, as well as meeting the differentiated needs for data granularity in different application scenarios. Attached Figure Description

[0037] Figure 1 This is a flowchart of a data anonymization access management method according to an embodiment of this application;

[0038] Figure 2 This is an architectural diagram of a data anonymization access management device according to an embodiment of this application;

[0039] Figure 3 This is a flowchart of another data anonymization access management method according to an embodiment of this application;

[0040] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of this application;

[0041] Figure 5 This is a schematic diagram of the structure of a data de-identification access management device according to an embodiment of this application;

[0042] Figure 6 This is a schematic diagram of the structure of a computer-readable storage medium according to an embodiment of this application. Detailed Implementation

[0043] To enable those skilled in the art to better understand the technical solution of this application, the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0044] It is understood that the specific embodiments and accompanying drawings described herein are merely for explaining this application and are not intended to limit this application.

[0045] It is understood that, without conflict, the various embodiments and features in the embodiments of this application can be combined with each other.

[0046] It is understood that, for ease of description, only the parts relevant to this application are shown in the accompanying drawings, while parts unrelated to this application are not shown in the drawings.

[0047] It is understood that each module or unit involved in the embodiments of this application may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple modules or units may be integrated into one entity structure.

[0048] It is understood that, without conflict, the functions and steps indicated in the flowcharts and block diagrams of this application may occur in a different order than that indicated in the accompanying drawings.

[0049] It is understood that the flowcharts and block diagrams of this application illustrate the possible architecture, functions, and operations of systems, apparatuses, devices, and methods according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, unit, program segment, or code, containing executable instructions for implementing the specified function. Furthermore, each block or combination of blocks in the block diagrams and flowcharts may be implemented using a hardware-based device to implement the specified function, or using a combination of hardware and computer instructions.

[0050] It is understood that the modules and units involved in the embodiments of this application can be implemented by software or by hardware. For example, the modules and units can be located in the processor.

[0051] Example 1:

[0052] like Figure 1 As shown, this application provides a data anonymization access management method, the method comprising:

[0053] S1. Obtain first stored data according to the data access request. The first stored data includes first raw data and its first sensitivity classification information.

[0054] S2. Identify the first access user, the first access environment, and the first access purpose based on the data access request, and obtain the first data reconstruction strategy based on the first access user, the first access environment, the first access purpose, and the first sensitivity classification information;

[0055] S3. Reconstruct the first original data according to the first data reconstruction strategy to obtain the first reconstructed data and output it as a response to the data access request.

[0056] In this embodiment, the provided method, by pre-setting sensitive classification information for the original data, obtains a reconstruction strategy for sensitive information based on the accessing user, environment, and purpose during the access process, and reconstructs the data to respond to the access based on the reconstruction strategy. This not only provides data that meets the user's purpose, but also avoids the leakage of sensitive information, achieving a balance between data value and privacy protection, as well as meeting the differentiated needs for data granularity in different application scenarios.

[0057] Specifically, traditional data access control systems, such as RBAC (Role-Based Access Control) systems, while providing granular access control based on roles and attributes, are designed based on "allow / deny" access to structured database records. This paradigm cannot be directly transferred and applied to unstructured natural language data and large-scale model training processes. This is because large models process continuous, high-dimensional text features, rather than discrete database records. Traditional access control cannot solve the technical challenge of "how to generate versions with different 'information densities' for different roles of the same text content."

[0058] To address the shortcomings of existing technologies, this embodiment proposes a large-scale model optimization method based on sensitive data classification and access control. By meticulously classifying and classifying data, combined with a flexible access control mechanism, a balance is achieved between data sensitivity and model performance, protecting privacy during the training and inference processes of the large model. This not only improves the security of data processing but also enhances the reliability of the large model when handling sensitive information.

[0059] More specifically, the following drawbacks exist in the application of large models:

[0060] The conflict between data utility and privacy protection: Existing data privacy protection technologies for large models cannot achieve a dynamic and refined balance between data utility and privacy protection. Either excessive desensitization impairs model performance, or insufficient protection leads to information leakage.

[0061] Insufficient granularity and scenario adaptability of access control: Traditional database access control systems (such as RBAC) cannot be applied to unstructured text data and large model training processes, and cannot generate differentiated and granular data versions of the same text content based on user roles and context.

[0062] The mismatch between static processing and dynamic environment: Most existing solutions are static rules, which cannot adapt and adjust according to the actual output effect of the model and user behavior, making it difficult to cope with the ever-evolving privacy risks and complex business scenarios.

[0063] Based on the above-mentioned shortcomings, this embodiment aims to solve the following technical problems:

[0064] How can we implement a dynamic, fine-grained data security classification method for large model training data, so that it can not only identify sensitive data types, but also assess their sensitivity in different contexts?

[0065] How can we deeply integrate user roles and access control policies with the data processing flow of large-scale language models, and design a method that can dynamically generate training data or inference inputs with different "information densities" or "knowledge concentrations" based on permissions?

[0066] How can we establish a closed-loop feedback optimization mechanism that enables the data processing strategy of large models to adaptively adjust and iterate based on the security of model outputs and user access logs, thereby achieving continuous security optimization?

[0067] The overall system architecture and data flow corresponding to the method in this embodiment are as follows: Figure 2 As shown, the system first obtains raw data from the original data source. The raw data is input into the dynamic data classification module to obtain sensitive classification information combining dynamic and static data. The context-aware permission adjudicator generates a differentiated data reconstruction strategy based on the sensitive classification information and the data application scenario. The differentiated data reconstruction engine executes the differentiated data reconstruction according to the strategy of the context-aware permission adjudicator. Response data is obtained through large model training / inference. User feedback is obtained through the feedback and optimization loop, and the large model is optimized based on the user feedback.

[0068] In one embodiment, wherein:

[0069] The first sensitivity classification information includes the first static sensitivity classification, the first dynamic sensitivity score, the first sensitive words and their contextual features;

[0070] The first user to access the data may be an internal or external person, the first environment may be an internal or external environment, and the purpose of the first access may be statistical analysis, content understanding, or direct access to data.

[0071] In this embodiment, the detailed process of dynamic data classification and permission adjudication is as follows: Figure 3 As shown, the process includes: the system receiving data access requests; obtaining static classification and dynamic security scores of user access data by querying an enhanced knowledge base; the operation of the permission adjudicator including: inputting user roles, attributes, data classification, dynamic scores, access context, etc., executing the policy model reasoning process, and outputting data reconstruction instructions; the data reconstruction engine selecting a processing path based on the instructions, including: if the instruction is to provide a statistical view, executing a score-based privacy aggregation algorithm; if the instruction is semantic desensitization, calling an entity generalization or text rewriting model; if the instruction is complete masking, applying strong encryption or masking techniques; and outputting securely reconstructed data.

[0072] In one embodiment, a first data reconstruction strategy is obtained based on the first accessing user, the first accessing environment, the first accessing purpose, and the first sensitivity classification information, specifically including:

[0073] Based on the first access user and the first static sensitivity classification, determine whether the first access user has access rights to the first raw data;

[0074] If the first accessing user has access to the first raw data, the reconstruction density of the first sensitive word and its contextual features is determined based on at least a portion of the first accessing user, the first access environment, the first access purpose, and the first dynamic sensitivity score.

[0075] In this embodiment, the user's access rights to the data are first confirmed based on the accessing user and static sensitivity classification information. At this time, an "allow / deny" access instruction can be output. Based on the allow access instruction, a data reconstruction instruction is further executed. This can achieve more granular sensitive data access permission management, and the user's access rights to the data can be appropriately relaxed compared to the existing static access rules.

[0076] In one embodiment, the reconstruction density of the first sensitive word and its contextual features is determined based on at least a portion of the first accessing user, the first accessing environment, the first accessing purpose, and the first dynamic sensitivity score, specifically including:

[0077] If the first user is an internal staff member and the purpose of the first visit is statistical analysis, then aggregate and statistically analyze the first sensitive word and its contextual features.

[0078] If the first user is an internal staff member and the purpose of the first visit is content understanding, then determine to extract the summary of the first sensitive word and its contextual features;

[0079] If the first access purpose is to directly access data, the first access environment is an internal environment, or the first dynamic sensitivity score is lower than the first threshold, then the entity generalization label will be used to replace the first sensitive word.

[0080] If the first access purpose is to directly access data, the first access environment is an external environment, and the first dynamic sensitivity score is lower than the first threshold, then the first sensitive word and its context features are to be encrypted with format preservation or homomorphic encryption.

[0081] In this embodiment, differentiated data reconstruction instructions can generate different information density versions suitable for user permissions, including but not limited to: data anonymization for statistical analysis within the enterprise: for example, when a manager requests departmental salary data for analysis, the engine provides aggregated statistical information; data anonymization for content understanding within the enterprise: when ordinary employees access documents containing sensitive items, the engine uses text summarization technology to generate a version that retains the core business logic but hides key financial figures; semantically preserved anonymization: for identified sensitive entities (such as names and locations), instead of using fixed [MASK], generalized tags (such as [PERSON-1], [LOCATION-A]) are used for replacement; context-based local encryption: for highly sensitive fragments, format-preserving encryption (FPE) or homomorphic encryption is used.

[0082] In one embodiment, the method further includes:

[0083] Raw data is obtained from the original data source, wherein the raw data is unstructured natural language data;

[0084] According to the preset data security classification and grading permission table, obtain the static sensitivity classification of the original data;

[0085] Using a pre-trained natural language model, sensitive words and their contextual features in the original data are identified, and a dynamic sensitivity score of the original data is obtained based on the sensitive words and their contextual features.

[0086] The static sensitivity classification, dynamic sensitivity score, sensitive words and their contextual features are used as sensitivity classification information. The original data and sensitivity classification information are combined to form the stored data.

[0087] In this embodiment, dynamic security assessment at the text fragment level is implemented, targeting unstructured natural language data and large model training processes. While obtaining static classification, a fine-tuned natural language model is used to analyze the context of the data content itself, generate dynamic security scores based on the context, and construct an enhanced knowledge base to form records in the format of <data identifier, static level, dynamic score, context features>.

[0088] In one embodiment, wherein:

[0089] The data security classification and grading permission table includes: an internal data security classification and grading permission table and a general data security classification and grading permission table;

[0090] The internal data security classification and hierarchical permission table sets different access permissions for different internal personnel roles;

[0091] The general data security classification and hierarchical permission table sets different access permissions for different external personnel roles;

[0092] The internal data security classification and grading permission table and the general data security classification and grading permission table each have different static sensitivity levels set according to different data types.

[0093] In this embodiment, examples of two types of permission tables are provided as follows:

[0094] Table 1 Internal Data Security Classification and Hierarchical Permission Table

[0095]

[0096] Table 2 General Data Security Classification and Hierarchical Permission Table

[0097]

[0098] In one embodiment, determining whether the first accessing user has access rights to the first raw data based on the first accessing user and the first static sensitivity classification specifically includes:

[0099] Identify whether the first accessing user is an internal or external user. Based on the internal or external user role, query the internal data security classification and grading permission table or the general data security classification and grading permission table. Based on the internal or general data security classification and grading permission table, determine whether the first accessing user has access rights to the first static sensitivity level of the first original data.

[0100] In this embodiment, the static classification of data sensitivity is obtained according to Table 1 or Table 2, and the user role's access permissions to the data are determined.

[0101] In one embodiment, the first original data is reconstructed according to a first data reconstruction strategy to obtain first reconstructed data and output as a response to the data access request, specifically including:

[0102] Based on the reconstruction density of the first sensitive word and its context features, obtain the first reconstructed data after data reconstruction of the first original data;

[0103] After outputting the first reconstructed data as a response to the data access request, the feedback from the first accessing user to the response is obtained, and the natural language model is trained again based on the feedback.

[0104] In this embodiment, the trained model can generate output that conforms to security specifications. The model output is monitored to see if it still contains sensitive information or if business users frequently request permissions due to insufficient information. Based on the feedback, the model is retrained so that it can learn which data reconstruction strategy can best balance security and utility in which situation.

[0105] Example 1:

[0106] This example provides a large-scale data optimization method for enterprises based on dynamic hierarchical and context-aware access control. The corresponding system mainly includes a dynamic hierarchical module, a context-aware access adjudicator, a data reconstruction engine, and a feedback learning loop. The corresponding method includes the following steps:

[0107] Step 1: Context-Aware Dynamic Data Classification

[0108] This step is not a simple static classification at the "table-field" level, but rather a dynamic security assessment at the text fragment level.

[0109] 1. Initial classification: First, the structured metadata (such as database field names) is initially tagged based on a predefined classification system (such as Table 1 and Table 2).

[0110] 2. Dynamic Risk Assessment: Using fine-tuned natural language models (such as variants of BERT), the model analyzes the context of the data content itself and generates dynamic safety scores based on the context. For example, the word "salary" in a text appears in the context of "the company's average salary is...", which has lower sensitivity than when it appears in the context of "employee Zhang San's salary is...". The model assigns dynamic score labels to data fragments based on the context.

[0111] 3. Construct an enhanced knowledge base: Store the initial hierarchical labels and dynamic security scores together in an enhanced knowledge base, forming records in the format of <data identifier, static level, dynamic score, contextual features>.

[0112] Step 2: Context-Aware Access Control and Data Restructuring

[0113] This transforms access control from "whether to allow access" to "what form of information is allowed to be accessed".

[0114] 1. Role and Attribute Definition: Similar to traditional RBAC / ABAC, define user roles (such as managers, ordinary employees) and attributes (departments, job levels).

[0115] 2. Context-Aware Adjudication: The permission adjudicator makes decisions based on roles, data static levels, dynamic security scores, and access context (such as time, purpose of access, etc.). The output is a data reconstruction instruction, not just "allow / deny".

[0116] 3. Differentiated Data Reconstruction: The data reconstruction engine generates different information density versions suitable for user permissions based on the adjudication instructions.

[0117] For enterprise-customized models, the goal is to maximize data value to serve the business. Simply anonymizing all salary data could prevent management from managing teams and HR from conducting compensation analysis. Current technology lacks a mechanism to dynamically control data sensitivity based on role and context while preserving the semantic utility of data (such as for statistical analysis).

[0118] For general-purpose large-scale models, the goal is to generate safe and compliant content. However, de-identification methods based on keyword filtering or simple classification can disrupt the coherence and logical integrity of the text, thereby affecting the quality of training data and the usability of generated content. For example, de-identifying all key entities in a text involving historical events can prevent the model from learning effective causal relationships.

[0119] Therefore, this example employs different data reconstruction strategies for these two models, for example:

[0120] 1) For internal enterprise models:

[0121] Data anonymization for statistical analysis: When a manager requests departmental salary data for analysis, the engine provides aggregated statistical information (such as departmental averages and quantiles) rather than individual details.

[0122] Content-understanding-based data anonymization: When ordinary employees access documents containing sensitive items, the engine uses text summarization technology to generate a version that retains the core business logic but hides key financial figures.

[0123] 2) For general large models:

[0124] Semantic-preserving desensitization: For identified sensitive entities (such as names and locations), instead of using fixed [MASK], use generalized tags (such as [PERSON-1], [LOCATION-A]) to replace them, and maintain consistency throughout the document.

[0125] Context-based local encryption: For highly sensitive fragments, format-preserving encryption (FPE) or homomorphic encryption is used, so that the model can still process its encrypted form, but cannot understand its plaintext meaning.

[0126] Step 3: Model Training and Output Generation

[0127] The datasets with different "information densities" processed in step two are used for incremental training or fine-tuning of large models. For general-purpose models, data that has been semantically anonymized is used; for enterprise models, dynamically reconstructed data incorporating access control policies is used. Thus, the trained model can generate outputs that comply with security standards.

[0128] Step 4: Closed-loop feedback and adaptive optimization

[0129] 1. Log Auditing and Anomaly Detection: Record all data access and model generation behaviors. Use unsupervised learning algorithms (such as Isolation Forest) to analyze user behavior sequences and detect abnormal patterns (such as large-scale access to sensitive data outside of working hours).

[0130] 2. Strategy effectiveness evaluation and reinforcement learning:

[0131] 1) Collect feedback: whether the monitoring model output still contains sensitive information, or whether business users frequently request permissions due to insufficient information.

[0132] 2) Optimize the strategy model: Combine the above feedback with the corresponding user roles, data, context, and reconstruction strategies to form training samples. Through a reinforcement learning framework, dynamically adjust the parameters of the strategy model in step two so that it can learn which data reconstruction strategy to use in which situation can optimally balance security and utility.

[0133] This example achieves the following technical effects:

[0134] 1. A dynamic data security classification method is based on natural language models to dynamically analyze the context of text data and generate a dynamic security score. This score, together with predefined static classification labels, constitutes the enhanced security attributes of the data.

[0135] 2. Context-aware data permission adjudication and reconstruction method: The permission adjudicator outputs reconstruction instructions for large model training data. The system includes a data reconstruction engine, which executes various differentiated reconstruction algorithms, including data aggregation, semantically preserved entity replacement / generalization, and text summarization, according to the adjudication instructions, to generate data versions with different information densities for users with different permissions.

[0136] 3. A closed-loop feedback optimization mechanism for large model training: The system collects user behavior data through log auditing and anomaly detection algorithms, and uses the model output security and user business feedback as reward signals. It uses reinforcement learning technology to continuously optimize the policy model in the permission arbitrator to achieve adaptive adjustment of data security policies.

[0137] 4. Starting from the application domain of the large model, the data is first classified according to the specified data level. Then, different security considerations are made for the internal large model and the general large model, focusing on specific business scenarios and data processing. For the internal large model, data is bound to permissions, and feedback is given in combination with the output of the large model. For the general large model, the data is then subject to content review and filtering to ensure that insecure content is not learned by the large model.

[0138] Example 2:

[0139] This example provides alternatives to Example 1, including:

[0140] 1. Alternative solutions for the "Dynamic Data Security Classification Module":

[0141] An ensemble learning and rule engine combination solution: This solution enables dynamic scoring without relying on a single pre-trained language model. It employs an ensemble learning framework that combines multiple lightweight classifiers specializing in different sensitive dimensions (such as personal identity, financial information, and health data) for parallel inference. A configurable rule engine then weights, fuses, and post-processes the outputs of each classifier to generate a dynamic security score. This solution improves the interpretability and adjustability of the module while maintaining evaluation accuracy.

[0142] A solution based on knowledge graphs and graph neural networks: For data with rich internal relationships within an enterprise, an enterprise knowledge graph can be constructed. Graph neural networks can then be used to analyze the strength and path of relationships between data entities (nodes) to assess the propagation and aggregation effects of their sensitivity. For example, node data directly associated with core personnel or projects will have a higher dynamic security score than indirectly associated data. This solution is particularly suitable for data environments with high levels of structure and complex relationships.

[0143] 2. Alternatives to the "Data Restructuring Engine":

[0144] A differential privacy-based statistical information generation scheme: When generating differentiated data views for different roles (such as providing statistical information for managers), differential privacy technology can be introduced in addition to traditional aggregation calculations. By adding strictly mathematically defined random noise that conforms to the differential privacy budget to the aggregation results, highly available data insights can be provided while mathematically preventing the possibility of individual information being restored, offering another theoretically solid technical path for the "utility-privacy balance".

[0145] A semantic rewriting scheme based on controllable text generation: For text anonymization that requires preserving semantics while hiding details, in addition to label replacement, a controllable text generation model can be used. This model takes the original text and a set of control attributes (such as "hiding specific names" and "generalizing monetary values") as conditions, and directly generates new text that is semantically coherent but has undergone anonymization rewriting. For example, "Zhang San's monthly salary is X yuan" can be rewritten as "The senior engineer's salary is at a high level in the market." This scheme can produce more natural and higher-quality anonymized text, which is beneficial for maintaining the language quality of downstream large-scale model training data.

[0146] 3. Alternative solutions to the "feedback optimization loop":

[0147] An online adaptive solution based on contextual multi-armed slot machines: This approach replaces offline reinforcement learning based on historical logs with an online learning paradigm using contextual multi-armed slot machines. The system treats the contextual features (user, data, environment) of each data access request as the state, and considers various selectable data reconstruction strategies as different "arms." Based on the immediate feedback obtained after each strategy application (such as successful operation, security alerts, and secondary user requests), the system updates the profit estimates of each strategy in real time, achieving fast and lightweight online adaptive optimization. This is particularly suitable for scenarios in the early stages of strategy exploration or where the environment changes rapidly.

[0148] Bayesian optimization-based hyperparameter tuning scheme: For optimization processes centered on adjusting policy model parameters, Bayesian optimization methods can be employed. The performance of the policy model (such as the combined security and utility scores) is treated as a black-box function. By intelligently selecting and experimenting with different parameter combinations, the optimal or near-optimal parameter configuration is found with as few iterations as possible. This method is efficient and suitable for scenarios with continuous parameter spaces but high evaluation costs.

[0149] Alternatively, when classifying data using a predefined data classification and grading table, predefined rule sets (such as those based on keywords or regular expressions) can be used for sensitivity labeling and classification, replacing machine learning methods. This method saves computational resources, but the labeling results are not flexible enough, limited to a given rule set. During the review and feedback phases, reinforcement learning algorithms can be used to train an automated review system, optimizing sensitive information detection and content filtering based on user access behavior and model output.

[0150] This example addresses the challenge of balancing data utility and privacy in large-scale model training and applications through dynamic, fine-grained data security processing and permission-aware mechanisms. It does not rely on any specific technical tool but provides a flexible technical framework, validating the feasibility and broad applicability of the solution from different technical perspectives.

[0151] Example 2:

[0152] like Figure 4 As shown, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes the data desensitization access management method as described in Embodiment 1.

[0153] The memory is connected to the processor. The memory can be flash memory, read-only memory or other types of memory. The processor can be a central processing unit or a microcontroller.

[0154] like Figure 5As shown, this computer device is specifically a data anonymization access management device, which includes:

[0155] Data acquisition module 1 is used to acquire first stored data according to a data access request. The first stored data includes first raw data and its first sensitivity classification information.

[0156] The reconstruction strategy module 2, connected to the data acquisition module 1, is used to identify the first access user, the first access environment, and the first access purpose according to the data access request, and to acquire the first data reconstruction strategy according to the first access user, the first access environment, the first access purpose, and the first sensitivity classification information.

[0157] The reconstruction response module 3, connected to the reconstruction strategy module 2, is used to reconstruct the first original data according to the first data reconstruction strategy, so as to obtain the first reconstructed data and output it as a response to the data access request.

[0158] In one embodiment, wherein:

[0159] The first sensitivity classification information includes the first static sensitivity classification, the first dynamic sensitivity score, the first sensitive words and their contextual features;

[0160] The first user to access the data may be an internal or external person, the first environment may be an internal or external environment, and the purpose of the first access may be statistical analysis, content understanding, or direct access to data.

[0161] In one embodiment, the reconstruction strategy module 2 specifically includes:

[0162] The permission determination unit is used to determine whether the first access user has access rights to the first original data based on the first access user and the first static sensitivity level.

[0163] A reconstruction density determination unit, connected to an access control unit, is used to determine the reconstruction density of a first sensitive word and its contextual features based on at least one of the first access user, the first access environment, the first access purpose, and the first dynamic sensitivity score, if the first access user has access rights to the first original data.

[0164] In one embodiment, determining the reconstruction density cell specifically includes:

[0165] The aggregation and statistics unit is used to determine the aggregation and statistics of the first sensitive word and its contextual features if the first accessing user is an internal person and the purpose of the first access is statistical analysis.

[0166] The summary extraction unit is used to determine the first sensitive word and its contextual features for summary extraction if the first accessing user is an internal person and the purpose of the first access is content understanding.

[0167] The tag replacement unit is used to determine to replace the first sensitive word with an entity generalization tag if the first access purpose is to directly access data, the first access environment is an internal environment, or the first dynamic sensitivity score is lower than the first threshold.

[0168] The encryption protection unit is used to determine whether to perform format-preserving encryption or homomorphic encryption on the first sensitive word and its context features if the first access purpose is to directly access data, the first access environment is an external environment, and the first dynamic sensitivity score is lower than the first threshold.

[0169] In one embodiment, the apparatus further includes a data storage unit for:

[0170] Raw data is obtained from the original data source, wherein the raw data is unstructured natural language data;

[0171] According to the preset data security classification and grading permission table, obtain the static sensitivity classification of the original data;

[0172] Using a pre-trained natural language model, sensitive words and their contextual features in the original data are identified, and a dynamic sensitivity score of the original data is obtained based on the sensitive words and their contextual features.

[0173] The static sensitivity classification, dynamic sensitivity score, sensitive words and their contextual features are used as sensitivity classification information. The original data and sensitivity classification information are combined to form the stored data.

[0174] In one embodiment, wherein:

[0175] The data security classification and grading permission table includes: an internal data security classification and grading permission table and a general data security classification and grading permission table;

[0176] The internal data security classification and hierarchical permission table sets different access permissions for different internal personnel roles;

[0177] The general data security classification and hierarchical permission table sets different access permissions for different external personnel roles;

[0178] The internal data security classification and grading permission table and the general data security classification and grading permission table each have different static sensitivity levels set according to different data types.

[0179] In one embodiment, the permission determination unit is specifically used for:

[0180] Identify whether the first accessing user is an internal or external user. Based on the internal or external user role, query the internal data security classification and grading permission table or the general data security classification and grading permission table. Based on the internal or general data security classification and grading permission table, determine whether the first accessing user has access rights to the first static sensitivity level of the first original data.

[0181] In one embodiment, the reconstructed response module 3 specifically includes:

[0182] The reconstruction unit is used to obtain the first reconstructed data after data reconstruction of the first original data based on the reconstruction density of the first sensitive word and its context features;

[0183] The feedback unit, connected to the reconstruction unit, is used to output the first reconstructed data as a response to the data access request, obtain the feedback from the first accessing user to the response, and then retrain the natural language model based on the feedback.

[0184] Example 3:

[0185] like Figure 6 As shown, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the data desensitization access management method as described in Embodiment 1.

[0186] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program units, or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), DVD or other optical disc storage, cartridges, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.

[0187] Embodiments 1-3 of this application provide a data anonymization access management method, device, and medium. By pre-setting sensitive classification information for the original data, during the access process, a reconstruction strategy for sensitive information is obtained based on the accessing user, environment, and purpose. The data is reconstructed according to the reconstruction strategy to respond to the access, which not only provides data that meets the user's purpose but also avoids the leakage of sensitive information, achieving a balance between data value and privacy protection, as well as meeting the differentiated needs for data granularity in different application scenarios.

[0188] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of this application, and this application is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this application, and these modifications and improvements are also considered to be within the scope of protection of this application.

Claims

1. A data anonymization access management method, characterized in that, The method includes: The first stored data is obtained according to the data access request. The first stored data includes the first raw data and its first sensitivity classification information. Identify the first access user, the first access environment, and the first access purpose based on the data access request, and obtain the first data reconstruction strategy based on the first access user, the first access environment, the first access purpose, and the first sensitivity classification information; The first original data is reconstructed according to the first data reconstruction strategy to obtain the first reconstructed data and output as a response to the data access request.

2. The method according to claim 1, characterized in that, in: The first sensitivity classification information includes the first static sensitivity classification, the first dynamic sensitivity score, the first sensitive words and their contextual features; The first user to access the data may be an internal or external person, the first environment may be an internal or external environment, and the purpose of the first access may be statistical analysis, content understanding, or direct access to data.

3. The method according to claim 2, characterized in that, The first data reconstruction strategy is obtained based on the first accessing user, the first accessing environment, the first accessing purpose, and the first sensitivity classification information, specifically including: Based on the first access user and the first static sensitivity classification, determine whether the first access user has access rights to the first raw data; If the first accessing user has access to the first raw data, the reconstruction density of the first sensitive word and its contextual features is determined based on at least a portion of the first accessing user, the first access environment, the first access purpose, and the first dynamic sensitivity score.

4. The method according to claim 3, characterized in that, Based on at least a portion of the first accessing user, the first accessing environment, the first accessing purpose, and the first dynamic sensitivity score, determine the reconstruction density of the first sensitive word and its contextual features, specifically including: If the first user is an internal staff member and the purpose of the first visit is statistical analysis, then aggregate and statistically analyze the first sensitive word and its contextual features. If the first user is an internal staff member and the purpose of the first visit is content understanding, then determine to extract the summary of the first sensitive word and its contextual features; If the first access purpose is to directly access data, the first access environment is an internal environment, or the first dynamic sensitivity score is lower than the first threshold, then the entity generalization label will be used to replace the first sensitive word. If the first access purpose is to directly access data, the first access environment is an external environment, and the first dynamic sensitivity score is lower than the first threshold, then the first sensitive word and its context features are to be encrypted with format preservation or homomorphic encryption.

5. The method according to claim 3, characterized in that, The method further includes: Raw data is obtained from the original data source, wherein the raw data is unstructured natural language data; According to the preset data security classification and grading permission table, obtain the static sensitivity classification of the original data; Using a pre-trained natural language model, sensitive words and their contextual features in the original data are identified, and a dynamic sensitivity score of the original data is obtained based on the sensitive words and their contextual features. The static sensitivity classification, dynamic sensitivity score, sensitive words and their contextual features are used as sensitivity classification information. The original data and sensitivity classification information are combined to form the stored data.

6. The method according to claim 5, characterized in that, in: The data security classification and grading permission table includes: an internal data security classification and grading permission table and a general data security classification and grading permission table; The internal data security classification and hierarchical permission table sets different access permissions for different internal personnel roles; The general data security classification and hierarchical permission table sets different access permissions for different external personnel roles; The internal data security classification and grading permission table and the general data security classification and grading permission table each have different static sensitivity levels set according to different data types.

7. The method according to claim 6, characterized in that, Based on the first accessing user and the first static sensitivity classification, determining whether the first accessing user has access rights to the first raw data specifically includes: Identify whether the first accessing user is an internal or external user. Based on the internal or external user role, query the internal data security classification and grading permission table or the general data security classification and grading permission table. Based on the internal or general data security classification and grading permission table, determine whether the first accessing user has access rights to the first static sensitivity level of the first original data.

8. The method according to any one of claims 5-7, characterized in that, The first original data is reconstructed according to a first data reconstruction strategy to obtain the first reconstructed data and output as a response to the data access request, specifically including: Based on the reconstruction density of the first sensitive word and its context features, obtain the first reconstructed data after data reconstruction of the first original data; After outputting the first reconstructed data as a response to the data access request, the feedback from the first accessing user to the response is obtained, and the natural language model is trained again based on the feedback.

9. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program. When the processor runs the computer program stored in the memory, the processor executes the data desensitization access management method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the data desensitization access management method as described in any one of claims 1-8.