Data desensitization processing method, device, equipment, storage medium and product
Patent Information
- Application Number
- CN202611157168.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本申请的主要目的在于提供一种数据脱敏处理方法、装置、设备、存储介质及产品,旨在解决现有的数据脱敏处理效率不高,导致存在商业秘密的违规使用或泄露风险的技术问题
[0014]本申请基于接收到的数据处理任务进行数据采集,得到待处理数据;通过预设商业秘密识别模型对所述待处理数据进行数据敏感识别,得到识别结果;基于所述识别结果通过分级脱敏引擎对所述待处理数据进行分级数据脱敏,得到数据脱敏结果;根据所述数据脱敏结果和所述数据处理任务进行风险识别,得到风险识别结果;基于所述风险识别结果响应所述数据处理任务。相对于现有的采用单一脱敏策略进行数据脱敏,本申请上述方式能够针对不同敏感等级的敏感信息,采用差异化脱敏策略,避免秘密泄露。
Smart Images

Figure CN122839431A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to data desensitization processing methods, apparatus, equipment, storage media and products. Background Technology
[0002] With the rapid development of large language models and agent technology, intelligent agents have been widely applied in enterprise office and business automation scenarios. During task execution, intelligent agents interact with various business systems, knowledge bases, and external tools, and these operational sessions often contain a large amount of sensitive information. Among related technologies, auditing of intelligent agent operational sessions mainly adopts rule-based log auditing, which matches sensitive content using preset keywords (such as "password" or "ID card") to trigger alarms. This method is simple and fast, but it has drawbacks such as a single de-identification strategy and the inability to identify unstructured or semantically sensitive information. Summary of the Invention
[0003] The main objective of this application is to provide a data anonymization processing method, apparatus, device, storage medium, and product, aiming to solve the technical problem that the existing data anonymization processing is inefficient, leading to the risk of unauthorized use or leakage of trade secrets.
[0004] To achieve the above objectives, this application proposes a data anonymization method, which includes: Data is collected based on the received data processing task to obtain the data to be processed; The data to be processed is subjected to data sensitivity identification using a preset trade secret identification model to obtain the identification results; Based on the identification results, the data to be processed is subjected to hierarchical data desensitization using a hierarchical desensitization engine to obtain the data desensitization result. Risk identification is performed based on the data anonymization results and the data processing task to obtain risk identification results; The data processing task is responded to based on the risk identification results.
[0005] Optionally, the step of performing risk identification based on the data anonymization result and the data processing task to obtain the risk identification result includes: The data anonymization result and the data processing task are input into a preset risk identification model to obtain the first risk identification result output by the preset risk identification model. Determine multi-source security information based on the data processing task; The second risk identification result is determined based on the multi-source security information; The risk identification result is determined based on the first risk identification result and the second risk identification result.
[0006] Optionally, determining the second risk identification result based on the multi-source security information includes: The verification rules are determined based on the data source of the multi-source security information. Based on the verification rules, the multi-source security information is validated to obtain a second risk identification result.
[0007] Optionally, determining the risk identification result based on the first risk identification result and the second risk identification result includes: If the first risk identification result indicates that there is a risk, then the existing target risk is determined; Determine the cross-validation risk results in the second risk identification results that correspond to the target risk; If the cross-validation risk result is no risk, a risk-free risk identification result is generated.
[0008] Optionally, the step of performing hierarchical data anonymization on the data to be processed using a hierarchical data anonymization engine based on the identification result to obtain the data anonymization result includes: The data confidentiality level is determined based on the identification results; The hierarchical data anonymization engine performs hierarchical data anonymization on the data to be processed based on the data confidentiality level, thereby obtaining the data anonymization result.
[0009] Optionally, after performing hierarchical data anonymization on the data to be processed using a hierarchical de-identification engine based on the identification result to obtain the data anonymization result, the method further includes: The data anonymization results are classified using a preset classification model to obtain the classification results; If the classification result is risk-free data, the data processing task is responded to based on the data anonymization result; If the classification result indicates the presence of risky data, risk identification is performed based on the data anonymization result and the data processing task to obtain the risk identification result.
[0010] Furthermore, to achieve the above objectives, this application also proposes a data desensitization processing apparatus, which includes: The receiving module is used to collect data based on the received data processing task to obtain the data to be processed. The data sensitivity identification module is used to perform data sensitivity identification on the data to be processed using a preset trade secret identification model, and obtain the identification result. The hierarchical data desensitization module is used to perform hierarchical data desensitization on the data to be processed based on the recognition result through a hierarchical desensitization engine, so as to obtain the data desensitization result. The risk identification module is used to identify risks based on the data anonymization results and the data processing task, and obtain risk identification results. The response module is used to respond to the data processing task based on the risk identification result.
[0011] In addition, to achieve the above objectives, this application also proposes a data desensitization processing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data desensitization processing method described above.
[0012] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the data desensitization processing method described above.
[0013] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the data desensitization processing method described above.
[0014] This application collects data based on a received data processing task to obtain data to be processed; it then performs data sensitivity identification on the data to be processed using a preset trade secret identification model to obtain identification results; based on the identification results, it performs hierarchical data desensitization on the data to be processed using a hierarchical desensitization engine to obtain data desensitization results; it performs risk identification based on the data desensitization results and the data processing task to obtain risk identification results; and it responds to the data processing task based on the risk identification results. Compared to existing methods that use a single desensitization strategy for data desensitization, the above method of this application can adopt differentiated desensitization strategies for sensitive information of different sensitivity levels, thus avoiding the leakage of secrets. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating an embodiment of the data anonymization method of this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the data anonymization processing method of this application; Figure 3 This is a flowchart illustrating Embodiment 3 of the data anonymization processing method of this application; Figure 4 This is a schematic diagram of the module structure of the data desensitization processing device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the data desensitization processing method in the embodiments of this application.
[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0021] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device, data desensitization processing device or intelligent agent that can realize the above functions.
[0022] Based on this, embodiments of this application provide a data desensitization processing method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the data anonymization method of this application.
[0023] In this embodiment, the data desensitization processing method includes the following steps: Step S10: Collect data based on the received data processing task to obtain the data to be processed; It should be understood that the executing entity in this embodiment can be a data desensitization processing device or an intelligent agent capable of data processing. The intelligent agent can be applied to scenarios such as enterprise office work and business automation. During task execution, the intelligent agent will interact with various business systems, knowledge bases, and external tools. The following uses an intelligent agent as an example to illustrate this embodiment and the subsequent embodiments.
[0024] It should be noted that the data processing task can be a specific data processing operation instruction triggered by a user and its associated session context. For example, the data processing task could be triggered by a user: "Analyze the financial report for this quarter and send me a summary of key indicators." Based on the received data processing task, data is collected to obtain the data to be processed. This data can be the raw data required for data processing determined according to the data processing task, and the data to be processed is obtained by querying the database. For example, if the data processing task is: "Analyze the financial report for this quarter and send me a summary of key indicators," the data to be processed could be the financial data and financial report for this quarter collected by the agent.
[0025] Step S20: Perform data sensitivity identification on the data to be processed using a preset trade secret identification model to obtain the identification result; It should be noted that the preset trade secret identification model can be a pre-trained natural language processing (NLP) model whose function is to understand the semantics of text, identify fragments that belong to the company's trade secrets, and determine their category and sensitivity level. The preset trade secret identification model can be a lightweight text classification model based on the BERT architecture. The identification results can include the identified trade secret type (e.g., financial data, customer information, technical secrets), sensitivity level (e.g., core, important, general internal data), confidence level, and the location information of the text fragment.
[0026] In practical implementation, a pre-defined trade secret identification model can identify the specific categories of the data to be processed, such as: financial data, operational data, strategic documents, customer information, trade secrets, and internal information. A sensitivity level (core, important, general internal data) can be predefined for each data type. For example: financial data directly affects a company's market capitalization, financing, or core interests; it can include revenue, profit, budgets, and unpublished financial reports, and its sensitivity level can be defined as core. Operational data reflects a company's operating status and competitiveness; it can include sales data, market share, growth rate, and procurement costs, and its sensitivity level can be defined as important. Strategic documents determine the company's core decisions and plans for future development; they can include merger and acquisition plans, new product plans, and market entry strategies, and their sensitivity level can be defined as core. Trade secrets constitute a company's technological barriers; they can include core algorithm source code, unpublished product architecture, and experimental data, and their sensitivity level can be defined as core. Internal information includes project progress, internal meeting minutes, and non-public organizational structures, and its sensitivity level can be defined as general internal data.
[0027] Step S30: Based on the recognition result, the data to be processed is subjected to hierarchical data desensitization using a hierarchical desensitization engine to obtain the data desensitization result; It should be noted that the hierarchical desensitization engine can be a rule executor or a strategy execution module. It is used to call different desensitization rules to process the corresponding original text fragments based on the sensitivity level in the identification results, thus obtaining the data desensitization results. For example, for content identified as "core" trade secrets: it is replaced with the [Trade Secret: Financial Data] type tag, without retaining any original text. For content identified as "important" trade secrets: generalized desensitization is used, for example, replacing "Company A's contract amount is 50 million yuan" with "Company X's contract amount is >10 million yuan". For content identified as "internal" sensitive information: complete desensitization is not required; the core is lightweight desensitization + preservation of business readability. Partial key information masking can be used to partially obscure a small number of key identifiers, retaining the main content and business context. For example: internal employee number EMP2026001, after desensitization, becomes: EMP****001.
[0028] Step S40: Based on the data desensitization results and the data processing task, risk identification is performed to obtain risk identification results; It should be noted that the risk identification results may include information such as risk type (e.g., unauthorized external transmission, unauthorized access, feeding external models) and risk level / confidence level. Risk identification based on the data anonymization results and the data processing task may involve identifying risks such as sending trade secrets to external email addresses, attempting to access unauthorized business data systems, or performing operations such as deleting or modifying critical configurations; determining the type and quantity of externally transmitted data based on the data anonymization results; and identifying risks such as uploading large amounts of sensitive data. Specific detection rules can be pre-configured according to business needs.
[0029] Furthermore, in order to improve task processing efficiency, after step S30, the method further includes: classifying the data anonymization results using a preset classification model to obtain classification results; If the classification result is risk-free data, the data processing task is responded to based on the data anonymization result; If the classification result indicates the presence of risky data, risk identification is performed based on the data anonymization result and the data processing task to obtain the risk identification result.
[0030] It should be noted that the classification of the data anonymization results using a preset classification model can be achieved through lightweight classification models such as FastText or DistilBERT, pre-classifying the anonymized audit logs (data anonymization results). Classification categories include casual conversation, routine business Q&A, high-risk operations, and operations related to trade secrets. If the classification result is casual conversation or routine business Q&A, the data is considered risk-free, and the data processing task is directly responded to based on the data anonymization result. If the classification result is high-risk operations or operations related to trade secrets, the process proceeds to the large-scale model risk identification step.
[0031] Step S50: Respond to the data processing task based on the risk identification result.
[0032] It should be noted that responding to the data processing task based on the risk identification result can be done by processing the task according to the data anonymization result when the risk identification result indicates no risk. For example, analyzing the data anonymization result according to the user's instructions and displaying the analysis result to the user. Alternatively, different responses can be triggered based on different risk levels in the risk identification result. For example: High risk / confirmed risk: intercept task execution, issue real-time alerts to security personnel, and generate an audit ticket. Medium risk: allow execution but record, or require secondary authentication. Low risk / no risk: process the data processing task normally.
[0033] This embodiment collects data based on a received data processing task to obtain data to be processed; it then performs data sensitivity identification on the data to be processed using a preset trade secret identification model to obtain identification results; based on the identification results, it performs hierarchical data desensitization on the data to be processed using a hierarchical desensitization engine to obtain data desensitization results; it then performs risk identification based on the data desensitization results and the data processing task to obtain risk identification results; and finally, it responds to the data processing task based on the risk identification results. Compared to existing methods that use a single desensitization strategy for data desensitization, the above method in this embodiment can employ differentiated desensitization strategies for sensitive information of different sensitivity levels, thus avoiding the leakage of secrets.
[0034] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating Embodiment 2 of the data anonymization processing method of this application. Step S40 further includes the following steps: Step S401: Input the data anonymization result and the data processing task into the preset risk identification model to obtain the first risk identification result output by the preset risk identification model; It should be noted that this embodiment uses a pre-set risk identification model deployed in a private enterprise environment to identify risk intent. The model's input can be anonymized operation session text (i.e., data anonymization result: the original text with trade secrets removed) and contextual information (operation time, agent ID, target system: the system or email address where the data processing task results need to be sent externally). The model output can be: risk type label, risk confidence level, and risk summary. Specific identification rules can be found in Table 1 - First Risk Identification Description Table. Table 1 - Description of First Risk Identification
[0035] Step S402: Determine multi-source security information based on the data processing task; It should be noted that the multi-source security information may include alarms or anomaly records from other independent security subsystems or log sources within the enterprise or intelligent entity, which are associated with the current data processing task. Examples include: Data Loss Prevention (DLP) system alarms, Zero Trust Network / Access Control System logs, Identity and Access Management (IAM) system logs, etc.
[0036] Step S403: Determine the second risk identification result based on the multi-source security information; It should be noted that the determination of the second risk identification result based on the multi-source security information can be based on information such as alarms, logs, and access in the multi-source security information to determine whether the data processing task has risks, such as verifying whether the data processing task has risks such as sending emails, abnormal logins, or uploading large amounts of data.
[0037] Furthermore, in order to avoid false alarms that the model (first risk identification result) may generate due to contextual understanding bias, step S403 of this embodiment may include: determining verification rules based on the data source of the multi-source security information; Based on the verification rules, the multi-source security information is validated to obtain a second risk identification result.
[0038] It should be noted that different data sources can be configured with different validation rules. Please refer to Table 2 - Cross-validation Description Table. Based on the data source, validation content and linkage logic in Table 2, risk identification can be performed again to obtain a second risk identification result.
[0039] Table 2 - Cross-validation Description Table
[0040] Step S404: Determine the risk identification result based on the first risk identification result and the second risk identification result.
[0041] It should be noted that determining the risk identification result based on the first and second risk identification results can involve first determining whether a risk exists based on the first risk identification result. If a risk exists, then determining whether a risk exists in the second risk identification result. If the second risk identification result also indicates a risk, the risk level is increased. If the second risk identification result indicates no risk, then the types of the first and second risk identification results are obtained. The credibility of the first and second risk identification results is determined based on their types, and the corresponding risk identification result is adopted based on the credibility. For example: if the risk identified by the large model (the first risk identification result) is data feeding to an external model, but there are no DLP alarms in the second risk identification result, it is determined to be a false alarm, and the overall judgment is no risk. If the risk type in the first risk identification result does not have relevant verifiable data in the second risk identification result, then the risk result of the large model shall prevail. If there is no risk in the first risk identification result, but there is risk in the second risk identification result, then the second risk identification result shall prevail.
[0042] In practical implementation, if the first risk identification result is unauthorized external transmission of trade secrets (confidence level 0.94), cross-validation is triggered: The DLP log is queried, revealing records of attachments being sent to other unknown domains at the same time; the zero-trust log is queried, showing that the agent's access to the financial database was normal. Comprehensive judgment: The DLP alarm provides strong corroboration. The final risk identification result is determined to be high-risk – unauthorized external transmission of trade secrets, and real-time interception is triggered. This embodiment does not rely solely on a single model for risk assessment, but introduces a decision-making mechanism combining preliminary model judgment with cross-validation of multi-source information. This significantly improves the accuracy of risk assessment and reduces false alarms.
[0043] Furthermore, to optimize the feedback loop and improve subsequent processing efficiency, this embodiment also includes a manual review process, where security operations personnel confirm the alarms in the risk identification results and mark them as real risks or false alarms. Real risk cases (including information such as trade secret type and risk type) are added to the training set, and false alarm cases are added to the negative sample set for incremental model training. Incremental model training includes training a preset trade secret identification model: weekly incremental fine-tuning to improve identification accuracy. A preset risk identification model: monthly fine-tuning using new samples to adapt to new attack methods. For high-frequency false alarm types, new filtering rules can be automatically generated or thresholds adjusted.
[0044] This embodiment inputs the data anonymization result and the data processing task into a preset risk identification model to obtain a first risk identification result output by the preset risk identification model; determines multi-source security information based on the data processing task; determines a second risk identification result based on the multi-source security information; and determines a final risk identification result based on the first and second risk identification results. This embodiment achieves a closed-loop security audit of the data processing task through the above method, completing data processing without allowing external systems to access the original sensitive data, fundamentally cutting off the leakage path in the audit process.
[0045] Based on the above embodiments of this application, in the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart illustrating Embodiment 3 of the data anonymization processing method of this application. Step S30 further includes the following steps: Step S301: Determine the data confidentiality level based on the identification results; It should be noted that the data confidentiality level can be the data sensitivity level, which can include core, important, and general internal data.
[0046] Step S302: The hierarchical data anonymization engine performs hierarchical data anonymization on the data to be processed based on the data confidentiality level to obtain the data anonymization result.
[0047] It should be noted that the hierarchical data anonymization engine performing hierarchical data anonymization on the data to be processed based on the data confidentiality level can be performed according to the data anonymization rules corresponding to the data confidentiality level. For example, for content identified as "core" trade secrets: replace it with the [Trade Secret: Financial Data] type tag, without retaining any original text. For content identified as "important" trade secrets: use generalized anonymization, for example, replace "Company A's contract amount of 50 million yuan" with "Company X's contract amount > 10 million yuan". For content identified as "internal" sensitive information: complete anonymization is not required; the core is lightweight anonymization + preservation of business readability. Partial key information masking can be used to partially mask a small number of key identifiers, while retaining the main content and business context. For example: internal employee number EMP2026001, after anonymization, becomes: EMP****001.
[0048] In practice, an internal AI agent named "Xiao Zhi" received a user instruction while assisting employees with reporting tasks: "Compile the revenue data of each business unit for this quarter into a table and send it to my external email address 123@exe.com for backup. This revenue data is a core trade secret. The processing flow includes:" Operation session collection: Automatically collects input commands, execution tool calls (reading the database, generating emails), and output content (data to be processed).
[0049] Trade secret identification and graded desensitization: The pre-set trade secret identification model performs semantic analysis on the collected content (data to be processed) to identify "revenue data of each business unit" as financial data, with sensitivity determination as the core.
[0050] The hierarchical desensitization engine performs full desensitization: the relevant content in the audit log is replaced with: "Trade Secret: Financial Data," without retaining the original text. The original content is encrypted and stored in controlled object storage.
[0051] Lightweight pre-filtering: The pre-filtering model identifies that the session involves the "data export + external sending" pattern, marks it as high risk, and sends it to the preset risk identification model.
[0052] Large-scale model risk intent identification: A preset risk identification model is used to identify risks and determine that there is a risk of "unauthorized external disclosure of trade secrets" with a confidence level of 0.96 and a summary of "sending core financial data to an external email address".
[0053] Multi-source cross-validation: Query the DLP system: During the same time period, the user has no other outgoing records, but the email behavior itself has been captured by DLP.
[0054] Querying the zero-trust logs: The agent container's access to the database is authorized normally, but the target email domain is not a company-authorized domain.
[0055] Cross-validation engine assessment: High risk confirmed.
[0056] Alarms and Responses: When a real-time alarm is triggered, the security operations personnel receive a work order. After manual review and confirmation that the risk is genuine, the agent's operation is blocked, and the email is not sent.
[0057] Technical results: Successfully identified and prevented the leakage of core trade secrets; the audit logs did not expose the original data, which complies with the principle of usability without visibility.
[0058] This embodiment determines the data confidentiality level based on the identification results; the hierarchical desensitization engine then performs hierarchical data desensitization on the data to be processed based on the data confidentiality level, obtaining the data desensitization result. This embodiment employs differentiated desensitization strategies for trade secrets with different sensitivity levels to improve data security.
[0059] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the data desensitization processing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0060] This application also provides a data desensitization processing device, please refer to... Figure 4 The data desensitization processing device includes: The receiving module 10 is used to collect data based on the received data processing task to obtain the data to be processed; The data sensitivity identification module 20 is used to perform data sensitivity identification on the data to be processed using a preset trade secret identification model, and obtain the identification result; The hierarchical data desensitization module 30 is used to perform hierarchical data desensitization on the data to be processed based on the recognition result through a hierarchical desensitization engine to obtain the data desensitization result. Risk identification module 40 is used to identify risks based on the data desensitization results and the data processing task, and obtain risk identification results; The response module 50 is used to respond to the data processing task based on the risk identification result.
[0061] This embodiment collects data based on a received data processing task to obtain data to be processed; it then performs data sensitivity identification on the data to be processed using a preset trade secret identification model to obtain identification results; based on the identification results, it performs hierarchical data desensitization on the data to be processed using a hierarchical desensitization engine to obtain data desensitization results; it then performs risk identification based on the data desensitization results and the data processing task to obtain risk identification results; and finally, it responds to the data processing task based on the risk identification results. Compared to existing methods that use a single desensitization strategy for data desensitization, the above method in this embodiment can employ differentiated desensitization strategies for sensitive information of different sensitivity levels, thus avoiding the leakage of secrets.
[0062] The data anonymization processing apparatus provided in this application, employing the data anonymization processing method described in the above embodiments, can solve the technical problem of low efficiency in existing data anonymization processing, leading to the risk of unauthorized use or leakage of trade secrets. Compared with the prior art, the beneficial effects of the data anonymization processing apparatus provided in this application are the same as those of the data anonymization processing method provided in the above embodiments, and other technical features in the data anonymization processing apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0063] This application provides a data desensitization processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data desensitization processing method in the first embodiment described above.
[0064] The following is for reference. Figure 5The diagram illustrates a structural schematic of a data anonymization processing device suitable for implementing embodiments of this application. The data anonymization processing device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The data desensitization processing device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0065] like Figure 5 As shown, the data anonymization processing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the data anonymization processing device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the data desensitization processing device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show data desensitization processing devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0066] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0067] The data anonymization processing device provided in this application, employing the data anonymization processing method described in the above embodiments, can solve the technical problem of low efficiency in existing data anonymization processing, leading to the risk of unauthorized use or leakage of trade secrets. Compared with the prior art, the beneficial effects of the data anonymization processing device provided in this application are the same as those of the data anonymization processing method provided in the above embodiments, and other technical features in this data anonymization processing device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0068] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0069] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0070] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the data desensitization processing method described in the above embodiments.
[0071] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0072] The aforementioned computer-readable storage medium may be included in the data desensitization processing device; or it may exist independently and not be assembled into the data desensitization processing device.
[0073] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Python, Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0075] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0076] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described data desensitization processing method. This solves the technical problem that existing data desensitization processing methods are inefficient, leading to the risk of unauthorized use or leakage of trade secrets. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the data desensitization processing method provided in the above embodiments, and will not be repeated here.
[0077] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the data desensitization processing method described above.
[0078] The computer program product provided in this application can solve the technical problem that the existing data anonymization processing is inefficient, leading to the risk of unauthorized use or leakage of trade secrets. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the data anonymization processing method provided in the above embodiments, and will not be repeated here.
[0079] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. A data anonymization processing method, characterized in that, The data anonymization method includes the following steps: Data is collected based on the received data processing task to obtain the data to be processed; The data to be processed is subjected to data sensitivity identification using a preset trade secret identification model to obtain the identification results; Based on the identification results, the data to be processed is subjected to hierarchical data desensitization using a hierarchical desensitization engine to obtain the data desensitization result. Risk identification is performed based on the data anonymization results and the data processing task to obtain risk identification results; The data processing task is responded to based on the risk identification results.
2. The data anonymization processing method as described in claim 1, characterized in that, The step of performing risk identification based on the data anonymization results and the data processing task to obtain risk identification results includes: The data anonymization result and the data processing task are input into a preset risk identification model to obtain the first risk identification result output by the preset risk identification model. Determine multi-source security information based on the data processing task; The second risk identification result is determined based on the multi-source security information; The risk identification result is determined based on the first risk identification result and the second risk identification result.
3. The data anonymization processing method as described in claim 2, characterized in that, The determination of the second risk identification result based on the multi-source security information includes: The verification rules are determined based on the data source of the multi-source security information. Based on the verification rules, the multi-source security information is validated to obtain a second risk identification result.
4. The data anonymization processing method as described in claim 2, characterized in that, The step of determining the risk identification result based on the first risk identification result and the second risk identification result includes: If the first risk identification result indicates that there is a risk, then the existing target risk is determined; Determine the cross-validation risk results in the second risk identification results that correspond to the target risk; If the cross-validation risk result is no risk, a risk-free risk identification result is generated.
5. The data anonymization processing method according to any one of claims 1-4, characterized in that, The step of performing hierarchical data anonymization on the data to be processed using a hierarchical data anonymization engine based on the identification results to obtain data anonymization results includes: The data confidentiality level is determined based on the identification results; The hierarchical data anonymization engine performs hierarchical data anonymization on the data to be processed based on the data confidentiality level, thereby obtaining the data anonymization result.
6. The data anonymization processing method according to any one of claims 1-4, characterized in that, After performing hierarchical data anonymization on the data to be processed using a hierarchical data anonymization engine based on the identification results to obtain the data anonymization results, the process further includes: The data anonymization results are classified using a preset classification model to obtain the classification results; If the classification result is risk-free data, the data processing task is responded to based on the data anonymization result; If the classification result indicates the presence of risky data, risk identification is performed based on the data anonymization result and the data processing task to obtain the risk identification result.
7. A data desensitization processing device, characterized in that, The data desensitization processing device includes: The receiving module is used to collect data based on the received data processing task to obtain the data to be processed. The data sensitivity identification module is used to perform data sensitivity identification on the data to be processed using a preset trade secret identification model, and obtain the identification result. The hierarchical data desensitization module is used to perform hierarchical data desensitization on the data to be processed based on the recognition result through a hierarchical desensitization engine, so as to obtain the data desensitization result. The risk identification module is used to identify risks based on the data anonymization results and the data processing task, and obtain risk identification results. The response module is used to respond to the data processing task based on the risk identification result.
8. A data anonymization processing device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data desensitization processing method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the data desensitization processing method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the data desensitization processing method as described in any one of claims 1 to 6.