Data desensitization method and device, computer equipment and storage medium
By combining entity recognition models, template libraries, semantic fidelity rewriting engines, and policy libraries, accurate sensitive entity recognition and controllable semantic replacement are achieved, solving the challenges of privacy protection and semantic integrity in corpus sharing and improving the security and effectiveness of data sharing and utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for anonymizing corpora pose a high risk of privacy breaches and cannot guarantee semantic integrity, thus affecting data sharing and utilization.
A multi-stage collaborative processing flow, consisting of entity recognition model, template library, semantic fidelity rewriting engine and strategy library, is adopted to generate desensitized target corpus data through accurate sensitive entity recognition, controllable semantic fidelity replacement and scientific desensitization quality assessment.
This effectively prevents the leakage of privacy information and improves the semantic integrity of the anonymized corpus data, ensuring the security and effectiveness of data sharing and utilization.
Smart Images

Figure CN121935959A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to the financial technology field, particularly to data desensitization methods, devices, computer equipment, and storage media. Background Technology
[0002] With the rapid acceleration of digitalization in the financial industry, financial institutions have accumulated massive amounts of textual data in their daily operations, including customer information, transaction records, and business operations. This data contains immense value and plays a crucial role in supporting the training of artificial intelligence (AI) models and business analysis. AI models trained using this data can help financial institutions achieve business goals such as precision marketing and risk assessment, while in-depth business analysis helps them optimize operational strategies and improve service quality.
[0003] However, sharing and using this financial corpus data presents numerous serious challenges. First, the risk of privacy breaches remains high. Financial corpora contain a large amount of sensitive data, such as personal identification information (e.g., ID number, name, contact information) and account information (e.g., bank card number, account number, password). If this data is improperly obtained or leaked during sharing, it will cause serious economic losses and privacy violations for customers, and financial institutions will also face legal risks and reputational damage. Second, semantic preservation is extremely difficult. Traditional desensitization methods are often simplistic and crude, frequently employing direct replacement or deletion of sensitive information. However, this approach easily damages the semantic integrity of the text, making the desensitized corpus unable to effectively support model training and business analysis. For example, in the financial insurance field, a text containing customer insurance information, if simply replacing or deleting sensitive information such as the customer's age and health status, may make the text semantically unclear, thus affecting the accuracy of the risk assessment model trained on that text. This would prevent accurate assessment of the customer's insurance risk and adversely affect the development of insurance business.
[0004] Therefore, there is an urgent need for a corpus anonymization method that can effectively solve the above problems in order to ensure the secure sharing and effective use of data. Summary of the Invention
[0005] The purpose of this application is to provide a data desensitization method, apparatus, computer equipment, and storage medium to solve the technical problems of existing corpus desensitization methods having a high risk of privacy leakage and failing to guarantee semantic integrity.
[0006] Firstly, a data anonymization method is provided, including: Receive input corpus data, and perform entity recognition on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; Obtain the business type corresponding to the corpus data, and obtain the target semantic template corresponding to the business type from the preset template library; Based on the target semantic template, a preset semantic fidelity rewriting engine is used to process the sensitive entities in the corpus data to generate corresponding replacement schemes. Obtain the sensitivity assessment information corresponding to the sensitive entity, and query the target desensitization strategy corresponding to the sensitivity assessment information from the preset strategy library; Based on the target desensitization strategy and the replacement scheme, the sensitive entities in the corpus data are desensitized to obtain the desensitized target corpus data. Data evaluation is performed on the target corpus data; If the target corpus data passes the data evaluation, then the target corpus data is output.
[0007] Secondly, a data anonymization device is provided, comprising: The recognition module is used to receive input corpus data and perform entity recognition on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; The acquisition module is used to acquire the business type corresponding to the corpus data and acquire the target semantic template corresponding to the business type from the preset template library; The first processing module is used to process the sensitive entities in the corpus data based on the target semantic template using a preset semantic fidelity rewriting engine to generate corresponding replacement schemes. The second processing module is used to obtain the sensitivity assessment information corresponding to the sensitive entity, and to query the target desensitization strategy corresponding to the sensitivity assessment information from the preset strategy library. The desensitization module is used to desensitize the sensitive entities in the corpus data based on the target desensitization strategy and the replacement scheme to obtain the desensitized target corpus data. The evaluation module is used to perform data evaluation on the target corpus data; The output module is used to output the target corpus data if the target corpus data passes the data evaluation.
[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described data desensitization method.
[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described data desensitization method.
[0010] In the above-mentioned data desensitization method, apparatus, computer equipment, and storage medium, the following steps are taken: First, input corpus data is received, and entity recognition is performed on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; the business type corresponding to the corpus data is obtained, and a target semantic template corresponding to the business type is obtained from a preset template library; then, based on the target semantic template, a preset semantic fidelity rewriting engine is used to process the sensitive entities in the corpus data to generate corresponding replacement schemes; next, sensitivity assessment information corresponding to the sensitive entities is obtained, and a target desensitization strategy corresponding to the sensitivity assessment information is queried from a preset strategy library; subsequently, the sensitive entities in the corpus data are desensitized based on the target desensitization strategy and the replacement scheme to obtain desensitized target corpus data; further, data evaluation is performed on the target corpus data; if the target corpus data passes the data evaluation, the target corpus data is output. Based on the above automated processing flow, this application constructs a multi-stage collaborative desensitization process by combining entity recognition models, template libraries, semantic fidelity rewriting engines, and policy libraries. Through accurate sensitive entity recognition, controllable semantic fidelity replacement, and scientific desensitization quality assessment, it solves the problem of balancing data privacy protection and semantic integrity in corpus sharing. It can effectively prevent privacy information leakage and effectively improve the semantic integrity of the target corpus data generated after desensitization. Attached Figure Description
[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the data desensitization method according to this application; Figure 3 This is a schematic diagram of a structure of an embodiment of the data desensitization device according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0019] Server 103 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal device 101.
[0020] It should be noted that the data anonymization method provided in this application is generally executed by a server / terminal device, and correspondingly, the data anonymization device is generally installed in the server / terminal device.
[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0022] Continue to refer to Figure 2 A flowchart illustrating an embodiment of the data anonymization method according to this application is shown. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The data anonymization method provided in this application embodiment can be applied to any scenario requiring data anonymization, and thus can be applied to products in these scenarios, such as data anonymization products in the financial insurance field. The data anonymization method includes the following steps: Step S201: Receive the input corpus data, and perform entity recognition on the corpus data based on a preset entity recognition model to obtain the sensitive entities in the corpus data.
[0023] In this embodiment, the data desensitization method operates on electronic devices (e.g., Figure 1The server / terminal device shown can acquire the input corpus data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods. The implementing entity of this application is specifically a data anonymization system, which can be simply referred to as the system. This application can be applied to business data anonymization scenarios in the fintech field. The aforementioned corpus data (or raw corpus) can be input financial text data, covering multiple aspects such as customer information, transaction information, and business operations. For example, customer information may include personal identification information: name, ID number, mobile phone number, address, etc.; transaction information may include account information: bank card number, account balance, transaction serial number, etc.; and business operations may include transaction operation information: transfer amount, transaction time, counterparty, etc.
[0024] This involves pre-constructing entity recognition models covering various types of information, including personal identity information, account information, and transaction information, based on financial knowledge graphs and domain dictionaries. The financial knowledge graph contains information on various entities and their relationships within the financial field, while the domain dictionary covers vocabulary and terminology specific to the financial domain. Then, the input corpus data is scanned and analyzed word by word and sentence by sentence. Using natural language processing techniques, such as lexical analysis and syntactic analysis, the corpus is decomposed into basic language units (such as words and phrases). Next, each language unit is matched and judged based on information from the financial knowledge graph and domain dictionary. For example, for words that may contain personal identity information, such as "ID number" or "name," the context is used to further confirm whether they are genuine personal identity entities; for account information, such as "bank card number" or "account number," accurate identification is also performed; and for transaction information, such as "transaction amount," "transaction time," or "counterpartie," identification is also performed according to relevant rules. In this way, entity recognition models covering multiple types of information, including personal identity information, account information, and transaction information, are constructed to identify all sensitive entities in the corpus.
[0025] Step S202: Obtain the business type corresponding to the corpus data, and obtain the target semantic template corresponding to the business type from the preset template library.
[0026] In this embodiment, the business type of the aforementioned corpus data can be analyzed to determine the relevant business type. The aforementioned template library refers to a pre-built financial template library. This library covers semantic templates for different financial businesses such as banking, securities, and insurance. These templates contain semantic structures and key business logic for various common business scenarios. Specifically, a semantic template corresponding to the aforementioned business type can be selected from this financial template library. For example, if the input corpus data is about bank loan business, the system will select a semantic template related to bank loans from the template library. This semantic template may contain semantic descriptions and key business logic for loan application, approval, and disbursement processes, such as the relationship between elements like loan amount, interest rate, and repayment period.
[0027] Step S203: Based on the target semantic template, use a preset semantic fidelity rewriting engine to process the sensitive entities in the corpus data to generate corresponding replacement schemes.
[0028] In this embodiment, the specific implementation process of processing sensitive entities in the corpus data and generating corresponding replacement schemes based on the target semantic template using a preset semantic fidelity rewriting engine will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0029] Step S204: Obtain the sensitivity assessment information corresponding to the sensitive entity, and query the target desensitization strategy corresponding to the sensitivity assessment information from the preset strategy library.
[0030] In this embodiment, the specific implementation process of obtaining the sensitive assessment information corresponding to the sensitive entity will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0031] The aforementioned strategy library (or configurable strategy library) is a pre-built database providing various de-identification strategies, including strong de-identification, weak de-identification, and compliance templates. Strong de-identification strategies are typically suitable for scenarios with extremely high data privacy requirements (corresponding to high sensitivity), such as completely replacing sensitive information with fictitious, meaningless information, for example, replacing a real ID number with a randomly generated string conforming to the ID number format. Weak de-identification strategies are more flexible and may only partially hide or obscure sensitive information (corresponding to medium sensitivity), such as replacing the middle digits of a bank card number with asterisks. Compliance template strategies are de-identification strategies formulated according to specific laws, regulations, and industry standards to ensure that the de-identified data complies with relevant requirements. Users can select and configure appropriate strategies from the strategy library based on their data usage scenarios and requirements.
[0032] In addition, the system is equipped with a multi-strategy desensitization execution module. This module allows users to select a desensitization strategy from a configurable strategy library that matches the aforementioned sensitivity assessment information. For example, if the sensitivity assessment information includes both high and medium sensitivity, the system will query the configurable strategy library to find a strong desensitization strategy that matches high sensitivity and a weak desensitization strategy that matches medium sensitivity.
[0033] Step S205: Based on the target desensitization strategy and the replacement scheme, the sensitive entities in the corpus data are desensitized to obtain the desensitized target corpus data.
[0034] In this embodiment, the replacement scheme generated by the semantic fidelity rewriting engine details how to perform replacement operations for different types of sensitive entities while maintaining financial logic and semantic structure. For example, it specifies which placeholder to use for names, how to hide bank card numbers, and how to obfuscate transaction amounts. Alternatively, the generated replacement scheme might specify replacing "transaction initiator's name" with "[transaction party's name]", "bank card number" with "[bank card number]", and transforming "transaction amount" according to a certain numerical relationship.
[0035] The system's differential replacement unit performs specific desensitization and replacement operations on sensitive entities in the original corpus (i.e., corpus data) according to the selected target desensitization strategy and the aforementioned replacement scheme, ultimately completing the desensitization task and generating desensitized target corpus data that meets the requirements. For example, if a sensitive entity is identified and the replacement scheme specifies that "ID number" should be strongly desensitized and replaced with a randomly generated fictitious number, then during the desensitization process, based on this scheme and the selected strong desensitization strategy, the real ID numbers in the original corpus data are replaced with randomly generated fictitious numbers that conform to the specified format, thus achieving desensitization. Furthermore, the system provides a function that allows users to preview the effect before data desensitization and supports strategy adjustments.
[0036] Step S206: Perform data evaluation on the target corpus data.
[0037] In this embodiment, the specific implementation process of data evaluation of the target corpus data described above will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0038] Step S207: If the target corpus data passes the data evaluation, then the target corpus data is output.
[0039] In this embodiment, the generated target corpus data can be sent to relevant business personnel via email, message, or interface display to complete the output processing of the target corpus data.
[0040] This application first receives input corpus data and performs entity recognition on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; it then obtains the business type corresponding to the corpus data and retrieves the target semantic template corresponding to the business type from a preset template library; next, based on the target semantic template, it uses a preset semantic fidelity rewriting engine to process the sensitive entities in the corpus data to generate corresponding replacement schemes; then, it obtains the sensitivity assessment information corresponding to the sensitive entities and queries the target desensitization strategy corresponding to the sensitivity assessment information from a preset strategy library; subsequently, it performs desensitization processing on the sensitive entities in the corpus data based on the target desensitization strategy and the replacement scheme to obtain desensitized target corpus data; further, it performs data evaluation on the target corpus data; if the target corpus data passes the data evaluation, it outputs the target corpus data. Based on the above automated processing flow, this application constructs a multi-stage collaborative desensitization process by combining entity recognition models, template libraries, semantic fidelity rewriting engines, and policy libraries. Through accurate sensitive entity recognition, controllable semantic fidelity replacement, and scientific desensitization quality assessment, it solves the problem of balancing data privacy protection and semantic integrity in corpus sharing. It can effectively prevent privacy information leakage and effectively improve the semantic integrity of the target corpus data generated after desensitization.
[0041] In some alternative implementations, step S203 includes the following steps: The semantically faithful rewriting engine calls a pre-defined controllable text generator and numerical relationship preserver.
[0042] In this embodiment, the aforementioned controllable text generator is a pre-built tool that maintains the accuracy of sentence structure and terminology when replacing sensitive information, based on identified sensitive entities and selected semantic templates. The aforementioned numerical relation maintainer is a pre-built tool that performs consistent transformations on financial numerical values in corpus data.
[0043] Based on the template specification information in the target semantic template, the controllable text generator is used to rewrite the sensitive entities in the corpus data to obtain the corresponding first processed data.
[0044] In this embodiment, the aforementioned template specification information refers to the sentence structure and terminology requirements in the target semantic template. Specifically, the controllable text generator operates based on a constrained text generation model. It uses the identified sensitive entities and the selected target semantic template as a foundation, maintaining the accuracy of sentence structure and terminology when replacing sensitive information. Specifically, for identified sensitive entities, such as names in personal identification information, the generator generates appropriate replacement content (i.e., the first processed data) based on the sentence structure and terminology requirements in the target semantic template. For example, in a bank loan business template, if the original sentence is "Applicant Zhang San's loan amount is 100,000 yuan," the generator might replace "Zhang San" with "[Applicant's Name]," while maintaining the sentence structure and terminology of "Loan amount is 100,000 yuan," generating "Applicant [Applicant's Name]'s loan amount is 100,000 yuan."
[0045] Based on the numerical relationship maintainer, a consistency transformation is performed on the specified values related to the specified business in the first processed data to obtain the corresponding second processed data.
[0046] In this embodiment, the specific implementation process of performing consistency transformation on the specified values related to the specified business in the first processed data based on the numerical relationship maintainer to obtain the corresponding second processed data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0047] The second processed data is used as the replacement scheme.
[0048] In this embodiment, the generated replacement scheme plans how to remove sensitive information while preserving the original semantics and financial logic of the corpus to the greatest extent possible. For example, if the sensitive entity "customer name" is identified in the corpus, it will be replaced with a placeholder form such as "[customer name]" based on the business scenario and template. This removes the real information while maintaining the integrity of the sentence structure and semantics, providing a detailed operational guide for subsequent actual desensitization.
[0049] This application utilizes a semantically faithful rewriting engine to invoke a pre-defined controllable text generator and a numerical relation maintainer. Then, based on template specification information in the target semantic template, the controllable text generator rewrites the sensitive entities in the corpus data to obtain the corresponding first processed data. Next, the numerical relation maintainer performs consistency transformation on specified values related to a specific business in the first processed data to obtain the corresponding second processed data. This second processed data is then used as a replacement scheme. Based on this processing flow, this application effectively ensures that key business logic remains unchanged during the desensitization process through template library matching, controllable text generator generation of replacement text, and numerical relation maintainer processing of numerical information. This also maintains the accuracy of sentence structure, terminology, and the dimensional relationships and statistical characteristics between numerical values. The resulting replacement scheme can remove sensitive information while preserving the original semantics and financial logic of the corpus to the greatest extent possible, providing crucial support for the subsequent generation of high-quality desensitized corpus.
[0050] In some optional implementations of this embodiment, the step of performing consistency transformation on the specified values related to the specified business in the first processed data based on the numerical relationship maintainer to obtain the corresponding second processed data includes the following steps: The numerical relationship maintainer identifies a specific value in the first processed data that is related to a specific business.
[0051] In this embodiment, the specified business refers to financial business, and the specified value refers to the financial value in the first processed data. All financial values in the first processed data, such as transaction amounts, interest rates, and stock prices, can be identified using a value relation maintainer.
[0052] Obtain the associated information of the specified value.
[0053] In this embodiment, the aforementioned correlation information includes dimensional relationships and statistical characteristics.
[0054] Based on a preset transformation strategy, the associated information is used to transform the specified value in the first processed data to obtain the processed transformed data.
[0055] In this embodiment, the transformation strategy includes: transforming data based on the dimensional relationships and statistical characteristics between numerical values. Specifically, if the corpus contains a set of stock price data, the relative proportions between these stock prices are maintained during the anonymization process. For example, if the original stock prices are 10 yuan, 20 yuan, and 30 yuan, after transformation they might become 15 yuan, 30 yuan, and 45 yuan, maintaining the 1:2:3 ratio. Simultaneously, for numerical values with specific statistical characteristics, such as the mean and standard deviation, these statistical characteristics are also ensured to remain unchanged during the transformation process.
[0056] The transformed data is used as the second processed data.
[0057] This application identifies a specified value related to a specific business in the first processed data using a numerical relation maintainer; then, it obtains the association information of the specified value, which includes dimensional relationships and statistical characteristics; subsequently, based on a preset transformation strategy, it uses the association information to transform the specified value in the first processed data, obtaining transformed data; this transformed data is then used as the second processed data. Based on this processing flow, this application identifies a specified value related to a specific business in the first processed data using a numerical relation maintainer, obtains the association information of the specified value, uses the association information to transform the specified value in the first processed data based on a transformation strategy, and uses the resulting transformed data as the corresponding second processed data. This achieves consistent transformation processing of the specified value, ensuring the accuracy of the obtained second processed data.
[0058] In some alternative implementations, step S204 includes the following steps: Invoke the preset sensitivity dynamic evaluator.
[0059] In this embodiment, the aforementioned sensitivity dynamic evaluator is a pre-built tool with sensitivity dynamic evaluation functionality.
[0060] The sensitivity dynamic evaluator is used to obtain the entity type, context, and data usage scenario of a specified entity; wherein, the specified entity is any one of all the sensitive entities.
[0061] In this embodiment, the designated entity is any one of all sensitive entities. The type of each entity is determined using a dynamic sensitivity assessor, such as a name or ID number in personal identification information, or a bank card number in account information. Then, the context in which the entity exists is analyzed. Different contexts may affect the privacy sensitivity of an entity. For example, a company name mentioned in a public financial report discussion may have low privacy sensitivity; while the same company name associated with an account in an internal document containing detailed customer account information would have higher privacy sensitivity. Furthermore, the data usage scenario is considered. If the data is used for internal risk assessment, the sensitivity of some entities may differ from that used for external marketing and promotion.
[0062] The entity type, the context, and the data usage scenario are quantified based on a preset quantization strategy to obtain the corresponding quantified data.
[0063] In this embodiment, the above-mentioned quantification strategy includes: 1) Quantification based on entity type. Different types are assigned base scores: Different types of sensitive entities are assigned different base sensitivity scores due to the varying importance of the privacy information they contain. For example, the ID card number in personal identity information can uniquely identify a person and contains a large amount of key privacy information, so the base sensitivity score may be set to a higher value, such as 8 points (out of 10). While name information is also personal identity information, its privacy sensitivity is slightly lower than that of the ID card number, so the base sensitivity score may be set to 5 points. The bank card number in account information involves fund security, so the base sensitivity score is set to 7 points. The transaction amount in transaction information may have relatively low privacy sensitivity when viewed alone, so the base score is set to 3 points. Further adjustments to sub-categories: For some sub-categories under a broad category, the scores will be further adjusted according to the specific circumstances. Taking account information as an example, the credit card number may have higher sensitivity than a regular savings card number because it involves sensitive businesses such as credit consumption, so 1-2 points are added to the base score.
[0064] 2) Context-Based Quantification. Analyzing Contextual Relevance: Examining the context in which a sensitive entity exists, analyzing its relevance to other information and the public nature of the scenario. If a sensitive entity is in a public, general context and has weak relevance to other sensitive information, its sensitivity score will decrease accordingly; conversely, if it is in a private, specific business scenario and is closely related to other sensitive information, its sensitivity score will increase. Setting Context Influence Coefficients: Setting influence coefficients for different contexts to adjust the basic sensitivity score. For example, in a discussion of a publicly available financial report, the context influence coefficient for a company name is set to 0.5. If the basic sensitivity score for the company name is 4 points, after adjustment, it becomes 4 × 0.5 = 2 points. However, in an internal document involving detailed customer account information, where the company name is also associated with the account, the context influence coefficient is set to 1.5, and after adjustment, the score becomes 4 × 1.5 = 6 points.
[0065] 3) Quantification based on data usage scenarios. Clearly define scenario-sensitive requirements: Different data usage scenarios have different privacy protection requirements. Internal risk assessment scenarios primarily focus on whether the data accurately reflects the risk situation, with relatively flexible requirements for protecting sensitive information; while external marketing and promotion scenarios require strict protection of customer privacy to prevent information leakage. Develop scenario adjustment rules: Develop corresponding adjustment rules based on different scenarios, and readjust the sensitivity scores after context adjustment. For example, for internal risk assessment scenarios, set the adjustment coefficient to 0.8 - 1.2. If an entity's score is 6 after context adjustment, the score may become between 6 × 0.8 = 4.8 and 6 × 1.2 = 7.2 in the internal risk assessment scenario; for external marketing and promotion scenarios, set the adjustment coefficient to 0.5 - 0.8, the same entity with a score of 6 will become between 6 × 0.5 = 3 and 6 × 0.8 = 4.8 after adjustment.
[0066] The quantified data is then comprehensively calculated and processed to obtain the corresponding quantified score data.
[0067] In this embodiment, the comprehensive calculation process includes: comprehensively processing the scores adjusted for three factors—entity type, context, and data usage scenario—to obtain the final privacy sensitivity quantification score (i.e., the quantified score data). This can be achieved using a weighted summation method, assigning a weight to each factor (e.g., entity type weight 0.4, context weight 0.3, and data usage scenario weight 0.3). The adjusted scores for each factor are multiplied by their respective weights and then summed to obtain the final quantified privacy sensitivity score. Alternatively, depending on the actual situation, professionals can comprehensively assess the impact of the three factors and provide a comprehensive quantified result. For example, after the above adjustments, an entity might have a final score of 5 out of 10 using the weighted summation method. This score represents the privacy sensitivity quantification result of the entity in the current situation, providing a basis for subsequent desensitization processing. A higher score may indicate a deeper level of desensitization.
[0068] The quantified score data is mapped based on a preset sensitivity mapping strategy to obtain the specified sensitivity assessment information of the specified entity.
[0069] In this embodiment, the aforementioned sensitivity mapping strategy is a pre-constructed strategy based on actual business needs, storing a one-to-one correspondence between quantified score data and sensitivity assessment information. For example, 0-3 points correspond to low sensitivity, 3-6 points to medium sensitivity, and 6-10 points to high sensitivity. The quantified score data can be mapped based on this correspondence, and the matched sensitivity can be used as the specified sensitivity assessment information for the designated entity.
[0070] This application invokes a preset sensitivity dynamic evaluator; then, based on the sensitivity dynamic evaluator, it obtains the entity type, context, and data usage scenario of a specified entity; wherein the specified entity is any one of all sensitive entities; and quantifies the entity type, context, and data usage scenario based on a preset quantification strategy to obtain corresponding quantified data; then, it performs comprehensive calculation processing on the quantified data to obtain corresponding quantified score data; subsequently, it maps the quantified score data based on a preset sensitivity mapping strategy to obtain the specified sensitivity assessment information for the specified entity. Based on the above processing flow, this application, through the combined use of the sensitivity dynamic evaluator, quantification strategy, and sensitivity mapping strategy, can intelligently and accurately complete the sensitivity assessment processing of the entity type, context, and data usage scenario of sensitive entities, ensuring the accuracy of the generated sensitivity assessment information.
[0071] In some alternative implementations, step S206 includes the following steps: Based on the corpus data, a preset semantic integrity evaluator is used to assess the degree of semantic preservation of the target corpus data.
[0072] In this embodiment, the semantic integrity evaluator is a pre-built tool that quantifies the degree of semantic preservation by comparing the differences in semantic representations before and after desensitization. Specifically, the semantic integrity evaluator uses natural language processing techniques and semantic analysis algorithms to perform in-depth analysis on the corpora before and after desensitization (i.e., the aforementioned corpus data and the target corpus data). First, the corpora before and after desensitization are converted into semantic vector representations that computers can understand. These vectors can capture the semantic information in the corpus. Then, the similarity between the semantic vectors before and after desensitization is calculated. The higher the similarity, the better the degree of semantic preservation. For example, a quantitative value is obtained to represent the degree of semantic preservation by calculating cosine similarity. If the similarity value is close to 1, it means that the desensitized corpus is very close to the original corpus semantically, and the semantic preservation is good. Therefore, the target corpus data is determined to have passed the semantic preservation assessment. If the similarity value is low, it means that the desensitization process may have caused some damage to the semantics. Therefore, the target corpus data is determined to have failed the semantic preservation assessment and needs further inspection and adjustment. In addition, if the target corpus data is detected as failing the semantic preservation assessment, it is directly determined that the target corpus data has failed the data evaluation.
[0073] If the target corpus data passes the semantic preservation assessment, then a preset task validator is used to perform a quality assessment on the target corpus data based on the corpus data.
[0074] In this embodiment, the specific implementation process of using a preset task verifier to perform quality assessment on the target corpus data based on the corpus data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0075] If the target corpus data passes the quality assessment, then the target corpus data is determined to have passed the data assessment.
[0076] In this embodiment, the target corpus data is considered to have passed the data evaluation only if it is detected that the target corpus data has passed both the semantic preservation degree evaluation and the quality evaluation.
[0077] If the target corpus data fails the quality assessment, then the target corpus data is deemed to have failed the data assessment.
[0078] In this embodiment, if the target corpus data is detected as failing the quality assessment, it is determined that the target corpus data has failed the data assessment. Similarly, if the target corpus data is detected as failing the semantic preservation assessment, it is also determined that the target corpus data has failed the data assessment.
[0079] This application assesses the semantic preservation level of target corpus data using a pre-defined semantic integrity evaluator. If the target corpus data passes the semantic preservation assessment, a pre-defined task validator is used to assess its quality. If the target corpus data passes the quality assessment, it is determined to have passed the data evaluation; otherwise, it is determined to have failed the quality assessment. Based on this process, this application achieves multi-dimensional data evaluation of the target corpus data by using a semantic integrity evaluator to assess semantic preservation and a task validator to assess its quality. This improves the intelligence and accuracy of the data evaluation process and ensures the accuracy of the obtained data evaluation results.
[0080] In some optional implementations of this embodiment, the step of using a preset task validator to perform quality assessment on the target corpus data based on the corpus data includes the following steps: Based on the task validator, the corpus data is used to train the downstream task model related to the specified business to obtain the corresponding first task model.
[0081] In this embodiment, the specified business refers specifically to financial business. The downstream task model can be selected based on actual needs and is closely related to financial business; for example, it may include financial risk assessment models and customer classification models. The financial risk assessment model aims to accurately assess the risk level of a customer or business through the analysis of various financial data (such as customer credit records and transaction data). This type of model has extremely high requirements for data completeness and accuracy, as any missing or incorrect key information may lead to deviations in risk assessment results, thereby affecting the financial institution's decision-making and business security. The customer classification model focuses on classifying customers into different categories based on their characteristics and behaviors, enabling financial institutions to develop personalized marketing strategies and service plans. This model also relies on rich and accurate data to achieve precise customer classification.
[0082] Furthermore, to provide a reliable benchmark for evaluating the quality of anonymization, the aforementioned task validator (or downstream task validator) first trains the selected downstream task model using the original corpus data. The corpus data contains complete, authentic data without any anonymization processing, fully reflecting the characteristics and patterns of the data in its original state. By training on the original corpus data, the model can learn various patterns and relationships within the data, thereby achieving a relatively stable and accurate performance level, which is the baseline model performance. Additionally, during training, the task validator employs appropriate training algorithms and parameter settings to ensure that the trained first-task model can fully extract information from the original corpus. Simultaneously, to evaluate the model's performance, a series of reasonable evaluation metrics are used, such as precision, recall, and F1 score (specific metrics depend on the characteristics and requirements of the downstream task). These metrics objectively measure the model's performance on the original corpus, providing an accurate basis for subsequent performance comparisons with models trained on anonymized corpus data.
[0083] The downstream task model is fine-tuned using the target corpus data to obtain the corresponding second task model.
[0084] In this embodiment, after obtaining the performance of the baseline model, the task validator fine-tunes the same model (i.e., the downstream task model) using the anonymized corpus (target corpus data, or anonymized corpus). The anonymized corpus is obtained by performing a series of anonymization operations on the original corpus, with the aim of removing or replacing sensitive information to protect data privacy. However, some information important to the downstream task may be inevitably lost during the anonymization process; therefore, it is necessary to fine-tune the model to verify whether the anonymized corpus can still support the normal operation of the downstream task.
[0085] During fine-tuning, the task validator uses similar training algorithms and parameter settings as when training on the original corpus, but makes appropriate adjustments based on the characteristics of the anonymized corpus. Since the data distribution and features of the anonymized corpus may differ from the original corpus, the model needs time to adapt to this change and relearn the patterns and relationships in the data. Through fine-tuning, the model can gradually adjust its parameters to achieve better performance on the anonymized corpus. After completing the training on the anonymized corpus and obtaining the second task model, the downstream task validator will obtain the performance of the model trained on the anonymized corpus.
[0086] Obtain the first model performance data of the first task model, and obtain the second model performance data of the second task model.
[0087] In this embodiment, the first model performance data refers to the baseline model performance of the first task model, and the second model performance data refers to the model performance of the second task model trained on the desensitized corpus.
[0088] A comparative analysis of the performance differences between the first model performance data and the second model performance data is performed to obtain the corresponding performance difference results.
[0089] In this embodiment, the performance difference results can be obtained by comparing the performance of the first model performance data (benchmark model performance) and the second model performance data (model performance). The results show that the model performance trained on the anonymized corpus is similar to the benchmark model performance, or that the model performance trained on the anonymized corpus differs significantly from the benchmark model performance.
[0090] The performance difference results are analyzed to generate quality assessment results corresponding to the target corpus data.
[0091] In this embodiment, if the performance of the model trained on the anonymized corpus is similar to that of the baseline model, it means that the anonymized corpus performs well in preserving key information. Despite the anonymization process, the corpus still contains sufficient important information to support downstream tasks. The model can learn patterns and relationships similar to those in the original corpus from the anonymized corpus, thus maintaining relatively stable performance. In this case, the anonymization quality can be considered high, and the anonymization process has achieved the expected results, protecting data privacy without significantly impacting downstream tasks. This leads to the generation of a first quality assessment result for the target corpus data, indicating that it has passed the quality evaluation.
[0092] Conversely, if the performance difference between the two models is significant, it indicates that too much information important to downstream tasks may have been lost during the desensitization process. Desensitization may excessively compromise data integrity and usability while removing sensitive information, preventing the model from learning effective information from the desensitized corpus and thus impacting model performance. In such cases, the desensitization quality can be deemed problematic, resulting in a second quality assessment result indicating that the target corpus data failed the quality evaluation. Furthermore, it is necessary to re-examine the desensitization strategy and processing methods, optimizing and adjusting the desensitization process to reduce the loss of key information and improve desensitization quality.
[0093] This application uses a task validator to train a downstream task model related to a specified business using corpus data to obtain a first task model; then, it fine-tunes the downstream task model using target corpus data to obtain a second task model. Next, it acquires the first model performance data of the first task model and the second model performance data of the second task model. Following this, it performs a comparative analysis of the performance differences between the first and second model performance data to obtain corresponding performance difference results. Subsequently, it analyzes these performance difference results to generate a quality assessment result corresponding to the target corpus data. Based on this processing flow, this application uses corpus data to train a downstream task model related to a specified business to obtain a first task model, and then uses de-identified target corpus data to fine-tune the downstream task model to obtain a second task model. The performance difference between the generated first and second task models is then used as a criterion for evaluating the de-identification quality, thereby enabling a scientific and objective evaluation of the de-identification effect and ensuring the accuracy of the quality assessment results for the generated target corpus data.
[0094] In some optional implementations of this embodiment, after step S207, the electronic device may further perform the following steps: Collect all operational information of the corpus data during the desensitization process.
[0095] In this embodiment, all operational information regarding the aforementioned corpus data during the de-identification process can be comprehensively obtained from the end-to-end audit tracker. The end-to-end audit tracker accurately captures every step of the data de-identification process. It meticulously records the execution time of each step, accurate to the second or even millisecond, to clearly understand the time span of the desensitization operation and the time intervals between each step; the operator information is clearly specified down to the individual, including their name, department, and position, which helps to clarify the responsible party; the desensitization strategy used is the core basis for the desensitization operation, and its name, version, scope of application, and specific parameter settings are recorded in detail. For example, for data replacement strategies, the replacement rules and replacement character sets are recorded; the types of sensitive information replaced cover personal identification information (such as name, ID number, contact information, etc.), account information (such as bank card number, account number, password, etc.), transaction information (such as transaction amount, transaction time, counterparty, etc.), and sensitive information specific to financial business, while also recording the specific replacement content, such as replacing the ID number "11010519900307XXXX" with "11010519900307****", for subsequent traceability and review.
[0096] All the aforementioned operation information is organized and categorized to obtain corresponding processing information.
[0097] In this embodiment, the above-mentioned sorting and classification process includes: classification by desensitization steps: the collected operation information is systematically sorted and classified, and the relevant information is arranged in an orderly manner according to the order of the desensitization steps. The operation information of the entity recognition model is classified separately, including the module's startup time, the types and number of sensitive entities identified, the recognition accuracy, etc.; the semantic fidelity rewriting engine's generation scheme information records in detail its generated rewriting rules, rewritten text examples, and semantic comparison analysis before and after rewriting; the desensitization operation information is subdivided according to different desensitization strategies, such as data masking strategy, data encryption strategy, data generalization strategy, etc., and the operation time, operators, and replaced sensitive information are recorded for each strategy; the evaluation results information of the target corpus data is classified separately to facilitate subsequent comprehensive analysis of the desensitization quality.
[0098] Sensitive Information Replacement Categorization: Detailed categorization and recording are maintained for different types of sensitive information replacement. Personal identification information replacement is listed separately, including the quantity and method of replacement (e.g., partial masking, complete replacement) of various personal identification information such as name, ID number, and contact information, along with before-and-after comparison examples. Account information replacement is also meticulously recorded, detailing the replacement of bank card numbers, account numbers, passwords, and other account-related information. Transaction information replacement covers the replacement of transaction amounts, transaction times, and counterparties. This detailed categorization allows for a clear understanding of the de-identification process for different types of sensitive information, facilitating focused review of the de-identification effectiveness of specific types of sensitive information during compliance checks.
[0099] Based on the processed information, a corresponding de-identified audit report is constructed.
[0100] In this embodiment, the process of constructing the aforementioned de-identification audit report includes: 1) Determining the report structure: Based on the organized and categorized information, design the structural framework of the audit report. This generally includes the report title, report purpose, de-identification project overview, detailed record of the de-identification process, quality assessment result analysis, compliance conclusion, and appendices. The report title should be concise and clear, accurately reflecting the core content of the report; the report purpose should clearly state the purpose of generating the audit report, such as for internal compliance checks or external regulatory audits; the de-identification project overview should briefly introduce the background, objectives, data scope, and sensitive information types involved in the de-identification project; the detailed record of the de-identification process should be elaborated according to the information categorized above; the data assessment result analysis should conduct an in-depth analysis of the semantic integrity assessment and downstream task verification results to assess whether the de-identification quality meets the requirements; the compliance conclusion should determine whether the de-identification process and results are compliant based on relevant laws, regulations, and industry standards; the appendix can include relevant original data, assessment model descriptions, operation manuals, and other supplementary materials. 2) Writing the report content: Write each part of the content according to the report structure framework. During the writing process, ensure that the language is accurate, clear, and concise, and avoid using vague or ambiguous expressions. Key data and information should be explained in detail, and charts and examples can be attached as needed. For example, when describing the detailed record of the desensitization process, a table can be used to list the operation time, operators, and operation content of each step, making the report more intuitive and easy to understand. When analyzing the quality assessment results, bar charts or line graphs can be drawn to show the changes in the performance indicators of different models, making it easier for readers to quickly understand the desensitization quality status. 3) Review the accuracy, completeness, and compliance of the report from different perspectives. During the review process, focus on whether the data in the report is accurate, whether the logic is clear and reasonable, and whether the conclusions are objective and fair. Based on the review comments, revise and improve the report to ensure that the report quality reaches a high level.
[0101] The de-identified audit report is stored.
[0102] In this embodiment, the final anonymized audit report can be securely stored by selecting appropriate storage media and methods to ensure the report's confidentiality, integrity, and availability. Encrypted storage can be used to encrypt the report document, preventing unauthorized access and tampering. Simultaneously, the report should be stored on a secure and reliable server or storage device and backed up regularly to prevent data loss.
[0103] This application collects all operational information from the data anonymization process; then, it organizes and categorizes this information to obtain corresponding processing information; subsequently, it constructs a corresponding anonymization audit report based on this processing information; and finally, it stores the anonymization audit report. Based on this process, this application automatically and intelligently generates a detailed, accurate, and compliant anonymization audit report, which provides strong support for compliance checks during the data anonymization process and ensures that the data anonymization work complies with relevant laws, regulations, and industry standards.
[0104] In some optional implementations, the system also includes a contextual relationship analyzer, which analyzes the business logic relationships between identified sensitive entities. It determines the contextual relationships between entities by analyzing the semantic associations, syntactic structures, and business rules in the financial field within the corpus. For example, in a transaction corpus, after identifying entities such as "transaction initiator," "transaction recipient," "transaction amount," and "transaction time," the system analyzes the business logic relationships between them, such as the "transaction initiator" conducting a transaction with the "transaction recipient" at the "transaction time" for the amount of "transaction amount." This ensures that the business logic relationships between these entities remain unchanged during subsequent de-identification processing, thereby guaranteeing the semantic relevance of the de-identified corpus.
[0105] Furthermore, the desensitization process of the corpus data can be optimized based on the contextual relationships between sensitive entities generated by the contextual relationship analyzer. During desensitization, each sensitive entity should not be desensitized in isolation; instead, it should be processed according to the business logic relationships between entities determined by the contextual relationship analysis. For example, if only "transaction initiator," "transaction recipient," "transaction amount," and "transaction time" are desensitized individually without considering their business logic connections, the desensitized corpus may disrupt the original semantic relevance, leading to information confusion and an inability to accurately convey the original transaction information. Only by maintaining these relationships unchanged during the desensitization process, based on the business logic relationships derived from the contextual relationship analysis, can the desensitized corpus, while removing sensitive information, still reasonably and accurately express the original business scenario and semantic content.
[0106] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0107] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0108] Furthermore, this application discloses an automatic desensitization and semantically faithful rewriting scheme for financial corpora, belonging to the fields of financial technology and data security technology. The core innovation lies in constructing a multi-stage collaborative desensitization pipeline. Through accurate sensitive entity recognition, controllable semantically faithful replacement, and scientific desensitization quality assessment, it solves the technical challenge of balancing data privacy protection and semantic integrity in financial corpus sharing. Technical highlights include: a financial-specific entity recognition engine that accurately identifies and classifies sensitive information in financial texts; a semantic template preserver that ensures the integrity of key business logic and financial terminology during the desensitization process; and a multi-strategy configurable desensitization framework that supports differentiated privacy protection needs in different scenarios. This invention achieves a sensitive entity recognition accuracy of ≥95% and a semantic integrity preservation rate of ≥90% while ensuring that the performance degradation of downstream tasks does not exceed 2%, providing reliable technical support for the secure sharing and compliant use of financial data.
[0109] Furthermore, through the solution proposed in this application, financial institutions can securely share data resources while fully protecting customer privacy, promoting financial technology innovation and cooperation, and simultaneously meeting increasingly stringent data compliance regulatory requirements. The beneficial effects of this application include: significant privacy protection: higher accuracy in identifying sensitive entities, effectively preventing the leakage of privacy information. Outstanding semantic preservation capabilities: the semantic integrity retention rate of the anonymized corpus exceeds 90%, and the performance degradation of downstream tasks is controlled within 2%.
[0110] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0111] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned target corpus data, the target corpus data can also be stored in a blockchain node.
[0112] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0113] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0114] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0115] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0116] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data anonymization device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0117] like Figure 3As shown, the data desensitization device 300 described in this embodiment includes: an identification module 301, an acquisition module 302, a first processing module 303, a second processing module 304, a desensitization module 305, an evaluation module 306, and an output module 307. Wherein: The recognition module 301 is used to receive input corpus data and perform entity recognition on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; The acquisition module 302 is used to acquire the business type corresponding to the corpus data and acquire the target semantic template corresponding to the business type from the preset template library; The first processing module 303 is used to process the sensitive entities in the corpus data based on the target semantic template using a preset semantic fidelity rewriting engine to generate corresponding replacement schemes. The second processing module 304 is used to obtain the sensitivity assessment information corresponding to the sensitive entity and query the target desensitization strategy corresponding to the sensitivity assessment information from the preset strategy library. The desensitization module 305 is used to desensitize the sensitive entities in the corpus data based on the target desensitization strategy and the replacement scheme to obtain the desensitized target corpus data. Evaluation module 306 is used to perform data evaluation on the target corpus data; The output module 307 is used to output the target corpus data if the target corpus data passes the data evaluation.
[0118] In some optional implementations of this embodiment, the first processing module 303 includes: The first calling submodule is used to call the preset controllable text generator and numerical relationship preserver based on the semantic fidelity rewriting engine; The rewriting submodule is used to rewrite the sensitive entities in the corpus data based on the template specification information in the target semantic template, using the controllable text generator to obtain the corresponding first processed data. The transformation submodule is used to perform consistency transformation processing on the specified values related to the specified business in the first processed data based on the numerical relationship maintainer, so as to obtain the corresponding second processed data. The first determining submodule is used to use the second processed data as the replacement scheme.
[0119] In some optional implementations of this embodiment, the transformation submodule includes: The identification unit is used to identify a specified value in the first processed data that is related to a specified business based on the numerical relationship maintainer. The first acquisition unit is used to acquire the correlation information of the specified value; wherein, the correlation information includes dimensional relationships and statistical characteristics; A transformation unit is used to transform the specified value in the first processed data using the associated information based on a preset transformation strategy, so as to obtain the processed transformed data. A determining unit is used to use the transformed data as the second processed data.
[0120] In some optional implementations of this embodiment, the second processing module 304 includes: The second calling submodule is used to call the preset sensitivity dynamic evaluator; The acquisition submodule is used to acquire the entity type, context, and data usage scenario of a specified entity based on the sensitivity dynamic evaluator; wherein, the specified entity is any one of all the sensitive entities; The quantization submodule is used to quantize the entity type, the context, and the data usage scenario based on a preset quantization strategy to obtain the corresponding quantized data. The calculation submodule is used to perform comprehensive calculations on the quantized data to obtain the corresponding quantized score data. The mapping submodule is used to map the quantized score data based on a preset sensitivity mapping strategy to obtain the specified sensitivity evaluation information of the specified entity.
[0121] In some optional implementations of this embodiment, the evaluation module 306 includes: The first evaluation submodule is used to evaluate the degree of semantic preservation of the target corpus data based on the corpus data using a preset semantic integrity evaluator; The second evaluation submodule is used to evaluate the quality of the target corpus data based on the target corpus data using a preset task validator if the target corpus data passes the semantic preservation degree evaluation. The first determination submodule is used to determine that the target corpus data passes the data evaluation if the target corpus data passes the quality evaluation. The second determination submodule is used to determine that the target corpus data has failed the data evaluation if the target corpus data fails the quality evaluation.
[0122] In some optional implementations of this embodiment, the second evaluation submodule includes: The training unit is used to train the downstream task model related to the specified business using the corpus data based on the task validator to obtain the corresponding first task model. The fine-tuning unit is used to fine-tune the downstream task model using the target corpus data to obtain the corresponding second task model; The second acquisition unit is used to acquire the first model performance data of the first task model and the second model performance data of the second task model. The analysis unit is used to perform comparative analysis on the performance differences between the first model performance data and the second model performance data to obtain the corresponding performance difference results. The generation unit is used to analyze the performance difference results to generate a quality assessment result corresponding to the target corpus data.
[0123] In some optional implementations of this embodiment, the data desensitization device further includes: The collection module is used to collect all operational information of the corpus data during the desensitization process; The third processing module is used to organize and classify all the operation information to obtain the corresponding processing information; The construction module is used to construct a corresponding de-identified audit report based on the processed information; The storage module is used to store and process the de-identified audit report.
[0124] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0125] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0126] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0127] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for data anonymization methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0128] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, such as executing computer-readable instructions for the data desensitization method.
[0129] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0130] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the data desensitization method described above.
[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0132] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A data anonymization method, characterized in that, Includes the following steps: Receive input corpus data, and perform entity recognition on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; Obtain the business type corresponding to the corpus data, and obtain the target semantic template corresponding to the business type from the preset template library; Based on the target semantic template, a preset semantic fidelity rewriting engine is used to process the sensitive entities in the corpus data to generate corresponding replacement schemes. Obtain the sensitivity assessment information corresponding to the sensitive entity, and query the target desensitization strategy corresponding to the sensitivity assessment information from the preset strategy library; Based on the target desensitization strategy and the replacement scheme, the sensitive entities in the corpus data are desensitized to obtain the desensitized target corpus data. Data evaluation is performed on the target corpus data; If the target corpus data passes the data evaluation, then the target corpus data is output.
2. The data anonymization method according to claim 1, characterized in that, The step of processing sensitive entities in the corpus data and generating corresponding replacement schemes based on the target semantic template using a preset semantic fidelity rewriting engine specifically includes: The semantically faithful rewriting engine calls a pre-defined controllable text generator and numerical relationship preserver. Based on the template specification information in the target semantic template, the controllable text generator is used to rewrite the sensitive entities in the corpus data to obtain the corresponding first processed data. Based on the numerical relationship maintainer, a consistency transformation is performed on the specified values related to the specified business in the first processed data to obtain the corresponding second processed data. The second processed data is used as the replacement scheme.
3. The data anonymization method according to claim 2, characterized in that, The step of performing consistency transformation on specified values related to a specified business in the first processed data based on the numerical relationship maintainer to obtain the corresponding second processed data specifically includes: Based on the numerical relationship maintainer, a specified value related to a specified business is identified in the first processed data; Obtain the correlation information of the specified value; wherein, the correlation information includes dimensional relationships and statistical characteristics; Based on a preset transformation strategy, the associated information is used to transform the specified value in the first processed data to obtain the processed transformed data. The transformed data is used as the second processed data.
4. The data anonymization method according to claim 1, characterized in that, The step of obtaining the sensitivity assessment information corresponding to the sensitive entity specifically includes: Invoke the preset sensitivity dynamic evaluator; The sensitivity dynamic evaluator is used to obtain the entity type, context, and data usage scenario of a specified entity; wherein, the specified entity is any one of all the sensitive entities. The entity type, the context, and the data usage scenario are quantified based on a preset quantization strategy to obtain the corresponding quantified data. The quantified data is comprehensively calculated and processed to obtain the corresponding quantified score data; The quantified score data is mapped based on a preset sensitivity mapping strategy to obtain the specified sensitivity assessment information of the specified entity.
5. The data anonymization method according to claim 1, characterized in that, The step of evaluating the target corpus data specifically includes: Based on the corpus data, a preset semantic integrity evaluator is used to evaluate the degree of semantic preservation of the target corpus data; If the target corpus data passes the semantic preservation assessment, then a preset task validator is used to perform a quality assessment on the target corpus data based on the corpus data. If the target corpus data passes the quality assessment, then the target corpus data is determined to have passed the data assessment. If the target corpus data fails the quality assessment, then the target corpus data is deemed to have failed the data assessment.
6. The data anonymization method according to claim 5, characterized in that, The step of evaluating the quality of the target corpus data using a preset task validator based on the corpus data specifically includes: Based on the task validator, the corpus data is used to train the downstream task model related to the specified business to obtain the corresponding first task model; The downstream task model is fine-tuned using the target corpus data to obtain the corresponding second task model; Obtain the first model performance data of the first task model, and obtain the second model performance data of the second task model; A comparative analysis of the performance differences between the first model performance data and the second model performance data is performed to obtain the corresponding performance difference results; The performance difference results are analyzed to generate quality assessment results corresponding to the target corpus data.
7. The data anonymization method according to claim 1, characterized in that, After the step of outputting the target corpus data, the method further includes: Collect all operational information of the corpus data during the desensitization process; All the aforementioned operation information is organized and categorized to obtain corresponding processing information; A corresponding de-identification audit report is constructed based on the processed information; The de-identified audit report is stored.
8. A data anonymization device, characterized in that, include: The recognition module is used to receive input corpus data and perform entity recognition on the corpus data based on a preset entity recognition model to obtain sensitive entities in the corpus data; The acquisition module is used to acquire the business type corresponding to the corpus data and acquire the target semantic template corresponding to the business type from the preset template library; The first processing module is used to process the sensitive entities in the corpus data based on the target semantic template using a preset semantic fidelity rewriting engine to generate corresponding replacement schemes. The second processing module is used to obtain the sensitivity assessment information corresponding to the sensitive entity, and to query the target desensitization strategy corresponding to the sensitivity assessment information from the preset strategy library. The desensitization module is used to desensitize the sensitive entities in the corpus data based on the target desensitization strategy and the replacement scheme to obtain the desensitized target corpus data. The evaluation module is used to perform data evaluation on the target corpus data; The output module is used to output the target corpus data if the target corpus data passes the data evaluation.
9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data desensitization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data desensitization method as described in any one of claims 1 to 7.