Iteration-based data privacy protection processing method and device and storage medium
By sharding and iteratively anonymizing the dataset, the challenge of adaptively adjusting data availability and privacy protection in data anonymization publishing is solved, and efficient data privacy protection and resource optimization are achieved.
Patent Information
- Application Number
- CN202510844670.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have difficulty adaptively adjusting data availability and privacy protection during the process of data anonymization and publication, and are particularly challenging in optimizing the efficiency of high-dimensional large data sets.
After the original dataset is sharded, it is anonymized iteratively, including data-sensitive attribute identification, k-anonymous data subset screening and generalization processing, to optimize the generalization model to meet the k-anonymity conditions, and finally aggregate the anonymization results.
It has achieved the goal of efficiently meeting the anonymization requirements of data privacy protection in the data sharing system of the big data center, optimizing the balance between resource consumption and data gains and losses, and ensuring the privacy protection and availability of the data set.
Smart Images

Figure CN120705909A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data security technology, and in particular to an iteration-based data privacy protection processing method, device, and storage medium. Background Art
[0002] In the era of big data, balancing data availability and privacy protection, based on actual application scenarios, is a key issue in data security. While maintaining a certain level of data availability and statistical accuracy, data desensitization can be used to reduce data sensitivity through distortion and other transformations; anonymization can be used to protect privacy through de-identification; differential privacy can be achieved through noise addition to resist differential attacks; and homomorphic encryption can even be used to directly encrypt sensitive personal information and then perform statistical analysis and machine learning on the ciphertext.
[0003] Currently, there are many mature research results in privacy protection for anonymized data release. Anonymization design principles such as k-anonymity, l-diversity, and t-approximation are widely used. K-anonymity requires that the released data contain at least k records that are indistinguishable by quasi-identifiers, preventing attackers from identifying the specific individual to whom the private information belongs, thereby protecting individual privacy. K-anonymity uses the parameter k to specify the maximum risk of information leakage that a user can tolerate.
[0004] However, some of these technologies still face many challenges when implemented in specific scenarios, such as how to achieve adaptive adjustment of data availability and privacy protection; how to optimize the efficiency of high-dimensional and large data sets, etc., which are issues worthy of in-depth study. Summary of the Invention
[0005] The present application provides an iterative data privacy protection processing method, device and storage medium to at least solve the above technical problems existing in the prior art.
[0006] According to a first aspect of the present application, an iterative data privacy protection processing method is provided, comprising the following steps: S1, shard the original data set to obtain several sharded data sets; S2, allocating the sharded data set to different nodes; S2, performing iterative anonymization data processing on the sharded data set; S3, aggregates the data sets output by each iteration of anonymized data processing.
[0007] In certain embodiments of the first aspect of the present application, the iterative anonymized data processing includes: identifying data sensitive attributes, screening k anonymous data subsets, generalizing the remaining data subsets and then screening out the data subsets that meet k anonymity, and iterating the above process until all data processing is completed.
[0008] In certain embodiments of the first aspect of the present application, the method for identifying data sensitive attributes is as follows: Use the metadata table of sensitive entities predefined in the metadata warehouse to identify sensitive entity attributes of the sharded dataset and mark the sensitive attribute sequence of the sharded dataset as {D1, D2, D3...Dd}.
[0009] In certain embodiments of the first aspect of the present application, the sharded data set is further subjected to dimensionality reduction processing to remove redundant attributes of the sharded data set; if a certain attribute in the sensitive attribute sequence of the sharded data set {D1, D2, D3...Dd} can be obtained by calculating other attributes, then the attribute is removed.
[0010] In certain embodiments of the first aspect of the present application, the method for screening the k anonymous data subsets is as follows: For the sharded data set, the number of tuples in the minimum cube composed of all attribute dimensions is calculated; when the number of tuples in a minimum cube is greater than or equal to the set k, the data subset composed of the tuple set in this minimum cube meets the k-anonymity condition, and all minimum cubes that meet the condition are filtered out.
[0011] In certain embodiments of the first aspect of the present application, the generalization method is as follows: Generalized attribute priority selection: Calculate the dissimilarity metric α for each column vector value of the sensitive attribute sequence {D1, D2, D3...Dd}; use the data processing grouping function to calculate the column vector, calculate the number of groups to be x, and calculate the size of each group to be m, denoted as sizeOf(group(i))=m; define α=x / l, where l is the total length of the column vector; sort the attribute sequences of the data subset from largest to smallest according to α as the pre-selected priority of the pre-selected generalization attributes, and select the attribute generalization with the highest pre-selected priority; Generalization processing method: For discrete-valued attributes, find the associated dimension table according to the definition of the sensitive entity metadata database, extract the hierarchical structure of the dimension table, replace the original attribute value with the attribute value of the previous level, and record the hierarchical level of the current attribute dimension. This completes the generalization of discrete-valued attributes. For numerical attributes, linear fuzzy or nonlinear fuzzy is used according to the characteristics of the numerical attributes to generalize the attribute value into an interval range; Save the generalized model: The selection of attributes corresponding to each generalization, the size of the fuzzy coefficient α, and the generalization interval parameter are saved as the generalization model corresponding to the shard data subset; Optimization of generalization model: Select a certain amount of sample data and optimize the generalization model of the entire original data set to the anonymized data set. Adjust the execution order of generalized attributes, the generalization fuzzy coefficient of numerical attributes, and the granularity of the dimensional hierarchy of discrete attributes. Calculate the data processing time and loss value of the anonymized sample data after each adjustment. When the data processing time and loss value meet the expected efficiency and profit and loss of the calculation, the model optimization is completed.
[0012] In certain embodiments of the first aspect of the present application, the linear blurring method is as follows: Specify the size of the fuzzy coefficient α and blur the attribute value w to the interval range of interval size R: [Mod(w / R)*R,(Mod(w / R)+1)*R], where R=(Max(value range of Di attribute)-Min(value range of Di attribute)) / α.
[0013] In certain embodiments of the first aspect of the present application, the nonlinear fuzzy method is as follows: Specify the size of the fuzzy coefficient α, sort the values, and divide them into R intervals according to the percentile of the values: R=mod(l / α); where l is the total number of records in the Di attribute column vector; Fuzzifies the attribute value into the percentile interval based on the upper and lower bounds of the interval.
[0014] According to a second aspect of the present application, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in this application.
[0015] According to a third aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the present application.
[0016] Compared with the prior art, this application has the following beneficial effects: 1. This application splits the original data set into shards and performs iterative anonymization data processing on each shard data, which can meet the requirements of online big data center data sharing system for data privacy protection and anonymization.
[0017] 2. The iterative anonymized data processing of this application includes data-sensitive attribute identification and dimensionality reduction, k-anonymous data subset screening, and the remaining subsets are screened out again after generalizing the preferred attributes to obtain the data subsets that meet the k-anonymity requirements. The above process is iterated until all data are processed. Finally, after the distributed task output data sets are aggregated, the result data set that meets the K-anonymity requirements is obtained. The processing model can tune the model parameters by evaluating the calculation data processing time and loss value to meet the expected balance between resource consumption and data profit and loss.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an illustrative and non-limiting manner, in which: In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0020] Figure 1 A schematic diagram of data publishing across different security domains of the present application is shown.
[0021] Figure 2 A schematic diagram of anonymized data processing in this application is shown.
[0022] Figure 3 A schematic diagram of the sensitive entity metadata database of this application is shown.
[0023] Figure 4 A schematic diagram of the iterative processing flow of this application is shown.
[0024] Figure 5 A schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] In order to make the purpose, features, and advantages of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0026] Example 1: Data processing begins with data collection at the source layer and goes through a series of cleaning, transformation, and implementation steps. The resulting datasets at each step may be output to different computing units for further analysis and mining. Data processing units may be located in different security domains. For security domains with lower privacy permissions, the datasets obtained must meet security standards that prevent inferences from being made to specific entities. Therefore, at the boundaries of security domains, upstream datasets must be processed for privacy protection before being output to downstream datasets.
[0027] This embodiment provides an iterative data privacy protection method. Figure 1 , including the following steps: S1, the original dataset of the data owner is of relatively large data size and involves many entity dimensions of privacy-preserving calculations. Therefore, the original dataset is first sharded using the map method to obtain several sharded datasets.
[0028] S2: Allocate the sharded data set to different data nodes.
[0029] S3, in different data nodes, performs iterative anonymization job processing on the sharded data set.
[0030] Please refer to Figure 2 The following is a detailed introduction to the iterative anonymization job processing process, which mainly includes: identifying data sensitive attributes, screening k anonymous data subsets, generalizing the remaining data subsets and then screening out the data subsets that meet the k anonymity requirements, and iterating the above process until all data are processed.
[0031] Specifically, the method for identifying data sensitive attributes is as follows: Please refer to Figure 3 For the n*d original dataset Set, which contains n data records and d attributes, each attribute is a column vector Di(i∈1,d) consisting of n data points. Using the metadata table of sensitive entities predefined in the metadata warehouse, sensitive entity attributes are identified for the sharded dataset S.
[0032] When a sharded dataset S contains predefined sensitive entity attributes, the sensitive attribute sequence of the sharded dataset S is marked as {D1, D2, D3...Dd}, which is recorded as the attribute subset to be determined in the original dataset Set. It should be noted that the original data of the data owner typically encrypts or removes sensitive attributes that are highly recognizable, especially those that uniquely identify entities.
[0033] In order to reduce the amount of data calculation in subsequent steps, this solution also performs dimensionality reduction on the sharded dataset S. The main step is to remove redundant attributes of the sharded dataset S. That is, if an attribute D1 in the sharded dataset S can be obtained by calculating other attributes, then the attribute D1 is removed and the data subset after dimensionality reduction is renamed S.
[0034] The method for filtering k anonymous data subsets is as follows: Perform a fast K-anonymity check on dataset S. The K-anonymity algorithm is defined as follows: If a dataset Si with the sensitive attribute sequence {Di, Di+1, Di+2...} satisfies the K-anonymity definition, then the attributes of each record in the dataset cannot be distinguished from at least (K-1) other data records. This step groups dataset S according to the sensitive attribute sequence {D1, D2, D3...Dd} group dimension and counts the number of identical values in the grouped records. If the number of identical values is greater than or equal to K, dataset S is considered to meet the K-anonymity definition.
[0035] For a sharded dataset Sn with n attributes, calculate the number of data tuples in the minimum cube formed by all attribute dimensions. Take three-dimensional attributes as an example, that is, a cube formed by Dim1, Dim2, and Dim3. Each minimum cube represents a data tuple (v1, v2, v3). When the number of tuples in a cube is greater than or equal to k, the data subsets cube1, cube2, etc. formed by the set of tuples in this minimum cube meet the K anonymity condition. All data cubes that meet the condition are screened out and recorded as the dataset S'1=cube1∪cube2∪cube3.
[0036] Let the remaining data subset be S'', that is, S''=S-S'1.
[0037] The remaining data subsets are generalized as follows: Generalized attribute priority selection Calculate the dissimilarity metric α for each column vector value in the sensitive attribute sequence {D1, D2, D3...Dd}. Specifically, use the data processing grouping function for the column vectors, calculate the number of groups x, and calculate the size of each group m, denoted by sizeOf(group(i))=m, (i∈0,x). Define α=x / l (α<=1), where l=sizeOf(Di) is the total length of the column vector. Sort the attribute sequence of the dataset S' by α from largest to smallest, and use this as the preselection priority for the preselected generalization attributes. Select the attribute generalization with the highest preselection priority.
[0038] Generalization processing method For discrete-valued attributes, find its associated dimension table (if it exists) according to the definition of the sensitive entity metadata database, retrieve the hierarchical structure of the dimension table, replace the original value of the attribute with the attribute value of the previous level, and record the hierarchical level of the current attribute dimension. This completes the generalization of the discrete-valued attributes.
[0039] For numerical attributes, linear fuzzy, nonlinear fuzzy and other methods are used according to the characteristics of the numerical attributes to generalize the attribute value to an interval range.
[0040] Linear fuzzy method: specify the size of the fuzzy coefficient α, and blur the attribute value w to an interval range of interval size R: [Mod(w / R)*R,(Mod(w / R)+1)*R], where R=(Max(value range of Di attribute)-Min(value range of Di attribute)) / α.
[0041] Nonlinear fuzzy method: The following is an optional fuzzy method based on distributed data: specify the size of the fuzzy coefficient α, sort the values and divide them into R intervals according to the percentile of the values: R=mod(l / α) where l is the total number of records in the Di attribute column vector.
[0042] Fuzzifies the attribute value into the percentile interval based on the upper and lower bounds of the interval.
[0043] Save the generalized model The selection of attributes, fuzzy coefficient size α, generalization interval and other parameters for each generalization are saved as the generalization model corresponding to the data set. The loss value θi of each step of generalization is calculated as follows: (numeric attribute) Linear fuzzy method: θi=1 / α Nonlinear fuzzy method, θi = length of Σ interval [ri, ri+1) * number of generalized data records in the interval / length of the value range of Di attribute * total number of records of Di attribute column vector (Discrete-valued attribute) θi=(Σ(1 / the number of data records before generalization corresponding to the generalized attribute target discrete value)) / the total number of records of the Di attribute column vector Optimization of generalization models Select a certain amount of sample data and optimize the generalization model from the original dataset to the anonymized dataset. Adjust the execution order of generalized attributes, the generalization fuzziness coefficient of numerical attributes, and the granularity of the dimensional hierarchy of discrete attributes. Calculate the data processing time T for each anonymized sample data adjustment, with the loss value θ = ∏θi(i∈0,n)n representing the total number of iterations. Model optimization is complete when the time T and θ meet the expected efficiency and profit / loss.
[0044] After generalization processing of the remaining data subsets, the data subsets that meet k-anonymity are screened out again. The above process is iterated until all data are processed. The content is as follows: Please refer to Figure 4 , each iterative process screens out subsets that meet the K anonymity conditions and records them as S'1, S'2, S'3...; When all data subsets S'' are screened out (the worst budget is that S'' is a one-dimensional data vector), the iteration terminates.
[0045] S4, after aggregating the datasets output by each iteration of the anonymization job processing task, all subsets that meet the K-anonymity requirements are merged to obtain the result set W that meets the privacy protection conditions, W = S'1 + S'2 + S'3. The final dataset meets the anonymization requirements.
[0046] Example 2: According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0047] Figure 5 A schematic block diagram of an example electronic device that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0048] like Figure 5 As shown, the device includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). The RAM can also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. An input / output (I / O) interface is also connected to the bus.
[0049] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.
[0050] The computing unit can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing units include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit performs the various methods and processes described above, such as the iteration-based data privacy protection method described in Example 1. For example, in some embodiments, the iteration-based data privacy protection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the computing unit, one or more steps of the iteration-based data privacy protection method described above can be performed. Alternatively, in other embodiments, the computing unit can be configured to perform the iteration-based data privacy protection method via any other suitable means (e.g., via firmware).
[0051] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0052] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0053] In the context of this application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0054] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0055] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0056] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0057] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
[0058] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0059] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data privacy protection processing method based on iteration, characterized in that: The following steps are involved: S1, shard the original data set to obtain several sharded data sets; S2, allocating the sharded data set to different nodes; S2, performing iterative anonymization data processing on the sharded data set; S3, aggregates the data sets output by each iteration of anonymized data processing.
2. The data privacy protection processing method based on iteration according to claim 1, characterized in that: The iterative anonymized data processing includes: identifying data sensitive attributes, screening k anonymous data subsets, generalizing the remaining data subsets and then screening out the data subsets that meet the k anonymity requirements, and iterating the above process until all data are processed.
3. The data privacy protection processing method based on iteration according to claim 2, characterized in that: The method for identifying data sensitive attributes is as follows: Use the metadata table of sensitive entities predefined in the metadata warehouse to identify sensitive entity attributes of the sharded dataset and mark the sensitive attribute sequence of the sharded dataset as {D1, D2, D3...Dd}.
4. The data privacy protection processing method based on iteration according to claim 3 is characterized in that: The sharded data set is further subjected to dimensionality reduction processing to remove redundant attributes of the sharded data set; if a certain attribute in the sensitive attribute sequence of the sharded data set {D1, D2, D3...Dd} can be obtained by calculating other attributes, the attribute is removed.
5. The data privacy protection processing method based on iteration according to claim 2, characterized in that: The method for screening the k anonymous data subsets is as follows: For the sharded data set, the number of tuples in the minimum cube composed of all attribute dimensions is calculated; when the number of tuples in a minimum cube is greater than or equal to the set k, the data subset composed of the tuple set in this minimum cube meets the k-anonymity condition, and all minimum cubes that meet the condition are filtered out.
6. The data privacy protection processing method based on iteration according to claim 5 is characterized in that: The generalization method is as follows: Generalized attribute priority selection: Calculate the dissimilarity metric α for each column vector value of the sensitive attribute sequence {D1, D2, D3...Dd}; use the data processing grouping function to calculate the column vector, calculate the number of groups to be x, and calculate the size of each group to be m, denoted as sizeOf(group(i))=m; define α=x / l, where l is the total length of the column vector; sort the attribute sequences of the data subset from largest to smallest according to α as the pre-selected priority of the pre-selected generalization attributes, and select the attribute generalization with the highest pre-selected priority; Generalization processing method: For discrete-valued attributes, find the associated dimension table according to the definition of the sensitive entity metadata database, extract the hierarchical structure of the dimension table, replace the original attribute value with the attribute value of the previous level, and record the hierarchical level of the current attribute dimension. This completes the generalization of discrete-valued attributes. For numerical attributes, linear fuzzy or nonlinear fuzzy is used according to the characteristics of the numerical attributes to generalize the attribute value into an interval range; Save the generalized model: The selection of attributes corresponding to each generalization, the size of the fuzzy coefficient α, and the generalization interval parameter are saved as the generalization model corresponding to the shard data subset; Optimization of generalization model: Select a certain amount of sample data and optimize the generalization model of the entire original data set to the anonymized data set. Adjust the execution order of generalized attributes, the generalization fuzzy coefficient of numerical attributes, and the granularity of the dimensional hierarchy of discrete attributes. Calculate the data processing time and loss value of the anonymized sample data after each adjustment. When the data processing time and loss value meet the expected efficiency and profit and loss of the calculation, the model optimization is completed.
7. The data privacy protection processing method based on iteration according to claim 6 is characterized in that: The linear fuzzy method is as follows: Specify the size of the fuzzy coefficient α and blur the attribute value w to the interval range of interval size R: [Mod(w / R)*R,(Mod(w / R)+1)*R], where R=(Max(value range of Di attribute)-Min(value range of Di attribute)) / α.
8. The data privacy protection processing method based on iteration according to claim 6 is characterized in that: The nonlinear fuzzy method is as follows: Specify the size of the fuzzy coefficient α, sort the values, and divide them into R intervals according to the percentile of the values: R=mod(l / α); where l is the total number of records in the Di attribute column vector; Fuzzifies the attribute value into the percentile interval based on the upper and lower bounds of the interval.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.