Method, device and electronic device for de-identifying personal information
By generating and segmenting biometric values, establishing a mapping relationship between sub-biological characteristics and data locations, and performing data location transformation, the shortcomings in data availability and privacy protection in the prior art are solved, and the security de-identification of personal information is realized.
Patent Information
- Application Number
- CN202010672803.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2040-07-14
AI Technical Summary
The existing personal information de-identification technology has shortcomings in protecting data availability and preventing re-identification attacks, resulting in a high risk of personal information leakage.
By obtaining the user's biometric information, generating biometric values, segmenting them according to the correlation relationship between data, establishing a mapping relationship between sub-biological characteristics and data location, and performing data position transformation, maintaining the correlation of associated data and cutting off the correlation of non-associated data.
On the premise of protecting data availability, it effectively prevents personal information leakage, resists heavy identification attacks, and ensures data privacy and security.
Smart Images

Figure CN113934999B_ABST
Abstract
Description
Technical field
[0001] The present application relates to the field of information security technology, and in particular to a method, device and electronic device for de-identifying personal information. [Background Technology]
[0002] With the rapid development of information technology and big data applications, more and more people are realizing the value of data and the significance of open data sharing. However, this open data sharing also raises concerns about the security of personal information. Data collected by government agencies, businesses, and other organizations often contains personal information such as names, phone numbers, and ID numbers. Directly publishing this raw data can lead to serious personal information leaks.
[0003] Personal information de-identification refers to the process of removing the association between a set of identifiable data and the individuals they correspond to, in order to prevent personal information leakage. Based on the data's attributes, personal information can be categorized into three categories: identifiers, quasi-identifiers, and sensitive data. Identifiers refer to information that can directly identify an individual, such as ID numbers and names; quasi-identifiers refer to information that can be used to identify an individual through association with external database tables, such as postal codes, birthdays, and gender; and sensitive data refers to information that users do not wish to be shared, such as salary, medical history, and purchasing preferences. Current personal information de-identification technologies primarily utilize methods such as generalization, suppression, encryption, and k-anonymity to address these three categories of information. However, existing technologies suffer from two major issues: First, while current de-identification technologies make it difficult to identify the individuals associated with the information, they also compromise the usability of quasi-identifiers and sensitive data, hindering data analysis; second, current de-identification technologies carry a high risk of re-identification, meaning that de-identified personal information is susceptible to re-identification attacks, such as linking attacks and background knowledge attacks, leading to personal information leakage. [Summary of the invention]
[0004] The embodiments of the present application provide a method, device, and electronic device for de-identifying personal information, which are used to de-identify personal information while protecting data availability, and effectively prevent the leakage of user personal information.
[0005] In a first aspect, an embodiment of the present application provides a method for de-identifying personal information, comprising: obtaining biometric information of a user accessing target data, and generating a biometric value based on the biometric information;
[0006] The biometric value is segmented based on the association relationship between the data included in the target data to obtain multiple sub-biometric values; a corresponding relationship is established between each data in the target data and the multiple sub-biometric values, wherein the sub-biometric values are mapped to the data positions; a data transformation position is determined based on the data position mapped to the sub-biometric value corresponding to each data in the target data, and the position of each data in the target data is transformed based on the data transformation position.
[0007] In one possible implementation, generating a biometric value based on the biometric information includes: converting the biometric information of the user accessing the target data into binary data;
[0008] A hash calculation is performed on the binary data or the transformed data after the binary data has been transformed at least once to obtain a biometric feature value of a preset length.
[0009] In one possible implementation, segmenting the biometric value based on the association relationship between the data contained in the target data includes: determining the data attributes of each data contained in the target data; and segmenting the biometric value based on the number of the data attributes and the association relationship between the data attributes.
[0010] In one possible implementation, segmenting the biometric value based on the number of data items contained in the target data and the association relationship between the data items includes: determining associated attributes and independent attributes in the target data; wherein the associated attributes include at least two data items in the target data, and the at least two data items have an association relationship; and the independent attributes have no association relationship with other data items in the target data; determining the number of segments of the biometric value based on the sum of the number of associated attributes and independent attributes in the target data, and segmenting the biometric value according to the number of segments.
[0011] In one possible implementation, segmenting the biometric value according to the number of segments includes: dividing the biometric value equally or randomly according to the number of segments based on the data length of the biometric value; wherein the number of segments is greater than or equal to the sum of the number of associated attributes and independent attributes in the target data.
[0012] In one possible implementation, the establishing of correspondences between each data in the target data and the multiple sub-biometric values includes: the at least two data attributes included in each associated attribute correspond to the same sub-biometric value; each independent attribute corresponds to a sub-biometric value; and each data in the target data corresponds to the same sub-biometric value as the data attribute to which it belongs.
[0013] In one possible implementation, determining the data transformation position based on the data position of the sub-biometric value mapping corresponding to each data in the target data includes: weighting the coordinate position of each data in the target data in the data table with the data position of the sub-biometric value mapping corresponding to each data to obtain the transformation position of each data in the target data.
[0014] In a second aspect, an embodiment of the present application provides a personal information de-identification device, comprising: a generation module for obtaining biometric information of a user accessing target data, and generating a biometric value based on the biometric information; a segmentation module for segmenting the biometric value based on the association relationship between the data contained in the target data to obtain multiple sub-biometric values; a determination module for determining the correspondence between each data in the target data and the multiple sub-biometric values, wherein the sub-biometric values are mapped to the data position; a position transformation module for determining a data transformation position based on the data position mapped to the sub-biometric value corresponding to each data in the target data, and transforming the position of each data in the target data according to the data transformation position.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method described above.
[0016] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method as described above.
[0017] In the above technical solution, biometric information of the user accessing the target data is obtained to generate a biometric value. Based on the associations between the data contained in the target data, the biometric value is segmented to obtain multiple sub-biometric values. Correspondences are established between each data item in the target data and the multiple sub-biometric values. Based on the data positions mapped to the sub-biometric values, the data transformation positions are determined, and the positions of each data item in the target data are transformed. This data position transformation maintains the association between associated data, thereby protecting data availability, while simultaneously breaking the association between unassociated data. This de-identifies personal information, effectively defends against re-identification attacks, and prevents the leakage of user personal information.
Brief Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of an embodiment of a method for de-identifying personal information for this application;
[0020] Figure 2 A diagram of a data sheet showing the method for de-identifying personal information for this application;
[0021] Figure 3 This is a schematic diagram of the structure of an embodiment of the personal information de-identification device for this application;
[0022] Figure 4 This is a structural diagram of an embodiment of the electronic device of the present application. [Specific implementation method]
[0023] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0024] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0025] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0026] Figure 1 This is a flowchart of an embodiment of the method for de-identifying personal information of this application.
[0027] In this embodiment, the personal information de-identification method of this application can be performed on user data contained in a data table. The user data in the data table may include name, age, gender, address, physical condition, income level, etc. Specifically, user data can be categorized into identifiers, quasi-identifiers, and sensitive data based on their attributes.
[0028] In a specific implementation process, the quasi-identifiers and sensitive data in the user data can be used as target data to execute the personal information de-identification method of this application. Figure 1 As shown, the personal information de-identification method may include:
[0029] Step 101: Obtain biometric information of a user accessing target data, and generate a biometric value based on the biometric information.
[0030] In this embodiment, the target data may be shared information stored in the system. When a user needs to obtain the shared information in the system, the user may input personal biometric information to access the shared information in the system.
[0031] In this embodiment, when the user accesses the target data, the user's biometric information can be obtained, including but not limited to: the user's DNA, fingerprints, voiceprints, facial features and other uniquely identifying biometric features, and the obtained biometric information can be converted into a biometric value.
[0032] Specifically, biometric information can be converted into binary data. Optionally, the resulting binary data can be transformed based on actual needs. Transformations may include, but are not limited to, negating the binary data, complementing the binary data, and adding or subtracting random numbers. Furthermore, to protect the user's personal information, the irreversibility of the hash algorithm can be leveraged to perform a hash calculation on the resulting binary data to obtain an encrypted biometric value. The hash algorithm can also compress the binary data, compressing the corresponding binary data of different accessing users into biometric values of equal length.
[0033] Step 102: Segment the biometric value according to the association relationship between the data included in the target data to obtain multiple sub-biometric values.
[0034] When segmenting the biometric feature values, first, the data attributes of each data in the target data are determined.
[0035] Furthermore, the associated attributes and independent attributes in the data attributes are determined.
[0036] Specifically, based on the specific application scenario of the information, data attributes that need to be associated with each other can be identified as associated attributes, and associated attributes must contain at least two data attributes. Data attributes that do not need to be associated with other data attributes can be identified as independent attributes. For example, when the application scenario of the information is to analyze the relationship between age and the incidence of heart disease, the two data attributes of age and whether or not the person has heart disease need to be associated and are associated attributes. However, data attributes such as gender, address, and educational background do not need to be associated and are independent attributes.
[0037] Next, the number of segments of the biometric value is determined based on the sum of the number of associated attributes and independent attributes, and the biometric value is segmented according to this number. It should be noted that when counting associated attributes, regardless of the number of data attributes contained in the associated attribute, it is counted as one associated attribute. Using the above example again, if age and heart disease status are associated attributes, the number of associated attributes is counted as one.
[0038] Step 103 : establishing corresponding relationships between each data in the target data and a plurality of sub-biometric feature values, wherein the sub-biometric feature values are mapped to data positions.
[0039] In this embodiment, at least two data attributes included in each associated attribute correspond to the same sub-biometric value; each independent attribute corresponds to a sub-biometric value; and each data item in the target data and the data attribute to which it belongs correspond to the same sub-biometric value. The sub-biometric value is mapped to the data location.
[0040] Specifically, mapping the sub-biometric value to the data location may be to map the sub-biometric value to the data location; or mapping the sub-biometric value to the data location, and obtaining the data location by looking up the table based on the sub-biometric value.
[0041] Step 104 : determining a data transformation position according to the data position of the sub-biometric feature value mapping corresponding to each data in the target data, and transforming the position of each data in the target data according to the data transformation position.
[0042] The coordinate position of each data in the target data in the data table is weighted with the data position mapped to the sub-biometric feature value corresponding to each data to obtain the transformed position of each data in the target data.
[0043] Specifically, first, the value of the sub-biometric characteristic value corresponding to each data in the target data can be used as the data position of the sub-biometric characteristic value mapping.
[0044] Then, the coordinate position of each data in the target data in the data table is weighted with the value of the sub-biometric feature value corresponding to each data to obtain the transformed position of each data in the target data.
[0045] Finally, the position of each data in the target data is transformed according to the obtained transformed position of the data.
[0046] In this embodiment, biometric information of a user accessing target data is obtained to generate a biometric value. Based on the associations between the data contained in the target data, the biometric value is segmented to obtain multiple sub-biometric values. Correspondences are established between each data item in the target data and the multiple sub-biometric values. Data transformation positions are determined based on the data positions mapped to the sub-biometric values, and the positions of the data in the target data are transformed. Because at least two data attributes contained in the associated attribute correspond to the same sub-biometric value, the target data contained in the associated attribute remains associated after the position transformation, thereby protecting data usability and enabling users to analyze useful information from de-identified information. Furthermore, because each independent attribute corresponds to a different sub-biometric value, the associations between the data contained in the independent attributes are severed, thereby protecting the privacy of the data subject and preventing users accessing the information from stealing the data subject's personal information.
[0047] In another embodiment of the present application, four methods for segmenting biometric feature values are provided.
[0048] In this embodiment, the biometric value is divided equally or randomly according to the number of segments, based on the data length of the biometric value. The number of segments is greater than or equal to the sum of the number of associated attributes and independent attributes in the target data. The specific operation method is as follows:
[0049] Method 1:
[0050] Based on the sum of the number of associated attributes and independent attributes, the number of splits of the biometric value is determined to be equal to the sum of the number of associated attributes and independent attributes. Afterwards, the biometric value can be divided equally according to the number of splits.
[0051] For example, when there are three associated attributes and 17 independent attributes, the sum of the number of associated attributes and independent attributes is 20. The number of biometric value segments is determined to be 20, and the biometric value is divided equally to obtain 20 sub-biometric values of equal length. When establishing the correspondence between the sub-biometric values and the associated attributes and independent attributes, the three sub-biometric values are respectively the three associated attributes and the data contained therein; and the 17 sub-biometric values are respectively the 17 independent attributes and the data contained therein. Specifically, for the three associated attributes and the data contained therein, at least two data attributes contained in each associated attribute and the data contained therein correspond to the same sub-biometric value.
[0052] Method 2:
[0053] According to the sum of the number of associated attributes and independent attributes, it is determined that the number of divisions of the biometric value is greater than the sum of the number of associated attributes and independent attributes, and the biometric value is equally divided according to the number of divisions.
[0054] For example, when there are 3 associated attributes and 17 independent attributes, the sum of the number of associated attributes and independent attributes is 20. The number of biometric value segments is determined to be 25, and the biometric value is divided equally to obtain 25 sub-biometric values of equal length. When establishing the correspondence between the sub-biometric values and the associated attributes and independent attributes, 20 sub-biometric values are selected from the 25 sub-biometric values. 3 sub-biometric values are respectively 3 associated attributes and the data contained therein; and 17 sub-biometric values are respectively 17 independent attributes and the data contained therein. Among them, for the 3 associated attributes and the data contained therein, at least two data attributes contained in each associated attribute and the data contained therein correspond to the same sub-biometric value. Specifically, when selecting 20 sub-biometric values from the 25 sub-biometric values, they can be selected randomly or in order, with the first 20 sub-biometric values being selected.
[0055] Method 3:
[0056] According to the sum of the number of associated attributes and independent attributes, the number of divisions of the biometric value is determined to be equal to the sum of the number of associated attributes and independent attributes, and the biometric value is unequally divided according to the number of divisions.
[0057] For example, when there are three associated attributes and 17 independent attributes, the sum of the number of associated attributes and independent attributes is 20. The number of segments of the biometric value is determined to be 20, and the biometric value is divided unequally to obtain 20 sub-biometric values of unequal lengths. When establishing the correspondence between the sub-biometric values and the associated attributes and independent attributes, the three sub-biometric values are respectively the three associated attributes and the data contained therein; and the 17 sub-biometric values are respectively the 17 independent attributes and the data contained therein. Specifically, for the three associated attributes and the data contained therein, at least two data attributes contained in each associated attribute and the data contained therein correspond to the same sub-biometric value.
[0058] Method 4:
[0059] According to the sum of the number of associated attributes and independent attributes, it is determined that the number of divisions of the biometric value is greater than the sum of the number of associated attributes and independent attributes, and the biometric value is unequally divided according to the number of divisions.
[0060] For example, when there are 3 associated attributes and 17 independent attributes, the sum of the number of associated attributes and independent attributes is 20. The number of segments of the biometric value is determined to be 25, and the biometric value is unequally divided to obtain 25 sub-biometric values of unequal length. When establishing the correspondence between the sub-biometric values and the associated attributes and independent attributes, 20 sub-biometric values are taken from the 25 sub-biometric values, and 3 sub-biometric values are respectively 3 associated attributes and the data contained therein; and 17 sub-biometric values are respectively 17 independent attributes and the data contained therein. Among them, for the 3 associated attributes and the data contained therein, at least two data attributes contained in each associated attribute and the data contained therein correspond to the same sub-biometric value. Specifically, when taking 20 sub-biometric values from the 25 sub-biometric values, they can be randomly selected or the first 20 sub-biometric values can be taken in order.
[0061] In another embodiment of the present application, a method for determining the transformation position of each data in the target data is provided.
[0062] First, determine the coordinate position of each data in the target data in the data table. Optionally, the coordinate position can be the row number of the row where each data is located in the data table, or the column number of the column where it is located, or both the row number and the column number can be used as the coordinate position. This application does not limit this.
[0063] Then, the data position of the sub-biometric feature value mapped to each data in the target data is determined. Optionally, the value of the sub-biometric feature value is used as the data position of the mapping.
[0064] Finally, the row number of each data item is added to the value of the corresponding sub-biometric feature. The resulting value is modulo the total number of rows in the data table to determine the transformed position of each data item. The position of each data item is then transformed. In particular, when the result of the modulo operation is 0, the row number of the transformed position of the data item is equal to the total number of rows.
[0065] Optionally, the column number of each data item is added to the value of the sub-biometric feature corresponding to each data item, and the resulting value is modulo the total number of columns in the data table to determine the transformed position of each data item, and the position of each data item is then transformed. Specifically, when the result of the modulo operation is 0, the column number of the transformed position of the data item is equal to the total number of columns.
[0066] Figure 2 This is a diagram of a data table of the personal information de-identification method of this application. In another embodiment of this application, a specific implementation process of using the personal information de-identification method of this application to achieve personal information de-identification is provided.
[0067] For example.
[0068] First, the fingerprint information M of the user accessing the target data is obtained, generating the fingerprint information binary data Get(M). The obtained binary data is encrypted using a hash algorithm to obtain the biometric value H(M). Here, H(M) is the binary number 111001011100001.
[0069] Secondly, determine the data attributes of each data contained in the target data, and determine the number of associated attributes and independent attributes in the data attributes based on the information application scenario.
[0070] like Figure 2 As shown in the figure, the gender, age, and height data in the target data are determined to be associated attributes, and the address and zip code data are determined to be independent attributes. Therefore, there is one associated attribute and two independent attributes, and the sum of the number of associated attributes and independent attributes is 3.
[0071] The biometric value H(M) is then segmented based on the sum of the number of associated and independent attributes. In this embodiment, H(M) is divided into three equal parts based on the sum of the number of associated and independent attributes (3), resulting in three sub-biometric values: 11100, 10111, and 00001.
[0072] A correspondence is established between the sub-biometric feature value and the associated attribute and the independent attribute, and between the data contained in the sub-biometric feature value and the associated attribute and the independent attribute.
[0073] In this embodiment, if Figure 2As shown, the associated attributes and the data contained therein are associated with 11100, that is, the gender attribute and the data contained therein, the age attribute and the data contained therein, and the height attribute and the data contained therein are associated with 11100. The address data in the independent attribute is associated with 00001, and the postal code data in the independent attribute is associated with 10111. Of course, different association methods can also be selected, which will not be described in detail.
[0074] Finally, according to the coordinate position of each data in the data table and the data position mapped by the sub-biometric feature value, the transformed position of each data is obtained, and the position of each data is transformed.
[0075] In this embodiment, the row number of each data in the target data in the data table is taken as its coordinate position, and the value of the sub-biometric feature value is taken as its mapped data position.
[0076] like Figure 2 As shown, the data table has a total of 10 rows. For example, we use row 3 of the data contained in the associated attribute as an example. The corresponding sub-biometric value 11100 is 28. For the address attribute among the independent attributes, we use row 4 of the address data as an example. The corresponding sub-biometric value 00001 is 1. For the postal code attribute among the independent attributes, we use row 5 of the postal code data as an example. The corresponding sub-biometric value 10111 is 23.
[0077] The row number of each data and the value of the corresponding sub-biometric feature value are added together, and the resulting value is modulo the total number of rows to obtain the transformed position of each data.
[0078] Among them, the transformation position of the data with row number 3 in the associated attribute is mod(3+28,10), that is, 1; the transformation position of the address data with row number 4 in the independent attribute is mod(4+1,10), that is, 5; the transformation position of the postal code data with row number 5 in the independent attribute is mod(5+23,10), that is, 8. Accordingly, the data with row number 3 in the associated attribute is transformed to the first row, the address data with row number 4 in the independent attribute is transformed to the fifth row, and the postal code data with row number 5 in the independent attribute is transformed to the eighth row.
[0079] In particular, when the result of the modulo operation is 0, the row number of the transformed position of the data is equal to the total number of rows in the data table, that is, the data is transformed to the 10th row.
[0080] for Figure 2 The transformation method of other data contained in is the same as the above method and will not be repeated here.
[0081] Figure 3This is a structural diagram of an embodiment of the personal information de-identification device of the present application. The personal information de-identification device in this embodiment can be used as a personal information de-identification device to implement the personal information de-identification method provided in the embodiment of the present application.
[0082] like Figure 3 As shown, the above-mentioned personal information de-identification device may include: a generation module 21, a segmentation module 22, a determination module 23 and a position transformation module 24.
[0083] The generating module 21 is configured to obtain biometric information of a user accessing target data and generate a biometric value based on the biometric information.
[0084] In specific implementation, the generating module 21 is used to convert the biometric information of the user accessing the target data into binary data; perform hash calculation on the binary data, or the transformed data after the binary data has been transformed at least once, to obtain a biometric value of a preset length.
[0085] The segmentation module 22 is configured to segment the biometric feature value according to the association relationship between the data contained in the target data to obtain a plurality of sub-biometric feature values.
[0086] In specific implementation, the data attributes of each data element in the target data are first determined. Based on the information application scenario, the associated and independent attributes within the data attributes are determined. Then, based on the sum of the associated and independent attributes, the number of segments of the biometric value is determined, and the biometric value is segmented according to the number of segments.
[0087] The determination module 23 is configured to determine a correspondence between each data in the target data and a plurality of sub-biometric feature values, wherein the sub-biometric feature values are mapped to data positions.
[0088] The position transformation module 24 is used to determine the data transformation position according to the data position mapped to the sub-biometric feature value corresponding to each data in the target data, and transform the position of each data in the target data according to the data transformation position.
[0089] Specifically, it is used to weight the coordinate position of each data in the target data in the data table and the data position mapped to the sub-biometric feature value corresponding to each data to obtain the transformed position of each data in the target data.
[0090] In this embodiment, the generation module 21 generates a biometric value based on the biometric information of the user accessing the target data. The segmentation module 22 segments the biometric value based on the number of associated attributes and independent attributes, generating multiple sub-biometric values. The determination module 23 determines the correspondence between the multiple sub-biometric values and the associated attributes and independent attributes in the target data. The position transformation module 24 transforms the data contained in the associated attributes and the data contained in the independent attributes in the target data based on the data locations mapped to the sub-biometric values. This achieves personal information de-identification while protecting data availability, preventing the leakage of user personal information.
[0091] Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of the present application, such as Figure 4 As shown, the above-mentioned electronic device may include at least one processor; and at least one memory communicatively connected to the above-mentioned processor, wherein: the memory stores program instructions that can be executed by the processor, and the above-mentioned processor calls the above-mentioned program instructions to execute the personal information de-identification method provided in the embodiment of the present application.
[0092] Among them, the above-mentioned electronic device can be a personal information de-identification device, and this embodiment does not limit the specific form of the above-mentioned electronic device.
[0093] Figure 4 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present application is shown. Figure 4 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0094] like Figure 4 As shown, the electronic device is in the form of a general-purpose computing device. Components of the electronic device may include, but are not limited to: one or more processors 31, memory 33, and a communication bus 34 connecting different system components (including memory 33 and processing unit 31).
[0095] The communication bus 34 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of such architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.
[0096] Electronic devices typically include a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.
[0097] The memory 33 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a Compact Disc Read Only Memory (hereinafter referred to as: CD-ROM), a Digital Video Disc Read Only Memory (hereinafter referred to as: DVD-ROM), or other optical media) may be provided. In these cases, each drive can be connected to the communication bus 34 via one or more data medium interfaces. The memory 33 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application.
[0098] A program / utility having a set (at least one) of program modules may be stored in memory 33. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. The program modules generally perform the functions and / or methods of the embodiments described herein.
[0099] The electronic device may also communicate with one or more external devices (e.g., keyboard, pointing device, display, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through the communication interface 32. In addition, the electronic device may also communicate with the network adapter ( Figure 4 The network adapter can communicate with other modules of the electronic device through the communication bus 34. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, disk arrays (Redundant Arrays of Independent Drives; hereinafter referred to as: RAID) systems, tape drives, and data backup storage systems.
[0100] The processor 31 executes various functional applications and data processing by running the programs stored in the memory 33, such as implementing the personal information de-identification method provided in the embodiment of the present application.
[0101] An embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions. The computer instructions enable the computer to execute the personal information de-identification method provided in an embodiment of the present application.
[0102] The above-mentioned non-temporary computer-readable storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (Read Only Memory; hereinafter referred to as: ROM), an erasable programmable read-only memory (ErasableProgrammable Read Only Memory; hereinafter referred to as: EPROM) or flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0103] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0104] The computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect via the Internet).
[0105] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0106] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0107] It should be noted that the terminals involved in the embodiments of the present application may include but are not limited to personal computers (Personal Computer; hereinafter referred to as: PC), personal digital assistants (Personal Digital Assistant; hereinafter referred to as: PDA), wireless handheld devices, tablet computers (Tablet Computer), mobile phones, etc.
[0108] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interface, indirect coupling or communication connection of the device or unit, which may be electrical, mechanical or other forms.
[0109] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0110] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program code.
[0111] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for de-identifying personal information, characterized in that: include: Obtaining biometric information of a user accessing target data, and generating a biometric value based on the biometric information; Segmenting the biometric value according to the association relationship between the data included in the target data to obtain a plurality of sub-biometric values; respectively establishing corresponding relationships between each data in the target data and the plurality of sub-biometric feature values, wherein the sub-biometric feature values are mapped to data positions; Determining a data transformation position according to a data position mapped to a sub-biometric feature value corresponding to each data in the target data, and transforming a position of each data in the target data according to the data transformation position; Segmenting the biometric value according to the association relationship between the data included in the target data includes: Determining data attributes of each data included in the target data; Segmenting the biometric feature value according to the number of the data attributes and the association relationship between the data attributes; Segmenting the biometric value according to the number of the data attributes and the association between the data attributes includes: Determining associated attributes and independent attributes among the data attributes; wherein the associated attributes include at least two data attributes, and the at least two data attributes have an associated relationship; and the independent attribute has no associated relationship with other data attributes; Determining the number of segments of the biometric feature value according to the sum of the number of associated attributes and independent attributes, and segmenting the biometric feature value according to the number of segments; The establishing of corresponding relationships between each data in the target data and the plurality of sub-biometric feature values includes: The at least two data attributes included in each of the associated attributes correspond to the same sub-biometric feature value; Each of the independent attributes corresponds to a sub-biometric value; Each data in the target data corresponds to the same sub-biometric feature value as the data attribute to which it belongs.
2. The method according to claim 1, characterized in that Generating a biometric value according to the biometric information includes: converting the biometric information of the user accessing the target data into binary data; A hash calculation is performed on the binary data or the transformed data after the binary data has been transformed at least once to obtain a biometric feature value of a preset length.
3. The method according to claim 1, characterized in that Segmenting the biometric value according to the number of segments includes: According to the data length of the biometric value, the biometric value is divided equally or randomly according to the number of divisions; The number of divisions is greater than or equal to the sum of the number of associated attributes and independent attributes in the target data.
4. The method according to claim 1, wherein Determine the data transformation position according to the data position of the sub-biometric feature value mapping corresponding to each data in the target data, including: The coordinate position of each data in the target data in the data table and the data position mapped to the sub-biometric feature value corresponding to each data are weighted to obtain the transformed position of each data in the target data.
5. A personal information de-identification device, characterized in that: include: a generating module, configured to obtain biometric information of a user accessing target data and generate a biometric value based on the biometric information; a segmentation module, configured to segment the biometric value according to the association relationship between the data included in the target data to obtain a plurality of sub-biometric value; a determination module, configured to respectively establish a correspondence between each data in the target data and the plurality of sub-biometric feature values, wherein the sub-biometric feature values are mapped to data positions; a position transformation module, configured to determine a data transformation position according to a data position mapped to a sub-biometric feature value corresponding to each data in the target data, and to transform the position of each data in the target data according to the data transformation position; Segmenting the biometric value according to the association relationship between the data included in the target data includes: Determining data attributes of each data included in the target data; Segmenting the biometric feature value according to the number of the data attributes and the association relationship between the data attributes; Segmenting the biometric value according to the number of the data attributes and the association between the data attributes includes: Determining associated attributes and independent attributes among the data attributes; wherein the associated attributes include at least two data attributes, and the at least two data attributes have an associated relationship; and the independent attribute has no associated relationship with other data attributes; Determining the number of segments of the biometric feature value according to the sum of the number of associated attributes and independent attributes, and segmenting the biometric feature value according to the number of segments; The establishing of corresponding relationships between each data in the target data and the plurality of sub-biometric feature values includes: The at least two data attributes included in each of the associated attributes correspond to the same sub-biometric feature value; Each of the independent attributes corresponds to a sub-biometric value; Each data in the target data corresponds to the same sub-biometric feature value as the data attribute to which it belongs.
6. An electronic device, characterized in that: include: at least one processor; as well as at least one memory in communication with the processor, wherein: The memory stores program instructions that can be executed by the processor, and the processor can execute the method according to any one of claims 1 to 4 by calling the program instructions.
7. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Database sensitive association attribute desensitization method based on invariant random response technology
CN110990876A