Method, apparatus, device and medium for generating logically true data

By grouping hash storage, tuple class decomposition and non-random matching of the original data in data mining, logical real data is generated, which solves the problems of inefficiency of existing privacy protection methods and loss of data attributes, and achieves efficient data privacy protection.

CN119312379BActive Publication Date: 2025-06-27海南省大数据发展中心
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411028828.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-06-27
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

During the data mining process, existing privacy protection methods such as privacy computing and federated learning require complex encryption and multi-participant data communication, resulting in high consumption of computing resources and low efficiency, and the desensitized data loses data attributes and correlation, reducing the value of data mining.

Method used

By extracting sensitive data sets from the original data set, grouping hash storage, a hash bucket set is obtained, and a standard tuple class set is generated through tuple class decomposition and non-random matching methods, and finally sensitive information is disassembled on the standard tuple class set to obtain logical real data.

Benefits of technology

This method improves data storage and search efficiency, disrupts the association between sensitive data and class names, ensures data privacy security, reduces the probability of attackers obtaining sensitive information, and effectively solves the problem of low privacy protection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312379B_ABST
    Figure CN119312379B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data privacy protection, and discloses a method, device, equipment and medium for generating logically true data. The method includes: extracting a sensitive data set from an original data set, and performing grouped hashing storage on the original data set according to the sensitive data set to obtain a set of hashed storage buckets; performing tuple class decomposition on the tuples in the set of hashed storage buckets to obtain a set of decomposed tuple classes and a set of remaining tuples; using a non-random matching method to allocate each tuple in the set of remaining tuples to each decomposed tuple class in the set of decomposed tuple classes to obtain a set of standard tuple classes; performing sensitive information disassembling on the set of standard tuple classes to obtain logically true data. Through the iterative multi-tuple decomposition, non-random allocation and sensitive data collection implemented by the present invention, the privacy protection of the original data set can be effectively realized, the data logic of the original data can be retained, and the efficiency of mining data privacy protection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data privacy protection, and in particular, to a method, device, equipment and medium for generating logically true data. Background Art

[0002] Data mining is a process of extracting valuable information and patterns from a large amount of data. Data mining can help all walks of life obtain valuable information and insights from data, thereby promoting the intelligent and refined management in various fields, improving efficiency and decision-making quality. However, in the process of data mining, in order to prevent the leakage of personal identity information or sensitive information and protect personal privacy, it is necessary to perform privacy protection on the original data.

[0003] Traditional privacy protection methods for data mining include data desensitization, privacy computing, data sandboxes, and federated learning. In actual use, technical solutions such as privacy computing and federated learning need to be carried out by multiple devices or computing nodes without sharing the original data. To ensure data security, complex encryption, decryption operations and data communication among multiple participants are required, which increases additional computing and communication overheads, consumes a high amount of computing resources, resulting in a relatively high overall computing cost and low computing efficiency. Moreover, in order to pursue the security of desensitized data, the desensitized data completely loses its data attributes and the relevance between data, thus reducing the mining value of the data and possibly leading to a low efficiency problem when performing privacy protection on the mined data. Summary of the Invention

[0004] The present invention provides a method, device, equipment and medium for generating logically true data, and its main purpose is to solve the problem of relatively low efficiency in privacy protection of mined data in related technologies.

[0005] To achieve the above object, a method for generating logically true data provided by the present invention includes: extracting a sensitive data set from an original data set, performing grouped hash storage on the original data set according to the sensitive data set to obtain a hash storage bucket set; performing tuple class decomposition on the tuples in the hash storage bucket set to obtain a decomposed tuple class set and a remaining tuple set; using a non-uniform random matching method to allocate each tuple in the remaining tuple set to each decomposed tuple class in the decomposed tuple class set to obtain a standard tuple class set; performing sensitive information disassembling on the standard tuple class set to obtain logically true data.

[0006] To solve the above problems, the present invention further provides a device for generating logically true data, the device comprising: a hash grouping module, configured to extract a sensitive data set from an original data set, and perform grouped hashing storage on the original data set according to the sensitive data set to obtain a set of hash storage buckets; an iterative grouping module, configured to perform tuple class decomposition on the tuples in the set of hash storage buckets to obtain a set of decomposed tuple classes and a remaining tuple set; a remaining grouping module, configured to use a non-random matching method to allocate each tuple in the remaining tuple set to each decomposed tuple class in the set of decomposed tuple classes to obtain a set of standard tuple classes; and an information disassembling module, configured to perform sensitive information disassembly on the set of standard tuple classes to obtain logically true data.

[0007] To solve the above problems, the present invention further provides an electronic device, the electronic device comprising:

[0008] at least one processor; and,

[0009] a memory communicatively connected to the at least one processor; wherein,

[0010] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned method for generating logically true data.

[0011] To solve the above problems, the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned method for generating logically true data is implemented.

[0012] In the embodiment of the present invention, by performing grouped hashing storage on the original data set according to the sensitive data set to obtain a set of hash storage buckets, grouped bucket storage of the sensitive data set can be realized according to the data values of the sensitive data, ensuring efficient data storage and search, and realizing primary tuple classification according to the sensitive data. By performing iterative multi-tuple decomposition and using a non-random matching method to generate a set of standard tuple classes, it can be ensured that each tuple in the remaining tuple set can also be allocated to a tuple class with different sensitive data, thereby disrupting the association between the sensitive data and the class name, ensuring data privacy and security. By performing sensitive information disassembly on the set of standard tuple classes to obtain logically true data, the original one-to-one correspondence between tuples and sensitive data is broken, so that each equivalence class corresponds to multiple sensitive data values. An attacker can only reconstruct the data by naturally joining the quasi-identifier data set and the sensitive information set, but according to the nature of natural join, the reconstruction will generate many records that do not exist in the original data set, resulting in a greatly reduced probability of obtaining sensitive information, thereby effectively ensuring user privacy. Therefore, the method, device, equipment and medium for generating logically true data proposed by the present invention can solve the problem of low efficiency in privacy protection of mined data. Brief Description of the Drawings

[0013] Figure 1 Schematic flowchart of the method for generating logically true data provided by an embodiment of the present invention;

[0014] Figure 2 Schematic diagram of an example of the original data set provided by an embodiment of the present invention;

[0015] Figure 3 Schematic diagram of an example of the bucket tuple number set provided by an embodiment of the present invention;

[0016] Figure 4 Schematic diagram of an example of the decomposed tuple class set provided by an embodiment of the present invention;

[0017] Figure 5 Schematic diagram of an example of the standard tuple class set provided by an embodiment of the present invention;

[0018] Figure 6 Schematic diagram of an example of the quasi-identifier data set provided by an embodiment of the present invention;

[0019] Figure 7 Schematic diagram of an example of the sensitive information set provided by an embodiment of the present invention;

[0020] Figure 8 Functional module diagram of the device for generating logically true data provided by an embodiment of the present invention;

[0021] Figure 9 Schematic structural diagram of an electronic device for implementing the method for generating logically true data provided by an embodiment of the present invention.

[0022] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0023] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] An embodiment of the present application provides a method for generating logically true data. The execution subject of the method for generating logically true data includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for generating logically true data can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0025] Referring to Figure 1 As shown, it is a schematic flowchart of the method for generating logically true data provided by an embodiment of the present invention. In this embodiment, the method for generating logically true data includes:

[0026] S1. Extract a sensitive data set from the original data set, and perform grouped hashing storage on the original data set according to the sensitive data set to obtain a set of hashed storage buckets.

[0027] Specifically, the original data set refers to the data set that needs to be mined for data and protected for privacy. The original data set can be a patient data form in a hospital, a student data form in a school, and a personnel data form in a company.

[0028] Specifically, the original data set can be a patient data form as Figure 2 shown, which contains headers such as serial number (ID), patient name (Name), patient age (Age), patient gender (Sex), zip code (Zipcode), and patient disease (Disease). The sensitive data set refers to the set of partial data in the original data set that needs to be protected for privacy. In the patient data form, the sensitive data set refers to the data set composed of all elements in the patient disease (Disease) column.

[0029] Specifically, extracting the sensitive data set from the original data set includes extracting the original header set from the original data set, performing sensitive keyword matching on the original header set to obtain sensitive header elements, generating a sensitive element query statement according to the sensitive header elements, and traversing and extracting the original data set according to the sensitive element query statement to obtain the sensitive data set.

[0030] Specifically, the SHOW COLUMNS query statement of MySQL or the INFORMATION_SCHEMA item of PostgreSQL can be used to extract the table headers. The original table header set is a data set composed of all table header elements in the original data set. For example, Figure 2 the corresponding original table header set in

[0031] is [ID, Name, Age, Sex, Zipcode, Disease]. Specifically, sensitive keyword matching refers to using the keyword matching method to extract the table headers in the original table header set that are related to sensitive or private elements as sensitive table header elements. The sensitive element query statement is a query statement used to traverse and extract all elements of the item where the sensitive table header element is located. The executeQuery method of the Statement object can be used to generate the sensitive element query statement, and the getString, getInt, etc. methods of the ResultSet object can be called using the sensitive element query statement for traversal and extraction.

[0032] Specifically, the original data set is grouped and hashed and stored according to the sensitive data set to obtain a hashed storage bucket set, including: selecting the sensitive data in the sensitive data set one by one as the target sensitive data, and taking the tuple where the target sensitive data is located in the original data set as the target data tuple; performing a hashing operation on the target sensitive data to obtain a target hash index; mapping a mapped storage bucket from a preset storage bucket list according to the target hash index; storing the target data tuple into the mapped storage bucket to obtain a hashed storage bucket, and aggregating the hashed storage buckets corresponding to all the target sensitive data in the sensitive data set into a hashed storage bucket set.

[0033] Specifically, a tuple refers to an element row in the original data set, and a target data tuple refers to the element row where the target sensitive data is located in the original data set. For example, referring to Figure 2 as shown, when the target sensitive data is gastritis, the corresponding target data tuple is [6, Jim, 65, F, 25120, gastritis].

[0034] Specifically, hashing operations can be performed using hashing algorithms such as MD5, SHA-256, CRC32, or FNV-1a. The target hash index is the hash value obtained after the target sensitive data undergoes a hashing operation.

[0035] Specifically, mapping a mapped bucket from a preset list of buckets according to a target hash index includes: performing a modulo operation on the target hash index to obtain a bucket index; matching the mapped bucket from the preset list of buckets according to the bucket index. Here, a bucket is a storage space for storing some specific types of data, and the list of buckets is part of a hash table for storing data items with similar hash values. When a new data item needs to be added to the hash table, a hash function is used to calculate the key of the data item, and then the data item is stored in the corresponding bucket.

[0036] In the embodiment of the present invention, by performing grouped hash storage on the original data set according to the sensitive data set, a set of hash storage buckets can be obtained, which can realize grouped bucket storage of the sensitive data set according to the data values of the sensitive data, ensuring efficient data storage and retrieval, and realizing primary tuple classification according to the sensitive data.

[0037] S2. Perform tuple class decomposition on the tuples in the set of hash storage buckets to obtain a set of decomposed tuple classes and a set of remaining tuples.

[0038] Specifically, each decomposed tuple class in the set of decomposed tuple classes is a combined class formed by aggregating multiple tuples. The tuples in the decomposed tuple class satisfy diverse classification characteristics, that is, the decomposed tuple class is a class composed of tuples corresponding to multiple different types of sensitive data. The set of remaining tuples is a set composed of the remaining tuples in the set of hash storage buckets except for the tuples corresponding to the set of decomposed tuple classes.

[0039] In the embodiment of the present invention, performing tuple class decomposition on the tuples in the set of hash storage buckets to obtain a set of decomposed tuple classes and a set of remaining tuples includes: counting the number of tuples in each hash storage bucket in the set of hash storage buckets to obtain a set of bucket tuple numbers; extracting a diversification coefficient from the set of bucket tuple numbers by using the method of maximum value matching; classifying the tuples in the set of hash storage buckets according to the diversification coefficient to obtain a set of decomposed tuple classes; and aggregating the remaining tuples in the set of hash storage buckets except for the set of decomposed tuple classes into a set of remaining tuples.

[0040] Specifically, the bucket tuple number statistics refers to counting the number of tuples in each hash storage bucket in the set of hash storage buckets and aggregating the numbers as the bucket tuple numbers into a set of bucket tuple numbers. Referring to Figure 3 as shown, it is the set of bucket tuple numbers corresponding to the patient data form when the original data set is a patient data form. Here, bucket-ID refers to the serial number ID of each hash storage bucket in the set of hash storage buckets, Disease is the sensitive table header element, and Count is the header of the set of same tuple numbers. For example, the hash storage bucket with serial number 1 is used to store tuples with the sensitive table header element pneumonia, and the bucket tuple number of the tuples with the sensitive table header element pneumonia in the hash storage bucket with serial number 1 is 2.

[0041] Specifically, the maximum value matching refers to matching the bucket tuple set with the largest number of bucket tuples as the diversification coefficient, and the diversification coefficient is used to infer the maximum probability of sensitive values. Refer to Figure 3 It can be seen that when the original data set is a patient data form, the diversification coefficient is 3.

[0042] Specifically, the tuples in the hash storage bucket set are classified according to the diversification coefficient to obtain a decomposed tuple class set, including: determining whether the number of non-empty storage buckets in the hash storage bucket set is less than the diversification coefficient; if not, screening out multi-tuple storage bucket groups from the hash storage bucket set according to the diversification coefficient; randomly screening out one tuple from each storage bucket in the multi-tuple storage bucket group, and generating a decomposed tuple class according to all the screened-out tuples; updating the hash storage bucket set with the multi-tuple storage bucket group from which the decomposed tuple class has been screened out, and returning to the step of determining whether the number of non-empty storage buckets in the hash storage bucket set is less than the diversification coefficient; if so, generating a decomposed tuple class set according to all the decomposed tuple classes.

[0043] Specifically, a non-empty storage bucket refers to a non-empty storage bucket in the hash storage bucket set, that is, a storage bucket storing tuples.

[0044] Specifically, screening out multi-tuple storage bucket groups from the hash storage bucket set according to the diversification coefficient includes: counting the number of bucket tuples in the hash storage bucket set to obtain a bucket tuple set; sorting the bucket tuple set in descending order to obtain a bucket tuple sequence; aggregating the first diversification coefficient number of bucket tuples in the bucket tuple sequence into a multi-bucket tuple array; aggregating the hash storage buckets corresponding to the multi-bucket tuple array in the hash storage bucket set into a multi-tuple storage bucket group.

[0045] Specifically, refer to Figure 3 As shown, when the original data set is a patient data form, the same tuple sequence is [2, 2, 2, 1, 1], the multi-bucket tuple array is [2, 2, 2], and the multi-tuple storage bucket group is the storage buckets corresponding to bucket-ID 1, 2, and 3.

[0046] Specifically, refer to Figure 4 As shown, for the case where the original data set is a patient data form, it is the tuple table corresponding to the decomposed tuple class set, where QI gcnt is the serial number of each decomposed tuple class in the decomposed tuple class set. The number of decomposed tuple classes is 2, and each decomposed tuple class contains 3 tuples.

[0047] Specifically, the remaining tuple set is a set composed of the remaining tuples in the hash storage bucket set except the decomposed tuple class set. Refer to Figure 3 and Figure 4 As shown, the remaining tuple set is the tuples corresponding to bucket-ID 4 and 5.

[0048] In the embodiments of the present invention, by performing iterative multi - tuple decomposition to obtain a set of decomposed tuple classes and a set of remaining tuples, each tuple in the hash storage bucket with the same sensitive data can be split into different tuple classes, thereby disrupting the association between the sensitive data and the class names and ensuring the privacy and security of the data.

[0049] S3. Use the method of non - random matching to allocate each tuple in the set of remaining tuples to each decomposed tuple class in the set of decomposed tuple classes to obtain a set of standard tuple classes.

[0050] In the embodiments of the present invention, using the method of non - random matching to allocate each tuple in the set of remaining tuples to each decomposed tuple class in the set of decomposed tuple classes to obtain a set of standard tuple classes includes: successively selecting a tuple in the set of remaining tuples as the target remaining tuple, and using the sensitive data in the target remaining tuple as the target remaining sensitive data; randomly selecting a decomposed tuple class in the set of decomposed tuple classes as the target decomposed tuple class, and using the corresponding multi - tuple storage bucket group of the target decomposed tuple class as the target multi - tuple storage bucket group; determining whether the target remaining sensitive data exists in the target multi - tuple storage bucket group; if so, returning to the step of randomly selecting a decomposed tuple class in the set of decomposed tuple classes as the target decomposed tuple class; if not, allocating the target remaining tuple to the target decomposed tuple class to obtain a secondary tuple class, and using the secondary tuple class to update the set of decomposed tuple classes into a set of secondary tuple classes until the target remaining tuple is the last tuple in the set of remaining tuples, and using the set of secondary tuple classes as the set of standard tuple classes.

[0051] Specifically, referring to Figure 5 as shown, it is the tuple table corresponding to the set of standard tuple classes when the original data set is a patient data form. Compared with Figure 4 , there are two more tuples corresponding to the sensitive data of gastritis and the sensitive data of bronchitis in the set of standard tuple classes, that is, the two tuples in the set of remaining tuples. Since the target multi - tuple storage bucket group corresponding to the decomposed tuple class serial number QI gcnt 1 does not contain the sensitive data gastritis, the tuple with the sensitive data gastritis can be added to the decomposed tuple class with the decomposed tuple class serial number QI gcnt 1.

[0052] In the embodiments of the present invention, by using the method of non - random matching to generate a set of standard tuple classes, it can be ensured that each tuple in the set of remaining tuples can also be allocated to tuple classes with different sensitive data, thereby disrupting the association between the sensitive data and the class names and ensuring the privacy and security of the data.

[0053] S4. Perform sensitive information disassembling on the set of standard tuple classes to obtain logically true data.

[0054] Specifically, logically true data refers to data with the true data logic of the original data to meet the data mining requirements. Logically true data includes identification information and sensitive information. Among them, the identification information refers to the quasi-identifiers (Quasi-Identifiers, abbreviated as QI) that can uniquely identify an individual, and the sensitive information (Sensitive Information, abbreviated as SI).

[0055] In the embodiment of the present invention, the sensitive information in the standard tuple class set is disassembled to obtain logically true data, including: decomposing the identification information in the standard tuple class set to obtain a quasi-identifier data set; decomposing the sensitive information in the standard tuple class set to obtain a sensitive information set; generating logically true data according to the quasi-identifier data set and the sensitive information set.

[0056] Specifically, decomposing the identification information in the standard tuple class set means using the class serial numbers of each tuple in the standard tuple class set to replace the sensitive data items in the tuple. The identification information in the standard tuple class set can be decomposed by using the keyword matching method and the character replacement method.

[0057] Specifically, referring to Figure 6 shown, when the original data set is a patient data form, it is the table corresponding to the quasi-identifier data set. The quasi-identifier data in the quasi-identifier data set is data composed of quasi-identifiers (Quasi-Identifiers, abbreviated as QI) that can uniquely identify an individual, and compared with Figure 5 the standard tuple class set of, the quasi-identifier data set replaces the sensitive data Disease item with the corresponding decomposed tuple class serial number QI gcnt .

[0058] Specifically, decomposing the sensitive information in the standard tuple class set to obtain a sensitive information set includes: extracting a sensitive data set from the standard tuple class set; selecting each sensitive data in the sensitive data set as the target sensitive data one by one, and using the standard tuple class corresponding to the target sensitive data in the standard tuple class set as the target standard tuple class; counting the number of sensitive data of the target sensitive data in the target standard tuple class; using the serial number of the target standard tuple class as the target class serial number; generating sensitive information according to the target class serial number, the target sensitive data and the number of sensitive data, and aggregating the sensitive information of all target sensitive data in the sensitive data set into a sensitive information set.

[0059] Specifically, referring to Figure 7 shown, when the original data set is a patient data form, it is the table corresponding to the sensitive information set. The sensitive information (Sensitive Information, abbreviated as SI) in the sensitive information set is sensitive data stored separately, where QI gcntThe item corresponds to the target class number, and the Count item corresponds to the number of sensitive data.

[0060] In the embodiment of the present invention, by disassembling sensitive information from the standard tuple class set, logical true data is obtained, breaking the one-to-one correspondence between the original tuples and sensitive data, so that each equivalence class corresponds to multiple values of sensitive data. An attacker can only reconstruct the data by natural-joining the quasi-identifier data set and the sensitive information set. However, according to the nature of natural join, many records that do not exist in the original data set will be generated during reconstruction, resulting in a significant reduction in the probability of obtaining sensitive information, thus effectively ensuring user privacy.

[0061] In the embodiment of the present invention, by grouping and hashing the original data set according to the sensitive data set, a set of hash storage buckets is obtained. Group bucket storage of the sensitive data set can be achieved according to the data values of the sensitive data, ensuring efficient data storage and retrieval. And primary tuple classification is achieved according to the sensitive data. By performing iterative multi-tuple decomposition and using the method of non-uniform random matching to generate a standard tuple class set, it can be ensured that each tuple in the remaining tuple set can also be assigned to a tuple class with different sensitive data, thus disrupting the association between the sensitive data and the class name, ensuring data privacy and security. By disassembling sensitive information from the standard tuple class set, logical true data is obtained, breaking the one-to-one correspondence between the original tuples and sensitive data, so that each equivalence class corresponds to multiple values of sensitive data. An attacker can only reconstruct the data by natural-joining the quasi-identifier data set and the sensitive information set. However, according to the nature of natural join, many records that do not exist in the original data set will be generated during reconstruction, resulting in a significant reduction in the probability of obtaining sensitive information, thus effectively ensuring user privacy. Therefore, the method for generating logical true data proposed by the present invention can solve the problem of low efficiency in privacy protection of mined data.

[0062] As Figure 8 shown, it is a functional module diagram of a device for generating logical true data provided by an embodiment of the present invention.

[0063] The device 800 for generating logical true data of the present invention can be installed in an electronic device. According to the functions implemented, the device 800 for generating logical true data can include a hash grouping module 801, an iterative grouping module 802, a remaining grouping module 803, and an information disassembling module 804. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0064] In this embodiment, the functions of each module / unit are as follows:

[0065] The hash grouping module 801 is used to extract a sensitive data set from the original data set, and group and hash store the original data set according to the sensitive data set to obtain a set of hash storage buckets;

[0066] The iterative grouping module 802 is used to perform tuple class decomposition on the tuples in the set of hash storage buckets to obtain a set of decomposed tuple classes and a set of remaining tuples;

[0067] The remaining grouping module 803 is used to allocate each tuple in the set of remaining tuples to each decomposed tuple class in the set of decomposed tuple classes by using a non-random matching method to obtain a set of standard tuple classes;

[0068] The information disassembling module 804 is used to disassemble sensitive information from the set of standard tuple classes to obtain logically true data.

[0069] Specifically, each module in the logically true data generation device 800 in the embodiments of the present invention adopts the same technical means as the above Figure 1 in the method for generating logically true data, and can produce the same technical effects, which will not be elaborated here.

[0070] As Figure 9 shown, it is a schematic structural diagram of an electronic device for implementing the method for generating logically true data provided by an embodiment of the present invention.

[0071] The electronic device 901 may include a processor 910, a memory 911, a communication bus 912, and a communication interface 913, and may further include a computer program stored in the memory 911 and executable on the processor 910, such as a logically true data generation program.

[0072] Among them, the processor 910 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 910 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 911 (such as executing the logically true data generation program, etc.), and calling data stored in the memory 911, to execute various functions of the electronic device and process data.

[0073] The memory 911 includes at least one type of readable storage medium, which includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 911 can be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 911 can also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 911 can also include both the internal storage unit and the external storage device of the electronic device. The memory 911 can be used not only to store application software installed in the electronic device and various types of data, such as the code of the program for generating logical real data, etc., but also to temporarily store the data that has been output or will be output.

[0074] The communication bus 912 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to enable the connection and communication between the memory 911 and at least one processor 910, etc.

[0075] The communication interface 913 is used for the communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is usually used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0076] Only an electronic device with components is shown in the figure. Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown in the figure, or combine some components, or have different component arrangements.

[0077] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to at least one processor 910 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0078] It should be understood that the embodiments are for illustrative purposes only and are not limited by this structure in the scope of the patent application.

[0079] Specifically, for the specific implementation method of the above instructions by the processor 910, reference may be made to the description of the relevant steps in the corresponding embodiments of the accompanying drawings, which will not be elaborated here.

[0080] Furthermore, if the modules / units integrated in the electronic device 901 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0081] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0082] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing reference signs in the claims should not be regarded as limiting the claimed rights.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for generating logical real data, characterized in that: The method comprises: Extracting a sensitive data set from the original data set, grouping and hashing the original data set according to the sensitive data set to obtain a hash storage bucket set; Performing tuple class decomposition on the tuples in the hash bucket set to obtain a decomposed tuple class set and a remaining tuple set; Selecting tuples in the remaining tuple set one by one as target remaining tuples, and using sensitive data in the target remaining tuples as target remaining sensitive data; Randomly select a decomposition tuple class in the decomposition tuple class set as a target decomposition tuple class, and use a tuple storage bucket group corresponding to the target decomposition tuple class as a target tuple storage bucket group; Determine whether there is target remaining sensitive data in the target multi-tuple storage bucket group; If yes, returning to the step of randomly selecting a decomposition tuple class from the decomposition tuple class set as a target decomposition tuple class; If not, the target remaining tuple is assigned to the target decomposition tuple class to obtain a secondary tuple class, and the decomposition tuple class set is updated into a secondary tuple class set by using the secondary tuple class, until the target remaining tuple is the last tuple in the remaining tuple set, and the secondary tuple class set is used as the standard tuple class set; Decomposing the identification information in the standard tuple class set to obtain a quasi-identifier data set; Decomposing the sensitive information in the standard tuple class set to obtain a sensitive information set; Logical real data is generated according to the quasi-identifier data set and the sensitive information set.

2. The method for generating logical real data according to claim 1, characterized in that: The step of extracting a sensitive data set from an original data set includes: Extracting headers from the original data set to obtain an original header set; Perform sensitive keyword matching on the original header set to obtain sensitive header elements; Generate a sensitive element query statement according to the sensitive header element; The original data set is traversed and extracted according to the sensitive element query statement to obtain a sensitive data set.

3. The method for generating logical real data according to claim 1, characterized in that: The grouping and hashing the original data set according to the sensitive data set to obtain a hash storage bucket set includes: Selecting sensitive data in the sensitive data set one by one as target sensitive data, and taking the tuple containing the target sensitive data in the original data set as the target data tuple; Performing a hash operation on the target sensitive data to obtain a target hash index; Mapping a mapping bucket from a preset bucket list according to the target hash index; The target data tuple is stored in the mapping bucket to obtain a hash bucket, and the hash buckets corresponding to all target sensitive data in the sensitive data set are aggregated into a hash bucket set.

4. The method for generating logical real data according to claim 1, characterized in that: The step of performing tuple class decomposition on the tuples in the hash bucket set to obtain a decomposed tuple class set and a remaining tuple set includes: Counting the number of bucket tuples on the hash storage bucket set to obtain a bucket tuple number set; Extracting a diversity coefficient from the bucket tuple set using a maximum value matching method; Classifying the tuples in the hash bucket set according to the diversity coefficient to obtain a decomposed tuple class set; The remaining tuples in the hash bucket set except the decomposed tuple set are assembled into a remaining tuple set.

5. The method for generating logical real data according to claim 4, characterized in that: The step of classifying the tuples in the hash bucket set according to the diversity coefficient to obtain a decomposed tuple class set includes: Determining whether the number of non-empty buckets in the hash bucket set is less than the diversity coefficient; If not, filtering out a multi-tuple bucket group from the hash bucket set according to the diversity coefficient; Randomly select a tuple from each storage bucket in the tuple storage bucket group, and generate a decomposed tuple class according to all the selected tuples; The hash bucket set is updated by using the tuple bucket group that has screened out the decomposed tuple class, and returning to the step of determining whether the number of non-empty buckets in the hash bucket set is less than the diversity coefficient; If so, generate a decomposition tuple class set based on all decomposition tuple classes.

6. The method for generating logical real data according to claim 1, characterized in that: The step of decomposing the sensitive information in the standard tuple class set to obtain a sensitive information set includes: Extracting a sensitive data set from the standard tuple set; Selecting sensitive data in the sensitive data set one by one as target sensitive data, and taking the standard tuple class corresponding to the target sensitive data in the standard tuple class set as the target standard tuple class; Counting the number of sensitive data of the target sensitive data in the target standard tuple class; Using the serial number of the target standard tuple class as the target class serial number; Sensitive information is generated according to the target class serial number, the target sensitive data and the amount of sensitive data, and the sensitive information of all target sensitive data in the sensitive data set is aggregated into a sensitive information set.

7. A device for generating logical real data, characterized in that: The device comprises: A hash grouping module is used to extract a sensitive data set from an original data set, and group and hash the original data set according to the sensitive data set to obtain a hash storage bucket set; An iterative grouping module, used for performing tuple class decomposition on the tuples in the hash bucket set to obtain a decomposed tuple class set and a remaining tuple set; The remaining grouping module is used to select tuples in the remaining tuple set one by one as target remaining tuples, and use the sensitive data in the target remaining tuple as target remaining sensitive data; randomly select a decomposition tuple class in the decomposition tuple class set as a target decomposition tuple class, and use the tuple storage bucket group corresponding to the target decomposition tuple class as a target tuple storage bucket group; determine whether there is target remaining sensitive data in the target tuple storage bucket group; if so, return to the step of randomly selecting a decomposition tuple class in the decomposition tuple class set as the target decomposition tuple class; if not, assign the target remaining tuple to the target decomposition tuple class to obtain a secondary tuple class, and use the secondary tuple class to update the decomposition tuple class set into a secondary tuple class set, until the target remaining tuple is the last tuple in the remaining tuple set, and use the secondary tuple class set as the standard tuple class set; The information disassembly module is used to decompose the identification information in the standard tuple class set to obtain a quasi-identifier data set; decompose the sensitive information in the standard tuple class set to obtain a sensitive information set; and generate logical real data based on the quasi-identifier data set and the sensitive information set.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for generating logical truth data according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating logical real data according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Logic address allocation method and device, electronic equipment and storage medium

    CN116633900A