A data privacy protection method, device and medium based on random noise

Through the data privacy protection method based on random noise, the limitations of the data privacy protection method in the prior art in terms of availability and efficiency are solved, and efficient data management and security control are realized to ensure that data meets business needs while protecting privacy.

CN118350042BActive Publication Date: 2025-09-02INSPUR GENERSOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410513925.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-26
Publication Date
2025-09-02
Estimated Expiration
2044-04-26

AI Technical Summary

Technical Problem

The existing data privacy protection methods have limitations in data availability, processing efficiency, etc., and cannot meet the needs of many aspects of data use.

Method used

The data privacy protection method based on random noise is adopted, and data classification and random matrix generation are obtained by obtaining the description information of the original collected data, adding random noise, generating noise injection obfuscated data packets, and data access control is carried out in combination with user permissions and query request information.

Benefits of technology

Significantly enhance data privacy protection capabilities, improve data management efficiency and security, ensure that data protects privacy while meeting business analysis and processing needs, achieve refined access control, and respond to complex privacy threats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118350042B_ABST
    Figure CN118350042B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose a data privacy protection method, device and medium based on random noise, which relate to the field of privacy protection technology. The method includes: obtaining multiple original collected data to determine data description information of each original collected data, where the data description information includes data subject, data structure, data noise probability and data privacy authority; classifying each original collected data according to the data subject and data structure to determine at least one data block; generating random matrix information corresponding to each data block, writing the data description information and random matrix information into corresponding metadata to generate at least one data packet; adding random noise to each data block based on the data noise probability and random matrix information corresponding to each data block to generate a noise injection obfuscated data packet for data storage; receiving query request information of a data query user to determine the query data through the query request information and the random matrix information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of privacy protection technology, and in particular to a data privacy protection method, device, and medium based on random noise. Background Art

[0002] With the development of information technology, data collection has become a critical application across various industries. As a crucial component of data collection systems, data collection interfaces help protect user privacy and assist data collection systems in complying with user information protection regulations, thereby preventing the inclusion of user-specific data within the data provided by the platform. Currently, the industry has made some progress in data privacy protection. Traditional methods of data privacy protection primarily include data encryption, anonymization, and access control. Data encryption technology encrypts raw data to ensure data security during transmission and storage. Anonymization technology desensitizes data to hide sensitive personal information and prevent the leakage of personal privacy. Access control technology restricts data access by setting appropriate access rights.

[0003] First, while traditional encryption technologies can protect data security, they can, in certain circumstances, reduce data availability and processing efficiency. Second, while anonymization technologies can conceal sensitive personal information, they can also compromise data accuracy and integrity. Furthermore, as attack vectors continue to evolve, a single privacy protection approach may not be effective against complex privacy threats. Therefore, existing data privacy protection methods still have limitations and cannot meet diverse data usage requirements, such as data availability and processing efficiency. Summary of the Invention

[0004] One or more embodiments of this specification provide a data privacy protection method, device, and medium based on random noise, which are used to solve the following technical problems: existing data privacy protection methods still have some limitations and cannot meet data usage requirements in many aspects such as data availability and processing efficiency.

[0005] One or more embodiments of this specification adopt the following technical solutions:

[0006] One or more embodiments of the present specification provide a data privacy protection method based on random noise, the method comprising: obtaining a plurality of original collected data to determine data description information of each of the original collected data, wherein the data description information includes a data subject, a data structure, a data noise probability, and a data privacy right; classifying the plurality of original collected data according to the data subject and the data structure of each of the original collected data to determine at least one data block, wherein each of the data blocks includes a plurality of specified original collected data; generating random matrix information corresponding to each of the data blocks, and writing the data description information and the random matrix information into corresponding metadata to generate at least one data packet, wherein the random matrix information includes an evenly distributed random number matrix, and the data packet includes metadata and a data block; adding random noise to each of the data blocks based on the data noise probability corresponding to each of the data blocks and the random matrix information to generate a noise-injected obfuscated data packet for data storage; receiving query request information from a data query user to determine query data through the query request information and the random matrix information, wherein the query request information includes user authority information of the data query user and query data attribute information.

[0007] Furthermore, generating random matrix information corresponding to each of the data blocks specifically includes: performing random number calculation on the specified original collected data in each of the data blocks to determine the specified random number corresponding to each of the specified original collected data; and constructing random matrix information according to the specified random number corresponding to each of the specified original collected data.

[0008] Furthermore, random matrix information is constructed based on the specified random number corresponding to each of the specified original collected data, specifically including: obtaining the data quantity of the specified original collected data in each of the data blocks, and dividing the multiple specified original collected data to determine the number of groups corresponding to each of the data blocks; based on the data quantity and the number of groups, generating an average distribution random number matrix corresponding to the specified random number.

[0009] Furthermore, based on the data noise probability corresponding to each data block and the random matrix information, random noise is added to each data block to generate a noise injection obfuscation data packet, specifically including: obtaining the average distribution random number matrix in the random matrix information; locating data in the average distribution random number matrix according to the data noise probability to determine at least one noise point; adding random noise to each noise point according to the average distribution random number matrix to generate noise injection obfuscation data corresponding to each noise point to determine the noise injection obfuscation data packet corresponding to each data block.

[0010] Furthermore, according to the data noise probability, data positioning is performed in the evenly distributed random number matrix to determine at least one noise point, specifically including: obtaining multiple specified random numbers in the evenly distributed random number matrix; and according to each of the specified random numbers and the data noise probability, determining the first noise point by using the first specified random number that is smaller than the data noise probability.

[0011] Furthermore, according to the evenly distributed random number matrix, random noise is added to each of the noise points to generate noise injection obfuscation data corresponding to each of the noise points, specifically including: determining the specified random number and original data corresponding to each of the noise points in the evenly distributed random number matrix; determining the noise data according to the specified random number corresponding to each of the noise points, so as to add random noise to the original data corresponding to each of the noise points through the noise data to generate noise injection obfuscation data corresponding to each of the noise points.

[0012] Furthermore, the query data is determined through the query request information and the random matrix information, specifically including: obtaining the user authority information of the data query user in the query request information to determine the query authority based on the data privacy authority and the user authority information, wherein the data privacy authority includes the data owner, and the query authority includes data owner query and non-data owner query; when the data query user is the data owner query, obtaining the query data packet corresponding to the query request information, wherein the query data packet includes metadata and a noise injection data block; performing an inverse operation on the noise injection data block through the random matrix information in the metadata to restore the original collected data and determine the query data; when the data query user is the non-data owner query, determining the query data according to the query data attribute information.

[0013] Furthermore, the query data is determined based on the query data attribute information, specifically including: obtaining the query data attribute information, wherein the query data attribute information includes summary data attributes and calculation data attributes; when the query data attribute information is the calculation data attribute, performing an inverse operation on the noise-injected data block through the random matrix information in the metadata to restore the original collected data and determine the query data; when the query data attribute information is the summary data attribute, obtaining the real summary result data corresponding to the query request information; obtaining the disturbance factor in the data description information to add noise to the real summary result data based on the disturbance factor to determine the query data.

[0014] One or more embodiments of this specification provide a data privacy protection device based on random noise, including:

[0015] at least one processor; and,

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0018] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to:

[0019] Acquire multiple original collected data to determine data description information of each original collected data, wherein the data description information includes data subject, data structure, data noise probability and data privacy authority; classify the multiple original collected data according to the data subject and data structure of each original collected data to determine at least one data block, wherein each data block includes multiple specified original collected data; generate random matrix information corresponding to each data block, write the data description information and the random matrix information into corresponding metadata to generate at least one data packet, wherein the random matrix information includes an evenly distributed random number matrix, and the data packet includes metadata and data blocks; based on the data noise probability corresponding to each data block and the random matrix information, add random noise to each data block to generate a noise injection obfuscation data packet for data storage; receive query request information of a data query user to determine query data through the query request information and the random matrix information, wherein the query request information includes user authority information of the data query user and query data attribute information.

[0020] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the above technical solution, by adding random noise to each original collected data, the privacy protection capability of the data can be significantly enhanced, and the injection of noise makes it difficult to directly identify or restore the original data, thereby effectively preventing data leakage and abuse; by classifying data and generating data packets, data with similar data themes and structures can be centrally managed, which not only improves the efficiency of data management, but also helps to perform more refined security control on the data; although random noise is added, this method allows inverse operations to be performed through the random matrix information in the metadata when necessary, thereby restoring the true value of the data or performing further analysis. Accurate calculations can be performed to ensure that data can meet the needs of business analysis and processing while protecting privacy; during the query process, combined with user permission information and query data attribute information, refined access control of data can be achieved, ensuring that only users with corresponding permissions can access specific data, further enhancing data security; by pre-generating data packets and random matrix information, data can be quickly located and processed during query, reducing the need for real-time calculation and processing, thereby improving data processing efficiency; due to the combination of multiple privacy protection methods (such as data classification, noise injection, metadata management, etc.), it can better cope with complex and changing privacy threats and provide more comprehensive data protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some of the embodiments described in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:

[0022] Figure 1 A flowchart of a data privacy protection method based on random noise provided in an embodiment of this specification;

[0023] Figure 2 This is an example diagram of the composition of a noise injection obfuscated data packet provided in an embodiment of this specification;

[0024] Figure 3 This is a schematic diagram of the structure of a data privacy protection device based on random noise provided in an embodiment of this specification. DETAILED DESCRIPTION

[0025] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this specification without creative work should fall within the scope of protection of this specification.

[0026] With the development of information technology, data collection has become a critical application across various industries. As a crucial component of data collection systems, data collection interfaces help protect user privacy and assist data collection systems in complying with user information protection regulations, thereby preventing the inclusion of user-specific data within the data provided by the platform. Currently, the industry has made some progress in data privacy protection. Traditional methods of data privacy protection primarily include data encryption, anonymization, and access control. Data encryption technology encrypts raw data to ensure data security during transmission and storage. Anonymization technology desensitizes data to hide sensitive personal information and prevent the leakage of personal privacy. Access control technology restricts data access by setting appropriate access rights.

[0027] First, while traditional encryption technologies can protect data security, they can, in certain circumstances, reduce data availability and processing efficiency. Second, while anonymization technologies can conceal sensitive personal information, they can also compromise data accuracy and integrity. Furthermore, as attack vectors continue to evolve, a single privacy protection approach may not be effective against complex privacy threats. Therefore, existing data privacy protection methods still have limitations and cannot meet diverse data usage requirements, such as data availability and processing efficiency.

[0028] The embodiments of this specification provide a data privacy protection method based on random noise. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 This is a flow chart of a data privacy protection method based on random noise provided in an embodiment of this specification, such as Figure 1 As shown, it mainly includes the following steps:

[0029] Step S101: Acquire a plurality of original collected data to determine data description information of each original collected data.

[0030] The data description information includes data subject, data structure, data noise probability and data privacy rights;

[0031] In one embodiment of this specification, a collection system first acquires multiple raw data sets. These data may come from various data sources, such as sensors, databases, and log files. It is important to ensure that the collected data is complete and unprocessed to enable accurate subsequent analysis of its characteristics. For each raw data set, its data subject is determined. A data subject refers to the primary content or object described by the data. For example, if the data is a temperature measurement, the data subject is temperature. Next, the data structure of each raw data set is parsed. This describes the data's organization, including data types, field names, and field lengths. Parsing the data structure provides insights into the data's composition and format. Data noise refers to random errors or deviations in the data. Statistical methods or machine learning algorithms can be used to analyze the data's distribution and outliers, determining the probability of data noise for each raw data set. By assessing this probability, the reliability and accuracy of the data can be assessed. Furthermore, different data sets have different data privacy permissions, and data privacy permissions must be determined for each raw data set. Data privacy permissions refer to the permissions allowed to access, use, and share data. Setting different privacy permission levels for different data sets, based on data sensitivity and business needs, helps protect data privacy and prevent unauthorized access and misuse. Through the above steps, we can obtain multiple original collected data and determine the data description information of each original collected data, including data description information such as data subject, data structure, data noise probability and data privacy rights.

[0032] Step S102 : classifying the plurality of original collected data according to the data subject and data structure of each original collected data, and determining at least one data block.

[0033] Each of the data blocks includes a plurality of specified original collected data;

[0034] In one embodiment of the present specification, multiple original collected data are classified according to the data subject and data structure of each original collected data, and data with the same data subject and the same data structure are output to a binary file to generate data blocks, and at least one data block is determined.

[0035] Step S103 : Generate random matrix information corresponding to each data block, and write the data description information and the random matrix information into corresponding metadata to generate at least one data packet.

[0036] The random matrix information includes an evenly distributed random number matrix, and the data packet includes metadata and data blocks;

[0037] In one embodiment of the present specification, the random matrix information corresponding to each data block is determined, the data description information and the random matrix information are written into the corresponding metadata, and a data packet is created. It should be noted that the data packet is divided into two parts, including metadata and data blocks.

[0038] Generating random matrix information corresponding to each data block specifically includes: performing random number calculation on the specified original collected data in each data block to determine the specified random number corresponding to each specified original collected data; and constructing random matrix information according to the specified random number corresponding to each specified original collected data.

[0039] According to the specified random number corresponding to each of the specified original collected data, random matrix information is constructed, specifically including: obtaining the data quantity of the specified original collected data in each of the data blocks, and dividing the multiple specified original collected data to determine the number of groups corresponding to each of the data blocks; based on the data quantity and the number of groups, generating an average distribution random number matrix corresponding to the specified random number.

[0040] In one embodiment of the present specification, a random number calculation is performed on the specified original collected data in each data block to determine a specified random number corresponding to each specified original collected data. It should be noted that a random number generation strategy is predetermined. The random number generation strategy here can be determined based on business requirements and data characteristics. Pseudo-random number generators, cryptographic hash functions, etc. need to ensure that the generated random numbers are uniformly distributed, and for the same input data, the random numbers generated each time should be consistent to ensure the repeatability and verifiability of the results. Based on the specified random number corresponding to each specified original collected data, random matrix information is constructed.

[0041] In one embodiment of the present specification, the number of data of the specified original collected data in each data block is obtained, assuming that there are m data, and the multiple specified original collected data are divided to determine the number of groups corresponding to each data block, for example, divided into n groups, recorded as A m,n , generate an m-row n-column average distribution random number matrix R m,n .

[0042] Step S104 , based on the data noise probability and random matrix information corresponding to each data block, random noise is added to each data block to generate a noise injection obfuscated data packet for data storage.

[0043] Based on the data noise probability corresponding to each data block and the random matrix information, random noise is added to each data block to generate a noise injection obfuscation data packet, specifically including: obtaining an average distribution random number matrix in the random matrix information; locating data in the average distribution random number matrix according to the data noise probability to determine at least one noise point; adding random noise to each noise point according to the average distribution random number matrix to generate noise injection obfuscation data corresponding to each noise point, so as to determine the noise injection obfuscation data packet corresponding to each data block.

[0044] In one embodiment of this specification, Figure 2 This is an example diagram of the composition of a noise injection obfuscated data packet provided in an embodiment of this specification, such as Figure 2 As shown, a data packet includes metadata and data blocks. Using an evenly distributed random number matrix and the data noise probability, at least one noise point is selected from the data block. A noise block, also called a noise block, is generated based on the evenly distributed random number matrix. For each noise point, noise-injected obfuscated data corresponding to the noise point is added to generate the corresponding noise-injected obfuscated data. The obfuscated data packet after random noise addition is then generated based on the noise-injected obfuscated data, the metadata, and the original data corresponding to the non-noise point.

[0045] According to the data noise probability, data positioning is performed in the evenly distributed random number matrix to determine at least one noise point, specifically including: obtaining multiple specified random numbers in the evenly distributed random number matrix; and according to each of the specified random numbers and the data noise probability, determining the first noise point by taking a first specified random number that is smaller than the data noise probability.

[0046]

[0047] The random number c is the noise probability. For example, if the data has a 1% noise rate, c = 0.01. If the random number at a point is 0.12, which is greater than 0.01, it is set to 0, indicating that this point is a non-noise point. If the random number at a point is 0.009, which is less than 0.01, it is set to 1, indicating that this point is a noise point. It should be noted that the noise point in this specification can also be referred to as a noise point.

[0048] According to the evenly distributed random number matrix, random noise is added to each noise point to generate noise injection obfuscation data corresponding to each noise point, specifically including: determining a specified random number and original data corresponding to each noise point in the evenly distributed random number matrix; determining noise data according to the specified random number corresponding to each noise point, and adding random noise to the original data corresponding to each noise point through the noise data to generate noise injection obfuscation data corresponding to each noise point.

[0049] In one embodiment of the present specification, for noise points in an evenly distributed random number matrix, the designated random number and original data corresponding to each noise point are determined by using the evenly distributed random number matrix. The value of the noise point in Am,n is the sum of the original data and the noise data. The formula example is as follows:

[0050] Vm,n'=Vm,n+Rm,n,

[0051] Among them, Vm,n' is the data after random noise is added, that is, the noise-injected obfuscated data, Vm,n is the original data, and Rm,n is the random number in the evenly distributed random number matrix, that is, the noise data.

[0052] The above technical solution obtains an evenly distributed random number matrix from the random matrix information, which helps understand the distribution of elements in the random matrix. An evenly distributed random number matrix means that the elements in the matrix are distributed with equal probability within a certain range. Based on the data noise probability, data is located in the evenly distributed random number matrix, and at least one noise point is identified. The data noise probability reflects the likelihood of noise in the data. By combining it with the evenly distributed random number matrix, potential noise can be more accurately identified. Random noise is added to each noise point based on the evenly distributed random number matrix, and noise-injected obfuscated data corresponding to each noise point is generated. This is to simulate the noise conditions in real-world data, making the obfuscated data more realistic and improving the authenticity of data processing. The corresponding noise-injected obfuscated data packets are determined for each data block. By applying these packets to the corresponding data blocks, data can be effectively obfuscated and anonymized while protecting data privacy, preventing unauthorized access and misuse.

[0053] Step S105 : receiving query request information from a data query user, and determining query data through the query request information and random matrix information.

[0054] The query request information includes the user authority information of the data query user and the query data attribute information.

[0055] Determining query data through the query request information and the random matrix information specifically includes: obtaining user authority information of the data query user in the query request information, and determining query authority based on the data privacy authority and the user authority information, wherein the data privacy authority includes the data owner, and the query authority includes data owner query and non-data owner query; when the data query user is the data owner query, obtaining a query data packet corresponding to the query request information, wherein the query data packet includes metadata and a noise injection data block; performing an inverse operation on the noise injection data block through the random matrix information in the metadata to restore the original collected data and determine the query data; when the data query user is the non-data owner query, determining the query data according to the query data attribute information.

[0056] In one embodiment of the present specification, the result of the collected data packet with random noise inserted is processed by a query engine. The query summary result reduces the risk of information disclosure through random noise, and the random noise function converts the query result as follows: First, the query request information of the data query user is received, wherein the query request information includes the user authority information of the data query user and the query data attribute information. The query authority is determined based on the data owner and the user authority information of the query user in the data privacy authority. The query authority here includes data owner query and non-data owner query. For the data owner, the data is visible to them. When performing queries and calculations, the query engine uses the noise matrix recorded in the metadata to perform inverse operations on the data, remove the noise pollution to the original data, restore the true value, and then process the data.

[0057] Determining the query data according to the query data attribute information specifically includes: obtaining the query data attribute information, wherein the query data attribute information includes summary data attributes and calculation data attributes; when the query data attribute information is the calculation data attribute, performing an inverse operation on the noise-injected data block through the random matrix information in the metadata to restore the original collected data and determine the query data; when the query data attribute information is the summary data attribute, obtaining the actual summary result data corresponding to the query request information; obtaining the disturbance factor in the data description information to add noise to the actual summary result data based on the disturbance factor to determine the query data.

[0058] The results of the collected data packets with random noise inserted are processed by the query engine. The main result query methods are: total number, maximum, minimum, average, grouping and standard deviation. For non-data owners, the total number, maximum and minimum queries for the stored data have a high probability of hitting the real data. In this case, the query engine confuses the real data by adding noise to each query result. The scale of the random noise is proportional to the upper and lower limits of the restricted value range. F'=(1+p)F, where F' represents the random noise summary result, F represents the real summary result data, and p is the disturbance factor preset by the original data metadata information when the data is stored. Each query fluctuates within the preset fluctuation range, such as setting 0.003-0.010. The main result query methods are: total number, maximum, minimum, average and standard deviation.

[0059] For aggregated data, including queries such as total number of entries, maximum, and minimum, and grouped queries, the system calculates the proportion of noise data used in each row and removes rows with too little noise data. For grouped data with reasonable noise, calculations running on the same dataset may randomly remove different noise rows, resulting in inconsistent query results. For calculated data such as mean and standard deviation, inverse operations are performed using the noise matrix recorded in the metadata, without affecting the overall calculation results.

[0060] The above technical solution significantly enhances data privacy protection by adding random noise to each piece of raw collected data. The noise injection makes the original data difficult to directly identify or recover, effectively preventing data leakage and abuse. By classifying data and generating data packages, data with similar data themes and structures can be centrally managed, which not only improves data management efficiency but also facilitates more refined data security control. Despite the addition of random noise, this method allows inverse operations to be performed when needed using the random matrix information in the metadata to restore the true value of the data or perform accurate calculations, ensuring that the data still meets the needs of business analysis and processing while protecting privacy. During the query process, the combination of user permission information and query data attribute information can achieve refined access control to the data, ensuring that only users with the corresponding permissions can access specific data, further enhancing data security. By pre-generating data packages and random matrix information, data can be quickly located and processed during queries, reducing the need for real-time calculation and processing, thereby improving data processing efficiency. Due to the combination of multiple privacy protection methods (such as data classification, noise injection, metadata management, etc.), it can better cope with complex and changing privacy threats and provide more comprehensive data protection.

[0061] The embodiment of this specification also provides a data privacy protection device based on random noise, such as Figure 3As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0062] An embodiment of the present specification also provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to: obtain multiple original collected data to determine data description information of each original collected data, wherein the data description information includes data subject, data structure, data noise probability, and data privacy rights; classify the multiple original collected data according to the data subject and data structure of each original collected data to determine at least one data block, wherein each data block includes multiple specified original collected data; generate random matrix information corresponding to each data block, write the data description information and the random matrix information into corresponding metadata to generate at least one data packet, wherein the random matrix information includes an evenly distributed random number matrix, and the data packet includes metadata and data blocks; based on the data noise probability corresponding to each data block and the random matrix information, add random noise to each data block to generate a noise injection obfuscated data packet for data storage; receive query request information of a data query user to determine query data through the query request information and the random matrix information, wherein the query request information includes user authority information of the data query user and query data attribute information.

[0063] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.

[0064] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0065] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0066] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0067] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0068] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0069] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0070] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0071] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0072] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0073] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0074] The foregoing description is merely one or more embodiments of this specification and is not intended to limit this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of one or more embodiments of this specification are intended to be within the scope of the claims of this specification.

Claims

1. A data privacy protection method based on random noise, characterized in that: The method comprises: Acquire a plurality of original collected data to determine data description information of each of the original collected data, wherein the data description information includes data subject, data structure, data noise probability, and data privacy rights; Classify the plurality of original collected data according to the data subject and data structure of each of the original collected data to determine at least one data block, wherein each of the data blocks includes a plurality of specified original collected data; Generating random matrix information corresponding to each of the data blocks, and writing the data description information and the random matrix information into corresponding metadata to generate at least one data packet, wherein the random matrix information includes an evenly distributed random number matrix, and the data packet includes the metadata and the data block; Based on the data noise probability corresponding to each data block and the random matrix information, random noise is added to each data block to generate a noise injection obfuscated data packet for data storage; Receive query request information from a data query user to determine query data through the query request information and the random matrix information, wherein the query request information includes user authority information of the data query user and query data attribute information.

2. The data privacy protection method based on random noise according to claim 1, characterized in that: Generating random matrix information corresponding to each data block specifically includes: Performing random number calculation on the designated original collected data in each of the data blocks to determine a designated random number corresponding to each of the designated original collected data; Random matrix information is constructed according to the specified random number corresponding to each of the specified original collected data.

3. The data privacy protection method based on random noise according to claim 2, characterized in that: Constructing random matrix information according to the specified random number corresponding to each of the specified original collected data, specifically including: Obtaining the number of specified original collected data in each of the data blocks, and dividing the plurality of specified original collected data to determine the number of groups corresponding to each of the data blocks; Based on the data quantity and the group number, an average distribution random number matrix corresponding to the specified random number is generated.

4. The data privacy protection method based on random noise according to claim 1, characterized in that: Adding random noise to each data block based on the data noise probability corresponding to each data block and the random matrix information to generate a noise injection obfuscated data packet specifically includes: Obtaining an average distribution random number matrix in the random matrix information; According to the data noise probability, data positioning is performed in the average distribution random number matrix to determine at least one noise point; According to the average distribution random number matrix, random noise is added to each noise point to generate noise injection obfuscation data corresponding to each noise point, so as to determine a noise injection obfuscation data packet corresponding to each data block.

5. The data privacy protection method based on random noise according to claim 4, characterized in that: According to the data noise probability, data positioning is performed in the average distribution random number matrix to determine at least one noise point, specifically including: Obtaining a plurality of specified random numbers from the evenly distributed random number matrix; According to each of the designated random numbers and the data noise probability, a first noise point is determined using a first designated random number that is smaller than the data noise probability.

6. The data privacy protection method based on random noise according to claim 4, characterized in that: Adding random noise to each noise point according to the average distribution random number matrix to generate noise injection obfuscated data corresponding to each noise point specifically includes: In the evenly distributed random number matrix, determining a designated random number and original data corresponding to each noise point; Noise data is determined according to a specified random number corresponding to each noise point, so as to add random noise to the original data corresponding to each noise point through the noise data to generate noise injection obfuscated data corresponding to each noise point.

7. The data privacy protection method based on random noise according to claim 1, characterized in that: Determining query data using the query request information and the random matrix information specifically includes: Obtaining user authority information of the data query user in the query request information, and determining query authority based on the data privacy authority and the user authority information, wherein the data privacy authority includes a data owner, and the query authority includes a data owner query and a non-data owner query; When the data query user is a data owner, obtaining a query data packet corresponding to the query request information, wherein the query data packet includes metadata and a noise injection data block; Performing an inverse operation on the noise-injected data block using the random matrix information in the metadata to restore the original collected data and determine the query data; When the data query user is not the data owner, the query data is determined according to the query data attribute information.

8. The data privacy protection method based on random noise according to claim 7, characterized in that: Determining the query data according to the query data attribute information specifically includes: Acquiring the query data attribute information, wherein the query data attribute information includes summary data attributes and calculated data attributes; When the query data attribute information is the calculated data attribute, performing an inverse operation on the noise injection data block using the random matrix information in the metadata to restore the original collected data and determine the query data; When the query data attribute information is the summary data attribute, obtaining the real summary result data corresponding to the query request information; A disturbance factor in the data description information is obtained, and noise is added to the real summary result data based on the disturbance factor to determine the query data.

9. A data privacy protection device based on random noise, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Acquire a plurality of original collected data to determine data description information of each of the original collected data, wherein the data description information includes data subject, data structure, data noise probability, and data privacy rights; Classify the plurality of original collected data according to the data subject and data structure of each of the original collected data to determine at least one data block, wherein each of the data blocks includes a plurality of specified original collected data; Generating random matrix information corresponding to each of the data blocks, and writing the data description information and the random matrix information into corresponding metadata to generate at least one data packet, wherein the random matrix information includes an evenly distributed random number matrix, and the data packet includes the metadata and the data block; Based on the data noise probability corresponding to each data block and the random matrix information, random noise is added to each data block to generate a noise injection obfuscated data packet for data storage; Receive query request information from a data query user to determine query data through the query request information and the random matrix information, wherein the query request information includes user authority information of the data query user and query data attribute information.

Citation Information

Patent Citations

  • Data processing method and device

    CN111143674A

  • Data analysis method, noise construction method, equipment and storage medium

    CN114282083A