A data dynamic publishing method and system based on differential privacy

By introducing sliding windows and DGIM algorithms to data processing for approximate statistics, and combining random perturbation algorithms for differential privacy protection, the problems of large time overhead, high spatial complexity and privacy vulnerability in the existing technology are solved, and efficient and secure dynamic data release is achieved.

CN115422236BActive Publication Date: 2025-06-06ANHUI UNIVERSITY OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211025116.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-06-06
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

The existing data processing and publishing methods based on differential privacy have problems such as high time overhead, high spatial complexity and susceptibility to third-party privacy attacks.

Method used

The dynamic data release method based on sliding window is adopted, and the approximate statistics are performed through the DGIM algorithm, and the data noise is added to the data by using the random perturbation algorithm to achieve differential privacy protection.

Benefits of technology

It improves the timeliness of data query, reduces spatial complexity, and effectively prevents data privacy leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115422236B_ABST
    Figure CN115422236B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of data processing technology, and discloses a method and system for dynamic data publishing based on differential privacy. The method includes: determining the publishing format of data; determining a sliding window of fixed length, and allowing real-time data flow to flow through the sliding window; at the same time, counting the sizes of all buckets in the sliding window at the current moment based on the DGIM algorithm to obtain the approximate statistical results of the data in the sliding window at the current moment; calculating the similarity results in the sliding window at the current moment and the previous moment, and adding probabilistic perturbations to the similarity measurement based on the random perturbation algorithm to obtain perturbation similarity results; if the perturbation similarity result is greater than a preset similarity threshold, then determining the interval to be published of the approximate statistical results of the data in the sliding window at the current moment to perform dynamic data publishing. The present invention not only has low time and space overheads, but also can effectively protect the privacy information of users and prevent third-party privacy attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for dynamic data publishing based on differential privacy. Background Art

[0002] In the current era of big data, in order for users to quickly and effectively obtain the information they need from a variety of complex data, it is necessary to first collect and classify various data and then publish the processed data to the public. For example, movie websites collect and process users' ratings and comments on movies, and publish them on the website platform to provide the public with references for watching movies. For another example, hospitals regularly collect information on seasonal infectious diseases, etc., and summarize and publish them to remind the public to take corresponding preventive measures. For another example, in the smart transportation system, the navigation system will collect travel information of each vehicle to predict traffic conditions, plan driving routes, and avoid traffic jams.

[0003] At the same time, since any data comes from a specific user end, it is bound to carry the privacy information corresponding to the user. For example, when a power supply website collects users' ratings and comments on movies, it needs to obtain the user's account information; from the current virtual account information, most of it is associated with the user's actual information. For another example, when a hospital collects disease information, it needs to obtain the patient's basic identity information including age, gender, name, etc. For another example, when collecting vehicle travel information, the driver's home address, work address and other address information will be obtained simultaneously. Therefore, when performing related processing such as data collection and classification, it is also necessary to use corresponding privacy protection such as anonymization, encryption or differential privacy for each data to prevent the leakage of users' personal privacy data.

[0004] However, in data processing and publishing, in order to improve the accuracy of information release, it is necessary to cache all data within a period of time and then process them uniformly. However, in the era of big data, data itself has extremely strong real-time characteristics; at the same time, in the process of continuous iteration and updating of data, the latest real-time data often has higher practical value. Therefore, the existing data processing and publishing methods of this type have the disadvantage of high time overhead, thereby reducing the actual reference role of the published results. Furthermore, since this type of method needs to store a large amount of historical data, it also leads to excessive spatial complexity of data processing, thereby reducing the efficiency of data processing.

[0005] Based on the above data processing and publishing methods, since each data is first cached uniformly, centralized differential privacy is often used to uniformly process its privacy. However, centralized differential privacy requires the introduction of a third party, so private data will inevitably be attacked by a third party, resulting in privacy leakage. Summary of the invention

[0006] The purpose of the present invention is to provide a method and system for dynamic data publishing based on differential privacy, which is used to solve the technical problems that the existing data processing and publishing based on differential privacy has high time overhead and high space complexity, and privacy protection is vulnerable to third-party privacy attacks.

[0007] To achieve the above object, the present invention proposes the following technical solutions:

[0008] The data dynamic publishing method based on differential privacy includes:

[0009] Determine a data publishing format; wherein the publishing format includes a plurality of publishing intervals;

[0010] Determine a sliding window of fixed length and make the real-time data flow through the sliding window; at the same time, count the sizes of all buckets in the sliding window at the current moment based on the DGIM algorithm to obtain the approximate statistical results of the data in the sliding window at the current moment; wherein, the size of the bucket is defined as the number of 1s counted therein; for the sliding windows at two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and by merging buckets with earlier timestamps, the number of buckets of the same size does not exceed a preset number;

[0011] Calculate the similarity results within the sliding window between the current moment and the previous moment, and add a probabilistic perturbation to the similarity measure based on a random perturbation algorithm to obtain a perturbation similarity result;

[0012] If the disturbance similarity result is greater than the preset disturbance similarity threshold, the interval to be published of the approximate statistical results of the data in the sliding window at the current moment is determined, and the data is updated and dynamically published after being processed based on the differential privacy algorithm.

[0013] Furthermore, obtaining the approximate statistical results of the data in the sliding window at the current moment includes:

[0014] Sum the sizes of all buckets to get a first count result;

[0015] Calculate half the size of the bucket with the earliest timestamp to get the second count result;

[0016] The difference between the first counting result and the second counting result is calculated to be the approximate statistical result of the data in the sliding window at the current moment.

[0017] Furthermore, the preset number of buckets of the same size is determined by the following steps:

[0018] The actual statistical result of the data in the sliding window at the i-th moment is H i =1+(r-1)(2 j-1); where r is the preset number of buckets of the same size to be determined, 2 j The size of the bucket with the earliest timestamp;

[0019] Determine the true statistical result H of the data in the sliding window at the i-th moment i The approximate statistical results The error between Among them, 2 j-1 is the size of the bucket adjacent to the bucket with the earliest timestamp;

[0020] Calculate the preset number of buckets of the same size:

[0021] Furthermore, the similarity results within the sliding window between the current moment and the previous moment are calculated, and a probabilistic perturbation is added to the similarity measure based on a random perturbation algorithm to obtain a perturbation similarity result; including:

[0022] Calculate the approximate statistical results within the sliding window at the current time i The final published result in the sliding window of the previous moment i-1 Similarity results

[0023] A random number is obtained based on the random perturbation algorithm. If the random number is less than or equal to the perturbation probability, the perturbation similarity result is determined to be any value in the interval (0,1); if the random number is greater than the perturbation probability, it is determined whether the similarity result T is greater than the similarity threshold T 0 ;

[0024] If T>T 0 , then the disturbance similarity result is determined to be 1; otherwise, the disturbance similarity result is determined to be 0;

[0025] Among them, the disturbance probability is Among them, ε is the privacy budget calculated by the M1 algorithm, and w is the length of the sliding window.

[0026] Furthermore, the determining of the interval to be published of the approximate statistical results of the data in the sliding window at the current moment includes:

[0027] A dynamic programming grouping algorithm is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment.

[0028] A data dynamic publishing system based on differential privacy, comprising:

[0029] A first building module is used to determine a data publishing format; wherein the publishing format includes a plurality of publishing intervals;

[0030] The data statistics module is used to determine a sliding window of fixed length and make the real-time data flow through the sliding window; at the same time, based on the DGIM algorithm, the sizes of all buckets in the sliding window at the current moment are counted to obtain the approximate statistical results of the data in the sliding window at the current moment; wherein, the size of the bucket is defined as the number of 1s counted therein; for the sliding windows at two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and the number of buckets of the same size does not exceed a preset number by merging buckets with earlier timestamps;

[0031] A random perturbation module, used to calculate the similarity results within the sliding window between the current moment and the previous moment, and add a probabilistic perturbation to the similarity metric based on a random perturbation algorithm to obtain a perturbation similarity result;

[0032] The dynamic publishing module is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment when the disturbance similarity result is greater than a preset disturbance similarity threshold, and to update and dynamically publish the data after processing it based on the differential privacy algorithm.

[0033] Further, including:

[0034] A first counting module, configured to sum the sizes of all the buckets to obtain a first counting result;

[0035] A second counting module, used for calculating half of the size of the bucket with the earliest timestamp to obtain a second counting result;

[0036] The third counting module is used to calculate the difference between the first counting result and the second counting result as an approximate statistical result of the data in the sliding window at the current moment.

[0037] Further, including:

[0038] The first calculation module is used to obtain the real statistical result of the data in the sliding window at the i-th moment as H i =1+(r-1)(2 j -1); where r is the preset number of buckets of the same size to be determined, 2 j The size of the bucket with the earliest timestamp;

[0039] The second calculation module is used to determine the true statistical result H of the data in the sliding window at the i-th moment i The approximate statistical results The error between Among them, 2 j-1 is the size of the bucket adjacent to the bucket with the earliest timestamp;

[0040] The third calculation module is used to calculate the preset number of buckets of the same size:

[0041] Further, including:

[0042] The fourth calculation module is used to calculate the approximate statistical results within the sliding window at the current time i The final published result in the sliding window of the previous moment i-1 Similarity results

[0043] The first judgment module is used to obtain a random number based on the random perturbation algorithm. If the random number is less than or equal to the perturbation probability, the perturbation similarity result is determined to be any value in the interval (0,1); if the random number is greater than the perturbation probability, it is determined whether the similarity result T is greater than the similarity threshold T 0 ;

[0044] The second judgment module is used to 0 When , the disturbance similarity result is determined to be 1; otherwise, the disturbance similarity result is determined to be 0;

[0045] Further, including:

[0046] The dynamic grouping module is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment by using a dynamic programming grouping algorithm.

[0047] Beneficial effects:

[0048] It can be seen from the above technical solutions that the technical solution of the present invention provides a method and system for dynamic data publishing based on differential privacy, which is used to simultaneously solve the defects of large time and space overhead and susceptibility to privacy attacks in the existing data publishing process.

[0049] The method includes: first, determining the data publishing format; wherein, the publishing format includes several publishing intervals. Secondly, determining a sliding window of fixed length, and allowing the real-time data stream to flow through the sliding window; at the same time, counting the sizes of all buckets in the sliding window at the current moment based on the DGIM algorithm to obtain the approximate statistical results of the data in the sliding window at the current moment. And the rules for updating the data in the sliding window at adjacent moments and counting it based on the DGIM algorithm are: in the sliding window of two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and by merging buckets with earlier timestamps, the number of buckets of the same size does not exceed the preset number.

[0050] It can be seen from the above steps that the technical solution uses data streams to collect data, and performs data queries through a sliding window added to the data stream. Based on the relativity of the movement of the sliding window on the data stream at different times, it is intuitively seen that the real-time updated data stream flows through the sliding window. Therefore, at any moment, the data in the sliding window are the latest collected data; thereby improving the timeliness of data query. At the same time, for the data stream, the sliding window at any moment includes a large amount of data. Based on the efficiency of data query, the technical solution also introduces the approximate statistics of the sliding window at any moment of the DGIM algorithm on the basis of the sliding window. And after analysis, when the DGIM algorithm is introduced for approximate statistics, its error does not exceed This ensures that the accuracy of data query results can be guaranteed by using the DGIM algorithm. At the same time, it can be seen that the process of approximate counting based on the sliding window is a dynamic statistical process that is carried out simultaneously with the update of the data stream; at this time, only the data in the sliding window at the current moment needs to be stored. In order to complete data collection in the prior art, a large amount of historical data needs to be stored. Therefore, the space complexity is effectively reduced.

[0051] Furthermore, the method also includes: then, calculating the similarity results in the sliding window at the current moment and the previous moment, and adding a probabilistic perturbation to the similarity measure based on a random perturbation algorithm to obtain a perturbation similarity result. Finally, if the perturbation similarity measure is greater than a preset perturbation similarity threshold, the interval to be published of the approximate statistical results of the data in the sliding window at the current moment is determined, and the data is updated and dynamically published after being processed based on the differential privacy algorithm.

[0052] It can be seen from the above steps that when the data is dynamically released, this technical solution uses a differential privacy algorithm to protect the privacy of the data. At the same time, based on the continuity of the data stream, although random responses are added to the data finally released. However, a third party can infer the true situation of the data through the difference between the data in the sliding window at two consecutive moments. Therefore, before the data is confirmed and released, this technical solution adds disturbances to the approximate statistical results obtained by direct query based on the random perturbation algorithm, that is, the similarity results of two adjacent moments are perturbed. Thereby effectively preventing the privacy leakage of data in the entire technical solution.

[0053] In summary, this technical solution designs a new LSHP algorithm to realize dynamic data release. For the entire LSHP algorithm, its processing process is mainly divided into: (1) approximate statistical counting based on the sliding window model; (2) noise grouping process based on approximate statistical results. As for the time cost, the time cost of approximate statistical counting is O(log w), and the time cost of noise grouping is O(w); therefore, the time cost of the LSHP algorithm is O(w). The time cost of existing differential privacy algorithms is generally: O(w 2 ), O(w log w), etc. Therefore, the LSHP algorithm has a smaller time overhead as a whole. As for the space overhead, the storage of timestamps in the approximate statistical counting process requires logw, and the size of the storage bucket requires O(log log w). The subsequent approximate statistical results and privacy budget are all constants, so the total space overhead is O((log w) 2 ). The space overhead of existing differential privacy algorithms is generally: O(w 2 ), O(w log w), etc. Therefore, the LSHP algorithm has a smaller space overhead as a whole.

[0054] It should be appreciated that all combinations of the foregoing concepts, as well as additional concepts described in greater detail below, may be considered to be part of the inventive subject matter of the present disclosure, provided such concepts are not mutually inconsistent.

[0055] The foregoing and other aspects, embodiments and features of the present invention can be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the present invention, such as the features and / or beneficial effects of the exemplary embodiments, will be apparent from the following description or learned from the practice of the specific embodiments according to the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in various figures may be represented by the same reference numeral. For clarity, not every component is labeled in every figure. Embodiments of various aspects of the present invention will now be described by way of example and with reference to the accompanying drawings, in which:

[0057] Figure 1 This is a flow chart of the data dynamic publishing method based on differential privacy described in this embodiment;

[0058] Figure 2 for Figure 1 Flowchart for approximate counting in each sliding window;

[0059] Figure 3 for Figure 1 A flow chart for determining a preset number of buckets of the same size;

[0060] Figure 4 for Figure 1 Flowchart for determining the perturbation similarity results;

[0061] Figure 5 for Figure 1 Flowchart of the final interval to be released. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be the common meanings understood by people with ordinary skills in the field to which the present invention belongs.

[0063] The words "first", "second" and similar words used in the patent application specification and claims of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, unless the context clearly indicates otherwise, the singular form of "a", "an" or "the" and other similar words do not indicate a quantitative limitation, but indicate the existence of at least one. "Include" or "comprise" and other similar words mean that the elements or objects appearing before "include" or "comprise" include the features, wholes, steps, operations, elements and / or components listed after "include" or "comprise", and do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0064] There are currently three main research directions for means of protecting data privacy: anonymization, encryption, and differential privacy. Compared with the first two, the latter defines an extremely strict attack model and conducts quantitative analysis on the background knowledge possessed by the attacker, that is, no matter how much background knowledge the attacker has, he cannot obtain user privacy from it. However, the commonly used centralized differential privacy is based on the premise of a trusted third party, which is not true in practical applications. At the same time, the existing dynamic publishing method for dynamic data sets also requires caching a large amount of historical data when storing and processing data, resulting in excessive spatial complexity and inability to achieve efficient data processing. In view of this, this embodiment provides a data dynamic publishing method based on differential privacy, and designs a new algorithm, namely the LSHP algorithm, for the large, fast, and real-time characteristics of data streams, as well as the need to pay attention to the latest data. To simultaneously solve the defects in time overhead, space overhead and privacy protection in the existing methods.

[0065] The following is a further detailed introduction to the data dynamic publishing method based on differential privacy disclosed in the present invention in conjunction with specific embodiments.

[0066] Combination Figure 1 As shown, the method comprises the following steps:

[0067] Step S102: determine a data publishing format; wherein the publishing format includes a plurality of publishing intervals.

[0068] In this step, the data publishing format may be a histogram.

[0069] Step S104, determine a sliding window of fixed length, and make the real-time data stream flow through the sliding window; at the same time, count the sizes of all buckets in the sliding window at the current moment based on the DGIM algorithm to obtain the approximate statistical results of the data in the sliding window at the current moment; wherein, the size of the bucket is defined as the number of 1s counted therein; for the sliding window at two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and by merging buckets with earlier timestamps, the number of buckets of the same size does not exceed a preset number.

[0070] In the specific implementation, a bucket is initialized at the initial moment, and then a preset number of buckets of the same size is set. As the data stream is updated, the number of buckets of the same size does not exceed the preset number, and the buckets in the sliding window are sorted from right to left in ascending order, with the leftmost bucket being the earliest bucket, which stores earlier data; and as new data flows in, new buckets are continuously set and buckets are merged.

[0071] Specifically, the timestamp is set at the end position of the corresponding bucket, which may be an index number.

[0072] As a specific implementation method, Figure 2 As shown, the approximate statistical results of the data in the sliding window at the current moment are performed through the following steps:

[0073] Step S202: sum the sizes of all buckets to obtain a first counting result.

[0074] Step S204: Calculate half the size of the bucket with the earliest timestamp to obtain a second counting result.

[0075] Step S206: Calculate the difference between the first counting result and the second counting result, which is the approximate statistical result of the data in the sliding window at the current moment.

[0076] In step S206, for the bucket with the earliest timestamp, since it is impossible to confirm how many bits in the bucket with the earliest timestamp are still in the current sliding window as the data stream and the data in the sliding window are updated in real time, it is assumed that the 0s and 1s in the earliest bucket are evenly distributed, and half of its size is the number of 1s in it. This can make the accuracy of the approximate statistical results obtained in any case the highest.

[0077] Since approximate counting is performed in this step, it is necessary to control the error of the approximate statistical result from the perspective of data release. However, it is shown in practice that the error is related to the preset number of buckets of the same size. Therefore, it needs to be accurately set.

[0078] Combination Figure 3 As shown, the preset number of buckets of the same size is determined by the following steps:

[0079] Step S302: Obtain the real statistical result of the data in the sliding window at the i-th moment as H i =1+(r-1)(2 j -1); where r is the preset number of buckets of the same size to be determined, 2 j The size of the bucket with the earliest timestamp.

[0080] Specifically, H i =1+(r-1)(1+2+3+…+2 j- 2+2 j-1 )=1+(r-1)(2 j -1).

[0081] Step S304: Determine the true statistical result H of the data in the sliding window at the i-th moment i The approximate statistical results The error between Among them, 2 j-1 is the size of the bucket adjacent to the bucket with the earliest timestamp.

[0082] Specifically, the approximate statistical results Obtained by steps S202 to S206.

[0083] Step S306: Calculate the preset number of buckets of the same size as:

[0084] The conclusion obtained from step S306 is that in order to ensure that the error does not affect the accuracy of the final data release, the preset number of buckets of the same size should not exceed Based on this, during the approximate statistics of data in the sliding window at different times, the preset number of buckets of the same size can be continuously adjusted based on the statistical results to optimize the statistical results.

[0085] As a specific implementation, when performing step S104, it also includes:

[0086] First, the real statistical results H in the sliding window at the corresponding time are calculated simultaneously according to the preset frequency i And approximate statistical results

[0087] Secondly, calculate If it is not less than the preset number of buckets of the same size currently actually used, the preset number is kept unchanged; otherwise, the preset number is adjusted.

[0088] Step S106: Calculate the similarity results within the sliding window between the current moment and the previous moment, and add a probabilistic perturbation to the similarity measure based on a random perturbation algorithm to obtain a perturbation similarity result.

[0089] Combination Figure 4 As shown, as a specific implementation, step S106 is specifically performed through the following steps:

[0090] Step S106.2: Calculate the approximate statistical results within the sliding window at the current time i The final published result in the sliding window of the previous moment i-1 Similarity results

[0091] Step S106.4: Obtain a random number based on the random perturbation algorithm. If the random number is less than or equal to the perturbation probability, determine the perturbation similarity result as any value in the interval (0,1); if the random number is greater than the perturbation probability, determine whether the similarity result T is greater than the similarity threshold T. 0 .

[0092] Step S106.6: If T>T 0, then the disturbance similarity result is determined to be 1; otherwise, the disturbance similarity result is determined to be 0.

[0093] Among them, the disturbance probability is Among them, ε is the privacy budget calculated by the M1 algorithm, and w is the length of the sliding window.

[0094] Specifically, the perturbation process in this step satisfies ε-differential privacy. The specific proof process is as follows:

[0095] Define v i is the similarity result T and similarity threshold T between the sliding window data at adjacent moments 0 The result after comparing the size, v i ' is the result output by the similarity calculation algorithm.

[0096] When the output v' i =1, we have:

[0097]

[0098] When the output v' i =0, we have:

[0099]

[0100] It can be determined that the disturbance satisfies Differential privacy.

[0101] Furthermore, according to the differential privacy property, for a sliding window with a fixed length of w, the privacy budget obtained by the M1 algorithm is:

[0102]

[0103] Step S108: If the disturbance similarity result is greater than a preset disturbance similarity threshold, determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment, and process it based on the differential privacy algorithm before updating and dynamically publishing the data.

[0104] Combination Figure 5 As shown, in order to reduce the error in the data release stage, the following method is used to determine the interval to be released:

[0105] Step S108.2: Use a dynamic programming grouping algorithm to determine the interval to be published for the approximate statistical results of the data in the sliding window at the current moment.

[0106] From step S102 to step S108, it can be seen that this embodiment designs a new algorithm, specifically defined as the LSHP algorithm, which realizes the dynamic release of data. In order to prove that the LSHP algorithm performs well in terms of space overhead and time overhead, the following proof is specifically performed:

[0107] The processing of the LSHP algorithm is mainly divided into: (1) approximate statistical counting based on a sliding window model; and (2) noise grouping based on approximate statistical results.

[0108] Regarding the time cost: the time cost of approximate statistical counting is O(log w), and the time cost of noise grouping is O(w); therefore, the time cost of the LSHP algorithm is O(w). The time cost of existing differential privacy algorithms is generally: O(w 2 ), O(w log w), etc. Therefore, the LSHP algorithm has a smaller time overhead as a whole.

[0109] Regarding the space cost: Assume that the size of the sliding window is w, and the data stored in the bucket created by the algorithm DGIM includes timestamps and statistical count values. Consider the extreme case: the data in the sliding window at the current moment is all 1, then the space required to store the timestamp is O(logw). Similarly, the space required to store the statistical count results in this case is O(log w). The data that a bucket needs to store includes timestamps and statistical count results, so the total space cost is O((log w) 2 ).

[0110] The space overhead of existing differential privacy algorithms is generally: O(w 2 ), O(w log w), etc. Therefore, the LSHP algorithm has a smaller space overhead as a whole.

[0111] The above program can be run in the processor, or it can also be stored in the memory (or computer-readable storage medium), which includes permanent and non-permanent, removable and non-removable media. Information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer-readable media does not include temporary computer-readable media, such as modulated data signals and carrier waves.

[0112] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more blocks can be implemented by different modules corresponding to different steps.

[0113] In this embodiment, such a system is provided, which can be called a data dynamic publishing system based on differential privacy. The system includes:

[0114] The first building module is used to determine a data publishing format; wherein the publishing format includes a plurality of publishing intervals.

[0115] The data statistics module is used to determine a sliding window of fixed length and make the real-time data flow through the sliding window; at the same time, based on the DGIM algorithm, the sizes of all buckets in the sliding window at the current moment are counted to obtain the approximate statistical results of the data in the sliding window at the current moment; wherein, the size of the bucket is defined as the number of 1s counted therein; for the sliding window at two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and by merging buckets with earlier timestamps, the number of buckets of the same size does not exceed a preset number.

[0116] The random perturbation module is used to calculate the similarity results within the sliding window between the current moment and the previous moment, and add probabilistic perturbation to the similarity metric based on the random perturbation algorithm to obtain a perturbation similarity result.

[0117] The dynamic publishing module is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment when the disturbance similarity result is greater than a preset disturbance similarity threshold, and to update and dynamically publish the data after processing it based on the differential privacy algorithm.

[0118] The system is used to implement the steps of the method in the above embodiment, which have been described and will not be repeated here.

[0119] For example, it also includes:

[0120] The first counting module is used to sum the sizes of all the buckets to obtain a first counting result.

[0121] The second counting module is used to calculate half of the size of the bucket with the earliest timestamp to obtain a second counting result.

[0122] The third counting module is used to calculate the difference between the first counting result and the second counting result, which is the approximate statistical result of the data in the sliding window at the current moment.

[0123] For example, it also includes:

[0124] The first calculation module is used to obtain the real statistical result of the data in the sliding window at the i-th moment as H i =1+(r-1)(2 j -1); where r is the preset number of buckets of the same size to be determined, 2 j The size of the bucket with the earliest timestamp.

[0125] The second calculation module is used to determine the true statistical result H of the data in the sliding window at the i-th moment i The approximate statistical results The error between Among them, 2 j-1 is the size of the bucket adjacent to the bucket with the earliest timestamp.

[0126] The third calculation module is used to calculate the preset number of buckets of the same size:

[0127] For example, it also includes:

[0128] The fourth calculation module is used to calculate the approximate statistical results within the sliding window at the current time i The final published result in the sliding window of the previous moment i-1 Similarity results

[0129] The first judgment module is used to obtain a random number based on the random perturbation algorithm. If the random number is less than or equal to the perturbation probability, the perturbation similarity result is determined to be any value in the interval (0,1); if the random number is greater than the perturbation probability, it is determined whether the similarity result T is greater than the similarity threshold T 0 .

[0130] The second judgment module is used to 0 When , the disturbance similarity result is determined to be 1; otherwise, the disturbance similarity result is determined to be 0.

[0131] For example, it also includes:

[0132] The dynamic grouping module is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment by using a dynamic programming grouping algorithm.

[0133] Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. A person with ordinary knowledge in the technical field to which the present invention belongs may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the definition of the claims.

Claims

1. A method for dynamic data publishing based on differential privacy, It is characterized in that include: Determine a data publishing format; wherein the publishing format includes a plurality of publishing intervals; Determine a sliding window of fixed length and make the real-time data flow through the sliding window; at the same time, count the sizes of all buckets in the sliding window at the current moment based on the DGIM algorithm to obtain the approximate statistical results of the data in the sliding window at the current moment; wherein, the size of the bucket is defined as the number of 1s counted therein; for the sliding windows at two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and by merging buckets with earlier timestamps, the number of buckets of the same size does not exceed a preset number; The acquisition of the approximate statistical result of the data in the sliding window at the current moment includes: first, summing the sizes of all buckets to obtain a first counting result; second, calculating half of the size of the bucket with the earliest timestamp to obtain a second counting result; then, calculating the difference between the first counting result and the second counting result as the approximate statistical result of the data in the sliding window at the current moment; The preset number of buckets of the same size is determined by the following steps: First, the actual statistical result of the data in the sliding window at the i-th moment is obtained as H i =1+(r-1)(2 j -1); where r is the preset number of buckets of the same size to be determined, 2 j is the size of the bucket with the earliest timestamp; secondly, determine the true statistical result H of the data in the sliding window at the i-th moment i The approximate statistical results The error between Among them, 2 j-1 is the size of the bucket adjacent to the bucket with the earliest timestamp; then, the preset number of buckets of the same size is calculated as: Calculate the similarity results within the sliding window between the current moment and the previous moment, and add probabilistic perturbations to the similarity measure based on the random perturbation algorithm to obtain perturbation similarity results; If the disturbance similarity result is greater than a preset disturbance similarity threshold, the interval to be published of the approximate statistical results of the data in the sliding window at the current moment is determined, and the data is updated and dynamically published after being processed based on the differential privacy algorithm.

2. According to the method for dynamic data publishing based on differential privacy according to claim 1, It is characterized in that The similarity results within the sliding window between the current moment and the previous moment are calculated, and a probabilistic perturbation is added to the similarity measure based on a random perturbation algorithm to obtain a perturbation similarity result; including: Calculate the approximate statistical results within the sliding window at the current time i The final published result in the sliding window of the previous moment i-1 Similarity results A random number is obtained based on the random perturbation algorithm. If the random number is less than or equal to the perturbation probability, the perturbation similarity result is determined to be any value in the interval (0,1); if the random number is greater than the perturbation probability, it is determined whether the similarity result T is greater than the similarity threshold T 0 ; If T>T 0 , then the disturbance similarity result is determined to be 1; otherwise, the disturbance similarity result is determined to be 0; Among them, the disturbance probability is Among them, ε is the privacy budget calculated by the M1 algorithm, and w is the length of the sliding window.

3. According to the method for dynamic data publishing based on differential privacy according to claim 1, It is characterized in that The determining of the interval to be published of the approximate statistical results of the data in the sliding window at the current moment includes: A dynamic programming grouping algorithm is used to determine the interval to be published for the approximate statistical results of the data in the sliding window at the current moment.

4. A data dynamic publishing system based on differential privacy, It is characterized in that include: A first building module is used to determine a data publishing format; wherein the publishing format includes a plurality of publishing intervals; The data statistics module is used to determine a sliding window of fixed length and make the real-time data flow through the sliding window; at the same time, based on the DGIM algorithm, the sizes of all buckets in the sliding window at the current moment are counted to obtain the approximate statistical results of the data in the sliding window at the current moment; wherein, the size of the bucket is defined as the number of 1s counted therein; for the sliding windows at two adjacent moments, only one bit is updated, and a new bucket is created when the data on the updated bit is 1; and the number of buckets of the same size does not exceed a preset number by merging buckets with earlier timestamps; The data statistics module includes the following components to obtain the approximate statistical results of the data in the sliding window at the current moment: A first counting module, configured to sum the sizes of all the buckets to obtain a first counting result; A second counting module, used for calculating half of the size of the bucket with the earliest timestamp to obtain a second counting result; A third counting module, used for calculating the difference between the first counting result and the second counting result as an approximate statistical result of the data in the sliding window at the current moment; The data statistics module includes the following components to determine the preset number of buckets of the same size: The first calculation module is used to obtain the real statistical result of the data in the sliding window at the i-th moment as H i =1+(r-1)(2 j -1); where r is the preset number of buckets of the same size to be determined, 2 j The size of the bucket with the earliest timestamp; The second calculation module is used to determine the true statistical result H of the data in the sliding window at the i-th moment i The approximate statistical results The error between Among them, 2 j-1 is the size of the bucket adjacent to the bucket with the earliest timestamp; The third calculation module is used to calculate the preset number of buckets of the same size: The random perturbation module is used to calculate the similarity results within the sliding window between the current moment and the previous moment, and add probabilistic perturbations to the similarity metric based on the random perturbation algorithm to obtain the perturbation similarity results; The dynamic publishing module is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment when the disturbance similarity result is greater than a preset disturbance similarity threshold, and to update and dynamically publish the data after processing it based on the differential privacy algorithm.

5. According to the data dynamic publishing system based on differential privacy according to claim 4, It is characterized in that include: The fourth calculation module is used to calculate the approximate statistical results within the sliding window at the current time i The final published result in the sliding window of the previous moment i-1 Similarity results The first judgment module is used to obtain a random number based on the random perturbation algorithm. If the random number is less than or equal to the perturbation probability, the perturbation similarity result is determined to be any value in the interval (0,1); if the random number is greater than the perturbation probability, it is determined whether the similarity result T is greater than the similarity threshold T 0 ; The second judgment module is used to 0 When , the disturbance similarity result is determined to be 1; otherwise, the disturbance similarity result is determined to be 0.

6. According to the differential privacy-based data dynamic publishing system of claim 4, It is characterized in that include: The dynamic grouping module is used to determine the interval to be published of the approximate statistical results of the data in the sliding window at the current moment by using a dynamic programming grouping algorithm.

Citation Information

Patent Citations

  • Localized differential privacy data stream publishing method for real-time data

    CN114662152A

  • Crowd-based scores for experiences from measurements of affective response

    US20160170996A1