Shuffling differential privacy-based data stream rapid histogram publishing method

By adopting a data flow fast histogram publishing method based on shuffled differential privacy in the data flow scenario, and using the sliding window model and the random response mechanism optimized by hash function, the privacy leakage risks, low computing efficiency and error accumulation problems in the data flow scenarios are solved, and efficient data privacy protection and histogram publishing are achieved.

CN120144641APending Publication Date: 2025-06-13ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318906.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional differential privacy methods have problems with privacy leakage risks, low computing efficiency and error accumulation in data flow scenarios, making it difficult to effectively protect data privacy and improve data processing efficiency.

Method used

The data flow rapid histogram publishing method based on shuffling differential privacy is adopted. Data is processed through a random response mechanism optimized by sliding window model and hash function, and disturbed data sets are generated, and data shuffling and sampling rate are dynamically adjusted at the shuffling end to ensure the improvement of data privacy protection and processing efficiency.

Benefits of technology

Without relying on trusted third parties, effectively protect user data privacy, control the time and space costs of data processing, reduce error accumulation, and improve the release availability and accuracy of data flow histograms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144641A_ABST
    Figure CN120144641A_ABST
Patent Text Reader

Abstract

The invention discloses a shuffling differential privacy-based data stream rapid histogram publishing method, and belongs to the technical field of computers. The method comprises the following steps: S1, a publishing end publishes a histogram query request, and a user end responds; s2, the user side uses a sliding window model to process a local data stream, and samples data in a sliding window; the method comprises the steps of S1, adopting data, S2, adopting the data, S3, carrying out disturbance and noise adding processing on the adopted data, sending the data to a shuffling end for shuffling, and generating a disturbance data set, S4, receiving the disturbance data set by a publishing end, carrying out unbiased statistical processing on the disturbance data set, drawing a histogram according to a processing result, and publishing the histogram, and S5, if the publishing end continuously keeps a request, circulating the steps S2 to S4 until the publishing end stops the request, and ending the histogram publishing process. The method provided by the invention can effectively solve the problems of privacy leakage risk, low calculation efficiency and error accumulation existing in a data stream scene in a traditional differential privacy method, improves the data processing efficiency and ensures the histogram availability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and more specifically, relates to a data stream fast histogram publishing method based on shuffle differential privacy. Background Art

[0002] With the rapid development of network technology and the widespread application of the industrial Internet, various applications in data flow scenarios, such as event monitoring, log flow analysis, and traffic management, have been greatly promoted, which has directly led to an unprecedented surge in the amount of data collected on various terminals. This data torrent has brought unprecedented challenges to data governance and privacy protection strategies, and has also prompted companies to increasingly rely on a wide range of data sets in their decision-making process.

[0003] In data stream analysis, histograms, as basic statistical summaries that capture data distribution, play a vital and fundamental role in various analysis tasks. However, in the field of streaming data, if the real-time generation of histograms is not properly handled, it may cause huge privacy risks. This is because as servers continue to collect and publish and share data that can reflect personal characteristics (such as location information, consumption habits, medical diagnosis, etc.), personal privacy is easily leaked, which may lead to serious consequences such as information leakage, discrimination, and unauthorized dangerous operations.

[0004] Therefore, how to effectively publish the real-time histogram of data streams while protecting privacy has become a major issue that needs to be solved urgently. In this context, differential privacy has received great attention and has been widely studied in the field of data stream protection because of its advantages of ignoring background knowledge and being able to quantitatively analyze the degree of privacy protection and availability.

[0005] At present, the histogram publishing methods based on differential privacy applied to data streams can be mainly divided into two categories: centralized differential privacy and localized differential privacy. However, these methods have some problems in practical applications. Centralized differential privacy relies on a trusted third party, which may be an unstable factor in the streaming data setting; although localized differential privacy does not require a trusted third party, as the user scale grows, the data error will become larger, affecting the accuracy of the histogram; and the existing data stream histogram publishing methods based on these two models require a large amount of memory to store data during data collection and processing, and sacrifice some processing efficiency to ensure the availability of the final histogram. Summary of the invention

[0006] 1. Problem to be solved

[0007] The purpose of the present invention is to provide a method for quickly publishing a histogram of data streams based on shuffle differential privacy, which is used to solve the problems of privacy leakage risk, low computational efficiency, and error accumulation existing in traditional differential privacy methods in the data stream scenario, effectively improving the data processing efficiency and ensuring the availability of the histogram.

[0008] 2. Technical solution

[0009] To solve the above problems, the technical solution adopted by the present invention is as follows:

[0010] The present invention provides a method for quickly publishing a histogram of data streams based on shuffle differential privacy, including the following steps:

[0011] S1. The publishing end publishes a histogram query request, and the user end responds;

[0012] S2. The user end initializes and sets a sliding window, collects the elements with the status of not expired in the sliding window at a sampling rate r, and processes the data using the random response mechanism optimized by the hash function to generate a perturbed data set;

[0013] S3. The shuffling end receives all the perturbed data sets from the user ends, calculates the error of the perturbed data sets to dynamically adjust the sampling rate r, and shuffles all the perturbed data sets to obtain a re-arranged perturbed data set;

[0014] S4. The publishing end receives the re-arranged perturbed data set from the shuffling end, performs unbiased statistical processing on it, and then draws and publishes a histogram based on the processing result;

[0015] S5. If the publishing end continuously maintains the request, loop steps S2 to S4 until the publishing end terminates the request, and the histogram publishing process ends.

[0016] Furthermore, in step S2, the sliding window model is used to extract the element e in the data stream D t , and the sliding window at time t is denoted as W t , W t = {e t-W+1 , e t-W+2 ,..., e t}; W represents the size of the sliding window;

[0017] The auxiliary space A in the sampling set is a set of triples where: ts i is the timestamp of the element, id i identifies the position in the auxiliary space, and size i is the size of the data field used for random response perturbation calculation; W + b is the size of the auxiliary space A;

[0018] Let \(M\) denote the size of the sampling set. For each element \(e\) in the data stream \(D\) t , perform screening according to the following requirements:

[0019] If there are expired elements among \(A[1]\) to \(A[W + b]\), replace the oldest expired element with \(e\) t ;

[0020] If there are no expired elements in \(A[1]\) to \(A[W + b]\), ignore \(e\) t and wait for the next element;

[0021] If a sampling set of size \(M\) needs to be selected from \(A\), randomly select \(M\) non - expired elements from \(A\).

[0022] Furthermore, in step S2, during the sampling process, each sliding window in the data stream \(D\) is iteratively processed, and the screened elements are processed to ensure that the sliding window \(W\) at time \(t\) t and the sampling set are kept up - to - date and relevant.

[0023] Furthermore, in step S2, the data is perturbed using a random response mechanism optimized based on a hash function. The original data remains unchanged with probability \(\gamma\), and with probability \(1-\gamma\), the original data is mapped to other values within the hash value range, where \(\gamma=\sum_{min}Pr[GRR(h(v)) = y]=g / (g + e\) ε \(- 1)\), \(v\) is the private value of the user - end, \(y\) is the output value after perturbation processing, and \(Pr[GRR(b)=y]\) represents the probability that the input value is \(v\) and the output value is \(y\); \(g\) is the size of the hash function value range, and \(\epsilon\) is the privacy budget.

[0024] Furthermore, the privacy budget \(n\) is a positive integer, \(\delta\in(0,1)\),

[0025] Furthermore,

[0026]

[0027] where \(d\) is the value range of the private value \(v\) of the user - end.

[0028] Furthermore, in step S3, the method for dynamically adjusting the sampling rate \(r\) is as follows:

[0029] When \(\vert r_e\) t \(-r_e\) t-1 \vert>T\), the shuffling end increases the sampling rate when sampling at the next moment;

[0030] When \(\vert r_e\) t \(-r_e\) t-1 \vert\leq T\), the shuffling end maintains or reduces the sampling rate when sampling at the next moment;

[0031] re t The relative error of the perturbation data set calculated for the current moment of the system, re t-1 is the relative error of the release result of the previous moment retained by the system. T is the error threshold, and the value of T ranges from 0% to 5%, excluding 0%.

[0032] Furthermore, when the difference in relative error is greater than the error threshold, increase the sampling rate by the actual difference in the calculated relative error; when the difference in relative error is equal to the error threshold, consider the size of the response time of the query request at the publishing end. If the response time is lower than the user tolerance, maintain the current sampling rate; otherwise, reduce the current sampling rate by 0.05%; when the difference in relative error is less than the error threshold, maintain the current sampling rate for sampling at the next moment.

[0033] Furthermore, the method for shuffling the perturbation data set at the shuffling end is as follows:

[0034] Start from the last element in the data vector of the perturbation data set using the rand() function, and swap each element with the element corresponding to the subscript generated by rand() one by one from back to front until all elements are processed.

[0035] Furthermore, in step S4, the unbiased statistical processing method is as follows:

[0036] First, the publishing end counts the frequency of the values of the private value v of all user ends according to the perturbation data set sent by the shuffling end to count the frequency of the values of all user - side private values v;

[0037] Then, use the following formula to estimate the unbiased estimate value of the private value v

[0038]

[0039] where m 1 represents the number of users who map the private value v to the first position in the hash value range. g is the size of the hash function value range, and n is the number of users.

[0040] 3. Beneficial effects

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] (1) A fast histogram publishing method for data streams based on shuffle differential privacy according to the present invention processes data streams by using a sliding window model, dynamically samples and collects data within the sliding window, and then the client perturbs and adds noise to the data within the sliding window and sends it to the shuffle end for shuffling to generate a shuffled perturbed data set, which can effectively protect the user's data privacy without the intervention of a trusted third party;

[0043] (2) A fast histogram publishing method for data streams based on shuffle differential privacy according to the present invention performs shuffling processing after receiving and adding noise to the data at the shuffle end, and then calculates the relative error between the current data and the original data. After that, the result is compared with the relative error at the previous time point, and the sampling rate of the local end is adjusted according to the comparison result. By controlling the sampling rate, the time and space costs required for processing each local end data are strictly controlled, and the privacy amplification effect brought by the shuffling operation effectively alleviates the influence of the number of users on the statistical results. In addition, controlling the sampling rate can optimize the privacy budget spent on perturbation at each local end, thereby achieving the purpose of improving usability;

[0044] (3) A fast histogram publishing method for data streams based on shuffle differential privacy according to the present invention performs unbiased statistical processing on the perturbed data set, and then draws and publishes a histogram based on the processing result. The technical solution of the present invention can effectively address the spatio-temporal challenges existing in data stream histogram publishing, and publish a real-time histogram of stream data with high usability while ensuring privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a system framework diagram of a fast data stream histogram publishing method based on shuffle differential privacy according to the present invention;

[0046] Figure 2 is a flowchart of shuffle differential privacy histogram publishing in Embodiment 1 of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] Combined with Figure 1, a method for quickly publishing a data stream histogram based on shuffle differential privacy in this embodiment. By using shuffle differential privacy (SDP) and random sampling technology, it can achieve efficient histogram publishing and effectively protect the data privacy of the user side. In view of the characteristics of the current data stream, such as persistence, real-time nature, and large scale, compared with the processing of static data, the research on dynamic data stream publishing has more profound practical significance. The publishing method in this embodiment not only breaks through the limitations of static simple statistical data protection but also can be extended to more complex publishing scenarios, thus greatly enriching the application scope of privacy protection and comprehensively improving its application scenarios.

[0049] Meanwhile, it is worth noting that in all data stream histogram publishing problems, spatio-temporal efficiency is a key indicator to measure the quality of the method. The adopted publishing method directly affects the data processing efficiency and the error size between the published data and the real data, thus determining the practicality and feasibility of the data publishing method. By introducing the shuffle differential privacy technology with sampling, the present invention can effectively cope with the existing spatio-temporal challenges and publish a reliable real-time histogram of stream data while ensuring privacy.

[0050] The following details the histogram publishing method of the present invention, which specifically includes the following steps:

[0051] S1. The publishing end issues a histogram query request, and the user end responds;

[0052] At a certain moment t, the publishing end initiates a histogram query request to all user ends by broadcasting. After receiving the request, the user ends prepare to upload data.

[0053] S2. The user end initializes and sets a sliding window, samples the elements with the status of not expired within the sliding window at the sampling rate r, and processes the data using the random response mechanism optimized by the hash function to generate a perturbed data set. In the data stream scenario, real-time publishing tasks account for a certain proportion. In the data processing stage, sampling is performed on the data within the sliding window, and by processing the sampled part of the data, the time and space costs can be effectively saved. Although the above optimization will affect the accuracy of the results to a certain extent, it can significantly improve the response speed of real-time tasks and the space utilization efficiency.

[0054] Specifically, the data stream owned by each user end locally is denoted as D, D = e 1 , e 2 ,..., e t (t≥1). For any user end U i The given data stream can be denoted as D i . For any user end, the sliding window model is used to extract the element e from the data stream D t。The client initializes and sets a sliding window. The size of the sliding window is denoted as W, and the sliding window at time t is represented as W t , W t ={e t-W+1 , e t-W+2 ,..., e t}.

[0055] At time t, when each client responds to the histogram request and uploads data, the client maintains a sampling structure based on the dataset within the sliding window W t to collect elements whose status in the sliding window is shown as not expired at a sampling rate of r; then uses a randomized response mechanism optimized based on a hash function to process the data, maintaining the original data unchanged with probability γ. Maps the original data to other values within the hash value range with probability 1 - γ. After processing, uploads the perturbed data to the shuffle side.

[0056] It should be noted that when sampling, set the initial sampling rate r and sample the elements within the sliding window at sampling rate r. First, determine the size M of the sampling set. The elements contained in the auxiliary space A of size W + b have three parts: Among them: ts i is the timestamp of the element, id i identifies the position in the auxiliary space, and size i is the size of the data field used for randomized response perturbation calculation; when selecting a sampling set of size M from the auxiliary space A, follow the following rules:

[0057] If there are expired elements among A[1] to A[W + b], replace the oldest expired element with e t .

[0058] If there are no expired elements among A[1] to A[W + b], ignore e t and wait for the next element.

[0059] If a sampling set of size M needs to be selected from A, randomly select M non - expired elements from A.

[0060] It should be added that the auxiliary space maintains the timestamp metadata ts through explicit display iTo track the time attributes of elements, and then the old and new states of elements can be determined based on the numerical relationship of timestamps. When the window slides, the timeliness of the data set is ensured by performing a dynamic replacement strategy on the elements in the auxiliary space A: when a new element enters the window, the system first checks whether there are expired elements that violate the timeliness constraint. If such elements exist, a replacement operation is performed, and the oldest element is eliminated using the timeliness advantage of the new element; if all elements in the window meet the current timeliness requirements, the insertion operation of new elements is suspended. This adaptive insertion control strategy effectively avoids redundant timestamp verification and storage operations caused by blindly appending elements, and significantly reduces the time complexity of the algorithm.

[0061] In terms of the sampling strategy, the system strictly follows the "time decay" principle of the data stream, excludes expired elements based on the timestamp threshold, and implements probability sampling among the valid elements that have not expired, ensuring that the final sample set can not only reflect the temporal characteristics of the data stream but also maintain representativeness in terms of statistics.

[0062] Therefore, during the adoption process, each sliding window in the data stream is iteratively processed, and the above rules are used to process new elements (denoted by e t ), which can ensure that the current window (denoted by W t ) and the sampling set remain up-to-date and relevant, while ensuring the accuracy and timeliness of the histogram publishing result.

[0063] In addition, for the technical solution of this embodiment, after obtaining the sampling data, in order to protect data privacy, the private values v of all user terminals participating in the histogram publishing need to use the GRR random response mechanism that optimizes the value range with hashing to perturb the data GRR(v), and this perturbation mechanism satisfies the following conditions:

[0064]

[0065] Among them, Pr[GRR(v) = y] represents the probability that the input value is v and the output value is y, d is the value range of the user terminal's private value v, ε is related to the privacy budget and the sampling rate r n is a positive integer, δ ∈ (0, 1),

[0066] When the user terminal uses the above perturbation mechanism, d is mapped to a smaller value range through a hash function. If h is the hash function, then finally it will be perturbed into y with a probability of 1 - γ (where γ = ∑min Pr[GRR(h(v)) = y] = g / (g + e ε -1), g is the size of the hash function value range, and the perturbed data stream is denoted as

[0067] S3. The shuffling end receives the perturbed datasets of all clients, calculates the error of the perturbed datasets to dynamically adjust the sampling rate r, and shuffles all the perturbed datasets to obtain the re-arranged perturbed datasets;

[0068] When the shuffling end performs shuffling, it uses a function rand() that randomly generates numbers from 1 to n, and then starts from the last element y of the data vector n and exchanges each element with the element corresponding to the subscript generated by rand() one by one from the back to the front until all elements are processed. This can ensure that the probability of each element at each position is equal, that is, it satisfies the uniform distribution, and at the same time ensures that the time complexity is O(n), where n is the data scale, and the efficiency is relatively high.

[0069] In this embodiment, the clients locally perturb and shuffle the data, effectively masking the connection between the data and the individual, thus protecting personal privacy. At the same time, it adopts a distributed setting of the client and the publishing end, and additionally adds a shuffling end to eliminate the need for a trusted third party, reducing the risk of privacy leakage. In addition, the privacy amplification effect provided by the shuffling operation performed by the shuffling end minimizes the damage to data integrity and avoids the problem of error peaks as the user group expands. Therefore, shuffled differential privacy has significant advantages in the release of data stream histograms, providing new ideas and methods for solving the balance problem between privacy protection and data availability.

[0070] Furthermore, before cleaning the perturbed datasets, the shuffling end of this embodiment calculates the error of the perturbed datasets, compares this error value with the relative error value of the release structure at the previous moment retained in the system, and dynamically adjusts the adoption rate r in real time according to the comparison result. Specifically, the adjustment method is as follows:

[0071] When |re t - re t-1 | > Τ, the shuffling end increases the sampling rate at the next moment of sampling;

[0072] When |re t - re t-1 | < Τ, the shuffling end maintains or decreases the sampling rate at the next moment of sampling;

[0073] re t is the relative error of the perturbed datasets calculated at the current moment of the system, re t-1 is the relative error of the release result at the previous moment retained in the system, and T is the error threshold. The value of T is 0% to 5%, but does not include 0%, because compared with processing all data, there will inevitably be errors in sampling processing, so the error threshold can never be equal to 0%.

[0074] When the difference of relative errors is greater than the error threshold, increase the sampling rate. The increased amplitude is: add the difference of relative errors to the current sampling rate, that is, r 调整后 = r 调整前 + α, where α is the actual difference of relative errors.

[0075] When the difference of relative errors is equal to the error threshold, consider introducing the time of the query request response at the publishing end. When the response time is lower than the user tolerance, generally lower than 100 ms, that is, maintain the current sampling rate. When the user tolerance, for example, exceeds 100 ms, the current sampling rate r can be appropriately reduced 调整后 = r 调整前 - 0.05% to further reduce the processing time;

[0076] When the difference of relative errors is less than the error threshold, maintain the current sampling rate for sampling at the next moment.

[0077] It should be noted that in most cases, the data in the data stream is unstable. Specifically, the variance or other metrics of each data point at two consecutive moments relative to its mean may vary greatly. If the same sampling rate is maintained before and after the data stream fluctuates, it will cause the collected data to deviate from the approximate distribution it should follow, resulting in poor usability of the final histogram result. Therefore, the sampling rate can be dynamically adjusted according to the relative error situation of the previous and next times, thereby improving the usability and processing efficiency, so that under the massive data stream, the histogram can not only be published in real time and quickly, but also its publishing accuracy is guaranteed.

[0078] S4. The publishing end receives the perturbed data set rearranged by the shuffling end, performs unbiased statistical processing on it, and then draws and publishes a histogram based on the processing result.

[0079] In this embodiment, the perturbed data set received by the publishing end is denoted as Then, the following formula is used to perform an unbiased estimate of each feature frequency, that is:

[0080]

[0081] where m 1 represents the number of users who map the private value v to the first position in the hash value domain. g is the size of the hash function value domain, and n is the number of users.

[0082] S5. If the publishing end continuously maintains the request, loop through steps S2 to S4 until the publishing end terminates the request, and the histogram publishing process ends.

[0083] The present invention introduces the shuffle differential privacy technology, dynamically adjusts the sampling rate in real time by combining the relative error magnitudes at the previous and current moments, and during the data publishing process, while ensuring that the model is not limited by a trusted third party, it also avoids the problems of excessive result variations caused by GRR with the data value range and the linear growth of the publishing error with the number of users, thereby ensuring the stability and reliability of the published histogram under big data.

[0084] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A data stream fast histogram publishing method based on shuffle differential privacy, characterized by: The steps include: S1. The publisher issues a histogram query request, and the user responds; S2, the user terminal initializes and sets the sliding window, collects the elements in the sliding window whose status is not expired at a sampling rate r, and processes the data using a random response mechanism optimized by a hash function to generate a perturbation data set; S3, the shuffling end receives the perturbation data sets from all user ends, calculates the errors of the perturbation data sets to dynamically adjust the sampling rate r, and shuffles all the perturbation data sets to obtain the rearranged perturbation data sets; S4, the publishing end receives the perturbed data set rearranged by the shuffling end, performs unbiased statistical processing on it, and then draws a histogram based on the processing result and publishes it; S5. If the publisher continues to maintain the request, steps S2 to S4 are repeated until the publisher terminates the request, and the histogram publishing process ends.

2. The histogram publishing method according to claim 1, characterized in that: In step S2, the sliding window model is used to extract the element e in the data stream D t , the sliding window of time t is recorded as W t , W t ={e t-W+1 ,e t-W+2 ,...,e t }; W represents the size of the sliding window; The auxiliary space A in the sampling set is a set of triples Among them: ts i is the timestamp of the element, id i Mark the position in the auxiliary space, size i is the size of the data domain used for random response disturbance calculation; W+b is the size of the auxiliary space A; M is the size of the sampling set. For each element e in the data stream D, t , filter by the following requirements: If there are expired elements in A[1] to A[W+b], use e t Replace the oldest expired element; If there is no expired element in A[1] to A[W+b], e is ignored. t and wait for the next element; If a sampling set of size M needs to be selected from A, then M non-expired elements are randomly selected from A.

3. The histogram publishing method according to claim 2, characterized in that: In step S2, the sampling process iterates each sliding window in the data stream D and processes the selected elements to ensure that the sliding window W at time t is t and sample sets are kept up to date and relevant.

4. The histogram publishing method according to claim 1, characterized in that: In step S2, the data is perturbed using a random response mechanism based on hash function optimization, the original data is maintained unchanged with a perturbation probability γ, and the original data is mapped to other values ​​in the hash value domain with a probability of 1-γ, where γ = ∑minPr[GRR(h(v)) = y] = g / (g+e ε -1), v is the user-side private value, y is the output value after perturbation processing, Pr[GRR(v)=y] represents the probability that the input value is v and the output value is y; g is the value domain size of the hash function, and ε is the privacy budget.

5. The histogram publishing method according to claim 4, characterized in that: Privacy Budget n is a positive integer, δ∈(0,1), 6. The histogram publishing method according to claim 4, characterized in that: Wherein, d is the value range of the user-side private value v.

7. The histogram publishing method according to any one of claims 1 to 6, characterized in that: In step S3, the method for dynamically adjusting the sampling rate r is as follows: when|re t -re t-1 |>Τ, the shuffle end increases the sampling rate when sampling at the next moment; when|re t -re t-1 |<=Τ, the shuffle end maintains or reduces the sampling rate when sampling at the next moment; re t is the relative error of the disturbance data set calculated by the system at the current moment, re t-1 It is the relative error of the result released at the last moment stored in the system. T is the error threshold. The value of T is 0% to 5%, but does not include 0%.

8. The histogram publishing method according to claim 7, characterized in that: When the relative error difference is greater than the error threshold, the sampling rate is increased by the actual calculated relative error difference; when the relative error difference is equal to the error threshold, the time it takes to respond to the query request from the publisher is considered. If the response time is lower than the user tolerance, the current sampling rate is maintained; otherwise, the current sampling rate is reduced by 0.05%; when the relative error difference is less than the error threshold, the current sampling rate is maintained for sampling at the next moment.

9. The histogram publishing method according to any one of claims 1 to 6, characterized in that: The method for shuffling the perturbed data set at the shuffling end is as follows: Use the rand() function to start from the last element in the perturbed data set data vector, and swap each element with the element corresponding to the subscript generated by rand() one by one from back to front until all elements are processed.

10. The histogram publishing method according to any one of claims 4 to 6, characterized in that: In step S4, the unbiased statistical processing method is as follows: First, the publisher sends the perturbed data set based on the shuffled data set. To count the frequency of all user-side private values ​​v; Then, the unbiased estimate of the private value v is estimated using the following formula Among them, m1 represents the number of users who map the private value v to the first position of the hash value domain, g is the size of the hash function value domain, and n is the number of users.