5g communication network data dynamic privacy publishing method and system based on reinforcement learning
By employing sliding window data stream sampling and reinforcement learning grouping noise addition methods in 5G communication networks, the shortcomings of data stream differential privacy methods in dynamic environments are addressed, enabling fast and accurate data publishing and privacy protection, and improving data availability.
Patent Information
- Application Number
- CN202311078272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-08-24
AI Technical Summary
Existing methods for publishing differential privacy data in data streams are ill-suited to adapting to the dynamic data characteristics and network environment in 5G communication networks. This results in less attention to recent elements, low data availability, and an inability to intelligently adjust differential privacy parameters.
We employ a sliding window-based data stream sampling method and a reinforcement learning-based grouping and noise-adding method. By using a sliding window model, we can quickly obtain approximate statistical results and optimize the decision-making process using reinforcement learning. We can also automatically adjust differential privacy parameters to publish data that meets user needs.
It enables the rapid and accurate release of data that meets user needs, improves data availability, and reduces privacy budgets, making it suitable for the dynamic environment of 5G communication networks.
Smart Images

Figure CN117135622B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer and information security technology, specifically to a method and system for dynamic privacy publishing of 5G communication network data based on reinforcement learning. Background Technology
[0002] The rapid development and widespread application of 5G communication networks have brought about a demand for large-scale data transmission and processing. However, this data contains a large amount of sensitive information and personal privacy, making the protection of data security and privacy an urgent task. Traditional data publishing methods cannot provide sufficient privacy protection, thus requiring a new approach to ensure data privacy during transmission and processing.
[0003] Differential privacy, as an effective method for privacy protection, protects data sensitivity by introducing noise when publishing data. However, in 5G communication networks, the characteristics of data and the network environment change dynamically, and traditional differential privacy methods often struggle to adapt to these changes.
[0004] The background of datastream differential privacy reflects the urgent need for personal privacy protection and the limitations of traditional privacy protection methods in datastream environments. It provides a new method and framework for privacy protection applicable to real-time datastreams, offering a feasible privacy protection solution for datastream analysis and applications. With the increasing application of datastreams, research and application of datastream differential privacy will continue to receive attention, providing a better balance and integration between privacy protection and datastream analysis. Currently, there are many technical solutions for data publishing, whether for static datasets or dynamic datastreams. However, the methods proposed in existing technologies are not suitable for datastream sliding window models. Specifically, when publishing histograms in a sliding window, the following drawbacks exist:
[0005] (1) Existing data stream differential privacy data publishing methods generally pay less attention to recent elements and less attention to the leakage of current data statistical results;
[0006] (2) Existing data flow histogram methods use data flow statistics that simply perform direct statistics and add noise, resulting in low data usability;
[0007] (3) Existing algorithms are not intelligent enough and cannot automatically adjust differential privacy parameters by learning the decision-making process of data publishers and the dynamic change model of the environment. Summary of the Invention
[0008] 1. The technical problem that the invention aims to solve
[0009] The purpose of this invention is to provide a method and system for dynamic privacy publishing of 5G communication network data based on reinforcement learning. This method is based on a sampling algorithm of a sliding window model and a noise-adding algorithm based on reinforcement learning. When publishing differential privacy histograms for sliding window data streams, it can quickly and accurately publish data that meets user needs, improve data availability and reduce privacy budget.
[0010] 2. Technical Solution
[0011] To achieve the above objectives, the technical solution provided by this invention is as follows:
[0012] The present invention provides a method for dynamic privacy publishing of 5G communication network data based on reinforcement learning, the steps of which are as follows:
[0013] Step 1: Determine the intervals of the histogram of data to be published in the data stream;
[0014] Step 2: For all intervals of the histogram to be published, use a sliding window-based data stream sampling method to sample the data within the sliding window at the current time, and obtain the current sampling set M;
[0015] Step 3: Based on the sampling set at the current time, obtain the approximate statistical results of all intervals within the sliding window at the current time; add noise based on the data statistics within the sliding window to obtain the noise-added statistical results of all intervals within the sliding window at the current time; then, based on the obtained approximate statistical results and noise-added statistical results of all intervals within the sliding window at the current time, use an error minimization algorithm based on the comparison of approximation and noise to obtain better results.
[0016] Step 4: Based on the better interval statistics obtained in Step 3, a grouping and noise-adding method based on reinforcement learning is adopted, and the local optimal grouping results are obtained for noise addition to obtain the final noisy histogram data that can be published.
[0017] The present invention provides a dynamic privacy publishing system for 5G communication network data based on reinforcement learning, comprising:
[0018] The determination module is used to determine the intervals of the histogram of data to be published in the data stream;
[0019] The sampling module is used to sample all intervals of the histogram to be published. It uses a data flow sampling algorithm based on a sliding window model to sample the data within the sliding window at the current time and obtain the sample set at the current time.
[0020] The acquisition module is used to obtain the statistical results of all intervals within the sliding window at the current time, based on the sampling set at the current time.
[0021] Comparison module: Based on the approximate statistical results and the noisy statistical results of all intervals within the sliding window at the current time, an error minimization algorithm based on the comparison of approximation and noise is adopted to obtain better results;
[0022] The publishing module is used to publish the histogram of the sliding window at the current time step using a grouping and noise-adding method based on reinforcement learning, based on the statistical results obtained by the comparison module. Based on the above grouping results, a privacy budget is added to each group to obtain the final noisy histogram data that can be published.
[0023] 3. Beneficial effects
[0024] Compared with existing known technologies, the technical solution provided by this invention has the following significant advantages:
[0025] (1) This invention stores the attribute statistics of each element in the data stream into a sliding window-based data stream sampling algorithm by scanning the data stream once, and then generates and publishes a histogram based on the data collected by the algorithm. This can quickly obtain an approximate statistical count of the sliding window interval. Moreover, by using the sliding window model to obtain the latest elements in the big data environment, it also overcomes the shortcomings of existing data stream differential privacy data publishing methods, which pay less attention to the most recent elements and less attention to the leakage of current data statistical results.
[0026] (2) In order to solve the problem that the existing data flow histogram method only directly counts and adds noise, resulting in low data availability, this invention adopts a data flow sampling algorithm for technical statistics, which improves the time and space overhead of the data flow sliding window histogram publishing algorithm. In addition, this invention proposes an error minimization algorithm based on the comparison of approximation and noise. This method compares the approximate value with the noise value and obtains the approximate value or noise value with smaller error than the true value, which further improves the availability of the data.
[0027] (3) In order to solve the problem that existing algorithms are not smart enough and cannot automatically adjust differential privacy parameters by learning the decision-making process of data publishers and the dynamic change model of the environment, this invention proposes a grouping noise method based on reinforcement learning. Based on reinforcement learning, this method obtains the local optimal grouping of the current sliding window and publishes privacy-protected data that meets the user's needs, further reducing the privacy budget and improving the availability of data.
[0028] (4) Compared with existing technologies, this invention can quickly and accurately publish data that meets user needs, improve data availability, and reduce privacy budgets. This indicates that this invention has broad application prospects in 5G communication networks and can provide users with safer and more reliable communication services. Attached Figure Description
[0029] Figure 1 This is a flowchart of the reinforcement learning-based differential privacy data publishing method of the present invention.
[0030] Figure 2 This is a flowchart of the data flow processing. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art to which this invention pertains.
[0032] While existing technologies offer various methods for data streaming based on sliding window histograms, they lack a fast way to compute data within the sliding window. Furthermore, current methods require scanning each sliding window while constructing the histogram, leading to high runtime and storage overhead. Therefore, this invention provides a reinforcement learning-based dynamic privacy publishing method and system for 5G communication networks. Addressing the characteristics of large volumes, high speed, and real-time data streaming, and the need to focus on the latest elements, this invention proposes a sliding window-based data streaming sampling method that can quickly obtain approximate statistical results within the current window. To further reduce privacy budgets and improve data usability, this invention proposes a reinforcement learning-based grouping and noise-adding method. This method, based on reinforcement learning, obtains the locally optimal grouping within the current sliding window and publishes privacy-protected data that meets user requirements. Reinforcement learning is a machine learning technique that enables agents to automatically learn and optimize decision-making processes through interaction and learning with the environment. Applying reinforcement learning-based methods to privacy publishing in 5G communication networks allows for automatic adjustment of differential privacy parameters by learning the decision-making process of the data publisher and the dynamic changes in the environment, thereby maximizing data usability and utility while protecting data privacy.
[0033] The present invention will be further described in detail below with reference to specific embodiments.
[0034] like Figure 1 As shown in this embodiment, the method for dynamic privacy publishing of 5G communication network data based on reinforcement learning includes the following steps:
[0035] Step 1, determine the intervals of the data to-be-published histogram in the data stream; and divide the total privacy budget into two parts (ε = ε1 + ε2). The first part of the privacy budget is used for data stream sampling based on a sliding window, and the second part is used for grouped noise addition based on reinforcement learning.
[0036] Step 2, for all intervals of the to-be-published histogram, use the data stream sampling method based on a sliding window to sample the data within the sliding window at the current moment, and obtain the sampling set M at the current moment;
[0037] Define a data stream D, D = {e1, e2,..., e t ,...} (t ≥ 0), the size of the sliding window is w, the current element at the current moment t is e t , the sampling set is M, and the size of the sampling set is m; the process of obtaining the sampling set M of the current sliding window at any moment by the data stream sampling method based on a sliding window is as follows:
[0038] For the current element e t ∈D at any moment, do the following: for the current element e t at the current moment t ≤ m, directly put the current element e t into the sampling set M; for the current element e t at the current moment m < t ≤ w, insert the current element e t into the sampling set M with a probability of m / t, because the range of the generated random number is [1, t], and it will be replaced into the sampling set M only when the number is less than or equal to m; for the current element e t at the current moment t > w, judge whether the oldest element in the current sampling set M has expired; if the oldest element in the sampling set has expired, delete the oldest element in the current sampling set M and insert the current element e t ; if the oldest element in the sampling set has not expired, insert the current element e t with a random probability of m / w, because the range of the generated random number is [1, w] (because the size of the sliding window set is w), and it will be replaced into the sampling set M only when the number is less than or equal to the size m of the sampling set.
[0039] Step 3, according to the sampling set at the current moment, obtain the approximate statistical results of all intervals within the sliding window at the current moment; add the noise randomly generated by the Laplace mechanism using the first part of the privacy budget according to the data statistical values within the sliding window, and obtain the noise-added statistical results of all intervals within the sliding window at the current moment; then, according to the approximate statistical results and the noise-added statistical results of all intervals within the sliding window obtained, use the error minimization algorithm based on approximate and noise comparison to obtain a better result; specifically:
[0040] Obtain the sampling results for all intervals of the histogram to be published from the sampling set M. This sampling result is denoted as H. M According to the sampling results H M Calculate the approximate statistical results for all intervals within the sliding window at the current time, and denote this approximate statistical result as H. W ,
[0041] Based on the privacy budget in Part 1, noise is added to the currently obtained true values (the true values of the intervals within the sliding window), and the noise-added statistics for all intervals within the sliding window at the current time are obtained. These noise-added statistics are denoted as...
[0042] According to the noise addition statistics and approximate statistical results H W The comparison process is as follows: For each interval, obtain the noisy statistical result or approximate statistical result with the smaller error between |Noisy Statistical Result - True Value| and |Approximate Statistical Result - True Value|, where the true value is the true value of the interval within the current window. Finally, obtain the statistical result with the smaller error within each sliding window.
[0043] Step 4: Based on the improved interval statistical results obtained above, a reinforcement learning-based grouping and noise-adding method is used, and the locally optimal grouping results are used for noise addition to obtain the final noisy histogram data that can be published; specifically:
[0044] First, initialization is required:
[0045] (1) Define the state space: The state consists of the current grouping status and the data to be processed;
[0046] (2) Define the action space: merge data from two different intervals;
[0047] (3) Initialize the Q table: used to store the Q values of state-action pairs, and initialize it to 0;
[0048] (4) Initialize hyperparameters: learning rate α, discount factor γ, exploration rate A.
[0049] Sort the statistical data (the better interval statistical results obtained in step 3) from smallest to largest according to the sorting method;
[0050] Repeat the following steps until all the data has been processed:
[0051] a. Observe the current state: Treat the previous group as the merged group, and the current group and the next data as data to be processed.
[0052] b. Select an action based on the current state and the Q table:
[0053] The A-greedy strategy selects actions based on the following: A random action is selected with probability A, and the action with the largest Q-value is selected with probability 1-A. If a random action is selected, the action to be merged or not merged is randomly chosen. If the action with the largest Q-value is selected, the action to be merged or not merged is chosen according to the Q-table.
[0054] c. Execute the action and observe the reward and next state: Based on the selected action, merge or not merge the two data points, and calculate the error after merging and the error without merging. (Currently, the error from merging the two data points will receive a reward score, and the error from not merging the data will also receive a reward score. This is because different rewards are given based on the different errors. The calculation of the merged error consists of two aspects: the reconstruction error and the Laplace error. For example, if there are two groups X and Y, and X and Y are merged, the total error is the result of the reconstruction error of X + Y and the privacy budget error generated by X + Y. The error without merging is the reconstruction error of X + the Laplace error generated by X + the reconstruction error of Y (generally, since there is only one data point, the reconstruction error is 0) + the Laplace error generated by Y.)
[0055] d. Update Q-values: Update the Q-values in the Q-table using Q-learning: Based on the next state and the Q-table, select the action with the maximum Q-value. Update the Q-values of the current state and action according to the Q-learning update rules:
[0056] Q(s,a)=(1-α)*Q(s,a)+α*(reward+γ*maxQ)
[0057] e. Proceed to the next step: Use the merged group as the preceding group for the next step, and perform a merge or non-merge operation between the current group and the next data. Based on the Q-table obtained from training, perform the final grouping of the data.
[0058] Based on the grouping results above, and by adding a privacy budget to each group according to the privacy budget in Part 2, we obtain the final noisy histogram data that can be published.
[0059] In this invention, the two errors are formed from two aspects: the sum of noise errors generated by privacy budget ε1 and privacy budget ε2 (firstly, for data stream sampling algorithms, the sliding window-based data stream sampling algorithm is equivalent to Laplace noise of privacy budget ε1, and the reinforcement learning-based grouping noise algorithm is equivalent to Laplace noise of privacy budget ε2). Privacy budget ε1 prevents statistical data from leaking privacy during the noise addition process, while privacy budget ε2 prevents real data from leaking privacy.
[0060] The core of this invention consists of two algorithms: a sliding window-based data stream sampling algorithm and a reinforcement learning-based grouping noise addition algorithm. Both algorithms satisfy the differential privacy condition. The sliding window-based data stream sampling algorithm is equivalent to Laplace noise with a privacy budget of ε1, and the reinforcement learning-based grouping noise addition algorithm is equivalent to Laplace noise with a privacy budget of ε2. Based on the parallel combination property of differential privacy, this invention satisfies ε-differential privacy (ε = ε1 + ε2), therefore, this invention satisfies the differential privacy condition.
[0061] The time and space complexity of this invention will be further analyzed below, including:
[0062] Given that the size of the sliding window sampling set is m and the size of the statistical interval is l, the time cost of this invention is O(m). The proof is as follows: For any sliding window, the processing of this invention mainly consists of: (1) data stream sampling based on the sliding window; (2) grouping and adding noise based on reinforcement learning. Among these, the main time cost lies in the processing time required to obtain the sampling set through the sliding window sampling algorithm and the time required to sort the counts in the current interval. Therefore, the time complexity of this invention is O(m+l), where the time required for the sliding window sampling algorithm to obtain a sampling set is the size of the sampling set, which is m, and the time required to sort the counts in the interval is l. Existing sliding window publishing algorithms need to cache the current sliding window data, so the time required by these algorithms is the size of the sliding window, which is O(w). Here, w is greater than the size of the sampling set. It is obvious that the time cost required by this invention is better than that of existing methods.
[0063] The space overhead of this invention is log₂l*m; the proof is as follows: For the current window sampling set M, the space required for data stream sampling based on the sliding window is mainly the space overhead of the sampling algorithm set M. Therefore, the space overhead required by the data stream sampling algorithm based on the sliding window is mainly the size of the space overhead of the sampling algorithm set M. Thus, the space overhead of the reinforcement learning-based differential privacy data publishing method is log₂l*m, where l is the number of intervals in the histogram to be published. The time required for the sliding window sampling algorithm to obtain a sampling set is the size of the sampling set, which is m. The time required to sort the counts in the intervals is l. Existing sliding window publishing algorithms need to cache the current sliding window data, so the time required by these algorithms is the sliding window size, i.e., O(w). Here, w is greater than the size of the sampling set. It is clear that the time overhead required by this invention is superior to existing methods.
[0064] Example 2
[0065] In this embodiment, a data stream sampling and publishing system with differential privacy is provided, the system comprising:
[0066] The determination module is used to determine the intervals of the histogram of data to be published in the data stream;
[0067] The sampling module is used to sample all intervals of the histogram to be published. It uses a data flow sampling algorithm based on a sliding window model to sample the data within the sliding window at the current time and obtain the sample set at the current time.
[0068] The acquisition module is used to obtain the statistical results of all intervals within the sliding window at the current time, based on the sampling set at the current time.
[0069] Comparison module: Based on the approximate statistical results and the noisy statistical results of all intervals within the sliding window at the current time, an error minimization algorithm based on the comparison of approximation and noise is adopted to obtain better results;
[0070] The publishing module is used to publish the histogram of the sliding window at the current time step using a grouping and noise-adding method based on reinforcement learning, based on the statistical results obtained by the comparison module. Based on the above grouping results, and according to the privacy budget in the second part, a privacy budget is added to each group to obtain the final noisy histogram data that can be published.
[0071] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A method for dynamic privacy-preserving data publishing in 5G communication networks based on reinforcement learning, characterized in that, The steps are as follows: Step 1: Determine the intervals of the histogram of data to be published in the data stream; Step 2: For all intervals of the histogram to be published, use a sliding window-based data stream sampling method to sample the data within the sliding window at the current time, and obtain the sample set at the current time. M ; Step 3: Based on the current sampling set, obtain approximate statistical results for all intervals within the sliding window at the current time; add noise to the statistical values within the sliding window to obtain noisy statistical results for all intervals within the sliding window at the current time; then, based on the obtained approximate and noisy statistical results for all intervals within the sliding window at the current time, use an error minimization algorithm based on the comparison between approximation and noise to obtain better results; specifically: From the sample set The sampling results of all intervals of the histogram to be published are obtained from the sample. These sampling results are denoted as... According to the sampling results Calculate the approximate statistical results for all intervals within the sliding window at the current time, and denote this approximate statistical result as... , ;in, w The size of the sliding window. Size of the sample set; Based on the privacy budget in Part 1, noise is added to the currently obtained true values, and the noise-added statistics for all intervals within the sliding window at the current time are obtained. These noise-added statistics are denoted as... ; According to the noise addition statistics and approximate statistical results Compare and obtain approximate statistical results or noisy statistical results with smaller errors from the true values in each interval, and obtain better statistical results within each sliding window; Step 4: Based on the improved interval statistical results obtained in Step 3, a reinforcement learning-based grouping and noise-adding method is used, and the locally optimal grouping results are used for noise addition to obtain the final publishable noisy histogram data; where, The process of obtaining the locally optimal grouping result is as follows: Sort the interval statistics obtained in step 3 from smallest to largest according to the sequential sorting method, and repeat the following steps until all the data has been processed: a. Observe the current state: Treat the previous group as the merged group, and the current group and the next data as data to be processed; b. Select an action based on the current state and the Q table: The A-greedy strategy selects actions as follows: a random action is selected with probability A, and the action with the maximum Q value is selected with probability 1-A; if a random action is selected, the action to merge or not merge is randomly selected; if the action with the maximum Q value is selected, the action to merge or not merge is selected according to the Q table. c. Perform the action and observe the reward and the next state: Based on the selected action, merge or not merge the two data points, and calculate the error after merging and the error without merging; d. Update Q-values: Update the Q-values in the Q-table using Q-learning: Select the action with the maximum Q-value based on the next state and the Q-table; Update the Q-values of the current state and action according to the Q-learning update rules; e. Proceed to the next step: Use the merged group as the previous group for the next step, and merge or not merge the current group with the next data; perform the final grouping of the data based on the Q-table obtained from training; Based on the grouping results, and by adding a privacy budget to each group according to the privacy budget in Part 2, the final noisy histogram data that can be published is obtained.
2. The method for dynamic privacy publishing of 5G communication network data based on reinforcement learning according to claim 1, characterized in that: The total privacy budget for the entire data release is divided into two parts: the first part is used for data stream sampling based on a sliding window, and the second part is used for grouping and adding noise based on reinforcement learning.
3. The method for dynamic privacy publishing of 5G communication network data based on reinforcement learning according to claim 2, characterized in that: Step 2 describes the use of a sliding window-based data stream sampling method to sample the data within the sliding window at the current moment. The specific process is as follows: Define data stream as D , If t ≥ 0, then the current element at time t is... For the current element at any given time... Perform the following processing: For the current time t≤ The current element , will the current element Directly put into the sampling set For the current moment The current element , will the current element by Probability of inserting sampling set For the current time t> w The current element Determine the current sampling set Check if the oldest element in the sample set has expired; if the oldest element in the sample set has expired, then in the current sample set... Remove the oldest element from the collection and insert the current element. If the oldest element in the sample set has not expired, then a random selection is used. m / w The probability of inserting the current element .
4. A 5G communication network data dynamic privacy publishing system based on reinforcement learning, characterized in that: Implementing the privacy publishing method according to any one of claims 1-3 includes: The determination module is used to determine the intervals of the histogram of data to be published in the data stream; The sampling module is used to sample all intervals of the histogram to be published. It uses a data flow sampling algorithm based on a sliding window model to sample the data within the sliding window at the current time and obtain the sample set at the current time. The acquisition module is used to obtain the statistical results of all intervals within the sliding window at the current time, based on the sampling set at the current time. Comparison module: Based on the approximate statistical results and the noisy statistical results of all intervals within the sliding window at the current time, an error minimization algorithm based on the comparison of approximation and noise is adopted to obtain better results; The publishing module is used to publish the histogram of the sliding window at the current time step using a grouping and noise-adding method based on reinforcement learning, based on the statistical results obtained by the comparison module. Based on the above grouping results, a privacy budget is added to each group to obtain the final noisy histogram data that can be published.
Citation Information
Patent Citations
Streaming histogram publishing method and system of weighted sliding window under differential privacy
CN114969656A
Online streaming sampling release method and system with differential privacy
CN115114584A