Track privacy protection method based on double-delay depth deterministic strategy gradient
Through a trajectory privacy protection method based on a dual-latency depth deterministic policy gradient, combined with deep convolutional neural network and dynamic privacy budget, the trajectory perturbation strategy is dynamically optimized, and the problems of insufficient protection of semantic sensitive information and time correlation and insufficient adaptability to the dynamic environment in the existing technology are solved, thereby achieving efficient trajectory privacy protection and balance of service quality.
Patent Information
- Application Number
- CN202510330924.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
AI Technical Summary
Existing trajectory privacy protection technologies are difficult to effectively protect semantic sensitive information and temporal correlations in user trajectories, and the static location privacy protection mechanism lacks robustness to dynamic environments.
The trajectory privacy protection method based on the dual-latency depth deterministic policy gradient is adopted, and the environment state information is captured through deep convolutional neural networks, combined with dynamic privacy budget and time correlation, dynamically optimized the trajectory perturbation strategy to protect user trajectory privacy in real time.
It significantly enhances the privacy and security performance of track data, effectively reduces the risk of exposure of sensitive location information, balances privacy protection and service quality, and adapts to different environments and privacy needs.
Smart Images

Figure CN120166360A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a trajectory privacy protection method based on double-delay deep deterministic policy gradient, belonging to the technical fields of location-based services and information security. Background Art
[0002] The rapid development of positioning technology and communication technology has promoted the wide application of location-based services (LBS) in mobile networks and trajectory scenarios. While LBS provides service convenience, the privacy of users also faces many threats ([1] S.A. Thompson and C. Warzel, Twelve Million Phones, One Dataset, Zero Privacy, New York Times, 2019-12-19). Attackers can infer users' sensitive information (such as living habits, health conditions or financial behaviors) by analyzing the temporal correlation and semantic tags in user trajectory data.
[0003] In recent years, the problem of trajectory privacy leakage in location-based services (LBS) has attracted the attention of scholars at home and abroad, but most of the trajectory privacy protection research ignores the protection of users' semantic sensitive information and temporal correlation. However, the exposure of semantic sensitive information and temporal correlation will significantly increase the risk of user privacy leakage. In the prior art, when a user frequently visits highly sensitive locations (such as hospitals or banks), it may expose their health conditions or financial activities, and attackers can more accurately infer privacy information such as users' living patterns and working hours by analyzing the temporal correlation in user trajectories. Currently, most trajectory privacy protection schemes such as ([2] C. Wang and Y. Zhao, A Novel Approach for Trajectory Partition Privacy in Location-Based Services[C], International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), 1464-1469) mainly focus on perturbations and protection at the trajectory level, lack in-depth consideration of the temporal correlation between specific locations in user trajectories, and fail to provide key protection for semantic sensitive location information.
[0004] Meanwhile, most current trajectory privacy protection mechanisms adopt a fixed privacy budget and it is difficult to dynamically adjust the privacy strength according to the actual characteristics of user trajectories. The paper ([3]M.Cao et al., Protecting Personalized Trajectory with Differential Privacy under Temporal Correlations[C], IEEE Wireless Communications and Networking Conference (WCNC), 2024, 1-6) studied a user-implemented trajectory privacy protection mechanism. This mechanism takes into account the temporal correlations between locations in the trajectory and generates pseudo-locations by perturbing the real locations in the trajectory, effectively protecting the privacy of user trajectories. Therefore, we further study the semantically sensitive locations in the trajectory based on deep reinforcement learning, and at the same time use a dynamic privacy budget combined with temporal correlations to provide key protection for semantically sensitive location information.
[0005] Furthermore, in the actual scenario, user trajectories are usually in a dynamic and ever-changing complex environment. Attackers may launch attacks on user privacy using inference attacks and various means, making it difficult for users to timely and comprehensively obtain environmental state information when optimizing and selecting location privacy protection strategies. However, most existing static location privacy protection mechanisms rely on given system models and attack models and lack robustness to dynamic environments. The exploration of dangerous strategies within a short period of time will lead to irreversible privacy leakage losses for users. The location privacy protection technology based on deep reinforcement learning (RL) can break through the dependence on information such as system models and attack models of existing mechanisms, realize the dynamic optimization of location privacy protection strategies for mobile users, and open up a new perspective for the research of trajectory privacy protection. ([4]D.Shi, J.Ding, S.M.Errapotu, H.Yue, W.Xu, X.Zhou, and M.Pan, Deep Q-Network-based route scheduling for TNC vehicles with passengers’ location differential privacy[J]. IEEE Internet of Things Journal, 2019, 6(5): 7681-7692.) Summary of the Invention
[0006] Aiming at the deficiencies of the existing technologies, a trajectory privacy protection method based on twin-delayed deep deterministic policy gradient is proposed. This method has simple steps and remarkable effects. It can perturb the user's trajectory position points in real time, effectively capture the environmental state information during location services with a deep convolutional neural network, fit the dynamic location service environment and attack patterns, and enhance the privacy security performance of the trajectory data in the location service system under inference attacks.
[0007] To achieve the above technical objectives, the disclosed trajectory privacy protection method based on twin-delayed deep deterministic policy gradient constructs an LBS system based on an LBS server and a user, including a two-dimensional spatial map marked with sensitivity levels; takes the user's current trajectory point position information and the sensitivity of the location where the user is located as the user state; sends the perturbed false position information to the LBS server using a perturbation strategy based on twin-delayed deep deterministic policy gradient for the current user state; the user relies on the service feedback of the LBS server to evaluate the degree of privacy protection and the loss of service quality, calculates the benefit of the LBS system of the user in the current time slot, and constructs the network state of the LBS system in the next time slot through a Markov transition process; based on the user's needs and the sensitivity of the current location, adopts a dynamic perturbation strategy to protect the user's trajectory privacy; uses transfer learning to recycle the determined perturbed position experiences, stores the network states of the LBS system in the current time slot and the next time slot, the dynamic perturbation strategy, and the evaluation results in an experience pool as experiences, and continuously optimizes the trajectory protection strategy by updating the weight parameters of the actor network and the critic network in the perturbation strategy; combines the user's real position, the degree of privacy protection, and the loss of service quality to obtain the benefit of the LBS server network system, and the user continuously updates the weight parameters of the actor network and the critic network until a stable trajectory perturbation strategy is obtained, finally achieving real-time trajectory privacy protection of the user in the LBS system.
[0008] The specific steps are as follows:
[0009] S1. Initialize the two-dimensional spatial map for the user to access the LBS system and the perturbation strategy parameters of the twin-delayed deep deterministic policy gradient network TD3;
[0010] S2. Mark the geographical location and semantic location currently accessed by the user in the two-dimensional space, and view the sensitivity level of the location where the user is located according to the user's own real position; form the network state of the LBS system where the current user is located, the possible location set of the information, and the coordinates of the geographical location based on the information obtained in the previous step;
[0011] S3. At time k, the user inputs its current state into the actor and critic networks of TD3, and the actor network outputs the action selected for the current state;
[0012] S4. Calculate the set of possible positions of the user at time k;
[0013] S5. Based on the lower bound of the expected inference error and the privacy budget obtained by using the perturbation strategy, determine the protection location set. Use the Hilbert search algorithm to search for the personalized location protection set, and then input the content of the personalized location protection set after search matching into the location perturbation mechanism Permute-and-Flip. Through the location perturbation mechanism Permute-and-Flip, perform the work of publishing the perturbed location to the LBS server;
[0014] S6. After the LBS server receives the user's location information, give service feedback to the user; the user judges the degree of privacy protection according to the feedback result, estimates the service quality loss and the network benefit of the LBS server at the user's place, and constructs the network state of the LBS system in the next time slot;
[0015] S7. Using the user's location transition probability matrix and the possible location set of the user at time t, use the posterior probability of the user at time t to infer the prior probability of the user at the next time k + 1, so as to infer multiple possible locations of the user at the next time;
[0016] S8. Substitute the prior probability at time k + 1 into the location point set to obtain the possible location set at time k + 1;
[0017] S9. Based on the experience replay technology, store the trajectory perturbation historical experience in the experience pool, and then select an experience from the experience pool. Based on the Critic1 network and the Critic2 network in TD3, calculate the target value;
[0018] S10. Use the gradient descent algorithm to minimize the error between the evaluation value and the target value, so as to update the parameters in the Critic1 network and the Critic2 network;
[0019] S11. According to the user's environment and privacy leakage situation, update the Critic1 network and the Critic2 network, and update the Actor network parameters through the deterministic policy gradient. Repeat this step for update until a stable trajectory perturbation strategy is learned.
[0020] Furthermore, when the LBS system starts to be used, it is first initialized. Simplify the coverage area of the LBS system into a two-dimensional space map and grid it, and perform semantic annotation and sensitivity level setting on the location functions of each grid; use the deep reinforcement learning algorithm to input the user's current location sensitivity level and the size of the possible location set as the network state parameters of the LBS system into the convolutional neural network to obtain the current trajectory location point perturbation strategy:
[0021] 3.1. Calibrate the geographical location d currently visited by the user on the two-dimensional space map based on the grid (k) and the semantic location c(k) , using the currently accessed geographical location d (k) to represent the real trajectory position point of the user, and using the semantic location c (k) is the location name indicating the current position point of the user, and is used to obtain the location sensitivity level of the user; the geographical location d (k) is a two-dimensional coordinate, including the horizontal and vertical coordinates x (k) , y (k) , k is the k-th time slot; view the sensitivity level l of the grid where the user is located according to the user's own real position (k) , according to the location sensitivity level l (k) provide dynamic privacy protection;
[0022] 3.2. According to the geographical location d accessed by the user currently (k) form the LBS system network state where the current user is located: s (k) = {x (k) , y (k) , l (k) , n (k)}, where n (k) represents the size of the possible location set, x (k) , y (k) respectively correspond to the horizontal and vertical coordinates of the geographical location d (k) , and the sensitivity level l corresponding to the semantic location (k) , and the possible location set Δχ of the user k ;
[0023] 3.3. At time k, the user inputs its LBS system network state into the actor network and critic network Actor-Critic in the twin-delayed deep deterministic policy gradient algorithm TD3, uses the Ciritc network to evaluate the actions of the Actor network and assist the Actor network to optimize the output actions, and the output of the Actor network is the action a selected according to the current state of the user (k) :
[0024] a (k) = π(s (k) ) + ε,
[0025] where, the Actor network structure is based on a fully connected neural network, ε is exploration noise, independently sampled from a truncated normal distribution, and the action a (k) includes the privacy budget of the current position and the lower bound of the expected inference error.
[0026] Furthermore, the perturbation strategy of the double-delay deep deterministic policy gradient adopted by the user is specifically as follows: The user first initializes a grid-based two-dimensional space map and the TD3 algorithm, determines the sensitivity of the user's current trajectory position point based on the two-dimensional space map information, and evaluates the set of possible positions Δχ of the current user trajectory position point k As the data input into the Actor network and the Critic network processed by the deep convolutional neural network CNN embedded in the Actor-Critic network for the network state of the LBS system, the CNN is used to extract features from the data, and the extracted features are used for subsequent calculations of the Actor network and the Critic network.
[0027] Furthermore, the user determines the perturbation position of the current trajectory point based on the perturbation strategy of the current time slot:
[0028] 5.1. The user determines the perturbation position of the current trajectory point using the privacy budget of the current time slot and the lower bound of the expected inference error: First, calculate the set of possible positions Δχ of the user trajectory position point according to the prior at the current moment k , and then determine the protected position set Φ of the user's current position using the privacy budget and the lower bound of the expected inference error k ;
[0029] Calculate the set of possible positions Δχ of the user at time k using the following formula k , and the set of possible positions Δχ k is obtained by screening through the prior probability and contains the true position of the user:
[0030]
[0031] where Δχ k represents the set of the smallest number of positions where the sum of the prior probabilities is less than 1 - δ, a i is a possible position of the user, δ is a constant (0 < δ < 1), represents the prior distribution at the current moment, and i represents the probability of the i-th region;
[0032] 5.2. Use the action output by the actor network as the lower bound of the expected inference error obtained by the strategy, and determine the personalized position protection set Φ using the privacy budget k , and the distance between any two positions in the personalized position protection set Φ k is less than or equal to D(Φ k ), that is and satisfies the following expression:
[0033]
[0034] In the formula, e € represents the privacy budget power of the exponential function e;
[0035] 5.3. The user searches for the personalized location protection set Φ through the Hilbert search algorithm k , and then inputs the content of the personalized location protection set Φ k matched by the search into the location perturbation mechanism Permute-and-Flip, and uses the location perturbation mechanism Permute-and-Flip to publish the perturbed location of the user to the LBS server. The perturbed location probability distribution matrix is as follows:
[0036]
[0037] where is used for normalization processing, ∈ is the privacy budget, is the set of all possible locations of the user.
[0038] Furthermore, the user transmits the perturbed location at the current moment to the LBS server. Based on the received user location information, the LBS server gives service feedback to the user; the user judges the privacy protection degree p (k) , and estimates the service quality loss q (k) . The privacy protection degree p (k) is determined by the difference between the sensitivity of the current location and the location inferred by the attacker. The service quality loss q (k) is the evaluation conclusion of the distance between the user's true location and the perturbed location. The farther the two distances are, the greater the service quality loss;
[0039] First, calculate the privacy protection degree
[0040] Then, obtain the service quality loss
[0041] In the formula, l (k) and respectively represent the sensitivity of the user's actual location and the sensitivity of the location inferred by the attacker. (x k , y k ) represents the user's original location, represents the perturbed location published by the user terminal;
[0042] Next, evaluate the benefit u of the LBS server network system where the user is located according to the following formula (k) , and construct the user state at the next moment. The specific expression of the benefit u of the LBS server network system (k) is as follows:
[0043] u (k) = p (k) - ωq (k)
[0044] In the formula, the weight parameter ω represents the importance degree of the quality - of - service loss relative to the trajectory privacy protection; η is a non - zero constant.
[0045] Furthermore, the time correlation between different positions in the user's trajectory at time k is utilized to obtain the user's true position at the next time k + 1 and the set of possible positions Δχ. k+1 Build the network state of the LBS system at the next moment:
[0046] 7.1: Construct the perturbed position transition probability matrix M using the user's historical trajectory data and behavior habits. The position transition probability matrix M is a two - dimensional n*n matrix, which represents the possibility of each position transitioning from the current position to the next corresponding position. Let m ij be the element in the i - th row and j - th column of the position transition probability matrix M, then m ij represents the probability that the user moves from area i to area j;
[0047] 7.2: According to the position information contained in the set of possible positions Δχ k at time t of the user, find the corresponding row vector m i in the position transition probability matrix M. m i represents the probability that the user moves from area i to other areas. These row vectors form the position transition probability matrix M k at time k. Through the transition probability matrix and combined with the Markov transition model, the user's true position at the next moment is obtained; the prior probability of the user at the next time k + 1 is inferred using the posterior probability of the user at time k Thus, multiple possible positions where the user may be at the next moment are inferred;
[0048] The posterior probability distribution of the user at time k satisfies the following formula:
[0049]
[0050] where d (k) and are the true position and the perturbed position of the user at time k respectively, is the position perturbation probability distribution. Then the prior probability distribution at time k + 1
[0051] 7.3: Considering the time correlation, substitute the prior probability at time k + 1 into the calculation formula of the set of position points Δχ k at time k to obtain the set of possible positions Δχ k+1 at time k + 1.
[0052] Further, the network status, dynamic perturbation strategy, and evaluation results of the current time slot and the next time slot are stored as experience in the experience pool. The specific steps to continuously optimize the trajectory protection strategy by updating the weight parameters of the actor network and the critic network are as follows:
[0053] 8.1: The historical experience of trajectory perturbation in time slot k is defined as e (k) , e (k) =(s (k) , a (k) , u (k) , s (k+1) ; Based on the experience replay technique, the historical experience of trajectory perturbation is stored in the experience pool Randomly select an experience e (k) from the memory pool, where 1 ≤ x ≤ k. Use the target Actor network to calculate the action a (k+1) in the state s (k+1) , and add noise to the target action to obtain a (k+1)′ . Then, based on the target Critic1 network and Critic2 network of the TD3 algorithm, calculate the target value using the following formula:
[0054]
[0055] 8.2: Input the state of the current time slot and the action a (k) output by the Actor network into two independent Critic1 networks and Critic2 networks respectively. The Critic network will evaluate the Q value of the action a (k) corresponding to the state in the current time slot. The larger the value of the action evaluation Q value, the higher the value of the action a (k) in the current state;
[0056] After obtaining the corresponding evaluation Q value, apply the gradient descent algorithm to minimize the error L (k) between the evaluation Q value and the target value y c , and update the parameters in the Critic1 network and the Critic2 network. The specific calculation is as follows:
[0057]
[0058] Among them, the structures of the Critic1 network and the Critic2 network are the same, and the two are set independently and in parallel;
[0059] 8.3: After updating the two Critic1 networks and Critic2 networks H times, update the Actor network parameters through the twin-delayed deep deterministic policy gradient as follows:
[0060]
[0061] 8.4: Finally, introduce the learning rate τ, and use the following formula to update the target network θ′ in the TD3 algorithm i , and the φ′ target network provides a stable target value for the system state at time k + 1. The specific update process of the target network is as follows:
[0062] θ′ i ← τθ i +(1 - τ)θ′ i (i = 1, 2),
[0063] φ′ ← τφ+(1 - τ)φ′.
[0064] According to the environment where the user is located and the privacy leakage situation, the mobile user repeats the above steps until the TD3 network learns a stable trajectory perturbation strategy.
[0065] A computer device includes a processor and a memory. The processor is electrically connected to the memory. The memory is used to store instructions and data, and the processor is used to execute the trajectory privacy protection method based on double delayed deep deterministic policy gradient.
[0066] Beneficial effects: The present invention innovatively introduces deep reinforcement learning technology, endowing the trajectory position perturbation strategy with powerful learning and adaptive capabilities, effectively compensating for the deficiencies of traditional research in terms of poor environmental adaptability and insufficient learning ability when dealing with complex and changing environments. The present invention highly values the sensitivity differences of semantic positions in the user's trajectory. By cleverly migrating the perturbation positions to adjacent semantic positions with lower sensitivity, while ensuring the privacy protection level, it takes into account the service quality that the user can obtain, greatly reducing the possibility of high-sensitivity semantic positions being exposed to risks.
[0067] The present invention adopts a unique deep reinforcement learning algorithm that does not rely on a specific model. It directly learns decision-making strategies through a convolutional neural network CNN, a TD3 network, and interaction with the environment, adapts to different environmental changes, and can handle model uncertainty and complexity, thus avoiding detailed modeling of the environment. It specifically solves the key problems commonly existing in current numerous location privacy protection studies, such as over-reliance on established environmental models and attack models, resulting in a lack of autonomous learning ability and effective adaptability to complex environments. In actual application scenarios, when it is difficult to accurately obtain detailed information about system environmental parameters and attack models, the convolutional neural network CNN used in the present invention accurately extracts environmental features, and uses the TD3 network to continuously interact with the environment and dynamically try and error to fully explore and determine the most optimized trajectory position privacy protection strategy.
[0068] The present invention ingeniously applies transfer learning technology, makes full use of the practical experience of position perturbation accumulated in the two-dimensional space under similar scenarios, and uses this to initialize the key parameters of the TD3 network. In the preliminary preparation stage of the algorithm, the position perturbation mechanism first conducts simulation experiments in N similar attack scenarios, performs operations for K time slots in each similar scenario, and generates the initial parameters of the neural network by means of the results output by the simulation experiments. Such a design enables users to effectively learn various unique environmental features in an environment facing different position inference attacks, and successfully transfer and apply these valuable learning results to the real two-dimensional space position privacy protection scenario. This process significantly improves the generalization ability of the learning algorithm, effectively reduces a large number of random exploration behaviors that occur in the initial stage of dynamic interaction, thus greatly accelerating the learning convergence speed, and also significantly improving the adaptability and operation efficiency of the algorithm in diverse complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a schematic diagram of the trajectory privacy protection method model based on double delayed deep deterministic policy gradient in the present invention.
[0070] Figure 2 It is the trajectory privacy protection method based on double delayed deep deterministic policy gradient in the present invention
[0071] Process schematic diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The following further describes the implementation of the present invention with reference to the drawings:
[0073] As Figure 1 and Figure 2 shown, a trajectory privacy protection method based on double delayed deep deterministic policy gradient (TD3) of the present invention first studies a personalized semantic trajectory privacy protection framework based on differential privacy and distortion privacy, models the trajectory privacy protection problem as a Markov decision process (MDP), and at the same time considers the situation where the attacker uses temporal correlation for inference. For each point on the trajectory, a personalized position protection set considering temporal correlation is established to protect the user's trajectory privacy. A position transition probability matrix is constructed with the user's historical trajectory data and behavior habits, which is used to dynamically obtain the possible positions of the user at the next moment, realizing a dynamic protection mechanism that changes with time and strictly ensuring the user's trajectory privacy. The user sends the perturbed trajectory position point information to the LBS server, and only the user himself knows the real trajectory position information.
[0074] Under the above technology, a trajectory privacy protection mechanism based on double-delayed deep deterministic policy gradient is proposed to improve the performance of semantic trajectory privacy protection in a dynamic spatial environment. This mechanism first conducts the initialization work of the location service system and the relevant parameters of the deep reinforcement learning algorithm, accurately determines the sensitivity of the current trajectory position point based on the information contained in the two-dimensional space, and evaluates the set of potential possible position points of the current user trajectory position point, and then inputs it as the network state data of the LBS system into the deep convolutional neural network. Due to the introduction of the double Q network and the delayed update mechanism, TD3 can effectively reduce the variance generated during the training process, thus significantly enhancing the stability and accuracy of the learning process. This mechanism dynamically selects the most suitable position perturbation strategy by means of deep reinforcement learning, thereby successfully resisting the position inference attack behavior from attackers and achieving a delicate balance between user trajectory privacy protection and service quality. In addition, this mechanism also makes full use of the experience replay technology, stores all data such as the state information before and after the system, the selected strategy, and the evaluation results as experience in the experience pool, and in this way greatly improves the update efficiency of the network weight parameters, and finally successfully realizes the dual key objectives of both enhancing the user privacy protection and effectively guaranteeing the service quality. And the personalized semantic trajectory privacy protection mechanism based on double-delayed deep deterministic policy gradient mainly includes the following specific steps:
[0075] The specific implementation steps are as follows:
[0076] Step 1: Select a user device with deep learning computing capabilities and supporting positioning functions, such as a mobile phone, a vehicle-mounted positioning device, etc.
[0077] Step 2: Build a user trajectory privacy protection platform based on a high-performance LBS server, USRP devices, and users (such as in-vehicle navigation, mobile phones, indoor positioning devices, etc.). Among them, the high-performance LBS server is the server side, and the USRP devices and users are mobile users. The mobile users initiate requests for trajectory location protection according to user needs. The server side is responsible for receiving the users' requests and providing trajectory location protection services. The client can obtain corresponding trajectory privacy protection after requesting location services. This protection method will first initialize the parameters related to the location service system and the deep reinforcement learning algorithm, determine the sensitivity of the current trajectory location point based on two-dimensional space information, evaluate the possible location set of the current user trajectory location point, use it as the network state of the LBS system to input into the deep convolutional neural network, select the trajectory privacy protection strategy to be used according to the current LBS system network state, and finally publish the generated perturbed trajectory location point to the LBS server. At this time, only the user himself knows the true trajectory location point, while the LBS server obtains the perturbed trajectory point information. Therefore, this solution not only effectively protects the user's trajectory privacy, but also achieves the goal of providing a good LBS service experience for users, and verifies the feasibility of the trajectory protection mechanism in actual scenarios.
[0078] Step 3: Determine the perturbed location of the current trajectory point based on the perturbation strategy of the current time slot, which is divided into two stages. In the first stage, the user first determines the possible location set Δχ through the value of δ and the prior probability at the current moment and then uses the privacy budget and the lower bound of the expected inference error to search for the protected location set in Δχ k . Then in the second stage, through the use of the privacy budget, the user publishes the perturbed location through the permutation and flipping (PF) mechanism. k
[0079] Step 4: The user transmits the perturbed location in the current time slot to the LBS server, and then the server gives service feedback based on the received user location information. After that, based on the feedback results, evaluate the degree of privacy protection and the loss of service quality, calculate the system benefit of the user in the current time slot, and construct the network state of the LBS system in the next time slot through the Markov transition process.
[0080] Step 5: Adopt a dynamic perturbation strategy to protect the user's trajectory privacy based on user needs and the sensitivity of the current trajectory location point. Using the idea of transfer learning, recycle the determined perturbed location experience, store the system's front and back states, perturbation strategy, and evaluation results as experience in the experience pool, and continuously optimize the location protection strategy by updating the network weight parameters. Combining the user environment and the privacy leakage situation, the mobile user continuously repeats the process until a stable trajectory perturbation strategy is obtained, so as to achieve real-time trajectory privacy protection of the user in the LBS.
[0081] Combined withFigure 2 Content, and the specific steps are as follows:
[0082] 1) Initialize the two-dimensional spatial map accessed by the user, the Critic network, the Actor network parameters θ1, θ2, φ, and the target Critic network, Actor network parameters θ1 ' , θ2 ' , φ'.
[0083] 2) Calibrate the geographical location d (k) and semantic location c (k) where the user is currently accessing in the two-dimensional space, and view the sensitivity level l (k) of the grid where the user is located according to the user's own real location, and form the current LBS system network state of the user based on the information obtained in the previous step: s (k) ={x (k) , y (k) , l (k) , n (k)}, where n (k) represents the size of the set of possible locations, and x (k) , y (k) correspond to the horizontal and vertical coordinates of the geographical location d (k) respectively.
[0084] 3) At time slot k, the user inputs their current state into the actor and critic networks. The output of the Actor network is the action a (k) selected according to the current state:
[0085] a (k) =π(s (k) +ε,
[0086] where ε is the exploration noise and is independently sampled from a truncated normal distribution.
[0087] 4) Calculate the set of possible locations Δχ k of the user at time k:
[0088]
[0089] where Δχ k represents the set of the smallest number of locations whose sum of prior probabilities is less than 1-δ, a i is the possible location area of the user, and δ is a constant (0 < δ < 1); if the true location d (k) of the user is excluded with a very small probability, then the location nearest to d (k) is used to replace d where is the (k) Euclidean distance between
[0090] 5) Determine the protected location set based on the lower bound of the expected inference error obtained by the strategy and the privacy budget in Δχ k On this basis, use the Hilbert search algorithm to search for the personalized location protection set Φx, and then input the personalized location protection set Φ k after search matching into the location perturbation mechanism Permute-and-Flip, and perform the work of publishing perturbed locations through the location perturbation mechanism Permute-and-Flip. The distance between any two locations in the personalized location protection set Φ k is less than or equal to D(Φ k ), that is satisfies the following expression:
[0091]
[0092] 6) After the user publishes the perturbed location to the LBS server, according to the received user location information, the LBS server will give service feedback to the user; the user judges the privacy protection degree p (k) based on the feedback result, and estimates the service quality loss q (k) , q (k) is measured by the Manhattan distance; evaluate the network benefit u (k) of the LBS server where the user is located according to the following formula, and construct the network state of the LBS system in the next time slot. The specific expression of u (k) is as follows:
[0093]
[0094] In the formula, the weight parameter ω represents the importance of the service quality loss relative to the trajectory privacy protection; η is a non-zero constant.
[0095] 7) Use the user's location transition probability matrix M and the possible location set Δχ k of the user at time t, and use the posterior probability of the user at time k to infer the prior probability of the user at the next time k + 1, so as to infer multiple possible locations of the user at the next time. The posterior probability distribution of the user at time k satisfies the following formula:
[0096]
[0097] where d (k) , are the true location and the perturbed location of the user at time k respectively, is the location perturbation probability distribution. Then the prior probability distribution
[0098] 8) The prior probability at time k + 1 Input the set of position points Δχ k Using the calculation formula, obtain the possible position set Δχ at time k+1 k+1 .
[0099] 9) Based on the experience replay technique, store the historical experience of trajectory perturbation in the experience pool The historical experience of trajectory perturbation within time slot k is defined as e (k) , e (k) =(s (k) , a (k) , u (k) , s (k+1) ). Randomly select an experience e (k) from the memory pool, where 1 ≤ x ≤ k. Use the target Actor network to calculate the action a (k+1) in the state s (k+1) , and add noise to the target action to obtain a (k+1)′ . Then, based on the idea of the TD3 double network, calculate the target value:
[0100]
[0101] 10) Apply the gradient descent algorithm to minimize the error L c between the evaluation value and the target value, and update the parameters in the Critic1 and Critic2 networks as follows:
[0102]
[0103] 11) After updating the two Critic networks H times, update the Actor network parameters through deterministic policy gradients as follows:
[0104]
[0105] 12) Finally, introduce the learning rate and update the target network through soft update as follows:
[0106] θ′ i ←τθ i +(1 - τ)θ′ i (i = 1, 2),
[0107] φ′←τφ+(1 - τ)φ′.
[0108] According to the user's environment and privacy leakage situation, the mobile user repeats the above steps until a stable trajectory perturbation strategy is learned.
[0109] This method acts on the user's trajectory, based on CNN and TD3, considers the spatio-temporal correlation between trajectory position points, and has the following advantages:
[0110] Can significantly reduce the exposure of sensitive location information: Dynamically explore the optimal privacy parameter selection strategy through continuous trial and error. In this process, as the weight parameters of the actor network and the critic network are updated, the model can more accurately determine appropriate privacy protection measures according to different scenarios. Search using the updated network parameters combined with the privacy budget and the expected inference error bound, making the published perturbed locations more reasonable, reducing the possibility of attackers obtaining highly sensitive location information, and reducing the exposure risk of sensitive location information.
[0111] Balance privacy and quality of service: By updating the network weight optimization strategy, the method finds a better balance between privacy protection and quality of service. When calculating the reward, the privacy protection level of each location on the trajectory (evaluated according to the sensitivity difference between the actual location and the inferred location) and the quality of service loss (evaluated using the Manhattan distance) are comprehensively considered. As the network parameters are continuously updated, while ensuring a certain level of privacy protection, the impact on the quality of service is minimized.
[0112] Adapt to different privacy requirements: Different users or the same user at different locations have different privacy protection requirements. By updating the network weight optimization strategy, it can better adapt to these personalized needs. Since the model is constantly learning and adjusting parameters, it can dynamically adjust the privacy protection strategy according to the specific situation of the user, including the sensitivity of the current location and the user's emphasis on privacy. When the user is in a sensitive area, more stringent privacy protection measures can be selected using the updated parameters, while in ordinary areas, it can be appropriately relaxed to ensure a better service experience.
Claims
1. A trajectory privacy protection method based on double-delayed deep deterministic policy gradient, characterized by: The LBS system is built based on the LBS server and the user, including a two-dimensional spatial map with sensitivity annotated; the user's current trajectory point location information and the sensitivity of the location are used as the user state; the current user state is sent to the LBS server with a perturbation strategy based on double-delay deep deterministic policy gradient to send the perturbed pseudo location information; the user relies on the service feedback of the LBS server to judge the degree of privacy protection and the service quality loss, calculates the LBS system benefit of the user in the current time slot, and constructs the LBS system network state of the next time slot through the Markov transition process; based on user needs and the sensitivity of the current location, a dynamic perturbation strategy is used to The dynamic strategy protects the privacy of user trajectories; uses transfer learning to recycle the determined post-disturbance location experience, stores the LBS system network status, dynamic perturbation strategy and evaluation results of the current time slot and the next time slot as experience in the experience pool, and continuously optimizes the trajectory protection strategy by updating the weight parameters of the actor network and the critic network in the perturbation strategy; combines the user's real location, privacy protection level and service quality loss to obtain the LBS server network system benefit, and the user continuously updates the weight parameters of the actor network and the critic network until a stable trajectory perturbation strategy is obtained, ultimately achieving real-time trajectory privacy protection for users in the LBS system.
2. According to claim 1, the trajectory privacy protection method based on double-delayed deep deterministic policy gradient is characterized in that: The specific steps are as follows: S1, initialize the two-dimensional spatial map of the user accessing the LBS system, and the perturbation strategy parameters based on the double-delayed deep deterministic policy gradient network TD3; S2. Calibrate the geographical location and semantic location of the user's current visit in two-dimensional space, and check the sensitivity level of the location based on the user's actual location; form the LBS system network status of the current user, the possible location set of information, and the coordinates of the geographical location based on the information obtained in the previous step; S3, at time k, the user inputs his current state into the actor and critic network of TD3, and the actor network outputs the action selected by the current state; S4, calculate the possible location set of the user at time k; S5. Determine the protected location set based on the expected inference error lower bound and privacy budget obtained by the perturbation strategy, use the Hilbert search algorithm to search for the personalized location protection set, and then input the personalized location protection set content after the search and matching into the location perturbation mechanism Permute-and-Flip, and execute the work of publishing the perturbed location to the LBS server through the location perturbation mechanism Permute-and-Flip; S6. The LBS server provides service feedback to the user based on the received user location information; The user judges the privacy protection degree according to the feedback results, estimates the service quality loss and the network benefit of the LBS server at the user's location, and constructs the LBS system network status for the next time slot; S7, using the user's position transfer probability matrix and the possible position set of the user at time t, and using the user's posterior probability at time k to infer the user's prior probability at the next time k+1, thereby inferring multiple possible positions of the user at the next time; S8. Substitute the prior probability at time k+1 into the set of position points to obtain the possible position set at time k+1; S9, based on the experience playback technology, the historical experience of trajectory disturbance is stored in the experience pool, and then one experience is selected from the experience pool, and the target value is calculated based on the Critic1 network and the Critic2 network in TD3; S10, using the gradient descent algorithm to minimize the error between the evaluation value and the target value, thereby updating the parameters in the Critic1 network and the Critic2 network; S11. Update the Critic1 network and Critic2 network according to the user's environment and privacy leakage, update the Actor network parameters through deterministic policy gradient, and repeat this step until a stable trajectory perturbation strategy is learned.
3. The trajectory privacy protection method based on double-delayed deep deterministic policy gradient according to claim 1 is characterized in that: When the LBS system is used, it is initialized first, and the LBS system coverage area is simplified into a two-dimensional space map and gridded. The location function of each grid is semantically labeled and the sensitivity level is set. The user's current location sensitivity level and the size of the possible location set are input into the Actor network as the LBS system network state parameters using a deep reinforcement learning algorithm to obtain the current trajectory location point perturbation strategy: 3.
1. Marking the user’s current geographical location in a two-dimensional space map based on a grid (k) and semantic position c (k) , using the current geographic location of the visit (k) Represents the user's real trajectory location point, using the semantic location c (k) It is the location name indicating the user's current location point, which is used to obtain the user's location sensitivity level; the geographic location d (k) It is a two-dimensional coordinate, including the horizontal and vertical coordinates x (k) ,y (k) , k is the kth time slot; check the sensitivity level l of the grid according to the user's actual location (k) , according to the position sensitivity level l (k) Provide dynamic privacy protection; 3.
2. Based on the user’s current geographic location (k) Form the LBS system network status of the current user: (k) ={x (k) ,y (k) , l (k) , n (k) }, where n (k) represents the size of the possible position set, x (k) ,y (k) Corresponding to geographical location d (k) The horizontal and vertical coordinates of the semantic position and the sensitivity level l (k) , and the possible location set of users Δχ k ; 3.
3. At time k, the user inputs the network state of his LBS system into the actor network and critic network Actor-Critic in the double-delayed deep deterministic policy gradient algorithm TD3. The Ciritc network is used to evaluate the actions of the Actor network and assist the Actor network in optimizing the output actions. The output of the Actor network is the action a selected according to the user's current state. (k) : a (k) =π(s (k) )+e, The Actor network structure is based on a fully connected neural network, ε is the exploration noise, which is independently sampled from a truncated normal distribution, and action a (k) This includes the privacy budget for the current location and the lower bound on the expected inference error.
4. The trajectory privacy protection method based on double-delayed deep deterministic policy gradient according to claim 3 is characterized in that: The perturbation strategy of the double-delayed deep deterministic policy gradient adopted by the user is as follows: the user first initializes the gridded two-dimensional space map and the TD3 algorithm, determines the sensitivity of the user's current trajectory position point based on the two-dimensional space map information, and evaluates the possible position set ΔX of the current user's trajectory position point k The deep convolutional neural network (CNN) embedded in the Actor-Critic network, which serves as the network state input of the LBS system, processes the data input to the Actor network and the Critic network, uses CNN to extract features from the data, and the extracted features are used for subsequent calculations of the Actor network and the Critic network.
5. The trajectory privacy protection method based on double-delayed deep deterministic policy gradient according to claim 4 is characterized in that: The user determines the disturbance position of the current trajectory point based on the disturbance strategy of the current time slot: 5.
1. The user uses the privacy budget of the current time slot and the lower bound of the expected inference error to determine the perturbation position of the current trajectory point: First, the possible position set ΔX of the user's trajectory position point is calculated based on the prior at the current time. k Then, the privacy budget and the expected inference error lower bound are used to determine the protected location set Φ of the user's current location k ; The possible location set Δχ of the user at time k is calculated using the following formula: k , the set of possible positions Δχ k Filtered by prior probability, including the user's real location: where Δχ k represents the minimum number of positions whose sum of prior probabilities is less than 1-δ, a i is the possible location of the user, δ is a constant (0<δ<1), represents the prior distribution at the current moment, i represents the probability of the i-th region; 5.
2. Using the actions output by the actor network as the lower bound of the expected inference error obtained by the strategy, and using the privacy budget to determine the personalized location protection set Φ k , personalized location protection set Φ k The distance between any two positions in the k ),Right now And satisfy the following expression: Where e € represents the privacy budget power of the exponential function e; 5.
3. Users search for the personalized location protection set Φ through the Hilbert search algorithm k Then, the personalized location protection set Φ after search matching is k The content is input into the location perturbation mechanism Permute-and-Flip, which uses the location perturbation mechanism Permute-and-Flip to publish the user's perturbed location to the LBS server. The perturbation location probability distribution matrix is: in is used for normalization, ∈ is the privacy budget, is the set of all possible locations of the user.
6. The trajectory privacy protection method based on double-delayed deep deterministic policy gradient according to claim 5 is characterized in that: The user transmits the current disturbance location to the LBS server, and the LBS server gives service feedback to the user based on the received user location information; Users judge the privacy protection level based on the feedback results (k) , and estimate the service quality loss q (k) , privacy protection level p (k) The service quality loss q is determined by the difference between the current location and the location sensitivity inferred by the attacker. (k) It is an evaluation conclusion of the distance between the user's real location and the disturbance location. The farther the distance between the two, the greater the loss of service quality. Calculate the privacy protection level first Then we get the service quality loss Where l (k) and They represent the sensitivity of the user’s actual location and the sensitivity of the attacker’s inferred location, respectively. k ,y k ) represents the user's original location, Indicates the disturbance position published by the user terminal; Then, the LBS server network system benefit u of the user is evaluated according to the following formula: (k) , and construct the user status at the next moment, LBS server network system benefits u (k) The specific expression is as follows: u (k) =p (k) -ωq (k) Where the weight parameter ω represents the importance of service quality loss relative to trajectory privacy protection; η is a constant that is not 0.
7. The trajectory privacy protection method based on double-delayed deep deterministic policy gradient according to claim 6 is characterized in that the time correlation between different positions in the user's trajectory at time k is used to obtain the user's real position at the next time k+1 and the possible position set Δx k+1 , construct the LBS system network status at the next moment: 7.1: Use the user's historical trajectory data and behavior habits to construct the perturbation position transfer probability matrix M. The position transfer probability matrix M is a two-dimensional n*n matrix, which represents the probability of each position transferring from the current position to the next corresponding position. Let m ij is the element in the i-th row and j-th column of the position transfer probability matrix M, then m ij represents the probability of a user moving from area i to area j; 7.2: Based on the possible location set Δχ of the user at time t k The position information contained in the position transfer probability matrix M is used to find the corresponding row vector m i , m i represents the probability of the user moving from area i to other areas. These row vectors constitute the position transfer probability matrix M at time k. k ,Through the transition probability matrix and the Markov transition model, the user’s real location at the next moment is obtained; Use the user's posterior probability at time k to infer the user's prior probability at the next time k+1 Thus, multiple possible locations of the user at the next moment can be inferred; The posterior probability distribution of the user at time k satisfies the following formula: where d (k) , are the user’s true position and disturbed position at time k, respectively. is the probability distribution of position disturbance. Then the prior probability distribution at time k+1 is 7.3: Considering time correlation, the prior probability at time k+1 Substitute the position point set Δχ at time k k Calculate the formula to get the possible position set Δχ at time k+1 k+1 .
8. The trajectory privacy protection method based on double-delayed deep deterministic policy gradient according to claim 7 is characterized by: The LBS system network status, dynamic perturbation strategy and evaluation results of the current time slot and the next time slot are stored in the experience pool as experience, and the trajectory protection strategy is continuously optimized by updating the weight parameters of the actor network and the critic network. The specific steps are as follows: 8.1: The historical experience of trajectory disturbance in time slot k is defined as e (k) , e (k) =(s (k) , a (k) ,u (k) ,s (k+1) ); Based on the experience playback technology, the trajectory disturbance historical experience is stored in the experience pool Randomly select an experience e from the memory pool (k) , where 1≤x≤k uses the target Actor network to compute the state s (k+1) Action a (k+1) , and add noise to the target action to get a (k +1)′ , then based on the target Critic1 network and Critic2 network of the TD3 algorithm, the target value is calculated using the following formula: 8.2: The state of the current time slot and the action output by the Actor network a (k) Input two independent Critic1 networks respectively and Critic2 network In the , the Critic network will be the action a corresponding to the state in the current time slot (k) The larger the evaluation Q value of the action, the better the action a (k) The higher the value in the current state; After evaluating the corresponding evaluation Q value, the gradient descent algorithm is applied to make the evaluation Q value and the target value y (k) The error between c Minimize and update the parameters in the Critic1 network and the Critic2 network. The specific calculation is as follows: Among them, the Critic1 network and the Critic2 network have the same structure, and the two are set up independently and in parallel; 8.3: After updating the two Critic1 networks and Critic2 networks H times, the Actor network parameters are updated as follows through the double-delayed deep deterministic policy gradient: 8.4: Finally, the learning rate τ is introduced, and the target network θ′ in the TD3 algorithm is calculated using the following formula i , the φ′ target network provides a stable target value for the system state at time k+1. The specific update process of the target network is as follows: θ′ i ←tth i +(1-τ)θ′ i (i=1,2), φ′←τφ+(1-τ)φ′. According to the user's environment and privacy leakage, the mobile user repeats the above steps until the TD3 network learns a stable trajectory perturbation strategy.
9. A computer device, characterized in that: It includes a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the trajectory privacy protection method based on double-delay deep deterministic policy gradient as described in any one of claims 1-8.
Citation Information
Cited By
Loop reinforcement learning-based position privacy protection method under three-dimensional positioning error
CN121665227A