A Location Privacy Protection Method for 3D Localization Errors Based on Recurrent Reinforcement Learning
By employing a recurrent reinforcement learning method in three-dimensional space, combined with a long short-term memory network and a dual-delay deep deterministic policy gradient algorithm, and dynamically selecting privacy parameters, the problem of location privacy protection under three-dimensional positioning errors is solved, achieving a balance between maximizing user utility and service quality, and adapting to complex and ever-changing environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA UNIV OF MINING & TECH
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-21
AI Technical Summary
Existing 3D positioning privacy protection methods suffer from reduced positioning accuracy when faced with factors such as multipath effects and signal interference, making it difficult for privacy protection mechanisms to function effectively. Furthermore, traditional methods are difficult to adapt dynamically in complex and ever-changing 3D environments, failing to achieve a balance between privacy protection and service quality.
A recurrent reinforcement learning-based approach is adopted to model 3D positioning error as a partially observable Markov decision process. By combining a long short-term memory network with a dual-delay deep deterministic policy gradient algorithm, privacy parameters are dynamically selected. Through user perturbation location and service feedback, a balance between location service and privacy protection is achieved.
It effectively protects user location privacy in three-dimensional space, adapts to complex and ever-changing environments, maximizes user utility through weight parameter adjustment, avoids dependence on environment and attack models, and improves the accuracy of location services and the effectiveness of privacy protection.
Smart Images

Figure CN121665227B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for protecting location privacy under three-dimensional positioning errors based on recurrent reinforcement learning, belonging to the field of location services and information security technology. Background Technology
[0002] Currently, various LPPM (Location Privacy Preserving Mechanisms) have been developed, including anonymized, encrypted, perturbation, and pseudonymous location privacy preservation methods ([1] H. Jiang, J. Li, P. Zhao, F. Zeng, Z. Xiao, and A. Iyengar, “Location privacy-preserving mechanisms in location-based services: A comprehensive survey,” ACM Computing Surveys (CSUR), vol. 54, no. 1, pp. 136, 2021.). However, most existing mechanical methods do not consider the height information of three-dimensional geolocation, and therefore cannot be applied to location privacy preservation in three-dimensional space. ([2]Y. Zhu and L. Zhai, “Location privacy in buildings: A 3-dimensional k-anonymity model,” in 2014 10th International Conference on Mobile Ad-hoc and Sensor Networks, pp. 195–200, IEEE, 2014.) studied the k-anonymity model 3d Clique Cloak to achieve location privacy protection in 3D indoor spaces. ([3]R. Kumar and S. Dawra, “Simulation of 3d privacy preservation and locationmonitoring approach,” International Research Journal of Engineering and Technology (IRJET), vol. 3, no. 5, pp. 1099–1103, 2016.) proposed another mechanism that uses 3D resource awareness and 3D quality awareness algorithms to protect user privacy in multi-story and multi-section buildings. However, both mechanisms rely on trusted third parties.The geo-indistinguishability mechanism will disturb the user's geolocation locally, and only the user knows the actual location. ([4]M. Min, L. Xiao, J. Ding, H. Zhang, S. Li, M. Pan, and Z. Han, “3d geo-indistinguishability for indoor location-based services,”IEEE Transactions on Wireless Communications, vol. 21, no. 7, pp. 4682 4694, 2021.) Extends the geo-indistinguishability mechanism to 3D space (3D-GI) and uses Laplace noise to disturb the user's 3D geolocation. However, in 3D space scenarios, factors such as multipath effects and signal interference will significantly reduce the accuracy of 3D positioning. In this context, when implementing location privacy protection, if the user ignores the positioning error when perturbing the location, it is easy to misjudge the wrong location as the real location, thereby exacerbating the risk of privacy leakage and making it difficult for the existing privacy protection mechanism to play an effective role. At the same time, traditional privacy protection methods have major defects. The complex and ever-changing 3D environment has high time-varying characteristics and multiple attack methods. Traditional methods apply a uniform privacy budget across all locations, making it difficult to dynamically adapt to time-varying 3D environments and achieve a better balance between privacy protection and quality of service. Reinforcement learning techniques can dynamically balance privacy protection and quality of service without knowing the attack model.
[0003] Partially observable Markov decision processes (POMDPs) have demonstrated good adaptability and optimization capabilities in solving partially observable state problems in application scenarios such as smart city path planning and residential resource switching ([6] X. Chen, Q. Hu, Y. Zhang, M. Song, Z. Wu, and J. Xiong, “Pomdp based dispatch scheme for residential distributed energy resources under customerfatigue consideration,” IEEE Transactions on Smart Grid, 2024.). The advantage of POMDP is that it avoids over-reliance on complete but hard-to-obtain information, thereby reducing the estimation cost of information collection and processing and improving the feasibility and practicality of the system. For example, ([7] D. Qiu, Y. Wang, T. Zhang, M. Sun, and G. Strbac, “Hierarchical multi agent reinforcement learning for repair crews dispatch control towards multi-energy microgrid resilience,” Applied Energy, vol. 336, p. 120826, 2023.) combines POMDP and reinforcement learning to solve the maintenance crew dispatch problem in multi-energy microgrids. By constructing a POMDP model, the system can dynamically optimize the path and maintenance strategy of the maintenance team, and make independent observations using local information, thereby increasing the system's benefits. ([8]Y. Xue and W. Chen, “Rloplanner: Combining learning and motion planner for uav safe navigation in cluttered unknown environments,” IEEE Transactions on Vehicular Technology, vol. 73, no. 4, pp.4904 4917, 2023.) proposed a method combining recurrent neural networks (RNN) and reinforcement learning algorithms to solve the task scheduling problem in partially observable environments. This method uses LSTM networks to utilize historical observation information and optimize decision-making strategies, which significantly reduces computational complexity compared to traditional belief state estimation methods. Summary of the Invention
[0004] To address the shortcomings of existing technologies, a location privacy protection method based on cyclic reinforcement learning for 3D localization errors is proposed. This method is simple to implement and highly effective. It fully considers the impact of localization errors on location privacy protection, models the privacy protection problem as a partially observable Markov decision process, and adopts a location privacy protection mechanism that combines long short-term memory networks and a dual-delay deep deterministic policy gradient algorithm to dynamically select privacy parameters. This achieves an effective balance between location services and privacy protection, maximizing user utility.
[0005] To achieve the above technical objectives, this invention discloses a method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning, the specific steps of which are as follows:
[0006] S1. Use the precise coordinates of landmarks around the user as anchor points;
[0007] S2. Users generate the perturbation location of the current time slot based on their own positioning error and upload it to the server to obtain an assessment of the degree of privacy protection and service quality;
[0008] S3. The perturbed position of the current time slot with positioning error is modeled as a partially observable Markov decision process. A dynamic location privacy protection mechanism combining long short-term memory network and dual-delay deep deterministic policy gradient algorithm is adopted to protect the user's location privacy when the user's own positioning has errors.
[0009] Furthermore, users move within a 3D spatial environment and request real-time user location services. The user location includes longitude, latitude, and altitude, forming a detailed 3D trajectory record. To protect privacy, users randomly perturb their location, generating a perturbed location that is then uploaded to the LBS server.
[0010] Furthermore, the user's environment is a 3D location privacy protection system, which includes the user, LBS server, and base station. The user uses a time difference of arrival (TDOA) positioning method for 3D positioning. The measurement signals sent by the user to different anchor points are used to obtain the time difference of arrival of the measurement signals at different anchor points. Combined with the known precise coordinates of the anchor points, the formula is used: The distance difference between the i-th anchor point and the first anchor point is calculated, where c represents the speed of light and t represents the time to reach the i-th anchor point. The user's location is located at the intersection of a hyperboloid with multiple anchor points as foci. Finally, the precise real location of the user is estimated using the Chan's positioning algorithm.
[0011] Furthermore, given the known information about the attacker's attack on the user, the user generates the perturbation location of the current time slot based on their own location error and uploads it to the server. The server then provides service feedback on the user's location after the perturbation. Based on the service feedback received by the user, the privacy protection level and service quality of the current time slot are evaluated.
[0012] Users assess location privacy based on the Euclidean distance between their actual location and the user's inferred location by an attacker. ;
[0013] In the formula, Representing user location privacy, Represents the user's precise real location. Represents the inferred location. The Euclidean distance between the user's actual location and the user's inferred location by the attacker;
[0014] The Sigmoid function is used to measure user satisfaction with privacy-preserving services. The Sigmoid function is related to the accuracy of location privacy-preserving services under user location errors.
[0015] ,
[0016] In the formula, Indicates service quality, This represents the upper bound of user satisfaction with service quality. Indicates the steepness of the service quality curve. The parameter represents the distance between the user's actual location and the disturbance location. Used to adjust service quality Sensitivity For service quality threshold;
[0017] By adjusting user weight parameters, a balance is struck between location privacy and quality of service, and user utility is used to represent the trade-off between location privacy and quality of service.
[0018] ,
[0019] In the formula, It is the weighting parameter, and k represents the time slot.
[0020] Furthermore, the steps for users to generate the disturbance position of the current time slot based on their own positioning errors are as follows:
[0021] Choose the action according to the following formula :
[0022] ,
[0023] ,
[0024] In the formula, To explore noise, samples are independently taken from a truncated normal distribution; The representative mean is variance is The normal distribution; and These are the upper and lower boundaries of the cutoff.
[0025] Calculate the location of the disturbance using the following formula :
[0026] ,
[0027] In the formula, Represents the disturbance distance. and These represent the polar angle and the azimuth angle, respectively. This indicates that the user obtained a location with errors through positioning. These represent the horizontal, vertical, and height measurements obtained by the user through positioning, which may contain errors.
[0028] Distance Based on privacy budget in gamma distribution The sampling yielded:
[0029] ,
[0030] In the formula, This refers to a privacy budget.
[0031] Furthermore, the Partially Observable Markov Decision Process (POMDP) includes information about the user's true state, actions, observed state, state transition function, observation function, reward function, and discount factor. It infers all possible sets of true states from the observation function, and the workflow is a loop of "perception-inference-decision-update".
[0032] Let the user's actual state in time slot k be represented as ,in, This indicates the attack result of the user in the previous time slot k; the value is 1 if the attack is successful and 0 otherwise.
[0033] The action represents the privacy protection strategy adopted by the user in time slot k, including privacy budget, polar angle, and azimuth angle. The action is specifically represented as follows: ;
[0034] The observed status is the user's location obtained through positioning. and the result of the attack ,Location Positioning errors exist due to multipath effects, signal interference, and hardware performance limitations; the observation state is specifically represented as follows: ;
[0035] The state transition function updates the user's state based on the user's previous state in the previous time slot and the privacy protection policy, including updates to the user's real location and the attack results;
[0036] The observation function is the function that obtains the observed state based on the actual state, i.e., the localization process;
[0037] The reward function is used to guide users to consider the relevant service quality while protecting location privacy, as a form of user utility;
[0038] The discount factor is used to weigh the importance of immediate rewards versus future rewards.
[0039] Furthermore, an estimate of the user's true location is generated by fusing historical data and user observation states through a Long Short-Term Memory (LSTM) network. The specific process is as follows:
[0040] Defining past history For including the previous Observation status and action sequence of each time slot :
[0041] ,
[0042] In the formula, and This represents the initial state when there is no user observation or action input. and This indicates the observation status and actions of the user in the previous time slot (k-slot);
[0043] The observation state and action sequence of the past Estimated user's actual state compared to the previous moment Inputting into an LSTM network, a gating mechanism is used to selectively retain historical information and update the user's true state estimate. :
[0044] ,
[0045] in These are the parameters for the LSTM network.
[0046] Furthermore, the location privacy protection strategy is continuously optimized using the twin-delayed deep deterministic policy gradient algorithm TD3 by updating the weight parameters of the actor network (Actor), critic networks (Critic1), and critic networks (Critic2). The process is as follows:
[0047] S3.1. The historical experience of positional perturbation within time slot k is represented as follows: Based on the experience replay method, historical experience of location perturbation is stored in the experience pool. From the experience pool A random empirical sample is selected from the data, denoted as: The user's actions are calculated using an Actor network. And add clipping noise to the target action to obtain Using Critic1 network as the target Q network Critic2 network target Q network Calculate the target value using the following formula. And serve as a supervisory signal for updating the Critic1 and Critic2 networks:
[0048] ,
[0049] Using the Critic1 and Critic2 networks to define the actions corresponding to the state in the current time slot. The action evaluation Q-score indicates the user's action. A higher Q-score indicates a more effective action. The higher the value under the current observation state; among which Indicates user utility. Represents the discount factor;
[0050] S3.2. Using the loss function Update the parameters of Critic1 and Critic2 networks :
[0051] ,
[0052] The parameters of Critic1 and Critic2 networks are updated by minimizing the loss function using the gradient descent algorithm. :
[0053] ,
[0054] in This represents the size of the batch sampled from the experience pool;
[0055] S3.3. Based on the outputs of Critic1 and Critic2 networks, adjust the Actor network parameters using the policy gradient ascent method to maximize the cumulative expected utility; perform an Actor network parameter update only once after Critic1 and Critic2 networks have completed H updates, using the following formula:
[0056] ,
[0057] in This represents the gradient of the objective function with respect to the parameters of the Actor network. This represents the gradient of the policy function with respect to the parameters, used for network parameter updates;
[0058] S3.4. Employ a soft update strategy, using the following formula to weightedly fuse the parameters of the current Critic1 and Critic2 networks with the parameters of the corresponding target network:
[0059] ,
[0060] The parameters of the target Actor network are also updated following the soft update rule:
[0061] ,
[0062] in and These represent the parameters of the target Critic and target Actor networks, respectively. This represents the soft update coefficient. Depending on the user's environment and privacy breach situation, mobile users repeat the above steps until the preset privacy protection requirements are met.
[0063] A computer device includes a processor and a memory, the processor being electrically connected to the memory for storing instructions and data, and the processor for executing a location privacy protection method for three-dimensional localization errors based on recurrent reinforcement learning.
[0064] Beneficial Effects: This invention fully considers the problem of limited positioning accuracy in 3D space due to user positioning errors. By modeling the problem as a partially observable Markov decision process and introducing recurrent reinforcement learning technology, it endows the location perturbation strategy with strong learning and adaptive performance, effectively compensating for the shortcomings of traditional research in terms of poor environmental adaptability and insufficient learning ability when dealing with complex and changing environments. This invention highly values the personalized needs of users when using location services, balancing the loss of location privacy and service quality by adjusting weight parameters to maximize user utility. This invention employs a unique recurrent reinforcement learning algorithm that does not rely on specific models. Through the interaction between the Long Short-Term Memory (LSTM) network, TD3 network, and the environment, it directly learns the decision strategy, adapting to different environmental changes and handling model uncertainty and complexity, thus avoiding detailed modeling of the environment. It specifically addresses the key problem common in many current location privacy protection studies: over-reliance on established environment models and attack models, resulting in a lack of autonomous learning ability and effective adaptability to complex environments. In practical applications, when it is difficult to accurately obtain user environment parameters, the Long Short-Term Memory (LSTM) network is used to deeply mine historical data, extract time features and fuse them with the current state to infer the user's potential state, providing comprehensive information for decision-making. The TD3 network is used to continuously interact with the environment and through a dynamic trial-and-error process, combined with the dual-Q network and delayed policy updates, to prevent overfitting and stabilize policy learning, and to fully explore and determine the optimal location privacy protection strategy. Attached Figure Description
[0065] Figure 1 This is a schematic diagram of the location privacy protection method model based on recurrent reinforcement learning in three-dimensional space under localization error in this invention.
[0066] Figure 2 This is a schematic diagram of the location privacy protection method for localization errors in three-dimensional space based on recurrent reinforcement learning in this invention. Detailed Implementation
[0067] The implementation of the present invention will be further described below with reference to the accompanying drawings:
[0068] like Figure 1 As shown, this invention discloses a location privacy protection method for 3D scene localization errors based on recurrent reinforcement learning. An LBS system is constructed based on an LBS server and users, including a 3D spatial map divided into different regions; a database of spatiotemporal related trajectory privacy protection features of mobile users is established, from which user behavioral location features are extracted. The user's current location information and the attack result are used as the user's state. A perturbation strategy combining a long short-term memory network and a dual-delay deep deterministic policy gradient algorithm is used to send perturbed false location information to the LBS server. The attacker is an untrusted location server or an external attacker who possesses certain prior knowledge. Specifically, by analyzing the user's historical location data and lifestyle habits, the attacker obtains the user's possible location set and the prior distribution of these locations, and is aware of the user's location perturbation mechanism. The attacker then calculates the posterior probability distribution and infers the user's actual location by minimizing the expected inference error based on the posterior distribution of the false location.
[0069] The system relies on service feedback from LBS servers to assess the degree of privacy protection and the extent of service quality loss. Users evaluate location privacy and service quality based on the distance between their actual location and the attacker's inferred location. Based on user needs, weight parameters are adjusted to balance location privacy and service quality loss. A partial Markov decision process is used to construct the LBS system network state for the next time slot. A dynamic perturbation strategy is employed to protect user location privacy. The current observation state, actions, user utility, next observation state, and termination flag are stored in an experience replay buffer. The trajectory protection strategy is continuously optimized by updating the weight parameters of the actor and critic networks in the perturbation strategy. The current observation state and actions are added to the historical data. A target value is obtained by sampling small batches of data from the experience replay buffer. Users continuously update the weight parameters of the actor and critic networks until a stable trajectory perturbation strategy is obtained, ultimately achieving effective location privacy protection for users under 3D scene positioning errors.
[0070] like Figure 2As shown, the specific steps of the location privacy protection method based on recurrent reinforcement learning in 3D space under localization error are as follows:
[0071] Step 1: Select users with 3D positioning capabilities and deep learning computing power, including smartphones and vehicle-mounted 3D positioning terminals that support high-precision BeiDou positioning. A 3D location privacy protection system is built based on a high-performance LBS server, 3D positioning anchor point devices (such as base stations with time-of-arrival positioning capabilities), and the user. The LBS server acts as the server, responsible for receiving perturbed locations uploaded by users and providing feedback. The precise coordinates of surrounding landmarks, including farms, hospitals, stadiums, gas stations, banks, and parks, are used as anchor points to help determine the user's true location information. The user acts as the client, running a location perturbation algorithm to initiate location service requests and obtain LBS services.
[0072] Step 2: Initialize the 3D location service environment, construct a three-dimensional spatial map containing longitude, latitude, and altitude information for the LBS system coverage area, calibrate the anchor point coordinates in the 3D space, and establish a 3D positioning model based on time difference of arrival positioning technology and Chan's positioning algorithm to estimate the user's location and positioning error.
[0073] Step 3: Determine the perturbation location of the current time slot based on the perturbation strategy of the current time slot. The perturbation location is generated based on 3D-GI, and the perturbation distance is sampled in the gamma distribution according to the privacy budget obtained by the perturbation strategy.
[0074] Step 4: The user transmits the location information of the current time slot, which has been perturbed and contains positioning errors, to the LBS server. The LBS server then provides service feedback based on the received user location information, evaluates the degree of privacy protection and service quality based on the feedback results, calculates the system benefits of the user in the current time slot, and constructs the LBS system network state for the next time slot using a Markov transition process.
[0075] Step 5: The user makes a location service request based on the perturbed location generated by 3D geographic indistinguishability. This includes perturbation policy generation, location perturbation at the current location, and location service request based on the perturbed location. Specifically, perturbation policy generation involves the user observing their own location information and estimating the attack outcome at the beginning of each iteration, selecting the optimal perturbation policy to balance location privacy and quality of service. Location perturbation involves the user running the perturbation policy on their local smart device, adding noise to the current location to generate a perturbed location to hide the real location. Location service request based on the perturbed location involves the user uploading the perturbed location to request location services, and the location service server responds based on the perturbed location.
[0076] Given information about attacks by external attackers against the user, the user generates a perturbed location for the current time slot based on their own location error and uploads it to the server. The server then provides service feedback based on the received perturbed user location. Combining the service feedback received by the user, the server assesses the level of privacy protection and service quality of the current time slot.
[0077] Users assess location privacy based on the distance between their actual location and the user's inferred location by the attacker: ;
[0078] Represents the user's precise real location. Represents the inferred location. The Euclidean distance between the user's actual location and the user's inferred location by the attacker;
[0079] The Sigmoid function is used to measure user satisfaction with privacy-preserving services. The Sigmoid function is related to the accuracy of location privacy-preserving services under user location errors.
[0080] ,
[0081] In the formula This represents the upper bound of user satisfaction with service quality. Indicates the steepness of the service quality curve. The parameter represents the distance between the user's actual location and the disturbance location. Used to adjust service quality Sensitivity For service quality threshold;
[0082] By adjusting user weight parameters, location privacy and service quality are balanced, and user utility is used to represent the trade-off between location privacy and service quality.
[0083] ,
[0084] Representing user location privacy, Represents service quality. It is the weighting parameter, and k represents the time slot.
[0085] Step 6: The steps for the user to generate the disturbance position of the current time slot based on their own positioning error are as follows: Select the action according to the following formula. :
[0086] ,
[0087] ,
[0088] in, To explore noise, samples are independently taken from a truncated normal distribution; This represents a normal distribution with a mean of 0 and a variance of δ. and These are the upper and lower boundaries of the cutoff.
[0089] The location of the disturbance is calculated using the following formula:
[0090] ,
[0091] in Represents the disturbance distance. and These represent the polar angle and the azimuth angle, respectively. This indicates that the user obtained a location with errors through positioning. These represent the horizontal, vertical, and height measurements obtained by the user through positioning, which may contain errors.
[0092] Distance Based on privacy budgeting, the following was sampled from the gamma distribution:
[0093] ,
[0094] In the formula, Privacy budget.
[0095] Step 7: Model the perturbed position of the current time slot with positioning error as a partially observable Markov decision process (POMDP). The POMDP includes the user's true state, actions, observed state, state transition function, observation function, reward function, and discount factor. It infers all possible sets of true states from the observation function. The workflow is a loop of "perception-inference-decision-update". Here, let the user's true state in time slot k be represented as... ,in, This indicates the attack result of the user in the previous time slot k; a value of 1 indicates a successful attack, and a value of 0 indicates an unsuccessful attack. The action represents the privacy protection strategy adopted by the user in time slot k, including privacy budget, polar angle, and azimuth angle. The user action is specifically represented as follows: The observed status is the user's location obtained through location services. and the result of the attack ,Location Positioning errors exist due to multipath effects, signal interference, and hardware performance limitations; the observation state is specifically represented as follows: The state transition function updates the user's state based on the user's previous state in the previous time slot and the privacy protection policy, including updates to the user's real location and the attack results; the observation function is the function that obtains the observed state based on the real state, i.e., the positioning process; the reward function is used to guide users to consider the relevant service quality while protecting location privacy, as a user utility.
[0096] Step 8: Employ a dynamic location privacy protection mechanism that combines a long short-term memory network with a dual-delay deep deterministic strategy gradient algorithm to protect user location privacy in the event of positioning errors.
[0097] An estimate of the user's true location is generated by fusing historical data and user observation states through a Long Short-Term Memory (LSTM) network. The specific process is as follows:
[0098] Defining past history For including the previous Observation status and action sequence of each time slot :
[0099] ,
[0100] In the formula, and This represents the initial state when there is no user observation or user action input. and This represents the user observation status and user actions in the previous time slot (k time slot);
[0101] The observation state and action sequence of the past Estimated user's actual state compared to the previous moment Inputting into an LSTM network, a gating mechanism is used to selectively retain historical information and update the user's true state estimate. :
[0102] ,
[0103] in These are the parameters for the LSTM network.
[0104] The following is a process of continuously optimizing the location privacy protection strategy using the twin-delayed deep deterministic policy gradient algorithm TD3 by updating the weight parameters of the actor network and the critic networks Critic1 and Critic2:
[0105] S3.1, the historical experience of positional perturbation within time slot k is represented as follows: Based on the experience replay method, historical experience of location perturbation is stored in the experience pool. From the experience pool A random empirical sample is selected from the data, denoted as: User actions are calculated using an Actor network. And add clipping noise to the target action to obtain Using Critic1 network as the target Q network Critic2 network target Q network Calculate the target value using the following formula. And serve as a supervisory signal for updating the Critic1 and Critic2 networks:
[0106] ,
[0107] Using the Critic1 and Critic2 networks to define the actions corresponding to the state in the current time slot. The Q-value for motion evaluation indicates that the action is effective. A higher Q-value indicates better performance. The higher the value under the current observation state; among which Indicates user utility. Represents the discount factor;
[0108] S3.2, Using the loss function Update Critic network parameters:
[0109] ,
[0110] The parameters of Critic1 and Critic2 networks are updated by minimizing the loss function using the gradient descent algorithm. :
[0111] ,
[0112] in This represents the size of the batch sampled from the experience pool;
[0113] S3.3. Based on the outputs of Critic1 and Critic2 networks, adjust the Actor network parameters using the policy gradient ascent method to maximize the cumulative expected utility; perform an Actor network parameter update only once after Critic1 and Critic2 networks have completed H updates, using the following formula:
[0114] ,
[0115] in This represents the gradient of the objective function with respect to the parameters of the Actor network. This represents the gradient of the policy function with respect to the parameters, used for network parameter updates;
[0116] S3.4. Employ a soft update strategy, using the following formula to weightedly fuse the parameters of the current Critic1 and Critic2 networks with the parameters of the corresponding target network:
[0117] ,
[0118] The parameters of the target Actor network are also updated following the soft update rule:
[0119] ,
[0120] in and These represent the parameters of the target Critic and target Actor networks, respectively. This represents the soft update coefficient. Depending on the user's environment and privacy breach situation, mobile users repeat the above steps until the preset privacy protection requirements are met.
[0121] This method operates on the user's 3D position trajectory, is based on an LSTM-TD3 network, and takes into account 3D scene localization errors, offering the following advantages:
[0122] Accurately adapts to the positioning error problem in 3D scenes: It fully considers the user's own positioning error caused by multipath effect, signal interference and hardware performance limitations in 3D space, avoids the risk of privacy leakage caused by traditional methods ignoring errors and misjudging the location, and builds a location privacy protection framework based on the indistinguishability of 3D geography. It models the privacy protection problem under positioning error as a partially observable Markov decision process (POMDP), which is more in line with the real 3D location service scenario.
[0123] Powerful historical information mining and policy learning capabilities: It innovatively combines Long Short-Term Memory (LSTM) network and Double Delayed Deep Deterministic Policy Gradient (TD3) algorithm. LSTM can deeply mine historical position and action data, extract time features and fuse them with the current state to provide comprehensive information for decision-making. TD3 uses a double Q network and delayed policy update to effectively prevent overfitting and stabilize policy learning, dynamically select privacy parameters, and avoid the shortcomings of traditional fixed privacy budgets that are difficult to adapt to time-varying 3D environments.
[0124] Achieving a dynamic balance between privacy and service quality: By balancing the loss of location privacy and service quality through user-adjustable weight parameters, assessing the degree of privacy protection based on the distance between the real location and the location inferred by the attacker, using the Sigmoid function to measure service quality, and defining a scientific method for calculating user utility, the system maximizes user utility while protecting user location privacy.
[0125] No need to rely on complex environments and attack models: It adopts a recurrent reinforcement learning algorithm that does not rely on specific models. Through continuous interaction and dynamic trial and error between the LSTM-TD3 network and the 3D location service environment, it directly learns the optimal decision strategy. It can handle model uncertainty and complexity, and solve the problems of traditional methods that rely too much on established environment models and attack models and lack autonomous learning and adaptability to complex environments. It can still play an effective role when it is difficult to accurately obtain system environment parameters.
Claims
1. A method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning, characterized in that: The specific steps are as follows: S1. Use the precise coordinates of landmarks around the user as anchor points; S2. Users generate the perturbation location of the current time slot based on their own positioning error and upload it to the server to obtain an assessment of the degree of privacy protection and service quality; S3. The perturbed position of the current time slot with positioning error is modeled as a partially observable Markov decision process. A dynamic location privacy protection mechanism combining long short-term memory network and dual-delay deep deterministic policy gradient algorithm is adopted to protect the user's location privacy when the user's own positioning has errors. Partially observable Markov decision processes (POMDPs) contain information about the user's true state, actions, observed state, state transition function, observation function, reward function, and discount factor. They infer all possible sets of true states from the observation function, and the workflow is a loop of "perception-inference-decision-update". Let the user's actual state in time slot k be represented as ,in, This indicates the attack result of the user in the previous time slot k; the value is 1 if the attack is successful and 0 otherwise. The action represents the privacy protection strategy adopted by the user in time slot k, including privacy budget, polar angle, and azimuth angle. The action is specifically represented as follows: ; The observed status is the user's location obtained through positioning. and the result of the attack ,Location Positioning errors exist due to multipath effects, signal interference, and hardware performance limitations; the observation state is specifically represented as follows: ; The state transition function updates the user's state based on the user's previous state in the previous time slot and the privacy protection policy, including updates to the user's real location and the attack results; The observation function is the function that obtains the observed state based on the actual state, i.e., the localization process; The reward function is used to guide users to consider the relevant service quality while protecting location privacy, as a form of user utility; Discount factors are used to weigh the importance of immediate rewards against future rewards; An estimate of the user's true location is generated by fusing historical data and user observation states through a Long Short-Term Memory (LSTM) network. The specific process is as follows: Defining past history For including the previous Observation status and action sequence of each time slot : , In the formula, and This represents the initial state when there is no user observation or action input. and This indicates the observation status and actions of the user in the previous time slot (k-slot); The observation state and action sequence of the past Estimated user's actual state compared to the previous moment Inputting into an LSTM network, a gating mechanism is used to selectively retain historical information and update the user's true state estimate. : , in These are the parameters for the LSTM network.
2. The method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning according to claim 1, characterized in that: Users move within a 3D spatial environment and request real-time user location services. The user's location includes longitude, latitude, and altitude, forming a detailed 3D trajectory record. To protect privacy, users randomly perturb their location, generating a perturbed location that is then uploaded to the LBS server.
3. The method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning according to claim 2, characterized in that: The user's environment is a 3D location privacy protection system, which includes the user, LBS server, and base station. The user uses a time-of-arrival (TOA) positioning method for 3D positioning. Measurement signals sent by the user to different anchor points are used to obtain the time difference between their arrival at those anchor points. Combined with the known precise coordinates of the anchor points, the formula is used: The distance difference between the i-th anchor point and the first anchor point is calculated, where c represents the speed of light and t represents the time to reach the i-th anchor point. The user's location is located at the intersection of a hyperboloid with multiple anchor points as foci. Finally, the precise real location of the user is estimated using the Chan's positioning algorithm.
4. The method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning according to claim 3, characterized in that, Given the attacker's information about the user's attack, the user generates a disturbed location for the current time slot based on their own location error and uploads it to the server. The server then provides service feedback on the user's location after the disturbance. Based on the service feedback received by the user, the privacy protection level and service quality of the current time slot are evaluated. Users assess location privacy based on the Euclidean distance between their actual location and the user's inferred location by an attacker. ; In the formula, Representing user location privacy, Represents the user's precise real location. Represents the inferred location. The Euclidean distance between the user's actual location and the user's inferred location by the attacker; The Sigmoid function is used to measure user satisfaction with privacy-preserving services. The Sigmoid function is related to the accuracy of location privacy-preserving services under user location errors. , In the formula, Indicates service quality, This represents the upper bound of user satisfaction with service quality. Indicates the steepness of the service quality curve. The parameter represents the distance between the user's actual location and the disturbance location. Used to adjust service quality Sensitivity For service quality threshold; By adjusting user weight parameters, a balance is struck between location privacy and quality of service, and user utility is used to represent the trade-off between location privacy and quality of service. , In the formula, It is the weighting parameter, and k represents the time slot.
5. The method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning according to claim 4, characterized in that, The steps for a user to generate the perturbation location of the current time slot based on their own positioning error are as follows: Choose the action according to the following formula : , , In the formula, To explore noise, samples are independently taken from a truncated normal distribution; The representative mean is variance is The normal distribution; and These are the upper and lower boundaries of the cutoff. Calculate the location of the disturbance using the following formula : , In the formula, Represents the disturbance distance. and These represent the polar angle and the azimuth angle, respectively. This indicates that the user obtained a location with errors through positioning. These represent the horizontal, vertical, and height measurements obtained by the user through positioning, which may contain errors. Distance Based on privacy budget in gamma distribution The sampling yielded: , In the formula, This refers to a privacy budget.
6. The method for protecting location privacy under 3D localization errors based on recurrent reinforcement learning according to claim 5, characterized in that, The following is a process of continuously optimizing the location privacy protection strategy using the twin-delayed deep deterministic policy gradient algorithm TD3 by updating the weight parameters of the actor network, critic1 network, and critic2 network: S3.
1. The historical experience of positional perturbation within time slot k is represented as follows: ; Based on the experience replay method, historical experience of location perturbation is stored in the experience pool. From the experience pool A random empirical sample is selected from the data, denoted as: The user's actions are calculated using an Actor network. And add clipping noise to the target action to obtain Using Critic1 network as the target Q network Critic2 network target Q network Calculate the target value using the following formula. And serve as a supervisory signal for updating the Critic1 and Critic2 networks: , Using the Critic1 and Critic2 networks to define the actions corresponding to the state in the current time slot. The action evaluation Q-score indicates the user's action. A higher Q-score indicates a more effective action. The higher the value under the current observation state; among which Indicates user utility. Represents the discount factor; S3.
2. Using the loss function Update the parameters of Critic1 and Critic2 networks : , The parameters of Critic1 and Critic2 networks are updated by minimizing the loss function using the gradient descent algorithm. : , in This represents the size of the batch sampled from the experience pool; S3.
3. Based on the outputs of Critic1 and Critic2 networks, adjust the Actor network parameters using the policy gradient ascent method to maximize the cumulative expected utility; perform an Actor network parameter update only once after Critic1 and Critic2 networks have completed H updates, using the following formula: , in This represents the gradient of the objective function with respect to the parameters of the Actor network. This represents the gradient of the policy function with respect to the parameters, used for network parameter updates; S3.
4. Employ a soft update strategy, using the following formula to weightedly fuse the parameters of the current Critic1 and Critic2 networks with the parameters of the corresponding target network: , The parameters of the target Actor network are also updated following the soft update rule: , in and These represent the parameters of the target Critic and target Actor networks, respectively. This represents the soft update coefficient. Depending on the user's environment and privacy breach situation, mobile users repeat the above steps until the preset privacy protection requirements are met.
7. A computer device, characterized in that, It includes a processor and a memory, the processor being electrically connected to the memory, the memory being used to store instructions and data, and the processor being used to execute the location privacy protection method for three-dimensional localization error based on recurrent reinforcement learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Privacy protection dynamic edge cache design method based on distributed reinforcement learning in mobile edge computing network
CN113364854A
Location privacy protection method in three-dimensional space LBS (Location Based Service) based on deep reinforcement learning
CN114117536A