Internet of vehicles track privacy protection method based on differential privacy and deep learning
By combining differential privacy and deep learning methods, gridded trajectories are generated and optimized using full convolutional neural networks, the problem of trajectory accuracy reduction in the existing technology is solved, and high-precision and high-availability trajectory data generation is achieved, while ensuring the protection of user privacy.
Patent Information
- Application Number
- CN202510233397.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art is difficult to generate high-precision and high-availability trajectory data while protecting user privacy, especially in complex traffic networks, where traditional differential privacy methods lead to a decrease in trajectory accuracy.
Using the Internet of Vehicles trajectory privacy protection method based on differential privacy and deep learning, a gridded trajectory generation method based on differential privacy is proposed, and a fully convolutional neural network (FCN) is used to optimize the trajectory.
While protecting user privacy, the generated trajectory data has high precision and high availability, which is significantly better than existing methods and can effectively solve the trajectory distortion problem caused by traditional differential privacy methods.
Smart Images

Figure CN120075263A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer science and technology, and particularly relates to the privacy protection technology of vehicle trajectory data, especially a method for generating privacy protection of trajectory data in the vehicle networking environment based on differential privacy technology and deep learning optimization method. Background Art
[0002] With the development of vehicle networking and intelligent transportation systems, more and more trajectory data is used in application fields such as traffic management, urban planning, and business analysis. These data usually contain a large amount of personal privacy information, such as the travel routes of users, departure and arrival locations, etc. Although the disclosure of these data is of great significance for system optimization and decision support, it also brings a great risk of privacy leakage. The trajectory data of users may disclose sensitive information such as personal living habits, residence, work location, etc., and in severe cases, it may even expose social relationships and health conditions, etc.
[0003] The existing privacy protection technologies are roughly divided into the following categories:
[0004] k-anonymity technology: This method merges the trajectory data of multiple users, making it impossible to associate a certain trajectory with a specific user. However, the k-anonymity technology may not be able to effectively resist data mining attacks when dealing with large-scale data, and it is difficult to maintain the usability of the data.
[0005] Differential Privacy (DP) technology: Differential privacy ensures the privacy by adding noise to the query result, ensuring that the addition or deletion of a single data will not significantly change the query result. However, the traditional differential privacy method is prone to cause a decrease in the accuracy of the generated trajectory. Especially in a complex traffic network, the generated trajectory may deviate greatly from the actual situation.
[0006] Deep learning methods: Deep learning can generate high-quality trajectory data in a large range, but these methods often lack built-in privacy protection mechanisms, and the generated trajectories cannot meet the strict requirements of differential privacy.
[0007] Therefore, how to balance privacy protection and the usability of trajectory data, especially generating real and privacy-protected trajectory data, has become a key challenge in the current technical field. Summary of the Invention
[0008] The purpose of the present invention is to provide a new method for protecting vehicle networking trajectory privacy based on differential privacy and deep learning. This method can generate high-precision and highly available trajectory data while protecting user privacy, and solve the problem of trajectory distortion caused by existing privacy protection methods.
[0009] The technical solution of the present invention proposes a grid-based trajectory generation method based on differential privacy by introducing a strategy that combines differential privacy and deep learning, and uses a fully convolutional neural network (FCN) to optimize the trajectory, ensuring privacy protection while improving the accuracy of the trajectory. The specific solution includes the following steps:
[0010] Step 1: Anonymous data collection;
[0011] Step 1.1: Initialization phase; Each user U i uses a key generation algorithm to generate a pair of public and private keys, and uploads the public key to the server; User U i generates a session key K session for encrypting the trajectory data to obtain encrypted data, encrypts the session key K session using the public key of the next-hop target user or server to obtain an encrypted session key, signs the encrypted data packet using the private key to obtain signed data, and then generates a data packet, which includes the encrypted data, the encrypted session key, and the signed data;
[0012] Step 1.2: Random forwarding phase; User U i generates a data packet and performs random forwarding. The intermediate user U j decrypts the session key of the data packet, re-encrypts and re-signs the data to generate a new data packet, and then continues random forwarding or directly uploads it; Finally, the data packet will pass through the final user U k and is sent to the server to end the forwarding process;
[0013] Step 1.3: Server processing phase; After receiving the data packet, the server decrypts the session key in the data packet using the private key, decrypts the trajectory data using the session key, verifies the data signature using the user's public key for the obtained trajectory data, and stores the data that passes the verification;
[0014] Step 2: Grid-based trajectory generation based on differential privacy;
[0015] Step 2.1: Divide the urban spatial area into dynamic grid cells, and further subdivide the high-density grids according to the trajectory point density to form multi-level grids;
[0016] Step 2.2: Construct a transition probability matrix between grids based on the Markov chain model, extract the transition patterns between grids by analyzing historical trajectory data, and obtain a preliminary transition probability matrix;
[0017] Step 2.3: Introduce Laplace noise to generate a noisy transition probability matrix that satisfies the differential privacy constraint;
[0018] Step 2.4: Generate a grid-based trajectory according to OD point pairs and the noisy transition probability matrix; where the OD point pairs are from the data collected through anonymous data collection and the statistical database;
[0019] Step 3: Optimize the trajectory generation with a fully convolutional neural network model; input the urban map image, the starting point of the trajectory, and the grid direction vector into the fully convolutional neural network, and output a trajectory path that conforms to the actual road network based on the grid-based trajectory;
[0020] Step 4: Evaluate the privacy protection and practicality of the generated trajectory.
[0021] Further, in the random forwarding stage, user U i sends the data packet directly to the server with a probability of 1 / x; or user U i forwards the data packet to another random user U j , j≠i; after the relay user U j receives the data packet, it follows the same forwarding probability as user U i .
[0022] Further, the transition probability matrix is represented by the following formula:
[0023]
[0024] where P(i,j) represents the transition probability from grid i to grid j, and ∑ j N i,j represents the number of trajectories from grid i to grid j;
[0025] Construct the state transition matrix based on the historical trajectory as follows:
[0026] M ij =Pr(T[i]=G j ∣T[i-1]=G i )
[0027] where M ij represents the probability of transferring from grid G i to grid G j , T is the trajectory path, and Pr represents the probability of transferring from grid G i to grid G j , given that the previous grid is G i ;
[0028] Introduce Laplace noise to the transition probability in the state transition matrix:
[0029]
[0030] where P noisy(i,j) represents the noisy transition probability, where Lap represents the Laplace noise function, λ is the noise amplitude, and ∈ is the privacy budget;
[0031] Finally, the noisy transition probability matrix is obtained:
[0032]
[0033] where Δf is the sensitivity.
[0034] Furthermore, the specific steps for generating the grid trajectory based on the OD point pairs and the noisy transition probability matrix are as follows:
[0035] Step 2.4.1: Initialize the start and end points of the trajectory: Let the start grid and end grid of the trajectory be respectively:
[0036] T syn [1] = G start , T syn [end] = G end
[0037] where T syn represents the preliminary grid trajectory, with the starting point being G start , and the end point being G end ;
[0038] Step 2.4.2: Randomly sample the trajectory length l from the length distribution L to determine the number of intermediate nodes of the trajectory path;
[0039] Step 2.4.3: Traverse each intermediate position i of the trajectory and perform the following steps:
[0040] For each candidate grid G cand in the grid set, calculate the transition probability prob1 from the current trajectory to G cand :
[0041] prob1 = Pr[T[i] = G cand |T[1]…T[i - 1]]
[0042] Calculate the probability prob2 from the current node to the end point G end :
[0043] prob2 = Pr[T[end] = G end |T[1]…T[i - 1], G cand
[0044] Assign the total probability of the candidate grid G cand as:
[0045] Pr(G cand ) = prob 1·prob 2
[0046] According to the calculated probability distribution, randomly sample from the candidate grids and select the current trajectory point G chosen ;
[0047] Update the trajectory:
[0048] T syn [i]=G chosen
[0049] Step 2.4.4: Repeat Step 2.4.3 until a complete gridded trajectory T is generated syn .
[0050] Furthermore, the fully convolutional neural network includes 3 parallel input branches: The first branch receives the urban map image, successively passes through a convolutional layer, a pooling layer, a convolutional layer, a pooling layer, and then passes through flattening to obtain a first one-dimensional vector; The second branch receives the image containing the position of the trajectory starting point, successively passes through a convolutional layer, a pooling layer, a convolutional layer, a pooling layer, and then passes through flattening to obtain a second one-dimensional vector; The third branch receives the direction vector of the trajectory; After being processed, the three branches are fused to obtain a unified feature representation, and finally the final trajectory path is output through a fully connected layer and a Sigmoid activation function.
[0051] Furthermore, the privacy protection and practicality of the trajectory generation in Step 4 are evaluated by one or several of the following indicators:
[0052] Trajectory query error: Evaluate the difference between the generated trajectory and the real trajectory, and calculate the query error using the following formula:
[0053]
[0054] where Q(Do) and Q(Ds) respectively represent the query results of the real trajectory and the generated trajectory, and b is the query threshold;
[0055] Path rationality: Use the map matching algorithm to evaluate the matching degree between the generated trajectory and the actual road network;
[0056] Frequent pattern retention: Calculate the difference in the support degrees of frequent patterns in the generated trajectory and the real trajectory through the frequent pattern mining algorithm;
[0057] Trajectory length distribution error: Compare the difference in the length distributions of the generated trajectory and the real trajectory;
[0058] Spatial coverage: Evaluate the spatial distribution and spatial coverage degree of the generated trajectory and the real trajectory.
[0059] The beneficial effects of the present invention are as follows:
[0060] 1. Privacy protection: By introducing differential privacy mechanisms and anonymous data collection methods, it ensures that private information of users will not be leaked during the generation and transmission of user trajectory data.
[0061] 2. Trajectory accuracy: Through the optimization of the FCN network, it solves the problem of rough and unreasonable paths generated by traditional differential privacy methods during trajectory generation, and provides synthetic trajectories highly consistent with actual trajectories.
[0062] 3. High practicality: Experimental results show that the trajectory data generated by the present invention is significantly superior to existing methods in terms of query error, path rationality, etc., and has high practical value.
[0063] 4. Wide applicability: This method can not only be applied in different urban environments, but also cope with the challenges of generating large-scale vehicle networking data, and has good adaptability and scalability. Brief Description of the Drawings
[0064] Figure 1 It is a flow chart of an anonymous data collection system.
[0065] Figure 2 It is a model diagram of an anonymous trajectory generation system based on differential privacy.
[0066] Figure 3 It is a structural diagram of a fully convolutional neural network (FCN) model.
[0067] Figure 4 It is a test diagram of trajectory generation by the FCN model.
[0068] Figure 5 It is a visualization diagram of real trajectories and generated trajectories.
[0069] Figure 6 It is a heat map of the spatial distribution of real trajectories and generated trajectories. Detailed Implementation Manner
[0070] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a method for protecting vehicle networking trajectory privacy based on differential privacy and deep learning of the present invention with reference to the drawings.
[0071] Step 1: Anonymous data collection
[0072] To solve the problem that users are reluctant to directly upload trajectory information to the server, the present invention designs an anonymous data collection mechanism to ensure anonymity and data integrity during the data collection process through encryption, signature and random forwarding. The whole process includes three stages: initialization, random forwarding and server processing, as Figure 1 shown.
[0073] Step 1.1: Initialization stage
[0074] Public and private key generation: For each user U i uses the key generation algorithm (KeyGen) to generate a pair of public and private keys (pk i , sk i ), and uploads the public key pk i to the server for subsequent data encryption and signature verification.
[0075] The user generates a session key K session for encrypting the trajectory data to ensure the security of data transmission.
[0076] Data encryption and signature: User U i encrypts its trajectory data and signs and packages the data packets, specifically including:
[0077] Trajectory data encryption: The user encrypts the trajectory data using the session key:
[0078] C Data = EncData(Data, K session )
[0079] where C Data is the encrypted trajectory data, EncData(·) is the symmetric encryption function, Data is the data to be transmitted, and K session is the session key.
[0080] Session key encryption: The user encrypts the session key K session using the public key of the target user (next hop) or the server:
[0081] C Key = EncKey(K session , pk target )
[0082] where C Key is the encrypted session key, EncKey is the function for encrypting the session key, and the session key is encrypted using the public key of the target user or the server, and pk target is the public key of the data recipient.
[0083] Data signature: The user signs the encrypted data packet using the private key to ensure data integrity and source verifiability:
[0084] SigData = SignData(C Data , sk i )
[0085] Among them, SigData is the signature data, SignData is the function for signing the encrypted data, which signs the data using the private key, C Data is the encrypted trajectory data, sk i is the private key of the current user.
[0086] Packet generation: The user generates a complete data packet Packet, including encrypted data, encrypted session key, and signature:
[0087] Packet = PacketGen(C Data , C Key , SigData)
[0088] Among them, PacketGen is the function for generating the data packet, which accepts three inputs: encrypted data C Data , encrypted session key C Key , and signature data SigData.
[0089] Step 1.2: Random forwarding stage;
[0090] During the data transmission process, user U i generates a data packet and performs random forwarding. The intermediate user U j decrypts the session key of the data packet, re-encrypts and re-signs the data, generates a new data packet, and then continues random forwarding or directly uploads it; finally, the data packet will be sent to the server through the final user U k , ending the forwarding process and ensuring anonymity. The specific steps are as follows:
[0091] 1) Random forwarding rule: Assume that the server needs to collect the trajectory data of n users: User U i sends the data packet directly to the server with a probability of 1 / x; or user U i forwards the data packet to another random user U j with a probability of (x - 1) / x (where j ≠ i).
[0092] Among them, x > 1 is a random parameter set by the system, and the user can adjust the size of x according to the system security requirements.
[0093] 2) Processing of intermediate users: When the intermediate user U j receives the data packet, it follows the same random mechanism as user U i : sends the data packet directly to the server with a probability of 1 / x; forwards the data packet to other random users U k with a probability of (x - 1) / x.
[0094] 3) End condition of forwarding: Since there is a probability of 1 / x for each forwarding to send the data to the server, after a finite number of forwardings, the data packet will eventually be sent to the server. This mechanism ensures the randomness of the data transmission path, and the server cannot trace the specific source of the data, thus achieving anonymous data collection.
[0095] 4) Schematic process: Initial user U i Generates a data packet and performs random forwarding. Intermediate user U j Decrypts the session key of the data packet, re-encrypts and re-signs the data, generates a new data packet, and then continues random forwarding or directly uploads it; finally, the data packet will pass through a certain user U k And is sent to the server to end the forwarding process.
[0096] Step 1.3: Server processing stage
[0097] After receiving the data packet, the server performs the following steps for processing:
[0098] 1) Data decryption: Use the private key to decrypt the session key in the data packet:
[0099] K session = DecKey(C Key , sk server )
[0100] Where DecKey is the function to decrypt the received encrypted session key C Key And it uses the server private key sk server To perform decryption.
[0101] Use the session key to decrypt the trajectory data:
[0102] Data = DecData(C Data , K session )
[0103] Where DecData is the function for the server to use the session key K session To decrypt the encrypted data C Data To obtain the original data Data.
[0104] 2) Signature verification: Use the public key pk i Of user U i To verify the data signature to ensure the integrity and authenticity of the data source:
[0105] Verify(C Data , SigData, pk i )
[0106] Where Verify is the verification function, pki is the public key of the sender, C Data is the data, and SigData is the signature data.
[0107] 3) Data storage: Store the data that has passed the verification to ensure the security and availability of the data.
[0108] Step 2: Grid-based trajectory generation based on differential privacy;
[0109] The goal of this step is to generate synthetic trajectories that conform to the statistical distribution and protect user privacy by introducing a differential privacy protection mechanism and combining a grid-based state transition matrix. The trajectory generation process includes two data sources: a real-time anonymous data collection mechanism and OD point pairs in a statistical database. The system model is as Figure 2 shown, and the specific steps are as follows:
[0110] Step 2.1: Divide the urban spatial area into dynamic grid cells, and further subdivide the high-density grids according to the trajectory point density to form multi-level grids; First, divide the target area (such as a city) into multiple grid cells (Grid), and each grid represents a spatial interval. The size of these grids is dynamically adjusted according to the trajectory data density. Specifically, the grid size GridSize is set by the following formula:
[0111]
[0112] where Area is the total area of the target area, and N is the number of grids; areas with higher density will be subdivided into smaller grids to ensure the accuracy of the data.
[0113] Subdivide the grids according to the density (Ni) of the trajectory data; for areas with higher density, divide the grids into more small grids, and the subdivision level is calculated by the following formula:
[0114]
[0115] where N i is the number of trajectory points contained in grid i, T is a balance factor that controls the upper limit of the density, this formula ensures the accuracy of trajectory generation in high-density areas, k i is the grid subdivision level.
[0116] Step 2.2: Construct the transition probability matrix between grids based on the Markov chain model, extract the transition patterns between grids by analyzing historical trajectory data, and obtain the preliminary transition probability matrix;
[0117] First, use the Markov chain model to simulate the natural transfer process of trajectories. The trajectory points within each grid depend on the state of the previous grid, and the transfer probability from the current grid to the next grid is calculated through the transfer probability matrix. The transfer probability matrix (P) is represented by the following formula:
[0118]
[0119] where P(i, j) represents the transfer probability from grid i to grid j, and ∑ j N i,j represents the number of trajectories from grid i to grid j.
[0120] Step 2.4: Construction of the state transition matrix M: Statistically analyze the transfer probabilities between grids from historical trajectory data to construct the Markov state transition matrix M, where:
[0121] M ij =Pr(T[i]=G j ∣T[i - 1]=G i )
[0122] where M ij represents the probability of transferring from grid G i to grid G j . T is the trajectory path, and Pr represents the transfer probability statistically analyzed between grids in historical trajectory data, referring to the probability of transferring from grid G i to grid G j , given that the previous grid is G i .
[0123] Step 2.3: Introduce Laplace noise to generate a noisy transfer probability matrix that satisfies differential privacy constraints;
[0124] To satisfy differential privacy constraints, Laplace noise is introduced to the probabilities in the state transition matrix M to obtain a noisy transfer probability matrix:
[0125]
[0126] where Lap represents the Laplace noise function, Δf is the sensitivity, and ∈ is the privacy budget.
[0127] Differential privacy protection: To ensure the privacy of the trajectory generation process, Laplace noise is added to each value of the transfer probability matrix, so as to ensure that the generated data cannot disclose the specific trajectories of users. The noise addition formula for each trajectory generation step is:
[0128]
[0129] where P noisy(i, j) represents the noisy transition probability, Lap represents the Laplace noise function, λ is the noise amplitude, and ∈ is the privacy budget.
[0130] Step 2.4: Generate a grid-based trajectory according to the OD point pairs and the noisy transition probability matrix; where the OD point pairs are from the data collected by the anonymous data collection and the statistical database;
[0131] Step 2.4.1: First, generate the starting point G of the trajectory start and the ending point G end , which have two sources:
[0132] 1) Real-time anonymous data collection mechanism: The starting point G start and the ending point G end come from the user trajectory data uploaded in the above anonymous data collection mechanism; this method is suitable for real-time data generation and can dynamically construct a real-time trajectory database to ensure the freshness and timeliness of the data.
[0133] 2) OD point pairs in the statistical database: Randomly sample based on the OD point pair probability distribution from the statistical database to select the starting point G start and the ending point G end ; this method can generate trajectory data that maintains statistical similarity with the original database and is suitable for offline trajectory generation tasks.
[0134] Step 2.4.2: Preliminary trajectory generation based on the Markov state transition matrix: According to the selected starting point G start and the ending point G end , use the noisy state transition matrix M' to generate a grid-based trajectory. The specific process is as follows:
[0135] Input: Grid G; length distribution L; noisy state transition matrix M'; starting point G start and the ending point G end
[0136] Output: Preliminary grid-based trajectory T syn
[0137] Among them, the steps of the trajectory generation algorithm are as follows:
[0138] Initialize the start and end points of the trajectory: Let the starting grid and the ending grid of the trajectory be respectively:
[0139] T syn [1] = G start , T syn [end] = G end
[0140] Determine the trajectory length: Randomly sample the trajectory length from the length distribution L Determine the number of intermediate nodes of the trajectory path.
[0141] Intermediate trajectory node generation: Traverse each intermediate position i of the trajectory (from 2 to ), and perform the following steps:
[0142] For each candidate grid G in the grid set cand , calculate the transition probability prob1 from the current trajectory to G cand :
[0143] prob1 = Pr[T[i] = G cand |T[1]…T[i - 1]]
[0144] Calculate the probability prob2 from the current node to the end point G end:
[0145] prob2 = Pr[T[end] = G end |T[1]…T[i - 1], G cand
[0146] Assign the total probability of the candidate grid G cand as:
[0147] Pr(G cand ) = prob 1 · prob 2
[0148] According to the calculated probability distribution, perform random sampling from the candidate grids and select the current trajectory point G chosen ;
[0149] Update the trajectory:
[0150] T syn [i] = G chosen
[0151] Complete trajectory generation: Repeat the above process until the complete grid-based trajectory T syn is generated.
[0152] Output the trajectory: Return the generated grid-based trajectory T syn .
[0153] Step 2.4.3: Construction of the trajectory database
[0154] Real-time trajectory database: Use the anonymous data collection mechanism to dynamically generate the starting point and the ending point, and continuously update and store the real-time grid-based trajectory.
[0155] Statistical trajectory database: Batch generate grid-based trajectories that satisfy the statistical distribution and meet the differential privacy constraints by statistically analyzing the OD point pairs and the state transition matrix in the database.
[0156] Step 3: FCN Model Optimization for Trajectory Generation;
[0157] To further improve the accuracy of the generated trajectory, the present invention uses a fully convolutional neural network (FCN) to optimize the grid-based trajectory. The FCN model maps the grid data of the trajectory to the actual geographical coordinates through a multi-layer convolutional network, realizing the conversion of the trajectory data from a coarse-grained grid to a high-precision actual geographical trajectory. The input of the FCN model includes the RGB image of the map, the starting point and direction information of the trajectory, and the output is the probability map of the trajectory. The model structure is as Figure 3 shown. Specifically, it includes three parallel input branches:
[0158] The first branch is the map image branch (map_input): This branch receives the RGB image of the city map (size 128×128×3), which is used to represent the geographical layout and environmental information of the city. Through the convolutional layer (Conv), different patterns in the image, such as roads, buildings, open areas, etc., are recognized. Then, through the pooling layer, the spatial size of the image is reduced to retain important features. Continuing with convolution, pooling, and then flattening into a one-dimensional vector for subsequent processing by fusing with other input data. The second branch is the starting point input branch (start_input): This branch receives a grayscale image of size 128×128×1, representing the starting point position of the trajectory. Each pixel value is usually 0 or 1, indicating whether it is the starting point position. Similar to the map input branch, it goes through convolution, pooling, convolution, pooling, and then flattening. The third branch is the direction input branch (direction_input): It inputs the direction vector (five-dimensional) of the grid in the grid-based trajectory corresponding to the starting point of the trajectory, guiding the movement direction of the trajectory (up, down, left, right, stop); specifically, the direction vector is the direction of the starting point relative to the ending point in the grid. These features are processed through convolutional and pooling layers, and finally flattened and fused into a unified feature representation. Finally, through the fully connected layer, the features are mapped to the probability map of the trajectory, representing the possibility of each pixel point being the trajectory path. The trajectory probability map generated by the FCN network is of size 128×128, representing the trajectory generation probability at each pixel position. The final trajectory path is output through the Sigmoid activation function.
[0159] When the FCN model generates the trajectory path, it first based on the starting point of the trajectory, corresponding to the grid where the starting point is located, obtains the direction vector of the grid. After being processed by the model, the ending point of this grid is obtained as the starting point of the next grid, and the path continues to be generated. After fusing with the city map, the final trajectory path is obtained.
[0160] As Figure 4The following shows the test data. After being trained with a small amount of data, when the model is given the specified starting point information, direction information, and map image input, the results generated by the model can already well indicate the possible trajectory paths of the data points. The goal of the model is to minimize the error between the generated trajectory and the true trajectory, and it is trained through the Binary Cross-Entropy Loss function:
[0161]
[0162] Among them, y i is the true trajectory label, and p i is the trajectory probability predicted by the model.
[0163] The trajectory data processed by the FCN model comes from the differential privacy protection mechanism and conforms to the post-processing property of differential privacy. The generated trajectory data will not disclose the private locations of users, so privacy can be guaranteed.
[0164] Step 4: Privacy protection and data availability evaluation; use the following metrics to evaluate the privacy protection and practicality of the generated trajectories:
[0165] (1) Trajectory query error: Evaluate the difference between the generated trajectory and the true trajectory, and calculate the query error using the following formula:
[0166]
[0167] Among them, Q(Do) and Q(Ds) respectively represent the query results of the true trajectory and the generated trajectory, and b is the query threshold.
[0168] (2) Path rationality: Use the map matching algorithm to evaluate the matching degree between the generated trajectory and the actual road network, and ensure that the generated trajectory conforms to the actual traffic network structure. As Figure 5 shown, this is the visualization graph of the results generated for a small trajectory dataset. It can be seen from the trajectory graph that the trajectories we generated conform to the road distribution and can also well reflect the trajectory density distribution of each section. As Figure 6 shown is the density heat map of the trajectory dataset. When we generate for a larger dataset, it can be seen that a new desensitized dataset that is statistically similar to the original trajectory dataset can already be generated.
[0169] (3) Frequent pattern retention: Through the frequent pattern mining algorithm (such as the Apriori algorithm), calculate the difference in the support degrees of the frequent patterns in the generated trajectory and the true trajectory to ensure that the generated trajectory can retain the common movement patterns of the true trajectory.
[0170] (4) Trajectory length distribution error: Compare the difference in the length distribution between the generated trajectory and the true trajectory;
[0171] (5) Spatial coverage rate: Evaluate the spatial coverage of the generated trajectory and the true trajectory.
[0172] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A vehicle network trajectory privacy protection method based on differential privacy and deep learning, characterized in that: The following steps are involved: Step 1: Anonymous data collection; Step 1.1: Initialization phase; each user U i Use the key generation algorithm to generate a pair of public and private keys, and upload the public key to the server; User U i Generate a session key K session , used to encrypt the trajectory data to obtain encrypted data, and the session key K session The encrypted data packet is signed using the public key of the next-hop target user or server to obtain the encrypted session key, and the encrypted data packet is signed using the private key to obtain the signature data, and then a data packet is generated, which includes the encrypted data, the encrypted session key and the signature data; Step 1.2: Random forwarding phase; User U i Generate data packets and perform random forwarding, transferring user U j The session key of the decrypted data packet is used to re-encrypt and re-sign the data packet, generate a new data packet, and then continue to forward it randomly or upload it directly; finally, the data packet will pass through the end user U k is sent to the server, ending the forwarding process; Step 1.3: Server processing stage; after receiving the data packet, the server uses the private key to decrypt the session key in the data packet, uses the session key to decrypt the trajectory data, verifies the data signature of the obtained trajectory data using the user's public key, and stores the verified data; Step 2: Gridded trajectory generation based on differential privacy; Step 2.1: Divide the urban space area into dynamic grid units, and further subdivide the high-density grid according to the density of trajectory points to form a multi-level grid; Step 2.2: Construct the transition probability matrix between grids based on the Markov chain model, extract the transition pattern between grids by analyzing the historical trajectory data, and obtain the preliminary transition probability matrix; Step 2.3: Introduce Laplace noise to generate a noisy transfer probability matrix that satisfies the differential privacy constraint; Step 2.4: Generate a gridded trajectory based on the OD point pairs and the noisy transition probability matrix; the OD point pairs come from the data collected from anonymous data and the statistical database; Step 3: Fully convolutional neural network model optimizes trajectory generation; input the city map image, trajectory starting point and grid direction vector into the fully convolutional neural network, and output the trajectory path that conforms to the actual road network based on the gridded trajectory; Step 4: Evaluate the privacy protection and practicality of the generated trajectory.
2. According to claim 1, a method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning is characterized in that: In the random forwarding phase, user U i Send the data packet directly to the server with a probability of 1 / x; or user U i Forward the packet to another random user U with probability (x-1) / x j , j≠i; transfer user U j After receiving the data packet, follow the user U i Same forwarding probability.
3. According to claim 2, a method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning is characterized in that: The transition probability matrix is expressed by the following formula: Where P(i,j) represents the transfer probability from grid i to grid j, ∑ j N i,j represents the number of trajectories from grid i to grid j; The state transfer matrix based on the historical trajectory is constructed as follows: M ij =Pr(T[i]=G j ∣T[i-1]=G i , Among them, M ij Represents the grid G i Transfer to Grid G j The probability of T is the trajectory path, Pr represents the probability of i Transfer to Grid G j probability; Introduce Laplace noise to the transition probability in the state transfer matrix: Among them, P noisy (i, j) represents the transition probability after adding noise, where Lap represents the Laplace noise function, λ is the noise amplitude, and ∈ is the privacy budget; Finally, the noisy transfer probability matrix is obtained: Where Δf is the sensitivity.
4. According to claim 3, a method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning is characterized in that: The specific steps of generating gridded trajectories based on OD point pairs and noise transfer probability matrix are as follows: Step 2.4.1: Initialize the start and end points of the trajectory: Let the start and end grids of the trajectory be: T syn [1]=G start ,T syn [end]=G end Where T syn Represents the initial gridded trajectory, starting at G start , the end point is G end ; Step 2.4.2: Randomly sample trajectory lengths from the length distribution L Determine the number of intermediate nodes of the trajectory path; Step 2.4.3: Traverse each intermediate position i of the trajectory and perform the following steps: For each candidate grid G in the grid set cand , calculate the current trajectory to G cand The transition probability prob1 is: prob1=Pr[T[i]=G cand ∣T[1]…T[i-1]] Calculate the distance from the current node to the end point G end The probability prob2 is: prob2=Pr[T[end]=G end ∣T[1]…T[i-1],G cand ] Assign the candidate grid G cand The total probability is: Pr(G cand )=prob 1·prob 2 According to the calculated probability distribution, random sampling is performed from the candidate grid to select the current trajectory point G chosen ; Update track: T syn [i]=G chosen Step 2.4.4: Repeat step 2.4.3 until a complete gridded trajectory T is generated. syn .
5. According to claim 4, a method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning is characterized in that: The fully convolutional neural network includes three parallel input branches: the first branch receives the city map image, passes through the convolution layer, the pooling layer, the convolution layer, the pooling layer in sequence, and then flattens to obtain the first one-dimensional vector; the second branch receives the image containing the starting point of the trajectory, passes through the convolution layer, the pooling layer, the convolution layer, the pooling layer in sequence, and then flattens to obtain the second one-dimensional vector; the third branch receives the direction vector of the grid in the gridded trajectory corresponding to the starting point of the trajectory, and the direction vector is specifically the direction of the starting point in the grid relative to the end point; After processing, the three branches are fused to obtain a unified feature representation, and finally the final trajectory path is output through a fully connected layer and a Sigmoid activation function.
6. According to claim 5, a method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning is characterized in that: The privacy protection and practicality of the trajectory generated in step 4 are evaluated by one or more of the following indicators: Trajectory query error: Evaluates the difference between the generated trajectory and the true trajectory, and calculates the query error using the following formula: Among them, Q(Do) and Q(Ds) represent the query results of the real trajectory and the generated trajectory respectively, and b is the query threshold; Path rationality: Use a map matching algorithm to evaluate how well the generated trajectory matches the actual road network; Frequent pattern preservation: The support difference of frequent patterns in the generated trajectory and the real trajectory is calculated through the frequent pattern mining algorithm; Trajectory length distribution error: compare the length distribution difference between the generated trajectory and the real trajectory; Spatial coverage: Evaluates the spatial distribution of generated trajectories and the spatial coverage of real trajectories.
7. The method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning according to claim 6 is characterized in that: The fully convolutional neural network is trained using a binary cross entropy loss function.
8. The method for protecting the privacy of vehicle network trajectory based on differential privacy and deep learning according to claim 7 is characterized in that: The multi-level grid is obtained by the following steps: The target area is divided into multiple grid cells, each grid represents a spatial interval, and the grid size GridSize is set by the following formula: Among them, Area is the total area of the target area, and N is the number of grids; The grid is subdivided according to the density of the trajectory data, and the level of subdivision is calculated by the following formula: Among them, N i is the number of trajectory points contained in grid i, T is the balance factor, k i The mesh subdivision level.
Citation Information
Patent Citations
Urban vehicle tracking method based on rapid region convolutional neural network
CN105868691A
Trajectory prediction method based on fusion inverse reinforcement learning
CN114445465A
Differential privacy-based vehicle trajectory data protection method and system, computer equipment and storage medium
CN115221557A
Track privacy protection method based on time increment
CN115828000A
Differential privacy trajectory data publishing method based on RNN (Recurrent Neural Network)
CN117113390A