Traffic optimal path planning method based on sparse completion collaborative Q deep learning

Through the method based on sparse completion collaborative Q deep learning, the problem of difficult balance of sparse road conditions data processing efficiency and model accuracy in the prior art is solved, and the privacy protection of sparse traffic data is realized, improving the computing efficiency and accuracy of path planning.

CN119940678AActive Publication Date: 2025-05-06湖南工商大学
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510426169.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

When existing cloud edge-end collaborative path planning technology processes sparse road conditions data, it is difficult to balance model accuracy and computing efficiency, and there are privacy and security risks.

Method used

Using a method based on sparse completion collaborative Q deep learning, efficient processing and privacy protection of sparse traffic data are achieved through adaptive selection of weighted error control, transformer's traffic graph feature privacy protection, KNN weighted average sparse traffic graph feature matrix completion, and Q deep learning adaptive gradient protection.

Benefits of technology

While ensuring the privacy protection of sparse traffic data, efficient traffic optimal path planning is achieved, computing efficiency and model accuracy are improved, and privacy and security risks are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940678A_ABST
    Figure CN119940678A_ABST
Patent Text Reader

Abstract

The invention provides a traffic optimal path planning method based on sparse completion collaborative Q deep learning, and relates to the field of traffic path planning, comprising the following steps: sampling and selecting a sparse region based on a weighted error control-based key region adaptive selection method; according to the traffic map feature privacy protection method based on transform, traffic map features are efficiently extracted and protected; a sparse traffic map feature matrix complementing method based on KNN weighted average is used for complementing missing parts; a sparse traffic map pre-training-oriented Q deep learning adaptive gradient protection method is used for protecting model parameters; a sparse traffic flow optimal path planning method based on a Kalman filter is used for selecting an optimal path; according to the protection method adopted by the invention, # imgabs0 # local differential privacy is obtained, and high-availability optimal path planning is realized on the premise of ensuring privacy protection of sparse traffic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of traffic path planning, and in particular to a traffic optimal path planning method based on sparse completion collaborative Q deep learning. Background Technology

[0002] With the continuous acceleration of global urban development, the challenges faced by transportation networks are becoming more and more complex, especially in urban path planning. Path planning is a key issue in traffic management. In smart cities and intelligent transportation systems, how to achieve efficient and accurate path planning has become the focus of researchers and engineers. Traditional path planning methods such as Algorithms, such as the Dijkstra algorithm, often rely on graph search strategies. These algorithms have high computational complexity, especially when the search space is large or the map is very complex. As the number of nodes increases and traffic flow continues to change, the time overhead of path planning will increase significantly, which is difficult to meet actual needs for real-time or large-scale application scenarios.

[0003] In the existing technology, the emergence of cloud-edge-end collaborative technology provides a new solution to the path planning problem. Under this architecture, the terminal device monitors the traffic conditions in real time through sensors and data acquisition devices and uploads the data to the edge node. The edge node is responsible for model training with the collected data and works with the cloud server to jointly optimize the path planning model. This cloud-edge-end collaborative path planning architecture has significant advantages in improving traffic efficiency, reducing congestion, and saving time.

[0004] However, the path planning technology of cloud-edge-end collaboration still has limitations in practical applications. On the one hand, when the terminal device monitors the road conditions through sensors and data acquisition devices, due to the huge amount of road condition data, the terminal device cannot identify the importance of the path and is insensitive to abnormal road section data. If all road condition data is uploaded, the path planning overhead will increase significantly. If only sparse road condition data is uploaded, the model accuracy will be affected. On the other hand, when the path selection model is trained at the edge node, the path selection model of multiple vehicles often needs to consider the local optimality of individual vehicles and the collaborative efficiency of the global traffic system at the same time. Traditional path planning models usually face a trade-off between individual optimization and global optimization, and cannot effectively solve the negative impact that individual path selection may have on overall traffic efficiency. At the same time, cloud-edge-end collaboration also faces security issues. When the terminal device uploads road condition data to the edge node, and when the edge node and the cloud server collaborate to train the model, untrusted edge nodes and cloud servers may infer the private information of traffic users, which will bring serious privacy and security risks. SUMMARY OF THE INVENTION

[0005] The present invention provides a traffic optimal path planning method based on sparse completion and Q deep learning, which relates to the field of traffic path planning, including: a key area adaptive selection method based on weighted error control to select sparse areas; a traffic map feature privacy protection method based on transformer to efficiently extract and protect traffic map features; a sparse traffic map feature matrix completion method based on KNN weighted average to complete the missing parts; a Q deep learning adaptive gradient protection method for sparse traffic map pre-training to protect model parameters; a sparse traffic flow optimal path planning method based on Kalman filter to select the optimal path; the protection method adopted by the present invention obtains Local differential privacy, while ensuring the privacy protection of sparse traffic data, achieves high-availability optimal path planning.

[0006] Specifically, the first aspect of the present application provides a traffic optimal path planning method based on sparse completion collaborative Q deep learning, including the following steps: Step 1: Based on the adaptive selection method of weighted error control, sparse sampling is performed on the key areas, and the sparsity is controlled to ensure quality by calculating the change rate of adjacent sampling data; Step 2: Based on the transformer-based traffic map feature privacy protection method, obtain the traffic image feature matrix and perform differential privacy protection; Step 3: Sparse traffic map feature matrix completion method based on KNN weighted average, which completes the feature matrix by calculating missing data through weighted average; Step 4: Q deep learning adaptive gradient protection method for sparse traffic map pre-training, differential privacy protection of model parameters according to gradient change trend; Step 5: Optimize the model parameters and perform optimal path planning based on the Kalman filter for sparse traffic flow.

[0007] Further, the step 1 specifically includes: Step 1.1: According to the adaptive collection method, dynamic weights are given to different paths at different times, the traffic flow and weights of the paths are calculated and normalized, and the weights of the paths selected in the current cycle are accumulated and summed to obtain the overall weight. The formula is as follows: ; ; ; Among them, is the traffic flow of path z; is the traffic density of path z; is the average speed of traffic on path z; is the weight of path z; is the total number of paths in the selected area; is the traffic flow of path i; is the number of selected paths; is the overall weight; Step 1.2: Use the overall weight to perform weighted error control to reduce the interference of outliers on the overall optimization process. Calculate the data change rate between two adjacent cycles. When the data change rate is less than the specified threshold, stop the current round of sparse sampling. The formula is as follows: ; ; Stop sampling conditions: ; Among them, is the overall weight; is the difference between the actual path data and the predicted path data in the kth cycle; is the actual path data of the kth cycle; is the predicted path data of the kth cycle; is the maximum allowed error threshold; is the maximum tolerance of relative error; is the actual path data of the k-1th cycle; is the specified threshold.

[0008] Further, the step 2 specifically includes: Step 2.1: After sampling, extract the feature map of the sparse sampling area, and then use the convolution layer of CNN to process it to obtain the preliminary feature map; The preliminary feature map is further processed using convolution kernels with expansion rates of 1, 2, and 3 to obtain local features: ; ; Among them, is the local feature of the i-th image after MIN-MAX normalization processing; The fusion feature represented by multiplying each feature map in the i-th feature map by its weight; is the 3 feature maps that the i-th feature map is split into along the channel; is the weight of the three feature maps divided along the channel in the i-th feature map; is the maximum value function; is the minimum function; is the total number of pictures; is the weight of the first feature map that the Softmax result is split into along the channel dimension; is the weight of the second feature map that the Softmax result is split into along the channel dimension; is the weight of the third feature map that the Softmax result is split into along the channel dimension; Step 2.2: Feature Map The feature map in each image is segmented along n dimensions using 1×1 convolution, and n importance score maps are Calculation and segmentation are performed to collect the position set of key positions, and the key position set is weighted by features to obtain the final global features. The formula is as follows: ; ; Among them, is the key position set in the image; is the total number of pictures; is the number of key features; is the nth feature of the first feature map; is the first feature of the feature of the Mth feature map; is the nth feature of the feature of the Mth feature map; is the global feature of the i-th image; is the feature map at the i-th position; is the feature map at the nth position; It is a global maximum pooling operation, which applies the sigmoid function to compress the score to between 0 and 1, indicating the importance of each position;​ is the score of n important features in the first group of positions in the i-th image in the range of 0 to 1; is the score of n important features in the range of 0 to 1 in the nth group of positions in the ith image; Step 2.3: Concatenate the feature maps of global features and local features, and then use the transformer's self-attention mechanism to find the output feature matrix for each position. The formula is as follows: ; Among them, is the feature matrix of the i-th image; is the self-attention mechanism of transformer; is the query in transformer; is the key in transformer; is the value in transformer; is the local feature of the i-th image after MIN-MAX normalization processing; is the global feature of the i-th image; Step 2.4: For the feature matrix , from Gaussian distribution Generate a Dimensional noise matrix , the feature matrix is ​​protected by the Gaussian mechanism of differential privacy, and the noise feature matrix is ​​obtained , and then the noise matrix set after adding noise is obtained by proof The Gaussian mechanism that satisfies differential privacy. The formula is as follows: ; ; Among them, is the number of samples; is the characteristic number; is the result of privacy protection for the i-th feature matrix; is the feature matrix of the i-th image; is generated from a Gaussian distribution dimensional noise vector; ​is the i-th noise feature matrix, with latitude ; is the value of the mth row and jth column in the ith feature matrix; Noise added to the mth row and jth column.

[0009] Further, the step three specifically includes: Step 3.1: According to the KNN algorithm, for the noise feature matrix f(i)' containing the missing data x, select the K data closest to it in the feature matrix and calculate the weights of the K data respectively. The formula is as follows: ; Of which: is the weight of data K in the i-th traffic map feature matrix; is the function used to calculate the distance between data; is a parameter used to adjust the weight ratio and prevent the distance from being 0; is the missing data in the feature matrix; is the neighboring data of the missing data in the feature matrix; is the total number of pictures; Step 3.2: Accumulate and sum the weights to obtain the K feature matrix data closest to the missing data x in the i-th traffic map feature matrix and calculate the weight sum , missing data are estimated by weighted average algorithm, the formula is as follows: ; ; Of which: is the weight sum of the nearest feature matrix data; is the weight of data K in the i-th traffic map feature matrix; is the complement value of missing feature matrix data; is the value of the adjacent feature matrix data in the i-th traffic map feature matrix; is the total number of pictures; Step 3.3: Construct a complete traffic matrix based on the missing feature matrix data , the set of traffic map feature matrices is , which means as follows: ; Among them: is the set of completed traffic map feature matrix; is the feature matrix after the first traffic map is completed; is the feature matrix after the second traffic map is completed; is the feature matrix after the Mth traffic map is completed; is the total number of pictures.

[0010] Further, the step 4 specifically includes: Step 4.1: Collection based on the complete traffic map feature matrix Construct individual Q networks and local Q networks. The individual Q network generates the optimal path action, and the local Q network generates the optimal joint action. After executing the action, the transfer data is stored in the experience buffer to provide training data for subsequent network updates and training samples for network updates. The Q value update formula of the individual Q network is: ; The Q value update formula of the local Q network is: ; Of which: is the updated Q value of the individual Q network; Q value calculated for individual Q network; is the learning rate, which determines the step size of the Q value update; is the individual reward value; is the discount factor, The larger the value, the more emphasis the model places on long-term benefits and global collaboration, whereas it places more emphasis on immediate rewards; Select the action that maximizes the Q value in the next observation and next path selection action of the individual Q network; Select an action for the path of the jth vehicle; is the observation value of the individual Q network for the jth vehicle; is the updated Q value of the local Q network; Q value calculated for the local Q network; is the local reward value; Select the action that maximizes the Q value for the next state of all vehicles in the local Q network and the next local joint path selection action; is a local joint selection action; is the local state, i.e. the traffic environment information in the edge node area; For all possible paths, ; Select a route combination for all vehicles, ; In order to achieve the balance between individual optimality and local cooperation in path planning, the network collects training data through T iterations, gradually optimizes the individual Q network and the local Q network, and each vehicle uses the individual Q network to observe its value Process and select the optimal path action: ; Of which: From the path collection Select Make Maximum; Then the edge node summarizes the optimal path selection actions of all vehicles , combine them into joint actions , evaluate the joint path selection action through the local Q network and select the joint path solution with the greatest benefit: ; Of which: To collect actions from a joint Select an optimal joint action Make Maximum; Finally, execute the action and record the state transfer data and , stored in the experience buffer for network updates.

[0011] Step 4.2: From the buffer and Randomly sample small batches of data to calculate the loss function of the individual Q network and the local Q network loss function. The formula is as follows: ; ; ; ; ​​Among them, is the target value of the individual Q network; is the discount factor, The larger the value, the more emphasis the model places on long-term benefits and global collaboration, whereas it places more emphasis on immediate rewards; is the individual reward value; Select the action that maximizes the Q value in the next observation and next path selection action of the individual Q network; is the local Q network target value; is the local reward value; Select the action that maximizes the Q value among the next state of all vehicles in the local Q network and the next local joint path selection action; is the loss function of the individual Q network; To provide a buffer from individual experience Sampling data, updating individual Q network; It is an individual experience buffer, which stores the historical experience and state-action pairs of individual vehicles; The Q value calculated by the individual Q network, that is, the estimated value of the current vehicle choosing a certain action in a certain state; is the loss function of the local Q network; For local experience buffer Sampling data in, updating local Q network; It is a local experience buffer that stores historical experience and joint state-action pairs of multiple vehicles; Q value calculated for the local Q network; Step 4.3: Introduce the loss function based on the cosine similarity coordination mechanism, obtain the pre-training loss function, and optimize the model parameters through the gradient descent algorithm. The formula is as follows: ; ; Optimize model parameters through gradient descent algorithm: ; ; Among them, ​is the loss function based on the cosine similarity coordination mechanism; Q value calculated for individual Q network; Q value calculated for the local Q network; is the number of vehicles; is a fully connected layer that maps the Q value to a 128-dimensional vector space to facilitate the calculation of cosine similarity.

[0012] is the pre-training loss function; is the loss function of the individual Q network; is the loss function of the local Q network; is the weight of the cosine similarity coordination mechanism, The larger it is, the more emphasis is placed on the constraints of the coordination mechanism on the consistency of individual and global strategies; the smaller it is, the less likely it is that the coordination mechanism will interfere excessively with the loss function; is the gradient descent learning rate, which controls the step size of each parameter update; is the model parameter of the optimized individual Q network; is the model parameter of the individual Q network; is the gradient of the individual Q network model parameters; is the model parameter of the optimized local Q network; is the model parameter of the local Q network; is the gradient of the local Q network model parameters; Step 4.4: Predict the gradient of the t+1th round of model training through the second-order exponential smoothing mechanism norm, and then according to the gradient of the t+1th round and the tth round The norm is used to obtain the gradient change trend of the t+1th round, and then the adjustment factor is calculated , the formula for calculating the adjustment factor is as follows: ; Among them, is the adjustment factor for round t; is the gradient of the ti-th round Norm; ​is the gradient of the ti-1 round Norm; is the gradient change trend weight, the smaller i is, The bigger; Step 4.5: Adjust the gradient clipping threshold according to the adjustment factor. If the gradient change trend increases, the clipping threshold will be increased and more noise will be added. Otherwise, the clipping threshold will be lowered and less noise will be added. Finally, the noisy model parameter vector is generated through adaptive gradient Laplace noise. The formula is as follows: ; ; ; Among them, For the first The mean of the round adjustment factors; is the adjustment factor of the i-th model parameter in the t-th round; is the number of model parameters; is the pruning threshold of the tth round; is the pruning threshold for the t-1th round; is the model parameter vector of the tth round; is the model parameter vector after the tth round of noise addition; is the Laplace noise vector, Each value in the vector is the Laplace noise of each parameter, , is the clipping threshold of the t-th round gradient, is the privacy budget for each parameter.

[0013] Further, the step five specifically includes: Step 5.1: The cloud receives noise parameters from N edge nodes , the cloud uses the federated average algorithm to aggregate multiple edge models and update the global model parameters. The formula is as follows: ; Among them, is the global model parameter; is the number of edge nodes; represents the noise parameter of the i-th edge node; The obtained global model parameters When the optimal path is sent to the edge node for optimal path planning, the existence of noise may cause errors in the optimal path planning. Kalman filter is used to Perform smoothing and optimization processing, which includes prediction and update steps. The prediction step predicts the current state and error covariance based on the state at the previous moment.

[0014] Step 5.2: Use Kalman filter to smooth and optimize the global model parameters, including prediction step and update step. The formula in the prediction step is as follows: ; ; ; Among them, is the global model parameter for prediction; is the global model parameter of the previous moment; is the predicted error covariance matrix; is the error covariance of the previous moment; is the process noise covariance matrix; is a small constant to avoid prediction error covariance Infinitely decrease, maintain Kalman gain effectiveness; is the identity matrix.

[0015] Step 5.3: In the update step, the predicted state is combined with the noise parameter to correct the predicted value, and finally the global model parameters optimized by Kalman filter are obtained , the formula is as follows: ; ; ; ; Among them, is the Kalman gain; is the predicted error covariance matrix; is the observation noise covariance matrix, which is used to represent the noise uncertainty in the actual observation value; is the updated global model parameter; is the global model parameter for prediction; is the global model parameter; is the updated error covariance; is the identity matrix; For the first The clipping threshold of the wheel; Budget for privacy.

[0016] Step 5.4: Global parameters optimized by Kalman filter Used for model update, define noise loss function , to reduce the impact of noise on training, and then use the weighted loss function to comprehensively consider the pre-training loss function and noise loss function , optimize the above weighted loss function through the gradient descent algorithm: ; Weighted loss function: ; Gradient descent algorithm: ; ; Among them, is the noise loss function; Q value calculated for individual Q network; Q value calculated for the local Q network; is the regularization term of parameter difference, which is used to limit the optimized Deviating Aggregate Results the extent; is the regularization coefficient, The larger the value, the stricter the limit on the deviation of the parameter from the aggregation result. The smaller it is, the more flexible the parameters can be adjusted according to local dynamics; is the number of edge nodes; 、 are all weighted coefficients of the weighted loss function, The higher the value, the more the model focuses on improving the accuracy of path planning. The higher the value, the stronger the model is in suppressing the noise in the aggregation parameters; ​is the pre-training loss function; is the gradient descent learning rate, which controls the step size of each parameter update; is the model parameter of the optimized individual Q network; is the model parameter of the individual Q network; is the gradient of the individual Q network model parameters; is the model parameter of the optimized local Q network; is the model parameter of the local Q network; is the gradient of the local Q network model parameters; Step 5.5: After the noise parameters are processed by Kalman filter and the noise loss function is optimized, the updated model is used to calculate the individual optimal path for each vehicle, and then the local optimal path planning is optimized according to the local value network to obtain the overall optimal path. The formula is as follows: ; ; Among them, is the optimal path for an individual; Based on the updated model parameters From the path collection Choose an optimal path Maximize the individual Q value; is the local optimal path; Based on the updated model parameters local from the joint action set Select an optimal joint action Make Max.

[0017] In the second aspect, the present application further provides a computing device, which has the function of implementing the method described in the first aspect above. The beneficial effects can be found in the description of the first aspect and will not be described in detail here. The functions can be implemented by hardware or by executing corresponding software through hardware. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the structure of the device includes an acquisition module, a training module, and optionally, a construction module. These modules can implement the functions of the training node in the method example of the first aspect above. For details, please refer to the detailed description in the method example and will not be described here. ​

[0018] In a third aspect, the present application further provides a computing device, which is used to implement the functions of the method described in the first aspect above. The beneficial effects can be found in the description of the first aspect and will not be repeated here. The structure of the computing device includes a processor and a memory, and the memory is used to store instructions and / or data. The memory is coupled to the processor, and when the processor executes the program instructions stored in the memory, the function of the training node in the example of the first aspect above can be implemented. The structure of the computing device also includes a communication interface for communicating with other devices.

[0019] In a fourth aspect, the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are executed on a computer, the computer executes the method in the first aspect and various possible designs of the first aspect.

[0020] In a fifth aspect, the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method in the first aspect and various possible designs of the first aspect.

[0021] In a sixth aspect, the present application further provides a computing chip, the chip is connected to a memory, the chip is used to read and execute the software program stored in the memory, and execute the method in the above-mentioned first aspect and each possible implementation of the first aspect. Brief Description of the Figures

[0022] In order to more clearly illustrate the embodiments of the drawings or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the drawings. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without creative work.

[0023] Figure 1 is a flow chart of the steps of the present invention; Figure 2 is a specific operation flow chart of the present invention; Figure 3 A comparison chart of the path planning accuracy of the embodiment of the present invention and the traditional sparse data sampling method under different maximum error thresholds; Figure 4 A comparison chart of the path planning accuracy of the embodiment of the present invention and the traditional image privacy protection method under different image feature dimensions; Figure 5 A comparison chart of the path planning accuracy of the embodiment of the present invention and the traditional model parameter protection method under different privacy budgets; Figure 6The figure is a comparison of the path planning accuracy of the embodiment of the present invention and the traditional Q-learning path planning method under different gradient learning rates; The purpose, features and advantages of this figure will be further described in conjunction with the embodiments and with reference to the figures. Specific implementation method

[0024] In order to make the purpose, technical solutions and advantages of this application more clear, the following describes and illustrates this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. Based on the embodiments provided by this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0025] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as the content disclosed in this application is insufficient.

[0026] If there is no special description, all embodiments and optional embodiments of the present application can be combined with each other to form a new technical solution.

[0027] If there is no special explanation, all technical features and optional technical features of this application can be combined with each other to form a new technical solution.

[0028] If there is no special explanation, all the steps of the present application can be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), which means that the method may include steps (a) and (b) performed sequentially, or may include steps (b) and (a) performed sequentially. For example, the method may also include step (c), which means that step (c) may be added to the method in any order. For example, the method may include steps (a), (b) and (c), or may include steps (a), (c) and (b), or may include steps (c), (a) and (b), etc.

[0029] If there is no special explanation, the "include" and "comprising" mentioned in this application are open-ended or closed-ended. For example, the "include" and "comprising" may mean that other components not listed may also be included or only the listed components may be included.

[0030] If not specifically stated, in this application, the term "or" is inclusive. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, any of the following conditions satisfies the condition "A or B": A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).

[0031] In order to better understand the solution of the embodiment of the present application, some relevant terms and concepts that may be involved in the embodiment of the present application are first introduced below.

[0032] (1) Q network is a deep neural network that outputs the Q value of each possible action by inputting the environment state (such as sensor data or images), that is, the value estimate of the state-action pair. For example, in DQN, the input of the Q network is the state vector, the number of neurons in the output layer is equal to the dimension of the action space, and each neuron corresponds to the Q value of an action. It usually includes an input layer, a hidden layer (such as a fully connected layer, an activation function), and an output layer. For example, a fully connected layer stacked structure using ReLU activation is often divided into an online network (real-time update, used for action selection) and a target network (regularly synchronized parameters, used to calculate TD target values) in DQN to stabilize training.

[0033] (2) Loss function. In the process of training a deep neural network, we hope that the output of the deep neural network will be as close as possible to the value we really want to predict. Therefore, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict a lower value, and continue to adjust until the deep neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0034] (3) Kalman Filter is an efficient recursive estimation algorithm used to estimate the system state from noisy observation data in a dynamic system. Its core idea is to combine prediction (based on the system model) and update (based on the observation data) to achieve optimal estimation by minimizing the mean square error (MSE).

[0035] If Figure 1 As shown, in this embodiment, the optimal traffic path planning method based on sparse completion collaborative Q deep learning includes the following steps: Step 1: Based on the adaptive selection method of weighted error control, sparse sampling is performed on the key areas, and the sparsity is controlled to ensure quality by calculating the change rate of adjacent sampling data; Step 2: Based on the transformer-based traffic map feature privacy protection method, obtain the traffic image feature matrix and perform differential privacy protection; Step 3: Based on the KNN weighted average sparse traffic map feature matrix completion method, the feature matrix is ​​completed by calculating the missing data through weighted average; Step 4: Q deep learning adaptive gradient protection method for sparse traffic map pre-training, differential privacy protection of model parameters according to gradient change trend; Step 5: Based on the Kalman filter, the optimal path planning method for sparse traffic flow is used to optimize the model parameters and perform optimal path planning.

[0036] In this embodiment, the specific operation flow chart of the present invention is as follows Figure 2 As shown.

[0037] Further, step one specifically includes: Assume that in the kth cycle sampling, there are a total of x paths in a certain area, and their set is expressed as: ; Sampling selected y paths, and the set is represented as: ; Set Y is a subset of set X. The traffic flow of path z is calculated by the traffic density K and the average traffic speed V. The larger the traffic flow, the more important the corresponding path is in this area and this period. The normalized result is used as the weight value of path z.

[0038] Step 1.1: According to the adaptive selection method, dynamic weights are given to different paths at different times, the traffic flow and weights of the paths are calculated and normalized, and the weights of the paths selected in the current cycle are accumulated and summed to obtain the overall weight. The formula is as follows: ; ; ; Among them, is the traffic flow of path z; is the traffic density of path z; is the average vehicle speed of path z; is the weight of path z; is the total number of paths in the selected area; is the traffic flow of path i; is the number of selected paths; is the overall weight; Step 1.2: Use the obtained overall weight to perform weighted error control to reduce the interference of outliers on the overall optimization process, and then calculate the data change rate between two adjacent cycles. When the data change rate is less than the specified threshold, stop this round of sparse sampling. The formula is as follows: ; ; Sampling stop condition: ; Among them, is the overall weight; is the difference between the actual path data and the predicted path data in the k-th cycle; is the actual path data in the k-th cycle; is the predicted path data in the k-th cycle; is the maximum allowable error threshold; is the maximum tolerance of relative error; is the actual path data in the (k - 1)-th cycle; is the specified threshold.

[0039] Furthermore, Step 2 specifically includes: The transformer-based traffic map feature privacy protection method uses the convolution layer of CNN to perform preliminary processing on the M images in the collected samples to obtain the feature map of each image, and then uses convolution kernels with different expansion rates to extract local features from the feature map of each image. At the same time, the global features of the feature map are extracted by means of a key area scoring mechanism. The local features and global features are fused as the input of the transformer to obtain the feature matrix of each image. Finally, the feature matrices of all the images are subjected to noise processing and transmitted to the edge nodes for model training.

[0040] Step 2.1: The input image set is I={I(1), I(2), ..., I(M)}, where , H is the length, W is the width, the number of channels is 3, represents the image information of the i-th image. Directly performing convolution operations on the image may cause some information loss. Therefore, the image data is first supplemented by zero padding to obtain a new image.

[0041] After the image is processed by the convolution layer, a feature map is obtained, and then these feature maps are spliced. Subsequently, the spliced ​​feature map is reduced in dimension using a 1×1 convolution with a channel size of 3, and the feature map of the i-th image is obtained , which is expressed as: ; Of which: is the feature map of the i-th image; The j-scale feature map for the i-th image input; To perform a splicing operation on all feature maps in the jth image; Then the feature map collection , the convolution kernels with different expansion rates of 1, 2, and 3 are used to process the feature map set, and three feature maps of the i-th picture are decomposed from the i-th feature map, which are , , , then these three feature maps are applied with a 1×1 convolution with 3 channels, and then Softmax is used to obtain the general distribution of the three feature maps. Finally, the Softmax result is split into three different feature maps along the channel dimension, and each feature map corresponds to a weight w, so we can get , the formula is as follows: ; Afterwards, in order to enhance the comparability between image data indicators, MIN-MAX normalization is used to process the obtained output features and extract local features in the preliminary sparse sampling traffic feature matrix. .

[0042] The formula is as follows: ; Step 2.2: Feature Map The feature map in each image is segmented along n dimensions using 1×1 convolution, and n importance score maps are Calculation and segmentation are performed to collect the position set of key positions, and the key position set is weighted by features to obtain the final global features. The formula is as follows: ; ; Get global features: When detecting traffic signs, characteristic areas such as intersections, near schools, and construction areas can be used as key locations for global features. By strengthening the information collection of these key locations, the required global features can be obtained and the recognition of key features in the image can be enhanced. The present invention uses a global maximum pooling operation to determine the maximum value in each convolution kernel channel to return the importance score of the position (x, y) and the channel. Subsequently, a key location set with n key features is obtained based on the importance score, and the collected features are highlighted and enhanced based on this key location set.

[0043] Step 2.3: Concatenate the feature maps of global features and local features, and then use the transformer's self-attention mechanism to find the output feature matrix for each position. The formula is as follows: ; In the present invention, the traffic flow feature is specifically manifested as the real-time number of vehicles in each lane recorded in a column of the feature matrix.

[0044] Step 2.4: Define Sensitivity is a function Maximum of adjacent data sets Norm changes. Set adjacent data sets and Only if an element differs by at most 1, then the matrix The norm difference is , where For norm, so The sensitivity of is .

[0045] For vectors ,​ , and is an adjacent data set, the adjacent matrix is ​​limited to a single element difference of 1, given a privacy budget and failure probability , , Privacy Budget Controls the intensity of privacy protection. The smaller the value, the higher the privacy protection. The probability of failure represents the probability that the mechanism may violate the definition of differential privacy in extreme cases.

[0046] For the feature matrix , from Gaussian distribution Generate a Dimensional noise matrix , the feature matrix is ​​protected by the Gaussian mechanism of differential privacy, and the noise feature matrix is ​​obtained , the formula is as follows: ; ; Set a feature matrix , there is only one pair , so that , and for all , both , The corresponding feature matrix is .

[0047] To ensure the privacy of the feature matrix set uploaded to the edge node, it is necessary to theoretically prove that the designed Gaussian mechanism based on differential privacy satisfies Differential privacy conditions: Proof: Prove that the privacy protection mechanism used in this invention satisfies the differential privacy condition, that is, prove that the noise scale added by the Gaussian mechanism to each element in each feature matrix is ​​ , , , and It is the feature matrix with noise added corresponding to the matrices of two adjacent data sets.

[0048] Proof required . Order , for any feature matrix , for any measurable set , Noise Standard Deviation Need to meet , upper bound analysis of the probability ratio of the output distribution: ; where​ , and , .

[0049] Consider neighboring datasets and The noise output of is and , the difference . Due to And the noise is independent, the probability ratio can be further simplified to: ; Among them .

[0050] Define the privacy loss random variable: ; Because , the moment generating function decomposes the expectation and uses the properties of the Gaussian moment generating function to obtain: ; For any , during the conversion process, the content of The exponential term of is about Take the derivative and set it to 0, and simplify it, when Verifiable About The inequality of holds, so we can conclude that: ; The tail probability of the Gaussian distribution satisfies: ; When due to the properties of Gaussian distribution, for any , there are: ; Order , then: ; When the noise parameter Satisfaction , for any observable collection , there are: ; Therefore, it can be proved that the above privacy protection process satisfies the formula of differential privacy and has the function of privacy protection.

[0051] Further, step three, in order to complete the above traffic Figure 2dimensional feature matrix set, using the KNN-based weighted average interpolation method to predict missing data by using the similarity of known data. Compared with traditional interpolation methods, the KNN weighted average interpolation method gives different weights to neighboring samples, making the interpolation results more accurate and making full use of the similarity between data. It has obvious advantages when facing a large number of missing values, including: Step 3.1: According to the KNN algorithm, for the noise feature matrix f(i)' containing missing data x, select the K data closest to it in the feature matrix, and calculate the weights of the K feature matrix data respectively. The formula is as follows: ; Step 3.2: Accumulate and sum the weights to obtain the K feature matrix data closest to the missing data x in the i-th traffic map feature matrix and calculate the weighted sum , the missing feature matrix data is estimated by weighted average algorithm, the formula is as follows: ; (i=1,2,3...,M); Step 3.3: Construct a complete traffic feature matrix based on the missing feature matrix data , the set of traffic map feature matrices is , which means as follows: ; Furthermore, in step 4, a deep Q-learning framework is used to combine individual and local value networks, with the feature matrix completed in step 3 as input. The individual network optimizes the single-vehicle path strategy, the local network evaluates the multi-vehicle collaborative benefits, and balances individual and global interests through the cosine similarity coordination mechanism, and then uses the gradient descent algorithm to optimize the model parameters. At the same time, differential privacy protection (such as adaptive gradient noise) is applied to the model parameters to prevent the cloud from inferring sensitive data on the edge, ensuring the safe upload of the model.

[0052] Step 4.1: Construct individual Q-network and local Q-network and process vehicle observations With regional status , the individual Q network generates the optimal path action, the local Q network generates the optimal joint action, and then records the state transition data and stores it in the buffer to provide training samples for network updates, For the traffic map feature matrix set Contains observations for individual vehicles and local states within the region . For each vehicle, its observation information ​, including the coordinates of the current location, speed, path cost and traffic status and other basic data, after standardized processing, are input into the individual Q network to generate individual path selection strategy, local state Indicates the traffic environment information in the edge node area, which is input into the local Q network after standardized processing.

[0053] Two core network architectures are designed in this invention: individual Q network and local Q network.

[0054] The main task of the individual Q network is to estimate the individual value of each vehicle under different path selection actions, and dynamically adjust its path selection strategy according to the local environment of the vehicle, thereby improving the flexibility and accuracy of path planning. Its input is the observation of the vehicle and individual actions , the output is the value estimate of the individual's corresponding action, and its Q value is updated by the following formula: ; The task of the local Q network is to evaluate the overall value of the joint actions of multiple vehicles in the local area, further optimize the global collaboration efficiency, and avoid congestion caused by multiple vehicles choosing the same path. Its input is the local state and joint path selection actions , the output is the benefit estimate of the local joint path selection action, and its Q value is updated by the following formula: ; Then design the reward function of the two Q networks, individual reward value and local reward value , The design of individual reward value combines factors such as safety, efficiency, and smoothness. Safety reward is used to ensure basic driving safety, efficiency reward is used to improve driving efficiency, and smoothness reward is used to improve driving stability. The individual comprehensive reward formula is: ; 、 and is the weight coefficient of the three reward items, if Larger, more emphasis is placed on security rewards, if Larger, more emphasis is placed on efficiency rewards, if Larger, more emphasis is placed on smoothness rewards.

[0055] Safety Rewards It is to prevent vehicles from choosing paths that may lead to collisions or dangers, encourage vehicles to maintain a safe distance, and ensure driving safety. The safety reward is defined by the following formula: ; Of which: For safety rewards; is the minimum distance to surrounding vehicles; is the safety distance threshold; is a smoothing term to prevent the reward value from approaching negative infinity; Efficiency Rewards Aims to shorten travel time and improve traffic efficiency. By quantifying path length, average speed and waiting time, it guides vehicles to choose more efficient paths. The efficiency reward is defined by the following formula: ; Of which: Reward for efficiency; is the traffic density of path segment p; Smoothness Rewards Reduce the negative impact of road conditions on the vehicle. The straighter and flatter the path, the smoother the vehicle will travel. Therefore, the higher the reward, the better the driving experience. The formula is as follows: ; Of which: is a smoothness reward; is the curvature of path segment p; is the flatness of path segment p; and is the weight coefficient, if is larger, more attention is paid to reducing the impact of path curvature, otherwise, Larger ones place more emphasis on path flatness.

[0056] Local Reward Value The design of combines local traffic efficiency and congestion to optimize the coordination of multi-vehicle joint path selection. The formula is as follows: ; Of which: is the local reward value; and are the weight coefficients of the reward items respectively, if Bigger, more attention is paid to improving overall efficiency, on the contrary, Bigger ones focus more on reducing congestion costs; Efficiency Gain Reward represents the improvement of local path efficiency after the execution of joint action a, the formula is as follows: ; where: is the efficiency gain reward; represents the path length of vehicle j; is the average speed of vehicle j before the action; is the average speed of vehicle j after the action; Congestion cost reward is used to punish the traffic congestion that may be caused by the joint action. The formula is as follows: ; where: is the traffic density of section p; is the total number of sections of the sections within the local range; To achieve the balance between individual optimality and local cooperation in path planning, the network collects training data through T iterations, gradually optimizing the individual Q-network and the local Q-network. Each vehicle uses the individual Q-network to process its observation value and selects the optimal path action: ; Then the vehicle calculates the priority according to the surrounding traffic conditions. The vehicle with a higher priority will execute the path selection action that is optimal for the individual, and the vehicle with a lower priority provides candidate actions for the local Q-network to optimize the joint action to avoid local conflicts. The priority formula is as follows: ; where: is the priority of vehicle j; is the average speed of vehicle j before the action; is the traffic density of the surrounding vehicles; are weight parameters respectively. If is larger, more attention is paid to the vehicle speed. On the contrary, is larger, more attention is paid to the surrounding traffic density; When , the priority is set to high, otherwise it is set to low. Assume that the initially set is always greater than the of the current round, is always less than , then .

[0057] Then the edge node aggregates the optimal path selection actions of all vehicles and combines them to form a joint action Evaluate the joint path selection action through the local Q-network, and select the joint path scheme with the largest benefit: ; Finally, execute the action and record the state transition data and , and store it in the experience buffer for network update.

[0058] Step 4.2: Randomly sample a small batch of data from the buffer and to calculate the loss functions of the individual Q-network and the local Q-network. The formulas are as follows: ; ; ; ; To update the model parameters of this network, the network is updated by sampling data from the buffer through T iterations, with P learning steps executed each time.

[0059] Step 4.3: To avoid conflicts between the individual optimal paths of vehicles and the global traffic efficiency, design a coordination mechanism based on cosine similarity to force the individual strategy to be consistent with the global strategy direction, obtain the pre-training loss function, and optimize the model parameters through the gradient descent algorithm. The formulas are as follows: ; ; Optimize the model parameters through the gradient descent algorithm: ; ; Step 4.4: Obtain the adjustment factor through the norm prediction value of the gradient of the (t + 1)-th round of model training. For the norm prediction of the gradient of the model parameters in the (t + 1)-th round, the norms of the first 3 rounds of gradients of this parameter in the (t + 1)-th round can be used , , norm , , ; Predict the norm of the gradient of the (t + 1)-th round of model training through the second-order exponential smoothing mechanism, and then obtain the trend of the gradient change in the (t + 1)-th round based on the norms of the gradients in the (t + 1)-th round and the t-th round, and then calculate the adjustment factor , and the formula for calculating the adjustment factor is as follows: ​​; Step 4.5: Adjust the gradient clipping threshold according to the adjustment factor. If the gradient change trend increases, the clipping threshold will be increased and more noise will be added. Otherwise, the clipping threshold will be lowered and less noise will be added. Finally, the noisy model parameter vector is generated through adaptive gradient Laplace noise. The formula is as follows: ; ; ; In differential privacy, function sensitivity S(f) describes the maximum impact that a change in a single data point may have on the function output. The clipping threshold C is used to limit the gradient contribution of each data point. Therefore, the intensity of Laplace noise is proportional to the clipping threshold, and the clipping threshold is adjusted by adjusting the factor , if the gradient trend increases, the clipping threshold will be increased and more noise will be added, otherwise the clipping threshold will be lowered and less noise will be added.

[0060] Privacy budget per parameter and clipping threshold , generate the Laplace noise vector Y that satisfies differential privacy, and finally use adaptive gradient Laplace noise according to the formula: , generate the model parameter vector after adding noise .

[0061] To ensure the privacy of uploaded model parameters, it is necessary to theoretically prove that the designed adaptive gradient noise protection mechanism meets Differential privacy conditions.

[0062] Proof: Prove that the adaptive gradient noise protection method proposed in this invention satisfies Differential privacy condition, that is, to prove that the Laplace mechanism adds noise scale to the model parameters uploaded to each edge node of noise, satisfying .

[0063] Given any adjacent dataset and , with model parameters and Make , ; Then each model parameter noise satisfies Differential privacy, based on the serial principle characteristics of differential privacy, the designed adaptive gradient noise protection method satisfies Differential privacy conditions.

[0064] Further, step five specifically includes: Step 5.1: The cloud receives noise parameters from N edge nodes , the cloud uses the federated average algorithm to aggregate multiple edge models and update the global model parameters. The formula is as follows: ; Getting the global model parameters When sending it to the edge node for optimal path planning, the presence of noise may cause errors in the optimal path planning. We need to use Kalman filter to Perform smoothing and optimization processing, which includes prediction and update steps. The prediction step predicts the current state and error covariance based on the state at the previous moment.

[0065] Step 5.2: Use Kalman filter to smooth and optimize the global model parameters, including prediction step and update step, set the state transfer matrix and observation matrix as the identity matrix, the formula in the prediction step is as follows: ; ; ; Step 5.3: In the update step, the predicted state is combined with the noise parameter to correct the predicted value, and finally the global model parameters optimized by Kalman filter are obtained , the formula is as follows: ; ; ; ; Step 5.4: Global parameters optimized by Kalman filter , used for model update, define a noise loss function to reduce the impact of noise on training, and then use the weighted loss function to comprehensively consider the pre-training loss function and the noise loss function, and optimize the above weighted loss function through the gradient descent algorithm. The formula is as follows: ; Weighted loss function: ; Gradient descent algorithm: ; ; Step 5.5: After the noise parameters are processed by Kalman filter and the noise loss function is optimized, the updated model is used to calculate the individual optimal path for each vehicle, and then the local optimal path planning is optimized according to the local value network to obtain the overall optimal path. The formula is as follows: ; ; In this embodiment, the path planning accuracy comparison diagram of the embodiment of the present invention and the traditional sparse data sampling method under different maximum error thresholds is as follows Figure 3 As shown; the comparison of the path planning accuracy of the embodiment of the present invention and the traditional image privacy protection method under different image feature dimensions is shown in Figure 4 As shown in the figure; the comparison of the path planning accuracy of the embodiment of the present invention and the traditional model parameter protection method under different privacy budgets is shown in Figure 5 As shown; the comparison of the path planning accuracy of the embodiment of the present invention and the traditional Q-learning path planning method under different gradient learning rates is shown in Figure 6 As shown; By Figure 3~Figure 6 It can be seen that under different maximum error thresholds, different image feature dimensions, different privacy budgets, and different gradient learning rates, the path planning accuracy of the embodiment of the present invention is higher than that of the currently commonly used traditional methods.

[0066] It should be noted that the present application is not limited to the above-mentioned embodiments. The above-mentioned embodiments are only examples, and the embodiments having substantially the same structure as the technical idea and exerting the same effect within the scope of the technical solution of the present application are all included in the technical scope of the present application. In addition, within the scope of the subject matter of the present application, various modifications that can be thought of by those skilled in the art to the embodiments and other methods of combining some of the constituent elements in the embodiments are also included in the scope of the present application.​

Claims

1. A traffic optimal path planning method based on sparse completion collaborative Q deep learning, characterized by: The following steps are involved: Step 1: Based on the adaptive selection method of weighted error control, sparse sampling is performed on the key area, and the sparsity is controlled to ensure quality by calculating the change rate of adjacent sampling data; Step 2: Based on the transformer-based traffic map feature privacy protection method, the traffic image feature matrix is ​​obtained and differential privacy protection is performed; Step 3: A sparse traffic map feature matrix completion method based on KNN weighted average is used to complete the feature matrix by calculating the missing data through weighted average; Step 4: Q deep learning adaptive gradient protection method for sparse traffic map pre-training, differential privacy protection of model parameters according to gradient change trend; Step 5: Optimize the model parameters and perform optimal path planning based on the Kalman filter optimal path planning method for sparse traffic flow.

2. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 1 is characterized in that: The step 1 specifically includes: Step 1.1: According to the adaptive collection method, dynamic weights are given to different paths at different times, the traffic flow and weights of the paths are calculated and normalized, and the weights of the paths selected in the current cycle are accumulated and summed to obtain the overall weight; Step 1.2: Use the overall weight to perform weighted error control to reduce the interference of outliers on the overall optimization process, perform relative error limit, calculate the data change rate between two adjacent cycles, and stop this round of sparse sampling when the data change rate is less than the specified threshold.

3. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 1 is characterized in that: The step 2 specifically includes: Step 2.1: After sampling, extract the feature map of the sparse sampling area, and then use the convolution layer of CNN to process it to obtain the preliminary feature map; The preliminary feature map is further processed based on the convolution kernel to obtain local features; Step 2.2: Perform convolution segmentation on the feature map in each image to obtain the importance score map; The importance score graph is then calculated and segmented to collect the position set of key positions, thereby obtaining a key position set; The final global features are obtained by weighting the key position set; Step 2.3: Concatenate the feature maps of global features and local features, and then use the transformer’s self-attention mechanism to find the output feature matrix for each position; Step 2.4: Use the Gaussian mechanism of differential privacy to protect the privacy of the feature matrix and obtain the noise feature matrix .

4. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 3 is characterized in that: The convolution kernel uses convolution kernels with expansion rates of 1, 2, and 3 respectively.

5. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 3 is characterized in that: The step 2.2 is specifically to obtain n importance score maps by segmenting the feature map in each image along n dimensions using 1×1 convolution, and then calculating and segmenting the n importance score maps to collect the position set of key positions, and obtain the final global feature by feature weighting the key position set.

6. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 3 is characterized in that: The step 2.4 is specifically for the feature matrix , from the Gaussian distribution Generate a The noise matrix , the feature matrix is ​​protected by the Gaussian mechanism of differential privacy, and the noise feature matrix is ​​obtained .

7. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 1 is characterized in that: The step three specifically includes: Step 3.1: For the noise feature matrix containing missing data x , select the K feature matrix data closest to the missing data x, and calculate the corresponding weights of these K feature matrix data; Step 3.2: Accumulate and sum the weights to calculate the sum of the weights of the K feature matrix data closest to the missing data x in the i-th traffic map feature matrix , missing data were estimated by weighted average algorithm; Step 3.3: Construct a complete traffic feature matrix based on the missing feature matrix data , the set of traffic map feature matrices is .

8. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 1 is characterized in that: The step 4 specifically includes: Step 4.1: Collection based on the complete traffic map feature matrix Construct individual Q networks and local Q networks. The individual Q network generates the optimal path action, while the local Q network generates the optimal joint action. After executing the action, the transfer data is stored in the experience buffer to provide training data for subsequent network updates. Step 4.2: Sample data from the experience buffer, set the target value, and calculate the loss function of the individual Q network and the local Q network loss function; Step 4.3: Design a cosine similarity coordination mechanism to calculate the consistency loss of individual and local Q values , weighted integration into the pre-training loss , optimize parameters through gradient descent to balance the conflicts between individual and global strategies; Step 4.4: Predict the gradient of the t+1th round of model training through the second-order exponential smoothing mechanism norm, and then according to the gradient of the t+1th round and the tth round The norm is used to obtain the gradient change trend of the t+1th round, and then the adjustment factor is calculated. ; Step 4.5: Adjust the gradient clipping threshold according to the adjustment factor, and then generate the noisy model parameter vector through adaptive gradient Laplace noise.

9. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 8 is characterized in that: The gradient clipping threshold is adjusted according to the adjustment factor. Specifically, if the gradient change trend increases, the clipping threshold is increased to add more noise. Otherwise, the clipping threshold is decreased to add less noise.

10. The traffic optimal path planning method based on sparse completion collaborative Q deep learning according to claim 1 is characterized in that: The step 5 specifically includes: Step 5.1: The cloud receives the noise model parameters from N edge nodes , a federated averaging algorithm is used to aggregate multiple edge models and update the global model parameters; Step 5.2: Use Kalman filter to smooth and optimize the obtained global model parameters, including prediction step and update step; Step 5.3: In the update step, the predicted state is combined with the noise parameter to correct the predicted value, and finally the global model parameters optimized by Kalman filter are obtained. ; Step 5.4: Use the global parameters optimized by Kalman filter Used for model update, define noise loss function , to reduce the impact of noise on training, and then use the weighted loss function to comprehensively consider the pre-training loss function And the noise loss function , optimize the above weighted loss function through the gradient descent algorithm; Step 5.5: After the noise parameters are processed by the Kalman filter and the noise loss function is optimized, the updated model is used to calculate the individual optimal path for each vehicle, and then the local optimal path planning is optimized according to the local value network to obtain the overall optimal path.

Citation Information

Patent Citations

  • Federal learning-oriented privacy protection method

    CN117294469A

  • Longitudinal federal k neighbor feature completion method for privacy protection

    CN117932242A

  • Cloud side-end collaborative sparse traffic flow prediction privacy protection method and device

    CN119066712A

  • Community discovery method and device for multi-modal spatio-temporal correlation privacy protection

    CN119760451A

  • Multi-intelligence federal reinforcement learning-based vehicle-road cooperative control system and method at complex intersection

    US11862016B1