Intelligent path navigation system based on deep learning and Q-learning algorithm

Through an intelligent path navigation system based on deep learning and Q-learning algorithm, the problem of path planning deviation in traditional navigation systems in complex urban traffic environments is solved, more efficient and accurate path planning is achieved, and user experience and traffic efficiency are improved.

CN120121071APending Publication Date: 2025-06-10FUJIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510277460.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When facing complex urban traffic environments, traditional navigation systems are difficult to accurately reflect traffic conditions, traffic signal control and dynamic changes in traffic conditions caused by weather, resulting in path planning deviations and unable to effectively alleviate traffic congestion.

Method used

An intelligent path navigation system based on deep learning and Q-learning algorithm is adopted, and through data acquisition, preprocessing, deep learning model construction and Q-learning algorithm path planning module, the path recommendations are updated in real time, taking into account factors such as traffic flow, traffic conditions, weather conditions and traffic lights.

Benefits of technology

It improves the timeliness and accuracy of path planning, effectively responds to unexpected traffic conditions, reduces traffic delays, provides more comprehensive intelligent path navigation services, and improves users' travel experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120121071A_ABST
    Figure CN120121071A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent path navigation system based on deep learning and a Q-learning algorithm, and the system comprises a data collection module which is used for collecting vehicle data, signal lamp data, traffic road condition data and meteorological data; the data preprocessing module is used for cleaning, integrating, normalizing and standardizing the data and generating a data set; the deep learning model construction module is used for training a long and short-term memory network through the data set generated by the data preprocessing module, constructing a traffic prediction model and obtaining a traffic prediction result in a period of time in the future; the Q-learning algorithm path planning module is used for comprehensively considering a plurality of influence factors and planning an optimal navigation path for the user; and the user interaction module is used for displaying an optimal navigation result to a user, collecting evaluation of the user and optimizing the deep learning model construction module and the Q-learning algorithm path planning module according to feedback data. According to the method, the advantages of deep learning and the Q-learning algorithm are combined, and the timeliness and accuracy of the intelligent path navigation system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly relates to an intelligent path navigation system based on deep learning and Q-learning algorithms. Background Art

[0002] With the acceleration of the urbanization process, traffic congestion has become the main factor restricting urban development, and solving urban traffic congestion is also an important part of realizing a strong transportation country. Traditional navigation systems generally perform path planning based on static map data, simple real-time traffic information, or historical road conditions. However, in the face of complex urban traffic environments, it is difficult to accurately reflect the dynamic changes in traffic conditions caused by traffic signal control and weather, resulting in deviations in the planned paths during actual driving, thus causing frequent urban traffic congestion.

[0003] In addition, traditional navigation systems still have deficiencies in the accuracy of predicting future traffic conditions, the timeliness of dynamic updates, and the provision of user personalized services, making it impossible to make reasonable driving routes in advance, and thus unable to effectively relieve traffic congestion. Deep learning technology can learn complex patterns and features from a large amount of data, providing high-precision predictions of multiple target factors. Q-learning can continuously interact with the environment and learn by designing complex reward functions and combining the prediction information provided by the deep learning model, and can gradually optimize the path planning strategy even in the case of insufficient data.

[0004] Therefore, considering various factors such as real-time traffic flow, traffic signals, traffic accidents, and weather conditions, developing an intelligent navigation system based on deep learning and Q-learning algorithms to overcome the limitations of traditional navigation systems, improve the timeliness and accuracy of path planning, and provide better personalized services for users is of great significance for improving urban traffic efficiency and relieving traffic congestion. Summary of the Invention

[0005] In view of the above problems, the present invention provides an intelligent path navigation system based on deep learning and Q-learning algorithms. This system can comprehensively consider various factors such as traffic flow, traffic conditions, weather conditions, and traffic signals, predict future traffic flow, weather conditions, and traffic signal switching states through deep learning technology, and calculate the optimal path in combination with the Q-learning algorithm, thereby improving the accuracy and efficiency of path planning.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0007] The present invention provides an intelligent path navigation system based on deep learning and Q-learning algorithms, including:

[0008] Data collection module, used to collect vehicle data, signal light data, traffic condition data and meteorological data;

[0009] The data preprocessing module is used to clean, integrate, normalize and standardize the collected data and then generate a data set;

[0010] The deep learning model building module is used to train the long short-term memory network through the data set generated by the data preprocessing module, build a traffic prediction model, and obtain traffic prediction results for a period of time in the future;

[0011] Q-learning algorithm path planning module, which is used to comprehensively consider multiple influencing factors and plan the optimal navigation path for users;

[0012] The user interaction module is used to display the optimal navigation results to users and collect user evaluations. The feedback data is used to optimize the deep learning model construction module and the Q-learning algorithm path planning module to continuously improve the accuracy of the navigation system and personalized services.

[0013] Furthermore, the vehicle data collection is to install vehicle sensors and video monitoring equipment, identify information such as license plate numbers or vehicle colors, and collect real-time statistics of vehicle traffic flow data passing through different road sections, wherein the real-time vehicle traffic flow data includes the number and speed of vehicles, etc.;

[0014] The signal light data collection is to install video monitoring equipment to record the status and switching frequency of traffic lights at each intersection in different time periods in real time during peak traffic hours;

[0015] The traffic condition data is to monitor the vehicle status data with the help of vehicle sensors in order to analyze the vehicle's driving conditions and road capacity; the vehicle status data includes acceleration, driving speed, etc.;

[0016] The meteorological data collection is to collect meteorological data of various areas of the city on a regular basis every day through meteorological sensors, and the meteorological data includes temperature, humidity, wind speed, etc.

[0017] Furthermore, the data cleaning process in the data preprocessing module includes: modeling the road section traffic by analyzing the topological relationship of multiple road intersections and their constituent sections; using sections with rich traffic flow data to supplement the missing traffic flow data of sections with strong correlation; using adjacent data to fill in the missing values ​​caused by sensor failure; deleting abnormal state data of vehicles; deleting or merging duplicate values ​​of collected data, etc.; it can maintain the continuity and integrity of the data, which is conducive to improving the accuracy of the data and the prediction performance of the deep learning model;

[0018] Data integration processing includes: integrating road section traffic flow and vehicle status data to obtain the vehicle driving speed of each road section; integrating traffic light status and switching frequency at different intersections to obtain the probability of a vehicle encountering a red light and the waiting time when encountering a red light;

[0019] Data normalization processing includes: normalizing data with different dimensions to the interval [0,1] through the maximum-minimum normalization method;

[0020] Data standardization processing includes: adopting the Z-score standardization method to make the mean of the real-time traffic flow data collected be 0 and the standard deviation be 1, which is convenient for model training.

[0021] Furthermore, the steps of constructing a traffic prediction model in the deep learning model construction module include data preparation, model construction, model training, and model verification; among them, the model uses LSTM, which can effectively process time series data and improve prediction accuracy;

[0022] The data preparation is to perform normalization processing on the numerical features in the dataset obtained by the data preprocessing module, perform one-hot encoding on data such as date and weather, and construct a multi-dimensional time series dataset, where 70% of the data is used as the training set, 20% of the data is used as the validation set, and 10% of the data is used as the test set;

[0023] The model construction includes defining the LSTM model structure, then using stacked LSTM units in multiple layers, adding a dropout mechanism between each layer of LSTM to prevent overfitting, and finally designing an attention mechanism layer to adjust the attention weight vector through the reward function of the Q-learning algorithm, so as to dynamically adjust each time point;

[0024] The model training is to use the training set data as input features, select the Stochastic Gradient Descent (SGD) optimization algorithm to train the model, and select an appropriate activation function according to the characteristics of the variables; in order to achieve a more accurate prediction effect, during the training process, the performance of the model is evaluated regularly using the validation set;

[0025] The model verification is to select the Root Mean Squared Error (RMSE) index to evaluate the performance of the model. The expression of the root mean square error RMSE is as follows:

[0026]

[0027] Among them, n is the number of samples, y i is the true value, is the predicted value. The smaller the RMSE, the higher the prediction accuracy of the model.

[0028] Furthermore, the Q-learning algorithm path planning module plans the optimal navigation path through the Q-learning algorithm;

[0029] Among them, planning the optimal navigation path through the Q-learning algorithm specifically includes: defining the state space, creating and initializing the Q-table, selecting and executing actions, and updating the Q-table;

[0030] The defining of the state space includes defining the current location, destination, current time, current weather, etc.;

[0031] The creating and initializing of the Q-table is to create a Q-table containing five arrays of the number of x coordinates, the number of y coordinates, the number of time periods, the number of weather conditions, and the number of actions, and initialize each element in the array to 0 for storing Q values;

[0032] The selecting and executing of actions is to select an action according to the current state using the ε-greedy strategy, randomly select an action with probability ε for exploration, select the action corresponding to the maximum Q value in the current Q-Table for this state with probability 1 - ε for application, and then execute this action in the current state to obtain a new state (s') and an immediate reward (r);

[0033] The updating of the Q-table includes: first defining a new reward function. Among them, introducing the vehicle travel time on the road section as part of the reward function can more accurately reflect the traffic efficiency of the road section. A shorter travel time means higher traffic efficiency, which helps to optimize route selection and traffic flow distribution. Introducing the probability of encountering a red light at an intersection can reduce the waiting time of vehicles in front of the traffic lights and improve traffic fluency. By optimizing the traffic light timing or providing prediction information to drivers, the number of stops can be reduced, fuel consumption and emissions can be reduced. Introducing the waiting time for red lights, as an indicator that can directly affect the driver's travel experience and the operating cost of the vehicle, reducing the waiting time can improve travel efficiency and reduce delays. Introducing the influence factor of weather can make the system adapt to traffic conditions under different weather conditions, adjust the driving speed and route selection when the weather changes to improve safety and traffic efficiency. Introducing the index of user preferences makes the system more personalized. Through the personalized reward function, a travel plan more in line with user expectations can be provided to meet the needs of different users. The expression of the reward function is: R(v, r, i, w, p, t) = -βT(v, r, t) - δP(v, i) - μW(v, i, t) - εM(w) + θU(p);

[0034] Among them, v represents the vehicle, r represents the road section, i represents the intersection, w represents the weather condition, p represents the user preference parameter, t represents the current time, R represents the reward function, T represents the vehicle driving time on the road section, P represents the probability of encountering a red light at the intersection, W represents the waiting time for the red light, M represents the influencing factor of the weather, U represents the user preference function, and β, δ, μ, ε, θ are the weights of T, P, W, M, U in the reward function respectively;

[0035] Then, update the Q value of the current state-action according to the update formula, which is expressed as:

[0036] Q(s,a)←Q(s,a)+α[R(s,a)+γmax a′ Q(s′,a′)-Q(s,a)];

[0037] Among them, Q(s,a) is the Q value of the current state-action; α is the learning rate; R(s,a) is the immediate reward obtained after executing the action; γ is the discount factor; max a′ Q(s′,a′) is the maximum Q value of all possible actions in the new state s';

[0038] Then update the current state to the new state, and finally repeat the above steps until the state is the destination entered by the user or reaches the threshold of the optimal target selected by the user.

[0039] Furthermore, the user interaction module includes a navigation page display module, a voice broadcast module, and a user evaluation module;

[0040] The navigation page display module and the voice broadcast module recommend the optimal route information to the user by inputting the starting point and the destination; the optimal route information includes the estimated driving time and road conditions of the recommended route, the estimated number of traffic lights to pass through and the waiting time for encountering a red light, weather influence, etc.;

[0041] The user evaluation module is used to collect the evaluation and feedback data of the user on the recommended route during actual navigation; the feedback data includes special situations such as sudden traffic accidents and road closures encountered; the intelligent path navigation system will re-plan the path according to the user feedback.

[0042] Furthermore, after the navigation ends, the user can feedback the satisfaction with the system by filling in text or selecting a score on aspects such as the accuracy, smoothness, and time saving of the system-recommended route.

[0043] Furthermore, by analyzing the user feedback data, the system optimizes system modules such as the deep learning model and the Q-learning path planning algorithm, can gradually adapt to the user's preferences and needs, improve the accuracy of path planning, and provide a navigation service that better meets the user's expectations, thereby reflecting the personalized service and user-friendliness of the system.

[0044] Further, the intelligent path navigation system collects real-time traffic data, road conditions, weather conditions, etc. through vehicle sensors, meteorological sensors, video monitoring devices, etc.; then preprocesses the collected data to form a data set; then constructs a deep learning model based on the data set, and predicts the traffic conditions in the next period of time based on the trained model; then uses the Q-learning algorithm for path planning, where the reward function incorporates multiple indicators such as the driving time of vehicles on the road section, the probability of encountering a red light at the intersection, the waiting time for the red light, the influence factors of the weather, and user preferences, with the goal of minimizing the total driving time and maximizing user satisfaction, so as to achieve real-time path optimization; finally, combining the model prediction and the algorithm recommendation path results, provides the user with the optimal path suggestion, and through the user feedback mechanism, collects the user's evaluation of the recommended path for optimizing and improving the model and algorithm.

[0045] The beneficial effects of the present invention are as follows:

[0046] The present invention adopts the above technical solutions to propose an intelligent path navigation system based on deep learning and Q-learning algorithm, which can comprehensively consider various factors such as traffic flow, traffic conditions, weather conditions, traffic lights, etc., use a deep learning model to predict the factors, and at the same time combine the Q-learning algorithm to update the path suggestion in real time, having the advantages of effectively coping with sudden traffic conditions, reducing traffic delays, providing a more comprehensive intelligent path navigation service, etc. The system also supports users to set different optimization goals according to their own preferences and common paths, improving the user's travel experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is the overall process structure diagram of the present invention;

[0049] Figure 2 It is the LSTM unit diagram of an embodiment of the present invention;

[0050] Figure 3 It is the Q-learning algorithm flow chart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0052] Refer to the attached Figure 1 As shown, the present invention provides an intelligent path navigation system based on deep learning and Q-learning algorithm, including: a data acquisition module, a data preprocessing module, a deep learning model construction module, a Q-learning algorithm path planning module, and a user interaction module;

[0053] The data acquisition module is used to collect vehicle data, signal light data, traffic condition data, and meteorological data; specifically including:

[0054] First, obtain map data information including road network, section speed limits, traffic signal positions, etc. through the API interface of map service providers; then install high-definition cameras on the roads and video monitoring devices at intersections or key urban sections to count the traffic flow and speed of vehicles and the red and green light status and switching frequency at each intersection, and at the same time analyze the videos captured by the cameras to identify the occurrence of traffic accidents and transmit them to the traffic management center in real time; then install acceleration sensors and speed sensors on the vehicles to monitor the acceleration and driving speed of the vehicles in real time to understand the driving conditions of the vehicles and the road passing capacity; finally, collect meteorological data such as temperature, humidity, and wind speed in each area of the city at regular intervals every day through various meteorological sensors on sections or buildings or meteorological stations distributed in different areas.

[0055] The data preprocessing module is used to perform data cleaning, integration, normalization, and standardization processing on the real-time data of the data acquisition module, and then generate a data set; specifically including:

[0056] First, data cleaning is performed. For the real-time vehicle traffic flow data, it is detected whether the vehicle speed on a certain section exceeds a certain percentage of the speed limit of that section. If so, it is regarded as an outlier and deleted, and the average speed of the section in the previous and subsequent time periods is used to approximately estimate the reasonable speed of the outlier. For the traffic light status and switching frequency data at intersections, it is detected whether the traffic light status remains unchanged for a long time or the switching frequency is too fast or too slow. If so, it is regarded as an outlier and deleted, and the historical data of the traffic light signals in the same time period is used for estimation. For the user's historical driving trajectory data, if there are missing values, the position information of the missing time period is inferred based on the driving speed and direction of the user in the previous and subsequent time periods. Then, data integration is carried out. The coordinate systems in the collected map data are unified into a common geographic coordinate system; the vehicle speed units collected by different sensors or video monitoring devices are unified into kilometers per hour; the traffic light status at intersections is converted into a digital coding format, for example, red light is 0, green light is 1, and yellow light is 2, which is convenient for data analysis. Next, data normalization is carried out. The data with different dimensions is normalized to the interval [0, 1]. By the maximum-minimum normalization method, the number of vehicles on a certain section is divided by the maximum number of vehicles that may appear on that section to obtain the normalized vehicle number value; the driving distance is divided by the total distance of the entire historical driving trajectory, and the driving time is divided by the total driving time to obtain the relative distance and time values; the traffic light switching frequency is divided by the maximum switching frequency within a certain time to obtain the relative switching frequency value. Finally, data standardization is carried out. For the vehicle speed in the real-time traffic flow data collected, the Z-score standardization method is used to make its mean 0 and standard deviation 1, which is convenient for model training.

[0057] The deep learning model construction module is used to train a long short-term memory network through the data set generated by the data preprocessing module, construct a traffic prediction model, and obtain the traffic prediction results for a future period of time; specifically including:

[0058] First, data preparation is carried out. The results obtained by the data preprocessing module are sorted and classified. The data such as time, date, and weather conditions are subjected to one-hot encoding processing. The date is converted into features such as day of the week, month, and quarter, the time is converted into features such as hour, minute, and second, and sunny, cloudy, rainy, snowy, etc. are represented by different vectors respectively; the traffic flow and historical traffic accident data are standardized and arranged according to the time series to form traffic flow sequences and historical traffic accident data sequences for different sections in different time periods. The sorted data is divided proportionally. 70% of the data is used as the training set, 20% of the data is used as the validation set, and 10% of the data is used as the test set. To avoid data deviation, in the above division, for the data in different time periods and different sections, it is necessary to divide according to their proportions in the overall data.

[0059] Then, model construction is carried out. The Pytorch deep learning framework is selected to design the LSTM model structure. The input layer receives multi-dimensional time series data including the current time and date, weather conditions, traffic flow, and historical traffic accident data, etc.; multiple LSTM layers are set to increase the depth and expressive power of the model; the fully connected layer integrates the outputs of the LSTM layers to further extract features, and the dropout mechanism is adopted to prevent overfitting. For the prediction of continuous variables such as traffic flow and travel time, a linear activation function is used; for the prediction of probability variables such as the probability of a traffic accident and the probability of encountering a red light, a Sigmoid activation function is used; for variables with a certain range, such as the encoded future weather conditions, a tanh activation function is used. The output layer outputs the predicted data, including the traffic flow on the road section, the travel time of the vehicle, the probability of a traffic accident, the probability of the vehicle encountering a red light and the waiting time at a specific intersection, and the future weather conditions. Different treatments are carried out for different output variables. For continuous variables such as traffic flow and travel time, the predicted values are directly output; for variables in the range of [0,1] such as the probability of a traffic accident and the probability of encountering a red light, a Sigmoid activation function is used for output; for the future weather conditions, a classification method can be used for output. For example, the weather is divided into categories such as sunny, cloudy, rainy, and snowy, and the Softmax activation function is used to output the probability of each category. At the same time, appropriate loss functions are selected according to different output variables. For example, the Mean Squared Error (MSE) loss function is used for the prediction of continuous variables; the Cross Entropy loss function is used for the prediction of probability variables; the categorical cross-entropy loss function is used for classification problems.

[0060] Then, model training is carried out. The stochastic gradient descent optimization algorithm is selected to train the model. During the model training process, the performance of the model is evaluated regularly using the validation set, and the optimal combination of hyperparameters is selected. Finally, model validation is carried out. The root mean square error is selected to evaluate the performance of the model, and its calculation process is as follows:

[0061]

[0062] where n is the number of samples, y i is the true value, is the predicted value. The smaller the RMSE, the higher the prediction accuracy of the model.

[0063] The Q-learning algorithm path planning module is used to comprehensively consider multiple influencing factors and plan the optimal navigation path for users; specifically including:

[0064] First, define the state space. Divide the urban road network into multiple grids, which are represented by two-dimensional coordinates. The center position of each grid is regarded as a state, and features such as the traffic flow, vehicle speed, road congestion, and distance to the destination at the current position are extracted for each state. Then, define the action space. Consider the path from state A directly to state B or the path from state A through multiple state turnovers to reach state B as an action, and features such as the path length, estimated travel time, and traffic congestion are extracted for each action. Next, design the reward function, comprehensively consider multiple factors such as travel time, traffic congestion, and traffic accident risk, and adjust and weight according to the actual needs of users to achieve different path planning goals. Finally, implement the Q-learning algorithm. Initialize the Q-value table, set algorithm parameters such as the learning rate and discount factor, iteratively update the Q-value according to the Q-value update formula, and adopt the ε-greedy strategy. Randomly select an action for exploration with probability ε, and select the current optimal action for exploitation with probability 1 - ε. Gradually decrease the value of ε as the algorithm progresses to increase the proportion of exploitation and improve the convergence speed of the algorithm. After the Q-value table converges, select the optimal path according to the Q-value table. Start from the starting state, select the action with the highest Q-value to reach the next state, and repeat this process until reaching the destination state. After finding the optimal path, further consider factors such as the vehicle's driving speed and fuel consumption according to the user's preferences, and adjust the reward function to select the optimal path. If the user believes that the vehicle's fuel consumption is high, select some relatively flat sections to reduce fuel consumption; if the vehicle's driving speed is restricted on a certain section, select some sections with higher speed limits to increase the driving speed.

[0065] As the algorithm progresses, gradually decrease the value of ε to increase the proportion of exploitation and improve the convergence speed of the algorithm. After the Q-value table converges, select the optimal path according to the Q-value table. Start from the starting state, select the action with the highest Q-value to reach the next state, and repeat this process until reaching the destination state. After finding the optimal path, further consider factors such as the vehicle's driving speed and fuel consumption according to the user's preferences, and adjust the reward function to select the optimal path. If the user believes that the vehicle's fuel consumption is high, select some relatively flat sections to reduce fuel consumption; if the vehicle's driving speed is restricted on a certain section, select some sections with higher speed limits to increase the driving speed.

[0066] The user interaction module is used to display the optimal navigation result to the user and collect the user's evaluation, and optimize the deep learning model construction module and the Q-learning algorithm path planning module with the feedback data; specifically including:

[0067] First, the system provides the user with detailed optimal path information, including the total estimated travel time from the starting point to the destination, the estimated traffic flow, traffic accident probability of each section along the way, the probability of encountering a red light and the waiting time at each intersection along the way, and the future weather conditions, etc. At the same time, the system also provides the user with various display methods. For example, use different colors to represent the traffic conditions of each section on the recommended path displayed on the map, and mark the waiting time of the traffic lights at the intersections, and provide detailed path text descriptions such as the turning instructions, traffic conditions, traffic accident occurrence frequency, and estimated travel time of each intersection.

[0068] Next is the user's evaluation and feedback on the navigation system. The user navigates according to the optimal path recommended by the system. After the navigation service ends, the user can give a text evaluation on aspects such as the accuracy, smoothness, and time-saving degree of the overall route, or choose a scoring method to express their satisfaction with the path. The system will continuously optimize the model and algorithm based on the user's feedback to improve the accuracy and personalized service of the navigation system. For special situations encountered during navigation, such as sudden traffic accidents and road closures, the user can also provide feedback through the system, and the system will adjust the path recommendation in real time according to the user's feedback. Through continuous interaction between the user and the system, the system can provide a path planning service that better meets the user's expectations, reflecting the user-friendliness of the navigation system.

[0069] Embodiment 1

[0070] This embodiment provides an intelligent path navigation system based on deep learning and Q-learning algorithm, including:

[0071] A data acquisition module for collecting vehicle data, signal light data, traffic condition data, and meteorological data; in this embodiment, the data acquisition module collects historical urban traffic flow data, including hourly weather statistical data, vehicle flow data collected every five minutes, traffic light status updated every minute, and vehicle position and speed information collected every 10 seconds, etc., covering data of multiple cycles.

[0072] A data preprocessing module for cleaning, integrating, normalizing, and standardizing the collected data, and then generating a data set; specifically including: modeling the traffic of road sections by analyzing the topological relationships of multiple road intersections and their constituent road sections; using road sections with rich traffic flow data to supplement the missing traffic flow data of strongly related road sections; using adjacent data to fill in the missing values caused by sensor failures; deleting abnormal state data of vehicles; deleting or merging duplicate values of the collected data, etc.; it can maintain the continuity and integrity of the data, which is beneficial to improving the accuracy of the data and the prediction performance of the deep learning model.

[0073] Integrate the traffic flow and vehicle status data of each road section to obtain the vehicle driving speed of each road section; integrate the traffic light status and switching frequency of different intersections to obtain the probability of a vehicle encountering a red light and the waiting time when encountering a red light; normalize the data with different dimensions to the [0, 1] interval through the maximum-minimum normalization method; adopt the Z-score standardization method to make the mean of the real-time traffic flow data collected be 0 and the standard deviation be 1, which is convenient for model training. In this embodiment, the data preprocessing module uses statistical methods to remove outliers, eliminate obviously incorrect traffic flow data and vehicle position information, check the integrity of the data, fill in missing values with historical data, integrate data sets from different sources into a unified time series data, ensure the consistency and integrity of the data, obtain historical urban traffic flow data and peak period and holiday data, and prepare the features required for the LSTM model to make predictions.

[0074] The deep learning model construction module is used to train a long short-term memory network through the data set generated by the data preprocessing module, construct a traffic prediction model, and obtain the traffic prediction results for a future period of time. In this embodiment, the LSTM model is used for prediction. LSTM (Long Short-Term Memory) is a commonly used recurrent neural network (RNN) model for modeling and predicting sequential data. The LSTM unit uses a memory cell to store the sequential information of the previous time period. Its network contains three key gating mechanisms: the forget gate, the input gate, and the output gate, which affect the operation of the sequential information. They jointly control the flow and update of information, allowing the model to selectively remember and forget the information in the input sequence.

[0075] Refer to the appendix Figure 2 As shown, the working principle of the LSTM model includes: the cell state C t is an information flow that runs through the entire sequence and is used to store long-term memory; the input gate calculates the output of the input gate through the sigmoid activation function and calculates the candidate cell state through the tanh activation function, which is responsible for determining which new information will be added to the cell state; the forget gate calculates the output of the forget gate through the sigmoid activation function and is responsible for determining which information will be discarded from the cell state; the output gate calculates the output of the output gate through the sigmoid activation function, and calculates the output of the cell state through the tanh activation function and element-wise multiplication, which is responsible for determining which part of the cell state will be output. Through the collaborative cooperation of these gating mechanisms, the LSTM can capture long-term dependencies in sequential data, avoid the problem of gradient vanishing or explosion, and thus perform well in processing sequential data.

[0076] Define the LSTM memory module value as C, the output as H, the input as X, and the factors affecting the calculation of the three gates include the memory module value C at the previous moment t-1 , the output value at the last moment is H t-1 and the current input X t , specifically, the calculation of the LSTM unit is as follows:

[0077] First, according to the formula Calculate candidate memory cells The tanh function is used as the activation function;

[0078] Then, according to formula I t =σ(X t W xi +H t-1 W hi +b i ) Calculate the input gate to weigh the impact of the current input on the real memory element;

[0079] Then, according to formula F t =σ(X t W xf +H t-1 W hf +b f ) Calculate the forget gate to evaluate the influence of the previous sequence information on the real memory element;

[0080] According to the formula Calculate the memory module value at the current moment, which is determined by the past memory element and the candidate memory element. The input gate and forget gate are used to measure the weights of these two parts.

[0081] Then according to formula O t =σ(X t W xo +H t-1 W ho +b o ) Calculate the output gate, which is used to adjust the output of the memory module;

[0082] Finally, according to the formula H t =O t ⊙tanh(C t )Calculate the output.

[0083] In the above formula, W xc is the weight matrix input to the candidate memory element, W xi is the weight matrix input to the input gate, W xf is the weight matrix input to the forget gate, W xo is the weight matrix input to the output gate, the dimensions of these weight matrices are: W xc , Wxi , W xf , W xi ∈R d×h ; W hc is the weight matrix from the previous hidden state to the candidate memory cell, W hi is the weight matrix from the previous hidden state to the input gate, W hf is the weight matrix from the previous hidden state to the forget gate, W ho is the weight matrix from the previous hidden state to the output gate. The dimensions of these weight matrices are: W hc , W hi , W hf , W ho ∈R h×h ; b c is the bias vector of the candidate memory cell, b i is the bias vector of the input gate, b f is the bias vector of the forget gate, b o is the bias vector of the output gate. The dimensions of these bias vectors are: b c , b i , b f , b o ∈R 1×h .

[0084] In this embodiment, the deep learning model construction module first obtains the processed result data from the data preprocessing module and organizes and classifies it according to different features. The time and date data are grouped into one category, the weather condition data are grouped into one category, the traffic flow data are grouped into one category, the historical traffic accident data are grouped into one category, etc. One-hot encoding is performed on the time, date, etc. data. After converting the date into features such as day of the week, month, quarter, etc., the day of the week is represented by a vector of length 7, where only the corresponding position is 1 and the rest are 0. For example, Monday is represented as [1, 0, 0, 0, 0, 0, 0]; the month is similarly represented by a vector of length 12; the quarter can be represented by a vector of length 4. For the weather condition, sunny, cloudy, rainy, snowy, etc. are represented by different vectors respectively. For example, [1, 0, 0, 0] represents sunny, [0, 1, 0, 0] represents cloudy, [0, 0, 1, 0] represents rainy, and [0, 0, 0, 1] represents snowy. Standardization processing is performed on the traffic flow and historical traffic accident data. For example, the Z-score standardization method is used to process the traffic flow data or the historical traffic accident occurrence frequency. The specific calculation is as follows:

[0085]

[0086] Among them, x is the original data, μ is the mean of the data, and σ is the standard deviation of the data. Arrange the processed various types of data in a time series. For traffic flow data, take each hour as a time period, and count the traffic flow of different sections to form time series data. Divide the sorted data proportionally, with 70% of the data as the training set, 20% of the data as the validation set, and 10% of the data as the test set. Then select the Pytorch deep learning framework. The input layer receives multi-dimensional time series data including the current time and date, weather conditions, traffic flow, and historical traffic accident data, etc. Set multiple LSTM layers to increase the depth and expressive ability of the model. The fully connected layer integrates the output of the LSTM layer to further extract features. At the same time, adopt the dropout mechanism to prevent overfitting, and set the dropout rate to 0.5, that is, randomly set the outputs of 50% of the neurons to 0 during each training. The output layer outputs prediction numbers including the traffic flow on the section, the driving time of the vehicle, the probability of a traffic accident occurring, the probability of the vehicle encountering a red light and the waiting time at a specific intersection, and the future weather conditions, etc. Then select the stochastic gradient descent optimization algorithm, set appropriate initial learning rates and momentum, use the training set to train the model, regularly use the validation set to evaluate the model performance, and adjust the hyperparameters according to the evaluation results. Finally, use the test set to calculate the root mean square error of the model to evaluate the performance of the model.

[0087] The Q-learning algorithm path planning module is used to comprehensively consider multiple influencing factors and plan the optimal navigation path for the user. Among them, the Q-learning algorithm process is as follows (as shown in the appendix Figure 3 shown):

[0088] First, perform initialization: Initialize the Q-Table that stores all state-action pairs, usually set to 0, and set learning parameters such as the learning rate, discount factor, and exploration rate;

[0089] Then select an action, observe the current state s, and select an action a according to the ε-greedy strategy; then execute the action a, and observe the next state s' and the immediate reward R obtained;

[0090] Next, update the Q-table, and at the same time perform a state transition on the current state s and update it to the next state s';

[0091] Finally, repeat the above steps until the state reaches the destination entered by the user or reaches the threshold of the optimal target selected by the user.

[0092] Specifically, it includes: defining the state space, creating and initializing the Q-table, selecting and executing actions, and updating the Q-table;

[0093] Defining the state space includes defining the current location, destination, current time, current weather, etc.

[0094] Create and initialize the Q - table: Create a Q - table that consists of five arrays for the number of x - coordinates, the number of y - coordinates, the number of time periods, the number of weather conditions, and the number of actions. Initialize each element in the arrays to 0 for storing Q - values.

[0095] Select and execute an action: Adopt the ε - greedy strategy to select an action according to the current state. With probability ε, randomly select an action for exploration, and with probability 1 - ε, select the action corresponding to the maximum Q - value in the current Q - Table for this state for application. Then execute this action in the current state to obtain a new state (s') and an immediate reward (r).

[0096] Update the Q - table: First, define a new reward function. The expression of the reward function is: R(v, r, i, w, p, t)= - βT(v, r, t)-δP(v, i)-μW(v, i, t)-εM(w)+θU(p);

[0097] where v is the vehicle, r is the road segment, i is the intersection, w is the weather condition, p is the user preference parameter, t is the current time, R is the reward function, T is the vehicle driving time on the road segment, P is the probability of encountering a red light at the intersection, W is the waiting time for the red light, M is the influencing factor of the weather, U is the user preference function, and β, δ, μ, ε, θ are the weights of T, P, W, M, U in the reward function respectively;

[0098] Then update the Q - value of the current state - action according to the update formula, which is expressed as:

[0099] Q(s,a)←Q(s,a)+α[R(s,a)+γmax a′ Q(s′,a′)-Q(s,a)];

[0100] where Q(s,a) is the Q - value of the current state - action; α is the learning rate; R(s,a) is the immediate reward obtained after executing the action; γ is the discount factor; max a′ Q(s′,a′) is the maximum Q - value of all possible actions in the new state s';

[0101] Then update the current state to the new state. Finally, repeat the above steps until the state is the destination entered by the user or reaches the threshold of the optimal goal selected by the user.

[0102] In this embodiment, the Q - learning algorithm path - planning module first represents the state space with a quadruple (x, y, t, w), where x and y are the coordinates of the current position, t is the current time, and w is the current weather; and represents an action with a triple {a 1 ,a 2 ,a 3} represents the action space, i.e., the direction the vehicle chooses at the next intersection, where a 1 represents going straight, a 2 represents turning left, a 3 represents turning right; by introducing elements such as the vehicle travel time T on the road section, the probability P of encountering a red light at the intersection, the waiting time W for the red light, the influencing factor M of the weather, and the user preference U, a new reward function is defined as R(v,r,i,w,p,t) = -βT(v,r,t) - δP(v,i) - μW(v,i,t) - εM(w) + θU(p). Then create a Q-table containing five arrays: the number of x coordinates, the number of y coordinates, the number of time periods, the number of weather conditions, and the number of actions. Initialize each element in the array to 0, and select the user's current location as the initial state. Then adopt the ε-greedy strategy to select actions and execute them. The next intersection obtained is used as the new state (s'), and the comprehensive reward calculated for factors such as travel time, red light waiting time, and weather influence is used as the immediate reward (r). Then, according to the update formula of the Q-learning algorithm Q(s,a) ← Q(s,a) + α[R(s,a) + γmax a′ Q(s′,a′) - Q(s,a)], where Q(s,a) is the Q value of the current state-action pair; α is the learning rate; R(s,a) is the immediate reward obtained after executing the action; γ is the discount factor; max a′ Q(s′,a′) is the maximum Q value of all possible actions in the new state s'. In this way, the current state can be updated to the new state (s'). Finally, repeat the above steps until the destination entered by the user is reached.

[0103] The user interaction module is used to display the optimal navigation results to the user and collect the user's evaluation, and optimize the deep learning model construction module and the Q-learning algorithm path planning module with the feedback data, continuously improving the accuracy and personalized service of the navigation system.

[0104] In this embodiment, the user interaction module inputs the starting point and destination according to the user in the navigation interface, selects the optimal target according to the user's preferences and common routes, and then the system provides the user with detailed optimal route information, including the estimated driving time, the traffic flow and accident frequency of the sections along the way, the number of red lights encountered at each intersection along the way and the estimated waiting time, the impact of future weather conditions on the driving time, etc. At the same time, the system provides display methods such as map display and text description. By using different colors on the map to represent the vehicle passing conditions of each section and marking the estimated waiting time of the traffic lights at the intersections, or providing detailed route descriptions such as the turning instructions, traffic conditions, accident occurrence frequency and estimated driving time of each intersection. Finally, the user navigates according to the recommended route of the system. After the navigation service ends, the user can evaluate the predicted indicators such as driving time, road conditions, and weather. The system will collect user feedback to continuously optimize the model and algorithm to improve the accuracy of the navigation system and user satisfaction.

[0105] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent path navigation system based on deep learning and Q-learning algorithm, characterized in that: include: Data collection module, used to collect vehicle data, signal light data, traffic condition data and meteorological data; The data preprocessing module is used to clean, integrate, normalize and standardize the collected data and then generate a data set; The deep learning model building module is used to train the long short-term memory network through the data set generated by the data preprocessing module, build a traffic prediction model, and obtain traffic prediction results for a period of time in the future; Q-learning algorithm path planning module, which is used to comprehensively consider multiple influencing factors and plan the optimal navigation path for users; The user interaction module is used to show the optimal navigation results to users and collect user evaluations, and use the feedback data to optimize the deep learning model construction module and the Q-learning algorithm path planning module.

2. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 1 is characterized in that: The vehicle data collection is to install vehicle sensors and video monitoring equipment to count the real-time data of vehicle traffic flow passing through different road sections in real time, and the real-time data of vehicle traffic flow includes the number and speed of vehicles; The signal light data collection is to install video monitoring equipment to record the status and switching frequency of traffic lights at each intersection in different time periods in real time during peak traffic hours; The traffic condition data is obtained by monitoring the vehicle status data with the help of vehicle sensors to analyze the vehicle's driving conditions and road capacity; the vehicle status data includes acceleration and driving speed; The meteorological data collection is to collect meteorological data of various areas of the city on a regular basis every day through meteorological sensors, and the meteorological data includes temperature, humidity, and wind speed.

3. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 1, characterized in that: The data cleaning process in the data preprocessing module includes: modeling the road section traffic by analyzing the topological relationship of multiple road intersections and their constituent sections; using sections with rich traffic flow data to supplement the missing traffic flow data of sections with strong correlation; using adjacent data to supplement the missing values ​​caused by sensor failure; deleting abnormal state data of vehicles; deleting or merging duplicate values ​​of collected data; Data integration processing includes: integrating the traffic flow and vehicle status data of the road section to obtain the vehicle speed of each road section; integrating the traffic light status and switching frequency of different intersections to obtain the probability of vehicles encountering red lights and the waiting time when encountering red lights; Data normalization processing includes: normalizing data of different dimensions to the interval [0,1] through the maximum and minimum normalization method; Data standardization processing includes: using the Z-score standardization method to make the mean of the collected real-time traffic flow data 0 and the standard deviation 1.

4. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 1, characterized in that: The steps of constructing a traffic prediction model in the deep learning model construction module include data preparation, model construction, model training and model verification; The data preparation includes normalizing the numerical features in the data set obtained by the data preprocessing module, performing one-hot encoding on the date and weather data, and constructing a multidimensional time series data set, in which 70% of the data is used as a training set, 20% of the data is used as a validation set, and 10% of the data is used as a test set; The model construction includes defining the LSTM model structure, then using multiple layers of stacked LSTM units, adding a dropout mechanism between each layer of LSTM to prevent overfitting, and finally designing an attention mechanism layer to adjust the attention weight vector through the reward function of the Q-learning algorithm, thereby dynamically adjusting each time point; The model training is to use the training set data as input features, select the stochastic gradient descent optimization algorithm to train the model, and select a suitable activation function according to the characteristics of the variables; during the training process, the validation set is regularly used to evaluate the performance of the model; The model verification selects the root mean square error (RMSE) indicator to evaluate the performance of the model. The expression of the root mean square error RMSE is as follows: Where n is the number of samples, y i is the true value, It is the predicted value. The smaller the RMSE is, the higher the prediction accuracy of the model is.

5. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 1, characterized in that: The Q-learning algorithm path planning module plans the optimal navigation path through the Q-learning algorithm; Among them, planning the optimal navigation path through the Q-learning algorithm includes: defining the state space, creating and initializing the Q table, selecting and executing actions, and updating the Q table; Defining the state space includes defining the current location, destination, current time, and current weather; The creating and initializing Q table is to create a Q table containing five arrays: the number of x-coordinates, the number of y-coordinates, the number of time periods, the number of weather conditions, and the number of actions, and initialize each element in the array to 0 for storing Q values; The selection and execution of the action is to select an action according to the current state using the ε-greedy strategy, randomly select an action for exploration with probability ε, select the action corresponding to the maximum Q value in the current Q-Table in this state with probability 1-ε for application, and then execute the action in the current state to obtain a new state (s') and immediate reward (r); The updating of the Q table includes: first defining a new reward function, the expression of the reward function is: R(v,r,i,w,p,t)=-βT(v,r,t)-δP(v,i)-μW(v,i,t)-εM(w)+θU(p); Where v is the vehicle, r is the road section, i is the intersection, w is the weather condition, p is the user preference parameter, t is the current time, R is the reward function, T is the vehicle driving time on the road section, P is the probability of encountering a red light at the intersection, W is the waiting time for the red light, M is the influencing factor of the weather, U is the user preference function, β, δ, μ, ε, θ are the weights of T, P, W, M, and U in the reward function respectively; Then the Q value of the current state-action is updated according to the update formula, expressed as: Q(s,a)←Q(s,a)+α[R(s,a)+γmax a′ Q(s′,a′)-Q(s,a)]; Among them, Q(s,a) is the Q value of the current state-action; α is the learning rate; R(s,a) is the immediate reward obtained after executing the action; γ is the discount factor; max a′ Q(s′,a′) is the maximum Q value of all possible actions in the new state s′; Then update the current state to the new state, and repeat the above steps until the state is the destination entered by the user or reaches the threshold of the optimal target selected by the user.

6. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 1, characterized in that: The user interaction module includes a navigation page display module, a voice broadcast module and a user evaluation module; The navigation page display module and the voice broadcast module recommend the optimal route information to the user through the user inputting the starting point and the destination; the optimal route information includes the estimated driving time and road conditions of the recommended route, the estimated number of traffic lights to be passed and the waiting time when encountering a red light, and the weather impact; The user evaluation module is used to collect the user's evaluation and feedback data on the recommended route during actual navigation; the feedback data includes unexpected traffic accidents and road closures encountered; the intelligent route navigation system will re-plan the route based on the user feedback.

7. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 6, characterized in that: After navigation, users can fill in text or choose to rate the accuracy, smoothness, and time saving of the system's recommended routes to provide feedback on their satisfaction with the system.

8. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 6, characterized in that: The system optimizes the deep learning model and Q-learning path planning algorithm module by analyzing user feedback data.

9. The intelligent path navigation system based on deep learning and Q-learning algorithm according to claim 1, characterized in that: The system collects real-time traffic data, road conditions and weather conditions through vehicle sensors, meteorological sensors and video surveillance equipment; The collected data is then preprocessed to form a data set; a deep learning model is then built based on the data set, and traffic conditions in the future are predicted based on the trained model; Then, the Q-learning algorithm is used for path planning. The reward function takes into account the vehicle travel time on the road section, the probability of encountering a red light at the intersection, the waiting time of the red light, the influencing factors of weather, and user preferences, with the goal of minimizing the total travel time and maximizing user satisfaction, thereby achieving real-time path optimization. Finally, the model prediction and algorithm recommendation path results are combined to provide users with optimal path suggestions, and through the user feedback mechanism, users' evaluations of the recommended paths are collected to optimize and improve the model and algorithm.

Citation Information

Cited By

  • Navigation method and system based on artificial intelligence

    CN120338236A