Unmanned vehicle real-time route planning and anomaly detection system based on deep Q network
By adopting a real-time route planning and abnormal detection system based on deep Q network in the unmanned vehicle navigation system, the problem of difficulty in dynamic adjustment and detection of road conditions in the existing system is solved, and a more efficient and safe driving of unmanned vehicles is achieved.
Patent Information
- Application Number
- CN202510117518.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The existing unmanned vehicle navigation systems lack dynamic adjustment capabilities, are difficult to adapt to real-time road conditions, and lack real-time detection and response mechanisms for abnormal road conditions, which pose safety hazards.
The real-time route planning and abnormal detection system for unmanned vehicles based on Deep Q Network (DQN) is adopted. By recording the length of time each unmanned vehicle passes through the road section in real time, dynamically calculate the optimal route, and promptly issue an alarm when road conditions are abnormal, and re-plan the route.
It significantly improves the driving efficiency and safety of unmanned vehicles, can dynamically adapt to environmental changes, timely avoid potential risks, reduce road network congestion, and provide safety guarantees.
Smart Images

Figure CN119984312A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation systems, and specifically relates to a real-time route planning and anomaly detection system for unmanned vehicles based on a deep Q network. Background Art
[0002] Existing unmanned vehicle navigation systems usually rely on preset maps and fixed route planning algorithms, lack dynamic adjustment capabilities, and are difficult to adapt to changes in real-time road conditions. When encountering congestion and emergencies, the existing system cannot adjust the route in time, increasing driving time and traffic pressure. At the same time, there is a lack of real-time detection and response mechanism for abnormal road conditions (such as accidents, fires, etc.), which poses a safety hazard. Summary of the invention
[0003] The purpose of the present invention is to provide a real-time route planning and anomaly detection system for unmanned vehicles based on a deep Q network, which can record in real time the length of time each unmanned vehicle passes through a road section, dynamically calculate the optimal route, and promptly issue an alarm when the road conditions are abnormal, prompting the re-planning of the route to improve driving efficiency and safety.
[0004] To achieve the above objectives, the technical solution of the present invention is: a real-time route planning and anomaly detection system for unmanned vehicles based on a deep Q network, which records the length of time each unmanned vehicle passes through a road section in real time, dynamically calculates the optimal route, and promptly issues an alarm when the road conditions are abnormal, and replans the route.
[0005] In one embodiment of the present invention, the system comprises:
[0006] The data collection and real-time analysis module records the time it takes each unmanned vehicle to pass through the section between two traffic lights and uploads it to the central server;
[0007] The optimal route planning module counts and searches all routes, calculates the total time of each route, and pushes the most time-saving route to the user;
[0008] The congestion detection and response module detects road congestion, sends an alert to unmanned vehicles about to enter the congested section, and replans the route;
[0009] The congestion cause identification module identifies the cause of the congested section and sends specific alerts to the unmanned vehicles that have entered the congested section;
[0010] The global coordination and optimization module realizes global traffic flow management and resource optimization to alleviate road network congestion.
[0011] In one embodiment of the present invention, the data collection and real-time analysis module uses vehicle-mounted sensors and V2X communication technology to collect GPS data and vehicle speed information of the unmanned vehicle in real time, store and dynamically update the average travel time of each road section. Specifically, GPS and speed sensors are used to record vehicle position and travel time in real time; data is uploaded to the central server through V2X communication technology; the central server uses Python language to process and analyze data, and dynamically updates the average travel time of each road section.
[0012] In one embodiment of the present invention, the optimal route planning module uses DQN to perform path optimization, uses the latest recorded time length of each road section to summarize the total time of the entire route, and recommends the most time-saving route.
[0013] In one embodiment of the present invention, the optimal route planning module is specifically implemented as follows:
[0014] 1) Initialization:
[0015] Initialize the urban road network environment; define the Q network and the target network; set hyperparameters, including learning rate, discount factor gamma, exploration rate epsilon, and batch size; initialize the experience replay buffer ReplayBuffer;
[0016] 2) Environment definition:
[0017] Environmental state State: It is represented by the current vehicle location, target location and the average travel time of all optional road sections. The environmental state can be represented by a vector S = {s1, s2, ..., s n}, where s j represents the average travel time of the jth road section, j∈[1,n], n is the total number of road sections in the road network;
[0018] Action space Actions: represents the next road segment that can be selected in the current state, defined as A = {a1, a2, ..., a m}, where a i Indicates the i-th road segment selected as the next action; m is the number of road segments that can be selected in the current state;
[0019] Reward function: The reward function is used to guide the vehicle to choose the optimal path. If a shorter route is chosen, the reward value is higher; the formula is R = -T (a i )+λC(s,a i ), where T(a i ) is the time to pass the road section, C(s,a i) is the negative increment of the target distance, λ is the coefficient of the weight between the balance time and distance, and s represents the average travel time vector of all optional sections in the current environment state;
[0020] 3) DQN training
[0021] The update formula of the Q-value function is Q(s,a)=Q(s,a)+α[r+γmaxQ(s',a')-Q(s,a)], where Q(s,a) is the Q-value of executing action a in the current state s, α is the learning rate, γ is the discount factor, r is the immediate reward for executing the action, and s' is the new state transferred to after executing action a; maxQ(s',a') represents the maximum Q-value of all possible actions a' under the new state s', which represents the best expected return in the new state.
[0022] 4) Model optimization
[0023] Using the experience replay mechanism, the state transition samples are stored as (s, a, r, s'), and training is performed by randomly sampling small batches of data to reduce time correlation; the loss function uses Huber loss:
[0024]
[0025] Where y reflects the long-term value of the agent's choice of action in the current state. y = r + γmaxQ (s', a') means that in state s', after taking action a', the immediate reward r that can be obtained plus the discounted return that may be obtained in the future.
[0026] 5) Dynamic programming strategy
[0027] In actual navigation, DQN predicts the current optimal action a * , the real-time update path is a * =argmaxQ(s,a).
[0028] In one embodiment of the present invention, the congestion detection and response module sets a road section time threshold. When the length of time that a certain unmanned vehicle passes through a road section exceeds the road section time threshold, the corresponding road section is determined to be congested, and an alarm is sent to the unmanned vehicle system through the Internet of Vehicles to re-plan the route. The specific implementation is as follows:
[0029] Set the segment time threshold T thresh , when the time T(a) for a vehicle to pass through section a satisfies T(a)>T thresh It is determined that road section a is congested;
[0030] When congestion occurs on road section a, an alarm message is sent to vehicles about to enter the congested section through V2X communication technology;
[0031] After the vehicle receives a congestion alert, it recalculates the optimal path and adjusts the driving direction based on the optimal route planning module.
[0032] In one embodiment of the present invention, the congestion cause identification module uses a deep learning model combined with the vehicle-mounted camera and sensor data to identify abnormal road conditions and send specific alarm information; the specific implementation is as follows:
[0033] Use on-board cameras to collect real-time image data and combine sensors to detect abnormal events on the road;
[0034] The use of deep learning models to detect road anomalies is expressed as P(cI) = f(I; θ), where I is the input image (visual information of the current road section), c is the category label (the identified cause of congestion, such as accidents, fires, etc.), and θ is the model parameter (which can effectively extract features from the input image and classify them after training); the model determines the specific cause of congestion based on the real-time image information of the road section, thereby providing timely warnings and decision support for vehicles. The deep learning model outputs the category c with the highest probability * As an abnormal cause;
[0035] Send specific information about the cause of congestion to vehicles.
[0036] In one embodiment of the present invention, the global coordination and optimization module uses DQN to implement dynamic route adjustment and intelligent scheduling of unmanned vehicles, and continuously optimizes system strategies based on historical data.
[0037] In one embodiment of the present invention, the global coordination and optimization module is specifically implemented as follows:
[0038] (1) Traffic flow modeling
[0039] Construct a global road network graph G = (V, E), where V is the set of road segment nodes and E is the set of edges connecting road segments; the road segment flow is represented by f(e), e∈E;
[0040] (2) Optimization goal
[0041] The optimization goal is to minimize the global average travel time. Where T(e) is the travel time of road section e;
[0042] (3) Dynamic Scheduling
[0043] Based on the DQN model, vehicle scheduling is adjusted in real time to avoid too many vehicles concentrating on the same road section. By setting a threshold v for the number of vehicles, when the number of vehicles on a certain road section exceeds the threshold, the system will automatically adjust the allocation of vehicles on other sections to reduce congestion on that section;
[0044] (4) Strategy Update
[0045] Update traffic optimization strategies based on historical data and real-time traffic data to improve overall efficiency.
[0046] In one embodiment of the present invention, the method is applied to the fields of complex urban road networks, logistics distribution, and emergency rescue.
[0047] Compared with the prior art, the present invention has the following beneficial effects: The present invention combines the DQN algorithm with machine learning technology to design and implement a real-time route planning and anomaly detection system for unmanned vehicles. During the driving process of the unmanned vehicle, the system can collect and analyze road condition information in real time, dynamically calculate the optimal driving route based on DQN, and timely avoid potential risks through congestion detection and anomaly identification modules, thereby significantly improving driving efficiency and safety. At the same time, the present invention has the following outstanding advantages:
[0048] 1) Dynamic adaptability. Through the DQN algorithm, the system can perceive environmental changes in real time and adjust the route to ensure that the unmanned vehicle always chooses the optimal path.
[0049] 2) Efficient resources. By leveraging the global coordination and optimization module, the system can effectively reduce overall road network congestion and optimize traffic resource allocation.
[0050] 3) Safety assurance. By combining multiple sensors with deep learning models, the system can quickly identify abnormal road conditions (such as accidents, fires, etc.) and issue targeted alarms to effectively avoid safety hazards.
[0051] 4) Wide application scenarios. This system is suitable for many fields such as complex urban road networks, logistics distribution, emergency rescue, etc., especially for navigation and path planning needs in dynamic and complex environments.
[0052] In summary, the present invention not only improves the driving efficiency and safety of unmanned vehicles, but also provides important technical support for the development of future intelligent transportation systems, and has broad application prospects and huge commercial value. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flow chart of the system of the present invention. DETAILED DESCRIPTION
[0054] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0055] The present invention provides a real-time route planning and anomaly detection system for unmanned vehicles based on a deep Q network, which records the length of time each unmanned vehicle passes through a road section in real time, dynamically calculates the optimal route, and issues an alarm in time when the road condition is abnormal to replan the route. The system includes:
[0056] The data collection and real-time analysis module records the time it takes each unmanned vehicle to pass through the section between two traffic lights and uploads it to the central server;
[0057] The optimal route planning module counts and searches all routes, calculates the total time of each route, and pushes the most time-saving route to the user;
[0058] The congestion detection and response module detects road congestion, sends an alert to unmanned vehicles about to enter the congested section, and replans the route;
[0059] The congestion cause identification module identifies the cause of the congested section and sends specific alerts to the unmanned vehicles that have entered the congested section;
[0060] The global coordination and optimization module realizes global traffic flow management and resource optimization to alleviate road network congestion.
[0061] The following is the specific implementation process of the present invention.
[0062] like Figure 1 As shown, the present invention provides a real-time route planning and anomaly detection system for unmanned vehicles based on a deep Q network, which specifically includes the following modules:
[0063] 1. Data collection and real-time analysis module
[0064] Function: Record the length of time each driverless car takes to pass through the section between two traffic lights in real time and upload it to the central server.
[0065] Implementation: Use vehicle-mounted sensors and V2X (Vehicle-to-Everything) communication technology to collect vehicle GPS data and speed information in real time, store and dynamically update the average travel time for each road section.
[0066] 2. Optimal route planning module
[0067] Function: Statistically search all routes, calculate the total time of each route, and push the route that saves the most time to the user.
[0068] Implementation: Use DQN for path optimization, use the latest recorded time length of each road segment to summarize the total time of the entire route, and recommend the most time-saving route.
[0069] 3. Congestion detection and response module
[0070] Function: Detect road congestion, send alerts to unmanned vehicles about to enter the congested section, and re-plan the route.
[0071] Implementation: Set a time threshold. When the time it takes for a driverless car to pass through a road section exceeds the threshold, the road section is considered congested. Send an alarm to the driverless car system through the Internet of Vehicles to prompt the driver to re-plan the route.
[0072] 4. Congestion Cause Identification Module
[0073] Function: Identify the cause of congestion and send specific alerts to unmanned vehicles that have entered the congested section.
[0074] Implementation: Use deep learning models combined with on-board camera and sensor data to identify abnormal road conditions (such as fires, people moving around) and send specific alarm information.
[0075] 5. Global coordination and optimization module
[0076] Function: Realize global traffic flow management and resource optimization to alleviate road network congestion.
[0077] Implementation: Use DQN to implement dynamic route adjustment and intelligent scheduling of unmanned vehicles, and continuously optimize system strategies based on historical data.
[0078] The functional modules of the system of the present invention are specifically implemented as follows:
[0079] 1. Implementation of data collection and real-time analysis module
[0080] First, GPS and speed sensors are used to record vehicle locations and travel times in real time. Then, the data is uploaded to the central server through V2X communication technology. Finally, the server uses Python language to process and analyze the data and dynamically update the average travel time for each road section.
[0081] 2. Implementation of the optimal route planning module
[0082] (1) Main functions
[0083] 1) Dynamically plan the route and calculate the shortest path.
[0084] 2) When the road conditions change, the model is updated in real time to re-recommend the optimal route.
[0085] 3) Balancing the ability to explore new paths and leverage existing experience in the path selection of unmanned vehicles.
[0086] (2) Key module description
[0087] 1) Environment definition
[0088] State: Current vehicle location information, target location, and average travel time on the current road segment.
[0089] Action: The next road segment that the vehicle can choose.
[0090] Reward function: guides the vehicle to choose the optimal route. Specific rewards include:
[0091] Positive reward: Choose the route with the shortest time.
[0092] Negative rewards: entering a congested section or deviating from the shortest path.
[0093] 2)DQN
[0094] The network architecture consists of an input layer (state dimension), two hidden layers (24 nodes, ReLU activation), and an output layer (action dimension).
[0095] Use two networks:
[0096] Main network: used to predict the Q value of the current state in real time.
[0097] Target network: used to calculate the target Q value and reduce the shock during training.
[0098] 3) Experience Replay
[0099] Store the state transitions (state, action, reward, next_state, done) during the navigation process of the unmanned vehicle in a buffer.
[0100] Randomly sample small batches of data for training to break temporal correlation and improve the generalization ability of the model.
[0101] 4) Reward Function
[0102] Reward design is key: actions that take a short time to travel on a road section receive high rewards; actions that are on a congested road section or impassable road section receive negative rewards; and the reward for reaching the end point is an additional high value.
[0103] 5) Dynamic route planning
[0104] The model predicts the best action based on the current state and dynamically adjusts the route.
[0105] If the road ahead is detected to be congested or abnormal (through the data analysis module), the model will re-plan the route.
[0106] (3) Steps
[0107] 1) Initialization
[0108] Initialize the urban road network environment, such as road sections, traffic lights, travel time, etc.
[0109] Define the Q network and target network (deep neural network).
[0110] Set hyperparameters: learning rate, discount factor gamma, exploration rate epsilon, batch size, etc.
[0111] Initialize the experience replay buffer ReplayBuffer.
[0112] 2) Environment definition
[0113] Environmental state: It is represented by the current vehicle location, target location, and the average travel time of all optional road sections. The environmental state can be represented by a vector S = {s1, s2, ..., s n}, where s j represents the average travel time of the jth road section, j∈[1,n], n is the total number of road sections in the road network;
[0114] Action space: represents the next possible path segment in the current state, defined as A = {a1, a2, ..., a m}, where a i Indicates the i-th road segment selected as the next action; m is the number of road segments that can be selected in the current state;
[0115] Reward function: The update formula of the Q-value function is Q(s,a)=Q(s,a)+α[r+γmaxQ(s',a')-Q(s,a)], where Q(s,a) is the Q-value of executing action a in the current state s, α is the learning rate, γ is the discount factor, r is the immediate reward for executing the action, and s' is the new state transferred to after executing action a; maxQ(s',a') represents the maximum Q-value of all possible actions a' in the new state s', which represents the best expected return in the new state.
[0116] 3) DQN training
[0117] The update formula of the Q-value function is Q(s,a)=Q(s,a)+α[r+γmaxQ(s',a')-Q(s,a)], where Q(s,a) is the Q-value of executing action a in the current state s, α is the learning rate, γ is the discount factor, r is the immediate reward for executing the action, and s' is the new state transferred to after executing action a; maxQ(s',a') represents the maximum Q-value of all possible actions a' under the new state s', which represents the best expected return in the new state.
[0118] 4) Model optimization
[0119] Using the experience replay mechanism, the state transition samples are stored as (s, a, r, s'), and training is performed by randomly sampling small batches of data to reduce time correlation. The loss function uses Huber loss:
[0120]
[0121] Where y reflects the long-term value of the agent's choice of action in the current state. y = r + γmaxQ (s', a') means that in state s', after taking action a', the immediate reward r that can be obtained plus the discounted return that may be obtained in the future.
[0122] 5) Dynamic programming strategy
[0123] In actual navigation, DQN predicts the current optimal action a * , the real-time update path is a * =argmaxQ(s,a).
[0124] 3. Implementation of congestion detection and response module
[0125] Real-time detection of road congestion status, reminding vehicles to avoid congested sections and re-plan routes. The specific steps are as follows:
[0126] (1) Congestion judgment
[0127] Set the segment time threshold T thresh , when the vehicle passing time T(a) satisfies T(a)>T thresh It is determined that road section a is congested.
[0128] (2) Alarm sending
[0129] Use V2X communication technology to send warning information to vehicles about to enter congested sections of road.
[0130] (3) Dynamic Adjustment
[0131] After the vehicle receives a congestion alert, it recalculates the optimal path based on DQN and adjusts the driving direction.
[0132] 4. Implementation of congestion cause identification module
[0133] Analyze the specific causes of congested road sections, such as accidents, fires, etc., and provide detailed alerts to vehicles. The specific steps are as follows:
[0134] (1) Data Collection
[0135] Use on-board cameras to collect real-time image data and combine with sensors to detect abnormal events on the road.
[0136] (2) Cause identification
[0137] The use of deep learning models to detect road anomalies is expressed as P(c|I)=f(I;θ), where I is the input image (visual information of the current road section), c is the category label (the identified cause of congestion, such as accidents, fires, etc.), and θ is the model parameter (which can effectively extract features from the input image and classify them after training); the model determines the specific cause of congestion based on the real-time image information of the road section, thereby providing timely warnings and decision support for vehicles. The deep learning model outputs the category c with the highest probability * as the cause of the exception.
[0138] The model outputs the category c with the highest probability * as the cause of the exception.
[0139] (3) Information Transmission
[0140] Send specific information about the cause of congestion to vehicles, such as "fire ahead, please take a detour."
[0141] 5. Implementation of global coordination and optimization module
[0142] Coordinate traffic flow globally, optimize resource utilization, and alleviate overall congestion. The specific steps are as follows:
[0143] (1) Traffic flow modeling
[0144] Construct a global road network graph G = (V, E), where V is the set of road segment nodes and E is the set of edges connecting road segments; the road segment flow is represented as f(e), e∈E.
[0145] (2) Optimization goal
[0146] The optimization goal is to minimize the global average travel time. Where T(e) is the travel time of road section e.
[0147] (3) Dynamic Scheduling
[0148] Based on the DQN model, vehicle scheduling is adjusted in real time to avoid too many vehicles concentrating on the same road section. By setting a threshold v for the number of vehicles, when the number of vehicles on a certain road section exceeds the threshold, the system will automatically adjust the allocation of vehicles on other sections to reduce congestion on that section.
[0149] (4) Strategy Update
[0150] Update traffic optimization strategies based on historical data and real-time traffic data to improve overall efficiency.
[0151] Example
[0152] In a major emergency in a city, such as a flood or earthquake disaster, the unmanned vehicle real-time route planning and anomaly detection system based on the present invention is deployed at the disaster site to assist volunteers in emergency rescue work.
[0153] 1. Application scenarios
[0154] Due to the flood, some areas of the city were seriously flooded, traffic was interrupted, and rescue resources were unevenly distributed. Intelligent means were urgently needed to assist in commanding and dispatching volunteers to complete rescue tasks. The system collects road condition data and volunteer activity information in real time, dynamically plans the optimal route, and guides volunteers to quickly reach the worst-hit areas while avoiding dangerous areas (such as sections with deep water or damaged bridges).
[0155] 2. System deployment and function realization
[0156] 1) Real-time data collection. First, using the sensors and cameras on the unmanned vehicle, the system collects real-time information about the road conditions in the disaster area, including the depth of water accumulation and whether the road is passable. Then, through vehicle-to-everything (V2X) communication technology, the system shares data with the command center and receives real-time positioning and status feedback from volunteers.
[0157] 2) Optimal route planning. First, based on the deep Q network (DQN), the system combines the real-time collected road condition data and disaster maps to calculate the optimal driving route for the volunteer rescue mission to avoid congestion and dangerous areas; then, the system dynamically updates the route. When it finds that the path is blocked (such as excessive water accumulation on the road), it will re-plan in time to ensure that the volunteers can arrive at the mission location smoothly.
[0158] 3) Abnormal detection and early warning. First, the system uses the unmanned vehicle's multimodal sensors (including thermal imagers and lidar) to monitor the disaster area in real time; then, when potential dangers (such as collapsed buildings, fires, and signs of landslides) are detected, the system sends an alert to volunteers and provides detour suggestions.
[0159] 3. Rescue effect
[0160] 1) The system assisted the unmanned vehicle in guiding the volunteers to complete the material distribution and personnel rescue tasks in the main disaster-stricken areas within 12 hours after the disaster.
[0161] 2) During one test, the system identified the depth of water and promptly warned volunteers to avoid dangerous areas, while also planning detours to prevent damage to the rescue team’s vehicles.
[0162] 3) For some areas where communication is interrupted, the system collects disaster data through the automatic inspection function of unmanned vehicles and feeds it back to the command center in real time to provide support for subsequent rescue decisions.
[0163] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A real-time route planning and anomaly detection system for unmanned vehicles based on deep Q network, characterized in that: By recording in real time the length of time each unmanned vehicle passes through a road section, the optimal route is dynamically calculated, and alarms are promptly issued and the route is replanned when road conditions are abnormal.
2. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 1 is characterized in that: include: The data collection and real-time analysis module records the time it takes each unmanned vehicle to pass through the section between two traffic lights and uploads it to the central server; The optimal route planning module counts and searches all routes, calculates the total time of each route, and pushes the most time-saving route to the user; The congestion detection and response module detects road congestion, sends an alert to unmanned vehicles about to enter the congested section, and replans the route; The congestion cause identification module identifies the cause of the congested section and sends specific alerts to the unmanned vehicles that have entered the congested section; The global coordination and optimization module realizes global traffic flow management and resource optimization to alleviate road network congestion.
3. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 2 is characterized in that: The data collection and real-time analysis module uses vehicle-mounted sensors and V2X communication technology to collect GPS data and vehicle speed information of unmanned vehicles in real time, store and dynamically update the average travel time of each road section. Specifically, GPS and speed sensors are used to record vehicle locations and travel times in real time. The data is uploaded to the central server through V2X communication technology; the central server uses Python language to process and analyze the data and dynamically update the average travel time of each road section.
4. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 2 is characterized in that: The optimal route planning module uses DQN for path optimization, uses the latest recorded time length of each road section to summarize the total time of the entire route, and recommends the most time-saving route.
5. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 4 is characterized in that: The optimal route planning module is specifically implemented as follows: 1) Initialization: Initialize the urban road network environment; define the Q network and the target network; set hyperparameters, including learning rate, discount factor gamma, exploration rate epsilon, and batch size; initialize the experience replay buffer ReplayBuffer; 2) Environment definition: Environmental state State: It is represented by the current vehicle location, target location and the average travel time of all optional road sections; the environmental state is represented by a vector S = {s1, s2, ..., s n }, where s j represents the average travel time of the jth road section, j∈[1,n], n is the total number of road sections in the road network; Action space: represents the next road segment that can be selected in the current state; it is defined as A = {a1, a2, ..., a m }, where a i Indicates the i-th road segment selected as the next action; m is the number of road segments that can be selected in the current state; Reward function: The reward function is used to guide the vehicle to choose the optimal path; if the road section with shorter time is selected, the reward value is higher; the formula is R = -T(a i )+λC(s,a i ), where T(a i ) is the time to pass the section, C(s,a i ) is the negative increment of the target distance, λ is the coefficient of the weight between the balance time and distance, and s represents the average travel time vector of all optional sections in the current environment state; 3) DQN training The update formula of the Q-value function is Q(s,a)=Q(s,a)+α[r+γmaxQ(s',a')-Q(s,a)], where Q(s,a) is the Q-value of executing action a in the current state s, α is the learning rate, γ is the discount factor, r is the immediate reward for executing the action, and s' is the new state transferred to after executing action a; max Q(s',a') represents the maximum Q-value of all possible actions a' under the new state s', indicating the best expected return under the new state; 4) Model optimization Using the experience replay mechanism, the state transition samples are stored as (s, a, r, s'), and training is performed by randomly sampling small batches of data to reduce time correlation; The loss function uses Huber loss: Where y reflects the long-term value of the agent's choice of action in the current state; y = r + γmax Q (s', a') represents the immediate reward r that can be obtained after taking action a' in state s' plus the discounted return that may be obtained in the future; 5) Dynamic programming strategy In actual navigation, DQN predicts the current optimal action a * , the real-time update path is a * =argmaxQ(s,a).
6. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 2 is characterized in that: The congestion detection and response module sets a road section time threshold. When the length of time a certain unmanned vehicle passes through a road section exceeds the road section time threshold, the corresponding road section is determined to be congested, and an alarm is sent to the unmanned vehicle system through the Internet of Vehicles to re-plan the route. The specific implementation is as follows: Set the segment time threshold T thresh , when the time T(a) for a vehicle to pass through section a satisfies T(a)>T thresh It is determined that road section a is congested; When congestion occurs on road section a, an alarm message is sent to vehicles about to enter the congested section through V2X communication technology; After the vehicle receives a congestion alert, it recalculates the optimal path and adjusts the driving direction based on the optimal route planning module.
7. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 2 is characterized in that: The congestion cause identification module uses a deep learning model combined with on-board camera and sensor data to identify abnormal road conditions and send specific alarm information; the specific implementation is as follows: Use vehicle-mounted cameras to collect real-time image data and combine sensors to detect abnormal events on the road; The use of deep learning models to detect road anomalies is expressed as P(c|I)=f(I;θ), where I is the input image, c is the category label, i.e. the identified cause of congestion, and θ is the model parameter. The model determines the specific cause of congestion based on the real-time image information of the road section, thereby providing timely alerts and decision support for vehicles. The deep learning model outputs the category c with the highest probability. * As an abnormal cause; Send specific information about the cause of congestion to vehicles.
8. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 2 is characterized in that: The global coordination and optimization module uses DQN to achieve dynamic route adjustment and intelligent scheduling of unmanned vehicles, and continuously optimizes system strategies based on historical data.
9. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 8 is characterized in that: The global coordination and optimization module is implemented as follows: (1) Traffic flow modeling Construct a global road network graph G = (V, E), where V is the set of road segment nodes and E is the set of edges connecting road segments; the road segment flow is represented as f(e), e∈E; (2) Optimization goal The optimization goal is to minimize the global average travel time. Where T(e) is the travel time of road section e; (3) Dynamic Scheduling Based on the DQN model, vehicle scheduling is adjusted in real time to avoid too many vehicles concentrating on the same road section. Specifically, by setting a threshold v for the number of vehicles, when the number of vehicles on a certain road section exceeds the threshold v, the system will automatically adjust the vehicle allocation on other sections to reduce congestion on that section. (4) Strategy Update Update traffic optimization strategies based on historical data and real-time traffic data to improve overall efficiency.
10. The unmanned vehicle real-time route planning and anomaly detection system based on deep Q network according to claim 1, characterized in that: The method is applied to complex urban road networks, logistics distribution, and emergency rescue fields.