An intelligent vehicle scheduling method and system for coping with extreme weather

Through real-time data-driven path planning and optimization, combined with deep reinforcement learning and multi-criteria decision-making analysis, the flexibility and adaptability of the vehicle scheduling system in extreme weather is solved, and efficient and safe rescue path selection is achieved.

CN119132067BActive Publication Date: 2025-07-04ELEPHANT CLOUD INTELLIGENCE DATA OPERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411404989.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-07-04
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

The existing vehicle dispatching system lacks flexibility and adaptability in extreme weather conditions, and cannot effectively and quickly pick up and drop off a large number of passengers, resulting in inefficient rescue efficiency and risk of traffic accidents.

Method used

By obtaining real-time traffic and weather data, using deep reinforcement learning models and multi-criteria decision analysis (MCDA) evaluation models, dynamically adjust path planning, optimize rescue paths, and combine Dijkstra, A* and Bellman-Ford algorithms to select the optimal rescue route to ensure safety and efficiency.

Benefits of technology

In extreme weather, the efficiency and safety of vehicle dispatching is improved, the risk of accidents is reduced, the effective utilization of resources is ensured, and a flexible and efficient rescue plan is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132067B_ABST
    Figure CN119132067B_ABST
Patent Text Reader

Abstract

The present application provides a vehicle intelligent scheduling method and system for coping with extreme weather, which relates to the technical field of digital traffic management. The method includes: obtaining the vehicle position of a target vehicle to be scheduled, a set of user positions corresponding to multiple platform users to be transported, real-time traffic data, and real-time weather data; determining at least one first candidate path based on a first path planning model, where the first candidate path passes through the vehicle position and all user positions; for each first candidate path, adjusting the path segments of the first candidate path based on a deep reinforcement learning model to obtain a corresponding corrected path; calculating the respective path comprehensive scores corresponding to each corrected path based on an MCDA evaluation model, and determining the rescue pick-up path of the target vehicle from each corrected path according to each path comprehensive score to sequentially pick up each platform user. Thus, through intelligent scheduling, it is ensured that resources can be most effectively utilized in extreme weather.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of digital traffic management, and particularly to a vehicle intelligent scheduling method and system for coping with extreme weather. Background Art

[0002] With the increasing complexity of urban traffic networks and the continuous growth of vehicle numbers, effective vehicle scheduling has become increasingly important. Existing vehicle scheduling systems have achieved remarkable results under normal weather conditions, being able to reasonably arrange vehicle driving routes, optimize traffic flow, and reduce congestion. However, when facing extreme weather conditions such as heavy rain, heavy snow, freezing rain, typhoons, etc., the effectiveness of these systems often drops significantly.

[0003] Extreme weather not only affects traffic flow and vehicle driving speed but also leads to poor road conditions and increases the risk of traffic accidents. For example, in heavy rain weather, roads are prone to waterlogging and visibility is low, causing the driver's line of sight to be blocked; heavy snow may cause roads to freeze, increasing the risk of vehicle skidding. In the event of extreme weather, the core of the vehicle scheduling system should no longer be limited to picking up and dropping off specific individual passengers, but rather to quickly reduce the number of passengers and their staying time in the affected area to ensure the life, health, and safety of the people.

[0004] Traditional scheduling systems lack flexibility and adaptability in extreme weather conditions and still usually adopt fixed scheduling logics. However, in extremely bad weather, there are fewer available social vehicles and rescue vehicles for scheduling, while the number of passengers to be served is numerous. How to quickly and effectively pick up and drop off passengers in the disaster area safely at a relatively low vehicle operation cost to implement effective rescue is an urgent problem in the industry at present. Summary of the Invention

[0005] This application provides a vehicle intelligent scheduling method and system for coping with extreme weather to at least solve the problem that the traditional vehicle scheduling scheme cannot meet the transportation needs of passenger users in extreme weather conditions.

[0006] The present application provides a vehicle intelligent dispatching method for coping with extreme weather, including: obtaining the vehicle position of a target vehicle to be dispatched, a set of user positions corresponding to a plurality of platform users to be transported, real-time traffic data, and real-time weather data; the distance between each user position in the set of user positions and the vehicle position is less than a preset distance threshold; the real-time traffic data includes: road closure information, traffic accident information, and road congestion information; in the case of determining that the real-time weather data includes extreme weather information, based on a first path planning model, determining at least one first candidate path, the first candidate path passing through the vehicle position and all of the user positions; the extreme weather information includes any one of the following: heavy rain, heavy snow, freezing rain, or typhoon; for each of the first candidate paths, based on a deep reinforcement learning model, adjusting at least one path segment of the first candidate path to obtain a corresponding corrected path; the path segment is determined according to any two user positions in the set of user positions; the state of the deep reinforcement learning model is defined according to the node position, real-time traffic data, and real-time weather data, and the action of the deep reinforcement learning model is defined according to the switching of the path segment; the node position is the vehicle position or the user position; calculating the path comprehensive score corresponding to each of the corrected paths based on the MCDA evaluation model, and determining the rescue pick-up path of the target vehicle from each of the corrected paths to sequentially pick up each of the platform users.

[0007] Optionally, after obtaining the vehicle position of the target vehicle to be dispatched, the set of user positions of a plurality of platform users to be transported, the real-time traffic data, and the real-time weather data, the method further includes: in the case of determining that the real-time weather data does not include extreme weather information, based on a second path planning model, determining at least one second candidate path, the second candidate path passing through the vehicle position and any one of the user positions; determining a passenger-carrying driving path from each of the second candidate paths to pick up and transport the platform user corresponding to the passenger-carrying driving path.

[0008] Optionally, the deep reinforcement learning model adopts a deep Q network, and the main network of the deep Q network includes a cascaded input layer, a hidden layer, and an output layer; the input layer is used for normalizing the data; the input data includes the real-time traffic data, the real-time weather data, and the first candidate path; the hidden layer includes a multi-layer convolutional neural network for extracting feature information from the input data; the output layer is used for determining the expected Q value corresponding to each action in the state space, and correcting the first candidate path according to the path segment corresponding to the action with the maximum expected Q value to determine a corresponding corrected path.

[0009] Optionally, the update formula of the deep Q network is as follows:

[0010] Q(s,a;θ) = Q(s,a;θ) + α[r + γmax a′ Q(s′,a′;θ′) - Q(s,a;θ)]

[0011] r = w time ×r time +w safety ×r safety

[0012] Wherein, Q(s,a;θ) represents the current Q value of taking action a in state s, θ represents the current network parameters, α represents the learning rate; γ represents the discount factor, which indicates the importance of future rewards; max a′ Q(s′,a′;θ′) represents the maximum Q value that can be generated by all possible actions a′ in the next state s′, θ′ represents the old network parameters; r represents the immediate reward, which is obtained according to the result of the current action; r time and r safety respectively represent the timeliness reward term and the safety reward term, and w time and w safety respectively represent the weight coefficients of the corresponding reward terms, and w safety > w time , where w time promotes the model to select a path segment with a shorter distance, and w safety promotes the model to select a path segment with a lower probability of road condition anomalies under real-time weather data.

[0013] Optionally, the loss function of the deep Q network is:

[0014] L(θ) = E[(y i - Q(s,a;θ)) 2

[0015] The formula for the temporal difference target y i is:

[0016] y i = λ(s,a)×r + γmax a′ Q(s′,a′;θ - )

[0017] Wherein, Q(s,a;θ) represents the Q value estimation of the current state-action pair, provided by the main network; Q(s′,a′;θ - ) represents the Q value of the next state, estimated by the target network; θ and θ - respectively represent the parameters of the main network and the target network; E[(y i - Q(s,a;θ)) 2 ​represents the expected value of the squared prediction error; λ(s,a) represents the environmental adaptability factor, which indicates the difficulty of taking a specific action a in a specific state s, and is used to adjust the weight of the immediate reward r.

[0018] Optionally, the first path planning model includes a Dijkstra algorithm module, an A* algorithm module, and a Bellman-Ford algorithm module, where the Dijkstra algorithm module, the A* algorithm module, and the Bellman-Ford algorithm module are connected in parallel to the deep reinforcement learning model.

[0019] Optionally, the MCDA evaluation model calculates the comprehensive path score corresponding to the corrected path in the following way:

[0020]

[0021] In the formula, S route1 represents the comprehensive path score of the corrected path route1, w i represents the weight of the i-th evaluation criterion, C i,route1 represents the normalized score of the corrected path route1 on the i-th evaluation criterion; the evaluation criteria include: driving timeliness evaluation criterion, path economic cost evaluation criterion, and road condition safety evaluation criterion;

[0022] The calculation formula for the normalized score corresponding to the driving timeliness evaluation criterion is:

[0023]

[0024] α = 1 + k × WL

[0025] In the formula, T route1 represents the expected driving time of the corrected path route1, T min and T max respectively represent the minimum and maximum values of the expected driving times of all the corrected paths; α represents the adjustment coefficient, k represents a positive constant, which is used to adjust the influence degree of the bad weather level on the adjustment coefficient; WL represents the bad weather level, which is determined by querying the real-time weather data with the bad weather level mapping table, and the bad weather level mapping table records multiple bad weather levels and the corresponding weather data intervals;

[0026] The calculation formula for the normalized score corresponding to the path economic cost evaluation criterion is:

[0027]

[0028] In the formula, E routel represents the expected economic cost of the corrected path route1, E maxrepresents the maximum of the expected economic costs of all the corrected paths; the expected economic cost of a path includes the energy consumption cost and the tolls for the toll sections involved in the path;

[0029] The calculation formula for the standardized score corresponding to the road condition safety assessment standard is:

[0030]

[0031] In the formula, m j is the weight of the j-th safety factor, and C j,route1,sf represents the score of the corrected path route1 for the j-th safety factor; the safety factors include at least one of the following: road wetness state, historical accident records of the path, path curve information, and path slope information.

[0032] This application also provides a vehicle intelligent dispatching system for coping with extreme weather, including: a data acquisition unit configured to acquire the vehicle position of the target vehicle to be dispatched, the set of user positions corresponding to multiple platform users to be transported, real-time traffic data, and real-time weather data; the distance between each user position in the set of user positions and the vehicle position is less than a preset distance threshold; the real-time traffic data includes: road closure information, traffic accident information, and road congestion information; a path planning unit configured to, when determining that the real-time weather data includes extreme weather information, determine at least one first candidate path based on a first path planning model, the first candidate path passing through the vehicle position and all the user positions; the extreme weather information includes any one of the following: heavy rain, heavy snow, freezing rain, or typhoon; a path correction unit configured to, for each of the first candidate paths, adjust at least one path segment of the first candidate path based on a deep reinforcement learning model to obtain a corresponding corrected path; the path segment is determined according to any two user positions in the set of user positions; the state of the deep reinforcement learning model is defined according to the node position, real-time traffic data, and real-time weather data, and the action of the deep reinforcement learning model is defined according to the switching of the path segment; the node position is the vehicle position or the user position; a pick-up path determination unit configured to calculate the path comprehensive score corresponding to each of the corrected paths based on the MCDA evaluation model, and determine the rescue pick-up path of the target vehicle from each of the corrected paths according to each of the path comprehensive scores to sequentially pick up each of the platform users.

[0033] This application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the vehicle intelligent dispatching method for coping with extreme weather as described in any one of the above.

[0034] The present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the vehicle intelligent scheduling method for coping with extreme weather as described in any one of the above.

[0035] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the vehicle intelligent scheduling method for coping with extreme weather as described in any one of the above.

[0036] Through a vehicle intelligent scheduling method and system for coping with extreme weather provided by the present application, to improve the vehicle scheduling efficiency and safety under extreme weather conditions, the following technical effects can be produced at least:

[0037] (1) By obtaining traffic and weather data in real time, the system can quickly respond to extreme weather conditions. Using a preset distance threshold ensures that vehicles can efficiently approach the users with the most urgent needs, thus accelerating the rescue speed and improving the emergency response speed.

[0038] (2) First, select candidate paths through a path planning model, and then use a deep reinforcement learning model to adjust these paths in real time. It can avoid situations such as road closures, traffic accidents or severe congestion caused by extreme weather while ensuring the maximum rescue efficiency, improve the practicality and safety of the rescue route, and enable the system to output the optimal rescue route.

[0039] (3) The deep reinforcement learning model can continuously adjust the path selection according to real-time data, enabling the system to maintain a high degree of adaptability and flexibility under extreme weather. Especially in extreme weather conditions, poor road conditions and low visibility will significantly increase the risk of traffic accidents. The system effectively reduces the accident risk during the rescue process by avoiding high-risk areas and choosing relatively safe routes. Based on the dynamic adjustment ability of the deep reinforcement learning model, enhance the adaptability and flexibility of the path, enable the vehicle to effectively respond to emergencies, and reduce rescue delays.

[0040] (4) Through the MCDA (Multi-Criteria Decision Analysis) evaluation model, various factors such as path length, expected time, safety, etc. can be comprehensively considered, score each corrected path, and by comparing these scores, the system can select the optimal rescue path to pick up platform users in the order of the smallest cost, ensuring the user service experience.

[0041] Through the embodiments of the present application, by combining real-time data, intelligent path planning, and dynamic adjustment capabilities, the efficiency, safety, and adaptability of vehicle scheduling under extreme weather conditions have been significantly improved. In the case of limited vehicles and rescue resources, the intelligent scheduling maximizes the rescue effect, ensuring that resources can be utilized most effectively under extreme weather, and providing an effective solution for vehicle scheduling under extreme weather. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the present application or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 FIG. shows a flowchart of an example of a vehicle intelligent scheduling method for coping with extreme weather according to an embodiment of the present application;

[0044] Figure 2 FIG. shows a schematic diagram of the effect of an example of candidate path planning according to an embodiment of the present application;

[0045] Figure 3 FIG. shows a schematic diagram of the working principle of an example of the state transition action in the reinforcement learning model;

[0046] Figure 4 FIG. shows a flowchart of an example of vehicle intelligent scheduling path planning operation according to an embodiment of the present application;

[0047] Figure 5 FIG. shows a schematic structural diagram of an example of a deep Q-network according to an embodiment of the present application;

[0048] Figure 6 FIG. shows a schematic block diagram of an example of a first path planning model according to an embodiment of the present application;

[0049] Figure 7 FIG. shows a schematic block diagram of an example of a vehicle intelligent scheduling system for coping with extreme weather according to an embodiment of the present application;

[0050] Figure 8 FIG. is a schematic structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0052] Figure 1 The flowchart of an example of a vehicle intelligent scheduling method for coping with extreme weather according to an embodiment of this application is shown.

[0053] Regarding the execution subject of the method of the embodiment of this application, it can be any electronic device with processing and computing capabilities, such as a computer terminal, a mobile phone terminal, etc. More specifically, in the business service scenario of vehicle centralized scheduling, its execution subject can be a central server to achieve the combination of real-time data, intelligent path planning, and dynamic adjustment capabilities. In the case of limited vehicles and rescue resources, through intelligent scheduling, the rescue effect is maximized, ensuring that resources can be most effectively utilized in extreme weather and providing an effective and new scheduling solution for extreme weather.

[0054] As Figure 1 shown, in step S110, the vehicle position of the target vehicle to be scheduled, the set of user positions corresponding to multiple platform users to be transported, real-time traffic data, and real-time weather data are acquired.

[0055] Here, the distance between each user position in the set of user positions and the vehicle position is less than a preset distance threshold. The real-time traffic data includes: road closure information, traffic accident information, and road congestion information. It should be understood that the preset distance threshold can be adjusted or set according to platform business requirements or the personalized requirements of the vehicle owner, for example, 2 - 5 kilometers, ensuring that the vehicle can efficiently approach the users with the most urgent needs, thus accelerating the rescue speed and improving the emergency response speed. In addition, the central server can update the real-time traffic data and real-time weather data by interacting with the third-party platform server.

[0056] In step S120, when it is determined that the real-time weather data includes extreme weather information, based on the first path planning model, at least one first candidate path is determined, and the first candidate path passes through the vehicle position and all user positions.

[0057] Specifically, the extreme weather information includes any one of the following: heavy rain, heavy snow, freezing rain, or typhoon. It should be understood that the extreme weather information can also be other types of information and can be set in the system, such as sandstorm weather, thick fog weather, etc.

[0058] Figure 2The figure shows a schematic diagram of the effect of an example of candidate path planning according to an embodiment of the present application.

[0059] As Figure 2 shown, within a preset distance threshold near vehicle position 20, there are multiple user positions to be transported (for example, 11, 12... 19). By planning a first candidate path that passes through the vehicle position and all user positions, rapid pick-up of all users can be achieved. It should be noted that this is different from the ride-hailing route planning for a single passenger (i.e., the pick-up route planning from the vehicle position to any user position), and is somewhat similar to the carpool pick-up route planning that needs to pick up multiple passengers simultaneously, so as to achieve rapid pick-up of all passengers in the nearby area and reduce the waiting time of passengers in extreme weather conditions.

[0060] In some embodiments, multiple parallel path planning modules are provided in the first path planning model to independently complete the planning and output of the first candidate path.

[0061] In step S130, for each first candidate path, at least one path segment of the first candidate path is adjusted based on the deep reinforcement learning model to obtain a corresponding corrected path.

[0062] In some embodiments, the path segments are determined according to any two user positions in the user position set (for example, 11 - 12, 12 - 15, 13 - 14, etc.).

[0063] Figure 3 The figure shows a schematic diagram of the working principle of an example of the state transition action in the reinforcement learning model.

[0064] As Figure 3 shown, the state transition schematic diagram involves a state space composed of multiple basic states f1 to f n . State transitions may occur between different basic states. For example, a1 represents the state transition from f1 to f2, a2 represents the state transition from f2 to f1, a3 represents the state transition from f1 to f3, and so on. Here, corresponding state transitions can occur based on the state transition strategy, and each state transition strategy can be used to occur different state transitions respectively. Exemplarily, based on the state transition strategy for basic state f1, state transitions a2 or a3 can occur.

[0065] In addition, the state range of another state (also called the transferable state) that a basic state in the state space can transfer to is generally restricted or conditional. For example, any one of f1 to f3 will not be related to f4 to f nA state transition occurs between them. For the state f1, the states it can transition to are f2 and f3, and so on. In the context of this technology scenario, nodes 12 - 15 can be adjusted to 12 - 13 or 12 - 14, but cannot be adjusted to 12 - 18.

[0066] In the example of the embodiment of this application, the state of the deep reinforcement learning model (DQN) is defined based on the node position, real - time traffic data, and real - time weather data, and the action of the deep reinforcement learning model is defined based on the switching of path segments. The node position is the vehicle position or the user position. Thus, through the deep reinforcement learning model, each path segment in the first candidate path can be optimized and adjusted, realizing adaptive state adjustment and transition according to the state environment variables (node position, real - time traffic data, and real - time weather data), so that the adjusted corrected path can effectively meet the pick - up requirements in extreme weather.

[0067] In some embodiments, the reward function of the deep reinforcement learning model needs to consider factors such as traffic safety and time efficiency to evaluate the effect of taking a certain action in a specific state through immediate rewards, and guide the deep reinforcement learning model to output the best action in a specific state.

[0068] In step S140, based on the MCDA evaluation model, calculate the comprehensive path scores corresponding to each corrected path, and determine the rescue pick - up path of the target vehicle from each corrected path according to each comprehensive path score to pick up each platform user in sequence.

[0069] In some embodiments, the scoring criteria involved in the MCDA evaluation model are diverse, such as including safety, time efficiency, economy, etc., and are also allowed to overlap with the reward function of DQN to guide the selection of the best path plan from multiple candidate path plans. By quantifying the value scores of each corrected path, the system can comprehensively evaluate and select the optimal rescue path, pick up platform users in sequence at the lowest cost, and ensure the user service experience.

[0070] Figure 4 Shows a flowchart of an example of the vehicle intelligent scheduling path planning operation according to an embodiment of this application.

[0071] As Figure 4 shown, in step S410, obtain the vehicle position of the target vehicle to be scheduled, the set of user positions corresponding to multiple platform users to be transported, real - time traffic data, and real - time weather data.

[0072] In step S420, detect whether the real - time weather data contains extreme weather information.

[0073] In step S431, when it is determined that the real-time weather data contains extreme weather information, based on the first path planning model, at least one first candidate path is determined, and this first candidate path passes through the vehicle position and all user positions.

[0074] In step S433, when it is determined that the real-time weather data does not contain extreme weather information, based on the second path planning model, at least one second candidate path is determined, and this second candidate path passes through the vehicle position and any one user position.

[0075] Here, the second path planning model can adopt the online car-hailing pick-up model to screen the best pick-up user position from multiple user positions and plan the corresponding pick-up path. For more details, reference can be made to the description in the current related technologies in part, and it will not be elaborated here.

[0076] In step S440, the passenger-carrying driving path is determined from each of the second candidate paths to pick up and transport the platform users corresponding to the passenger-carrying driving path.

[0077] Through the embodiments of the present application, the central server monitors the current weather state in real time. When the current weather state does not belong to the extreme state, the passenger-carrying pick-up path planning is used, and when the current weather state belongs to the extreme state, the rescue pick-up path planning is used, providing a new vehicle scheduling function and improving the riding service experience of the vehicle owner and passengers.

[0078] In some examples of the embodiments of the present application, the deep reinforcement learning model adopts a deep Q-network. Figure 5 FIG. shows a schematic structural diagram of an example of the deep Q-network according to the embodiments of the present application.

[0079] As Figure 5 shown, the deep Q-network 500 includes a main network 510 and a target network 520. The main network 510 is responsible for generating the Q value of the current action, that is, Q(s,a;θ), where during the training process, the parameters θ of the main network will be continuously updated. The target network 520 is used to estimate the Q value of the future state, that is, Q(s′,a′;θ - ), where the parameters θ of the target network - are copied from the main network regularly, but remain unchanged between two copies.

[0080] It should be noted that in the deep Q-network (DQN), the target network mainly plays a role in the training stage and is not directly used in the actual prediction or decision-making stage. During the training process of the DQN, the target network provides a stable target value for calculating the loss function and updating the parameters of the main network.

[0081] More specifically, the main network 510 includes a cascaded input layer 511, a hidden layer 513, and an output layer 515. Here, the input layer 511 is used to normalize the data, and the input data includes real-time traffic data, real-time weather data, and the first candidate path. The hidden layer 513 includes a multi-layer convolutional neural network for extracting feature information from the input data. The output layer 515 is used to determine the expected Q value corresponding to each action in the state space, and correct the first candidate path according to the path segment corresponding to the action with the maximum expected Q value to determine the corresponding corrected path.

[0082] Thus, by using the double-network architecture of DQN, the feedback loop that may occur during the self-update process is avoided, thereby improving the stability of the learning process. When estimating the TD (temporal difference) target, if the same network is used to evaluate the Q values of the current and future states, it may lead to a large deviation.

[0083] In some examples of the embodiments of the present application, the update formula of the deep Q network is as follows:

[0084] Q(s,a;θ) = Q(s,a;θ) + α[r + γmax a′ Q(s′,a′;θ′) - Q(s,a;θ)] Equation (1)

[0085] r = w time ×r time +w safety ×r safety Equation (2)

[0086] In the formula, Q(s,a;θ) represents the current Q value of taking action a in state s, θ represents the current network parameters, α represents the learning rate; γ represents the discount factor, which indicates the importance of future rewards; max a′ Q(s′,a′;θ′) represents the maximum Q value that can be generated by all possible actions a′ in the next state s′, θ′ represents the old network parameters; r represents the immediate reward, which is obtained according to the result of the current action; r time and r safety respectively represent the timeliness reward term and the safety reward term, and w time and w safety respectively represent the weight coefficients of the corresponding reward terms, and w safety > w time , where w time promotes the model to select a path segment with a shorter distance, and w safety promotes the model to select a path segment with a lower probability of road condition anomaly risk under real-time weather data.

[0087] For example, assume that there are three second user locations where path segments can be formed for the first user location. Through real-time traffic, it is found that there is a traffic accident in one path segment, and another path segment is blocked due to heavy rain. Furthermore, the DQN model will evaluate the Q-values of each path, which is not only based on the length and estimated time of the path, but also takes into account the impact of real-time data (such as road closures caused by heavy rain), to decide to select a detour but safer and more reliable path, even if it is not the shortest. Thus, effective decision-making for intelligent vehicle scheduling under extreme weather conditions is achieved, improving the adaptability and robustness of path planning.

[0088] In some examples of the embodiments of the present application, the loss function of the deep Q-network is:

[0089] L(θ) = E[(y i - Q(s,a;θ)) 2 Equation (3)

[0090] The temporal difference target y i has the formula:

[0091] y i = λ(s,a) × r + γmax a′ Q(s′,a′;θ - ) Equation (4)

[0092] In the formula, Q(s,a;θ) represents the Q-value estimation of the current state-action pair, provided by the main network; Q(s′,a′;θ - ) represents the Q-value of the next state, estimated by the target network; θ and θ - represent the parameters of the main network and the target network respectively; E[(y i - Q(s,a;θ)) 2 represents the expected value of the square of the prediction error; λ(s,a) represents the environmental adaptability factor, which indicates the difficulty of taking a specific action a in a specific state s, and is used to adjust the weight of the immediate reward r.

[0093] It should be noted that DQN usually uses diverse variants of temporal difference learning as the loss function. In the examples of the embodiments of the present application, in order to adapt to the special requirements under extreme weather conditions, an environmental adaptability factor λ is introduced into the TD target. This factor dynamically adjusts the weight of the immediate reward according to the complexity and difficulty of the current environment (such as weather conditions, traffic conditions), so as to complete the optimization and adjustment of the loss function. More specifically, as λ(s,a) in the above Equation (4) represents a function based on the current state and action, used to evaluate the risk and difficulty of the current selection. For example, path selection under extreme weather conditions or in high traffic congestion areas may receive a higher weight because they have higher complexity.

[0094] More specifically, an example of the functional formula for calculating λ(s,a) is as follows:

[0095] λ(s,a) = λ W ×λ T ×λ R Equation (5)

[0096] In the formula, λ W , λ T and λ R respectively represent the weather condition influence coefficient, the traffic flow influence coefficient, and the road surface condition influence coefficient. The value range of λ W is [1, 2], where 1 represents clear weather and 2 represents extremely bad weather. The value range of λ T is [1, 1.5], where 1 represents smooth traffic and 1.5 represents highly congested traffic. The value range of λ R is [1, 1.5], where 1 represents good condition and 1.5 represents bad condition (such as icing, waterlogging).

[0097] Through the above Equation (5), the comprehensive influence of different environmental factors on the immediate reward under the given state s and action a can be reflected. Exemplarily, assume that in a certain state, the weather is moderate rain (λ W = 1.5), the traffic is slightly congested (λ T = 1.2), and the road surface is wet but drivable (λ R = 1.1). In this way, λ(s,a) = 1.5×1.2×1.1 = 1.98, which means that the immediate reward in this state and action will increase by 98% to reflect the additional difficulty brought by adverse weather and road conditions.

[0098] Through the embodiments of the present application, the above loss function is used to train the DQN, and the optimization goal is to minimize the loss function, so that the predicted Q value of the network is as close as possible to the TD target weighted by environmental adaptability, aiming to make the network more focused on effective decision-making under extreme weather conditions and improve the adaptability and robustness of the system in complex environments.

[0099] Figure 6 Shows a structural schematic diagram of an example of the first path planning model according to the embodiments of the present application.

[0100] As Figure 6 shown, the first path planning model 600 includes a Dijkstra algorithm module 610, an A* algorithm module 620, and a Bellman-Ford algorithm module 630.

[0101] It should be noted that the Dijkstra algorithm is a path search method based on the greedy algorithm, which is used to find the shortest paths from a node to all other nodes in a weighted graph. It repeatedly selects the node closest to the starting point among the unvisited nodes and updates the distances of its adjacent nodes until all nodes are visited. In the context of this technology scenario, the Dijkstra algorithm is suitable for finding the shortest paths. Especially when there are no negative-weight edges in the road network, it can provide an optimal path from the starting point to the end point, and it is particularly efficient when there are many path options.

[0102] The Bellman-Ford algorithm calculates the shortest paths from a single source node to all other nodes by performing multiple relaxation operations on all the edges in the graph. And even in the presence of negative-weight edges, this algorithm is still effective. In the context of this technology scenario, under extreme weather conditions, the toll or time cost of some roads may suddenly increase due to emergencies (which can be regarded as negative weights). At this time, the Bellman-Ford algorithm can handle such negative-weight edges and find a reasonable path.

[0103] The A* algorithm is a heuristic search algorithm that finds the most efficient path by comprehensively considering the known path cost (such as the actual distance to the current node) and an estimated remaining cost (such as the estimated distance from the current node to the destination). In the context of this technology scenario, under extreme weather conditions, the A* algorithm can effectively balance the actual state of the path and the predicted future conditions to find an efficient driving route.

[0104] The Dijkstra algorithm module 610, the A* algorithm module 620, and the Bellman-Ford algorithm module 630 are connected in parallel to the deep reinforcement learning model, enabling the deep reinforcement learning model to analyze and correct the path segments in the candidate paths output by the path planning algorithm module to ensure that the corrected paths can at least meet the safety and economy under extreme weather.

[0105] Furthermore, the MCDA evaluation model is used to comprehensively score and rank each corrected path output by the deep reinforcement learning model, so as to select the final rescue pick-up path.

[0106] More specifically, the MCDA evaluation model calculates the comprehensive path score corresponding to the corrected path in the following way:

[0107]

[0108] In the formula, S route1 represents the comprehensive path score of the corrected path route1, w i represents the weight of the i-th evaluation criterion, C i,route1Represents the standardized score of the corrected path route1 on the i-th evaluation criterion; the evaluation criteria include: driving timeliness evaluation criterion, path economic cost evaluation criterion, and road condition safety evaluation criterion.

[0109] The calculation formula for the standardized score corresponding to the driving timeliness evaluation criterion is:

[0110]

[0111] α = 1 + k × WL Formula (8)

[0112] In the formula, T route1 represents the expected driving time of the corrected path route1, T min and T max respectively represent the minimum and maximum values among the expected driving times of all corrected paths; α represents the adjustment coefficient, k represents a positive constant used to adjust the influence degree of the bad weather level on the adjustment coefficient; WL represents the bad weather level, which is determined by querying the real-time weather data with the bad weather level mapping table. The bad weather level mapping table records multiple bad weather levels and corresponding weather data intervals.

[0113] Regarding the above adjustment coefficient α, for example, assume k = 0.1, which means that for each increase in one level of bad weather level, the adjustment coefficient increases by 0.1. If the weather condition is heavy rain (level 3), then α = 1 + 0.1 × 3 = 1.3. In this way, in heavy rain weather, the original driving time score will increase by 30% to reflect the additional difficulty brought by the bad weather.

[0114] In the embodiment of the present application, when calculating the standardized score of the driving timeliness evaluation criterion, the adjustment coefficient α that is in a positive proportional linear relationship with the bad weather level is comprehensively considered, effectively taking into account the severity of the weather conditions, making the score of the driving time more fair and in line with the actual situation.

[0115] The calculation formula for the standardized score corresponding to the path economic cost evaluation criterion is:

[0116]

[0117] In the formula, E routel represents the expected economic cost of the corrected path route1, E max represents the maximum value among the expected economic costs of all corrected paths; the expected economic cost of the path includes energy consumption costs and toll road passage fees involved in the path.

[0118] In this way, the economic evaluation can help the system compare different routes from the perspective of cost-benefit, select an economically efficient path under extreme weather conditions, and contribute to controlling operating costs and improving service efficiency.

[0119] The formula for the standardized score corresponding to the road condition safety evaluation standard is:

[0120]

[0121] In the formula, m j is the weight of the j-th safety factor, and C j,route1,sf represents the score of the modified path route1 for the j-th safety factor. The safety factors include at least one of the following: road wetness condition, path historical accident record, path curve information, and path slope information.

[0122] In the embodiments of the present application, each path is scored according to its safety factors, and the scores of different safety factors are combined to obtain the total safety score of each path. Thus, the system can comprehensively evaluate the safety of each path and ensure that the selected route under extreme weather conditions is as safe as possible.

[0123] It should be noted that in some cases, there is an overlap between the evaluation criteria of the MCDA evaluation model and the reward items of the DQN reward function, for example, both involve considerations for safety and time efficiency. However, on the one hand, in terms of methods, DQN focuses on learning and optimizing the dynamic decision-making process, while MCDA focuses more on comprehensive evaluation and decision-making, and they can achieve effective complementarity. On the other hand, DQN is used to dynamically adjust the path segments of candidate paths to obtain modified paths from the candidate paths, while MCDA further evaluates and verifies each modified path on the basis of DQN to ensure that the final rescue pick-up path can fully and comprehensively meet various evaluation criteria.

[0124] The vehicle intelligent scheduling system for coping with extreme weather provided by the present application will be described below. The vehicle intelligent scheduling system for coping with extreme weather described below can be mutually referred to with the vehicle intelligent scheduling method for coping with extreme weather described above.

[0125] Figure 7 The structural block diagram of an example of the vehicle intelligent scheduling system for coping with extreme weather according to an embodiment of the present application is shown.

[0126] As Figure 7 shown, the vehicle intelligent scheduling system 700 for coping with extreme weather includes a data acquisition unit 710, a path planning unit 720, a path correction unit 730, and a pick-up path determination unit 740.

[0127] The data acquisition unit 710 is configured to acquire the vehicle position of the target vehicle to be scheduled, the set of user positions corresponding to multiple platform users to be transported, real-time traffic data, and real-time weather data; the distance between each user position in the set of user positions and the vehicle position is less than a preset distance threshold; the real-time traffic data includes: road closure information, traffic accident information, and road congestion information.

[0128] The path planning unit 720 is configured to, when determining that the real-time weather data includes extreme weather information, determine at least one first candidate path based on the first path planning model, where the first candidate path passes through the vehicle position and all of the user positions; the extreme weather information includes any one of the following: heavy rain, heavy snow, freezing rain, or typhoon.

[0129] The path correction unit 730 is configured to, for each of the first candidate paths, adjust at least one path segment of the first candidate path based on the deep reinforcement learning model to obtain a corresponding corrected path; the path segment is determined according to any two user positions in the set of user positions; the state of the deep reinforcement learning model is defined according to the node position, real-time traffic data, and real-time weather data, and the action of the deep reinforcement learning model is defined according to the switching of the path segment; the node position is the vehicle position or the user position.

[0130] The pick-up path determination unit 740 is configured to calculate the path comprehensive score corresponding to each of the corrected paths based on the MCDA evaluation model, and determine the rescue pick-up path of the target vehicle from each of the corrected paths according to each path comprehensive score to sequentially pick up each of the platform users.

[0131] In some embodiments, the embodiment of the present application provides a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the vehicle intelligent scheduling method for coping with extreme weather in the present application.

[0132] In some embodiments, the embodiment of the present application further provides a computer program product, where the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the vehicle intelligent scheduling method for coping with extreme weather.

[0133] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a vehicle intelligent scheduling method for coping with extreme weather.

[0134] Figure 8 FIG. is a schematic hardware structure diagram of an electronic device for executing a vehicle intelligent scheduling method for coping with extreme weather provided by another embodiment of the present application. As Figure 8 shown, the device includes:

[0135] One or more processors 810 and a memory 820, Figure 8 Taking one processor 810 as an example.

[0136] The device for executing the vehicle intelligent scheduling method for coping with extreme weather may further include: an input device 830 and an output device 840.

[0137] The processor 810, the memory 820, the input device 830, and the output device 840 may be connected through a bus or other means, Figure 8 Taking connection through a bus as an example.

[0138] The memory 820, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the vehicle intelligent scheduling method for coping with extreme weather in the embodiments of the present application. The processor 810 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 820, that is, implements the vehicle intelligent scheduling method for coping with extreme weather in the above method embodiments.

[0139] The memory 820 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device. In addition, the memory 820 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 820 may optionally include a memory remotely provided with respect to the processor 810, and these remote memories may be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0140] The input device 830 can receive input digital or character information and generate signals related to user settings and function controls of the electronic device. The output device 840 can include a display device such as a display screen.

[0141] The one or more modules are stored in the memory 820 and, when executed by the one or more processors 810, perform the vehicle intelligent scheduling method for coping with extreme weather in any of the above method embodiments.

[0142] The above product can execute the vehicle intelligent scheduling method for coping with extreme weather provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0143] The electronic device in the embodiments of the present application exists in various forms, including but not limited to:

[0144] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0145] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0146] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.

[0147] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A vehicle intelligent scheduling method for coping with extreme weather, comprising: Obtaining the vehicle position of the target vehicle to be scheduled, the set of user positions corresponding to multiple platform users to be transported, real-time traffic data, and real-time weather data; The distance between each user position in the set of user positions and the vehicle position is less than a preset distance threshold; the real-time traffic data includes: road closure information, traffic accident information, and road congestion information; When it is determined that the real-time weather data includes extreme weather information, based on the first path planning model, determining at least one first candidate path that passes through the vehicle position and all the user positions; the extreme weather information includes any one of the following: heavy rain, heavy snow, freezing rain, or typhoon; For each of the first candidate paths, adjusting at least one path segment of the first candidate path based on a deep reinforcement learning model to obtain a corresponding corrected path; The path segment is determined according to any two user positions in the set of user positions; the state of the deep reinforcement learning model is defined according to the node position, real-time traffic data, and real-time weather data, and the action of the deep reinforcement learning model is defined according to the switching of the path segment; the node position is the vehicle position or the user position; Calculating the path comprehensive score corresponding to each of the corrected paths based on the MCDA evaluation model, and determining the rescue pick-up path of the target vehicle from each of the corrected paths to sequentially pick up each of the platform users.

2. The method according to claim 1, wherein After obtaining the vehicle position of the target vehicle to be scheduled, the set of user positions of multiple platform users to be transported, real-time traffic data, and real-time weather data, the method further includes: When it is determined that the real-time weather data does not include extreme weather information, based on the second path planning model, determining at least one second candidate path that passes through the vehicle position and any one of the user positions; Determining a passenger-carrying driving path from each of the second candidate paths to pick up and transport the platform user corresponding to the passenger-carrying driving path.

3. The method according to claim 1, wherein, The deep reinforcement learning model uses a deep Q network, and the main network of the deep Q network includes a cascaded input layer, hidden layer, and output layer; The input layer is used to perform normalization processing on the input data; the input data includes the real-time traffic data, the real-time weather data, and the first candidate path; The hidden layer includes a multi-layer convolutional neural network for extracting feature information from the input data; The output layer is used to determine the expected Q value corresponding to each action in the state space, and correct the first candidate path according to the path segment corresponding to the action with the maximum expected Q value to determine the corresponding corrected path.

4. The method according to claim 3, wherein The update formula of the deep Q network is as follows: Q(s,a;θ) = Q(s,a;θ) + α[r + γmax a′ Q(s′,a′;θ′) - Q(s,a;θ)] r = w time × r time + w safety × r safety Wherein, Q(s,a;θ) represents the current Q value of taking action a in state s, θ represents the current network parameters, α represents the learning rate; γ represents the discount factor, which indicates the importance of future rewards; max a′ Q(s′,a′;θ′) represents the maximum Q value that can be generated by all possible actions a′ in the next state s′, θ′ represents the old network parameters; r represents the immediate reward, which is obtained according to the result of the current action; r time and r safety represent the time-limited reward term and the safety reward term respectively, and w time and w safety represent the weight coefficients of the corresponding reward terms respectively, and w safety > w time , where w time promotes the model to select a path segment with a shorter distance, and w safety promotes the model to select a path segment with a lower probability of road condition anomaly risk under real-time weather data.

5. The method according to claim 4, wherein, The loss function of the deep Q network is: Temporal difference target y i The formula is as follows: y i = λ(s,a) × r + γ max a′ Q(s′, a′; θ - ) Wherein, Q(s,a;θ) represents the Q-value estimation of the current state-action pair, provided by the main network; Q(s′,a′;θ - ) represents the Q-value of the next state, estimated by the target network; θ and θ - respectively represent the parameters of the main network and the target network; represents the expected value of the square of the prediction error; λ(s,a) represents the environmental adaptability factor, which indicates the difficulty of taking a specific action a in a specific state s, and is used to adjust the weight of the immediate reward r.

6. The method according to any one of claims 1-5, wherein, The first path planning model includes a Dijkstra algorithm module, an A* algorithm module, and a Bellman-Ford algorithm module, wherein the Dijkstra algorithm module, the A* algorithm module, and the Bellman-Ford algorithm module are connected in parallel to the deep reinforcement learning model.

7. The method according to claim 6, wherein, The MCDA evaluation model calculates the comprehensive path score corresponding to the corrected path in the following manner: Where S route1 represents the comprehensive path score of the corrected path route1, w i represents the weight of the i-th evaluation criterion, and C i,route1 represents the normalized score of the corrected path route1 on the i-th evaluation criterion; the evaluation criteria include: driving timeliness evaluation criterion, path economic cost evaluation criterion, and road condition safety evaluation criterion; The calculation formula for the standardized score corresponding to the driving timeliness evaluation criterion is: α = 1 + k × WL Wherein, T route1 represents the expected travel time of the correction path route1, T min and T max respectively represent the minimum and maximum values among the expected travel times of all the correction paths; α represents an adjustment coefficient, and k represents a positive constant, which is used to adjust the influence degree of the bad weather level on the adjustment coefficient; WL represents the bad weather level, which is determined by querying the real-time weather data with the bad weather level mapping table. The bad weather level mapping table records multiple bad weather levels and corresponding weather data intervals; The calculation formula for the standardized score corresponding to the path economic cost evaluation criterion is: where E routel represents the expected economic cost of the correction path route1, and E max represents the maximum value among the expected economic costs of all correction paths; the expected economic cost of a path includes the energy consumption cost and the toll for the toll road sections involved in the path; The calculation formula for the standardized score corresponding to the road condition safety evaluation criterion is: Where m j is the weight of the j-th safety factor, and C j,route1,sf represents the score of the modified path route1 for the j-th safety factor; the safety factors include at least one of the following: road wetness condition, path historical accident records, path bend information, and path slope information.

8. A vehicle intelligent dispatching system for coping with extreme weather, comprising: A data acquisition unit configured to acquire the vehicle position of a target vehicle to be dispatched, a set of user positions corresponding to a plurality of platform users to be transported, real-time traffic data, and real-time weather data; the distance between each user position in the set of user positions and the vehicle position is less than a preset distance threshold; the real-time traffic data includes: road closure information, traffic accident information, and road congestion information; A path planning unit configured to, when determining that the real-time weather data includes extreme weather information, determine at least one first candidate path based on a first path planning model, the first candidate path passing through the vehicle position and all of the user positions; the extreme weather information includes any one of the following: heavy rain, heavy snow, freezing rain, or typhoon; A path correction unit configured to, for each of the first candidate paths, adjust at least one path segment of the first candidate path based on a deep reinforcement learning model to obtain a corresponding corrected path; The path segment is determined according to any two user positions in the set of user positions; the state of the deep reinforcement learning model is defined according to the node position, real-time traffic data, and real-time weather data, and the action of the deep reinforcement learning model is defined according to the switching of the path segment; the node position is the vehicle position or the user position; A pick-up path determination unit configured to calculate the comprehensive path score corresponding to each of the corrected paths based on an MCDA evaluation model, and determine the rescue pick-up path of the target vehicle from each of the corrected paths according to each of the comprehensive path scores to sequentially pick up each of the platform users.

Citation Information

Patent Citations

  • Sanitation robot vehicle scheduling method and system based on vehicle infrastructure cooperation and reinforcement learning

    CN116611635A

  • Vehicle path planning method and system based on reinforcement learning

    CN117711173A