A road transport route planning method and system with weather adaptability

Through deep reinforcement learning algorithms and importance sampling technology, combined with weather data and transportation cost information, dynamic path optimization based on real-time data is achieved, solving the problem that existing systems are difficult to cope with complex weather conditions, significantly improving the reliability and accuracy of transportation decisions, and reducing the total transportation cost.

CN119721426BActive Publication Date: 2025-05-13HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221610.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing road transport route planning system is difficult to effectively deal with complex and changeable weather conditions, resulting in transportation delays or cargo damage, affecting supply chain stability and operational efficiency.

Method used

Deep reinforcement learning algorithms and importance sampling technology are adopted, combined with weather data and transportation cost information, dynamic path optimization based on real-time data is achieved, and the weight of data similar to the current situation is improved through neural networks, and transportation costs and time of different paths are accurately predicted.

Benefits of technology

It significantly improves the reliability and accuracy of decision-making, realizes effective response to complex weather conditions, reduces the total transportation cost, and provides an intelligent and efficient route planning method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721426B_ABST
    Figure CN119721426B_ABST
Patent Text Reader

Abstract

The present invention discloses a highway transportation route planning method and system with weather adaptability, which relates to the technical field of deep reinforcement learning. The method includes the following steps: S1. Input environmental information, where the environmental information includes basic information, weather prediction information, and historical data; S2. Initialize a training network, a target network, and an experience replay pool; S3. Initialize a state space, an action space, and a reward function; S4. Train the training network, and select an action a using a greedy strategy t ; Calculate the predicted reward value r t and the predicted transportation time h t of executing the action a by using the reward function in combination with the historical data and the weight matrix λ t ; S5. Sample extraction and training: Calculate the target value of the sample, and update the training network by using the target value and the first loss function; S6. Repeat S4 and S5, and synchronize the training network to the target network regularly; S7. Generate an optimal path. Finally, dynamic path optimization based on real-time data is realized to minimize the total transportation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep reinforcement learning technology, and in particular to a road transport route planning method and system with weather adaptability. Background Art

[0002] With the acceleration of globalization, logistics supply chains are becoming increasingly complex, especially in road transportation, which faces more and more challenges. Traditional transportation methods usually rely on fixed routes or simple cost optimization strategies. However, these methods often ignore the impact of external factors such as weather. For many goods that are sensitive to moisture and humidity, such as auto parts, food, medicine, textiles, etc., weather uncertainties may cause transportation delays or serious damage, which in turn affects the stability of the supply chain and the operational efficiency of enterprises.

[0003] At present, some transportation route planning systems have tried to introduce weather data, but most of them are still at the static analysis level based on historical data and fail to achieve dynamic response to real-time weather changes. Some solutions use simple rule models to use weather as auxiliary information and only re-plan routes when encountering extreme weather. For example, if a road is closed due to rain or snow, the system will switch to an alternative route according to preset rules. These rules rely on historical experience and the summary of common scenarios, and are suitable for transportation tasks with relatively stable environments or less weather changes. However, this method fails to deeply explore the laws of weather changes and cannot cope with complex and changing transportation environments.

[0004] Intelligent optimization methods based on reinforcement learning dynamically optimize transportation routes by utilizing a large amount of historical data and real-time feedback. Such methods usually learn patterns from historical weather data information and transportation results to predict the performance of different routes under specific conditions. For example, by analyzing past weather and traffic data, the system can estimate the travel time and potential risks of a certain route. In terms of technical implementation, the method is usually divided into two stages. In the first stage, a regression model or a classification model is used to model the environment, taking weather, traffic flow, and transportation time as input variables to generate preliminary route recommendations. This step mainly extracts environmental features through the analysis of historical data to provide support for subsequent decision-making. In the second stage, by building a reinforcement learning framework, the system experiments with different route selection options in a simulated environment and interacts with the environment to obtain feedback. Through continuous experimentation and adjustment, the model can learn the optimal strategy under complex conditions.

[0005] Reinforcement learning methods rely on the integration of multi-source data, including real-time weather information, traffic conditions, and transportation costs. However, their practical application is limited by data and computing resources. Such methods require high-quality and diverse data sets for training. If the data is insufficient or noisy, the model may be biased, affecting the reliability of decision-making. In addition, the training process of reinforcement learning models is complex and time-consuming, especially in the face of dynamically changing environments, making it difficult to quickly deploy and respond to emergencies. At the same time, when the environment changes too quickly, the model may not be able to adapt to new conditions in time, resulting in a decrease in the accuracy of its decision-making. Summary of the invention

[0006] In response to the problems existing in the prior art, the present invention provides a road transport route planning method and system with weather adaptability. By utilizing deep reinforcement learning algorithm and importance sampling technology, combined with weather data and transportation cost information, dynamic path optimization based on real-time data is achieved to minimize the total transportation cost.

[0007] The technical solution of the present invention is achieved in this way:

[0008] A road transport route planning method with weather adaptability comprises the following steps:

[0009] S1. Input environmental information; the environmental information includes basic information, weather forecast information and historical data; the historical data includes historical weather; the historical data includes serialized transportation records; the transportation records are represented as [first node, second node, historical weather, historical transportation time, historical transportation cost]; a transportation record represents the transportation time and transportation cost required to reach the second node from the first node under a certain historical weather constraint;

[0010] S2. Initialize the training network Q ω , target network Q ω - and experience replay pool; Q ω and Q ω - All are deep neural networks;

[0011] S3, initialize the state space S, action space A and reward function; the state space S includes multiple state nodes; the state node is composed of the basic information, including a starting node, several intermediate nodes and a destination node; that is, it represents the starting node and the destination node, as well as the intermediate nodes that may be passed from the starting node to the destination node; the action space A includes multiple actions; the action represents from one state node to another state node; the reward function includes a preset weight matrix λ; the weather forecast information is the predicted weather of the state node; the weather forecast information is expressed as sunny, light rain, heavy rain, etc.; in specific applications, various types of weather are represented by coding;

[0012] S4. Training the training network Q ω :For the current state node s t , using the greedy strategy to select action a t ; Execute action a t , get the next state node s t+1 ; Generate or optimize the weight matrix λ according to the historical data, and the reward function combines the historical data and the weight matrix λ to calculate the execution action a t The predicted reward value r t and the predicted transport time h t ; The weather forecast information participates in the optimization of the weight matrix λ;

[0013] The data(s t , a t , r t , s t+1 ) is stored in the experience replay pool; (s t , a t , r t , s t+1 ) can describe the process of state transfer.

[0014] S5, extracting samples for training: selecting several samples from the experience replay pool; calculating the target value of each sample respectively; updating the training network Q using the target value and the first loss function ω ;

[0015] S6. Update the target network Q ω - : Repeat steps S4 and S5, and regularly train the network Q ω Synchronize to the target network Q ω - ;

[0016] S7, generating the optimal path: using the target network Q ω - An optimal path from the starting node to the destination node is generated.

[0017] The calculation of the reward function is the application of importance sampling technology in this solution. Based on historical data, the current weather forecast information is fully considered. By increasing the weight of data similar to the current situation through the neural network, the model can more accurately predict the transportation cost and time of different routes, thereby significantly improving the reliability and accuracy of decision-making.

[0018] Predicted transportation time h t Although it is not stored in the state transfer pool, it is also stored and participates in subsequent calculations.

[0019] As a further optimization of the above solution, the basic information includes the delivery period.

[0020] The weather forecast information includes the forecasted weather of the status node during the delivery period.

[0021] As a further optimization of the above solution, in step S4, the training process sets the starting node as the initial state node;

[0022] The greedy strategy is expressed as:

[0023] ;

[0024] Wherein, ε represents the preset exploration rate; |A| represents the size of the action space; Indicates that the The biggest a.

[0025] ε gradually decays from 1 to close to 0 during training, which can be achieved by simulated annealing or other decay methods.

[0026] As a further optimization of the above scheme, the predicted reward value r t The calculation is to select all the historical data with (s t , s t+1 ) Combine the associated transport records and extract the time vector and the reward value vector ; Among them, s t and t+1 Map the first node and the second node respectively; j Corresponding to the historical transportation time, r j Corresponding to the historical transportation cost; j is the index, m is the length of the time vector; then:

[0027] ;

[0028] Among them, λ j That is the data of the weight matrix λ.

[0029] Through the neural network training weight matrix λ, historical samples containing matching predicted weather are given a higher sampling probability to be selected, and other unmatched historical samples are selected with a lower probability. In this scheme, the matrix involved is a one-dimensional matrix, that is, a vector representation.

[0030] As a further optimization of the above scheme, the second loss function is used Optimize the weight matrix λ, that is:

[0031] ;

[0032] Among them, y i Indicates whether the historical weather of the i-th transport record conforms to the weather forecast information; y i =1 means compliance, y i =0 means not in compliance.

[0033] L 正则化 Avoid excessive weight concentration on a single sample in historical selected data. β is a preset hyperparameter used to control the weight of the regularization term. L 标签 This means that a higher weight should be assigned to data that matches weather forecast information.

[0034] As a further optimization of the above scheme, the training network Q ω Includes input layer, hidden layer and output layer;

[0035] The input layer receives the state node and the action, that is, receives the encoding features of the current state s and the action a;

[0036] The hidden layer adopts a multi-layer fully connected network to extract the high-order relationship between the historical reward value and the historical weather label, and the hidden layer includes an activation function;

[0037] The output layer outputs the calculation results, specifically, the Q values ​​of all possible actions, that is, the expected value of each action selected in the current state;

[0038] The activation function adopts ReLU, which is used to calculate the historical transportation time h j and the historical transportation cost r j Calculate the output z j ;

[0039] The initialization calculation of the weight matrix λ is:

[0040] ;

[0041] Among them, exp() represents exponential function calculation.

[0042] ReLU stands for Rectified Linear Unit, which has the advantages of simple calculation, alleviating the problem of gradient disappearance, accelerating convergence, and sparse activation.

[0043] As a further optimization of the above scheme, in step S5, the sample is represented as (s, a, r, s'); for each of the samples, the target value q is calculated, and then the training network Q is updated. ω :

[0044] ;

[0045] Where γ is a preset discount factor used to weigh the impact of immediate rewards and future rewards; N is the number of samples extracted; a' represents the number of samples in the training network Q ω In, the action corresponding to s'; Indicates that the The largest ; s is the previous state of s'.

[0046] L(ω) is the first loss function.

[0047] As a further optimization of the above solution, in step S7, the target network generates several paths; among them, the paths whose total predicted transportation time exceeds the delivery period are discarded, and the remaining paths are the optimal paths.

[0048] That is, for multiple h on the same path t , if ∑h t >delivery period, the path is abandoned.

[0049] As a further optimization of the above scheme, in step S7, the target network generates several paths; among them, the paths whose total predicted transportation time exceeds the delivery period are discarded, and among the remaining paths, the path with the highest total predicted reward value is selected, which is the optimal path.

[0050] That is, for multiple h on the same path t , if ∑h t > delivery period, the path is discarded. Among the remaining paths, argmax(∑r t ), and obtain the optimal path.

[0051] The present invention also provides a road transport route planning system with weather adaptability, which applies the road transport route planning method with weather adaptability as described above.

[0052] Compared with the prior art, the present invention achieves the following beneficial effects:

[0053] The present invention proposes to combine importance sampling technology in path planning, based on historical data, and fully consider the current weather forecast information. By increasing the weight of data similar to the current situation through a neural network, the model can more accurately predict the transportation cost and time of different paths, thereby significantly improving the reliability and accuracy of decision-making. Specifically, the present invention combines weather data and transportation cost information to construct a fast-response system architecture, providing an intelligent and efficient route planning method for road transportation. Ultimately, while achieving dynamic path optimization based on real-time data, the total transportation cost is minimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a flow chart of a method for planning a road transport route with weather adaptability provided by an embodiment of the present invention;

[0055] Figure 2 It is a route topology diagram of basic information provided by an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of the exploration process of a neural network provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present invention more clear, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] like Figure 1 , Figure 2 As shown, this embodiment provides a road transport route planning system with weather adaptability, and applies a road transport route planning method with weather adaptability. The method includes the following steps:

[0059] S1. Input environmental information; environmental information includes basic information, weather forecast information and historical data;

[0060] In this embodiment, the basic information includes the delivery period, as well as the starting node, the destination node, the possible intermediate nodes and other related information; for example, the starting point of the transportation task is X, the end point is Y, and there are several possible intermediate nodes 1, 2, 3, 4, 5.

[0061] Weather forecast information includes the forecast weather of the state node during the delivery period; if the delivery period is 5 days, the forecast weather of a certain node for each of the 5 days. Specifically, the delivery period can also limit the total amount of transportation driving time per day, such as 8 hours or 10 hours per day.

[0062] Weather forecast information is expressed as sunny, light rain, heavy rain, etc.; in specific applications, various types of weather are represented by codes, such as "1: sunny", "2: light rain", "3: heavy rain".

[0063] Historical data includes serialized transportation records; transportation records are expressed as [first node, second node, historical weather, historical transportation time, historical transportation cost]. For example, "[X, 1, 2, 4, 150]", "[X, 3, 3, 5, 250]", "[3, Y, 1, 2, 100]", etc. "[X, 1, 2, 4, 150]" means that it takes 4 hours from node X to node 1, the cost is "150", and the weather during transportation is "light rain". The "150" here is just a number and can be customized according to needs.

[0064] A transportation record represents the transportation time and cost required to reach the second node from the first node under certain historical weather constraints.

[0065] S2. Initialize the training network Q ω , target network Q ω - and experience replay pool; Q ω and Q ω - All are deep neural networks; in this embodiment, the neural network includes an input layer, a hidden layer and an output layer;

[0066] The input layer receives state nodes and actions, that is, receives the encoding features of the current state s and action a;

[0067] The hidden layer uses a multi-layer fully connected network to extract the high-order relationship between historical reward values ​​and historical weather labels. The hidden layer includes an activation function.

[0068] The output layer outputs the calculation results. Specifically, it outputs the Q value of all possible actions, that is, the expected value of each action selected in the current state.

[0069] like Figure 3 As shown, in the input layer, state s will choose different directions, that is, action a n ; (s,a n ) sequence will train the neural network parameters, namely ω n ; Each (s, a n ) have their own corresponding reward value r, which is reflected in the output layer, namely Q (s, a n ).

[0070] S3, initialize the state space S, action space A and reward function; the state space S includes multiple state nodes; the state node is composed of basic information, including a starting node, several intermediate nodes and a destination node; that is, it represents the starting node and the destination node, as well as the intermediate nodes that may be passed from the starting node to the destination node; the action space A includes multiple actions; the action represents from one state node to another state node; the reward function includes a preset weight matrix λ;

[0071] S4. Training network Q ω ,During the training process, the starting node is set as the initial state node;

[0072] For the current state node s t , using the greedy strategy to select action a t .

[0073] In this embodiment, the greedy strategy is expressed as:

[0074] ;

[0075] Among them, ε represents the preset exploration rate; |A| represents the size of the action space; Indicates that the The biggest a.

[0076] ε gradually decays from 1 to close to 0 with training, which can be achieved by simulated annealing. For example, assuming that at node X, the neural network calculates that the Q value of going to node 1 is the largest, the system will choose to go to node 1 with the highest probability, and randomly choose to go to other nodes ("2" or "3") with a smaller probability.

[0077] Execute action a t , get the next state node s t+1 ; The reward function combines historical data and weight matrix λ to calculate the execution action a t The predicted reward value r t and the predicted transport time h t ;like Figure 3 As shown in the figure, in the input layer, the reward value r and constraint condition h in the relevant historical data are extracted according to (s, a), and the hidden layer extracts its high-order relationship to update the optimized weight matrix. In the output layer, the optimized weight matrix is ​​combined with the reward value r to calculate Q (s, a n ).

[0078] In this embodiment, the predicted reward value r t and the predicted transport time h t The calculation is to select all the historical data with (s t , s t+1 ) Combine the associated transport records and extract the time vector and the reward value vector ; Among them, s t and t+1 Map the corresponding first node and second node respectively; h j Corresponding to the historical transportation time, r j Corresponding to the historical transportation cost; j is the index, m is the length of the time vector; then:

[0079] ;

[0080] Among them, λ j That is the data of the weight matrix λ.

[0081] Through the neural network training weight matrix λ, historical samples containing matching predicted weather are given a higher sampling probability to be selected, and other unmatched historical samples are selected with a lower probability. In this scheme, the matrix involved is a one-dimensional matrix, that is, a vector representation.

[0082] Before the first prediction reward value calculation, the weight matrix λ needs to be initialized. Specifically, the activation function uses ReLU to calculate the historical transportation time h. j and historical transportation cost r j Calculate the output z j ;

[0083] The initialization calculation of the weight matrix λ is:

[0084] ;

[0085] Among them, exp() represents exponential function calculation.

[0086] ReLU stands for Rectified Linear Unit, which has the advantages of simple calculation, alleviating the problem of gradient disappearance, accelerating convergence, and sparse activation.

[0087] Use weather forecast information to update and optimize the weight matrix λ;

[0088] In this embodiment, the second loss function is used Optimize the weight matrix λ, that is:

[0089] ;

[0090] Among them, y i Indicates whether the historical weather of the i-th transport record is consistent with the weather forecast information; y i =1 means compliance, y i =0 means not in compliance.

[0091] L 正则化 Avoid excessive weight concentration on a single sample in historical selected data. β is a preset hyperparameter used to control the weight of the regularization term. L 标签 It means that for data that meets the weather forecast information, a higher weight should be assigned. The second loss function gradient is used to optimize λ.

[0092] Transfer the state (s t , a t , r t , st+1 ) Deposit into the experience replay pool;

[0093] In this embodiment, the experience replay pool has a maximum capacity. When the maximum capacity is reached, new samples will replace the oldest samples.

[0094] S5, sample extraction training: select several samples from the experience replay pool, the samples are represented as (s, a, r, s'); calculate the target value of each sample respectively; use the target value and the first loss function to update the training network Q ω ; Specifically, for each sample, calculate the target value q, and then update the training network Q ω :

[0095] ;

[0096] Among them, γ is the preset discount factor, which is used to weigh the impact of immediate rewards and future rewards; N is the number of samples extracted; a' represents the number of samples in the training network Q ω In, the action corresponding to s'; Indicates that the The largest ; s is the previous state of s'; L(ω) is the first loss function.

[0097] S6. Update target network Q ω - : Repeat steps S4 and S5, and regularly train the network Q ω Synchronize to target network Q ω - .

[0098] S7. Generate the optimal path: Using the target network Q ω - Generate the optimal path from the starting node to the destination node. In this embodiment, the target network generates several paths; among them, the paths whose total predicted transportation time exceeds the delivery period are discarded, and the remaining paths are the optimal paths. That is, for multiple h t , if ∑h t >delivery period, the path is abandoned.

[0099] The optimal route is not necessarily the shortest route, but it is the optimal balance between transportation distance and cargo damage after fully considering the impact of the cost of cargo damage caused by special weather (such as "rainy days").

[0100] The calculation of the reward function is the application of importance sampling technology in this solution. Based on historical data, the current weather forecast information is fully considered. By increasing the weight of data similar to the current situation through the neural network, the model can more accurately predict the transportation cost and time of different routes, thereby significantly improving the reliability and accuracy of decision-making.

[0101] The trained network parameters can be used in subsequent new training for further optimization. That is, the parameters of the neural network to be trained next time are initialized with the parameters of the currently saved neural network.

[0102] According to the disclosure and teaching of the above description, those skilled in the art to which the present invention belongs may also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the scope of protection of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for the convenience of description and do not constitute any limitation to the present invention.

Claims

1. A road transport route planning method with weather adaptability, characterized in that: The steps include: S1. Input environmental information; the environmental information includes basic information, weather forecast information and historical data; the historical data includes serialized transportation records; the transportation records are represented as [first node, second node, historical weather, historical transportation time, historical transportation cost]; S2. Initialize the training network Q ω , target network Q ω - and experience replay pool; S3, initialize the state space S, action space A and reward function; the state space S includes multiple state nodes; the state node is composed of the basic information, including a starting node, several intermediate nodes and a destination node; the action space A includes multiple actions; the action represents from one state node to another state node; the reward function includes a preset weight matrix λ; the weather forecast information is the predicted weather of the state node; S4. Training the training network Q ω :For the current state node s t , using the greedy strategy to select action a t ; Execute action a t , get the next state node s t+1 ; Generate or optimize the weight matrix λ according to the historical data, and the reward function combines the historical data and the weight matrix λ to calculate the execution action a t The predicted reward value r t and the predicted transport time h t ; The weather forecast information participates in the optimization of the weight matrix λ; The data (s t , a t , r t , s t+1 ) is stored in the experience replay pool; S5, extracting samples for training: selecting several samples from the experience replay pool; calculating the target value of each sample respectively; updating the training network Q using the target value and the first loss function ω ; S6. Update the target network Q ω - : Repeat steps S4 and S5, and regularly train the network Q ω Synchronize to the target network Q ω - ; S7, generating the optimal path: using the target network Q ω - An optimal path from the starting node to the destination node is generated.

2. A method for planning a road transport route with weather adaptability according to claim 1, characterized in that: The basic information includes the delivery period; the weather forecast information includes the forecasted weather of the status node within the delivery period.

3. A method for planning a road transport route with weather adaptability according to claim 1, characterized in that: In step S4, the training process sets the starting node as the initial state node; The greedy strategy is expressed as: ; Wherein, ε represents the preset exploration rate; |A| represents the size of the action space; Indicates that the The biggest a.

4. A method for planning a road transport route with weather adaptability according to claim 2, characterized in that: The predicted reward value r t and the predicted transport time h t The calculation is to select all the historical data with (s t , s t+1 ) Combine the associated transport records and extract the time vector and the reward value vector ; Among them, s t and t+1 Map the first node and the second node respectively; j Corresponding to the historical transportation time, r j Corresponding to the historical transportation cost; j is the index, m is the length of the time vector; then: ; Among them, λ j That is the data of the weight matrix λ.

5. A method for planning a road transport route with weather adaptability according to claim 4, characterized in that: Using the second loss function Optimize the weight matrix λ, that is: ; Among them, y i Indicates whether the historical weather of the i-th transport record conforms to the weather forecast information; y i =1 means compliance, y i =0 means non-compliance; β is the preset hyperparameter; L 正则化 Avoid excessive weighting on a single sample in historical selected data; L 标签 This means that a higher weight should be assigned to data that matches weather forecast information.

6. A method for planning a road transport route with weather adaptability according to claim 4, characterized in that: The training network Q ω Includes input layer, hidden layer and output layer; The input layer receives the state node and the action; The hidden layer adopts a multi-layer fully connected network, including an activation function; The output layer outputs the calculation result; The activation function adopts ReLU, which is used to calculate the historical transportation time h j and the historical transportation cost r j Calculate the output z j ; The initialization calculation of the weight matrix λ is: ; Among them, exp() represents exponential function calculation.

7. A method for planning a road transport route with weather adaptability according to claim 1, characterized in that: In step S5, the sample is represented as (s, a, r, s'); for each sample, the target value q is calculated, and then the training network Q is updated. ω : ; Wherein, γ is the preset discount factor; N is the number of samples extracted; a' represents the number of samples in the training network Q ω In , the action corresponding to s'; s is the previous state of s'; Indicates that the The largest ; L(ω) is the first loss function.

8. The method for planning a road transport route with weather adaptability according to claim 2, characterized in that: In step S7, the target network generates several paths; among them, the paths whose total predicted transportation time exceeds the delivery period are discarded, and the remaining paths are the optimal paths.

9. A method for planning a road transport route with weather adaptability according to claim 2, characterized in that: In step S7, the target network generates several paths; among them, the paths whose predicted transportation time sum exceeds the delivery period are discarded, and among the remaining paths, the path with the highest predicted reward value sum is selected, which is the optimal path.

10. A road transport route planning system with weather adaptability, characterized in that: A road transport route planning method with weather adaptability as described in any one of claims 1 to 9 is applied.

Citation Information

Patent Citations

  • Deep-sea mining robot path planning method based on deep reinforcement learning

    CN116339316A

  • Vehicle route planning method based on preference-driven multi-objective reinforcement learning

    CN118195457A