Multimodal transport path optimization method for deep reinforcement learning

Through deep reinforcement learning agents to optimize multimodal transport paths, the problem of difficult to adapt to real-time traffic changes and insufficient trade-offs in the existing technology is solved, and efficient and dynamic path optimization and response efficiency are achieved.

CN120069723AActive Publication Date: 2025-05-30ZHEJIANG SIGANG LINKAGE DEV CO LTD

Patent Information

Application Number
CN202510550210.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing multimodal transport methods are difficult to adapt to real-time traffic changes, lack of multi-objective trade-offs, low response efficiency, and insufficient data utilization, resulting in delayed decision-making.

Method used

The multimodal transport path optimization method of deep reinforcement learning is adopted. By constructing a multimodal transport path network topology model, a deep reinforcement learning agent is constructed based on state space, action space and reward functions. The agent adopts the Sum Tree priority sampling mechanism combined with the uniform sampling + heavy weight sampling mechanism to select samples from the training samples, generate the optimal path, and dynamically update the optimal path during the intermodal process.

Benefits of technology

It has achieved dynamic adaptation to real-time traffic changes, improved the ability of multi-objective collaborative optimization, improved the response efficiency of emergencies, used historical transportation data and real-time information, reduced manual intervention, and improved the real-time and accuracy of decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069723A_ABST
    Figure CN120069723A_ABST
Patent Text Reader

Abstract

The invention discloses a deep reinforcement learning-based multimodal transport path optimization method, which belongs to the field of logistics management, and comprises the following steps of S1, forming a plurality of feasible paths from a starting point to an end point as training samples, calculating transportation time, freight cost, carbon emission, accident rate and maximum load index of each sub-path as quality evaluation indexes of the corresponding feasible path; s2, constructing a deep reinforcement learning agent based on the state space, the action space and the reward function; s3, selecting a sample by the intelligent agent by adopting a Sum Tree priority sampling mechanism, and generating an optimal path; and S4, in the intermodal process, monitoring the quality evaluation index of the path, and dynamically generating an optimal path according to the real-time path state data. The method can adapt to dynamic factors such as real-time traffic changes, manual intervention is reduced, the emergency response efficiency is high, the occurrence probability of accidents can be reduced, and carbon emission and emission of other pollutants are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of logistics management, and particularly relates to a multimodal transportation path optimization method based on deep reinforcement learning. Background Art

[0002] In a globalized economic system, effective logistics management is crucial for the competitiveness of enterprises.

[0003] With the increasing complexity of the global supply chain, multimodal transportation has become the core mode of modern logistics due to its flexibility and cost advantages. However, the existing technologies have the following problems: limitations of static models: traditional optimization methods (such as linear programming and heuristic algorithms) rely on fixed parameters and are difficult to adapt to dynamic factors such as real-time traffic changes; insufficient multi-objective trade-off: existing solutions mostly focus on a single objective (such as cost or time) and lack coordinated optimization of comprehensive indicators such as carbon emissions and transportation risks; low response efficiency: path adjustment in case of emergencies relies on manual intervention and cannot achieve automated dynamic replanning; insufficient data utilization: the integrated application of historical transportation data and real-time information is insufficient, resulting in lagged decision-making. Summary of the Invention

[0004] The purpose of the present invention is to provide a multimodal transportation path optimization method based on deep reinforcement learning to solve the problems existing in the existing multimodal transportation methods, such as difficulty in adapting to dynamic factors such as real-time traffic changes, insufficient multi-objective trade-off, and low response efficiency to emergencies.

[0005] To achieve the above purpose, the technical solution of the present invention is as follows: The present invention relates to a multimodal transportation path optimization method based on deep reinforcement learning, which includes the following steps: S1. Construct a multimodal transportation path network topology model, abstract the path network into a graph structure, and form multiple feasible paths from the starting point to the ending point as training samples, and calculate the transportation time, freight cost, carbon emissions, accident incidence rate, and maximum load index of each sub-path as the quality evaluation indicators of the corresponding feasible path; S2. Construct a deep reinforcement learning agent based on the state space, action space, and reward function; S3. The agent uses a combination of the Sum Tree priority sampling mechanism and the uniform sampling + reweighting sampling mechanism to select samples from the training samples and generate the optimal path; S4. During the multimodal transportation process, monitor the quality evaluation indicators of the path. If the quality evaluation indicators exceed the threshold, dynamically generate the optimal path according to the real-time path status data.

[0006] Preferably, the index of the transportation time of each path in S1 is calculated by the following formula: , Among them, d represents the transportation time, Psd represents the feasible path from the starting point s to the end point d ; eij represents the sub-path between adjacent path nodes i and j in the feasible path; The index of the freight cost of each path is obtained by calculating through the following formula: , wherein, b represents the freight cost; The index of the carbon emission of each path is obtained by calculating through the following formula: , wherein, l represents the carbon emission index; The index of the accident rate of each path is obtained by calculating through the following formula: , wherein, m represents the accident rate; The index of the maximum load of each path is obtained by calculating through the following formula: , wherein, n represents the maximum load.

[0007] Preferably, the specific steps for constructing the deep reinforcement learning agent of S2 based on the state space, action space and reward function are as follows: S2.1. Establish the state space of the agent. The state space includes two items of information, namely business request information and feasible path state matrix, which are represented by vectors, and the vector representation is: , wherein, is a vector representing the state at the current moment t , D is the business request information, including the starting point and the end point, TM is the feasible path state matrix corresponding to the business request information. The feasible path state matrix includes freight cost index, transportation time index, carbon emission index, accident rate index and maximum load index; The said feasible path state matrix is represented as: , where k is the number of feasible paths; S2.2. Establish the action space of the agent. The action space is used to represent the set of feasible paths, and it is represented by ε -[[]]END]]greed Make corresponding path decisions according to the strategy, and the decision-making method is as follows: , wherein, is the path optimization strategy made under the current path network state , a represents the action, x is a random number within the range of [0,1], ε is a variable that controls the random exploration probability with an initial value of 1, θ is the path network parameter of the common part; β and α are the unique parameters of the value function and the advantage function respectively, α is the learning rate; S2.3. Establish the reward function of the intelligent agent, and the reward function is expressed as: , wherein, R is the reward function, , , , and represent the weights corresponding to the freight cost index, transportation time index, carbon emission index, accident incidence index, and maximum load index respectively, and the sum of the weights is 1; S2.4. Define the way for the intelligent agent to select the optimal path. The specific way is: use the network state , path optimization strategy and reward to iteratively update the action value function, and select the optimal path strategy under each path set state with the goal of maximizing the expected reward value. The expression of the action value function is: , wherein, Q represents the expected reward value, is the immediate reward value for selecting the action in the state , represents the maximum reward value for selecting each action in the state , and γ is the discount factor, which reflects the importance of future rewards.

[0008] Preferably, the elements in the feasible path state matrix are normalized by the Min-Max method, and the normalization formula is: , wherein, x iis an element in the feasible path status matrix, corresponding to a certain transportation time index, freight cost index, carbon emission index, accident incidence rate index, and maximum load index in the matrix. is the minimum value in the corresponding index item. is the maximum value in the corresponding index item. is the index after normalization processing.

[0009] Preferably, the variable controlling the random exploration probability gradually decreases as the number of iterations increases, and the update formula is: , where, is the initial value of ε at the beginning of the iteration; represents the minimum value of ε; is the attenuation factor.

[0010] Preferably, the S2 constructs a deep reinforcement learning agent based on the Dueling DQN algorithm. The calculation of the expected reward value Q is divided into a value function and an advantage function, and its expression is: , where s represents the state, a represents the action, θ is the path network parameter of the common part; β and α are the unique parameters of the value function and the advantage function respectively. V is the value function, representing the value inherent in the state s, and A is the advantage function, representing the outstanding value of selecting the action a in the state s, is an index variable that traverses all possible actions.

[0011] Preferably, the specific way for the agent to select samples from the training samples using the Sum Tree priority sampling mechanism and the uniform sampling + reweighting mechanism is: S3.1. The agent uses the Sum Tree priority sampling mechanism to select samples: Calculate the priority of each sample, and select a set number of samples in descending order of priority. The formula for calculating the priority is: , where, i is the sample number, and each sample corresponds to a feasible path, p 1 is the priority, is the temporal difference error of the Sum Tree priority sampling mechanism; The formula for calculating the temporal difference error is: , where, The reward value generated for the target neural network; S3.2. The agent selects samples from the training samples using the uniform sampling + reweighting mechanism: All feasible path samples are stored in a circular queue, and a set number of samples are randomly selected uniformly each time training. The probability of each sample being selected is: , where N is the total capacity of the sample pool, i is the sample number, p 2 is the probability of the sample being selected; Then, for each selected sample, calculate its temporal difference error. The calculation formula is: , where, is the temporal difference error of the uniform sampling + reweighting mechanism, is the path reward estimate output by the target neural network, I is the current network estimate value, is the discount factor, is the sample i corresponding state, is the index variable for traversing all sample states, is an index variable for traversing all possible actions; Then, reweight the loss function. During the backpropagation process, dynamically adjust the loss weight according to the sample temporal difference error. The loss function is defined as: , where, φ is the loss weight, M is the number of batch samples, is the smoothing constant, used to prevent zero-error samples from having no weight, is the priority strength coefficient, w represents the weight, j represents a traversal of all values of an index label; S3.3. Combine the samples obtained in S3.1 and S3.2 to form a sample set.

[0012] Preferably, in S4, by calculating the real-time reward value and comparing it with the set reward value threshold, when the real-time reward value is less than the reward value threshold, it is determined that the quality evaluation index exceeds the threshold.

[0013] Preferably, the expression for updating the optimal path in S4 is: , where, represents Function p is the optimal path before update, and P is the set of feasible paths is the optimal path after update

[0014] Preferably, the feasible paths in S1 include the following constraints Time window constraint: the total path time does not exceed a preset threshold ; Load constraint: the weight of the goods ; Risk constraint: the accident rate , where h is the threshold of the goods load represents the threshold of the accident incidence rate

[0015] Adopting the technical solution provided by the present invention, compared with the prior art, it has the following beneficial effects 1. The multimodal transport path optimization method based on deep reinforcement learning involved in the present invention constructs a deep reinforcement learning agent based on the state space, action space and reward function, trains the feasible path samples by using the deep reinforcement learning agent to find the optimal path, and monitors the quality evaluation index of the path during the multimodal transport process. If the quality evaluation index exceeds the threshold, the optimal path is dynamically updated according to the real-time path state data, which can adapt to dynamic factors such as real-time traffic changes, reduce manual intervention, and has a high response efficiency to emergencies

[0016] 2. The multimodal transport path optimization method based on deep reinforcement learning involved in the present invention takes the transport time, freight cost, carbon emission, accident incidence rate and maximum load index as the quality evaluation indexes of the corresponding feasible paths, and incorporates the transport risk and carbon emission into the consideration of the path optimization scheme, which can identify potential risks in advance so as to take preventive measures in time, reduce the occurrence probability and its impact of accidents, reduce the carbon emission and the emission of other pollutants, promote green logistics and sustainable development, and achieve the purpose of reducing costs and improving service quality Description of the Drawings

[0017] Figure 1 is a flowchart of the multimodal transport path optimization method based on artificial intelligence involved in the present invention Detailed Embodiments

[0018] To further understand the content of the present invention, the present invention will be described in detail in combination with the embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention

[0019] Referring to the attached Figure 1 As shown, the present invention relates to a multimodal transport path optimization method based on deep reinforcement learning, which includes the following steps S1. Construct a multimodal transport path network topology model, abstract the path network as a graph structure G=(V, E), where V={v1, v2, …, vm} represents the set of all routing nodes, and |V| = m represents the number of path nodes; E=(e12, e23, ……, eij) represents the set of sub-paths between all adjacent nodes. Based on the starting point in the path network, form multiple feasible paths from the starting point to the ending point as training samples. Among them, each feasible path should follow the following constraints: Time window constraint: The total path time does not exceed a preset threshold ; Load constraint: The cargo weight ; Risk constraint: Accident rate , where h is the cargo load threshold, represents the accident incidence threshold.

[0020] After obtaining the feasible paths, calculate the transportation time, freight cost, carbon emissions, accident incidence rate, and maximum load index of each sub-path as the quality evaluation index (QE index) of the corresponding feasible path; The index of the transportation time of each path is calculated by the following formula: , where, d represents the transportation time, Psd represents the feasible path from the starting point s to the ending point d , eij represents the sub-path between adjacent path nodes i and j in the feasible path; The index of the freight cost of each path is calculated by the following formula: , where, b represents the freight cost; The index of the carbon emissions of each path is calculated by the following formula: , where, l represents the carbon emissions index; The index of the accident incidence rate of each path is calculated by the following formula: , where, m represents the accident incidence rate; The index of the maximum load of each path is calculated by the following formula: , Among them, n represents the maximum load.

[0021] S2. Construct a deep reinforcement learning agent based on the state space, action space, and reward function. The specific steps are as follows: S2.1. Establish the state space of the agent. The state space is a set of state information obtained by the agent from the path network. Each state contains two aspects of information, namely service request information (starting point - destination) and its corresponding feasible path state matrix. The state space includes two items of information, service request information and feasible path state matrix, which are represented by the vector s t as follows: , Among them, is a vector representing the state at the current moment t . D is the service request information, including the starting point and the destination. TM is the feasible path state matrix corresponding to the service request information. The feasible path state matrix includes freight cost indicators, transportation time indicators, carbon emission indicators, accident incidence indicators, and maximum load indicators. The described feasible path state matrix is represented as: , where k is the number of feasible paths.

[0022] Since the element values of the state matrix vary greatly and cannot objectively reflect the influence of each network path state information, the intelligent path selection algorithm fluctuates too much during the training process and is difficult to converge. Therefore, the elements in the feasible path state matrix are normalized to the range of [0, 1] through the Min - Max method. The normalization formula is: , Among them, x i is an element in the feasible path state matrix, corresponding to a certain transportation time indicator, freight cost indicator, carbon emission indicator, accident incidence indicator, and maximum load indicator in the matrix. is the minimum value in the corresponding item indicator. is the maximum value in the corresponding item indicator. is the indicator after normalization processing.

[0023] S2.2. Establish the action space of the agent. The action space is used to represent the set of feasible paths P = { P 1 , P 2 , …, Pk}, which makes corresponding path decisions through ε - greed the policy, and the decision-making method is as follows: , where is the path optimization strategy made under the current path network state , a represents the action, x is a random number within the range of [0, 1], ε is a variable that controls the random exploration probability with an initial value of 1, θ are the path network parameters of the common part; β and α are the unique parameters of the value function and the advantage function respectively, α is the learning rate; The variable that controls the random exploration probability gradually decreases as the number of iterations increases, and the update formula is: , where is the initial value of ε at the beginning of the iteration, represents the minimum value of ε, is the decay factor; This way enables the transportation entity to conduct random exploration with a high probability in the initial stage of iteration, exploring all possible situations as much as possible. As the number of iterations increases, the transportation entity will select the action with the maximum expected value with a higher probability.

[0024] S2.3. Establish the reward function of the agent. The reward function is the weighted QE index, expressed as: , where R is the reward function, , , , and represent the weights corresponding to the freight cost index, transportation time index, carbon emission index, accident incidence index, and maximum load index respectively, and the sum of the weights is 1; S2.4. By constructing the state space, action space, and reward function, design the way for the agent to interact with the path set, that is, define the way for the agent to select the optimal path. The specific method is: using the network state , path optimization strategy and reward to iteratively update the action value function, aiming to maximize the expected reward value, and select the optimal path strategy under each path set state . The expression of the action value function is: , wherein, Q represents the expected reward value, is the immediate reward value for selecting action in state , represents the maximum reward value for each action selected in state . γ is the discount factor, reflecting the importance of future rewards. In this embodiment, a deep reinforcement learning agent is constructed through the Dueling DQN algorithm. The calculation of the expected reward value Q is divided into a value function and an advantage function, and its expression is: , wherein, s represents the state, a represents the action, θ is the path network parameter of the common part; β and α are the unique parameters of the value function and the advantage function respectively. V is the value function, representing the value inherent in state s itself. A is the advantage function, representing the outstanding value of selecting action a in state s, is an index variable that traverses all possible actions; The improvement of Dueling DQN compared with the traditional DQN lies in dividing the Q-value calculation into a value function V(s; θ, β) and an advantage function A(s, a; θ, α), which respectively represent the value inherent in state s itself and the outstanding value of selecting action a in state s, and reducing the estimation bias of the Q-value by centralizing the advantage function to improve the stability of the algorithm.

[0025] S3. The agent selects samples from the training samples by combining the Sum Tree priority sampling mechanism with the uniform sampling + reweighting sampling mechanism to generate the optimal path. The specific method is as follows: S3.1. The agent selects samples by using the Sum Tree priority sampling mechanism: Calculate the priority of each sample, and select a set number of samples in descending order of priority. The calculation formula of the priority is: , wherein, i is the sample number, and each sample corresponds to a feasible path, p 1 is the priority, is the temporal difference error of the Sum Tree priority sampling mechanism; The calculation formula of the said temporal difference error is: , wherein, is the reward value generated by the target neural network; S3.2. The agent selects samples from the training samples using the uniform sampling + reweighting mechanism: All feasible path samples are stored in a circular queue. Each time training is performed, a set number of samples are randomly selected uniformly. The probability of each sample being selected is: , where N is the total capacity of the sample pool, i is the sample number, p 2 is the probability of the sample being selected; Then, for each selected sample, calculate its temporal difference error. The calculation formula is: , where, is the temporal difference error of the uniform sampling + reweighting mechanism, is the path reward stock index output by the target neural network, I is the current network estimated value, is the discount factor, is the sample i corresponding state, is the index variable for traversing all sample states, is an index variable for traversing all possible actions; Then, reweight the loss function. During the backpropagation process, dynamically adjust the loss weight according to the sample temporal difference error. The loss function is defined as: , where, φ is the loss weight, M is the number of batch samples, is the smoothing constant used to prevent samples with zero error from having no weight, is the priority strength coefficient, w represents the weight, j represents an index label for traversing all values; S3.3. Combine the samples obtained in S3.1 and S3.2 to form a sample set.

[0026] During the intermodal transportation process, monitor the quality evaluation index of the path. Specifically, calculate the real-time reward value and compare it with the set reward value threshold. When the real-time reward value is less than the reward value threshold, it is determined that the quality evaluation index exceeds the threshold. When it exceeds the threshold, generate an optimal path dynamically according to the real-time path status data; The expression for updating the optimal path is: , where, represents function,p is the optimal path before update, and P is the set of feasible paths, is the optimal path after update.

[0027] The present invention has been described in detail in conjunction with the embodiments, but the above content is only the preferred embodiments of the present invention and cannot be considered as limiting the scope of implementation of the present invention. Any equivalent changes and improvements made within the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.

Claims

1. A multimodal transport path optimization method based on deep reinforcement learning, characterized in that: It includes the following steps: S1. Construct a multimodal transport path network topology model, abstract the path network into a graph structure, and form multiple feasible paths from the starting point to the end point as training samples. Calculate the transportation time, freight cost, carbon emissions, accident rate and maximum load index of each sub-path as the quality evaluation index of the corresponding feasible path; S2. Construct a deep reinforcement learning agent based on state space, action space and reward function; S3. The agent uses the Sum Tree priority sampling mechanism combined with the uniform sampling + heavy weight sampling mechanism to select samples from the training samples and generate the optimal path; S4. During the intermodal transport process, the quality evaluation index of the route is monitored. If the quality evaluation index exceeds the threshold, the optimal route is dynamically generated based on the real-time route status data.

2. The multimodal transport path optimization method based on deep reinforcement learning according to claim 1, characterized in that: The transport time index of each path in S1 is calculated by the following formula: , in, d Indicates the transportation time, Psd Indicates starting point s To the end d The feasible path of eij Represents the adjacent path nodes in the feasible path i and j The subpaths between The freight cost index for each route is calculated using the following formula: , in, b represents the freight cost; The carbon emission index of each path is calculated by the following formula: , in, l Indicates the carbon emission index; The accident rate index for each path is calculated using the following formula: , in, m represents the accident rate; The maximum load index of each path is calculated by the following formula: , in, n Indicates the maximum load.

3. The multimodal transport path optimization method based on deep reinforcement learning according to claim 2, characterized in that: The specific steps of constructing a deep reinforcement learning agent based on state space, action space and reward function in S2 are: S2.

1. Establish the state space of the agent. The state space includes two pieces of information: business request information and feasible path state matrix, which are represented by vectors. The vector representation is: , in, is a vector representing the current time t The state of D is the business request information, including the starting point and the end point. TM is the feasible path state matrix corresponding to the business request information, and the feasible path state matrix includes freight cost index, transportation time index, carbon emission index, accident rate index and maximum load index; The feasible path state matrix is ​​expressed as: , Where k is the number of feasible paths; S2.

2. Establish the action space of the agent. The action space is used to represent the set of feasible paths. ε - greed The strategy makes corresponding path decisions in the following ways: , in, Current path network status The path optimization strategy made under a Indicates action, x is a random number in the range [0,1], ε is a variable with an initial value of 1 that controls the probability of random exploration. θ The path network parameters for the public part; β and α are the unique parameters of the value function and advantage function, respectively. α is the learning rate; S2.

3. Establish the reward function of the agent. The reward function is expressed as: , Among them, R is the reward function, , , , and They represent the weights of freight cost index, transportation time index, carbon emission index, accident rate index and maximum load index respectively, and the sum of the weights is 1; S2.

4. Define the way for the agent to select the optimal path, specifically: using the network status , Path Optimization Strategy and rewards Iteratively update the action value function to maximize the expected reward value and select each path set state The optimal path strategy under , the expression of the action value function is: , Among them, Q represents the expected reward value, For the status Next select action The immediate reward value, Indicates status The maximum reward value of each action is selected, and γ is the discount factor, which reflects the importance of future rewards.

4. The multimodal transport path optimization method based on deep reinforcement learning according to claim 3 is characterized by: The elements in the feasible path state matrix are normalized by the Min-Max method, and the normalization formula is: , in, x i is an element in the feasible path state matrix, corresponding to a transportation time indicator, freight cost indicator, carbon emission indicator, accident rate indicator and maximum load indicator in the matrix. is the minimum value of the corresponding index, is the maximum value of the corresponding index, It is the normalized index.

5. The multimodal transport path optimization method based on deep reinforcement learning according to claim 3, characterized in that: The variable controlling the random exploration probability gradually decreases as the number of iterations increases, and the update formula is: , in, is the initial value of ε at the beginning of the iteration; represents the minimum value of ε; is the attenuation factor.

6. The multimodal transport path optimization method based on deep reinforcement learning according to claim 3, characterized in that: The S2 constructs a deep reinforcement learning agent based on the Dueling DQN algorithm. The calculation of the expected reward value Q is divided into a value function and an advantage function, which are expressed as follows: , Among them, s represents the state, a Indicates action, θ The path network parameters for the public part; β and α They are the unique parameters of the value function and advantage function, V is the value function, which indicates the value of state s itself, and A is the advantage function, which indicates the action selected under state s. a The outstanding value of An index variable that iterates through all possible actions.

7. The multimodal transport path optimization method based on deep reinforcement learning according to claim 3, characterized in that: The specific way in which the agent selects samples from training samples using the Sum Tree priority sampling mechanism and the uniform sampling + reweighting mechanism is as follows: S3.

1. The agent uses the Sum Tree priority sampling mechanism to select samples: calculate the priority of each sample, and select a set number of samples in descending order of priority. The priority calculation formula is: , in, i is the sample number, each sample corresponds to a feasible path, p 1 is the priority, is the timing difference error of the Sum Tree priority sampling mechanism; The calculation formula of the timing differential error is: , in, The reward value generated by the target neural network; S3.

2. The agent uses uniform sampling + reweighting mechanism to select samples from training samples: all feasible path samples are stored in a circular queue, and a set number of samples are uniformly randomly selected during each training. The probability of each sample being selected is: , Where N is the total capacity of the sample pool, i is the sample number, p 2 is the probability of sample selection; Then, for each selected sample, the timing difference error is calculated using the following formula: , in, is the time difference error of uniform sampling + reweighting mechanism, is the path reward estimate output by the target neural network, I is the current network estimate, is the discount factor, For sample i Corresponding status, is the index variable for traversing all sample states, is an index variable that traverses all possible actions; Then, the loss function is reweighted. During the back propagation process, the loss weight is dynamically adjusted according to the sample time difference error. The loss function is defined as: , in, φ is the loss weight, M is the number of batch samples, is a smoothing constant used to prevent zero error samples from having no weight, is the priority intensity coefficient, w represents the weight, j Represents a traversal of all An index label for the value; S3.

3. Combine the samples obtained in S3.1 and S3.2 to form a sample set.

8. The multimodal transport path optimization method based on deep reinforcement learning according to claim 3, characterized in that: In S4, the real-time reward value is calculated and compared with the set reward value threshold. When the real-time reward value is less than the reward value threshold, it is determined that the quality evaluation index exceeds the threshold.

9. The multimodal transport path optimization method based on deep reinforcement learning according to claim 3, characterized in that: The expression of the optimal path updated by S4 is: , in, express function, p is the optimal path before updating, P is the set of feasible paths, is the updated optimal path.

10. The multimodal transport path optimization method based on deep reinforcement learning according to claim 1, characterized in that: The feasible path in S1 includes the following constraints: Time window constraint: the total path time does not exceed the preset threshold ; Loading restriction: cargo weight ; Risk Constraint: Accident Rate , Where h is the cargo load threshold, Represents the accident rate threshold.

Citation Information

Patent Citations

  • Dynamic power control method for deep double-Q network based on Sum tree sampling

    CN113795050A

  • Multi-target path planning method based on improved SAC algorithm

    CN116858248A

  • Data center network energy saving method and system based on reinforcement learning

    CN117240636A

  • Traveling salesman problem solving method combining deep reinforcement learning and heuristic algorithm

    CN119106778A

  • Unmanned ship dynamic path planning method and system based on deep reinforcement learning

    CN119396146A

Cited By

  • Carbon reduction path combination scheduling method based on reinforcement learning and related equipment

    CN120579792A

  • Multimodal transport path selection method based on multi-agent reinforcement learning

    CN120765154A

  • Disaster early warning method and device for large model in integrated delivery network, and storage medium

    CN121029970A