Reinforcement learning based asynchronous multi-agent multi-modal incentive joint optimization method
By constructing a hierarchical intelligent agent architecture and an asynchronous multi-agent reinforcement learning algorithm, the low-carbon path guidance and travel mode transfer incentives in a multimodal transportation system are optimized. This solves the problem of insufficient incentive design in existing traffic management methods, improves the accuracy and real-time performance of incentive schemes, and reduces the financial burden on the government.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2026-03-31
AI Technical Summary
Existing traffic management methods are unable to dynamically respond to the differentiated needs of different types of travelers in multimodal transportation networks. They suffer from inadequate incentive scheme design, inaccurate carbon emission reduction assessments, high computational complexity, and poor real-time performance, resulting in limited incentive effects and significant financial pressure.
We adopt a reinforcement learning-based asynchronous agent multi-mode incentive joint optimization method, construct a hierarchical agent architecture, design an asynchronous multi-agent reinforcement learning algorithm, optimize low-carbon path guidance and travel mode transfer incentives, and combine an individual carbon trading mechanism to dynamically adjust the incentive scheme to optimize the multi-modal transportation system.
It has improved the accuracy of incentive schemes in multimodal transportation systems, increased carbon emission reduction benefits, enhanced real-time response capabilities, reduced reliance on government incentive budgets, and improved the sustainability of incentive implementation.
Smart Images

Figure CN120299255B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent transportation and artificial intelligence, and more specifically to a multi-mode incentive joint optimization method for asynchronous intelligent agents based on reinforcement learning. Background Technology
[0002] With the acceleration of urbanization and the continuous growth of transportation demand, urban transportation systems face severe challenges: increased road congestion, decreased commuting efficiency, and rising carbon emissions. Traditional traffic management methods often focus on optimizing a single transportation mode, such as alleviating road congestion through traffic light control or guiding public transportation use through fare adjustments.
[0003] However, such methods have significant limitations in the collaborative optimization of multimodal transportation networks (such as roads, subways, and buses). In particular, in the design of incentive schemes, existing technologies typically employ static rules or single-dimensional incentive strategies (such as uniform carbon credit rewards), which are difficult to dynamically respond to the differentiated needs of different types of travelers (fuel vehicle users, electric vehicle users, and public transportation users), resulting in limited incentive effects.
[0004] Furthermore, existing carbon emission reduction assessment methods largely rely on fixed emission factors and do not fully consider the real-time impact of dynamic changes in traffic flow on carbon emission intensity. For example, in congested road sections, the carbon emissions per unit mileage of gasoline vehicles increase significantly, but traditional models often use average values for estimation, leading to errors in emission reduction calculations.
[0005] Meanwhile, the lack of effective linkage between incentive program budget allocation and carbon trading mechanisms has created considerable financial pressure. Although some studies have attempted to introduce reinforcement learning to optimize traffic control, most of them adopt a synchronous multi-agent framework, where all agents share the same policy network. This makes it difficult to coordinate the complex interaction between upper-level incentive policy formulation and lower-level user path selection, easily leading to policy conflicts and low training efficiency.
[0006] Another shortcoming of existing technologies is that most models only consider time and cost, neglecting the inconvenience costs of travel (such as the impact of transfers and walking distance on user experience) and individual differences in the value of time. For example, commuters are generally more sensitive to time than leisure travelers, but existing solutions do not dynamically adjust cost weights according to user categories, leading to a discrepancy between incentive strategies and actual behavioral responses. Furthermore, traditional optimization algorithms (such as dynamic programming) face bottlenecks of high computational complexity and poor real-time performance when solving bi-level programming models for large-scale multimodal transportation networks, making it difficult to support minute-level decision-making requirements.
[0007] The aforementioned problems collectively lead to limitations in existing incentive schemes in multimodal transportation systems, such as low accuracy, limited carbon emission reduction benefits, and insufficient real-time response capabilities. Therefore, there is an urgent need for a new optimization method that can deeply integrate the characteristics of multimodal networks, heterogeneous user behavior, and dynamic carbon trading mechanisms. Summary of the Invention
[0008] In view of this, in order to at least partially solve the above-mentioned technical problems and make up for the shortcomings of existing methods, this invention provides an asynchronous agent multi-mode incentive joint optimization method based on reinforcement learning. It aims to consider the diverse travel decisions of individual travelers in a multi-modal transportation system and the complex interweaving between various travel incentive schemes. It designs two different travel incentive schemes: mode shift and low-carbon travel path guidance, and performs synchronous joint optimization in the same model. Furthermore, considering the complex structure of the bi-level programming model and the complex interaction mechanism between the upper and lower levels, it designs an asynchronous multi-agent reinforcement learning algorithm and solves the model to solve the problem of insufficient solution space exploration in traditional heuristic solving algorithms.
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] A multi-mode incentive joint optimization method for asynchronous agents based on reinforcement learning, comprising the following steps:
[0011] S1. Construct a hierarchical intelligent agent architecture, which includes upper-layer model intelligent agents and lower-layer model intelligent agents;
[0012] The upper-layer model agent corresponds to a multi-modal transportation network topology model. The action is defined as a continuous travel mode transfer incentive value and a low-carbon path inducement incentive value. The state is defined as the carbon emission of the road segment, saturation and total travel time. The reward value is defined as the weighted sum of the road segment travel cost, carbon emission penalty and budget penalty function.
[0013] The lower-level model agent corresponds to an individual travel participant. The action is defined as a discrete travel mode or low-carbon route selection decision. The state depends on the incentive value of the upper-level agent, and the reward value is the individual's generalized travel cost.
[0014] S2. With the goal of minimizing total cost, the upper-layer agent dynamically adjusts the incentive scheme through iterative training, and the lower-layer agent responds to the incentive and reports the changes in traffic status.
[0015] The multi-mode incentive joint optimization method for asynchronous agents based on reinforcement learning disclosed in this invention has the following advantages compared with existing technologies:
[0016] (1) Optimizing two types of incentives, low-carbon path incentives and travel mode shift incentives, in a multimodal transportation system can comprehensively evaluate the effectiveness of incentive schemes and their impact on individual travel decisions compared to optimizing a single incentive scheme in a single transportation system.
[0017] (2) Taking into full account the unique hierarchical structure and sequential solution paradigm of bi-level programming, we design a joint optimization method for travel incentive schemes using asynchronous multi-agent reinforcement learning algorithms, which can dynamically and quickly generate travel incentive schemes in scenarios of dynamic changes in travel demand, electrification rate, and carbon trading price.
[0018] Another advantage of this application is that, compared with previous methods of designing travel incentives, it integrates an individual carbon trading mechanism. By trading carbon emission reductions generated by changes in individual behavior, it reduces the reliance of incentive budgets on government financial input and improves the sustainability of incentive implementation. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 A schematic diagram of the topology design for a multimodal transportation network;
[0021] Figure 2 Design diagram of agent action-comment network in upper and lower layer models;
[0022] Figure 3 This is a flowchart of the multi-mode incentive joint optimization method for asynchronous intelligent agents based on reinforcement learning, as described in this invention.
[0023] Figure 4 A schematic diagram of a simulation environment for dynamic multimodal traffic allocation in Beijing.
[0024] Figure 5 A graph showing the Q-value changes during the training process of an asynchronous multi-agent reinforcement learning algorithm;
[0025] Figure 6 A graph showing the changes in travel costs and carbon emissions after the joint optimization of Beijing's multimodal transportation and multi-type travel incentive schemes;
[0026] Figure 7 The solution results of the algorithm proposed in this invention under the scenario of changing electrification rate;
[0027] Figure 8 The solution results of the algorithm proposed in this invention are given under the scenario of carbon trading price changes. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0030] This invention discloses a multi-mode incentive joint optimization method for asynchronous agents based on reinforcement learning, comprising the following steps:
[0031] S1. Construct a hierarchical intelligent agent architecture, which includes upper-layer model intelligent agents and lower-layer model intelligent agents;
[0032] The upper-level model aims to optimize two types of incentive schemes: mode-of-travel shift and low-carbon route incentives; the lower-level model aims to optimize the generalized cost of individual travel and make optimal travel mode and route decisions; among them,
[0033] The upper-level model agent corresponds to the multi-modal transportation network topology model. The action is defined as the continuous travel mode transfer incentive value and the low-carbon path inducement incentive value. The state is defined as the carbon emission of the road segment, saturation and total travel time. The reward value is defined as the weighted sum of the road segment travel cost, carbon emission and budget penalty function.
[0034] The lower-level model agent corresponds to an individual travel participant. The action is defined as a discrete decision on the choice of travel mode or low-carbon route. The state depends on the incentive value of the upper-level agent, and the reward value is the individual's generalized travel cost.
[0035] S2. With the goal of minimizing total cost, the upper-layer agent dynamically adjusts the incentive scheme through iterative training, and the lower-layer agent responds to the incentive and reports the changes in traffic status.
[0036] This invention is a joint optimization method for multimodal transportation and multi-type travel incentive schemes based on asynchronous multi-agent reinforcement learning. First, a multimodal transportation network model is constructed using geographic information data of roads, subways, and bus networks. Second, considering the differences in travel decisions of different types of individuals in the multimodal transportation system, a generalized travel cost function for individuals is customized. Third, a method for calculating the baseline carbon emissions of the multimodal transportation system is defined, and a method for calculating the carbon emission reduction of the multimodal transportation system under the intervention of multiple travel incentive mechanisms is proposed.
[0037] Then, a model for joint optimization of multi-mode traffic and multi-type travel incentive schemes was constructed using a two-level programming model. In conjunction with the established multi-mode traffic dynamic simulation environment, an asynchronous multi-agent reinforcement learning algorithm was designed to jointly optimize multi-type travel incentive schemes, thereby obtaining the final optimal joint incentive scheme.
[0038] In one alternative embodiment:
[0039] Geographic information data of road networks, subway networks, and bus networks were obtained from OpenStreetMap to construct a multimodal transportation network topology model.
[0040] The traffic nodes are abstracted as bus stops, subway stations, and road intersections, and the traffic links are abstracted as road segments and bus / subway line connecting edges. A buffer zone is constructed with each subway station as the center and the maximum walking distance for passenger transfers as the radius. Transfer connecting edges are generated between the subway station and the bus stops within the buffer zone. For example, if the walking distance between a bus stop and a subway station is less than 200 meters, an edge is used to connect them. Finally, a multi-modal transportation network supporting car, bus, and subway travel is obtained, and its topology diagram is shown below. Figure 1 As shown.
[0041] G = (A s A b A r A bs N r N b N s )
[0042] In the formula, A s A b A r These are the internal connecting edges of the subway network, bus network, and road network, respectively. bs N represents the dependent connection edge between buses and bus stops. r N b N s These are the sets of nodes in the road network, bus network, and subway network, respectively.
[0043] In one alternative embodiment:
[0044] When constructing a hierarchical intelligent agent architecture, all road network segments and multimodal public transportation networks in the upper-layer model are first defined as upper-layer model intelligent agents, aiming to formulate travel incentive schemes for multimodal transportation systems and decide the reward value that participants in each road segment of the road network can obtain by adopting low-carbon routes and the reward value that participants can obtain per kilometer by changing their travel mode to use public transportation; individual participants in the lower-layer model are defined as lower-layer model intelligent agents, aiming to make travel mode and low-carbon route decisions.
[0045] Secondly, the actions of the upper-level model agent are defined as incentive values in the incentive scheme, which are continuous real numbers greater than 0. Since the agent in the upper-level model mainly focuses on the carbon emissions and traffic status of the multimodal transportation network, the state of the agent in the upper-level model is defined as a triple of road segment carbon emissions, saturation, and total travel time. The reward value is defined as a weighted sum of road segment travel cost, carbon emissions, and penalty function.
[0046] The actions of the lower-level model agent are defined as 0-1 decision variables, representing the agent's decision to change its travel mode or choose a low-carbon route. That is, when the agent's action is 1, it means that the participant represented by the agent chooses to change its travel mode and use public transportation; otherwise, it means that the participant represented by the agent chooses a low-carbon route. Since the agent in the lower-level model mainly focuses on travel costs and the incentive scheme formulated by the upper-level agent, the state is defined as the reward of the upper-level agent, the travel time of the lower-level model agent, and carbon emissions. The reward value is defined as the individual's generalized travel cost.
[0047] Furthermore, during asynchronous multi-agent reinforcement learning training, due to the large number of agents in both the upper and lower layers, creating an independent action-comment network for each agent is quite difficult, resulting in extremely high algorithm efficiency and training space complexity. Therefore, during training, independent action-comment networks were created for all agents in both the upper and lower layers. The unique agent encoding was used as the input to both the action and comment networks to represent the different decisions made by each agent.
[0048] The action-comment network for the upper and lower layer intelligent agents designed in this application has the following structure: Figure 2 As shown;
[0049] The action network of the upper-layer agent is designed as a four-layer stacked fully connected network. The input is the state generated in the previous step, and the output is a continuous stimulus value. The first three layers have 32, 64, and 128 hidden neurons, respectively, and the last layer outputs the action of the upper-layer agent. The comment network of the upper-layer agent is designed as a three-layer fully connected network. It takes the action output from the action network and the agent's state obtained from the multi-modal traffic simulation environment as input, and outputs a reward Q-value. The first two layers of this network have 128 and 64 hidden neurons, respectively. Furthermore, the target network parameters and structure used by the upper-layer agent during training are consistent with the parameters and structure of the comment network in the upper-layer model.
[0050] The lower-level model agent action network is designed as a 5-layer stacked fully connected network. Its inputs are the agent's previous state and the agent's action in the upper-level model, and its output is the next discrete action. The first four layers of this network have 64, 64, 32, and 32 hidden neurons, respectively, and the last layer outputs the action of the lower-level agent. The lower-level model agent comment network is designed as a 3-layer fully connected network. Its inputs are the agent's state obtained from the multi-modal traffic simulation environment, which is the output of the agent action networks in both the upper and lower levels, and its output is the Q-value of the lower-level agent. The first two hidden layers of this network have 16 and 16 hidden neurons, respectively. Furthermore, the target network parameters and structure used by the agent in the lower-level model during training are consistent with the parameters and structure of the agent comment network in the lower-level model.
[0051] In one alternative embodiment:
[0052] After using the objective functions of the upper and lower layer models as the reward values of the agents in the upper and lower layer models, the reward function of each agent is obtained.
[0053] In this embodiment, the reward function of the upper-layer agent satisfies the following formula:
[0054]
[0055] In the formula, Let τ1, τ2, and τ3 be the reward value for the upper-layer agent, τ1, τ2, and τ3 be the scaling factors of the reward function to ensure that the values of each part of the reward function are within the same dimension, and w1 and w2 represent the weight parameters. Indicates road segment l i Traffic, l i Let represent the i-th road segment, where i is the road segment index, b represents the b-th type of commuter, r represents the commuter's starting point, and s represent the commuter's ending point. Indicates flow rate Cost, EM b (l i ) indicates road segment l iCarbon emissions, ρ represents the penalty term coefficient, P(l i ) represents the penalty function; and A r These represent the number of connecting edges within the road network, respectively, and are represented by the set of road network segments.
[0056] and These represent the average traffic flow and average travel cost of a road segment in a multimodal public transport network, respectively. and The average carbon emissions per segment in the bus and subway network are represented by PTN, where PTN represents the public transportation network. bs and A sb These represent road segments within the surface public transport network and road segments within the subway network, respectively.
[0057] The penalty function is determined based on emission reductions before and after the incentive implementation. Its purpose is to minimize the difference between the gains from carbon emission reduction transactions resulting from individual behavioral changes and the total incentive budget, i.e., to maximize the substitution effect of the individual carbon trading mechanism on the government's fiscal expenditure on travel incentive programs. The expression is:
[0058]
[0059] In the formula, For each segment of the multimodal transportation network i The length, r ptn The reward value earned by participants for each kilometer of public transportation used to switch from car travel to public transportation. When participants use low-carbon routes for travel, in road network segments l i The rewards obtained on This indicates that the mode of occurrence has shifted to commuters on segments of the public transport network (PTN). i The traffic generated on the road network, where rn represents the road network. Indicates road segment l i The reward value is given by fc, where fc is the carbon trading price, RCE is the carbon emission reduction in the multimodal transport network, and γ is the scaling factor for the reward value.
[0060] In this embodiment, the carbon emission reduction is calculated based on the difference between the carbon emissions in the equilibrium state before and after the incentive implementation. Specifically, the carbon emissions of the multimodal transportation network when it reaches equilibrium before the incentive is implemented are used as the baseline carbon emissions; the carbon emissions of the multimodal transportation network when it reaches a new equilibrium state after the incentive is implemented are used as the carbon emissions after the incentive. The carbon emission reduction in the multimodal transportation network affected by the incentive is obtained based on the difference between the baseline carbon emissions and the carbon emission reduction affected by the incentive. The calculation formula is as follows:
[0061] RCE = E0 - E1
[0062]
[0063] In the formula, E0 and E1 represent the total carbon emissions of the multimodal transportation network when it reaches equilibrium before and after the implementation of the incentives, respectively. and These represent the traffic flow on road segments when the multi-modal transportation network reaches equilibrium before and after the implementation of incentives. and These represent the carbon emissions per car when the road network reaches equilibrium before and after the incentive implementation, EF. bs and EF sb These represent the carbon emission factors for public transport and subway travel, respectively, while RCE represents the carbon emission reductions in a multimodal transport network.
[0064] In one alternative embodiment:
[0065] The reward function of the lower-level agent satisfies the following formula:
[0066]
[0067] In the formula, This indicates a path selection indicator variable. This represents the travel costs for non-participants. This represents the generalized travel cost for participants using low-carbon routes. This represents the generalized travel cost of participants using public transportation; NINC represents participants who did not change their travel routes and modes of transportation, RINC represents participants who changed their travel routes, and MINC represents participants who changed their modes of transportation.
[0068] In this embodiment, individuals traveling by private car are first categorized into four groups based on their participation in the incentive program and the type of car they own: non-participants using gasoline-powered cars, non-participants using electric cars, participants using gasoline-powered cars, and participants using electric cars. Individuals who originally used public transportation are categorized into a fifth group. Then, considering the travel time, travel cost, travel convenience, and participant rewards for all travelers in the multimodal transportation system, a generalized travel cost function is constructed for each category of traveler. The corresponding expression is as follows:
[0069]
[0070]
[0071] In the formula, These represent the travel costs for individuals in categories M2, MGC1, and MEV1, respectively. M2 represents travelers on the public transportation network, while MGC1 and MEV1 represent private car travelers on the road who are not incentivized. and t(x) represents the generalized travel costs for participants using public transportation and low-carbon routes, respectively. MGC3 and MEV3 represent participants owning gasoline-powered cars and electric vehicles, respectively. ptn (l i )) and t(x rn (l i These refer to multimodal public transport networks and road segments within the road network. i Travel time on x ptn (l i ) and x rn (l i These refer to multimodal public transport networks and road segments within the road network. i Travel volume on the road and These are the travel inconvenience costs for each segment of the multimodal public transport network and the road network, respectively. These costs are related to the traffic flow of links in the multimodal transport network and are a monotonically non-decreasing function of their costs. For each segment of the multimodal transportation network i Length, g represents the average fare cost per kilometer that an individual needs to pay for using a multimodal public transport network. m The fuel cost per kilometer for car travel. When participants use low-carbon routes for travel, in road network segments l i The reward obtained above, r ptn The reward value for each kilometer of public transportation used by participants who switch from car travel to public transportation is η, where η is the time value.
[0072] In one alternative embodiment:
[0073] A bi-level programming model is used to jointly optimize two types of incentive schemes: low-carbon path-inducing incentives and mode-shifting incentives, in order to minimize travel costs and carbon emissions in a multi-modal transportation network. The objective function is expressed as:
[0074]
[0075] In the formula, This refers to the path flow of user type b under the equilibrium state of a multimodal transportation network. Represents path flow The cost, E1 is the total carbon emissions after incentives; and
[0076]
[0077] In the formula, Let r be the expected path cost between the starting point r and the ending point s. This represents the travel needs of user b between origin r and destination s.
[0078] The above constraints include:
[0079]
[0080] Where fc is the carbon trading price and γ is the scaling factor for the reward value, used to represent how integrated mobility operators expand the actual value of incentives through commercial promotion.
[0081] In one exemplary embodiment:
[0082] A multi-modal traffic dynamic allocation simulation environment was built and asynchronous reinforcement learning training was performed. The process is described in reference to... Figure 3 This includes: building a dynamic traffic simulation environment based on SUMO simulation software; using an empirical replay pool (D) and asynchronous update mechanism for reinforcement learning training according to a predetermined asynchronous neural network structure, parameters, and loss function; and optimizing the policy gradient and comment network loss.
[0083] The loss of the comment network is determined by the loss between the agent's actual reward value in the upper and lower layers of the model and the output loss of the comment network. The calculation method is as follows: the target value is composed of the current reward value multiplied by a discount factor and the predicted Q value of the next state; the comment network parameters are updated by minimizing the mean square error between the predicted Q value and the target value; the expression is:
[0084]
[0085] In the formula, and These are the loss functions for the upper and lower layers of the comment network, respectively. and These represent the reward target values for the agents in the upper and lower layers of the model, respectively.
[0086] The training of the agent action network in both upper and lower layers is based on the policy gradient algorithm, which maximizes the expected value of the long-term cumulative reward through gradient ascent. The expression is:
[0087]
[0088] In the above formula, and These are the gradients of the agent's objective function in the upper and lower layers of the model, respectively. and denoted by , where represents the distribution of the agent states in the upper and lower layers of the model, and D represents the experience replay pool. and These are the deterministic policy networks for the upper and lower layers of the model, respectively. and These are the Q-functions of the upper and lower layer models, respectively. and These represent the states of the agents in the upper and lower layers of the model, respectively. and These represent the actions of the agents in the upper and lower layers of the model, respectively.
[0089] One application of this application is as follows:
[0090] Taking 369 bus routes, 10 subway lines, 285 bus stops, 43 subway stations, and 2310 road sections in Tongzhou District, Beijing as an example, a multi-modal traffic simulation environment is built based on SUMO, such as... Figure 4 As shown;
[0091] Furthermore, the performance of the proposed asynchronous multi-agent reinforcement learning (ASMADDPG) algorithm is evaluated using several benchmark reinforcement learning algorithms, including soft actor-critic (SAC), multi-agent deep policy gradient (MADDPG), hierarchical action-critic network (HAC), and double-delay deep policy gradient (TD3). The results are as follows: Figure 5 and Figure 6 As shown; the evaluation data is shown in Table 1;
[0092] Table 1
[0093]
[0094]
[0095] from Figure 5 and Figure 6 As shown in the table, the ASMADDPG algorithm achieves a carbon emission of 156.37 tons for the multimodal transportation system upon iterative convergence, which is 9.52%, 1.16%, 4.04%, and 7.7% lower than the baseline algorithms SAC, TD3, MADDPG, and HAC, respectively. Furthermore, the ASMADDPG algorithm also achieves a lower travel cost for the multimodal transportation system upon iterative convergence compared to the baseline algorithms SAC, TD3, MADDPG, and HAC. Additionally, ASMADDPG converges after 2178 training iterations, exhibiting shorter total training time and fewer iterations compared to other algorithms, demonstrating higher learning efficiency. These results indicate that the ASMADDPG algorithm outperforms other baseline algorithms in both policy stability and optimization effectiveness. The reinforcement learning method designed in this invention surpasses other types of reinforcement learning algorithms in both performance and efficiency, achieving the lowest carbon emissions and travel costs in the multimodal transportation system upon convergence.
[0096] Furthermore, a pre-trained asynchronous multi-agent reinforcement learning algorithm was used to evaluate the changes in carbon emission reductions and system travel costs brought about by a joint incentive scheme of travel mode shift incentives and low-carbon route inducement incentives within a multimodal transportation system under the scenarios of changes in electrification rate and carbon trading price. The results are as follows: Figure 7 and Figure 8 As shown. Figure 7 The results show that as the penetration rate of electric vehicles increases, the carbon emission reductions and travel cost savings obtained by the ASMADDPG algorithm gradually decrease. However, when the electric vehicle penetration rate does not exceed 80%, the incentive scheme obtained by this dynamic solution still induces more than 5% of participants to change their travel modes and more than 2% of participants to change their travel routes, thus participating in low-carbon travel. This indicates that the ASMADDPG algorithm can obtain reliable solutions even under the scenario of dynamically changing electric vehicle penetration rates. Figure 8 The results show that, in the scenario without an incentive budget, the travel cost and carbon emissions of the multimodal transportation system under the optimal incentive scheme gradually decrease as the carbon trading price increases. This indicates that a higher carbon trading price can significantly improve the incentive effect obtained by the ASMADDPG algorithm. Overall, the asynchronous multi-agent reinforcement learning algorithm designed in this invention can achieve fast solutions in various dynamic scenarios and obtain effective multi-class travel incentive schemes.
[0097] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0098] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-mode incentive joint optimization method based on reinforcement learning asynchronous agent, characterized in that, S1, a hierarchical agent architecture is constructed, including an upper-layer model agent and a lower-layer model agent; The upper-layer model agent corresponds to a multi-mode traffic network topology model, the action is defined as continuous travel mode transfer incentive value and low-carbon path induction incentive value, the state is defined as road segment carbon emission, saturation and total travel time, and the reward value is defined as the weighted sum of road segment travel cost, carbon emission and budget penalty function; The reward function of the upper-layer agent satisfies the following formula: ; In the formula, is the reward value of the upper agent, represents the upper model agent, , and is the scaling factor of the reward function, and is the weight coefficient, represents the traffic of the road segment , represents the road segment, represents the road segment index, represents the commuter category, represents the commuter starting point, represents the commuter destination, represents the cost of traffic , represents the carbon emissions of the road segment , represents the penalty term coefficient, represents the penalty function; represents the number of internal connection edges of the road network, represents the road segment set of the road network; and denote the average flow and the average travel cost of a link in the multimodal public transport network, respectively; and denote the average carbon emission of a link in the bus and metro network, denote the public transport network, denote the internal links of the bus network, denote the internal links of the metro network; The penalty function is determined based on the emission reduction amount before and after the implementation of the incentive, and the expression is: ; In the formula, For each segment of the multimodal transportation network Length, The reward value earned by participants for each kilometer of public transportation used to switch from car travel to public transportation. When participants use low-carbon routes for travel, on road network segments The rewards obtained on Indicates the road network. This indicates a shift in the mode of occurrence to commuters on public transport networks. Middle section The traffic generated on Indicates road segment The reward value on the screen, For carbon trading prices, For carbon emission reductions in multimodal transportation networks, This is a scaling factor for the reward value; The carbon emission reduction amount is calculated based on the difference between the equilibrium state carbon emission before and after the implementation of the incentive, and the formula is: ; ; ; wherein, and are the total carbon emissions of the multi-modal transportation network when reaching equilibrium state before and after the implementation of the incentive, and are the link flows of the multi-modal transportation network when reaching equilibrium state before and after the implementation of the incentive, and are the carbon emissions of each car when the road network reaches equilibrium state before and after the implementation of the incentive, and are the carbon emission factors of bus and subway travel, respectively. The lower-layer model agent corresponds to a travel participant, the action is defined as a discrete travel mode or low-carbon path selection decision, the state depends on the incentive value of the upper-layer agent, and the reward value is the generalized travel cost of the individual; The reward function of the lower-layer agent satisfies the following formula: ; wherein, denotes the path selection indicator variable, denotes the travel cost of non-participants, denotes the total flow on link , denotes the flow generated by non-incentive program participants on link in the road network, denotes the flow generated by incentive program participants on link in the road network, denotes the flow generated by incentive program participants on link in the public transit network, denotes the non-incentive program participants, denotes the incentive program participants, denotes the generalized travel cost of participants using low-carbon paths, denotes the generalized travel cost of participants using public transit; denotes the participants who do not change travel paths and travel modes, denotes the participants who change travel paths, denotes the participants who change travel modes; S2, the upper-layer agent dynamically adjusts the incentive scheme by iterative training to minimize the total cost, and the lower-layer agent responds to the incentive and feeds back the traffic state change.
2. The optimization method of claim 1, wherein, An action-comment network structure is designed for the upper-layer and lower-layer agents; The action network of the upper-layer agent is a four-layer fully connected neural network, the input is the current state including road segment carbon emission, saturation and total travel time, and the output is continuous incentive value; The comment network is a three-layer fully connected neural network, the input is the current state and the corresponding action, and the output is the reward value; The action network of the lower-layer agent is a five-layer fully connected neural network, the input is the current state of the agent and the action of the agent in the upper-layer model, and the output is a discrete action; The comment network is a three-layer fully connected neural network, the input includes the upper-layer action and the current state, and the output is the value evaluation.
3. The optimization method of claim 1, wherein, The generalized travel cost includes travel inconvenience cost, time and consumption cost.
4. The optimization method of claim 1, wherein, The objective function with the minimum total cost as the target is: ; wherein is the path flow in the equilibrium state of the multimodal transport network, denotes the path flow of the cost is the total carbon emission after the incentive.
5. The optimization method of claim 2, wherein, The loss function of the comment network is calculated in the following way: the target value is composed of the current reward value and the predicted value of the next state Q value multiplied by the discount factor; Update the comment network parameters by minimizing the mean square error of the predicted Q value and the target value; The training of the action network is based on the policy gradient algorithm, which maximizes the expected value of the long-term cumulative reward by gradient ascent.
Citation Information
Patent Citations
Intelligent network connection automobile carbon emission reduction method based on C-V2X vehicle-road cooperation
CN118197072A
Multi-lane scene integrated energy-saving driving strategy optimization method based on deep reinforcement learning algorithm
CN118707849A