Asynchronous agent multi-mode excitation joint optimization method based on reinforcement learning

By building a hierarchical agent architecture and asynchronous multi-agent reinforcement learning algorithm, low-carbon path induction and travel mode transfer incentives in multi-modal transportation networks are optimized, and the problem of insufficient incentive design in the existing traffic management methods is solved, achieving more efficient carbon emission reduction and travel cost optimization.

CN120299255AActive Publication Date: 2025-07-11BEIHANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510603793.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-11
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing traffic management methods are difficult to dynamically respond to the differentiated needs of different categories of travelers in multi-modal transportation networks. The incentive plan is insufficient, the carbon emission reduction assessment is inaccurate, the calculation complexity is high, and the real-time performance is poor, resulting in limited incentive effects and significant financial pressure.

Method used

The asynchronous agent multi-mode incentive joint optimization method based on reinforcement learning is adopted to build a hierarchical agent architecture, design an asynchronous multi-agent reinforcement learning algorithm, optimize low-carbon path induction and travel mode transfer incentives, and combine individual carbon trading mechanisms to dynamically adjust the incentive plan to reduce government fiscal investment.

Benefits of technology

The accuracy of incentive plans in multi-modal transportation systems has been improved, the benefits of carbon emission reduction have been enhanced, and the real-time response has been enhanced, which has reduced government fiscal pressure and optimized travel costs and carbon emissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299255A_ABST
    Figure CN120299255A_ABST
Patent Text Reader

Abstract

The invention relates to an asynchronous agent multi-mode excitation joint optimization method based on reinforcement learning, and the method comprises the steps: firstly constructing a layered agent architecture, enabling an upper-layer model agent to correspond to a multi-mode traffic network topology model, enabling a lower-layer model agent to correspond to a trip participant individual, and enabling the upper-layer model agent to correspond to a multi-mode traffic network topology model; actions, states and reward values of intelligent agents of the upper-layer model and the lower-layer model are defined respectively; and enabling the upper-layer intelligent agent to dynamically adjust the excitation scheme through iterative training by taking the minimum total cost as a target, and enabling the lower-layer intelligent agent to respond to excitation and feed back traffic state changes. According to the method, the effectiveness of the incentive scheme and the influence of the incentive scheme on individual travel decisions can be comprehensively evaluated, and the travel incentive scheme can be dynamically and quickly generated in the dynamic change situations of travel demands, the motorization rate and the carbon transaction price.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent transportation and artificial intelligence, and more specifically, to an asynchronous agent multi-mode incentive joint optimization method based on reinforcement learning. Background Art

[0002] With the acceleration of urbanization and the continuous growth of traffic demand, urban transportation systems face severe challenges: increased road congestion, decreased commuting efficiency, and rising carbon emissions. Traditional traffic management methods mostly focus on the optimization of a single traffic mode. For example, road congestion is alleviated through signal control, or bus travel is guided through fare adjustment.

[0003] However, such methods have significant deficiencies in the collaborative optimization of multi-mode transportation networks (such as roads, subways, and buses). Especially in the design of incentive schemes, existing technologies usually adopt static rules or single-dimensional incentive strategies (such as unified carbon credit rewards), which are difficult to dynamically respond to the differentiated needs of different types of travelers (fuel vehicle users, electric vehicle users, public transportation users), resulting in limited incentive effects.

[0004] In addition, existing carbon emission reduction assessment methods mostly rely on fixed emission factors and do not fully consider the real-time impact of traffic flow dynamics on carbon emission intensity. For example, in congested sections, the per-unit mileage carbon emissions of fuel vehicles increase significantly, but traditional models often use average values for estimation, resulting in calculation deviations in emission reduction amounts.

[0005] Meanwhile, there is a lack of effective linkage between the budget allocation of incentive schemes and carbon trading mechanisms, causing considerable financial pressure. Although some studies have attempted to introduce reinforcement learning to optimize traffic control, most of them use synchronous multi-agent frameworks, where all agents share the same policy network, making it difficult to coordinate the complex interactions between upper-layer incentive policy formulation and lower-layer user route selection, and prone to policy conflicts and low training efficiency problems.

[0006] Another defect of the existing technologies is that most models only consider time and cost costs, ignoring the costs of travel inconvenience (such as the impact of the number of transfers and walking distance on the user experience) and individual differences in the value of time. For example, commuters are usually more sensitive to time than leisure travelers, but existing schemes do not dynamically adjust the cost weights according to user categories, resulting in deviations between incentive policies and actual behavior responses. In addition, traditional optimization algorithms (such as dynamic programming) face bottlenecks of high computational complexity and poor real-time performance when solving the bi-level programming model of large-scale multi-mode transportation networks, and it is difficult to support minute-level decision-making requirements.

[0007] The above problems jointly lead to limitations in the existing incentive schemes in multi - modal transportation systems, such as low accuracy, limited carbon emission reduction benefits, and insufficient real - time response capabilities. Therefore, there is an urgent need for a new optimization method that can deeply integrate the characteristics of multi - modal networks, heterogeneous user behaviors, and dynamic carbon trading mechanisms. Summary of the Invention

[0008] In view of this, to at least partially solve the above - mentioned technical problems and make up for the deficiencies of existing methods, the present invention provides an asynchronous agent multi - modal incentive joint optimization method based on reinforcement learning. The aim is to consider the diverse travel decisions of individual travelers in a multi - modal transportation system and the complex interweaving between various travel incentive schemes, design two different travel incentive schemes, namely travel mode transfer and low - carbon travel path induction, and perform synchronous joint optimization in the same model; and considering the complex structure of the bilevel programming model and the complex interaction mechanism between the upper and lower layer models, design an asynchronous multi - agent reinforcement learning algorithm to solve the model, so as to solve the problem of insufficient exploration of the solution space by traditional heuristic solution algorithms.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] An asynchronous agent multi - modal incentive joint optimization method based on reinforcement learning, the steps include:

[0011] S1. Construct a hierarchical agent architecture, including an upper - layer model agent and a lower - layer model agent;

[0012] The upper - layer model agent corresponds to the multi - modal transportation network topology model. The action is defined as continuous travel mode transfer incentive values and low - carbon path induction incentive values. The state is defined as the carbon emissions, saturation, and total travel time of road segments. The reward value is defined as the weighted sum of the travel cost of road segments, carbon emission penalty, and budget penalty function;

[0013] The lower - layer model agent corresponds to individual travel participants. The action is defined as discrete travel mode or low - carbon path selection decisions. The state depends on the incentive values of the upper - layer agent, and the reward value is the individual generalized travel cost;

[0014] S2. With the goal of minimizing the total cost, the upper - layer agent dynamically adjusts the incentive scheme through iterative training, and the lower - layer agent responds to the incentive and feedbacks the changes in traffic states.

[0015] The asynchronous agent multi - modal incentive joint optimization method based on reinforcement learning publicly provided by the present invention, compared with the prior art, has the advantages that:

[0016] (1) Optimize the two types of incentives, namely low-carbon path incentive and travel mode transfer incentive, in the multi-mode transportation system. Compared with optimizing a single incentive plan in a single transportation system, it can comprehensively evaluate the effectiveness of the incentive plan and its impact on individual travel decisions.

[0017] (2) Fully consider the unique hierarchical structure and sequential solution paradigm of bilevel programming, and design a joint optimization method for travel incentive plans based on asynchronous multi-agent reinforcement learning algorithm, which can dynamically and quickly generate travel incentive plans in scenarios where travel demand, electrification rate, and carbon trading price change dynamically.

[0018] Another advantage of this application is that, compared with previous methods for travel incentive design, it integrates the individual carbon trading mechanism, and uses the carbon emission reduction volume trading generated by individual behavior changes to reduce the dependence of the incentive budget on government financial investment, and improves the sustainability of incentive implementation. Brief Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of the topological design of a multi-mode transportation network;

[0021] Figure 2 It is a design diagram of the agent action-critic network in the upper and lower layer models;

[0022] Figure 3 It is a flowchart of the asynchronous agent multi-mode incentive joint optimization method based on reinforcement learning of the present invention;

[0023] Figure 4 It is a schematic diagram of the multi-mode transportation dynamic allocation simulation environment in Beijing;

[0024] Figure 5 It is a diagram of the Q-value change during the training process of the asynchronous multi-agent reinforcement learning algorithm;

[0025] Figure 6 It is a diagram of the changes in travel cost and carbon emissions after the joint optimization of multiple travel incentive plans for multi-mode transportation in Beijing;

[0026] Figure 7 It is the solution result of the algorithm proposed by the present invention under the scenario of changing electrification rate;

[0027] Figure 8 It is the solution result of the algorithm proposed by the present invention under the scenario of changing carbon trading price. Detailed implementation manners

[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0030] The embodiments of the present invention disclose a method for jointly optimizing multi-mode incentives of asynchronous agents based on reinforcement learning. The steps include:

[0031] S1. Construct a hierarchical agent architecture, including an upper-layer model agent and a lower-layer model agent;

[0032] The upper-layer model aims to optimize two types of incentive schemes: travel mode transfer and low-carbon path induction incentives; the lower-layer model aims to optimize the generalized cost of individual travel and make optimal travel mode and path decisions. Among them,

[0033] The upper-layer model agent corresponds to the multi-mode traffic network topology model. The actions are defined as continuous travel mode transfer incentive values and low-carbon path induction incentive values. The states are defined as section carbon emissions, saturation, and total travel time. The reward value is defined as the weighted sum of section travel costs, carbon emissions, and budget penalty functions;

[0034] The lower-layer model agent corresponds to individual travel participants. The actions are defined as discrete travel mode or low-carbon path selection decisions. The states depend on the incentive values of the upper-layer agent. The reward value is the individual's generalized travel cost;

[0035] S2. With the goal of minimizing the total cost, the upper-layer agent dynamically adjusts the incentive scheme through iterative training, and the lower-layer agent responds to the incentives and feedbacks the changes in traffic states.

[0036] The present invention is a method for jointly optimizing multi-mode traffic multi-class travel incentive schemes based on asynchronous multi-agent reinforcement learning. First, a multi-mode traffic network model is constructed using the geographical information data of road, subway, and bus line networks. Secondly, considering the differences in travel decisions of different classes of travel individuals in the multi-mode traffic system, a customized generalized travel cost function for individuals is defined. Thirdly, a calculation method for the benchmark carbon emissions of the multi-mode traffic system is defined, and based on this, a calculation method for the carbon emission reduction of the multi-mode traffic system under the intervention of multi-class travel incentive mechanisms is proposed.

[0037] Then, a model for jointly optimizing multi-mode traffic multi-class travel incentive schemes is constructed with a bilevel programming model; and in combination with the established multi-mode traffic dynamic simulation environment, a joint optimization algorithm for multi-class travel incentive schemes of asynchronous multi-agent reinforcement learning algorithm is designed to obtain the final optimal joint incentive scheme.

[0038] In an alternative embodiment:

[0039] Geographical information data of road networks, subway line networks, and bus line networks are obtained from OpenStreetMap to construct a multi-mode traffic network topology model.

[0040] That is, traffic nodes are abstracted into bus stops, subway stations, and road intersections, and traffic links are abstracted into road segments and bus / subway line connection edges; with each subway station as the center and the maximum transfer walking distance of passengers as the radius, a buffer zone is constructed to generate transfer connection edges between subway stations and bus stops within the buffer zone. For example, when the walking distance between a bus stop and a subway station is less than 200 meters, they are connected with an edge; finally, a multi-mode traffic network for car, bus, and subway travel can be obtained, and its topological architecture schematic diagram is as Figure 1 shown.

[0041] G=(A s ,A b ,A r ,A bs ,N r ,N b ,N s )

[0042] In the formula, A s ,A b ,A r are the internal connection edges of the subway line network, bus line network, and road network respectively, A bs is the dependent connection edge between buses and bus stops, and N r ,N b ,N s are the node sets in the road network, bus network, and subway line network respectively.

[0043] In an alternative embodiment:

[0044] When constructing a hierarchical agent architecture, first, all road network segments and multi-modal public transportation networks in the upper-layer model are defined as upper-layer model agents, aiming to formulate travel incentive programs for the multi-modal transportation system. Participants who decide to take a low-carbon path for each segment in the road network can obtain a reward value, and participants who change their travel mode to take public transportation can obtain a reward value per kilometer; the individual participants in the lower-layer model are defined as lower-layer model agents, aiming to make decisions on travel mode and low-carbon path.

[0045] Secondly, the action of the upper-layer model agent is defined as the incentive value in the incentive program, which is a continuous real number space greater than 0; since the agents in the upper-layer model mainly focus on the carbon emissions and traffic conditions of the multi-modal transportation network, the state of the agents in the upper-layer model is defined as a triple of segment carbon emissions, saturation, and total travel time; the reward value is defined as a weighted sum composed of segment travel cost, carbon emissions, and penalty function.

[0046] The action of the lower-layer model agent is defined as a 0-1 decision variable, representing the travel mode change or low-carbon path decision of the agent. That is, when the action of the agent is 1, it means that the participant represented by the agent chooses to change the travel mode and take public transportation; otherwise, it means that the participant represented by the agent chooses to take a low-carbon path. Since the agents in the lower-layer model mainly focus on travel cost and the incentive program formulated by the upper-layer agents, the state is defined as the reward of the upper-layer agent, the travel time and carbon emissions of the lower-layer model agent; the reward value is defined as the generalized travel cost of the individual.

[0047] Furthermore, when performing asynchronous multi-agent reinforcement learning training, since the number of agents in the upper and lower layer models is large, it is difficult to create independent actor-critic networks for each agent for training, and the efficiency of the algorithm and the space complexity of training are extremely high. Therefore, during the training process, independent actor-critic networks are created for all agents in the upper-layer model and all agents in the lower-layer model respectively, and the unique encoding of the agent is used as the input of the actor network and the critic network to represent the different decisions of different agents.

[0048] The actor-critic networks of the upper-layer agent and the lower-layer agent designed in this application are structured as Figure 2 shown;

[0049] The action network of the upper-layer model agent is designed as a 4-layer stacked fully-connected network. The input is the state generated in the previous step, and the output is the continuous excitation value. The number of neurons in the hidden layers of the first three neural network layers is 32, 64, and 128 respectively, and the output of the last layer is the action of the upper-layer agent; the critic network of the upper-layer model agent is designed as a 3-layer fully-connected network. The input is the action output by the action network and the agent state obtained from the multi-modal traffic simulation environment, and the output is the reward Q value; the number of hidden neurons in the first two layers of this network is 128 and 64 respectively. In addition, the target network parameters and structure used by the upper-layer agent during training are the same as those of the critic network parameters and structure in the upper-layer model.

[0050] The action network of the lower-layer model agent is designed as a 5-layer stacked fully-connected network. The input is the state of the agent in the previous step of the lower-layer model and the action of the agent in the upper-layer model, and the output is the next discrete action; the number of neurons in the hidden layers of the first four neural network layers of this network is 64, 64, 32, and 32 respectively, and the output of the last layer is the action of the lower-layer agent; the critic network of the lower-layer model agent is designed as a 3-layer fully-connected network. The input is the output of the action networks of the agents in the upper and lower layer models and the agent state obtained from the multi-modal traffic simulation environment, and the output is the Q value of the lower-layer agent. The number of neurons in the first two hidden layers of this network is 16 and 16, and 9) the target network parameters and structure used by the agent in the lower-layer model during training are the same as those of the critic network parameters and structure in the lower-layer model.

[0051] In an alternative embodiment:

[0052] After taking the objective functions of the upper and lower layer models as the reward values of the agents in the upper and lower layer models, the reward function of each agent is obtained;

[0053] In this embodiment, the reward function of the upper-layer agent satisfies the following formula:

[0054]

[0055] In the formula, is the reward value of the upper-layer agent, τ1, τ2, and τ3 are the scaling factors of the reward function to ensure that the values of each part of the reward function are within the same dimension range, w1 and w2 represent weight parameters, represents the traffic flow of section l i l i represents the i-th section, i represents the section index, b represents the b-th type of commuter, r represents the starting point of the commuter's trip, s represents the ending point of the commuter's trip, represents the traffic flow cost of, EM b (l i ) represents section l iThe carbon emissions, ρ represents the penalty coefficient, and P(l i ) represents the penalty function; and A r represent the number of internal connecting edges of the road network and the set of road network sections respectively.

[0056] and represent the average flow of sections and the average travel cost of sections in the multimodal public transport network respectively; and represent the average carbon emissions of sections in the bus and subway network. ptn represents the public transport network, A bs and A sb represent the internal sections of the ground bus network and the internal sections of the subway network respectively.

[0057] Among them, the penalty function is determined based on the carbon emission reduction before and after the implementation of the incentive, aiming to minimize the difference between the income from carbon emission reduction trading generated by individual behavior changes and the total incentive budget, that is, to maximize the substitution of the individual carbon trading mechanism for the financial expenditure of the government's travel incentive plan. The expression is:

[0058]

[0059] In the formula, is the length of each section l of the multimodal transportation network i , r ptn is the reward value obtained per kilometer of public transport travel when the participant switches from car travel to using public transport, is the reward obtained on the road network section l when the participant travels on a low-carbon path i , represents the flow generated on section l of the commuter in the public transport network ptn when the mode transfer occurs, rn represents the road network, i represents the reward value on section l i , fc is the carbon trading price, RCE is the carbon emission reduction in the multimodal transportation network, and γ is the scaling factor of the reward value.

[0060] In this embodiment, the carbon emission reduction is calculated based on the difference in the equilibrium state carbon emissions before and after the implementation of the incentive, that is, the carbon emissions when the multimodal transportation network reaches the equilibrium state before the implementation of the incentive are used as the benchmark carbon emissions; the carbon emissions when the multimodal transportation network reaches a new equilibrium state after the implementation of the incentive are the carbon emissions affected by the incentive; according to the difference between the benchmark carbon emissions and the carbon emissions reduced by the incentive, the carbon emissions reduced in the multimodal transportation network affected by the incentive are obtained. The calculation formula is:

[0061] RCE = E0 - E1

[0062] ​

[0063] In the formula, \(E_0\) and \(E_1\) are respectively the total carbon emissions of the multi - modal transportation network when it reaches the equilibrium state before and after the implementation of the incentive, and are respectively the link flows of the multi - modal transportation network when it reaches the equilibrium state before and after the implementation of the incentive, and are respectively the carbon emissions per car of the road network when it reaches the equilibrium state before and after the implementation of the incentive, \(EF\) bs and \(EF\) sb are respectively the carbon emission factors of bus and subway trips, and \(RCE\) is the carbon emission reduction in the multi - modal transportation network.

[0064] In an alternative embodiment:

[0065] The reward function of the lower - layer agent satisfies the following formula:

[0066]

[0067] In the formula, represents the path - selection indicator variable, represents the travel cost of non - participants, represents the generalized travel cost of participants using low - carbon paths, represents the generalized travel cost of participants using public transportation; \(NINC\) represents participants who do not change their travel paths and travel modes, \(RINC\) represents participants who change their travel paths, and \(MINC\) represents participants who change their travel modes.

[0068] In this embodiment, first, the individual travel by car is divided into four categories according to the status of participating in the incentive project and the type of car owned: non - participants traveling by fuel - powered cars, non - participants traveling by electric cars, participants traveling by fuel - powered cars, participants traveling by electric cars, and individuals originally traveling by public transportation are regarded as the fifth - type individuals; then, considering the travel time, travel cost, travel convenience, and participant rewards of all travelers in the multi - modal transportation system, the generalized travel cost functions of different types of travelers are constructed, and the corresponding expressions are:

[0069]

[0070]

[0071] In the formula, are respectively the travel costs of individuals of types \(M2\), \(MGC1\), \(MEV1\). \(M2\) is the traveler in the public transportation network, and \(MGC1\) and \(MEV1\) are respectively the groups of cars traveling on the road that do not participate in the incentive, and are the generalized travel costs for participants traveling by public transportation and using low-carbon paths respectively. MGC3 and MEV3 are participants with fuel cars and electric cars respectively, t(x ptn (l i )) and t(x rn (l i )) are the travel times on section l i in the multi-modal public transportation network and the road network respectively. x ptn (l i ) and x rn (l i ) are the travel volumes on section l i in the multi-modal public transportation network and the road network respectively. and are the inconvenience costs for each section on the multi-modal public transportation network and the road network respectively. This cost is related to the traffic flow on the links in the multi-modal transportation network and is a monotonically non-decreasing function of its cost; is the length of each section l i in the multi-modal transportation network. is the average fare cost that an individual needs to pay per kilometer for traveling by the multi-modal public transportation network. g m is the fuel consumption cost per kilometer for car travel. is the reward obtained by the participant on section l i of the road network when using the low-carbon path. r ptn is the reward value obtained by the participant for each kilometer of public transportation travel when switching from car travel to public transportation. η is the value of time.

[0072] In an optional embodiment:

[0073] A bi-level programming model is used to jointly optimize two types of incentive schemes, namely low-carbon path induction incentive and travel mode transfer incentive, to minimize the travel cost and carbon emissions in the multi-modal transportation network. The objective function is expressed as:

[0074]

[0075] In the formula, is the path flow of the b-th type of user in the equilibrium state of the multi-modal transportation network. represents the cost of the path flow . E1 is the total carbon emissions after incentives; and

[0076]

[0077] In the formula, is the expected path cost between the origin r and the destination s. The travel demand of the b - type users between the starting point r and the ending point s.

[0078] The above - mentioned constraint conditions include:

[0079]

[0080] Among them, fc is the carbon trading price, and γ is the scaling factor of the reward value, which is used to represent that the integrated travel operator expands the actual value of the incentive through commercial promotion.

[0081] In an exemplary embodiment:

[0082] Build a multi - mode traffic dynamic allocation simulation environment and conduct asynchronous reinforcement learning training. The process refers to Figure 3 , including: building a dynamic traffic simulation environment based on the SUMO simulation software, and adopting an experience replay pool (D) and an asynchronous update mechanism for reinforcement learning training according to the established asynchronous neural network structure, parameters, and loss function to optimize the policy gradient and the loss of the critic network.

[0083] Among them, the loss of the critic network is determined by the real reward value of the agent in the upper and lower layer models and the loss output by the critic network. The calculation method is: the target value is composed of the current reward value and the predicted value of the next - state Q value multiplied by the discount factor; the critic network parameters are updated by minimizing the mean square error between the predicted Q value and the target value; the expression is:

[0084]

[0085] In the formula, and are the loss functions of the upper and lower layer critic networks respectively, and are the agent reward target values in the upper and lower layer models respectively.

[0086] The training of the agent action networks in the upper and lower layer models is based on the policy gradient algorithm, and the expected value of the long - term cumulative reward is maximized through gradient ascent. The expression is:

[0087]

[0088] In the above formula, and are the gradients of the agent objective functions in the upper and lower layer models respectively, and are the distributions of the agent states in the upper and lower layer models respectively, D is the experience replay pool, and are the deterministic policy networks of the upper and lower layer models respectively, and are the Q - functions of the upper and lower layer models respectively, and are the states of the agents in the upper and lower layer models respectively, and are the actions of the agents in the upper and lower layer models respectively.

[0089] One application mode of this application is as follows:

[0090] Taking 369 bus lines, 10 subway lines, 285 bus stops, 43 subway stops and 2310 road segments in Tongzhou District, Beijing as an example, a multi-modal traffic simulation environment is built based on SUMO, as Figure 4 shown;

[0091] Furthermore, using several benchmark reinforcement learnings such as the soft actor-critic (SAC) algorithm, the multi-agent deep policy gradient (MADDPG) algorithm, the hierarchical action-critic network (HAC), and the twin delayed deep deterministic policy gradient algorithm (TD3), the performance of the asynchronous multi-agent reinforcement learning (ASMADDPG) algorithm proposed in the present invention is evaluated, and the results are as Figure 5 and Figure 6 shown; the evaluation data is shown in Table 1;

[0092] Table 1

[0093]

[0094]

[0095] From Figure 5 and Figure 6 , and the data in the table, it can be seen that when the iteration converges, the carbon emissions of the multi-modal traffic system obtained by the ASMADDPG algorithm are 156.37 tons, which are 9.52%, 1.16%, 4.04% and 7.7% less than the carbon emissions when the benchmark algorithms SAC, TD3, MADDPG and HAC algorithms converge respectively. In addition, when the iteration converges, the travel cost of the multi-modal traffic system obtained by the ASMADDPG algorithm is also lower than the values when the benchmark algorithms SAC, TD3, MADDPG and HAC algorithms converge. At the same time, ASMADDPG achieves convergence after 2178 trainings. Compared with other algorithms, the total training time and the number of iterations are both shorter, showing a higher learning efficiency. These results all show that the ASMADDPG algorithm is superior to other benchmark algorithms in improving policy stability and optimization effect. The reinforcement learning designed in the present invention is superior to other types of reinforcement learning algorithms in terms of performance and efficiency. When the algorithm converges, the carbon emissions and travel costs in the multi-modal traffic system are the lowest.

[0096] Furthermore, the trained asynchronous multi-agent reinforcement learning algorithm is used to evaluate the carbon emission reduction and the change in the system travel cost brought about by the joint incentive scheme of travel mode transfer incentive and low-carbon path induction incentive in the multi-mode transportation system under the scenarios of electrification rate and carbon trading price changes. The results are as Figure 7 and Figure 8 shown. Figure 7 As can be seen from the shown results, as the penetration rate of electric vehicles and others increases, the carbon emission reduction and the travel cost savings obtained by the ASMADDPG algorithm gradually decrease. However, when the penetration rate of electric vehicles does not exceed 80%, the incentive scheme obtained by this dynamic solution still induces more than 5% of the participants to change their travel modes and more than 2% of the participants to change their travel paths to participate in low-carbon travel. This shows that the ASMADDPG algorithm can obtain reliable solutions under the scenario of dynamic changes in the electric penetration rate. Figure 8 As can be seen from the shown results, in the scenario without incentive budget, as the carbon trading price increases, the travel cost and carbon emissions of the multi-mode transportation system under the optimal incentive scheme gradually decrease. This shows that a higher carbon trading price can significantly improve the incentive effect obtained by the ASMADDPG algorithm. Generally speaking, the asynchronous multi-agent reinforcement learning algorithm designed in the present invention can achieve fast solution under various dynamic scenarios and obtain effective multi-category travel incentive schemes.

[0097] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0098] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An asynchronous agent multi-mode incentive joint optimization method based on reinforcement learning, characterized in that S1. Construct a hierarchical agent architecture, including an upper-layer model agent and a lower-layer model agent; The upper-layer model agent corresponds to a multi-mode traffic network topology model. The actions are defined as continuous travel mode transfer incentive values and low-carbon path induction incentive values. The states are defined as road section carbon emissions, saturation, and total travel time. The reward value is defined as the weighted sum of road section travel costs, carbon emissions, and budget penalty functions; The lower-layer model agent corresponds to individual travel participants. The actions are defined as discrete travel mode or low-carbon path selection decisions. The states depend on the incentive values of the upper-layer agent, and the reward value is the individual generalized travel cost; S2. With the goal of minimizing the total cost, the upper-layer agent dynamically adjusts the incentive scheme through iterative training, and the lower-layer agent responds to the incentives and feedbacks the changes in traffic states.

2. The optimization method according to claim 1, characterized in that Design an action-critic network structure for the upper and lower layer agents; The action network of the upper-layer agent is a four-layer fully connected neural network. The input is the current state, including road section carbon emissions, saturation, and total travel time, and the output is a continuous incentive value; the critic network is a three-layer fully connected neural network. The input is the current state and the corresponding action, and the output is the reward value; The action network of the lower-layer agent is a five-layer fully connected neural network. The input is the previous state of the current agent and the action of the agent in the upper-layer model, and the output is a discrete action; the critic network is a three-layer fully connected neural network. The input includes the upper-layer action and the current state, and the output is the value evaluation.

3. The optimization method according to claim 1, wherein The reward function of the upper-layer agent satisfies the following formula: In the formula, is the reward value of the upper-level agent, and L i represents the i-th agent of the upper-level model. τ1, τ2, and τ3 are the scaling factors of the reward function, and w1 and w2 are the weight coefficients. represents the traffic flow of section l i , and l i represents the i-th section. i represents the section index, b represents the type of commuter, r represents the starting point of the commuter's trip, and s represents the ending point of the commuter's trip. represents the traffic flow cost, and EM b (l i ) represents the carbon emission of section l i . ρ represents the penalty term coefficient, and P(l i ) represents the penalty function. represents the number of internal connection edges of the road network, and A r represents the set of road network sections. and respectively represent the average flow and average travel cost of sections in the multi-modal public transport network; and represent the average carbon emissions of sections in the bus and subway network. ptn represents the public transport network, and A bs represents the sections within the surface bus network, and A sb represents the sections within the subway network.

4. The optimization method according to claim 3, wherein The penalty function is determined based on the carbon emission reduction before and after the implementation of the incentive, and the expression is: Wherein, is the length of each section l of the multi - modal transportation network i r ptn is the reward value obtained per kilometer of public transportation travel when a participant switches from car travel to using public transportation is the reward obtained by a participant when traveling on a low - carbon path on section l of the road network i The subscript rn represents the road network indicates the traffic flow generated on section l of the public transportation network ptn when the travel mode transfer occurs to the commuter i The subscript l is the reward value on section l i c is the carbon trading price, RCE is the carbon emission reduction amount in the multi - modal transportation network, and γ is the scaling factor of the reward value 5. The optimization method according to claim 4, wherein The carbon emission reduction is calculated based on the difference in the equilibrium state carbon emissions before and after the implementation of the incentive. The formula is: RCE = E0 - E1 Where \(E_0\) and \(E_1\) are the total carbon emissions of the multi - modal transportation network at the equilibrium state before and after the implementation of the incentive respectively, and are the link flows of the multi - modal transportation network at the equilibrium state before and after the implementation of the incentive respectively, and are the carbon emissions per car of the road network at the equilibrium state before and after the implementation of the incentive respectively, \(EF\) bs and \(EF\) sb are the carbon emission factors for bus and subway trips respectively.

6. The optimization method according to claim 1, wherein The reward function of the lower-layer agent satisfies the following formula: In the formula, represents the route selection indicator variable, represents the travel cost of non - participants, x(l i ) represents the total flow on link l i ; represents the flow generated by non - incentive project participants on link l i in the road network, represents the flow generated by incentive project participants on link l i in the road network, represents the flow generated by incentive project participants on link l i in the public transport network. non represents non - incentive project participants, inc represents incentive project participants, represents the generalized travel cost for participants traveling on a low - carbon route, represents the generalized travel cost for participants traveling by public transport; NINC represents participants who do not change their travel routes and travel modes, RINC represents participants who change their travel routes, and MINC represents participants who change their travel modes.

7. The optimization method according to claim 4, characterized in that The generalized travel cost includes travel inconvenience cost, time, and consumption cost.

8. The optimization method according to claim 1, wherein The objective function with the goal of minimizing the total cost is: In the formula, is the path flow under the equilibrium state of the multi-mode transportation network, represents the path flow cost, and E1 is the total carbon emission after incentive.

9. The optimization method according to claim 2, wherein The loss function of the critic network is calculated as follows: The target value is composed of the current reward value and the predicted value of the next state Q value multiplied by the discount factor; the critic network parameters are updated by minimizing the mean square error between the predicted Q value and the target value; the training of the action network is based on the policy gradient algorithm, and the expected value of the long-term cumulative reward is maximized through gradient ascent.

Citation Information

Patent Citations

  • Intelligent network connection automobile carbon emission reduction method based on C-V2X vehicle-road cooperation

    CN118197072A

  • Multi-lane scene integrated energy-saving driving strategy optimization method based on deep reinforcement learning algorithm

    CN118707849A

  • Multi-mode public transport travel incentive strategy optimization method based on reinforcement learning

    CN118822059A

  • Method for predicting carbon dioxide concentration on roads and computing apparatus thereof

    KR102797919B1