Urban road network subregion division method based on game-reinforcement learning cooperative control
By integrating multi-source data and employing game-reinforcement learning collaborative control, the sub-regional division of the urban road network is dynamically optimized, solving the problems of rigid regional division and unstable multi-agent decision-making in traffic control, and achieving efficient traffic management and collaborative control.
Patent Information
- Application Number
- CN202510877769.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-28
AI Technical Summary
Existing traffic control methods lack a global perspective and cannot adapt to real-time changes in traffic conditions, leading to rigid regional divisions and the spread of traffic congestion. Furthermore, reinforcement learning faces learning difficulties and instability issues in multi-regional control.
The system employs multi-source data fusion technology to perceive the road network status in real time. It combines Markov decision process and game-reinforcement learning collaborative control to dynamically optimize the regional division strategy. It introduces mean-field game theory to simplify multi-agent interaction and achieves collaborative decision-making between regions through policy gradient algorithm.
It significantly improves road network efficiency, reduces traffic delays, prevents congestion from spreading, enhances regional traffic balance and control flexibility, and has real-time adaptability and robustness.
Smart Images

Figure CN121034067A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent traffic control, and particularly relates to a city road network sub-region division method based on game-reinforcement learning collaborative control. BACKGROUND
[0002] With the acceleration of urban motorization, the problem of urban road traffic congestion is becoming increasingly serious. Traditional traffic control methods mainly focus on single points or local areas, lack a global perspective, and are difficult to respond to congestion spreading in a larger range in a timely manner. In recent years, the concept of active traffic control has gradually emerged, that is, by actively adjusting the control strategy before congestion forms to improve traffic flow. Existing adaptive signal control systems such as SCATS and SCOOT have introduced the concept of regional division, and the city is divided into several control sub-areas according to the physical topology of the road network. However, such division is usually static or manually set, and cannot adapt to real-time changes in traffic conditions. When the traffic pattern changes, the fixed regional boundary may lead to poor control effect, and even the phenomenon of queue spreading between regions and overflow at intersections.
[0003] The regional traffic management method based on the macroscopic traffic flow fundamental diagram shows that within a properly divided city sub-region, there is a stable relationship curve between the average travel speed of vehicles and the average density, which can be used to evaluate the congestion degree of the region and guide the regional flow control. However, if the internal traffic conditions are uneven, the MFD relationship will have large dispersion, making it difficult to accurately characterize the overall characteristics. Therefore, dynamically dividing the city road network into several sub-regions with relatively uniform internal traffic conditions can help obtain a clear MFD curve, thereby improving the effectiveness of regional traffic control. Existing research has proposed using community discovery algorithms, clustering analysis and other methods to dynamically partition the road network according to the correlation degree of intersections or the similarity of traffic flow. For example, some literature considers the traffic flow correlation coefficient and physical distance of adjacent intersections, establishes an index matrix and uses spectral clustering to divide the road network into sub-regions with similar traffic characteristics; some research uses fuzzy C-means or modularity maximization to divide the regions. These methods to some extent alleviate the problem of internal heterogeneity of the region, but still have shortcomings: on the one hand, most methods only use single source data, and do not fully integrate multi-source information to comprehensively characterize the traffic state; on the other hand, the division algorithm is usually based on pre-set rules or static optimization, and lacks a mechanism for closed-loop feedback adjustment according to traffic dynamics.
[0004] Meanwhile, reinforcement learning techniques have made progress in the field of traffic control. Reinforcement learning learns strategies autonomously through interaction with the environment, and can gradually approach optimal control without the need for an accurate model. In particular, deep reinforcement learning combines neural networks and can handle complex large-scale state spaces, showing potential advantages over traditional methods in problems such as traffic signal control. However, there are challenges in directly applying reinforcement learning to large-scale regional division and coordinated control: the traffic network has multi-agent interaction characteristics, and the control decisions of each region will affect other regions, making the environment non-stationary; in addition, the state and action space dimensions increase exponentially with the number of regions, making the learning process difficult. A single agent cannot effectively learn a large-scale coordinated control strategy, while multiple agents learning independently will be affected by each other, leading to unstable learning and even non-convergence.
[0005] To solve the above problems, researchers have begun to explore methods that combine game theory and reinforcement learning. One approach is to view multi-region control as a multi-agent game, where each agent adjusts its strategy by considering the strategies of other agents, thereby optimizing overall performance at the equilibrium point. Mean-field game theory is an effective method for handling large-scale games, which approximates the original multi-agent game by having a single representative agent play against the mean field, greatly reducing the complexity of analysis. Introducing mean-field theory into multi-agent reinforcement learning allows each agent to consider the impact of other agents as an average effect during the learning process, thereby reducing environmental non-stationarity, improving learning efficiency, and promoting convergence. In recent years, scholars in the field of traffic signal control have attempted mean-field multi-agent reinforcement learning methods, and the results show that they can improve the performance of multi-intersection coordinated control. It is inferred that incorporating game theory into the reinforcement learning framework is expected to enable the learning of coordinated control strategies for multiple traffic sub-regions. In summary, there is still room for improvement in existing technologies for urban road network regional division and control: a method is needed that can fuse multi-source data to perceive traffic status in real time, dynamically optimize regional division to improve traffic uniformity within regions, and automatically generate control strategies using reinforcement learning, while considering the mutual influence between multiple regions to achieve coordinated control. To address this technical gap, the present invention proposes a method for dynamic division of urban road network sub-regions based on multi-source fusion and game-reinforcement learning coordinated control. SUMMARY
[0006] The present invention aims to overcome the shortcomings of existing technologies and provide a method for dynamic division of urban road network sub-regions based on game-reinforcement learning coordinated control. By fusing multi-source traffic data to perceive road network status in real time, the method uses Markov decision processes and reinforcement learning algorithms to dynamically optimize regional division strategies. At the same time, the method introduces a game theory mechanism to enable coordinated decision-making by multiple regional control agents, simplifies complex interactions using mean-field games, and can adjust regional boundaries in real time according to traffic status changes, significantly improving road network traffic efficiency and control effectiveness.
[0007] To achieve the above object, the application provides a city road network sub-region division method based on game-reinforcement learning collaborative control, comprising the following steps:
[0008] 1) Obtain multi-source traffic data, fuse multiple data from detectors, floating cars and camera equipment, and estimate traffic state parameters of each road section of the road network in real time;
[0009] 2) Construct a multi-objective optimization model for dynamic division of city road network sub-regions to minimize traffic delay and regional non-uniformity, and the objective function is defined as follows:
[0010]
[0011] In the formula, F1 is a global traffic delay index, q i represents the number of queued vehicles on the i-th road section of the road network; F2 is a regional non-uniformity index, ρ i represents the vehicle density of the i-th road section, represents the average vehicle density in the r-th sub-region; α and β are weight coefficients; N is the total number of road sections; R is the number of sub-regions; y i,r is a distribution variable, y i,r = 1 when the i-th road section is divided into the r-th region, otherwise 0;
[0012] 3) Model the multi-objective dynamic division problem as a Markov decision process, define state space S, action space A, state transition probability P, immediate reward function R and discount factor γ; wherein the state includes traffic flow indicators of each sub-region, the action includes adjustment operations on the boundaries of the sub-regions, and the reward function is set as the negative value of the current multi-objective function value:
[0013] r(s,a)=-[αF1(s)+βF2(s)]
[0014] In the formula, s represents the current state, a represents the action, and the remaining symbols are the same as those in formula (2). The environment dynamics and decision objectives are described by the above MDP model.
[0015] 4) For the MDP, a reinforcement learning algorithm is used to solve the optimal division strategy, and a game theory collaborative mechanism is introduced to solve the multi-agent decision interaction problem; wherein each regional control agent independently learns the optimal action strategy using the policy gradient method, and is corrected according to the average strategy of adjacent regions after each round of decision. The policy gradient update is based on the following formula:
[0016]
[0017] In the formula, θ is the policy parameter, π θ(a|s) is the parameterized policy, J(θ) is the expected return, and R(s, a) is the reward value obtained by performing action a.
[0018] 5) The influence between multiple agents is described using the mean field game theory, the joint behavior of other agents is abstracted as the average effect; when updating the strategy, each regional control agent assumes that the average strategy of other regions remains unchanged, and calculates its optimal response to the average field. The mean field game equilibrium satisfies the following conditions:
[0019]
[0020] where Q(s, a) is the action value function, V(s') is the state value function, representing the average behavior of other agent strategies; μ(s) is the state distribution, and π * (a|s) is the probability distribution of selecting action a in state s under the current strategy. At equilibrium, the best strategy π * of each agent produces a state distribution μ' consistent with μ.
[0021] 6) Based on the macroscopic fundamental diagram theory, a sub-region traffic flow model is established, and the model function is obtained by fitting the relationship curve between the regional vehicle density and the outflow; a quadratic polynomial is used to fit the MFD curve, and the expression is as follows:
[0022]
[0023] where K r represents the average vehicle density of the rth region, Q r represents the out-of-region traffic flow of the rth region, and a and b are fitting coefficients which can be determined by the least squares method according to historical traffic data.
[0024] 7) The reinforcement learning decision is periodically executed to monitor the road network and adjust the sub-region division with a period T. When the non-equilibrium index of regional traffic is detected to exceed the threshold H th , the re-division calculation is triggered to update the regional boundary and generate new control strategy parameters.
[0025] Further, the multi-source traffic data includes fixed detector data, floating car trajectory data, and video detection data, and different source data is calculated using a weighted fusion method to calculate the road segment traffic state, wherein the fusion calculation formula of the road segment average driving speed is:
[0026]
[0027] where, is the speed of the ith road segment measured by the kth type of data source, and θ kis the corresponding data credibility weight, and
[0028] Further, the state S of the Markov decision process includes macroscopic traffic flow state quantities such as vehicle average speed, vehicle queue length, vehicle density, etc. of each sub-region, the action A includes an operation set of adjusting regional division, specifically, re-allocating the belonging sub-region for one or more road segments at the sub-region boundary, the reward function R is defined according to the formula in claim 1 to reflect the multi-objective optimization goal, and the discount factor γ is in the range of 0 to 1 to balance the short-term and long-term rewards.
[0029] Further, the reinforcement learning algorithm adopts a policy gradient reinforcement learning, uses a neural network to parameterize the policy, estimates the policy gradient based on the sampled trajectory data, and updates the policy parameters through gradient ascent; preferably, an actor-critic architecture is introduced to reduce variance and improve convergence speed, wherein the critic network estimates the state value function V(s), and the actor network outputs the action probability distribution π θ (a|s), the gradient update of which uses a baseline (the baseline is in the form of V(s):
[0030]
[0031] wherein V(s) represents the state value function baseline, and the remaining symbols are the same as the formula in claim 1.
[0032] Further, the multiple regional control agents make decisions and learn in parallel, each agent only interacts with the agents of adjacent regions through game playing, and the game analysis is simplified through mean field approximation; specifically, a neighborhood average action is introduced:
[0033]
[0034] wherein, is the neighborhood (set of adjacent sub-regions) of the i-th regional control agent, is the average value of the actions of other agents in the neighborhood. Each agent substitutes into its reinforcement learning algorithm as an approximate environment input to replace the global multi-agent influence.
[0035] Further, the fitting of the macroscopic traffic flow fundamental diagram MFD model uses historical data or offline simulation data, and uses a polynomial, logarithmic model or exponential model for curve fitting. In addition to the quadratic polynomial expression form in claim 1, the following Logistic model form can also be selected:
[0036]
[0037] wherein Qr,max K is the maximum flow rate of the rth sub-region, K r,c is the critical vehicle density when the maximum flow rate is reached, λ is the fitting coefficient. By fitting the above model parameters, the congestion evolution characteristics of each sub-region can be accurately described.
[0038] Further, the method further comprises setting a criterion for sub-region division adjustment, triggering a re-division when the traffic state difference between regions exceeds a predetermined threshold; the criterion is based on a regional traffic equilibrium index, such as triggering the controller to re-optimize when the global non-equilibrium index H exceeds the threshold H th , where H is defined as the average of the standard deviation of the traffic state of each region:
[0039]
[0040] where n r is the number of road segments in the rth sub-region, and the remaining symbols have the same meaning as in the formula in claim 1. The threshold H th is calibrated according to historical data to represent the allowable upper limit of the difference in traffic state within the region.
[0041] Further, at the beginning of each decision cycle, the latest traffic data is obtained and the MDP state is updated, at the end of each cycle, the division result of each sub-region is updated according to the executed action, and the policy network is trained through the reinforcement learning algorithm, so that it gradually approaches the optimal division strategy in the long-term operation; the decision cycle T is selected according to the traffic change rate, a smaller value is taken when the traffic state fluctuates frequently to improve the response speed, and a larger value is taken when the traffic is stable to reduce the computational load.
[0042] Further, the regional boundary control is further combined with signal timing optimization to achieve coordinated flow regulation between regions; specifically, after determining the sub-region division, the optimal critical vehicle number of each region is determined based on the MFD curve, and the actual vehicle number N r of each region is maintained near by peripheral signal timing control, and the control law can use proportional control:
[0043]
[0044] where u r (t) is the adjustment amount of the inflow of the rth region, K p is the control gain. When , the green ratio limit inflow is reduced to enter the region, otherwise the green ratio is increased to improve the inflow, thereby keeping each sub-region in the best operating state.
[0045] Further, the regional boundary control is further combined with signal timing optimization to achieve inter-regional traffic regulation in a coordinated manner; specifically, after determining the sub-region division, the optimal critical vehicle number of each region is determined based on the MFD curve and the peripheral signal timing control is adopted to adjust the actual vehicle number N r is maintained nearby, and the control law can adopt proportional control.
[0046] Compared with the prior art, the present application has at least the following beneficial effects:
[0047] (1) Real-time dynamic adaptability: by fusing multi-source traffic data to realize real-time perception of road network state, combining Markov decision process and game-reinforcement learning cooperative mechanism, the sub-region division strategy is dynamically optimized. Compared with the traditional static or artificially set region division, the present application can adjust the boundary in real time according to the traffic state fluctuation (such as peak period, accident, etc.), significantly improve the control flexibility and response speed, and effectively suppress the congestion diffusion.
[0048] (2) Multi-source data fusion improves the perception accuracy: the weighted fusion method is adopted to integrate multi-source data such as fixed detectors, floating cars and video detectors, to overcome the noise interference and coverage limitation of a single data source, accurately estimate the traffic parameters such as speed and density of the road section, provide reliable input for reinforcement learning decision, and ensure the scientificity and robustness of the division strategy.
[0049] (3) Multi-region cooperative control optimization: the mean field game theory is introduced to simplify the multi-agent interaction, and the complex multi-region decision problem is converted into an approximate model of a single agent and mean field game. Through the strategy gradient algorithm and the neighborhood average action mechanism, the strategy coordination between regions is realized, the strategy conflict and oscillation phenomenon of traditional multi-agent reinforcement learning are avoided, the calculation complexity is reduced, the overall traffic efficiency of the road network is significantly improved, and the global delay is reduced.
[0050] (4) Regional traffic flow balance and stability: based on the basic graph theory of macroscopic traffic flow, a sub-region model is established, combined with periodic dynamic division and signal timing optimization, to ensure that the vehicle density and flow in each sub-region maintain an optimal relationship. By adjusting the regional boundary and the entrance flow in real time, the local oversaturation is prevented, the internal traffic homogeneity of the region is improved, and the number of vehicle stops and queuing overflow is reduced.
[0051] These effects collectively solve the problems of one-sided data perception, rigid region division, multi-agent decision conflict and insufficient control stability in the prior art, and realize the efficient cooperative control of the urban road network. BRIEF DESCRIPTION OF DRAWINGS
[0052] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the application and, together with the specification, serve to explain the principles of the application.
[0053] The present application can be more clearly understood by referring to the following detailed description in conjunction with the accompanying drawings, in which:
[0054] Figure 1 is a flowchart of a city road network sub-region division method based on game-reinforcement learning collaborative control provided by an embodiment of the present application; the figure shows, in time sequence, steps such as multi-source data fusion, state evaluation, strategy calculation, region division adjustment and feedback update, and the relationship between the steps;
[0055] Figure 2 is a simulation schematic diagram of an embodiment of the present application; the figure takes a city road network as an example to show a comparison between the initial sub-region boundary and the new boundary after dynamic adjustment by the method of the present application during a peak period, different colors represent different control sub-regions, and arrows represent the vehicle flow direction between regions. DETAILED DESCRIPTION
[0056] The technical solutions of the present application will be described in detail below by means of the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, and not limitations of the technical solutions of the present application. In the case of no conflict, the technical features in the embodiments and the embodiments can be combined with each other. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and cannot limit the protection scope of the present application.
[0057] The term "and / or" in this paper is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents that the associated objects before and after are in an "or" relationship.
[0058] As shown in Figure 1 The present application provides a city road network sub-region division method based on game-reinforcement learning collaborative control, which comprises:
[0059] Step S1, obtaining and fusing multi-source traffic data, and estimating the traffic state parameters of each road section of the road network in real time as the basic input for the subsequent optimization modeling and decision-making process;
[0060] Step S2, based on the traffic state parameters obtained in step S1, constructing a multi-objective optimization model to minimize traffic delay and regional non-uniformity, obtaining the objective function and constraint conditions for region division optimization;
[0061] Step S3, the multi-objective optimization model constructed in step S2 is further modeled as a Markov decision process, a state space, an action space, a state transition probability, an immediate reward function and a discount factor are defined, and a reinforcement learning training environment is formed;
[0062] Step S4, an optimal partition strategy in the MDP established in step S3 is solved by using a reinforcement learning algorithm, a game theory cooperation mechanism is introduced to solve the multi-agent decision interaction problem, the influence among the multi-agents is described by using the mean field game theory, and the learning of the regional partition adjustment action strategy is guided;
[0063] Step S5, based on the traffic state parameters obtained in step S1 and the regional partition results obtained in step S4, a sub-regional traffic flow model is established, the macro traffic flow basic graph theory is used to fit the traffic flow characteristics in the region, and a basis is provided for subsequent control decisions;
[0064] Step S6, based on the sub-regional traffic flow model established in step S5 and the partition strategy learned in step S4, a periodic reinforcement learning decision is executed, when it is detected that the regional traffic non-equilibrium index exceeds a threshold, a re-partition calculation is triggered, the regional boundary is updated, and a new control strategy parameter is generated.
[0065] The above steps realize the real-time adaptability of the sub-regional partition of the road network, improve the state perception accuracy through multi-source data fusion, balance the global and regional optimization objectives by combining the game-reinforcement learning mechanism, and significantly improve the road network passing efficiency and balance.
[0066] Further, the multi-source traffic data in step S1 includes fixed detector data, floating car trajectory data and video detection data, different source data uses a weighted fusion method to calculate the road segment traffic state, effectively integrates the credibility of different source data, eliminates the noise interference of a single data source, and improves the accuracy and robustness of the road segment speed estimation, wherein the fusion calculation formula of the road segment average driving speed is:
[0067]
[0068] In the formula, is the speed of the ith road segment measured by the kth data source, θ k is the corresponding data credibility weight, and
[0069] Further, the objective function of the multi-objective optimization model is:
[0070] minαF1+βF2,
[0071] Wherein:
[0072] The constraint condition is:
[0073] where F1 is the global traffic delay indicator, q i denotes the number of queued vehicles on the i-th link; F2 is the regional imbalance indicator, p i denotes the vehicle density on the i-th link, denotes the average vehicle density within the r-th sub-region; a and b are weight coefficients; N is the total number of links; R is the number of sub-regions; y i,r is the assignment variable, y i,r = 1 if the i-th link is assigned to the r-th sub-region, otherwise 0.
[0074] Further, the immediate reward function of the Markov decision process is:
[0075] r(s, a) = - [aF1(s) + bF2(s)]
[0076] where s denotes the current state, a denotes the action, F1(s) and F2(s) are the global traffic delay indicator and the regional imbalance indicator under the current state, respectively, and a and b are weight coefficients. The multi-objective optimization problem is converted into a single-objective reward signal of reinforcement learning, ensuring that the algorithm optimizes the regional traffic state homogeneity while reducing the delay.
[0077] Further, other elements of the Markov decision process include:
[0078] The state space S is defined as the traffic state vector of all sub-regions in the road network within the current period, including but not limited to the average vehicle speed, vehicle density, number of queued vehicles, and imbalance indicator of each sub-region, which is obtained by real-time estimation of multi-source traffic data fusion results;
[0079] The action space is defined as the set of fine-tuning operations on the sub-region boundary structure, including operations such as dividing the boundary link to adjacent regions, merging or splitting sub-regions, etc. The action space is a discrete finite set, representing the structure adjustment actions that can be executed by each sub-region;
[0080] The state transition probability P(s' | s, a) is modeled by offline traffic simulation or historical traffic evolution data, which describes the probability distribution of the system transitioning to a new state s' after performing action a under the current state s. It can be updated by sampling during the training process;
[0081] The discount factor g e (0, 1) is a hyperparameter in the reinforcement learning algorithm, used to balance the immediate reward and long-term cumulative reward. Its value can be set in the interval [0.8, 0.99] according to the degree of attention to future rewards.
[0082] Further, the reinforcement learning algorithm adopts a policy gradient method, and the policy gradient update formula is:
[0083]
[0084] wherein, represents the state distribution induced by the current policy θ and the expected value is calculated under the joint distribution of actions selected by the policy θ in each state, and θ is the policy parameter, and is the parameterized policy, and
[0085] is the expected return, and is the return value obtained by performing the action a.
[0086]
[0087] wherein, represents the average behavior of other agent policies; and * is the probability distribution of selecting action a in state s under the current policy. Simplifying the multi-agent interaction into a neighborhood average field game reduces the computational complexity and avoids the oscillation phenomenon caused by inter-regional policy conflicts.
[0088] Further, the macroscopic traffic flow basic graph model adopts a quadratic polynomial fitting, and the expression is:
[0089]
[0090] wherein, r represents the average vehicle density of the rth region, and r represents the out-of-region traffic flow of the rth region, and a and b are fitting coefficients. By fitting the density-flow relationship with a low-order polynomial, the model accuracy and computational efficiency are considered, and an analytical mathematical basis is provided for regional flow control.
[0091] Further, the regional traffic disequilibrium index is:
[0092]
[0093] wherein, r is the number of road segments in the rth sub-region, and i represents the vehicle density of the ith road segment, represents the average vehicle density in the rth sub-region, and R is the number of sub-regions.
[0094] Further, the decision cycle T is adjusted according to the traffic variation rate, and a smaller value is taken when the traffic state fluctuates frequently, and a larger value is taken when the traffic is more stable.
[0095] Further, the area boundary control is combined with signal timing optimization, and the control law adopts proportional control:
[0096]
[0097] In the formula, u r (t) is the inlet flow regulation amount of the rth area, K p is the control gain, is the optimal critical vehicle number of the rth area, N r (t) is the actual vehicle number of the rth area.
[0098] The application provides a kind of urban road network sub-area dynamic division method, and the technical scheme of combining the technical scheme of the combination of multi-source data fusion, reinforcement learning and game collaborative control is adopted.The overall architecture includes three main parts of data fusion module, decision optimization module and execution feedback module.
[0099] In the data fusion module, the flow, speed data from fixed detectors, travel time data of floating car and queue length information obtained by video detection are collected together.Through data cleaning and fusion algorithm, the unified key traffic state parameters are obtained, including the average speed, flow, density and queue length of each road and the like.These multi-source fusion data provide comprehensive and accurate real-time traffic flow state description for subsequent decision.
[0100] The decision optimization module first calculates the parameters of the current multi-objective optimization model based on the fused data. This model considers two main optimization objectives: first, reducing global traffic delay, represented by the total number of queued vehicles (F1) across all road segments; and second, improving regional uniformity, represented by the sum of squares of vehicle density deviations within each sub-region (F2). The two objectives are combined into a single objective function using weighting coefficients α and β, thus transforming the problem into a standard single-objective optimization J = αF1 + βF2. The optimization variable is the regional division scheme, i.e., the sub-region identifier of each road, which must satisfy the constraint that each road belongs to only one region. To avoid the complexity of combinatorial optimization, this invention transforms the optimization problem into a sequential decision problem, modeled as a Markov decision process. In this model, the system state includes indicators such as average speed, density, and queue length in each region, and the action is defined as adjusting the region to which one or more boundary roads belong. After executing an action, the traffic state will evolve to a new state over a period of time, and the state transition can be observed through simulation models or real-time data feedback. Each action yields a reward value *r*. This invention designs the reward as the inverse of the objective function *J*, so as *J* decreases, the reward *r* increases. Reinforcement learning tends to generate policies that maximize cumulative reward. The core of the decision optimization module is to use reinforcement learning algorithms to solve the above MDP to obtain the optimal decision policy. Considering that the problem of this invention involves simultaneous decision-making by multiple regional control units, we adopt a multi-agent reinforcement learning framework, treating the control of each sub-region as an agent. These agents share the same reward, have a common objective, but their respective action spaces are different. To simplify the influence between multiple agents, this invention introduces the concept of mean-field game theory: assuming that each regional control agent plays against a "virtual opponent" representing the average behavior of all other agents. During policy learning, each agent perceives not the individual actions of other specific agents, but a comprehensive average action. As part of the environment, this is equivalent to compressing the originally high-dimensional multi-agent policy space into an average influence factor, thus significantly reducing dimensionality and randomness. Based on this idea, we designed a game-reinforcement learning collaborative algorithm: each agent simultaneously runs policy gradient RL, and at each policy update, the average action of each agent's neighborhood is first calculated based on the current policy set. Then, each agent assumes that its neighborhood behavior is fixed. Under these conditions, for their own policy parameters θ i The policy gradient is calculated and updated. Due to the mean-field approximation, each agent's policy update actually solves for the optimal response in a game with the average policy of its neighborhood, which can be viewed as approximating a Nash equilibrium. With multiple iterations, these policies will gradually converge to an equilibrium policy combination under which no single regional control can further improve global performance by unilaterally changing its policy.
[0101] In the execution of the feedback module, the application applies the learned control strategy to the actual traffic control, and continuously improves the model through feedback. Specifically, in each decision-making cycle, the system adjusts the regional division according to the current strategy. If the strategy output suggests keeping the existing division unchanged, it proceeds to the next cycle of monitoring; if it suggests adjustment, the boundaries of one or more sub-regions are modified accordingly. After the adjustment is completed, the signal timing on the regional boundaries is updated to control the flow between regions. To ensure the continuity of traffic flow, the application uses a peripheral traffic control strategy for related boundary intersections while adjusting the regional division: for example, based on the MFD curve of each region, the threshold N opt of the critical number of vehicles in each region is calculated, and by changing the green light timing of the boundary signals, the number of vehicles flowing into each region is limited or increased to maintain the actual number of vehicles N opt in each region near N . This can avoid excessive congestion or excessive idleness in a region, helping to improve overall balance. The above process continues to run online in a closed loop. It is worth noting that the reinforcement learning algorithm of the application can use heterogeneous strategy training, that is, while executing the current optimal strategy, the strategy network is continuously trained and improved using real environment data collected. In the initial stage, the system may divide based on historical experience or simple rules, and then gradually take over and optimize by reinforcement learning. After sufficient exploration and training, the strategy network will tend to be stable, at which point the system can automatically adjust the regional division for different traffic periods and events, such as appropriately expanding the boundaries of congested regions to control vehicle influx during morning and evening rush hours, merging regions to reduce unnecessary control overhead during off-peak hours, or quickly isolating the region where the affected road segment is located in the event of a sudden accident, etc.
[0102] The application realizes adaptive optimization of sub-regional division and multi-regional collaborative control of urban road network through the method of fusing multi-source data, reinforcement learning and game theory, and has the following beneficial effects:
[0103] 1. Real-time adaptability: Compared with traditional static division schemes, the application can adjust the regional boundaries in real time according to changes in traffic conditions, ensuring that the traffic flow characteristics within the region remain stable. This is particularly important when traffic patterns change, as dynamic division can prevent congestion from becoming overly concentrated or spreading, improving the flexibility of traffic control.
[0104] 2. Accurate perception of multi-source data: By fusing traffic data from multiple sources, the application has a more comprehensive and accurate perception of road network conditions. Compared with a single data source, multi-source fusion reduces perception errors and blind spots, providing a reliable foundation for decision-making algorithms, enabling the division decision to be optimized for actual traffic conditions.
[0105] 3. Good cooperative control effect: using the game-reinforcement learning cooperative mechanism, the sub-region control agent not only pays attention to the congestion situation of itself, but also considers the influence on other regions through the mean field game, so as to realize the overall optimization. Compared with the isolated control method of each region, the application significantly reduces the conflict between regions and improves the global traffic performance index. According to the simulation test results, the method of the application can reduce the average delay of the whole network by about 10-20%, effectively preventing the queue overflow caused by over-saturation of a region.
[0106] 4. Robustness and scalability: the introduction of the mean field theory makes the complexity of the algorithm approximately independent of the number of agents, so the application can be extended to large-scale road networks. The self-learning ability of the reinforcement learning framework enables it to adapt to different cities, different road network structures and variable traffic demand without manual parameter adjustment. Even in the case of partial data loss or abnormality, the algorithm can still be corrected through continuous learning, and has good robustness.
[0107] Embodiment 1
[0108] The embodiment provides a specific implementation process of a sub-region division method for urban road network cooperative control based on game-reinforcement learning:
[0109] Step one: multi-source data acquisition and fusion, specifically including:
[0110] 1) Multi-source traffic data collection and preprocessing. First, real-time traffic data of urban road network is collected through multiple sources such as fixed detectors, floating car GPS and video monitoring, such as traffic flow, speed, queue length, trip OD information, etc. Synchronize and filter different time resolution and spatial distribution data, and eliminate outliers; then, map the data to a unified road network model to obtain the key traffic state parameters of each road segment in the current period.
[0111] 2) Data preprocessing and time-space alignment. Align the original data of different time granularity according to the unified time axis, for example, take 1 minute or 5 minutes as the basic time period; match the information from different spatial resolution to the unified road network model to ensure that the same road segment has consistent ID identification. Filter or interpolate missing or noise values to eliminate the influence of missing or noise.
[0112] 3) Multi-source data fusion. On the basis of aligned data, weighted average or Kalman filtering method is used to fuse multiple data sources of the same road segment to obtain more accurate traffic indicators. For example:
[0113]
[0114] In the formula, is the speed of the i-th road segment measured by the k-th data source, and θ kfor corresponding data credibility weight, and Finally, a more comprehensive "road network traffic state table" is obtained in each decision cycle, including the key information of each road segment such as traffic flow, speed, occupancy rate or vehicle density, which lays the data foundation for subsequent steps.
[0115] Step two: constructing a macroscopic basic graph and extracting sub-regional candidate boundaries, specifically including:
[0116] 1) Preliminary division of geographical boundaries of traffic sub-regions. In combination with urban road structure and administrative division information, the road network can be preliminarily divided into several "large regions" in geographical sense, such as according to urban trunk roads, ring roads or natural division lines such as rivers. This can narrow the search range in the subsequent clustering or optimization on a macro level, while avoiding too dispersed sub-regional boundaries. The preliminary division is only auxiliary to the subsequent technical division.
[0117] 2) Collecting and generating historical traffic data. Before implementing the present application, the vehicle distribution and outflow in each "large region" can be counted through historical traffic data or offline traffic simulation for a period of time, forming sampling points under different congestion levels. If the historical data is insufficient, as shown in the simulation platform, the urban road network can be simulated under different demand levels. Figure 2
[0118] 3) Fitting MFD curve. Map the sampled data (N, Q) to the coordinate system, where represents the total number of vehicles or the average density of a region, represents the out-of-region flow of the region. Use a polynomial or Logistic function to fit, to obtain the macroscopic basic graph of the region. If it is found that a "large region" presents high dispersion or multiple mapping on MFD, it means that the interior is not homogeneous and needs to be refined or further split in subsequent steps.
[0119] 4) Extracting candidate boundaries. Identify several congestion hotspots or road segments with significant traffic pattern differences within the region, and consider the surrounding roads as potential sub-regional boundaries. According to traffic state variance, road grade and other indicators, combined with clustering algorithms, the sub-regional candidate boundaries can be preliminarily determined from the geographical and flow perspectives. These candidate boundaries will be used as input or constraint conditions for the formal regional division algorithm in the next step, making the subsequent search more targeted.
[0120] Step three: initial division of sub-regions and multi-objective optimization, specifically including:
[0121] 1) Multi-objective optimization modeling of sub-regional division. Establish a division decision variable y i,r for all road segments (1 when the ith road segment is divided into the rth sub-region, otherwise 0). Define the following optimization objectives:
[0122] 1. Minimize network-wide latency: Let F1 = ∑q i , where q i This represents the number of vehicles queuing in the i-th road segment during the observation period;
[0123] 2. Uniformity within a sub-region: (Note: The original text contains some inconsistencies and unclear grammatical structure. A more accurate translation would require the full context.) Where ρ i Let the density of the i-th segment be... Let be the average density of the subregion r.
[0124] The multiple objectives are combined into a weighted sum αF1 + βF2. At the constraint level, ∑... r y i,r =1 and each sub-region must be geographically connected to avoid fragmentation.
[0125] 2) Heuristic Algorithm Solution. Since this is a discrete combinatorial optimization problem, genetic algorithms, simulated annealing, or other metaheuristic algorithms can be used to solve it. The initial population can randomly generate several feasible solutions based on the candidate boundaries identified in the previous step, and then gradually approach the optimal solution through iterative selection, crossover, and mutation operations. If the obtained partitioning results show good MFD consistency on the experimental data while also reducing latency, then this partitioning scheme is established as the initial sub-region partitioning scheme.
[0126] 3) Partitioning Results and Boundary Verification. Visualize the partitioned results on the road network map and check whether each sub-region meets requirements such as topological connectivity and homogeneity of traffic characteristics. If any sub-region is too large or has excessive internal heterogeneity, it can be further subdivided at this stage. If any sub-region is too small or isolated, it can be considered for merging with adjacent regions. After this correction, an initial partitioning scheme that can be dynamically adjusted in subsequent steps is obtained.
[0127] Step 4: Dynamic adjustment based on multi-agent game theory and reinforcement learning, specifically including:
[0128] 1) Establish a multi-agent (MDP) model. Based on the initial partitioning, configure each sub-region as an agent. For intelligent agents Its state s r This may include: the average density ρ within the sub-region r Number of cars in queue Q r MFD curve parameters, etc.; its action a r To fine-tune the boundaries or internal road segments of sub-regions, such as assigning certain road segments at the boundaries to adjacent sub-regions or merging originally divided sub-regions, the immediate reward r(s,a) is defined as the negative sum of delay and non-uniformity:
[0129] r(s,a)=-(αF1+βF2)
[0130] where s denotes the current global state (combining all regional s r ), a = {a1,..., a R} denotes the joint action of all agents. The Markov transition is driven by traffic simulation or real-time observation, describing the evolution of traffic flow and congestion in the next period after performing the action.
[0131] 2) Mean-field game mechanism is introduced. When multiple agents interact, let each agent consider the joint strategy of other agents as an average behavior which can be defined as:
[0132]
[0133] where is the set of sub-regions adjacent to or interacting with agent r through direct boundary. When updating, agent r assumes unchanged, and the incremental r is calculated for its own strategy parameter r . With multiple rounds of iteration, each r will tend to approach the solution that satisfies the game equilibrium.
[0134] 3) Policy gradient reinforcement learning. Each agent performs reinforcement learning training individually. The actor-critic structure can be used, and the actor network outputs the action distribution The critic network estimates the value function The gradient update formula can be written as:
[0135]
[0136] Here s' r denotes the next state, and γ is the discount factor. By repeatedly sampling trajectories in the simulation or online running environment, the reward is obtained after each action is performed, and the strategy parameters are updated, eventually enabling each sub-region agent to learn the optimal boundary adjustment strategy under various traffic conditions.
[0137] 4) Periodic dynamic adjustment. In real operation, every TTT minutes is a decision period, and the system inputs the multi-source data fusion results into the reinforcement learning module. Each agent determines whether to fine-tune its own boundary according to the learned optimal strategy, such as merging or splitting specific road segments. When the decision result is "adjustment", the boundary ownership change is performed; if it is "unchanged", the current division is retained. This cycle is repeated, and after several periods, the sub-region boundaries gradually adapt and approach stability; if there is a major mutation in traffic patterns (such as sudden accidents, peak tides), a larger range of sub-region redivision is performed according to the latest round of reinforcement learning output, achieving adaptive dynamic adjustment.
[0138] Step five: Implement sub-region flow control and closed-loop feedback, including:
[0139] 1) Inter-region MFD value control. After sub-region division, to prevent excessive saturation in the region, the application can combine the macro basic map MFD to achieve entry flow gating. For each sub-region r, fit the MFD function Q r = f r (N r ), determine its critical vehicle number By limiting the flow entering at the boundary signal light or increasing the flow exiting, maintain When , reduce the green light duration, and moderately gate the external traffic flow; when and the adjacent sub-region is more congested, then appropriately increase the green light duration to guide the traffic shift.
[0140] 2) Boundary signal and flow quota. While performing division adjustment, the signal timing scheme can be configured for the new sub-region boundary. For example, if sub-region A incorporates a high congestion section, the traffic flow entering A region needs to be moderately reduced; if sub-region B splits out part of the section, its traffic demand needs to be recalculated and the corresponding signal cycle needs to be adjusted to achieve coordinated control across regions. This process can be further included in the action space of reinforcement learning or considered as an additional module.
[0141] 3) Closed-loop feedback and long-term optimization. After each decision cycle, the system re-collects traffic state information, compares the current network delay, sub-region congestion degree and the changes from the last cycle, and uses this as a training sample for the reinforcement learning algorithm, so as to continuously update the policy network of each agent. When there is a major change in the actual situation, the multi-objective optimization of step three or the multi-agent training of step four can be triggered to continue to maintain the optimal effect in the new traffic environment.
[0142] The above only describes the preferred embodiments of the application. It should be noted that for those skilled in the art, without departing from the technical principles of the application, several improvements and modifications can be made, and these improvements and modifications should also be considered as the protection scope of the application.
Claims
1. A method for sub-regional division of urban road networks based on game-reinforcement learning collaborative control, characterized in that, The method includes: Step S1: Acquire multi-source traffic data and perform fusion processing to estimate the traffic state parameters of each road segment in the road network in real time; Step S2: Based on the traffic state parameters obtained in Step S1, construct a multi-objective optimization model with the goal of minimizing traffic delay and regional imbalance. Step S3: Further model the multi-objective optimization model constructed in step S2 into a Markov decision process, and define the state space, action space, state transition probability, immediate reward function and discount factor for this process; Step S4: To solve the Markov decision process defined in step S3, a reinforcement learning algorithm is used to calculate the optimal partitioning strategy, and a cooperative mechanism based on mean field game theory is introduced to solve the interaction problem between multi-agent decisions. Step S5: Based on the traffic state parameters obtained in Step S1 and the regional division results obtained by the optimal division strategy in Step S4, a sub-region traffic flow model is established using the macro-traffic flow basic graph theory to describe the macro-traffic flow characteristics of the sub-region. Step S6: Based on the sub-regional traffic flow model established in Step S5, the optimal partitioning strategy obtained in Step S4 is periodically applied for decision-making. When the regional traffic imbalance index is detected to exceed the preset threshold, a re-partitioning calculation is triggered to update the regional boundary and generate new control strategy parameters.
2. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S1, the multi-source traffic data includes fixed detector data, floating car trajectory data, and video detection data, and a weighted fusion method is used to calculate the average speed of the road segment: In the formula, v i Let be the average driving speed after merging for the i-th road segment; The speed of the i-th road segment is measured by the k-th type of data source; θ k Let be the credibility weight corresponding to the k-th type of data source, and satisfy . K represents the total number of categories in the data source.
3. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S2, the objective function of the multi-objective optimization model is defined as: minJ(p)=αF1(p)+βF2(p) in: The constraints are: In the formula, p represents a specific regional division scheme; F1 is the global traffic delay index; q i ρ represents the number of vehicles queuing on the i-th road segment of the road network; F2 is the regional imbalance index; i This represents the vehicle density of the i-th road segment; The value represents the average vehicle density within the r-th sub-region; α and β are weighting coefficients; N is the total number of road segments in the road network; R is the number of sub-regions; y i,r As an assignment variable, when the i-th road segment is assigned to the r-th region, y i,r =1, otherwise 0.
4. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S3, the immediate reward function of the Markov decision process is defined as: Rew(s,a)=-[αF1(s)+βF2(s)] In the formula, Rew(s,a) is the immediate reward obtained by performing action a in state s; s represents the current system state; a represents the action performed; F1(s) and F2(s) are the global traffic delay index and regional imbalance index in the current state s, respectively; α and β are weighting coefficients.
5. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S4, the reinforcement learning algorithm employs the policy gradient method, and its policy gradient update formula is as follows: In the formula, θ represents the parameters of the policy network; π θ (a|s) is the parameterized policy, representing the probability of choosing action a in state s; J(θ) is the expected reward of the policy; Rew(s,a) is the reward value obtained after performing action a in state s; It represents the expectation of the distribution of states and actions.
6. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S4, the equilibrium of the mean-field game theory satisfies the following condition: In the formula, Q(s,a) is the action value function; It takes into account the average behavior of other agents. The reward afterwards; γ represents the average behavior of other agents' policies; γ is the discount factor; V(s') is the value function of the next state s'; The corresponding state transition probability; μ(s) is the current state distribution; μ'(s') is the state distribution at the next moment; π * (a|s) represents the optimal strategy in the equilibrium state; S and A represent the state space and action space, respectively.
7. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S5, the sub-region traffic flow model is fitted using a quadratic polynomial, and its expression is: In the formula, Q r K represents the outbound traffic flow of the r-th sub-region; r c1 and c2 represent the average vehicle density of the r-th sub-region; c1 and c2 are model coefficients obtained by fitting historical data.
8. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S6, the regional traffic imbalance index H used to trigger the rezoning is defined as: In the formula, n r ρ represents the number of road segments within the r-th sub-region; i This represents the vehicle density of the i-th road segment; Let R represent the average vehicle density in the r-th subregion; R is the total number of subregions.
9. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, In step S6, the decision period T is dynamically adjusted according to the rate of change of traffic conditions: when traffic conditions fluctuate frequently, a smaller value is taken to improve the response speed; when traffic conditions are relatively stable, a larger value is taken to reduce the computational load.
10. The urban road network sub-region division method based on game-reinforcement learning collaborative control according to claim 1, characterized in that, The method also includes, after step S6, combining area boundary control with traffic light timing optimization, using proportional control as the control law: In the formula, u r (t) represents the inlet flow rate adjustment for the r-th region; K p The gain is controlled proportionally. N represents the optimal critical number of vehicles in the r-th region determined based on the macroscopic traffic flow basic map. r (t) represents the actual number of vehicles in region r at time t.
Citation Information
Cited By
Urban parking-oriented Internet of Things mobile charging system and dynamic pricing method
CN121366451A