Intelligent management and control method for large-range urban road network based on mobile phone signaling data
By constructing a multi-objective macro-micro collaborative two-layer planning model based on mobile phone signaling data and deep reinforcement learning, the intersection signal timing and variable lane strategy are optimized, solving the efficiency and safety problems of traditional road network management methods in complex traffic scenarios, and realizing dynamic and precise management and efficiency improvement of urban road networks.
Patent Information
- Application Number
- CN202511015135.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional road network management methods are difficult to accurately adapt to complex traffic scenarios, leading to efficiency and safety issues. Existing technologies are also unable to handle the challenges of dynamic and complex traffic control.
A multi-objective macro-micro collaborative two-layer planning model is constructed based on mobile phone signaling data. Combined with deep reinforcement learning, the optimal strategy is learned through the Actor-Critic architecture to optimize mixed decision variables such as intersection signal timing and variable lane strategy. The deep neural network is used for iterative training in a micro traffic simulation environment, and the correlation between control strategy and road network performance is clarified by combining the causal graph framework.
It enables dynamic and precise control of large-scale urban road networks, ensuring operational safety while significantly improving operational efficiency, and provides an innovative solution.
Smart Images

Figure CN120932440A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traffic management and control technology, and in particular to a method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data. Background Technology
[0002] With the dynamic changes in urban traffic demand, traditional road network management methods struggle to accurately adapt to complex traffic scenarios, leading to efficiency and safety issues. Mobile signaling data, with its real-time and wide-coverage characteristics, can accurately invert traffic flow parameters (such as OD matrix and speed field), providing core data support for management optimization. Deep reinforcement learning, as an emerging intelligent technology, possesses the ability to handle complex dynamic decisions. Based on this, this method constructs a multi-objective macro-micro collaborative two-layer programming model using traffic parameters inverted from mobile signaling data, incorporating discrete-continuous hybrid decision variables such as intersection signal timing and variable lane strategies into the deep reinforcement learning framework. By constructing a Markov decision process, the optimal strategy is learned using a deep neural network Actor-Criti architecture, iteratively trained in a constructed micro-traffic simulation environment, and optimized with real-time feedback. This process integrates a causal graph framework, clarifying the relationship between management strategies and road network performance. While ensuring safety indicators such as speed dispersion and conflict index, it improves operational efficiency, providing an innovative solution for urban intelligent traffic management and effectively addressing the dynamic and complex traffic control challenges that traditional methods struggle to handle. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a large-scale intelligent management and control method for urban road networks based on mobile phone signaling data. By combining multi-source data fusion, macro-micro dual-layer planning models and deep reinforcement learning, dynamic and precise management and control of urban road networks can be achieved, significantly improving operational efficiency while ensuring operational safety.
[0004] To achieve the above objectives, in a first aspect, the present invention provides a method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data, comprising the following steps:
[0005] Step 1: Construct a micro-level traffic simulation environment. Import open-source map data into the simulation environment and make local adjustments. After preprocessing the mobile phone signaling data, convert it into OD matrix files and vehicle route files. Dynamically adjust the control parameters of the urban road network through the written traffic control strategy code interface to provide a simulation basis for the macro-micro collaborative two-level planning model in Step 3.
[0006] Step 2: Establish an evaluation index system based on the coupling characteristics of multi-source data and spatiotemporal dynamic response, including operational safety index and operational efficiency index. After standardization, construct a multi-objective comprehensive evaluation function to provide optimization objectives for the upper-level model in Step 3 and the calculation basis for the reward function in Step 4.
[0007] Step 3: Construct a macro-micro collaborative bi-layer planning model, which includes an upper-layer model and a lower-layer model. The lower-layer model generates a balanced traffic flow state based on the micro-traffic simulation environment and the real-time OD matrix. The upper-layer model uses the multi-objective comprehensive evaluation function constructed in Step 2 as the optimization objective and uses continuous-discrete hybrid decision variables such as intersection signal timing parameters, variable lane switching strategies, and bus priority duration as control parameters.
[0008] Step 4: Optimize the mixed decision variables using a deep reinforcement learning algorithm, construct a Markov decision process, map the continuous-discrete mixed decision variables in Step 3 into an action space, use the cumulative reward generated by the weighted evaluation index system in Step 2 as the optimization guide, perform policy learning through an Actor-Critic architecture, and perform iterative optimization by combining an experience replay mechanism and a policy gradient algorithm to output the optimal combination of decision variables that includes the signal timing scheme and lane function control strategy.
[0009] Furthermore, the method for constructing the microscopic traffic simulation environment in step 1 includes: supplementing the intersection signal phase configuration parameters, marking the physical location and conversion rules of variable lanes, adding priority indicators for bus lanes and ordinary lanes, calibrating the speed limit values of road sections to match the actual road design parameters, and forming a digital twin model that is completely consistent with the road network structure of the target city. The digital twin model provides a scenario basis for the dynamic traffic flow simulation of the lower-level model in step 3.
[0010] Furthermore, the mobile signaling data preprocessing method includes: reconstructing user movement trajectories through cellular grid division; inferring continuous driving paths based on signal strength attenuation models and road network constraints; cleaning data with spatiotemporal jumps and abnormal base station correlations; filling missing trajectory segments using spatiotemporal interpolation algorithms; desensitizing user identifiers using k-anonymity technology; and providing OD matrix files and vehicle route files for the microscopic traffic simulation environment in step 1.
[0011] Furthermore, the operational safety indicators include speed dispersion, emergency braking frequency, and intersection conflict index; the operational efficiency indicators include average travel time, intersection traffic efficiency, and road network throughput, which together constitute the evaluation indicator system described in step 2.
[0012] Furthermore, the emergency braking frequency F b The calculation formula is:
[0013]
[0014] Where T is the statistical time, a i (t) represents the acceleration of the i-th vehicle at time t, a th I is the acceleration threshold, and I(·) is the indicator function. This formula quantifies the frequency of emergency braking behavior in the road network by counting the number of times the vehicle acceleration is lower than the safety threshold per unit time, providing quantitative support for the operational safety indicators in step 2. The lower the frequency of emergency braking, the fewer sudden dangerous situations the road network will have, and the better the operational safety.
[0015] Furthermore, the velocity discreteness D v The calculation formula is:
[0016]
[0017] Where N is the total number of vehicles, v i Let be the speed of the i-th vehicle. The average speed is given; the formula for calculating the emergency braking frequency is: Where T is the statistical time, a i (t) represents the acceleration of the i-th vehicle at time t, a th Acceleration threshold (e.g., -3 m / s²) 2 I(·) is an indicator function, which takes the value 1 when the condition is met and 0 otherwise. This formula quantifies the consistency of vehicle speeds within the road network by calculating the degree of deviation between vehicle speed and average speed, and provides a quantitative basis for the operational safety indicators in step 2.
[0018] Furthermore, the intersection conflict index C int The calculation formula is:
[0019]
[0020] Where K is the number of time steps, p i (t k Let ) represent the time of the i-th vehicle at time t. k Position, d safe The safe distance threshold; the average travel time T avg The calculation formula is:
[0021]
[0022] Where M represents the total number of vehicles that completed the trip. and These are the start and end times of the m-th vehicle, respectively. This formula quantifies the overall traffic efficiency of the road network by calculating the average time taken for all vehicles to complete their journeys, providing a quantitative basis for the operational efficiency indicators in step 2. The shorter the average journey time, the higher the traffic efficiency of the road network.
[0023] Furthermore, the formula for calculating the intersection's traffic efficiency is as follows:
[0024]
[0025] Where C represents the total number of vehicles, PCU c Let T be the standard vehicle equivalent of the c-th vehicle. g The green light time; the formula for calculating the road network throughput is:
[0026]
[0027] Where S is the total number of critical sections of the road network, T is the unit time, and N is the number of key sections. s (t) represents the total number of vehicles passing through the key section s of the road network within a unit time t.
[0028] Furthermore, the multi-objective comprehensive evaluation function is the core of the evaluation index system constructed in step 2, and is used for the optimization objective calculation of the upper-level model in step 3. The calculation formula of the multi-objective comprehensive evaluation function Q is as follows:
[0029] Q=α·(1-D v ′)+β·Q int ′+γ·Q net ′
[0030] Among them, D v ′ represents the standardized velocity dispersion, Q int Q' represents the standardized intersection throughput efficiency. net ′ represents the standardized network throughput. Each indicator is standardized using the range method; for positive indicators, the formula is: For negative indicators, the formula is: The weighting coefficients α, β, and γ are determined using the entropy weighting method, and satisfy α + β + γ = 1. First, the entropy value of the j-th index is calculated. Then calculate the difference coefficient g. j =1-e j The weights of each indicator are obtained by normalization. The weight combination can be determined by solving the following multi-objective optimization problem: max α,β,γ {Q|D v ≤D th Q int ≥Q min Q net ≥Q base}, where D th Q min Q baseThese are the preset thresholds for each indicator. This formula integrates multi-dimensional evaluation indicators into a single comprehensive evaluation value by weighting and fusing standardized safety and efficiency indicators. This provides a quantitative standard for the optimization direction of the upper-level model in step 3, and realizes the coordinated consideration of road network safety and efficiency.
[0031] Furthermore, the continuous-discrete hybrid decision variables are the control parameters of the upper-level model proposed in step 3, used for the optimization of the deep reinforcement learning algorithm in step 4. Specifically, the continuous-discrete hybrid decision variables include:
[0032] Continuous variable: Intersection signal cycle duration T cycle ∈[T min ,T max Green credit ratio allocation coefficient λ k ∈[0,1], Bus priority phase duration T bus ∈[0,T cycle ];
[0033] Discrete variables: signal phase pattern φ∈{four-phase, eight-phase}, variable lane switching time interval Δt lane The variables ∈{15min,30min,60min} and the area restriction combination scheme Ω∈{ABC,ABD,BCD} cover various control measures such as intersection signal control, dynamic adjustment of lane function and traffic control, providing specific optimization objects for the deep reinforcement learning algorithm in step 4. Through collaborative optimization, dynamic adaptation of road network control strategies is achieved.
[0034] Furthermore, the lower-level model is a dynamic traffic flow balance simulation, based on the real-time OD matrix derived from mobile signaling data, and performs dynamic user equilibrium (DUE) allocation in a micro-traffic simulation environment. The model is as follows:
[0035]
[0036] Where, x a Let f be the traffic flow of road segment a, where A is the set of road segments. p δ is the flow of path p, where P is the set of paths. If segment a is on path p, then δ ap The value is 1 if the value is not specified, and 0 otherwise. The travel time for a road segment is calculated using the BPR impedance function.
[0037]
[0038] Among them, t a (x a () represents the travel time for road segment a. For free-flow time, c aLet be the capacity of road segment a, and α and β be calibration parameters. The equilibrium condition requires that all used paths have the same minimum generalized travel cost:
[0039]
[0040] Among them, C p C is the cost of path p. min It has the lowest travel cost.
[0041] Furthermore, the macro-micro collaborative two-layer model is coupled through a feedback mechanism. The traffic flow state generated by the lower-layer model is fed back to the upper layer in real time to calculate the comprehensive evaluation value Q. The gradient of the traffic flow allocation result of the lower-layer model to the upper-layer decision variable u is passed to the Actor-Critic network for backpropagation through the chain rule.
[0042] Furthermore, the specific implementation of the deep reinforcement learning algorithm includes:
[0043] A state space is constructed, which is based on the road network performance indicators fed back in real time from the microscopic traffic simulation environment in step 1 as the state input. The multidimensional road network performance indicator vector x (such as speed dispersion, intersection conflict index, and average travel time) fed back in real time from the microscopic traffic simulation environment is mapped to a low-dimensional state vector s through a fully connected layer. The calculation formula is as follows:
[0044] s = ReLU(W·x + b)
[0045] Where W is the weight matrix, b is the bias vector, and ReLU is the activation function;
[0046] Construct an action space, which maps the continuous-discrete mixed decision variables from step 3. The continuous variables are parameterized and encoded using a Gaussian distribution.
[0047]
[0048] Wherein, mean μ cycle and variance Output from the Actor network, Represents a normal distribution;
[0049] Discrete variables are encoded using the Softmax probability distribution:
[0050]
[0051] Among them, z i It is the unnormalized log probability output by the Actor network; design a reward function to construct a reward signal based on the real-time changes in the comprehensive evaluation value to avoid training instability: r t=clip(Q t -Q t-1 ,r min ,r max ), where r t It is the reward at the current moment, Q. t and Q t-1 These are the current and previous time values, r. min and r max To determine the upper and lower bounds of reward scaling and cropping, the goal is to maximize the cumulative reward Q.
[0052] Furthermore, the parameter update process of the deep reinforcement learning algorithm includes: randomly sampling batch data from the experience replay buffer and calculating the target Q value y = r. t +γQ w′ (s t+1 ,π θ′ (s t+1 )),
[0053] Where γ is the discount factor, Q w′ and π θ′ These are the target Critic network and the target Actor network, respectively. The loss function of the Critic network is... The Critic network parameters w are updated by minimizing this loss. The Actor network updates the policy parameters θ using gradient ascent, with the objective function being... The training termination condition is when the average reward value of the test set fluctuates less than a set threshold for N consecutive generations (i.e., ...). (or reach the maximum number of training rounds T) max When the optimal strategy is reached, training is terminated and the combination of decision variables corresponding to the optimal strategy is output.
[0054] Secondly, the present invention also provides a large-scale urban road network intelligent management and control system based on mobile phone signaling data, comprising:
[0055] Simulation environment building module: used to build an urban traffic flow simulation environment based on microscopic traffic simulation;
[0056] Evaluation index establishment module: used to establish an evaluation index system for urban traffic network performance based on multi-source data coupling characteristics and spatiotemporal dynamic response;
[0057] Macro-micro collaborative bi-level programming model construction module: used to construct the framework of a macro-micro bi-level programming model;
[0058] Optimal Decision Variable Selection Module: Used to optimize the decision variables of a two-level model using deep reinforcement learning methods.
[0059] Compared with the prior art, the present invention has at least the following beneficial effects:
[0060] (1) By constructing a micro-traffic simulation environment based on mobile phone signaling data, establishing a macro-micro collaborative dual-layer planning model, and using deep reinforcement learning algorithms for strategy optimization, this invention can accurately respond to complex traffic conditions in a large-scale urban road network, overcome the drawbacks of static and lagging traditional control methods, and significantly improve traffic efficiency while ensuring operational safety.
[0061] (2) By constructing a high-fidelity digital twin model of the road network, combining refined preprocessing of mobile phone signaling data, and performing precise mathematical quantification of safety and efficiency indicators, this invention ensures the authenticity of the simulation environment and the objectivity of the evaluation system, providing a reliable data foundation and clear quantitative targets for subsequent strategy optimization.
[0062] (3) By establishing a comprehensive evaluation system that takes into account both safety and efficiency and unifying it into a single multi-objective comprehensive evaluation function, and by incorporating continuous variables such as signal timing and discrete strategies such as lane function adjustment into a unified hybrid decision space, this invention scientifically constructs an optimization problem that can collaboratively handle diverse control measures in the real world and achieve balanced optimization of multiple objectives.
[0063] (4) By designing an action space encoding scheme and reward function that adapts to the mixed decision variables for deep reinforcement learning algorithms and encapsulating them in a modular system architecture, this invention not only solves the optimization problem of complex decision spaces in terms of technology, but also provides a systematic solution that is easy to deploy and maintain, ensuring that the algorithm can converge to the optimal control strategy stably and efficiently, and has strong engineering application value. Attached Figure Description
[0064] The accompanying drawings, which form part of this specification, illustrate embodiments of the invention and, together with the specification, serve to explain the principles of the invention.
[0065] The invention will be more clearly understood with reference to the accompanying drawings and the following detailed description, wherein:
[0066] Figure 1 This is a flowchart of the intelligent management and control method for large-scale urban road networks based on mobile phone signaling data according to the present invention;
[0067] Figure 2 This is the iterative optimization graph based on deep reinforcement learning strategy in this invention;
[0068] Figure 3 This invention constructs a macro-micro collaborative two-level planning model. Detailed Implementation
[0069] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments and specific features in the embodiments are detailed descriptions of the technical solution of the present application, and not limitations thereof. Where there is no conflict, the embodiments and technical features in the embodiments can be combined with each other. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0070] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0071] This invention proposes a method and system for large-scale intelligent management and control of urban road networks based on mobile phone signaling data. The core idea of this invention is to leverage the real-time and wide-coverage advantages of mobile phone signaling data to accurately invert traffic flow parameters (such as OD matrix and velocity field), addressing the dynamic changes in urban traffic demand and the limitations of traditional management methods, thus providing core data support for management and control optimization. Based on this, a macro-micro collaborative two-layer planning model is constructed, and deep reinforcement learning, an emerging intelligent technology, is employed to handle complex dynamic decision-making problems.
[0072] Specifically, this invention uses a mixture of discrete and continuous decision variables, such as intersection signal timing, variable lane strategies, and bus priority duration, as optimization objects for deep reinforcement learning. It constructs a Markov Decision Process (MDP) and utilizes an Actor-Critic architecture of a deep neural network to learn the optimal strategy. The entire learning and optimization process is iteratively trained in a high-fidelity microscopic traffic simulation environment, which provides real-time feedback on the effectiveness of the control strategy. This process also incorporates a causal graph framework to clarify the intrinsic relationship between the control strategy and road network performance, thereby significantly improving operational efficiency while ensuring road network safety (e.g., reducing speed dispersion and conflict index). This provides an innovative intelligent solution for addressing the complex and ever-changing traffic control challenges of modern cities.
[0073] The specific implementation steps and modules of this invention will be described in detail below.
[0074] Example 1
[0075] like Figures 1-3 As shown, this invention provides a method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data, comprising the following steps:
[0076] Step 1: Construct a high-fidelity microscopic traffic simulation environment
[0077] To accurately simulate real-world traffic conditions and test the effectiveness of traffic control strategies, it is first necessary to build a microscopic traffic simulation environment capable of dynamically simulating real-world traffic flow in urban road networks in real time. This step is fundamental for subsequent model training and strategy optimization, and mainly includes the following:
[0078] 1.1 Road Network File Construction and Localized Deployment: First, import basic road network data for the target city from an open-source map platform (such as OpenStreetMap), including road geometric parameters, intersection attributes, and traffic sign information, to form a basic road network file. Second, perform refined localized adjustments and calibrations on this basic road network file to construct a "digital twin model" that is completely consistent with the actual road network structure and operating rules. Specific adjustments include: supplementing signal phase configuration parameters for intersections (such as phase timing and green light duration); accurately marking the physical location of reversible lanes and their transition rules; adding priority indicators for bus lanes and regular lanes; and calibrating speed limits for each road segment based on actual road design parameters.
[0079] 1.2 Mobile Signaling Data Preprocessing and Fusion: Mobile signaling data is the core of the simulation. Preprocessing begins by reconstructing the user's movement trajectory through cellular grid partitioning. Next, combining the signal strength attenuation model and the constructed road network topology constraints, the continuous driving path of vehicles is inferred. The raw data is cleaned to remove noisy data with spatiotemporal jumps and abnormal base station correlations. Spatiotemporal interpolation algorithms (e.g., cubic spline interpolation) are used to fill in missing segments in the cleaned trajectory. To protect user privacy, k-anonymization technology is used to desensitize user identifiers. Finally, the massive, high-precision traffic flow information obtained after the above preprocessing is converted into OD (origin-endpoint) matrix files and vehicle route files that are recognizable by the simulation environment, thereby driving the simulation environment to generate dynamically changing traffic flow.
[0080] 1.3 Dynamic Simulation of Proactive Traffic Control Strategies: In a simulation environment, the traffic control strategy code interface (e.g., using Python) is used to dynamically adjust the control parameters of the road network to simulate the implementation effects of various proactive traffic control strategies. These strategies include:
[0081] Intersection signal control: Based on real-time traffic flow, direction, and congestion conditions, dynamically adjust the signal cycle duration, green light ratio, and phase difference at the intersection.
[0082] Variable lane control: By combining historical signaling data to analyze traffic flow characteristics at different times, the driving direction or function of lanes can be flexibly changed, such as activating tidal lanes during peak hours.
[0083] Bus priority strategy: By combining GPS data of buses, the signal priority mechanism of bus lanes is triggered in real time (such as extending the green light) to ensure the right-of-way of buses.
[0084] Regional traffic restriction strategy: By configuring road network weight parameters, the effect of restricting specific vehicle types (such as trucks) in specific areas and at specific times is simulated.
[0085] Speed limit control: Dynamically set the speed limit value of road sections based on factors such as road type and time of day.
[0086] Step 2: Establish a multi-dimensional, spatiotemporally dynamic evaluation index system.
[0087] To scientifically evaluate the effectiveness of control strategies, it is necessary to establish an evaluation index system that can comprehensively reflect the performance of the transportation network. This system is based on the coupling characteristics and spatiotemporal dynamic response of multi-source data (especially parameters derived from mobile signaling data), and mainly includes two dimensions: operational safety and operational efficiency.
[0088] 2.1 Quantification of Operational Safety Indicators:
[0089] Velocity Dispersion (D) v The standard deviation (SD) reflects the stability and safety of the road network and is defined as the standard deviation of vehicle speeds on a road segment. The smaller the dispersion, the less acceleration and deceleration behavior of vehicles, resulting in a smoother ride. The calculation formula is as follows:
[0090]
[0091] Where N is the total number of vehicles, v i Let be the speed of the i-th vehicle. The average speed of all vehicles.
[0092] Emergency braking frequency (F) b This reflects the risk of sudden braking, specifically when the vehicle's acceleration falls below a certain safety threshold (e.g., -3 m / s²) per unit time. 2 The number of times the signal is triggered. The smaller the value, the lower the risk of a rear-end collision. The calculation formula is:
[0093]
[0094] Where T is the statistical time, a i (t) represents the acceleration of vehicle i at time t, a th I is the acceleration threshold, and I(·) is the indicator function.
[0095] Intersection Conflict Index (C) int This reflects the rationality of the intersection design. Based on a spatiotemporal trajectory overlap detection algorithm, it calculates the number of potential spatiotemporal conflicts between vehicles from different directions in key areas of the intersection per unit time. The calculation formula is as follows:
[0096]
[0097] Where K is the total number of time steps, p i (t k Let i be vehicle i at time t. k Position, d safe This is the safe distance threshold.
[0098] 2.2 Quantification of Operational Efficiency Indicators:
[0099] Average travel time (T) avg This refers to the average travel time of all vehicles within a road network from the origin to the destination. A smaller value indicates higher traffic efficiency. The calculation formula is as follows:
[0100]
[0101] Where M represents the total number of vehicles that completed the trip. and These are the start and end times for the m-th vehicle, respectively.
[0102] Intersection traffic efficiency (Q) int ( ): This refers to the standard number of vehicles passing through the intersection within a unit of time period. The calculation formula is:
[0103]
[0104] Among them, PCU c T is the standard vehicle equivalent of vehicle c. g The duration of the green light.
[0105] Network throughput (Q) net This refers to the total number of vehicles passing through key sections of the road network per unit time. The calculation formula is:
[0106]
[0107] Where, N s (t) represents the number of vehicles passing through section s within a unit time t.
[0108] 2.3 Design of Multi-Objective Comprehensive Evaluation Function:
[0109] Indicator standardization: The range method is used to map heterogeneous indicators to intervals. For positive indicators (the larger the value, the better), the formula is as follows: For negative indicators (the smaller the value, the better), the formula is:
[0110] Dynamic weighting: The entropy weighting method is used to dynamically determine the weights of each indicator. First, the entropy value of the j-th indicator is calculated. Then calculate the difference coefficient g. j =1-e j The final weights are obtained by normalization.
[0111] Calculation of the comprehensive evaluation value: A linear weighted model is constructed to obtain the comprehensive evaluation value Q. The weighting coefficients α, β, and γ can be determined by solving a multi-objective optimization problem: max α,β,γ {Q|[cite s tart]D v ≤D th Q int ≥Q min Q net ≥Q base The comprehensive evaluation function formula is as follows:
[0112] Q=α·(1-D′ v )+β·Q′ int +γ·Q′ net
[0113] Among them, D′ v ,Q′ int ,Q′ net These represent the standardized speed dispersion, intersection throughput, and network throughput, respectively, with α+β+γ=1.
[0114] Step 3: Construct a macro-micro collaborative two-level programming model
[0115] This invention constructs a two-layer planning model, including an upper-layer model and a lower-layer model. By combining macro-level strategy optimization with micro-level traffic flow simulation, a closed-loop iterative mechanism of "strategy optimization - traffic flow simulation - performance evaluation" is formed.
[0116] 3.1 Lower-level model, used for dynamic traffic flow balance simulation. The lower-level model, based on the real-time OD matrix, performs dynamic user equilibrium (DUE) allocation in a microscopic traffic simulation environment, simulating vehicle route selection behavior. The dynamic user equilibrium model can be represented as... Its equilibrium condition requires that all used paths have the same minimum generalized travel cost:
[0117] This process approximates the actual traffic flow distribution by iteratively adjusting the impedance function of the road segment. This embodiment uses the BPR function to calculate the travel time t of the road segment. a :
[0118]
[0119] in For free-flow time, x a For the traffic flow of the road segment, ca For traffic capacity.
[0120] The balanced traffic flow state generated by the lower-level model will be fed back to the upper-level model to calculate the comprehensive evaluation value Q.
[0121] 3.2 Upper-level model, used for collaborative optimization of mixed decision variables. The upper-level model aims to maximize the comprehensive evaluation value Q: max u Q=α·(1-D v )+β·Q int +γ·Q net The object of optimization, namely the decision variable, is a complex combination of both continuous and discrete variables.
[0122] Continuous variables: including the intersection signal cycle duration T cycle ∈[T min ,T max Green credit ratio allocation coefficient λ k ∈[0,1], Bus priority phase duration T bus ∈[0,T cycle ].
[0123] Discrete variables: including signal phase pattern φ∈{four-phase, eight-phase}, and variable lane switching time interval Δt lane ∈{15min,30min,60min}, and the area restriction combination scheme Ω∈{ABC,ABD,BCD}.
[0124] 3.3 Coupling and Iteration of Macro-Micro Models
[0125] Traffic flow status generated in the lower layer is fed back to the upper layer in real time, driving policy optimization. The gradient of the traffic flow allocation results from the lower-layer model with respect to the upper-layer decision variables is passed to the Actor-Critic network via a chain rule, achieving end-to-end optimization.
[0126]
[0127] Step 4: Use deep reinforcement learning for policy iteration and optimization.
[0128] To solve the aforementioned complex optimization problem, this invention employs a deep reinforcement learning algorithm based on the Actor-Critic architecture. By constructing the problem as a Markov Decision Process (MDP), closed-loop learning is performed in a simulation environment, gradually converging to the optimal policy combination.
[0129] 4.1 Markov Decision Process (MDP) Construction:
[0130] State space: The multi-dimensional road network performance indicators (such as speed dispersion, average travel delay, etc.) output from the simulation are mapped to a low-dimensional state vector s through a fully connected layer. Let the input indicator vector be x, then the state vector s = f(W·x+b), where f is the activation function (such as ReLU).
[0131] Action Space: A composite action space design is employed. For continuous variables, a Gaussian policy is used for output, for example... The mean and variance are output by the Actor network. For discrete variables, probabilities are output through a Softmax strategy, for example... Where z i It is the logarithmic probability output by the Actor network.
[0132] The reward function (Reward) constructs a reward signal based on the real-time changes in the comprehensive evaluation value Q, and avoids training instability through scaling and cropping. The formula is r. t =clip(Q t -Q t-1 ,r min ,r max ).
[0133] 4.2 Strategy Learning and Network Parameter Update:
[0134] Data collection and storage: The Actor network generates actions based on the current state, inputs them into the simulation environment for execution, obtains the next state and reward, and forms a quadruple (s,a,r,s′) which is stored in the experience replay buffer.
[0135] Network update: Randomly sample batches of data from the buffer to update network parameters, and calculate the target Q value: y i =r i +γQ′(s′ i ,a′ i ), where Q′ is the target Critic network's estimate of the next state-action pair.
[0136] The loss function of the Critic network is the mean squared error: The Critic network parameters are updated by minimizing this loss.
[0137] The Actor network updates policy parameters using gradient ascent, and its objective function is to maximize the expected Q value. Where action a is generated by the Actor network a = π(s|θ) actor )generate.
[0138] 4.3 Iteration Termination Judgment:
[0139] When the average reward value on the test set fluctuates less than a preset threshold for N consecutive generations, or when the maximum number of training rounds is reached, training is terminated, and the combination of decision variables corresponding to the current optimal strategy is output.
[0140] Example 2
[0141] The present invention also provides a system for implementing the above method, which can be integrated into an urban traffic management center and includes the following functional modules:
[0142] Simulation Environment Setup Module: Responsible for executing the operation in step 1 of Example 1. It is used to build and maintain a high-fidelity urban traffic flow simulation environment, dynamically simulating the traffic operation status of the urban road network under active control measures such as intersection signal control, speed limit control, variable lane control, bus priority strategy, and area traffic restriction strategy. It achieves high-fidelity reproduction of real traffic flow scenarios through mobile signaling data.
[0143] Evaluation index establishment module: Responsible for executing the operation in step 2 of Example 1. It is responsible for establishing an urban traffic network performance evaluation index system based on multi-source data coupling characteristics and spatiotemporal dynamic response, including operational safety dimensions (speed dispersion, emergency braking frequency, intersection conflict index) and operational efficiency dimensions (average travel time, intersection throughput, road network throughput). A multi-objective comprehensive evaluation function is constructed after standardizing the units of measurement using standardized methods.
[0144] The macro-micro collaborative two-layer planning model construction module is responsible for executing step 3 in Example 1. It is used to construct the macro-micro two-layer planning model framework. The upper-layer model uses a comprehensive evaluation function as the optimization objective, while the lower-layer model, based on balanced traffic flow data generated from micro-traffic simulation, uses discrete-continuous mixed variables such as intersection signal control parameters (cycle duration, green light ratio), variable lane switching rules, and bus priority phase duration as decision variables. Simulation feedback enables bidirectional iteration between traffic flow and control strategies.
[0145] The optimal decision variable selection module is responsible for executing step 4 in Example 1. This module is the core decision-making unit of the system. It uses deep reinforcement learning to optimize the decision variables in the two-layer model. By constructing a deep reinforcement learning framework based on Markov decision processes, it uses a deep neural network Actor-Critic architecture for policy learning. Combined with mechanisms such as experience replay, policy gradient optimization, and exploration-exploitation balance, it performs multiple rounds of iterative training in the constructed microscopic traffic simulation environment. Finally, it outputs the optimal combination of decision variables, including signal timing schemes, variable lane control strategies, and bus lane management rules.
[0146] In summary, this invention provides a complete, advanced, and intelligent urban road network management method and system by organically combining mobile signaling big data, micro-traffic simulation, two-layer planning model, and deep reinforcement learning. It can effectively cope with the dynamic control needs under complex traffic scenarios and improve efficiency while ensuring safety.
[0147] Example 3
[0148] This embodiment, as a preferred embodiment of the present invention, takes the road network of the core urban area of a city in a province as the application object, and the specific steps are as follows:
[0149] Step 1: Building a Microscopic Traffic Simulation Environment
[0150] Step one, the digital twin modeling of the road network includes:
[0151] 1) Obtain road network data for the target area from OpenStreetMap, including road geometry parameters (number of lanes, length), intersection attributes (phase configuration), and traffic signs (speed limit, lane function).
[0152] 2) Construct a road network file. Based on the actual road network situation, modify the topology of the road network file to make it exactly the same as the real road network:
[0153] ① Configure four-phase timing for signalized intersections in the road network (green light duration 20-40s, cycle 60-120s);
[0154] ② Mark reversible lanes in the urban road network (for two-way six-lane main roads, several lanes will be set up for entering the city during the morning peak from 7:00 to 9:00);
[0155] ③ The speed limits of the calibrated road sections (60km / h for main roads and 40km / h for secondary roads) are set to form a simulation model that maps to the real road network in a 1:1 ratio.
[0156] In step one, the preprocessing of mobile signaling data includes:
[0157] 1) Clean the original signaling data: reconstruct the trajectory using a cellular grid (500m×500m), remove spatiotemporal jump data, and fill the missing trajectory using cubic spline interpolation - remove segments with a missing rate >10%;
[0158] 2) The user ID is anonymized using k-anonymization technology, and the processed data is converted into OD matrix and vehicle route files compatible with the construction of a microscopic traffic simulation platform.
[0159] In step one, the proactive control strategy is simulated by dynamically adjusting control parameters based on the corresponding interface using written Python code, including:
[0160] 1) Signal control: Receive real-time traffic flow data from each approach lane from the microscopic traffic simulation environment and adjust the green light ratio accordingly;
[0161] 2) Reversible lanes: Based on historical signaling data, the lane switching interval is set to 30 minutes.
[0162] 3) Based on bus GPS data, the green light is extended when a bus approaches an intersection.
[0163] Step 2: Construction of an evaluation index system based on multi-source data coupling characteristics and spatiotemporal dynamic response (corresponding to claims 4 and 5)
[0164] In step two, safety indicators are calculated:
[0165] Speed dispersion: Vehicle speeds are collected in real time every 5 seconds on the main road, and the standard deviation is calculated.
[0166] Threshold setting D v ≤8km / h;
[0167] Emergency braking frequency: Statistical acceleration below -3 m / s² per unit time 2 The number of times is given by the formula:
[0168] Target control at F b ≤5 times / min;
[0169] Intersection Conflict Index: Calculated by detecting spatiotemporal trajectory overlap, the number of vehicle conflicts within a critical area (radius 5m) is determined using the following formula:
[0170]
[0171] In step two, the efficiency index is calculated:
[0172] Average travel time: This is the total travel time for all vehicles from the origin to the destination, calculated using the following formula:
[0173] Threshold setting D v ≤8km / h;
[0174] Intersection throughput efficiency: the number of standard vehicles (PCUs) passing through per unit green light time, expressed by the formula:
[0175]
[0176] Road network throughput: the number of vehicles passing through a key section per unit time, calculated using the formula:
[0177] The target is 2,000 vehicles per hour during peak hours.
[0178] In step two, the comprehensive evaluation value is calculated:
[0179] Standardize the indicators using the range method; positive indicators include Q. int Q net The formula is
[0180]
[0181] Negative indicators such as D v F b C int T avg The formula is:
[0182]
[0183] The weights were determined using the entropy weight method (safety weight α = 0.4, efficiency weight β = 0.3, throughput weight γ = 0.3), and the comprehensive evaluation value is:
[0184] Q = 0.4·(1-D) ′ v )+0.3·Q ′ int +0.3·Q ′ net .
[0185] Step 3: Construction of the macro-micro collaborative two-level programming model includes:
[0186] Lower-level model: Dynamic traffic flow assignment:
[0187] Based on the preprocessed OD matrix, dynamic user equalization assignment is performed in a microscopic traffic simulation environment, and the segment impedance is calculated using the BPR function.
[0188]
[0189] Upper-level model: Mixed variable optimization:
[0190] 1) Continuous variable: signal period duration T cycle ∈[60,120]s, Green ratio allocation coefficient λ k ∈[0,1] (k is the phase number), bus priority duration bus∈[0,T] cycle ];
[0191] 2) Discrete variables: signal phase pattern (four-phase / eight-phase), variable lane switching interval (15 / 30 / 60 minutes), area restriction scheme (combination of A, B, and C zones, with restrictions on trucks from 7:00 to 20:00);
[0192] 3) The objective function is to maximize the comprehensive evaluation value Q. The traffic flow parameters output by the lower-level model are obtained in real time through the written Python code based on the corresponding interface.
[0193] Step 4: The deep reinforcement learning strategy is optimized as follows:
[0194] Markov Decision Process (MDP) Modeling:
[0195] 1) State space: a four-dimensional vector a t =[D v (t),C int (t),T avg (t),Q net [t] is mapped to a low-dimensional state vector through a fully connected layer (ReLU activation);
[0196] 2) Action space: Continuous variables are generated using a Gaussian strategy: the mean μ and variance σ are output by the Actor network; discrete variables are selected using a Softmax strategy: the probability distribution is output by the network.
[0197] 3) Reward function: calculated based on incremental evaluation values to avoid training fluctuations.
[0198] Actor-Critic network training:
[0199] 1) Initialization: The Actor network (2 fully connected layers, 256 neurons per layer) outputs continuous variable parameters and discrete variable log probabilities; the Critic network (2 fully connected layers) estimates the state value;
[0200] 2) Data collection: Execute the current strategy in the microscopic traffic simulation environment and generate interactive data (s,a,r,s′) which is stored in the experience replay buffer;
[0201] 3) Parameter Update: Critic Network: Minimize TD Error:
[0202] 4) Exploration strategy: An ε-greedy strategy is adopted, with an initial exploration rate ε = 0.9, which decays exponentially to ε = 0.1.
[0203] Training termination conditions:
[0204] The optimal strategy is output when the average reward of the test set fluctuates by less than 0.1 for 10 consecutive rounds, or when the maximum number of training rounds of 2000 is reached.
[0205] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data, characterized in that, Includes the following steps: Step 1: Construct a microscopic traffic simulation environment. Import open-source map data into the simulation environment and make localized adjustments. Preprocess the mobile signaling data and convert it into OD matrix files and vehicle route files. Step 2: Establish operational safety indicators and operational efficiency indicators, and adopt standardized processing to construct a multi-objective comprehensive evaluation function; Step 3: Based on the microscopic traffic simulation environment and the real-time OD matrix, generate a balanced traffic flow state to construct a lower-level model. Based on the multi-objective comprehensive evaluation function constructed in Step 2, use the continuous-discrete mixed decision variables as control parameters to construct an upper-level model. Construct a macro-micro collaborative two-level planning model through the upper-level model and the lower-level model. The control parameters include intersection signal timing parameters, variable lane conversion strategies, and bus priority duration. Step 4: Optimize the mixed decision variables using a deep reinforcement learning algorithm, construct a Markov decision process, map the continuous-discrete mixed decision variables in Step 3 into an action space, use the cumulative reward generated by the weighted evaluation index system in Step 2 as the optimization guide, perform policy learning through an Actor-Critic architecture, and perform iterative optimization by combining an experience replay mechanism and a policy gradient algorithm to output the optimal combination of decision variables that includes the signal timing scheme and lane function control strategy.
2. The method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data according to claim 1, characterized in that, The construction of the micro-traffic simulation environment in step 1 includes: supplementing the signal phase configuration parameters of intersections, marking the physical location and conversion rules of variable lanes, adding priority labels for bus lanes and ordinary lanes, calibrating the speed limit values of road sections to match the actual road design parameters, and forming a digital twin model that is completely consistent with the road network structure of the target city. The digital twin model provides a scenario basis for the dynamic traffic flow simulation of the lower-level model in step 3.
3. The method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data according to claim 1, characterized in that, The mobile phone signaling data preprocessing includes: User movement trajectories are reconstructed by dividing the data into cellular grids. Continuous driving paths are inferred based on signal strength attenuation models and road network constraints. Data with spatiotemporal jumps and abnormal base station correlations are cleaned. Missing trajectory segments are filled using spatiotemporal interpolation algorithms. User identifiers are desensitized using k-anonymization technology. The processed data provides OD matrix files and vehicle route files for the microscopic traffic simulation environment in step 1.
4. The method for intelligent management and control of large-scale urban road networks based on mobile phone signaling data according to claim 1, characterized in that, The operational safety indicators mentioned in step 2 include speed dispersion, emergency braking frequency, and intersection conflict index; the operational efficiency indicators include average travel time, intersection throughput, and road network throughput, which together constitute the evaluation indicator system mentioned in step 2. The velocity dispersion D v The calculation formula is: Where N is the total number of vehicles, v i Let be the speed of the i-th vehicle. The formula calculates the degree of deviation between the vehicle speed and the average speed, quantifies the consistency of vehicle speeds within the road network, and provides a quantitative basis for the operational safety indicators in step 2. The emergency braking frequency F b The calculation formula is: Where T is the statistical time, a i (t) represents the acceleration of the i-th vehicle at time t, a th I is the acceleration threshold, and I(·) is the indicator function; The intersection conflict index C int The calculation formula is: Where K is the number of time steps, p i (t k Let ) represent the time of the i-th vehicle at time t. k Position, d safe This is the safe distance threshold; The average travel time T avg The calculation formula is: Where M represents the total number of vehicles that completed the trip. and These are the start and end times of the m-th vehicle, respectively.
5. The method for large-scale intelligent management and control of urban road networks based on mobile phone signaling data according to claim 1, characterized in that, The multi-objective comprehensive evaluation function is used to calculate the optimization objective of the upper-level model in step 3. The calculation formula for the multi-objective comprehensive evaluation function Q is as follows: Q=α·(1-D v ′)+β·Q int ′+γ·Q net ′ Where α, β, and γ are weighting coefficients and α + β + γ = 1, D v ′ represents the standardized velocity dispersion, Q int Q' represents the standardized intersection throughput efficiency. net ′ represents the standardized network throughput.
6. The method for large-scale intelligent management and control of urban road networks based on mobile phone signaling data according to claim 1, characterized in that, The specific implementation of the deep reinforcement learning algorithm includes: Construct a state space, which is based on the road network performance indicators fed back in real time from the microscopic traffic simulation environment in step 1 as state input; Construct an action space, which maps the continuous-discrete mixed decision variables in step 3. The continuous variables are parameterized and encoded using a Gaussian distribution, and the discrete variables are encoded using a Softmax probability distribution. Design a reward function to construct a reward signal based on the real-time changes in the comprehensive evaluation value, thereby maximizing the cumulative reward Q.
Citation Information
Cited By
Automatic parallel traffic simulation analysis method and device
CN121211978A
Automated parallel traffic simulation analysis method and device
CN121211978B
Multi-mode traffic simulation method based on multi-agent sequential decision
CN121600715A
A multi-modal traffic simulation method based on multi-agent sequential decision
CN121600715B