Traffic signal dynamic optimization method and system for user space-time behavior tracking
Through user temporal and spatial behavior tracking technology and reinforcement learning model, traffic signals are dynamically adjusted, which solves the problem that traditional traffic signal control methods are difficult to adapt to complex traffic environments, and achieves more efficient traffic flow management.
Patent Information
- Application Number
- CN202510272448.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The traditional fixed-duration traffic signal control method is difficult to adapt to the complex and changeable traffic environment, resulting in congestion easily occur during peak periods. Although the induction control method can dynamically adjust the signal duration, it is difficult to optimize globally.
Through user temporal and spatial behavior tracking technology, the spatial and temporal characteristics of users' group driving behavior in the lane are captured, combined with the complexity of the dynamic road network status of the lane, and the reinforcement learning method is used to build a dynamic optimization model of traffic signals, dynamically adjust traffic signals, and realize smarter signal timing.
It effectively reduces the complexity of the road network status, reduces the waiting time of pedestrians and the probability of slowing down and stopping of vehicles, and improves road traffic efficiency.
Smart Images

Figure CN120220392A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic signal optimization, and particularly to a traffic signal dynamic optimization method and system for tracking user spatio-temporal behavior. Background Art
[0002] With the development of the Intelligent Transportation System (ITS), the traditional fixed-duration traffic signal control method has been difficult to adapt to the complex and changeable traffic environment. The combination of user spatio-temporal behavior tracking and data-driven methods can accurately depict the changes in traffic flow, thereby optimizing traffic signal control strategies, improving road traffic efficiency, and reducing traffic congestion. The traditional traffic signal control method has limitations. Under fixed-time control, the red and green light durations are fixed and do not consider the real-time traffic conditions. It is suitable for sections with relatively stable traffic flow, but it is prone to congestion during peak periods. The inductive control method uses geomagnetic sensors or cameras to detect vehicles and dynamically adjusts the signal duration, but it is still relatively localized and difficult to optimize globally. In order to optimize traffic signals more accurately, user spatio-temporal behavior tracking technology can be used, that is, combining multi-source data to analyze the spatio-temporal evolution characteristics of individual and overall traffic flow, so as to optimize traffic signal control. Summary of the Invention
[0003] The present invention provides a traffic signal dynamic optimization method and system for tracking user spatio-temporal behavior, captures the spatio-temporal characteristics of the collective driving behavior of users in the lane, combines the dynamic road network state complexity of the lane to evaluate the priority of traffic signals, aims to reduce the road network state complexity, generates actions to adjust the traffic signal length vector, dynamically adjusts traffic signals, realizes more intelligent signal timing, and reduces the pedestrian waiting time and the probability of vehicle deceleration and stop.
[0004] To achieve the above object, a traffic signal dynamic optimization method for tracking user spatio-temporal behavior provided by the present invention includes the following steps:
[0005] S1: Obtain the driving behavior data of users and the road network state data;
[0006] S2: Use the driving behavior prediction model to analyze the behavior characteristics of the driving behavior data of users to obtain the spatio-temporal characteristics of the collective driving behavior;
[0007] S3: Combine the collective driving behavior characteristics and the road network state data to evaluate the priority of traffic signal optimization to obtain the priority of traffic signal optimization;
[0008] S4: Use the reinforcement learning method to construct a traffic signal dynamic optimization model;
[0009] S5: Take the traffic signals with priorities higher than the preset priority threshold as the traffic signals to be optimized, and use the traffic signal dynamic optimization model to generate the dynamic optimization results of the traffic signals to be optimized.
[0010] As a further improvement method of the present invention:
[0011] Optionally, the driving behavior data of the user includes the position, speed, acceleration, and driving direction of the vehicle driven by the user at different times, where the user is the driver of the vehicle;
[0012] The road network state data is the road network state data of different lanes, including the average number of instantaneous passing vehicles, the average vehicle speed, and the average number of pedestrians waiting at the intersections associated with the lanes within a period of time. The lane is the road on which the vehicle travels, and the intersections at both ends of the lane are the intersections associated with the lane, and traffic lights are set at the intersections.
[0013] Optionally, the driving behavior prediction model includes an input layer, a lane information conversion layer, a spatio-temporal encoding layer, and a spatio-temporal feature output layer. The input layer is used to receive the collected driving behavior data. The lane information conversion layer is used to convert the driving behavior data into a spatio-temporal information matrix representing lane driving information. The spatio-temporal encoding layer is used to perform spatio-temporal encoding on the time and position in the spatio-temporal information matrix to obtain the spatio-temporal encoding vectors in the spatio-temporal information matrix, and correct the spatio-temporal information matrix. The spatio-temporal feature output layer is used to extract features from the corrected spatio-temporal information matrix to obtain the spatio-temporal features of the group driving behavior of the user on different roads.
[0014] Optionally, the corrected spatio-temporal information matrix represents the corrected lane driving information of the nth lane, n ∈ [1, N], where N represents the total number of lanes. The spatio-temporal feature output layer is a long short-term memory neural network structure. The process of extracting features from the corrected spatio-temporal information matrix is as follows:
[0015] Extract the corrected lane driving information of different lanes, and input the corrected lane driving information into the forget gate, input gate, and output gate in sequence to obtain the lane driving features of different lanes;
[0016] Adopt a local attention mechanism to perform local attention calculation on all lane driving features to obtain the local attention of the lane driving features, and use the local attention to perform attention weighting on the lane driving features;
[0017] Input the weighted lane driving features into the forget gate, input gate, and output gate in sequence to obtain the spatio-temporal features of the group driving behavior of different lanes. The spatio-temporal feature of the group driving behavior of the nth lane is f n :
[0018]
[0019] Wherein:
[0020] The nth lane is evenly divided into G lane segments, representing the spatio-temporal characteristics of driving behavior of the gth lane segment, where g ∈ [1, G], and the spatio-temporal characteristics of driving behavior include the position of the gth lane segment, the acceleration probability, deceleration probability, parking probability, speed probability distribution, and acceleration probability distribution of the gth lane segment at different time periods, where the speed probability distribution is the probability value of different speeds, and the acceleration probability distribution is the probability value of different accelerations.
[0021] Optionally, the traffic signal is the control signal of a traffic light. The traffic lights are located at both ends of the lane. The number of traffic lights is M, and the traffic signal of the mth traffic light is E m , where m ∈ [1, M]. A priority evaluation formula is constructed, and the road network state data and group driving behavior characteristics of the lanes associated with the traffic lights are obtained to evaluate the priority of traffic signal optimization.
[0022] Optionally, the policy matrix in the traffic signal dynamic optimization model is a parameter to be solved. A reward function is constructed to calculate the policy value between the state and the action in the policy matrix, where the reward function is:
[0023] R(Y,A) = Load(Y) - Load(Y;A)
[0024]
[0025] Wherein:
[0026] R(Y,A) represents the reward value of taking action A in state Y. Load(Y) represents the complexity of the road network state of the lane corresponding to state Y. Load(Y,1), Load(Y,2), Load(Y,3) are the road network state data of the lanes associated with the traffic lights in state Y, representing the average value of the instantaneous passing vehicles, the average vehicle speed, and the average number of pedestrians waiting at the intersection associated with the lane within a period of time in sequence;
[0027] max1 is the preset maximum value of the instantaneous passing vehicles, max2 is the preset maximum vehicle speed, and max3 is the preset maximum number of pedestrians waiting.
[0028] Optionally, a loss function for solving the policy matrix is constructed based on the reward function, and the loss function is solved to obtain the policy matrix in the traffic signal dynamic optimization model, where the representation form of the loss function is F(θ):
[0029]
[0030] θ = [θ(Y h ; A q )] H×Q
[0031]
[0032] where:
[0033] θ represents the policy matrix to be solved, Y h represents the h-th state in the state space, h ∈ [1, H], where H represents the number of states in the state space, A q represents the q-th action in the action space, q ∈ [1, Q], where Q represents the number of actions in the action space, θ(Y h ; A q ) represents the policy value between state Y h and action A q ;
[0034] represents the advantage policy value of action A q for state Y h ;
[0035] represents the logarithmic gradient of θ(Y h ; A q ).
[0036] Optionally, the traffic signal dynamic optimization model is used to receive the road network state data of the traffic signal to be optimized, the spatio-temporal characteristics of the collective driving behavior, and the current traffic signal length vector, use the received data as the state, and based on the policy matrix, select the action with the highest policy value between the states to optimize and adjust the current traffic signal length vector of the traffic signal to be optimized, and use the optimization and adjustment result as the dynamic optimization result of the traffic signal to be optimized.
[0037] To solve the above problems, the present invention provides a traffic signal dynamic optimization system for user spatio-temporal behavior tracking. The traffic signal dynamic optimization system for user spatio-temporal behavior tracking includes a server and a data sensing device. The server includes a priority evaluation module and a signal optimization module:
[0038] The priority evaluation module is used to evaluate the priority of traffic signal optimization by combining the collective driving behavior characteristics and the road network state data to obtain the priority of traffic signal optimization;
[0039] The signal optimization module is used to receive the road network state data, spatio-temporal characteristics of group driving behavior, and the current traffic signal length vector of the lanes associated with the traffic signal to be optimized by using the traffic signal dynamic optimization model, and generate a dynamic optimization result of the traffic signal;
[0040] The data perception device is used to collect the driving behavior data of users and the road network state data, and perform behavior feature analysis on the driving behavior data of users to obtain the spatio-temporal characteristics of group driving behavior.
[0041] To solve the above problems, the present invention also provides an electronic device, which includes:
[0042] A memory that stores at least one instruction;
[0043] A communication interface for realizing the communication of the electronic device; and
[0044] A processor that executes the instructions stored in the memory to implement the above-mentioned traffic signal dynamic optimization method for user spatio-temporal behavior tracking.
[0045] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned traffic signal dynamic optimization method for user spatio-temporal behavior tracking.
[0046] Compared with the prior art, the present invention proposes a traffic signal dynamic optimization method for user spatio-temporal behavior tracking, and this technology has the following advantages:
[0047] First of all, this solution proposes a method for extracting spatio-temporal characteristics of group driving behavior, including spatio-temporal coding correction and feature extraction processing based on local attention. During the local attention calculation of lane driving features, a mask matrix is introduced for local weighting. The mask matrix is an N×N matrix composed of the distance measurement results between different lane driving features. The distance measurement method between lane driving features is cosine similarity. If the cosine similarity between the spatio-temporal coding vectors associated with different lane driving features is higher than a preset similarity threshold, the distance measurement result between different lane driving features is 1, otherwise it is 0, further emphasizing the correlation influence between similar spatio-temporal information, realizing the local attention correlation influence, and obtaining the group driving behavior features representing spatio-temporal correlation.
[0048] Meanwhile, this solution proposes a dynamic priority calculation method, which evaluates the priority of traffic signal optimization based on the road network status data and group driving behavior characteristics of the lanes associated with the traffic lights. During the priority evaluation process, the user performance and lane performance of the lanes are evaluated respectively. The user performance is characterized by the deceleration probability and stopping probability of the user, and the lane performance is characterized by the vehicle driving speed, the number of vehicles passing through instantaneously, and the number of pedestrians waiting. The higher the deceleration probability and stopping probability of the user, the more pedestrians waiting, and the lower the driving speed, the higher the degree of lane congestion and the higher the priority. During the calculation of the road network status complexity, a dynamic optimization weight is introduced, and the road network status complexity is dynamically weighted according to the number of traffic signal optimizations and the number of pedestrians waiting, reasonably adjusting the proportion of user performance and lane performance in the priority evaluation process, selecting the traffic signals to be optimized, and using the reinforcement learning method to generate actions to adjust the traffic signal length vector with the goal of reducing the road network status complexity, dynamically adjusting the traffic signals, realizing more intelligent signal timing, and reducing the pedestrian waiting time and the vehicle deceleration and stopping probability. Brief Description of the Drawings
[0049] Figure 1 It is a schematic flow chart of a traffic signal dynamic optimization method for user spatio-temporal behavior tracking provided by an embodiment of the present invention;
[0050] Figure 2 It is a functional module diagram of a traffic signal dynamic optimization system for user spatio-temporal behavior tracking provided by an embodiment of the present invention;
[0051] Figure 2 In the figure: 100 is the traffic signal dynamic optimization system for user spatio-temporal behavior tracking, 101 is the priority evaluation module, 102 is the signal optimization module, and 103 is the data perception device;
[0052] The realization, functional characteristics, and advantages of the purpose of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0053] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] An embodiment of the present application provides a dynamic optimization method for traffic signals based on user spatio-temporal behavior tracking. The execution subject of the dynamic optimization method for traffic signals based on user spatio-temporal behavior tracking includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the dynamic optimization method for traffic signals based on user spatio-temporal behavior tracking can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0055] Referring to Figure 1 , Embodiment 1 of the present invention is as follows:
[0056] S1: Obtain the driving behavior data of the user and the road network status data.
[0057] The driving behavior data of the user includes the position, speed, acceleration, and driving direction of the vehicle driven by the user at different times, where the user is the vehicle driver; specifically, the driving behavior data is collected by intelligent perception devices, where the intelligent perception devices include a GPS positioning system, in-vehicle sensors, etc.;
[0058] The road network status data is the road network status data of different lanes, including the average number of instantaneous passing vehicles, the average vehicle speed, and the average number of pedestrians waiting at the intersections associated with the lanes within a period of time, where the lane is the road on which the vehicle travels, and the intersections at both ends of the lane are the intersections associated with the lane, and traffic lights are set at the intersections.
[0059] Specifically, the average number of instantaneous passing vehicles of the lane is collected by a geomagnetic coil, the average vehicle speed of the lane is collected by a multi-target radar, and the number of pedestrians waiting at the intersection is collected by a camera and a target detection algorithm. The collection time interval of the driving behavior data and the road network status data is greater than time, where time represents the preset shortest traffic signal optimization interval, to avoid multiple optimization adjustments of the traffic signal in a short time and affect traffic safety.
[0060] S2: Use the driving behavior prediction model to perform behavior feature analysis on the driving behavior data of the user to obtain the spatio-temporal characteristics of the group driving behavior.
[0061] The driving behavior prediction model includes an input layer, a lane information conversion layer, a spatio-temporal encoding layer, and a spatio-temporal feature output layer. The input layer is used to receive the collected driving behavior data. The lane information conversion layer is used to convert the driving behavior data into a spatio-temporal information matrix representing lane driving information. The spatio-temporal encoding layer is used to perform spatio-temporal encoding on the time and position in the spatio-temporal information matrix to obtain spatio-temporal encoding vectors in the spatio-temporal information matrix, and correct the spatio-temporal information matrix. The spatio-temporal feature output layer is used to extract features from the corrected spatio-temporal information matrix to obtain the spatio-temporal features of the group driving behavior of users on different roads;
[0062] The representation form of the spatio-temporal information matrix is S:
[0063] S = [S1, S2,..., S n ,..., S N T
[0064]
[0065] Where:
[0066] T represents transpose, and S n represents the lane driving information of the nth lane, n ∈ [1, N], and N represents the total number of lanes, represents that the position of the driving vehicle is at the num n th driving behavior data of the nth lane, represents that the position of the driving vehicle is at the ith driving behavior data of the nth lane, i ∈ [1, num n , and num n represents the total number of driving behavior data of the driving vehicle at the nth lane;
[0067] successively represent the time, position, speed, acceleration, and driving direction in the driving behavior data ;
[0068] The correction result of the spatio-temporal information matrix S is
[0069]
[0070] Where:
[0071] represents the correction result of the lane driving information S n , represents the correction result of the driving behavior data , represents the correction result of the driving behavior data The spatio-temporal coding vector, where ε represents the spatio-temporal coding coefficient, represents the position The horizontal direction difference from the center position of the nth lane, represents the position The vertical direction difference from the center position of the nth lane;
[0072] represents vector concatenation.
[0073] The corrected spatio-temporal information matrix represents the corrected lane driving information of the nth lane, n ∈ [1, N], where N represents the total number of lanes. The spatio-temporal feature output layer is a long short-term memory neural network structure, and the process of extracting features from the corrected spatio-temporal information matrix is as follows:
[0074] Extract the corrected lane driving information of different lanes, and sequentially input the corrected lane driving information into the forget gate, input gate, and output gate to obtain the lane driving features of different lanes;
[0075] Adopt a local attention mechanism to perform local attention calculation on all lane driving features to obtain the local attention of the lane driving features, and use the local attention to perform attention weighting on the lane driving features;
[0076] Sequentially input the weighted lane driving features into the forget gate, input gate, and output gate to obtain the spatio-temporal features of the group driving behavior of different lanes. The spatio-temporal features of the group driving behavior of the nth lane are f n :
[0077]
[0078] Where:
[0079] The nth lane is evenly divided into G lane segments, represents the spatio-temporal features of the driving behavior of the gth lane segment, g ∈ [1, G]. The spatio-temporal features of the driving behavior include the position of the gth lane segment, the acceleration probability, deceleration probability, parking probability, speed probability distribution, and acceleration probability distribution of the gth lane segment at different time periods. The speed probability distribution is the probability value of different speeds, and the acceleration probability distribution is the probability value of different accelerations.
[0080] As a preferred embodiment of the present invention, in the process of calculating the local attention of the lane driving characteristics, a mask matrix is introduced for local weighting. The mask matrix is an N×N matrix composed of the distance measurement results between different lane driving characteristics. The distance measurement method between the lane driving characteristics is cosine similarity. If the cosine similarity between the spatio-temporal encoding vectors associated with different lane driving characteristics is higher than the preset similarity threshold, the distance measurement result between different lane driving characteristics is 1, otherwise it is 0, further emphasizing the correlation impact between similar spatio-temporal information and realizing the local attention correlation impact;
[0081] Hierarchical feature distillation is introduced in the calculation process of the input gate. The hierarchical feature distillation includes two feature distillation methods. The first feature distillation method is feature transformation, normalization, and max pooling processing. The second feature distillation method is the memory of important feature information based on Sigmoid gating and historical information splicing, further retaining key feature information.
[0082] S3: Combine the group driving behavior characteristics and the road network state data to evaluate the priority of traffic signal optimization, and obtain the priority of traffic signal optimization.
[0083] The traffic signal is the control signal of the traffic light. The traffic lights are located at both ends of the lane. The number of traffic lights is M, and the traffic signal of the m-th traffic light is E m , m ∈ [1, M]. Construct a priority evaluation formula, and obtain the road network state data and the group driving behavior characteristics of the lane associated with the traffic light, and evaluate the priority of traffic signal optimization. Among them, the traffic signal E m The priority of optimization is Pri(m):
[0084]
[0085]
[0086] Where:
[0087] β represents the priority regulation coefficient, and Pri(·) is the priority evaluation formula;
[0088] g ∈ [1, G], representing the content of the group driving behavior characteristics of the lane associated with the m-th traffic light, represents the deceleration probability vector of the g-th lane segment of the lane associated with the m-th traffic light at different time periods, represents the parking probability vector of the g-th lane segment of the lane associated with the m-th traffic light at different time periods. w1 represents the deceleration control coefficient vector, w2 represents the parking control coefficient vector, ||·||2 represents the L2 norm, L grepresents the lane weight coefficient of the g-th lane segment;
[0089] λ m represents the dynamic optimization weight of traffic signal E m λ m (-1) represents the dynamic optimization weight of traffic signal E m in the previous traffic signal optimization process, represents the dynamic learning rate of traffic signal E, K0 represents the preset initial learning rate, and count m represents the number of optimization times of traffic signal E m in a day, and C(·) represents the queuing saturation function; m in a day, and C(·) represents the queuing saturation function;
[0090] represents the driving behavior complexity of the lane associated with the m-th traffic signal, and Load m represents the dynamic road network state complexity of the lane associated with the m-th traffic signal;
[0091] successively represent the average value of the instantaneous passing vehicle numbers, the average vehicle speed, and the average number of pedestrians waiting at the intersection associated with the lane of the m-th traffic signal within a period of time. max1 is the preset maximum value of the instantaneous passing vehicle numbers, max2 is the preset maximum vehicle speed, and max3 is the preset maximum number of pedestrians waiting;
[0092] Specifically, the lane associated with the traffic signal is the lane adjacent to the traffic signal and in the opposite direction to the traffic signal orientation.
[0093] S4: Construct a traffic signal dynamic optimization model by using the reinforcement learning method.
[0094] The policy matrix in the traffic signal dynamic optimization model is a parameter to be solved. Construct a reward function to calculate the policy value between the state and the action in the policy matrix, where the reward function is:
[0095] R(Y,A) = Load(Y) - Load(Y; A)
[0096]
[0097] where:
[0098] R(Y, A) represents the reward value for taking action A in state Y, Load(Y) represents the road network state complexity of the lane corresponding to state Y, and Load(Y, 1), Load(Y, 2), Load(Y, 3) are the road network state data of the lanes associated with the traffic lights in state Y, representing the average number of instantaneous passing vehicles, the average vehicle speed, and the average number of pedestrians waiting at the intersection associated with the lane within a period of time, respectively;
[0099] max1 is the preset maximum value of the instantaneous passing vehicles, max2 is the preset maximum vehicle speed, and max3 is the preset maximum number of pedestrians waiting.
[0100] Construct a loss function for solving the policy matrix based on the reward function, and solve the loss function to obtain the policy matrix in the traffic signal dynamic optimization model, where the representation form of the loss function is F(θ):
[0101]
[0102] θ = [θ(Y h ; A q )] H×Q
[0103]
[0104] Where:
[0105] θ represents the policy matrix to be solved, Y h represents the h-th state in the state space, h ∈ [1, H], H represents the number of states in the state space, A q represents the q-th action in the action space, q ∈ [1, Q], Q represents the number of actions in the action space, and θ(Y h ; A q ) represents the policy value between state Y h and action A q ;
[0106] represents the advantage policy value of action A q for state Y h ;
[0107] represents the logarithmic gradient of θ(Y h ; A q ). As an embodiment of the present invention, one or more algorithms among the gradient descent algorithm, the Newton iteration method, and the Adam optimizer are used to solve the loss function.
[0108] S5: Take the traffic signals with priorities higher than the preset priority threshold as the traffic signals to be optimized, and use the traffic signal dynamic optimization model to generate the dynamic optimization results of the traffic signals to be optimized.
[0109] Use the traffic signal dynamic optimization model to receive the road network state data, spatio-temporal characteristics of group driving behavior, and the current traffic signal length vector of the traffic signals to be optimized. Take the received data as the state, and based on the policy matrix, select the action with the highest policy value between the states to optimize and adjust the current traffic signal length vector of the traffic signals to be optimized, and take the optimization and adjustment result as the dynamic optimization result of the traffic signals to be optimized. Specifically, the road network state data associated with traffic signal E m is The road network state data associated with traffic signal E m The spatio-temporal characteristics of group driving behavior associated with it are the spatio-temporal characteristics of group driving behavior of the lanes associated with the mth traffic signal.
[0110] Example 2:
[0111] This solution conducts a comparative experiment on the traffic signal dynamic optimization method, fixed-time control method, and inductive control method for tracking the spatio-temporal behavior of users, evaluates the degree of lane congestion and the average waiting time of pedestrians after traffic signal optimization, and the comparative experiment results are shown in Table 1:
[0112] Table 1
[0113]
[0114] As shown in Table 1, the traffic signal dynamic optimization method for tracking the spatio-temporal behavior of users adopts an overall optimization strategy for multiple lanes, which can significantly reduce the number of vehicle stops and the waiting time of pedestrians. The traffic signal optimization method can improve the traffic efficiency of vehicles, but fails to reduce the waiting time of pedestrians.
[0115] Example 3:
[0116] As Figure 2 shown, it is the functional module diagram of the traffic signal dynamic optimization system 100 for tracking the spatio-temporal behavior of users provided by an embodiment of the present invention, which can implement the traffic signal dynamic optimization method for tracking the spatio-temporal behavior of users in Example 1.
[0117] According to the implemented functions, the traffic signal dynamic optimization system 100 for tracking the spatio-temporal behavior of users may include a priority evaluation module 101, a signal optimization module 102, and a data perception device 103. The modules of the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions.
[0118] The priority evaluation module 101 is used to evaluate the priority of traffic signal optimization by combining the characteristics of group driving behavior and road network status data, and obtain the priority of traffic signal optimization;
[0119] The signal optimization module 102 is used to use the traffic signal dynamic optimization model to receive the road network status data, the spatio-temporal characteristics of group driving behavior, and the current traffic signal length vector of the lane associated with the traffic signal to be optimized, and generate a dynamic optimization result of the traffic signal;
[0120] The data perception device 103 is used to collect the driving behavior data of users and the road network status data, and perform behavior feature analysis on the driving behavior data of users to obtain the spatio-temporal characteristics of group driving behavior.
[0121] Specifically, each module in the traffic signal dynamic optimization system 100 for user spatio-temporal behavior tracking in the embodiments of the present invention adopts the same technical means as those in the Figure 1 traffic signal dynamic optimization method for user spatio-temporal behavior tracking described above, and can produce the same technical effects, which will not be elaborated here.
[0122] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of patent applications.
[0123] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variant thereof in this article are intended to cover a non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including the element.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0125] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A traffic signal dynamic optimization method based on user spatiotemporal behavior tracking, characterized in that: The method comprises: S1: Obtain the user's driving behavior data and road network status data; S2: Use the driving behavior prediction model to analyze the user's driving behavior data and obtain the spatiotemporal characteristics of group driving behavior; S3: Priority evaluation of traffic signal optimization is performed based on group driving behavior characteristics and road network status data to obtain the priority of traffic signal optimization; S4: constructing a traffic signal dynamic optimization model by using reinforcement learning, wherein the traffic signal dynamic optimization model includes a state space, an action space, and a strategy matrix; The state space includes all states of traffic signals, each group of states includes road network state data, spatiotemporal characteristics of group driving behavior, and current traffic signal length vector, wherein the traffic signal length vector is composed of the red light time length, yellow light time length, and green light time length of the traffic signal, the action space includes all actions of the traffic signal, each group of actions is an optimized adjustment vector of the traffic signal, used to adjust the traffic signal length vector, and the strategy matrix is composed of strategy values between states and actions; S5: The traffic signal with a priority higher than the preset priority threshold is taken as the traffic signal to be optimized, and a dynamic optimization result of the traffic signal to be optimized is generated by using the traffic signal dynamic optimization model.
2. A traffic signal dynamic optimization method for user spatiotemporal behavior tracking as claimed in claim 1, characterized in that: The user's driving behavior data includes the position, speed, acceleration and driving direction of the vehicle driven by the user at different times; The road network status data are the road network status data of different lanes, including the average number of vehicles passing through different lanes within a period of time, the average speed of vehicles, and the average number of pedestrians waiting at the intersections associated with the lanes, where the lanes are the roads on which vehicles travel, the intersections at both ends of the lanes are the intersections associated with the lanes, and traffic lights are set at the intersections.
3. The traffic signal dynamic optimization method for user spatiotemporal behavior tracking according to claim 1, characterized in that: The driving behavior prediction model includes an input layer, a lane information conversion layer, a spatiotemporal coding layer and a spatiotemporal feature output layer. The input layer is used to receive the collected driving behavior data. The lane information conversion layer is used to convert the driving behavior data into a spatiotemporal information matrix representing lane driving information. The spatiotemporal coding layer is used to spatiotemporally encode the time and position in the spatiotemporal information matrix to obtain the spatiotemporal coding vector in the spatiotemporal information matrix, and correct the spatiotemporal information matrix. The spatiotemporal feature output layer is used to output the corrected spatiotemporal information matrix. Feature extraction is performed to obtain the spatiotemporal characteristics of group driving behaviors of users on different roads.
4. A traffic signal dynamic optimization method for user spatiotemporal behavior tracking as claimed in claim 4, characterized in that: The corrected spatiotemporal information matrix T stands for transpose, represents the corrected lane driving information of the nth lane, n∈[1,N], N represents the total number of lanes, and the spatiotemporal feature output layer is a long short-term memory neural network structure. The process of feature extraction is: Extract the corrected lane driving information of different lanes, input the corrected lane driving information into the forget gate, the input gate, and the output gate in sequence, and obtain the lane driving features of different lanes; The local attention mechanism is used to calculate the local attention of all lane driving features, obtain the local attention of the lane driving features, and use the local attention to weight the lane driving features; The weighted lane driving features are sequentially input into the forget gate, the input gate, and the output gate to obtain the spatiotemporal features of group driving behaviors in different lanes. The spatiotemporal features of group driving behaviors in the nth lane are f n : in: The nth lane is divided into G lane segments. represents the spatiotemporal characteristics of driving behavior in the g-th lane segment, g∈[1,G], It includes the position of the g-th lane segment, the acceleration probability, deceleration probability, parking probability, speed probability distribution and acceleration probability distribution of the g-th lane segment in different time periods.
5. The method for dynamic optimization of traffic signals by tracking user spatiotemporal behavior according to claim 1, characterized in that: The traffic signal is a control signal of a traffic light. The traffic lights are located at both ends of the lane. The number of the traffic lights is M, and the traffic signal of the mth traffic light is E. m , m∈[1,M], construct a priority evaluation formula, obtain the road network status data of the lanes associated with the traffic lights and the group driving behavior characteristics, and evaluate the priority of traffic signal optimization. m The optimization priority is Pri(m): in: β represents the priority control coefficient, Pri(·) is the priority evaluation formula; represents the group driving behavior characteristics of the lane associated with the mth traffic light, represents the deceleration probability vector of the g-th lane segment associated with the m-th traffic light in different time periods, represents the parking probability vector of the g-th lane segment associated with the m-th traffic light in different time periods, w1 represents the deceleration control coefficient vector, w2 represents the parking control coefficient vector, ||·||2 represents the L2 norm, L g represents the lane weight coefficient of the g-th lane segment; λ m Indicates traffic signal E m Dynamic optimization weight, λ m (-1) indicates traffic signal E m The dynamic optimization weights during the last traffic signal optimization process, Indicates traffic signal E m Dynamic learning rate, K0 represents the preset initial learning rate, count m Indicates traffic signal E m The number of optimizations in a day, C(·) represents the queue saturation function; represents the driving behavior complexity of the lane associated with the mth traffic light, Load m Represents the dynamic road network state complexity of the lane associated with the mth traffic light; They represent respectively the average number of vehicles passing through the lane associated with the mth traffic light within a period of time, the average speed of vehicles, and the average number of pedestrians waiting at the intersection associated with the lane. max1 is the preset maximum number of vehicles passing through, max2 is the preset maximum speed of vehicles, and max3 is the preset maximum number of pedestrians waiting.
6. The method for dynamic optimization of traffic signals by tracking user spatiotemporal behavior according to claim 1, characterized in that: The policy matrix in the traffic signal dynamic optimization model is the parameter to be solved, and a reward function is constructed to calculate the policy value between the state and the action in the policy matrix, where the reward function is: in: R(Y,A) represents the reward value for taking action A in state Y, Load(Y) represents the road network state complexity of the lane corresponding to state Y, Load(Y,1), Load(Y,2), and Load(Y,3) are the road network state data of the lane associated with the traffic light in state Y, which respectively represent the average number of vehicles passing through the lane in a period of time, the average speed of the vehicles, and the average number of pedestrians waiting at the intersection associated with the lane.
7. A traffic signal dynamic optimization method for user spatiotemporal behavior tracking as claimed in claim 6, characterized in that: A loss function for solving the strategy matrix is constructed based on the reward function, and the loss function is solved to obtain the strategy matrix in the traffic signal dynamic optimization model, wherein the loss function is expressed as F(θ): in: θ represents the strategy matrix to be solved, Y h represents the hth state in the state space, h∈[1,H], H represents the number of states in the state space, A q represents the qth action in the action space, q∈[1,Q], Q represents the number of actions in the action space, θ(Y h ; A q ) indicates state Y h With action A q The strategy value between Indicates action A q For state Y h Advantage strategy value of represents θ(Y h ; A q ) is the logarithmic gradient of .
8. The method for dynamic optimization of traffic signals by tracking user spatiotemporal behavior according to claim 1, characterized in that: The traffic signal dynamic optimization model is used to receive the road network state data of the traffic signal to be optimized, the spatiotemporal characteristics of group driving behavior, and the current traffic signal length vector, the received data is used as the state, and based on the strategy matrix, the action with the highest strategy value between the states is selected to optimize and adjust the current traffic signal length vector of the traffic signal to be optimized, and the optimization adjustment result is used as the dynamic optimization result of the traffic signal to be optimized.
9. A traffic signal dynamic optimization system for tracking user spatiotemporal behavior, characterized in that: The traffic signal dynamic optimization system for tracking user spatiotemporal behavior includes a server and a data sensing device, wherein the server includes a priority evaluation module and a signal optimization module: The priority evaluation module is used to evaluate the priority of traffic signal optimization in combination with group driving behavior characteristics and road network status data to obtain the priority of traffic signal optimization; The signal optimization module is used to receive the road network status data of the lane associated with the traffic signal to be optimized, the spatiotemporal characteristics of the group driving behavior and the current traffic signal length vector by using the traffic signal dynamic optimization model, and generate the dynamic optimization result of the traffic signal; The data sensing device is used to collect the user's driving behavior data and road network status data, perform behavior feature analysis on the user's driving behavior data, and obtain the spatiotemporal characteristics of group driving behavior; To implement a traffic signal dynamic optimization method for tracking user spatiotemporal behavior as described in any one of claims 1-9.