A vehicle speed recommendation method and device for diverse dynamic signal lamp patterns

By modeling vehicles and traffic lights as collaborative agents, using advanced machine learning methods to predict signal light phase changes and optimize vehicle speed recommendations, the problem of vehicle speed recommendations under intelligent signal lights is solved, and a better travel experience is achieved.

CN116259175BActive Publication Date: 2025-06-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211719951.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-06-13
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

The prior art is difficult to provide effective vehicle speed recommendations under the intelligent signal light control strategy, and traditional methods rely heavily on the control of the front vehicle and the fixed traffic light mode, and cannot adapt to the diversified dynamic signal light mode.

Method used

By modeling the vehicle and traffic lights into collaborative heterogeneous collaborative agents, using phase-aware attention mechanisms, imitation learning and strategic gradient reinforcement learning methods, predict phase changes of the signal lights and combine multiple sensor information to optimize the optimal speed recommendation of green lights.

Benefits of technology

It realizes the multi-objective optimal speed recommendation in diversified dynamic signal light mode, improves the comfort, speed and safety of users' travel, and overcomes the limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116259175B_ABST
    Figure CN116259175B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle speed recommendation method and device for diverse dynamic signal light patterns. First, according to the traffic conditions of adjacent intersections and its own intersection, the phase perception attention mechanism is used to infer the proportion of the number of vehicles in each lane of this intersection after a few seconds; then, according to the proportion of the number of vehicles in each lane of this intersection in the previous K time periods and the proportion of the number of vehicles in each lane of this intersection inferred after a few seconds, imitation learning is used to approximately estimate the optimal preference of the signal lights, and the LSTM model is used to infer the signal light phase change sequence in the next period of time; finally, according to the predicted signal light phase change sequence, combined with the traffic conditions obtained by a variety of sensors, the method of policy gradient reinforcement learning is used to give the optimal green light speed that optimizes multiple targets under the influence of multi-dimensional data. The present invention can be applied to any scenario of urban roads, and the recommended speed is the multi-objective optimal speed, ensuring the safety, efficiency and time-saving characteristics of the optimal green light speed recommendation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of vehicle-road collaboration and optimal green light speed recommendation, and particularly to a vehicle speed recommendation method and device for diverse dynamic signal light patterns. Background Art

[0002] The optimal green light speed recommendation algorithm aims to provide appropriate speed suggestions, enabling vehicles to pass through intersections when the signal lights are green, thereby reducing stop-and-wait time, improving overall traffic efficiency, and reducing vehicle fuel consumption and carbon dioxide emissions. This algorithm needs to comprehensively consider intersection signal light information and the road conditions around the vehicle, so that the recommended optimal speed can approach the optimal solution under a multi-index evaluation system.

[0003] The related concept of optimal green light speed recommendation was proposed as early as the 20th century and can be classified into the following two categories according to traffic signal light patterns and speed recommendation forms:

[0004] (1) Recommend the optimal speed on a section of the path based on known traffic signal light information. Traditional traffic signal lights are basically in a fixed polling mode, and the signal lights switch orderly between several phases, and the relevant information is fixedly known. At this time, the constant driving speed on the entire section of the road can be given so that the vehicle can pass through the intersection when the signal light is green. However, this method is too idealistic. It is difficult for the vehicle to ensure precise uniform motion on the entire section of the road, and there will also be problems with excessive speed differences that are difficult to control when switching speeds.

[0005] (2) Recommend the optimal speed in real time based on known traffic signal light information. This method gives real-time speed suggestions according to real-time environmental information when making speed recommendations, and generally uses algorithms related to deep reinforcement learning for model training. Therefore, the recommended speed is the optimal speed for the current environment, and the long-term impact of the recommendation strategy is also emphasized. However, in order to improve traffic efficiency, some cities and regions have adopted intelligent signal light control strategies, making it impossible to predict traffic signal light information in advance. In this case, this method can only recommend following the speed of the vehicle in front intelligently, and the quality of regulation is greatly affected by the vehicle in front, and it cannot adapt to the restrictions of urban scene traffic signal lights.

[0006] The intelligent signal light control strategy can be divided into the following two categories: (1) The rule-based signal light control strategy, which can adaptively change the phase sequence and their respective durations according to the dynamic traffic environment. (2) The signal light control strategy based on deep reinforcement learning, which can implicitly predict traffic changes based on the original data, thereby giving a signal light control strategy that is better than the rule-based method.

[0007] It can be seen from this that the intelligent signal light control strategy is closely related to the spatio-temporal changes of vehicles. At present, there is no research on the vehicle speed recommendation method for intelligent signal lights.

[0008] Therefore, in view of the above problems, the present invention proposes an optimal speed recommendation algorithm based on diverse dynamic signal light patterns. By modeling vehicles and traffic signal lights as heterogeneous cooperative intelligent agents that cooperate with each other, it is possible to predict signal lights in both fixed and intelligent modes, and consider the relevance of multi-dimensional information, such as traffic conditions, signal light conditions, and surrounding road conditions, so as to give real-time optimal green light speed suggestions for the entire driving path. Summary of the Invention

[0009] The object of the present invention is to propose a vehicle speed recommendation method and device for diverse dynamic signal light patterns. The vehicle speed recommendation method based on heterogeneous cooperation is mainly used to provide reasonable vehicle speed suggestions for vehicles driving on urban roads, enabling users to obtain a better travel experience and improvements in aspects such as comfort, speed, and safety, overcoming the shortcomings and limitations of traditional optimal green light speed suggestion methods, and cooperating with vehicle speed regulation to optimize traffic conditions through spatio-temporal prediction of intelligent signal lights at intersections.

[0010] Since the current traffic flow conditions affect the changes of signal lights at intersections in the present invention, traffic flow data can be used to predict signal light changes. Then, by using various sensors such as millimeter-wave radars and GPS to obtain the perception information around the vehicle, the above multi-factor information can be integrated to provide users with optimal speed planning for the entire section of the path, while taking into account the optimization of multiple objectives. Through the dual reminders of mobile phone voice broadcasts and interface displays, the purpose of safely and efficiently assisting users in traveling can be achieved.

[0011] To achieve the above object, the present invention provides the following technical solutions:

[0012] In the first aspect, the present invention provides a vehicle speed recommendation method for diverse dynamic signal light patterns, including the following steps:

[0013] S1. According to the traffic conditions of adjacent intersections and its own intersection, use the phase-aware attention mechanism to infer the proportion of the number of vehicles in each lane of this intersection after a few seconds;

[0014] S2. According to the proportion of the number of vehicles in each lane of this intersection in the previous K time periods and the inferred proportion of the number of vehicles in each lane of this intersection after a few seconds, use imitation learning to approximately estimate the optimal preference of the signal light, and use the LSTM model to infer the signal light phase change sequence in the next period of time;

[0015] S3. Based on the predicted signal light phase change sequence, combined with the traffic conditions obtained by multiple sensors, use the method of policy gradient reinforcement learning to give the optimal green light speed for optimizing multiple targets under the influence of multi-dimensional data.

[0016] Further, the specific process of step S1 is as follows:

[0017] S11. Input the state information of this intersection and its adjacent intersections into the fully connected network to extract the hidden feature H j ; The state information includes the number of waiting vehicles in each incoming lane and the current signal light phase;

[0018] S12. Calculate the phase-aware attention score α ji :

[0019]

[0020] where r ji is the correlation coefficient of the phase-aware attention score, and the calculation formula is as follows:

[0021]

[0022] where, connected and non-connected indicate whether it is passable due to the influence of the signal light from I j to I i β is a hyperparameter used to balance the influence between different intersections on this intersection, and the set value is 0.5; is the average value of the passing times from all adjacent intersections of this intersection I i to this intersection, represents the set of intersections that have an adjacent relationship with this intersection I i geographically; T ji is the average passing time from the adjacent intersection I j to this intersection I i and is calculated through the distance between the adjacent intersection I j and this intersection I i and the average driving speed;

[0023] S13. Combine the hidden feature H j of each intersection and the corresponding phase-aware attention score α ji :

[0024]

[0025] where W q and W c are weight matrices, and b q is a bias vector;

[0026] S14. Use a fully connected network to obtain the proportion of the number of vehicles in each incoming lane at the current intersection

[0027] Furthermore, the specific process of step S2 is as follows:

[0028] S21. Output the predicted signal light phase sequence after t' seconds at intervals of n seconds

[0029] S22. Train each prediction value separately, and each training model obtains the i probability distribution of each signal light phase at this intersection The final predicted signal light phase is:

[0030]

[0031] S23. Through one layer of LSTM network and two layers of fully connected networks, with the final activation function being the softmax function, obtain the probability distribution of each signal light phase.

[0032] Furthermore, the multiple sensors in step S3 include:

[0033] A millimeter-wave radar deployed at the front end of the vehicle bumper and a Bluetooth converter deployed in the vehicle to connect the millimeter-wave radar, which are used to continuously monitor the relative position and relative speed with the vehicle ahead;

[0034] The GPS and accelerometer in the driver's mobile phone, which are used to obtain the position information and speed of the vehicle itself in real time;

[0035] And cameras of road infrastructure, which are used to capture the global traffic conditions.

[0036] Furthermore, the traffic conditions obtained in step S3 include local information and global information The local information includes: vehicle speed v t 、relative speed Δv with the vehicle ahead t and relative distance Δs t ; The global information includes the distance d from the current position to the signal light ahead t and the estimated arrival time ΔT t , as well as the predicted signal light phase sequence

[0037] Furthermore, step S3 optimizes multiple objectives, including setting a reward function from three aspects: the travel time, safety, and green light passing rate of the vehicle, where:

[0038] The formula for setting the reward function of the travel time of the vehicle is as follows:

[0039]

[0040] Among them, v max represents the speed limit on urban roads, that is, the maximum allowable speed; v t represents the vehicle's own speed value at the current moment, and k is a hyperparameter in this formula, such that the maximum value of kv t is 1;

[0041] The collision time is used to measure the occurrence probability of potential dangerous behaviors, and the safety reward function is set as follows:

[0042]

[0043] Among them, η is a hyperparameter in this formula, meaning the safety distance, which is set to 0.8;

[0044] The reward function for the green light passing rate of the vehicle is R 3 , and it is calculated based on the predicted time and predicted phase of the vehicle passing through the intersection whether the vehicle will pass through the intersection with a green light in the future. It is 1 when the light is green, otherwise it is -1;

[0045] The final reward function is R = R 1 +R 2 +R 3 .

[0046] Furthermore, the training process of the policy gradient reinforcement learning model in step S3 uses the Bellman equation to obtain the optimal cumulative discounted reward value.

[0047] Furthermore, the training process of the policy gradient reinforcement learning model in step S3 is as follows: The state S t , the action a t , the reward r t and the state S t+1 at the next time step are stored in the memory pool in the form of a tuple [S t , a t , r t , S t+1 . Each time, a batch of data is randomly sampled for training. The gradient update direction of the actor network is to improve its advantage value, and the loss function is:

[0048]

[0049] Among them, is the ratio between the actor behavior network p θ (s, a) and the actor target network p′ θ (s, a). The clip() function is used to limit the update amplitude within [1 - ε, 1 + ε]. The advantage value is:

[0050]

[0051] Among them, γ represents the discount value, and are the value function values of the critic behavior network and the critic policy network respectively. The loss function of the critic network is:

[0052]

[0053] Furthermore, the optimal green - light speed given in step S3 is output with an acceleration a t The value range of the acceleration is set to [-4m / s 2 , 2m / s 2 .

[0054] Second, the present invention also provides a vehicle - speed recommendation device for diverse dynamic signal - light patterns. The device includes the following modules to implement the steps of the vehicle - speed recommendation method for diverse dynamic signal - light patterns described in any one of the above:

[0055] A traffic - flow spatio - temporal relationship reasoning module, which is used to infer the proportion of the number of vehicles in each lane of this intersection a few seconds later by using a phase - aware attention mechanism according to the traffic conditions of adjacent intersections and its own intersection;

[0056] A signal - light behavior approximation module, which is used to approximately estimate the optimal preference of the signal - light by using imitation learning according to the proportion of the number of vehicles in each lane of this intersection in the previous K time periods and the proportion of the number of vehicles in each lane of this intersection inferred a few seconds later, and infer the signal - light phase change sequence in the next period of time by using an LSTM model;

[0057] A speed recommendation module, which is used to give the optimal green - light speed that optimizes multiple targets under the influence of multi - dimensional data by using the method of policy - gradient reinforcement learning according to the predicted signal - light phase change sequence and combining the traffic conditions obtained by multiple sensors.

[0058] Third, the present invention also provides an electronic device, which is characterized by including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0059] The memory is used to store a computer program;

[0060] The processor is used to implement the steps of the vehicle - speed recommendation method for diverse dynamic signal - light patterns described in any one of the above when executing the program stored in the memory.

[0061] Fourthly, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned vehicle speed recommendation methods for diverse dynamic signal lamp patterns are realized.

[0062] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0063] The existing green light optimal speed recommendation algorithm has the following disadvantages: on the one hand, the user participation is not high and the scenarios are limited. In order to ensure that the recommended speed approaches the optimal solution, previous algorithms often perform speed recommendation based on exact and fixed signal lamp control strategies. However, currently, some cities and regions have started to apply intelligent signal lamps for traffic regulation, and the characteristic is that the signal lamp information cannot be predicted in advance, which greatly affects the user participation to a large extent; some other methods focus on using surrounding information for speed recommendation. For example, the regulated vehicle follows the vehicle in front at the best speed smoothly. This method that only considers the information of the vehicle in front strongly depends on the regulation quality of the vehicle in front. In the actual application process, users often cannot guarantee the use efficiency due to limited scenarios. On the other hand, the speed recommendation efficiency is low and the recommendation target is single. The existing speed recommendation methods only focus on providing the speed of the user under a single target, such as the fastest speed in the time-saving mode, the most fuel-saving speed in the fuel-saving mode, or the smooth speed based on closed-loop in the comfort mode. The logic of these strategies is too simple to meet the various needs of users.

[0064] In contrast, the vehicle speed recommendation method and device for diverse dynamic signal lamp patterns proposed by the present invention can be applied to any scenario of urban roads, and the recommended speed is the multi-objective optimal speed. Through inverse reasoning based on the fact that the intelligent signal lamp control strategy is closely related to the traffic flow condition, the signal lamp phase sequence in the future period of time is predicted, so that the vehicle speed recommendation scenario is no longer limited to the fixed traffic light regulation scenario; at the same time, the present invention takes the information around the vehicle and the overall traffic information as comprehensive consideration factors, and is no longer limited to the scenario of following the vehicle in front. In addition, through the setting of the reward function in deep reinforcement learning, the safety, efficiency and time-saving characteristics of the green light optimal speed recommendation algorithm are also ensured. Description of the Drawings

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0066] Figure 1 It is a schematic diagram of the system architecture of a vehicle speed recommendation method and device for diverse dynamic signal lamp patterns provided by an embodiment of the present invention.

[0067] Figure 2 This is the speed recommendation interface for the mobile device provided by the embodiments of the present invention.

[0068] Figure 3 This is a schematic structural diagram of an electronic device for implementing a vehicle speed recommendation method and device for diverse dynamic signal light patterns provided by the embodiments of the present invention. Detailed implementation manners

[0069] To better understand the technical solution, the method of the present invention will be described in detail below with reference to the accompanying drawings.

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described examples are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope of protection of the present invention.

[0071] The overall system architecture of the vehicle speed recommendation method and device for diverse dynamic signal light patterns proposed by the present invention is as Figure 1 shown, including a traffic flow spatio-temporal relationship reasoning module, a signal light behavior approximation module, and a speed recommendation module. The traffic flow spatio-temporal relationship reasoning module estimates the percentage of the number of vehicles in each lane of this intersection a few seconds later based on the traffic conditions of surrounding intersections and its own intersection. The signal light behavior approximation module takes the output of the signal light spatio-temporal relationship reasoning module as input and uses imitation learning to approximately estimate the optimal preference of the signal light, thereby inferring the signal light phase change sequence in the next period of time. Next, the speed recommendation module uses the method of policy gradient reinforcement learning according to the signal light phase change sequence predicted by the previous module and combines the traffic conditions obtained by multiple sensors to give the optimal green light speed that optimizes multiple targets under the influence of multi-dimensional data.

[0072] (1) Traffic flow spatio-temporal relationship reasoning module

[0073] The traffic flow spatio-temporal relationship reasoning module needs to rely on the information of adjacent intersections that are related in time and space to speculate on the traffic flow changes at this intersection in the future. In terms of spatial relationship, it is necessary to consider the spatio-temporal correlation between multi-agent intersections to reason about the vehicles that may enter this intersection in the future; in terms of time relationship, it is necessary to combine the signal light changes and the vehicle conditions in the lanes affected by the signal lights to reason about the vehicle changes at this intersection in the future. Although there have been many previous works that can predict road conditions, these works generally focus on road changes at the regional level and minute level, and cannot achieve fine-grained accurate prediction reasoning at the lane level and second level. Due to the connection relationship between road networks, the graph attention mechanism can reflect the influence degree between each intersection through attention scores. However, the change frequency of signal lights and the instantaneous differences formed in traffic flow make it impossible to achieve by the simple graph attention mechanism. Therefore, the present invention proposes a phase-aware attention mechanism.

[0074] First, input the status information of this intersection and its adjacent intersections (the number of waiting vehicles in each incoming lane and the current signal light phase) into the fully connected network to extract the hidden feature H. j . Then calculate the phase-aware attention score. Here, this method calculates the average travel time T from intersection I j to intersection I i through the distance and average travel speed between adjacent intersections I j and I i . And the average value of the travel times from all adjacent intersections of this intersection I ji to this intersection is i . Among them, where represents the set of intersections that are geographically adjacent to this intersection I i . From this, the correlation coefficient of its attention score can be calculated:

[0075]

[0076] where, connected and unconnected indicate whether it is passable due to the influence of signal lights from I j to I i . β is the hyperparameter of this formula, used to balance the influence between different intersections on this intersection. Here, the value is set to 0.5. Next, the softmax() function can be used to normalize the correlation coefficient to obtain the final attention score:

[0077]

[0078] The phase-aware attention mechanism reflects the influence of different intersections on this intersection. It should be noted that the influence of this intersection on itself is the strongest.

[0079] Next, the hidden feature H of each intersection j and the corresponding attention score α ji are combined:

[0080]

[0081] where W q and W c are weight matrices, and b q is a bias vector. Then, a fully connected network is used to obtain the final proportion of the number of vehicles in each incoming lane at this intersection

[0082] (2) Signal Light Behavior Approximation Module

[0083] The signal light behavior approximation module aims to predict the phase change of the signal light in the future for a period of time. For intelligent signal lights, their control strategies need to use the vehicle distribution in the incoming lanes and the current signal light phase conditions to determine the future regulation direction. Generally, an offline supervised learning method can be used to infer the strategy with the help of a large amount of data resources, such as image classification technology. However, the signal light control strategy has a long-tail effect. Its strategy is a real-time sequence decision-making, which will have an impact on the future for a period of time. Therefore, here it is necessary to fit the signal light control strategy through historical observation data with spatio-temporal correlation. Imitation learning can extract the logic and preferences of the strategy, while the LSTM (Long Short-Term Memory) model can extract the time correlation of the phase sequence. The specific usage methods of these two technologies are as follows:

[0084] First, the input is the historical observation data (signal light phase and the proportion of the number of vehicles in the incoming lanes of the intersection) in the previous K time periods and the predicted proportion of the number of vehicles in the incoming lanes of the intersection obtained by the first module. The output is the signal light phase after t's, at intervals of 20s (i.e., 20s, 40s, 60s). Currently, the predicted phase sequence within the next minute is output The three prediction values are trained separately, that is, there are three independent models. Each model will obtain the probability distribution of each signal light phase at intersection I i The final predicted phase is:

[0085]

[0086] The network part goes through a layer of LSTM network and two layers of fully connected networks in sequence. The final activation function is the softmax() function, and the probability distribution of each phase can be obtained. Other settings such as using the Adam optimizer for the optimizer and the cross-entropy loss function for the loss function.

[0087] (3) Speed Recommendation Module

[0088] The speed recommendation module uses the PPO (proximal policy optimization) algorithm in the reinforcement learning method. This method has been proven to achieve better results in multiple scenarios and is more suitable for dealing with continuous space control problems. The PPO algorithm is an algorithm in the AC (Actor-Critic) system. This system generally uses multiple actors for information collection and a centralized critic for policy control to ensure the comprehensiveness of information collection and the unity of regulation. The agent design and model training are introduced separately below.

[0089] (A) Agent design

[0090] First, it is necessary to accurately, comprehensively, and completely obtain the real-time environmental information of the vehicle. Therefore, this algorithm extracts the following information as the input S of the reinforcement learning model t , which is divided into local information and global information The local information includes: the vehicle speed v t , the relative speed Δv with the vehicle ahead t and the relative distance Δs t . The global information includes the distance d from the current position to the traffic light ahead t and the estimated arrival time ΔT t , as well as the predicted traffic light sequence obtained through the second module Here, it is necessary to normalize the above information to eliminate the influence of weight distribution caused by information dimension differences, and then obtain the output of the model - acceleration a through the neural network t . Compared with directly outputting the speed, the acceleration pays more attention to the smoothness of the vehicle driving trajectory and eliminates the influence caused by the difficult-to-control sharp change in speed difference. The acceleration value range in this invention is set to [-4m / s 2 , 2m / s 2 . Another highlight of the optimal speed recommendation algorithm in this invention also lies in the setting of the reward function.

[0091] The setting of the reward function in reinforcement learning is a key link, which determines the optimization direction of the model, and the ultimate goal is to maximize the long-term cumulative reward function value. The goal of this algorithm is to provide the vehicle with the optimal speed based on the current environment, so that users can obtain a satisfactory travel experience. In terms of setting the reward function, it is necessary to pay attention to optimizing multiple goals simultaneously to ensure the global optimality of the vehicle speed recommendation. Here, this method considers three aspects: the travel time, safety, and green light passing rate of the vehicle. The reward function is set as follows:

[0092] First, in order to make the vehicle travel at a normal speed and minimize the travel time consumption, this method sets a reward function related to the vehicle speed, and the formula is as follows:

[0093]

[0094] where v max represents the speed limit on urban roads, that is, the maximum allowable speed. v t represents the vehicle's own speed value at the current moment, and k is a hyperparameter in this formula, such that the maximum value of kv a is 1.

[0095] The second term focuses on the safety of the vehicle during driving. Here, only the danger caused by the other vehicles on the road is considered. Therefore, this invention uses the Time To Collision (TTC) to measure the occurrence probability of potential dangerous behaviors, and the formula is as follows:

[0096]

[0097] η is a hyperparameter in this formula, which means the safety distance and is set to 0.8.

[0098] The third term R 3 is the green light passing rate of the vehicle. Whether the vehicle just passes through the traffic light when the signal is green is a low-probability event in the real-time regulation process, which will bring the impact of sparse rewards. Therefore, this invention calculates whether the vehicle will pass through the intersection in the future with a green light according to the predicted time and predicted phase of the vehicle passing through the intersection. It is 1 when the light is green, otherwise it is -1.

[0099] In summary, the final reward function can be obtained as R = R 1 +R 2 +R 3 .

[0100] (B) Model training

[0101] The key idea of model training is to use the Bellman equation to obtain the optimal cumulative discounted reward value. In this algorithm, multiple actors collect information, and the state S t , action a t , reward r t and the state S t+1 at the next time step are stored in the memory pool in the form of a tuple [S t , a t , r t , S t+1 . Each time, a batch of data is randomly sampled for training, usually 128. The gradient update direction of the actor network is to improve its advantage value, and the loss function is:

[0102]

[0103] Among them, is the ratio between the actor behavior network p θ (s, a) and the actor target network p' θ (s, a). To ensure progressive update, this method uses the clip() function to limit the update amplitude between [1 - ε, 1 + ε]. The advantage value is:

[0104]

[0105] Among them, γ represents the discount value, and are the value function values of the critic behavior network and the critic policy network respectively. The loss function of the critic network is:

[0106]

[0107] So far, the optimal speed recommended by the algorithm can be obtained.

[0108] In addition, for the convenience of use during actual driving, the present invention takes advantage of the portability of smartphones to develop a prototype system to help users obtain real-time optimal speed recommendations and optimize the travel experience. All the required information can be obtained through sensing devices. For example, millimeter-wave radars can be deployed at the front end of the vehicle bumper, and the connected Bluetooth converters can be deployed inside the vehicle to continuously monitor the relative position and relative speed with the vehicle ahead; the GPS and accelerometers in the mobile phone can also obtain the position information and speed of the vehicle itself in real time; road infrastructure such as cameras can also capture the global traffic conditions. All the above information can be uploaded to the cloud server through communication technologies such as 5G and WiFi for information processing and utilization.

[0109] During actual use, the user needs to turn on the Bluetooth setting when using the APP, input the destination for navigation, and the APP will give real-time speed recommendations to the user through voice announcements and interface displays. The speed recommendation interface is as Figure 2 shown. This interface uses relevant components of Amap Navigation to display information such as the current driving route and the current speed.

[0110] Corresponding to the device for generating millimeter-wave radar data based on video provided in the above embodiments of the present invention, embodiments of the present invention also provide an electronic device.

[0111] As Figure 3As shown, the electronic device includes a processor 201, a communication interface 202, a memory 203, and a communication bus 204. Among them, the processor 201, the communication interface 202, and the memory 203 complete communication with each other through the communication bus 204.

[0112] The memory 203 is used to store computer programs.

[0113] When the processor 201 is used to execute the program stored on the memory 203, it implements the steps of any one of the vehicle speed recommendation methods for diverse dynamic signal lamp modes provided in the embodiments of the present invention above.

[0114] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0115] The communication interface is used for communication between the above electronic device and other devices.

[0116] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0117] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0118] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the vehicle speed recommendation methods for diverse dynamic signal lamp patterns provided by the embodiments of the present invention are implemented.

[0119] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, the computer is made to execute the steps of any one of the vehicle speed recommendation methods for diverse dynamic signal lamp patterns provided by the embodiments of the present invention.

[0120] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0121] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0122] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, the electronic device embodiment, the computer-readable storage medium embodiment and the computer program product embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the description of the method embodiment.

[0123] The above is only the preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A vehicle speed recommendation method for diverse dynamic signal light patterns, characterized in that, it includes the following steps: S1. According to the traffic conditions of adjacent intersections and its own intersection, use the phase-aware attention mechanism to infer the proportion of the number of vehicles in each lane of this intersection after a few seconds; The specific process of step S1 is: S11. Input the status information of this intersection and its adjacent intersections into the fully connected network to extract the hidden feature H j ; The status information includes the number of waiting vehicles in each incoming lane and the current signal light phase; S12. Calculate the phase perception attention score α ji : where r ji is the correlation coefficient of the phase-aware attention score, and the calculation formula is as follows: Among them, connected and non - connected indicate from I j to I i Whether it is passable due to the influence of traffic lights. p is a hyperparameter used to balance the influence between different intersections on this intersection, and the set value is 0.5; For this intersection I i is the average of the travel times from all adjacent intersections of this intersection to this intersection, represents the set of intersections that have an adjacency relationship with this intersection I i geographically; T ji is the average travel time from the adjacent intersection I j to this intersection I i and is calculated through the distance and average driving speed between the adjacent intersection I j and this intersection I i ; S13. Combine the hidden feature H of each intersection j with the corresponding phase-aware attention score α ji as follows: where W q and W c are weight matrices, and b q is a bias vector; S14. Use a fully connected network to obtain the proportion of the number of vehicles in each incoming lane at this intersection S2. According to the proportion of the number of vehicles in each lane of this intersection in the previous K time periods and the inferred proportion of the number of vehicles in each lane of this intersection after a few seconds, use imitation learning to approximately estimate the optimal preference of the signal lights, and use the LSTM model to infer the signal light phase change sequence in the next period of time; The specific process of step S2 is: S21. Output the predicted signal light phase sequence after t' seconds at an interval of n seconds S22. Train each predicted value separately, and each training model obtains the probability distribution of each signal light phase at this intersection I i Probability distribution of each signal light phase The final predicted phase of the signal light is: S23. Through one layer of LSTM network and two layers of fully connected networks, with the final activation function being the softmax function, obtain the probability distribution of each signal light phase; S3. According to the predicted signal light phase change sequence, combined with the traffic conditions obtained by multiple sensors, use the method of policy gradient reinforcement learning to give the optimal green light speed that optimizes multiple targets under the influence of multi-dimensional data; The multiple sensors in step S3 include: A millimeter-wave radar deployed at the front end of the vehicle bumper and a Bluetooth converter deployed in the vehicle for connecting the millimeter-wave radar, used to continuously monitor the relative position and relative speed with the vehicle in front; The GPS and accelerometer in the driver's mobile phone, used to obtain the position information and speed of the vehicle itself in real time; And cameras of road infrastructure, used to capture the global traffic conditions. The traffic conditions obtained in step S3 include local information and global information The local information includes: vehicle speed v t , relative speed Δv with the vehicle ahead t and relative distance Δs t ; The global information includes the distance d from the current position to the traffic signal ahead t and the estimated arrival time ΔT t , as well as the predicted phase sequence of the traffic signal Step S3 optimizes multiple targets, including setting reward functions from three aspects: the travel time of the vehicle, safety, and green light passing rate, where: The formula for setting the reward function of the travel time of the vehicle is as follows: Among them, v max represents the speed limit on urban roads, that is, the maximum allowable speed; v t represents the vehicle's own speed value at the current moment, and k is a hyperparameter in this formula, such that the maximum value of kv t is 1; Use the time to collision to measure the occurrence probability of potential dangerous behaviors, and the formula for setting the safety reward function is as follows: Among them, η is a hyperparameter in this formula, meaning the safety distance, which is set to 0.8; The reward function for the green light passing rate of the vehicle is R 3 , and it is calculated whether the vehicle will pass through the intersection under a green light in the future based on the predicted time and predicted phase of the vehicle passing through the intersection. It is 1 when it is a green light, otherwise it is -1; The final reward function is R = R 1 + R 2 + R 3 ; The training process of the policy gradient reinforcement learning model in step S3 uses the Bellman equation to obtain the optimal cumulative discounted reward value; The training process of the policy gradient reinforcement learning model in step S3 is as follows: The state S t , the action a t , the reward r t and the state S t+1 at the next time step are stored in the memory pool in the form of a tuple [S t , a t , r t , S t+1 . Each time, a batch of data is randomly sampled for training. The gradient update direction of the actor network is to increase its advantage value, and the loss function is: Among them, is the actor behavior network p θ (s, a) and the actor target network p' θ The ratio between (s, a) is limited to [1 - ε, 1 + ε] using the clip() function, and the advantage value is: where γ represents the discount value, and are the value function values of the critic behavior network and the critic policy network respectively. The loss function of the critic network is:

2. The vehicle speed recommendation method for diverse dynamic signal light patterns according to claim 1, characterized in that, The optimal speed of the green light given in step S3 is at an acceleration of a t Output, and the value range of the acceleration is set to [-4 m / s 2 , 2 m / s 2 .

3. A vehicle speed recommendation device for diverse dynamic signal light patterns, characterized in that, The device includes the following modules to implement the method described in any one of claims 1-2: A traffic flow spatio-temporal relationship reasoning module, used to infer the proportion of the number of vehicles in each lane of this intersection after a few seconds according to the traffic conditions of adjacent intersections and its own intersection, using the phase-aware attention mechanism; A signal light behavior approximation module, used to approximately estimate the optimal preference of the signal lights using imitation learning according to the proportion of the number of vehicles in each lane of this intersection in the previous K time periods and the inferred proportion of the number of vehicles in each lane of this intersection after a few seconds, and use the LSTM model to infer the signal light phase change sequence in the next period of time; A speed recommendation module, used to give the optimal green light speed that optimizes multiple targets under the influence of multi-dimensional data according to the predicted signal light phase change sequence, combined with the traffic conditions obtained by multiple sensors, using the method of policy gradient reinforcement learning.

Citation Information

Patent Citations

  • Network connection vehicle signal lamp control intersection economic passing method based on reinforcement learning

    CN113269963A

  • Traffic signal optimization control method

    CN115171408A