A traffic driving strategy optimization method based on unmanned aerial vehicle observation joint modeling
By using UAV observations and joint modeling to acquire traffic data, and generating route and traffic light optimization strategies, the problem of lack of global optimization in traffic driving strategies is solved, thus improving the efficiency of road network operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, traffic driving strategies lack global optimization, and route planning is disconnected from traffic light control, resulting in low road network operating efficiency.
By using UAV observation and joint modeling, current and historical traffic data are acquired. Target path optimization algorithms and multi-level network structures are then used to generate path and traffic light optimization strategies, achieving information sharing and collaborative optimization.
It improves the overall operational efficiency of the road network, ensures the real-time nature of traffic strategies and data integrity, and achieves integrated optimization of route planning and traffic light control.
Smart Images

Figure CN120412277B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traffic flow optimization, in particular to a traffic driving strategy optimization method based on joint modeling of unmanned aerial vehicle observation. BACKGROUND
[0002] In the technical field of traffic flow optimization, traffic flow is analyzed to recommend traffic driving strategies for vehicle drivers.
[0003] In related technologies, the recommended traffic driving strategy only involves the optimal path provided for a single vehicle. On the one hand, there is a lack of global traffic driving strategy for all vehicles. On the other hand, path planning and signal light control are disconnected from each other, lacking effective information sharing and collaborative optimization mechanisms, which reduces the overall operational efficiency of the road network. SUMMARY
[0004] Therefore, it is necessary to provide a traffic driving strategy optimization method based on joint modeling of unmanned aerial vehicle observation, a device, a computer equipment and a computer readable storage medium to effectively and accurately improve the overall operational efficiency of the road network according to effective information sharing and collaborative optimization mechanisms.
[0005] In a first aspect, the present application provides a traffic driving strategy optimization method based on joint modeling of unmanned aerial vehicle observation, comprising:
[0006] In the current traffic environment, current image traffic data collected by a preset image capturing device in a current period and historical text traffic data stored in a historical period are obtained, and the image capturing device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to jointly observe the current traffic environment;
[0007] A target path optimization algorithm matching the complexity of the current traffic environment is determined, and a path optimization strategy is obtained by data analysis and processing of current traffic flow characteristics obtained from the current image traffic data based on the target path optimization algorithm, wherein the complexity of the current traffic environment is determined according to the current image traffic data;
[0008] Based on a preset multi-level structure network, the current traffic flow characteristics and predicted traffic flow characteristics obtained by combining the current image traffic data and the historical text traffic data are processed layer by layer to obtain a signal light optimization strategy;
[0009] The path optimization strategy and the signal light optimization strategy are used as traffic driving strategies for each vehicle in the current traffic environment, and the traffic driving strategies are used to provide driving path suggestions and driving speed suggestions for each vehicle in the current traffic environment.
[0010] In a second aspect, the application further provides a traffic driving strategy optimization device based on unmanned aerial vehicle observation joint modeling, comprising:
[0011] An acquisition module is configured to acquire current image traffic data collected by a preset image shooting device in a current time period and historical text traffic data stored in a historical time period in a current traffic environment, wherein the image shooting device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to jointly observe the current traffic environment.
[0012] A first optimization module is configured to determine a target path optimization algorithm matched with a complexity of the current traffic environment, perform data analysis and processing on current traffic flow features obtained from the current image traffic data based on the target path optimization algorithm, and obtain a path optimization strategy, wherein the complexity of the current traffic environment is determined according to the current image traffic data.
[0013] A second optimization module is configured to perform layer-by-layer data analysis and processing on the current traffic flow features and predicted traffic flow features obtained by combining the current image traffic data and the historical text traffic data based on a preset multi-level structure network, and obtain a signal light optimization strategy.
[0014] An execution module is configured to take the path optimization strategy and the signal light optimization strategy as traffic driving strategies of each vehicle in the current traffic environment, wherein the traffic driving strategies are used to provide driving path suggestions and driving speed suggestions for each vehicle in the current traffic environment.
[0015] In a third aspect, the application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above steps when executing the computer program.
[0016] In a fourth aspect, the application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above steps.
[0017] The traffic driving strategy optimization method, device, computer equipment and computer readable storage medium based on unmanned aerial vehicle observation joint modeling have the following advantages. Firstly, current image traffic data is collected through an image shooting device, and the current traffic environment is comprehensively perceived in combination with historical text traffic data, so that the subsequent optimization of the traffic driving strategy has real-time performance and data integrity. Secondly, a target path optimization algorithm matched to the complexity of the current traffic environment is used to analyze current traffic flow characteristics obtained from the current image traffic data to obtain a path optimization strategy, so that the driving path suggestion has individuality and global adaptability. Thirdly, a multi-level network is used to perform layer-by-layer analysis on current traffic flow characteristics and predicted traffic flow characteristics obtained in combination with the current image traffic data and the historical text traffic data to obtain a signal lamp optimization strategy, so that fine adjustment of the signal lamp setting time is realized. Based on this, the path optimization strategy and the signal lamp optimization strategy are used as the traffic driving strategy of each vehicle in the current traffic environment, so that fusion control and collaborative optimization of path planning and signal lamp control can be realized to effectively and accurately improve the overall operation efficiency of the road network according to an effective information sharing and collaborative optimization mechanism. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 A flowchart of a traffic driving strategy optimization method based on unmanned aerial vehicle observation joint modeling in an embodiment;
[0020] Figure 2 A flowchart of generating a path optimization strategy by a path optimization algorithm based on a time sequence network in an embodiment;
[0021] Figure 3 A structural diagram of a path optimization algorithm based on a time sequence network enhanced by Transformer relationship modeling in an embodiment;
[0022] Figure 4 A flowchart of training a time sequence network enhanced by Transformer relationship modeling in an embodiment;
[0023] Figure 5 A structural diagram of training a time sequence network enhanced by Transformer relationship modeling in an embodiment;
[0024] Figure 6 A flowchart of a path optimization strategy generation process based on improved ant algorithm combined with feedback of UAV in an embodiment;
[0025] Figure 7 A flowchart of a path optimization algorithm of improved ant algorithm combined with feedback of UAV in an embodiment;
[0026] Figure 8 A flowchart of a transition probability calculation process of improved ant algorithm combined with feedback of UAV in an embodiment;
[0027] Figure 9 A flowchart of a signal light optimization strategy generation process based on multi-level structure in an embodiment;
[0028] Figure 10 A structural diagram of a signal light optimization strategy generation process based on multi-level structure in an embodiment;
[0029] Figure 11 A flowchart of a real-time dynamic path induction strategy algorithm in an embodiment;
[0030] Figure 12 A flowchart of a real-time dynamic path induction strategy algorithm in another embodiment;
[0031] Figure 13 A structural block diagram of a traffic driving strategy optimization device based on UAV observation joint modeling in an embodiment. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0033] In an embodiment, as shown in Figure 1 , a traffic driving strategy optimization method based on UAV observation joint modeling is provided, and the present embodiment takes the method applied to a server as an example. It should be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and can be realized through the interaction of the terminal and the server. In the present embodiment, the method includes the following steps S101 to S104.
[0034] Among them, the traffic driving strategy represents the decision scheme such as path and speed planning, signal light duration planning formulated according to traffic data, which is used to guide the vehicle to select a more suitable driving mode in a specified traffic environment, so as to reduce congestion risk, shorten travel time or avoid high-risk road sections.
[0035] In the current traffic environment, current image traffic data collected by a preset image shooting device in a current period and historical text traffic data stored in a historical period are acquired, and the image shooting device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to jointly observe the current traffic environment.
[0036] The current traffic environment represents a traffic running state in a specified area, i.e., a comprehensive traffic condition for describing vehicles, pedestrians, road facilities, signal light control time length, etc. in the specified area.
[0037] The image shooting device represents an unmanned aerial vehicle device deployed in the current traffic environment to acquire real-time traffic pictures, i.e., to acquire traffic data in the form of image data related to vehicles, pedestrians, road facilities, etc. The image shooting device can represent a tethered unmanned aerial vehicle satellite hovering in the air, which is configured with a high-pixel camera and can achieve real-time centimeter-level resolution of ground objects.
[0038] The current image traffic data represents traffic data collected by the image shooting device in the current period, reflecting the traffic running state in the current period. The historical text traffic data represents traffic data stored in a text format, reflecting the traffic running state in the historical period.
[0039] Exemplarily, in a city traffic map, a physical space corresponding to a certain map area can be specified as the current traffic environment. For example, the entire city corresponding to the physical space in the city traffic map can be specified as the current traffic environment, or a plurality of interconnected trunk roads corresponding to the physical space in the city traffic map can be specified as the current traffic environment. Furthermore, as time goes by, a large amount of historical data reflecting the traffic conditions of each road in the current traffic environment at different historical time points can be generated based on the network model or algorithm for optimizing the traffic driving strategy in the embodiment, and the historical data is stored in a database in the form of text and used as historical text traffic data of the historical period in the scene of optimizing the traffic driving strategy in the current period.
[0040] The image capturing device represents a tethered unmanned aerial vehicle deployed in a city, that is, a number of tethered unmanned aerial vehicles are evenly arranged in the air above the city, and the image data with centimeter-level accuracy is obtained in real time and continuously through the camera unit in the tethered unmanned aerial vehicle, so as to realize comprehensive visual coverage and global traffic monitoring of the entire city traffic area. The number, distribution position, and spatial distance between the tethered unmanned aerial vehicles arranged in the air above the city can be determined in combination with the road distribution condition, city boundary contour, area size, and other factors of the city and the field of view of the tethered unmanned aerial vehicle, so as to comprehensively cover the entire city traffic area through the global field of view composed of a number of tethered unmanned aerial vehicles. The field of view of one tethered unmanned aerial vehicle can represent a full-range circular area with the position of the tethered unmanned aerial vehicle as the center and the farthest visible distance as the radius.
[0041] Furthermore, the structure and corresponding functions of the above tethered unmanned aerial vehicle can refer to a kind of sustained endurance internet unmanned aerial vehicle (publication number CN105652886A) in related technology, which is a tethered unmanned aerial vehicle satellite configured with high-pixel camera and the field of view can reach 5 square kilometers. On the one hand, the unmanned aerial vehicle can realize sustained endurance capability based on continuous power supply, so as to continuously maintain working state, on the other hand, based on higher wind resistance and rainproof capability, the unmanned aerial vehicle can continuously work under general weather conditions, based on this, image data with centimeter-level accuracy is obtained in real time and continuously.
[0042] In step S102, a target path optimization algorithm matching the complexity of the current traffic environment is determined, and the current traffic flow characteristics obtained from the current image traffic data are analyzed and processed based on the target path optimization algorithm to obtain a path optimization strategy. The complexity of the current traffic environment is determined according to the current image traffic data.
[0043] The complexity of the current traffic environment represents the dynamic change characteristics and structural complexity in the current traffic environment, for example: at the level of road facilities, the complexity of the current traffic environment is determined based on factors such as the size of the area involved in the traffic area, whether there is a construction section, whether there are complex roads with multiple branches, whether the signal lights at the intersection are too many, etc.; at the level of vehicle traffic, the complexity of the current traffic environment is determined based on factors such as the number and density of vehicles, whether the speed difference is obvious, whether the time and space distribution is uneven, etc.; at the level of pedestrian traffic, the complexity of the current traffic environment is determined based on factors such as the number and density of pedestrians, whether the pedestrian flow direction is diversified, whether there are non-motor vehicles, etc.
[0044] The target path optimization algorithm represents a specific calculation method for generating a path optimization strategy, and the computing power thereof matches the complexity of the current traffic environment, for example, a simple computing power path optimization algorithm is used as the target path optimization algorithm in a simple traffic environment, and an enhanced computing power path optimization algorithm is used as the target path optimization algorithm in a complex traffic environment.
[0045] The path optimization strategy represents a strategic result for guiding each vehicle in the current traffic environment to select a passing path, that is, it includes a recommended path set of different vehicles from the starting point to the terminal point.
[0046] The current traffic flow feature represents a representative statistical or distribution index extracted from the current image traffic data, which is used to describe the traffic running state in the current period, for example, the distance between each vehicle and the corresponding terminal point in the current period, the number of passing vehicles per unit time on each road, the average speed, the vehicle density, the lane saturation, the intersection queue length, and the like, and for example, the number and density of pedestrians passing through each road in the current period, the flow direction and trend, and the like.
[0047] Exemplarily, after obtaining the current image traffic data, the complexity of the current traffic environment is determined, so as to select a target path optimization algorithm with computing power suitable for the complexity of the current traffic environment; further, the current traffic flow feature extracted from the current image traffic data is taken as the input of the target path optimization algorithm, and in the processing process of the target path optimization algorithm on the current traffic flow feature, a plurality of reachable paths of the same vehicle are comprehensively compared and analyzed in terms of passing efficiency, delay time, path load, and the like, and finally a path optimization strategy with optimal passing performance of each vehicle under the current traffic condition is output.
[0048] In step S103, based on the preset multi-level structure network, the current traffic flow feature and the predicted traffic flow feature obtained by combining the current image traffic data and the historical text traffic data are processed layer by layer to obtain a signal light optimization strategy.
[0049] The multi-level structure network represents a calculation model with a hierarchical organization form for processing the current traffic flow feature and the predicted traffic flow feature, so as to sequentially process the input data in different calculation layers, thereby gradually analyzing the signal light optimization strategy.
[0050] The signal light optimization strategy represents a strategic result for guiding the setting time of each traffic signal light in the current traffic environment, that is, it includes a set of green light time and red light time of different traffic signal lights.
[0051] The prediction traffic flow feature represents a representative statistical or distribution index extracted from the current image traffic data and the historical text traffic data, and is used to describe the traffic operation state in the future period, such as the number of vehicles passing through per unit time on each road, the average vehicle speed, the vehicle density, the lane saturation, the intersection queue length, and the like in the future period, and the number and density of pedestrians passing through on each road, the flow direction and trend, and the like in the future period.
[0052] Exemplarily, the prediction traffic flow feature is extracted by combining the current image traffic data and the historical text traffic data, and is taken as an input of the multi-level structure network. In the processing of the prediction traffic flow feature by the multi-level structure network, the initial layer mainly focuses on the classification and aggregation of the original features, the intermediate layer mainly models the dynamic connection and interaction between different road segments, and the high layer is used to perform traffic decision analysis based on the information abstracted from the previous layer, and finally outputs the signal lamp optimization strategy with the optimal passing performance of each traffic signal lamp under the current traffic condition.
[0053] In step S104, the path optimization strategy and the signal lamp optimization strategy are taken as the traffic driving strategy of each vehicle in the current traffic environment, and the traffic driving strategy is used to provide the driving path suggestion and the driving speed suggestion for each vehicle in the current traffic environment.
[0054] The driving path suggestion represents the recommended information of the passing route of the specific vehicle generated based on the path optimization strategy, and is used to indicate the road combination that should be preferentially selected by the vehicle from the current position to the destination. The driving speed suggestion represents the reference information of the passing speed of the specific vehicle generated based on the path optimization strategy and the signal lamp optimization strategy under the given driving path suggestion, and is used to indicate the driving speed required by the vehicle from the current position to the destination, or from a certain intersection node to the next intersection node.
[0055] Exemplarily, the path optimization strategy and the signal lamp optimization strategy are integrated and applied as the traffic driving strategy of each vehicle in the current traffic environment. The traffic driving strategy contains two aspects of instruction contents. On the one hand, the global planning of the driving path of each vehicle is performed, so as to provide the driving path suggestion based on the current position and the destination of each vehicle, thereby improving the overall operation efficiency of the road network. On the other hand, the setting time of the signal lamp on each path segment is analyzed and controlled, and the global planning of the driving path of each vehicle is combined, so as to provide the driving speed suggestion required by each vehicle from the current position to the destination, or from a certain intersection node to the next intersection node, thereby reducing the number of signal lamp waiting and passing interruption.
[0056] In the traffic driving strategy optimization method based on unmanned aerial vehicle observation joint modeling, firstly, current image traffic data is collected through an image shooting device, and the current traffic environment is comprehensively perceived in combination with historical text traffic data, so that the subsequent optimization of the traffic driving strategy has real-time and data integrity; secondly, according to a target path optimization algorithm matched to the complexity of the current traffic environment, the current traffic flow characteristics obtained from the current image traffic data are analyzed to obtain a path optimization strategy, so that the driving path suggestion has individuality and global adaptability; thirdly, according to a multi-level network, the current traffic flow characteristics and the predicted traffic flow characteristics obtained in combination with the current image traffic data and the historical text traffic data are analyzed layer by layer to obtain a signal lamp optimization strategy, so as to realize fine adjustment of the setting time of the traffic signal lamp; based on this, the path optimization strategy and the signal lamp optimization strategy are used as the traffic driving strategy of each vehicle in the current traffic environment, so as to realize fusion control and collaborative optimization of path planning and signal lamp control, so as to effectively and accurately improve the overall operation efficiency of the road network according to an effective information sharing and collaborative optimization mechanism.
[0057] In one exemplary embodiment, as shown in Figure 2 The target path optimization algorithm matched to the complexity of the current traffic environment is determined, the current traffic flow characteristics obtained from the current image traffic data are analyzed based on the target path optimization algorithm, and a path optimization strategy is obtained, including steps S201 to S203.
[0058] In step S201, if the complexity of the current traffic environment meets a preset first determination condition, a path optimization algorithm based on a time sequence network is used as the target path optimization algorithm, and the time sequence network is a network combining deep reinforcement learning and time sequence modeling structure.
[0059] The first determination condition represents a logical rule for determining that the complexity of the current traffic environment is in a high complexity state, for example, the first determination condition can include that the traffic area involved in the current traffic environment exceeds a preset threshold, the number or density of vehicles exceeds a preset threshold, the vehicle queue length at multiple road intersections exceeds a warning level, the average speed of vehicles per unit time exceeds a certain standard, or the traffic flow direction appears frequent fluctuations, etc.
[0060] The path optimization algorithm based on the time sequence network represents a method for performing path planning in a complex traffic environment, which processes the current traffic flow characteristics evolving over time through a preset time sequence network, so as to identify the passing trend and behavior pattern of each vehicle or pedestrian in the current traffic environment, and output a path optimization strategy.
[0061] The time sequence network represents a calculation model capable of modeling traffic flow information in a time dimension and predicting vehicle behavior in a time dimension, which combines deep reinforcement learning and time sequence modeling structure.
[0062] In step S202, any vehicle in the current traffic environment is taken as a target vehicle, and the current traffic flow feature is processed for time sequence feature analysis based on the time sequence network, so as to predict the driving action of the target vehicle at each intersection node from the starting point to the ending point and the driving speed between each intersection node.
[0063] The driving action of each intersection node represents the recommended passing behavior of the vehicle when passing through a road intersection, such as straight, left turn, right turn, deceleration waiting, stay or detour, etc.; and the driving speed between each intersection node represents the recommended passing speed of the vehicle between two adjacent road intersections.
[0064] In step S203, the path optimization strategy is obtained based on the driving action of each vehicle in the current traffic environment at each intersection node from the corresponding starting point to the corresponding ending point and the driving speed between each intersection node.
[0065] Exemplarily, Figure 3 A structure diagram of a path optimization algorithm based on a time sequence network enhanced by Transformer relationship modeling is shown, and the implementation manner is as follows: any vehicle in the current traffic environment is taken as a target vehicle and the starting point and the ending point of the target vehicle are determined, the current traffic flow feature obtained based on the current image traffic data and the city traffic map is determined, the state of the target vehicle at the current i-th intersection node is determined , i.e., the time sequence feature data related to the target vehicle in the current traffic flow feature, and then the state is input into the time sequence network for processing.
[0066] Further, in the time sequence network, the local time sequence state relationship modeling of the state is performed by Conv Block1 (i.e., the first convolution block), and then the different states in a larger spatial range are subjected to first-stage larger spatial range nonlinear modeling by Transformer1 (i.e., the first time sequence modeling structure) based on the local time sequence state relationship modeling, so as to obtain sequence state features; further, in order to learn more effective time sequence relationship, the sequence state features are sequentially subjected to larger spatial range nonlinear modeling in the second stage by Conv Block2 (i.e., the second convolution block) and Transformer2 (i.e., the second time sequence modeling structure), so as to obtain the final enhanced sequence state features.
[0067] Further, the sequence state features are respectively input into the state evaluator, the action classifier, and the speed prediction predictor to obtain the recommended target vehicle driving action at the current i-th intersection node output by the action classifier , the recommended target vehicle driving speed from the current i-th intersection node to the next intersection node output by the speed prediction predictor , the reward obtained by the target vehicle at the i-th intersection node based on the driving action and the driving speed output by the action evaluator . The reward can be used to fine-tune the parameters of the time sequence network in the application process to adapt to environmental changes.
[0068] Further, in the current traffic environment, the next state obtained by the target vehicle based on the driving action and the driving speed is predicted to update the state of the target vehicle; and is input as the new state of the target vehicle and re-input into the time sequence network to re-predict the driving action and the driving speed of the target vehicle at the next intersection node, until the driving action at each intersection node and the driving speed between each intersection node of the target vehicle from the starting point to the ending point are obtained. Thus, the predicted driving action and the corresponding speed of each intersection node passed through by the target vehicle are integrated into the recommended driving path suggestion for the target vehicle to achieve path planning and speed suggestion based on the enhanced deep reinforcement learning of the Transformer time sequence modeling.
[0069] For example, the driving action at the i-th intersection node and the driving speed from the i-th intersection node to the next intersection node can be fed back to the driver as driving suggestion information through an in-vehicle terminal or a mobile terminal carried by the driver during vehicle driving to drive the driver to make corresponding driving behavior at the corresponding intersection node and the corresponding path segment according to the received driving suggestion information. Further, the fed-back driving speed is the average driving speed from the i-th intersection node to the next intersection node, that is, the actual vehicle speed during the driving of the vehicle on the path segment between the adjacent two intersection nodes can be maintained at the average driving speed, and if the front vehicle is slow to affect the vehicle speed, the vehicle speed can be restored to the average driving speed after accelerating to overtake the front vehicle, until the next intersection node is reached, so that the actual driving time on the path segment is not much different from the expected driving time.
[0070] On one hand, based on deep reinforcement learning, the agent can simulate the learning of the optimal strategy in the interaction process with the preset traffic environment, thereby maximizing the cumulative reward. On the other hand, based on the time sequence modeling tool Transformer, the dependence relationship in the long sequence data can be effectively captured, and the traffic data changing over time involved in the path planning has a significant advantage.
[0071] In the embodiment, firstly, in the case that the complexity of the current traffic environment meets the first determination condition, the path optimization algorithm based on the time sequence network is selected as the target path optimization algorithm, so that the target path optimization algorithm has the adaptability to the dynamic traffic state; secondly, the driving actions of each intersection node in the vehicle path and the driving speed between adjacent intersection nodes are predicted according to the time sequence network, so as to improve the response accuracy of the path decision to the actual traffic changes; thirdly, the global path optimization strategy is generated according to the node-level driving behavior of each vehicle, so as to realize the rationality of path allocation and the balance of traffic flow; based on this, dynamic path optimization under complex traffic environment can be realized, and the accuracy of path planning and the overall scheduling ability of the traffic network can be improved.
[0072] In one exemplary embodiment, as shown in Figure 4 if the complexity of the current traffic environment meets the preset first determination condition, the method further includes steps S301 to S304 before the path optimization algorithm based on the time sequence network is selected as the target path optimization algorithm.
[0073] Step S301, acquiring a time sequence network to be trained and a training set, constructing a target time sequence network matched with the network structure of the time sequence network to be trained, the training set including traffic flow characteristics of a road where a vehicle object is located, and the network parameters of the target time sequence network at each iteration round are obtained according to the momentum update of the network parameters of the time sequence network to be trained at the corresponding iteration round.
[0074] The training set represents a data set for training the time sequence network model, and the data set contains feature samples in multiple traffic environments, that is, the corresponding traffic data and traffic flow characteristics of different vehicle objects at different times and different road positions.
[0075] The target time sequence network represents an auxiliary network with the same structure as the time sequence network to be trained but with independent parameter update, which is used to improve the training stability and optimization convergence speed.
[0076] In step S302, in the current iteration round, the traffic flow features in the training set are processed based on the to-be-trained time sequence network to predict the predicted driving behavior information of the vehicle object at the current intersection node, and the predicted driving behavior information includes the driving action, the driving speed, and the reward of the vehicle object at the current intersection node.
[0077] The predicted driving behavior information represents the behavior output result of the vehicle object predicted by the to-be-trained time sequence network based on the traffic flow features at the current intersection node, and includes the driving action at the current intersection node, the driving speed from the current intersection node to the next intersection node, and the reward based on the driving action and the driving speed.
[0078] In step S303, the updated traffic flow features are obtained based on the predicted driving behavior information, and the updated traffic flow features are processed based on the target time sequence network to predict the target driving behavior information of the vehicle object at the next intersection node, and the target driving behavior information includes the driving action, the driving speed, and the reward of the vehicle object at the next intersection node.
[0079] The target driving behavior information represents the behavior output result of the vehicle object predicted by the target time sequence network based on the updated traffic flow features at the next intersection node, and includes the driving action at the next intersection node, the driving speed from the next intersection node to the next intersection, and the reward based on the driving action and the driving speed.
[0080] In step S304, based on the difference between the predicted driving behavior information and the target driving behavior information, and in combination with the maximization of the cumulative reward of the vehicle object from the starting point to the ending point, the network parameters of the to-be-trained time sequence network are updated in the current iteration round until the trained time sequence network is obtained by traversing each iteration round.
[0081] Exemplarily, Figure 5 A structural schematic diagram of training of a time sequence network based on Transformer relationship modeling enhancement is shown, and the implementation manner is as follows: based on the traffic flow features corresponding to any vehicle object in the training set, the state of the vehicle object at the i-th intersection node is determined The state is input into the to-be-trained time sequence network for processing.
[0082] Further, in the to-be-trained time sequence network, the input state is sequentially processed through Conv Block1, Transformer1, Conv Block2, and Transformer2, and the processing manner of each network layer can refer to Figure 3The temporal network shown is used to obtain the final sequence state features.
[0083] Furthermore, the sequence state features are input into the state evaluator, action classifier, and pace predictor, respectively, to obtain predicted driving behavior information. This predicted driving behavior information includes the driving actions of the recommended vehicle at the current i-th intersection node, as output by the action classifier. The recommended driving speed of the vehicle object output by the pace predictor from the current i-th intersection node to the next intersection node. The vehicle object output by the motion evaluator at the i-th intersection node is based on its driving action. With driving speed The reward received .
[0084] Furthermore, within the pre-defined traffic environment in which the vehicle object is located, the prediction of the vehicle object's driving actions is performed. With driving speed The next state obtained To determine the driving action of the vehicle object at the current i-th intersection node. Driving speed Traffic signal light configuration duration ,award Next state Forming a six-tuple sample The samples are then added to the reward pool, allowing for the random selection of several six-tuple samples from the reward pool during training to iterate the network parameters of the temporal network.
[0085] Furthermore, the next state The new state of the vehicle object is then re-input into the temporal network to be trained, thereby re-predicting the vehicle object's driving action and speed at the next intersection node, resulting in a new six-tuple sample.
[0086] Furthermore, the next state The data is synchronously input into the target temporal network. Based on the network structure of the target temporal network, which is consistent with the temporal network, the target temporal network determines the next state. The target driving behavior information is obtained by analyzing the data using a processing method consistent with that of temporal networks. This information includes the target driving action of the recommended vehicle at the next intersection node, output by the action classifier; the target driving speed of the recommended vehicle from the next intersection node to the next intersection node after that, output by the pace predictor; and the target reward obtained by the vehicle at the next intersection node based on the target driving action and target driving speed, output by the action evaluator.
[0087] Further, based on minimizing the difference Loss between the predicted driving behavior information and the target driving behavior information, and combining the maximization of the cumulative reward of the vehicle object from the starting point to the end point, the network parameters of the state evaluator, the action classifier and the speed prediction in the time sequence network are iteratively updated until the trained time sequence network is obtained by traversing each iteration round.
[0088] In each iteration round, the network parameters of the target time sequence network can be updated based on the network parameters of the time sequence network in the current iteration through momentum updating, that is, the network parameter updating manner of the target time sequence network can refer to the following formula:
[0089] (1)
[0090] wherein, represents the network parameters of the time sequence network in the current iteration, represents the network parameters of the target time sequence network in the current iteration, represents the network parameters of the target time sequence network in the next iteration, and mu represents the momentum coefficient. As the time sequence network is continuously iteratively trained, the value of the momentum coefficient mu gradually increases, for example, mu can gradually increase in the corresponding numerical set.
[0091] Further, in the initial stage of training of the time sequence network, because the network parameters are randomly initialized, the driving action and the driving speed obtained according to the current state are not accurate, and therefore, as shown in Figure 5 , random driving actions a and random driving speeds v can be randomly generated and added to the experience pool to quickly explore the feasible path and speed through the random driving actions a and the random driving speeds v; as the number of training increases, the random driving actions a and the random driving speeds v in the experience pool are gradually reduced, so that the driving actions and the driving speeds predicted by the time sequence network are used more in the later training.
[0092] In the embodiment, first, the target time sequence network matched with the network structure to be trained is constructed, and the momentum updating mechanism is introduced, so as to enhance the parameter stability and the reliability of model convergence in the training process; further, the traffic flow characteristics are analyzed in time sequence by the time sequence network to be trained, and the predicted driving behavior information is output, and then the updated traffic flow characteristics are analyzed in time sequence by the target time sequence network, and the target driving behavior information is output, based on the difference between the predicted driving behavior information and the target driving behavior information, so as to optimize the network parameters of the time sequence network in the dimension of analyzing the time sequence evolution relationship of the vehicle driving behavior between the continuous nodes, and in the dimension of maximizing the cumulative reward, so as to improve the global path quality evaluation and learning ability of the time sequence network, so as to realize the high-performance time sequence network training method with behavior prediction ability and path optimization ability in the dynamic traffic environment.
[0093] In one exemplary embodiment, as shown in Figure 6 the target path optimization algorithm matching the complexity of the current traffic environment is determined, and based on the target path optimization algorithm, the current traffic flow characteristics obtained from the current image traffic data are analyzed and processed to obtain a path optimization strategy, including steps S401 to S404.
[0094] Step S401, if the complexity of the current traffic environment meets the preset second determination condition, the path optimization algorithm based on the ant algorithm is used as the target path optimization algorithm.
[0095] The second determination condition represents a logical rule for determining that the complexity of the current traffic environment is in a low complexity state, for example, the second determination condition can include that the traffic area involved in the current traffic environment is less than a preset threshold, the number or density of vehicles is less than a preset threshold, the vehicle queue length at multiple road intersections is less than a warning level, the average speed of vehicles per unit time is less than a certain standard, or the traffic flow direction is relatively smooth.
[0096] The improved ant algorithm based on the feedback of the unmanned aerial vehicle represents an improved path optimization algorithm combining the ant algorithm inspired by the foraging behavior of ants in nature and the real-time data collected by the unmanned aerial vehicle, that is, in traffic path planning, multiple virtual ants (i.e. simulated vehicles) are simulated to select paths on the traffic network in combination with real-time traffic data obtained by the unmanned aerial vehicle.
[0097] Step S402, any vehicle in the current traffic environment is taken as a target vehicle, and the current traffic flow characteristics and pheromones between the target vehicle and each adjacent intersection node in the current intersection node are analyzed and processed based on the improved ant algorithm based on the feedback of the unmanned aerial vehicle, and the transition probability between the current intersection node and each adjacent intersection node is calculated.
[0098] The pheromones between each adjacent intersection node represent an experiential index value on the path segment between two adjacent intersection nodes for guiding path selection decisions, that is, the pheromones reflect the use frequency, passing efficiency or success rate of the path segment in historical path selection, and the higher the value, the more virtual ants choose the path segment in the past, and the higher the passing priority.
[0099] The transition probability between the current intersection node and each adjacent intersection node represents the selection probability of the simulated vehicle turning to each adjacent intersection node from the current intersection node when making a path selection decision, which is used to quantify the passing priority of the current intersection node turning to a certain adjacent intersection node.
[0100] Step S403, the adjacent intersection node corresponding to the maximum transition probability is taken as the next intersection node corresponding to the current intersection node, and the transition probability calculation is repeated for the next intersection node until the intersection nodes between the target vehicle from the current intersection node to the terminal point are determined.
[0101] Step S404, the path optimization strategy is obtained based on the intersection nodes between the respective starting points and the respective terminals of the respective vehicles in the current traffic environment.
[0102] Exemplarily, Figure 7 A flowchart of a path optimization algorithm based on an improved ant algorithm combined with feedback of a UAV is shown, and the implementation is as follows: first, the pheromone content on the corresponding path segment between each adjacent intersection node is initialized, in the current iteration round N, the path walked by each virtual ant (i.e. the target vehicle) from the starting point is recorded, and whether each virtual ant reaches the destination is determined according to the recorded path; if a virtual ant does not reach the destination, the current traffic flow characteristics and pheromone between the current intersection node and each adjacent intersection node are analyzed and processed, so as to calculate the transition probability between the current intersection node and each adjacent intersection node.
[0103] Further, the adjacent intersection node corresponding to the maximum transition probability is taken as the next intersection node for the virtual ant to turn from the current intersection node, so as to realize the path planning calculation of the virtual ant in the current iteration round N; if the current iteration round does not reach the maximum iteration number, the pheromone on each path segment is updated based on the movement of each virtual ant in the current iteration round N; if the current iteration round reaches the maximum iteration number, the path planning calculation of each virtual ant is ended.
[0104] In the next iteration round N+1, the path walked by each virtual ant from the starting point is recorded, if it is determined that a virtual ant reaches the destination according to the path walked by the virtual ant, the path planning calculation of the virtual ant is ended, and the optimal path output by the virtual ant in the iteration calculation of the ant algorithm is obtained.
[0105] If it is determined that a virtual ant does not reach the destination according to the path walked by the virtual ant, the path planning calculation of the virtual ant is restarted in the next iteration round N+1 until the virtual ant reaches the destination or the iteration round reaches the maximum iteration number.
[0106] Based on this, the intersection nodes of the recommended route of the target vehicle from the starting point to the terminal point are obtained, and are integrated into the driving path recommendation recommended for the target vehicle, thereby realizing the improved ant colony algorithm for local optimal path rapid planning.
[0107] In this embodiment, first, when the complexity of the current traffic environment meets the second determination condition, an improved ant algorithm based on unmanned aerial vehicle joint feedback is selected as the path optimization algorithm, so that the path planning has the ability of distributed search and adaptation to complex traffic structure; secondly, the transition probability is calculated according to the pheromone and the current traffic flow characteristics between the current intersection node and the adjacent intersection node, so as to realize the joint evaluation of historical experience and real-time flow in path selection; thirdly, the intersection node sequence corresponding to the passing path is gradually constructed according to the maximum transition probability, and the path optimization strategy is generated according to the intersection node sequence of all vehicles, so as to realize the reasonable distribution and global coordination among multiple vehicle paths; based on this, dynamic path optimization under simple traffic environment can be realized, and the accuracy of path planning and the overall scheduling ability of traffic network can be improved.
[0108] In one exemplary embodiment, as shown in Figure 8 According to the improved ant algorithm based on unmanned aerial vehicle joint feedback, the current traffic flow characteristics and pheromone of the unmanned aerial vehicle real-time feedback between the target vehicle at the current intersection node and each adjacent intersection node are analyzed and processed, and the transition probability between the current intersection node and each adjacent intersection node is calculated, including steps S501 to S503.
[0109] Step S501, the pheromone of each path segment is calculated by combining the pheromone retention amount, the pheromone new amount and the global pheromone compensation amount of the path segment between the current intersection node and each adjacent intersection node.
[0110] Among them, the pheromone retention amount represents the experiential identification value retained by the specific path segment after being selected by multiple virtual ants combined with the pheromone evaporation rate; the pheromone new amount represents the experiential identification value newly added to the specific path segment after being selected by multiple virtual ants within a preset period; the global pheromone compensation amount represents the pheromone enhancement value actively applied to some path segments that are not frequently selected but have reasonable structure or high passing potential.
[0111] Step S502, the heuristic factor corresponding to each path segment is calculated by combining the path segment length and the traffic congestion degree of each path segment in the current traffic flow characteristics.
[0112] Among them, the path segment length represents the actual physical length of the path segment between two adjacent intersection nodes; the traffic congestion degree represents the traffic flow saturation on the path segment between two adjacent intersection nodes, which is determined by factors such as the number and density of vehicles, the average speed change rate, the length of vehicle queue, etc.
[0113] The heuristic factor represents an evaluation value formed by comprehensively combining the current traffic flow characteristics such as path segment length and traffic congestion degree, and is used to guide the real-time behavior preference in the path selection process.
[0114] In step S503, the transition probability between the current intersection node and each adjacent intersection node is calculated by combining the pheromone corresponding to each path segment and the heuristic factor.
[0115] For example, the pheromone concentration update formula on each path segment can refer to the following formula:
[0116] (2)
[0117] wherein, represents the evaporation rate of the pheromone, represents the retention rate of the pheromone, wherein ; represents the pheromone concentration of the path segment between the intersection node i and the intersection node j at time t+1 obtained after updating at time t.
[0118] wherein, represents the original pheromone corresponding to the path segment between the intersection node i and the intersection node j at time t, represents the pheromone retention amount corresponding to the path segment between the intersection node i and the intersection node j at time t.
[0119] wherein, represents the newly added amount of pheromone released by all virtual ants on the path segment between the intersection node i and the intersection node j from time t to time t+1.
[0120] wherein, represents the global compensation factor, represents the pheromone retention amount corresponding to all path segments at time t combined with the global compensation factor to calculate the global pheromone compensation amount.
[0121] In formula (2), the newly added amount of pheromone is calculated as follows:
[0122] (3)
[0123] wherein, represents the newly added amount of pheromone released by the mth virtual ant on the path segment between the intersection node i and the intersection node j from time t to time t+1; represents the newly added amount of pheromone released by all virtual ants on the path segment between the intersection node i and the intersection node j from time t to time t+1.
[0124] In formula (3), the newly added amount of pheromone released by the mth virtual ant on the path segment between intersection node i and intersection node j from time t to time t+1 The calculation method is as follows:
[0125] (4)
[0126] Wherein, represents the total amount of pheromone carried by the mth virtual ant, represents the path segment length corresponding to the path segment between intersection node i and intersection node j.
[0127] In formula (2), the calculation method of the remaining amount of pheromone corresponding to all path segments at time t is as follows:
[0128] (5)
[0129] Wherein, represents the original pheromone corresponding to the path segment between each pair of adjacent intersection nodes at time t, until the global pheromone corresponding to all path segments at time t is obtained.
[0130] Exemplarily, the calculation method of the transition probability is as follows:
[0131] (6)
[0132] Wherein, represents the transition probability of the mth virtual ant from intersection node i to intersection node j; α and β are weight coefficients of pheromone concentration and heuristic factor respectively; represents the pheromone concentration corresponding to the path segment between intersection node i and intersection node j at time t combined with α, which can be the updated pheromone concentration based on formula (2); represents the heuristic factor corresponding to the path segment between intersection node i and intersection node j at time t combined with β.
[0133] Wherein, represents the path segment attraction corresponding to the path segment between intersection node i and intersection node j; represents the total path segment attraction accumulated by all path segments in all adjacent intersection nodes j corresponding to intersection node i.
[0134] Wherein, denotes a set of intersection nodes which have not been passed by the m-th virtual ant; if the intersection node j is an intersection node which has not been passed by the virtual ant, a transition probability of turning to the intersection node j is calculated based on the pheromone concentration and the heuristic factor; if the intersection node j is an intersection node which has been passed by the virtual ant, the transition probability of turning to the intersection node j is 0.
[0135] In formula (6), the heuristic factor is calculated in the following manner:
[0136] (7)
[0137] wherein, denotes a path segment length corresponding to a path segment between the intersection node i and the intersection node j, denotes a traffic congestion degree of the intersection node j.
[0138] In the embodiment, firstly, the pheromone concentration of each path segment is adaptively and effectively updated in multiple dimensions based on the pheromone reservation amount, the pheromone new addition amount and the global pheromone compensation amount of each path segment; secondly, the heuristic factor is calculated according to the path segment length and the traffic congestion degree, so as to guide the path selection to be more in line with the traffic efficiency demand of the current traffic environment; thirdly, the transition probability is calculated according to the combination result of the pheromone and the heuristic factor, so as to realize the double consideration of historical experience and real-time state in the path selection; based on this, a path segment attraction modeling method which takes into account global experience accumulation and current traffic state perception can be realized, thereby providing stable and time-effective transition decision basis for subsequent path planning.
[0139] In one exemplary embodiment, the multi-level structure network comprises a long-short time sequence modeling layer, a nonlinear time sequence modeling layer and a regression analysis layer; as shown in Figure 9 based on the preset multi-level structure network, the current traffic flow characteristics and the predicted traffic flow characteristics obtained by combining the current image traffic data and the historical text traffic data are processed layer by layer to obtain the signal light optimization strategy, including steps S601 to S603.
[0140] The long-short time sequence modeling layer represents a modeling structure capable of capturing short-term fluctuations and long-term trends in traffic flow characteristics at the same time, so as to analyze the evolution law of the input characteristics in the time sequence dimension; the nonlinear time sequence modeling layer represents a modeling structure for processing complex and nonlinear relationships between long-short time sequence characteristics, so as to mine the nonlinear evolution law of the input characteristics; the regression analysis layer represents a calculation structure for mapping the nonlinear time sequence characteristics to specific numerical values, so as to establish a functional relationship between the input characteristics and the signal light setting time.
[0141] In step S601, the current traffic flow feature and the predicted traffic flow feature are respectively subjected to data analysis and processing based on the long-short time sequence modeling layer to obtain a first long-short time sequence feature corresponding to the current traffic flow feature and a second long-short time sequence feature corresponding to the predicted traffic flow feature.
[0142] The first long-short time sequence feature represents a time sequence feature extracted from the current traffic flow feature after processing by the long-short time sequence modeling layer, and is used to describe the structural change of the current traffic flow feature in the time dimension, such as the growth rate of the traffic density of the current period, the trend of the traffic direction, the fluctuation trend of the traffic speed, etc.
[0143] The second long-short time sequence feature represents a time sequence feature extracted from the predicted traffic flow feature after processing by the long-short time sequence modeling layer, and is used to describe the structural change of the predicted traffic flow feature in the time dimension, such as the growth rate of the traffic density of the future period, the trend of the traffic direction, the fluctuation trend of the traffic speed, etc.
[0144] In step S602, the first long-short time sequence feature and the second long-short time sequence feature are subjected to data analysis and processing based on the nonlinear time sequence modeling layer to obtain a nonlinear modeling enhanced time sequence feature.
[0145] The nonlinear modeling enhanced time sequence feature represents a deep-level expression result generated after the first long-short time sequence feature and the second long-short time sequence feature are processed by the nonlinear time sequence modeling layer, and is used to describe the high-dimensional feature structure of the nonlinear dependence, cross-influence and time relationship existing in the traffic flow evolution process.
[0146] In step S603, the nonlinear modeling enhanced time sequence feature is subjected to data analysis and processing based on the regression analysis layer to obtain the signal light setting time of each intersection node in different directions in the current traffic environment, as the signal light optimization strategy.
[0147] Exemplarily, Figure 10 A structural schematic diagram of a network generating a signal light optimization strategy based on a multi-level structure is shown, and the implementation manner is as follows: in a preset traffic environment, current image traffic data of a current period and historical text traffic data of a historical period are obtained, the current image traffic data is processed based on a first feature extraction network to obtain a current traffic flow feature , and the current image traffic data and the historical text traffic data are processed based on a second feature extraction network to obtain a predicted traffic flow feature .
[0148] Further, the current traffic flow feature and the predicted traffic flow feature The current traffic flow features are input into the network based on the reinforcement learning architecture, and the LSTM1 (i.e., the first long short-term modeling layer) performs long short-term modeling operation on the current traffic flow features to obtain the first long short-term features The LSTM2 (i.e., the second long short-term modeling layer) performs long short-term modeling operation on the predicted traffic flow features to obtain the second long short-term features, so as to sufficiently learn the temporal context information.
[0149] Furthermore, the Transformer (i.e., the nonlinear time series modeling layer) performs nonlinear time series modeling operation on the first long short-term features and the second long short-term features to obtain the nonlinear modeling enhanced time series features; and the FC (i.e., the regression analysis layer) performs regression operation on the nonlinear modeling enhanced time series features to obtain the traffic signal light setting time of a certain intersection node in four inlet direction.
[0150] Furthermore, the obtained traffic signal light setting time is taken as feedback information and acts on the preset traffic environment to obtain updated current traffic flow features and predicted traffic flow features.
[0151] The calculation method of the traffic signal light setting time of a certain intersection node in four directions is as follows:
[0152] (8)
[0153] wherein, represents the long short-term modeling operation, represents the nonlinear time series modeling operation, represents the regression operation; the symbol represents the element-wise multiplication; represents the setting time of the traffic signal light j corresponding to the left turn of the intersection from the i-th inlet direction of the specified intersection node, represents the setting time of the traffic signal light j corresponding to the straight driving of the intersection from the i-th inlet direction of the specified intersection node, represents the setting time of the traffic signal light j corresponding to the right turn of the intersection from the i-th inlet direction of the specified intersection node.
[0154] wherein, i can take values 1, 2, 3 or 4 to represent one of the four inlet directions of the specified intersection node; and the value range of i can be adaptively determined according to the actual number of inlets of the intersection node and the actual configuration of the traffic signal lights corresponding to each inlet.
[0155] wherein, j can take values 1 or 2, 1 representing red light and 2 representing green light; and the value range of j can be adaptively determined according to the signal type of the traffic signal light.
[0156] In this embodiment, firstly, the current traffic flow feature and the predicted traffic flow feature are respectively extracted through the long-short time sequence modeling layer to obtain the first long-short time sequence feature and the second long-short time sequence feature, so as to realize modeling of the continuous change of the traffic flow feature in different time scales; secondly, the first long-short time sequence feature and the second long-short time sequence feature are extracted through the nonlinear time sequence modeling layer to obtain the nonlinear modeling enhanced time sequence feature, so as to enhance the recognition ability of the complex traffic evolution rule; thirdly, the nonlinear modeling enhanced time sequence feature is output through the regression analysis layer to obtain the signal lamp setting time length of each entrance direction of the specified intersection node, so as to realize the structured representation of the signal lamp setting time length adjustment result; on this basis, the time sequence rule of the dynamic traffic flow can be comprehensively modeled, and the refined signal lamp optimization strategy can be generated accordingly, so as to improve the adaptability and scheduling accuracy of the signal lamp control.
[0157] In one exemplary embodiment, as shown in FIG. 7, after the path optimization strategy and the signal lamp optimization strategy are used as the traffic driving strategy of each vehicle in the current traffic environment, the method further includes steps S701 to S703. Figure 11
[0158] Step S701, after each vehicle in the current traffic environment drives according to the traffic driving strategy, target traffic data collected in a target period is obtained.
[0159] The target traffic data represents the traffic data reflecting the traffic running state of the target period collected by the image shooting device in the target period after the traffic driving strategy is implemented; the target period can represent a time window set in advance for implementing the traffic driving strategy and monitoring the effect thereof.
[0160] Step S702, based on the target traffic data, the position information of each vehicle is obtained, and if a vehicle that has not arrived at the destination is detected based on the position information of each vehicle, the starting point and the driving timing point corresponding to the vehicle that has not arrived at the destination are updated to obtain the updated starting point and the updated driving timing point corresponding to the vehicle that has not arrived at the destination.
[0161] The driving timing point represents a time reference point at which a certain vehicle starts or restarts path planning calculation, for example, the vehicle takes its current position and the current time as the starting point and the driving timing point when the path is recalculated, as the starting condition of the passing path.
[0162] Step S703, according to the updated starting point and the updated driving timing point corresponding to the vehicle that has not arrived at the destination, the path optimization strategy calculation is performed again for the vehicle that has not arrived at the destination until each vehicle in the current traffic environment arrives at the corresponding destination.
[0163] Exemplarily, the path optimization strategy and the signal lamp optimization strategy can be used as the traffic driving strategy of each vehicle in the current traffic environment. Figure 12 A flowchart based on a real-time dynamic path induction strategy algorithm is shown for performing the steps in the above embodiments: after each vehicle in the current traffic environment travels according to the traffic travel strategy, target traffic data collected in a target time period is obtained, and the current position of each vehicle is obtained based on the target traffic data. Further, whether each vehicle reaches the destination is determined according to the current position of each vehicle.
[0164] Further, if a vehicle reaches the destination, the execution of the traffic travel strategy for the vehicle is ended. Further, if a vehicle does not reach the destination, the current traveled path of the vehicle and the current time are obtained; the starting point of the vehicle is updated according to the current traveled path to obtain an updated starting point, and the travel timing point is updated according to the current time to obtain an updated travel timing point; a new traffic travel strategy for the vehicle is recalculated based on the updated starting point and the updated travel timing point, until the vehicle reaches the destination.
[0165] Based on this, by gradually reducing the prediction step in the traffic travel strategy calculation process, the changes in the traffic conditions can be captured more finely in a shorter time interval, so that the accuracy of the predicted vehicle travel path gradually increases.
[0166] In this embodiment, first, the actual passing trajectory of each vehicle is obtained according to the target traffic data collected in the target time period, so as to realize dynamic monitoring of the execution effect of the traffic travel strategy; further, the starting point and the travel timing point of the vehicle that does not reach the destination are updated according to the position information of the vehicle, so as to ensure that the path optimization calculation can be performed in real time based on the current position and the current time of the vehicle; further, the traffic travel strategy is recalculated according to the updated starting point and the updated travel timing point, until all vehicles complete passing, so as to realize continuous closed-loop adjustment of the path execution process; based on this, an adaptive path optimization control mechanism for dynamic road conditions can be realized, and the completion rate of the vehicle passing task and the real-time performance of the system scheduling are improved.
[0167] It should be understood that although each step in the flowchart involved in each of the above embodiments is shown in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0168] Based on the same inventive concept, the embodiments of the present application also provide a device for implementing the traffic driving strategy optimization method based on the unmanned aerial vehicle observation joint modeling. The device provides a solution similar to the solution described in the above method, and therefore the specific limitations in one or more device embodiments based on the unmanned aerial vehicle observation joint modeling of the traffic driving strategy optimization device can be referred to the limitations of the traffic driving strategy optimization method based on the unmanned aerial vehicle observation joint modeling described above, which will not be repeated here.
[0169] In one exemplary embodiment, as shown in Figure 13 a device for traffic driving strategy optimization based on unmanned aerial vehicle observation joint modeling is provided, comprising: an acquisition module 101, a first optimization module 102, a second optimization module 103 and an execution module 104, wherein:
[0170] The acquisition module 101 is configured to acquire current image traffic data collected by a preset image shooting device in a current time period and historical text traffic data stored in a historical time period in a current traffic environment;
[0171] The first optimization module 102 is configured to determine a target path optimization algorithm matched with the complexity of the current traffic environment, perform data analysis and processing on current traffic flow characteristics obtained from the current image traffic data based on the target path optimization algorithm, and obtain a path optimization strategy, wherein the complexity of the current traffic environment is determined according to the current image traffic data collected by the image shooting device;
[0172] The second optimization module 103 is configured to perform layer-by-layer data analysis and processing on the current traffic flow characteristics and the predicted traffic flow characteristics obtained by combining the current image traffic data and the historical text traffic data based on a preset multi-level structure network, and obtain a signal light optimization strategy;
[0173] The execution module 104 is configured to use the path optimization strategy and the signal light optimization strategy as traffic driving strategies of each vehicle in the current traffic environment, and the traffic driving strategies are used to provide driving path suggestions and driving speed suggestions for each vehicle in the current traffic environment.
[0174] In an example embodiment, the first optimization module 102 further comprises a first optimization unit configured to: if the complexity of the current traffic environment satisfies a preset first determination condition, set a path optimization algorithm based on a time sequence network as a target path optimization algorithm, the time sequence network being a network combining deep reinforcement learning and a time sequence modeling structure; set any vehicle in the current traffic environment as a target vehicle, perform time sequence feature analysis and processing on the current traffic flow feature based on the time sequence network, and predict the driving actions of the target vehicle at each intersection node between the start point and the end point and the driving speed between each intersection node; and obtain a path optimization strategy based on the driving actions of each vehicle in the current traffic environment at each intersection node between the respective start point and the respective end point and the driving speed between each intersection node.
[0175] In an example embodiment, the first optimization unit further comprises a training unit configured to: obtain a time sequence network to be trained and a training set, the training set comprising traffic flow features of a road on which a vehicle object is located, construct a target time sequence network matching the network structure of the time sequence network to be trained, the network parameters of the target time sequence network at each iteration round being obtained according to momentum updating of the network parameters of the time sequence network to be trained at the corresponding iteration round; in a current iteration round, perform time sequence feature analysis and processing on the traffic flow features in the training set based on the time sequence network to be trained, and predict the predicted driving behavior information of the vehicle object at a current intersection node, the predicted driving behavior information comprising the driving action, the driving speed and the reward of the vehicle object at the current intersection node; obtain updated traffic flow features based on the predicted driving behavior information, perform time sequence feature analysis and processing on the updated traffic flow features based on the target time sequence network, and predict the target driving behavior information of the vehicle object at a next intersection node, the target driving behavior information comprising the driving action, the driving speed and the reward of the vehicle object at the next intersection node; and based on the difference between the predicted driving behavior information and the target driving behavior information, and in combination with maximizing the cumulative reward of the vehicle object between the start point and the end point, update the network parameters of the time sequence network to be trained in the current iteration round until the trained time sequence network is obtained by traversing each iteration round.
[0176] In an example embodiment, the first optimization module 102 further comprises a second optimization unit configured to: if the complexity of the current traffic environment satisfies a preset second determination condition, take the path optimization algorithm based on the improved ant algorithm combined with the feedback of the UAV as a target path optimization algorithm; take any vehicle in the current traffic environment as a target vehicle, and according to the improved ant algorithm combined with the feedback of the UAV, perform data analysis and processing on the current traffic flow characteristics and pheromones in real time fed back by the UAV between the target vehicle and each adjacent intersection node respectively, to calculate the transition probability between the current intersection node and each adjacent intersection node; take the adjacent intersection node corresponding to the maximum transition probability as the next intersection node corresponding to the current intersection node, and repeat the transition probability calculation for the next intersection node until each intersection node between the target vehicle from the current intersection node to the terminal point is determined; and based on each intersection node between each vehicle in the current traffic environment from the corresponding starting point to the corresponding terminal point, obtain a path optimization strategy.
[0177] In an example embodiment, the second optimization unit is further configured to: combine the pheromone retention amount, the pheromone addition amount, and the global pheromone compensation amount of each path segment between the current intersection node and each adjacent intersection node, to calculate the pheromone corresponding to each path segment; combine the path segment length and the traffic congestion degree of each path segment in the current traffic flow characteristics, to calculate the heuristic factor corresponding to each path segment; and combine the pheromone and the heuristic factor corresponding to each path segment, to calculate the transition probability between the current intersection node and each adjacent intersection node.
[0178] In an example embodiment, the second optimization module 103 is further configured to: perform data analysis and processing on the current traffic flow characteristics and the predicted traffic flow characteristics based on the long-short time sequence modeling layer, to obtain a first long-short time sequence feature corresponding to the current traffic flow characteristics and a second long-short time sequence feature corresponding to the predicted traffic flow characteristics; perform data analysis and processing on the first long-short time sequence feature and the second long-short time sequence feature based on the nonlinear time sequence modeling layer, to obtain a nonlinear modeling enhanced time sequence feature; and perform data analysis and processing on the nonlinear modeling enhanced time sequence feature based on the regression analysis layer, to obtain the signal light setting time of each intersection node in the current traffic environment in different directions, as a signal light optimization strategy.
[0179] In an example embodiment, the execution module 104 is further configured to: after each vehicle in the current traffic environment travels according to the traffic travel strategy, acquire target traffic data collected in a target time period; obtain position information of each vehicle based on the target traffic data, and if a vehicle that has not arrived at a destination is detected based on the position information of each vehicle, update a starting point and a travel timing point corresponding to the vehicle that has not arrived at the destination to obtain an updated starting point and an updated travel timing point corresponding to the vehicle that has not arrived at the destination; and perform path optimization strategy calculation again for the vehicle that has not arrived at the destination according to the updated starting point and the updated travel timing point corresponding to the vehicle that has not arrived at the destination until each vehicle in the current traffic environment arrives at a corresponding destination.
[0180] The modules in the traffic travel strategy optimization device based on unmanned aerial vehicle observation joint modeling described above can be realized by software, hardware, and combinations thereof, in whole or in part. The modules described above can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.
[0181] In an example embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in any of the above embodiments when executing the computer program.
[0182] In an example embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in any of the above embodiments when executing the computer program.
[0183] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0184] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0185] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A traffic driving strategy optimization method based on joint modeling of UAV observations, characterized in that, The method includes: In the current traffic environment, current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in historical time periods are acquired. The image capturing device consists of multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment. A target path optimization algorithm matching the complexity of the current traffic environment is determined. Based on the target path optimization algorithm, the current traffic flow characteristics obtained from the current image traffic data are analyzed and processed to obtain a path optimization strategy. The complexity of the current traffic environment is determined based on the current image traffic data. Based on a pre-defined multi-layered network, the current traffic flow characteristics and the predicted traffic flow characteristics obtained by combining the current image traffic data and the historical text traffic data are analyzed layer by layer to obtain a traffic light optimization strategy. The route optimization strategy and the traffic light optimization strategy are used as traffic driving strategies for each vehicle in the current traffic environment. The traffic driving strategies are used to provide driving route suggestions and driving speed suggestions for each vehicle in the current traffic environment. The step of determining a target path optimization algorithm that matches the complexity of the current traffic environment involves performing data analysis and processing on the current traffic flow characteristics obtained from the current image traffic data based on the target path optimization algorithm to obtain a path optimization strategy, including: If the complexity of the current traffic environment meets a preset first judgment condition, then the path optimization algorithm based on a temporal network is used as the target path optimization algorithm. The temporal network is a network that combines deep reinforcement learning and temporal modeling structure. Any vehicle in the current traffic environment is taken as the target vehicle. Based on the temporal network, the current traffic flow characteristics are analyzed and processed to predict the driving actions of the target vehicle at each intersection node from the starting point to the ending point and the driving speed between each intersection node. Based on the driving actions of each vehicle in the current traffic environment at each intersection node from its corresponding starting point to its corresponding ending point and the driving speed between each intersection node, a path optimization strategy is obtained. If the complexity of the current traffic environment meets the preset second judgment condition, then the path optimization algorithm based on the improved ant colony algorithm with UAV joint feedback is used as the target path optimization algorithm; any vehicle in the current traffic environment is taken as the target vehicle, and according to the improved ant colony algorithm based on UAV joint feedback, the current traffic flow characteristics and pheromones fed back by the UAV in real time between the target vehicle at the current intersection node and each adjacent intersection node are analyzed and processed to calculate the transition probability between the current intersection node and each adjacent intersection node; the adjacent intersection node corresponding to the maximum transition probability is taken as the next intersection node corresponding to the current intersection node, and the transition probability calculation is repeated for the next intersection node until the intersection nodes between the target vehicle and the destination are determined; based on the intersection nodes between each vehicle in the current traffic environment from the corresponding starting point to the corresponding destination, the path optimization strategy is obtained.
2. The method according to claim 1, characterized in that, Before using the time-series network-based path optimization algorithm as the target path optimization algorithm if the complexity of the current traffic environment meets a preset first judgment condition, the method further includes: Obtain the temporal network to be trained and the training set, and construct a target temporal network that matches the network structure of the temporal network to be trained. The training set includes traffic flow characteristics of the road where the vehicle object is located. The network parameters of the target temporal network in each iteration are obtained by updating the momentum of the network parameters of the temporal network to be trained in the corresponding iteration. In the current iteration, the traffic flow features in the training set are analyzed and processed based on the temporal network to be trained, and the predicted driving behavior information of the vehicle object at the current intersection node is predicted. The predicted driving behavior information includes the driving action, driving speed and reward of the vehicle object at the current intersection node. Based on the predicted driving behavior information, the updated traffic flow characteristics are obtained. Based on the target temporal network, the updated traffic flow characteristics are subjected to temporal feature analysis and processing to predict the target driving behavior information of the vehicle object at the next intersection node. The target driving behavior information includes the driving action, driving speed and reward of the vehicle object at the next intersection node. Based on the difference between the predicted driving behavior information and the target driving behavior information, and by maximizing the cumulative reward of the vehicle object from the starting point to the end point, the network parameters of the temporal network to be trained are updated in the current iteration round until all iteration rounds are traversed to obtain the trained temporal network.
3. The method according to claim 1, characterized in that, The improved ant colony algorithm based on UAV joint feedback is used to analyze and process the current traffic flow characteristics and pheromones fed back by UAVs in real time between the target vehicle at the current intersection node and each adjacent intersection node, and to calculate the transition probability between the current intersection node and each adjacent intersection node, including: By combining the pheromone retention, pheromone addition, and global pheromone compensation of the path segments between the current intersection node and each of the adjacent intersection nodes, the pheromone corresponding to each path segment is calculated. By combining the path segment length and traffic congestion level of each path segment in the current traffic flow characteristics, the heuristic factors corresponding to each path segment are calculated. By combining the pheromone and heuristic factor corresponding to each path segment, the transition probabilities between the current intersection node and each of its adjacent intersection nodes are calculated.
4. The method according to claim 1, characterized in that, The multi-level network includes a long and short time series modeling layer, a nonlinear time series modeling layer, and a regression analysis layer. The network based on a preset multi-layered structure performs layer-by-layer data analysis on the current traffic flow characteristics and the predicted traffic flow characteristics obtained by combining the current image traffic data and the historical text traffic data to obtain a traffic light optimization strategy, including: The current traffic flow feature and the predicted traffic flow feature are respectively analyzed and processed based on the long and short time series modeling layer to obtain the first long and short time series feature corresponding to the current traffic flow feature and the second long and short time series feature corresponding to the predicted traffic flow feature. The first long and short time series features and the second long and short time series features are subjected to data analysis and processing based on the nonlinear time series modeling layer to obtain nonlinear modeling enhanced time series features; The time-series features enhanced by the nonlinear modeling are processed by the regression analysis layer to obtain the signal light setting durations of each intersection node in different directions in the current traffic environment, which can be used as a signal light optimization strategy.
5. The method according to claim 1, characterized in that, After using the route optimization strategy and the traffic light optimization strategy as the traffic driving strategies for each vehicle in the current traffic environment, the method further includes: After each vehicle in the current traffic environment travels according to the traffic driving strategy, target traffic data collected during the target time period is acquired. Based on the target traffic data, the location information of each vehicle is obtained. If a vehicle that has not reached the destination is detected based on the location information of each vehicle, the starting point and driving time point corresponding to the vehicle that has not reached the destination are updated to obtain the updated starting point and updated driving time point corresponding to the vehicle that has not reached the destination. Based on the updated starting point and updated travel time point corresponding to the vehicles that have not reached the destination, the route optimization strategy is recalculated for the vehicles that have not reached the destination until all vehicles in the current traffic environment reach their respective destinations.
6. A traffic driving strategy optimization device based on joint modeling of UAV observations, characterized in that, The device includes: The acquisition module is used to acquire current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in the historical time period in the current traffic environment. The image capturing device is a combination of multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment. The first optimization module is used to determine a target path optimization algorithm that matches the complexity of the current traffic environment, and to perform data analysis and processing on the current traffic flow characteristics obtained from the current image traffic data based on the target path optimization algorithm to obtain a path optimization strategy. The complexity of the current traffic environment is determined based on the current image traffic data. The second optimization module is used to perform layer-by-layer data analysis and processing on the current traffic flow characteristics and the predicted traffic flow characteristics obtained by combining the current image traffic data and the historical text traffic data, based on a preset multi-layer network structure, to obtain a traffic light optimization strategy. An execution module is used to use the route optimization strategy and the traffic light optimization strategy as traffic driving strategies for each vehicle in the current traffic environment. The traffic driving strategies are used to provide driving route suggestions and driving speed suggestions for each vehicle in the current traffic environment. The first optimization module is further configured to: if the complexity of the current traffic environment meets a preset first judgment condition, then use a path optimization algorithm based on a temporal network as the target path optimization algorithm, wherein the temporal network is a network that combines deep reinforcement learning and temporal modeling structure; take any vehicle in the current traffic environment as the target vehicle, perform temporal feature analysis processing on the current traffic flow characteristics based on the temporal network, and predict the driving actions of the target vehicle from the starting point to the ending point at each intersection node and the driving speed between each intersection node; and obtain a path optimization strategy based on the driving actions of each vehicle in the current traffic environment from the corresponding starting point to the corresponding ending point at each intersection node and the driving speed between each intersection node. The first optimization module is further configured to: if the complexity of the current traffic environment meets a preset second judgment condition, then use the improved antagonist algorithm based on UAV joint feedback as the target path optimization algorithm; take any vehicle in the current traffic environment as the target vehicle, and according to the improved antagonist algorithm based on UAV joint feedback, perform data analysis and processing on the current traffic flow characteristics and pheromones fed back by UAVs between the target vehicle at the current intersection node and each adjacent intersection node, respectively, to calculate the transition probability between the current intersection node and each adjacent intersection node; take the adjacent intersection node corresponding to the maximum transition probability as the next intersection node corresponding to the current intersection node, and repeat the transition probability calculation for the next intersection node until the intersection nodes between the target vehicle and the destination are determined; and obtain a path optimization strategy based on the intersection nodes between each vehicle in the current traffic environment from the corresponding starting point to the corresponding destination.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Internet unmanned aerial vehicle capable of achieving continuous endurance
CN105652886A
Multi-dimensional intelligent driving path planning system
CN117824695A
Big data intelligent real-time processing and analysis method and system
CN119152686A