Method for predicting, regulating and controlling traffic flow

By adopting traffic flow prediction and regulation methods based on Transformer and A3C algorithms in intelligent transportation systems, combined with multi-path AGV planning, the system's response problems in complex environments and emergencies are solved, and the prediction accuracy and real-time nature of path planning are improved.

CN120199096APending Publication Date: 2025-06-24JIANGSU SECOND NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510498741.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing intelligent transportation systems do not respond quickly and accurately to complex environments and emergencies, and the popularization of autonomous driving technology is slow, which limits the overall efficiency of the system.

Method used

Traffic flow prediction and regulation methods based on Transformer, A3C algorithm and multi-path AGV planning are adopted. By constructing and improving Transformer model and A3C model, collaborative multi-AGV path planning is realized, quickly adapt to traffic signal changes and optimize AGV paths.

Benefits of technology

It improves the accuracy of traffic flow prediction and the real-time nature of path planning, enhances the adaptability and stability of the system, and can respond to complex spatiotemporal and spatial characteristics and emergencies more quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199096A_ABST
    Figure CN120199096A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic flow prediction and regulation and control method based on Transform, an A3C algorithm and multi-path AGV (Automatic Guided Vehicle) planning. According to the method, the mixed attention is introduced into the Transform model, so that high-precision traffic flow prediction is realized. A traffic signal control method based on an A3C algorithm is constructed, and a strategy gradient method is introduced to optimize a signal scheduling strategy so as to reduce vehicle waiting time and relieve traffic congestion. Through AGV (Automatic Guided Vehicle) multi-path planning, cooperative scheduling of signal control and path optimization is realized, and the traffic efficiency is improved. The method can effectively deal with a complex traffic environment, optimizes a signal control strategy, and improves the adaptability and stability of an intelligent traffic system. The method is suitable for the scenes of urban road management, intelligent traffic scheduling, automatic driving systems and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent transportation, and relates to a traffic flow prediction and regulation method based on Transformer, A3C algorithm, and multi-path AGV planning. Background Art

[0002] With the continuous development of intelligent transportation systems, the importance of traffic prediction and regulation has become increasingly prominent. As the core of advanced traffic management systems, it is crucial for realizing traffic planning, management, and control. Traffic flow prediction and regulation can help traffic managers anticipate traffic congestion in advance, make preparations for traffic restrictions in advance, and enable residents to choose the optimal travel route, thereby improving travel efficiency. However, since traffic prediction needs to handle complex spatio-temporal dependencies, it remains a challenging task.

[0003] There are still some drawbacks and deficiencies in the practical application of intelligent transportation systems in our country that need to be solved urgently. The adaptability of intelligent transportation systems to complex environments is poor, especially in adverse weather and emergencies, and the response is not fast and accurate enough. Over-reliance on prediction models of artificial intelligence and machine learning may also lead to improper handling of emergencies. The lack of in-depth cooperation between vehicles and infrastructure affects the effectiveness of vehicle networking technology, and the slow popularization process of autonomous driving technology restricts the overall efficiency of intelligent transportation systems. In the past few decades, various segmentation technical solutions have been proposed, whether region-based or edge-based. However, unfortunately, most methods are affected by operating conditions such as lighting, weather, object reflection, and other random changes when applied to natural scenes. Extracting the exact shape of vehicles is not an easy task, especially when vehicles appear overlapping or far from the camera. Various schemes for detecting vehicles in traffic scenes have been reported, and most of them utilize static and temporal image attributes to overcome the difficulties of segmentation.

[0004] Therefore, it is necessary to design effective traffic flow prediction and regulation methods to capture the periodicity and dependencies of traffic data, combine long-short-term encoders and attention mechanisms, so as to improve the prediction accuracy and the real-time performance of path planning. Summary of the Invention

[0005] The present application provides a method for traffic flow prediction and regulation. By constructing an improved Transformer model and an A3C model, and performing collaborative multi-AGV path planning, it solves the current problems of AGV collision, deadlock, and congestion, can quickly adapt to changes in traffic signals and optimize AGV paths, and realizes fast response and real-time adjustment under complex spatio-temporal feature conditions.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a traffic flow prediction method, including:

[0008] S1: Obtain road traffic data and construct a traffic flow data set;

[0009] S2: Build a traffic flow prediction network, and use the traffic flow data set for training to obtain a traffic flow prediction model;

[0010] S3: Use the traffic flow prediction model to achieve real-time prediction of traffic flow.

[0011] Preferably, in step S2, the traffic flow prediction network includes a temporal attention mechanism module, a spatial attention mechanism module, a channel attention mechanism module, and a Transformer model. The steps of building the traffic flow prediction network include:

[0012] S2.1: Based on the intersection topology of the target city, abstract a graph structure; in the graph structure, nodes represent intersections of urban traffic roads, and edges represent lanes connecting intersections;

[0013] S2.2: Take three-dimensional traffic flow data X ∈ R N×T×F as the inputs of the temporal attention mechanism module, the spatial attention mechanism module, the channel attention mechanism module, and the Transformer respectively, where N is the number of nodes in the graph structure, T is the number of time steps, and F is the feature dimension;

[0014] S2.3: After the output of the temporal attention mechanism module passes through residual connection, layer normalization, and linear projection transformation, obtain the feature X time ;

[0015] S2.4: After the output of the spatial attention mechanism module passes through residual connection, layer normalization, and linear projection transformation, obtain the feature X space ;

[0016] S2.5: After the output of the channel attention mechanism module passes through normalization and linear projection transformation, obtain the feature X channel ;

[0017] S2.6: After concatenating the features X time , X space and X channel , input them into two multi-head attention modules in the Transformer model respectively.

[0018] Preferably, in the temporal attention mechanism module, obtain the attention score matrix T (T) between time steps, and reshape T (T) to an N×T×F structure; the element in the t-th row and t'-th column of T (T) T tt′ represents the degree of attention of the \(t\)-th time step to the \(t'\)-th time step, where \(Q\) T (t) is the query vector at the \(t\)-th time step, and \(K\) T (t′) is the key vector at the \(t'\)-th time step, and \(K\) T (k) is the key vector at the \(k\)-th time step.

[0019] Preferably, in the spatial attention mechanism module, an attention matrix \(S\) between nodes is obtained (S) , and \(S\) is (S) reshaped into an \(N\times T\times F\) structure; the element at the \(i\)-th row and \(i'\)-th column in \(S\) (S) represents the degree of attention of the \(i\)-th node to the \(i'\)-th node, where \(Q\) \(S\) ii′ represents the degree of attention of the \(i\)-th node to the \(i'\)-th node, where \(Q\) s (i) is the query vector of the \(i\)-th node, and \(K\) s (i′) is the key vector of the \(i'\)-th node, and \(K\) s (c) is the key vector of the \(c\)-th node.

[0020] Preferably, the channel attention mechanism module includes a cascaded global pooling module, a non-linear mapping module, and a channel reweighting module.

[0021] Preferably, the positional encoding in Transformer includes:

[0022]

[0023] where \(PE(t, 2m)\) and \(PE(t, 2m + 1)\) respectively represent the even and odd position encodings of the \(m\)-th dimension feature at the \(t\)-th time step.

[0024] Preferably, the traffic flow data includes traffic volume, vehicle speed, driving route, traffic density, signal light status, and congestion situation.

[0025] In a second aspect, a traffic signal control method is further provided, and the control method includes:

[0026] Construct a hierarchical critic-actor structure, where: the actor structure adopts a two-level hierarchical modeling method, the high-level actor takes the traffic state matrix as input and the signal cycle vector of each intersection as output; the low-level actor takes the single intersection state vector as input and the phase selection probability distribution as output; the critic structure adopts a distributed value function estimation mechanism, including a global critic and a local critic, and at the same time sets a cross-level gradient compensation mechanism; the estimation method of the advantage function adopts a multi-time scale hybrid strategy;

[0027] Train the constructed hierarchical critic-actor structure, and use the trained hierarchical critic-actor structure for traffic signal control;

[0028] where the traffic state matrix is d is the dimension of the state vector of the intersection, r i = [v i,1 , v i,2 , …, v i,num , q i,1 , q i,2 , … q i,num , v n,p represents the total number of vehicles on the p-th lane of intersection i, q n,p represents the number of vehicles queuing for passage on the p-th lane of intersection i. The elements in the traffic state matrix are obtained based on the traffic flow prediction method described above.

[0029] In a third aspect, a collaborative multi-AGV path planning is also provided. The path planning method sets the reward function as follows based on the traffic flow prediction result obtained by the traffic flow prediction method described above: where Flow u (t) represents the vehicle throughput of lane u at the t-th time step, QueueLength u (t) represents the queue length of lane u at the t-th time step, and Delay u (t) represents the vehicle delay time of lane u at the t-th time step.

[0030] Compared with the prior art by adopting the above technical solutions, the present invention has the following technical effects:

[0031] 1. Compared with the prior art of traffic flow prediction, after introducing three kinds of attention, the Transformer model can, on the premise of the ability to mine long-term historical impacts, directly model the road network structure by combining the spatial attention mechanism, realizing dynamic spatio-temporal fusion. Using the channel attention structure, the model can automatically learn to select important features during training and assign higher weights, enabling the model to have stronger generalization ability and scalability. Using the function of dynamically selecting features, the present invention can significantly improve the prediction stability in scenarios such as emergencies or holidays that are "quite different from normal times". And in the splicing of time attention, the attention ability of multi-head attention is more comprehensive. Compared with a single Transformer model that only emphasizes global modeling, the present invention is more sensitive to local traffic anomalies and has a stronger response ability.

[0032] 2. In signal light regulation, the present invention introduces a hierarchical policy gradient to optimize the A3C algorithm. Through hierarchical reinforcement learning, the high-level policy makes macro decisions, and the low-level policy controls specific actions at the micro level according to the instructions of the high-level. Using this optimization method, the model can decouple the target task training and be more stable. Compared with the machine learning in use, it can take into account both macro-strategic and micro-execution capabilities, and the regulation accuracy is higher. At the same time, the hierarchical policy expands the learning strategy expression ability of the algorithm and can adapt to various traffic patterns. Compared with the existing traffic signal light regulation technology, the model can also achieve coordinated regulation between multiple intersections.

[0033] 3. In realizing path planning, the AGV algorithm of the present invention combines the target planning and the advantage function mechanism, comprehensively considers the traffic flow, vehicle waiting time, and passing efficiency, and optimizes the shortest path selection through weighted regression by the comprehensive advantage function. Compared with the prior art, it can dynamically optimize the path selection by combining real-time traffic data and historical experience, and improve the passing efficiency. And the advantage weighted regression method not only minimizes the driving time, but also comprehensively considers factors such as signal light waiting and road risks. Brief Description of the Drawings

[0034] Figure 1 It is a flowchart of the method of the embodiment of the present application.

[0035] Figure 2 It is a schematic structural diagram of the traffic flow prediction network in the embodiment of the present application. Detailed Embodiments

[0036] The following details the embodiments of the present invention, and the examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0037] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.

[0038] The present invention discloses a traffic flow prediction and regulation method based on Transformer, A3C algorithm, and multi-path AGV planning, as Figure 1As shown, this method realizes high-precision traffic flow prediction by introducing hybrid attention on the Transformer model. A traffic signal control method based on the A3C algorithm is constructed, and the policy gradient method is introduced to optimize the signal scheduling strategy to reduce vehicle waiting time and alleviate traffic congestion. Through the multi-path planning of AGV (Automated Guided Vehicle), the collaborative scheduling of signal control and path optimization is realized to improve the traffic efficiency. The present invention can effectively cope with complex traffic environments, optimize signal control strategies, and enhance the adaptability and stability of intelligent transportation systems. This method is applicable to scenarios such as urban road management, intelligent traffic scheduling, and autonomous driving systems. The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0039] I. Traffic Flow Prediction Process

[0040] Select the main traffic nodes in the target city, and determine the weight relationship of vehicles based on the number of vehicles in each lane and the number of vehicles queuing for passage in each lane. Use three-dimensional traffic flow data X ∈ R N×T×F as the unified input, where N represents the number of nodes, T is the number of time steps, and F is the feature dimension.

[0041] Set the action space at the input end of the model to 4 phases or 8 phases according to the topological structure of the traffic nodes; and select different signal phases according to the specific extension direction of the road.

[0042] In actual traffic flow prediction, the ST-GCN structure is generally used for short-term traffic conditions, and the Transformer structure is generally used for long-term traffic conditions.

[0043] 1. ST-GCN

[0044] Construct a graph convolutional network GCN, and abstract the graph structure according to the topological structure of the intersections in the selected city. The input feature matrix is X ∈ R N×F , where N is the number of intersection nodes and F is the feature dimension. The formula for graph convolution operation: H (l) is the node feature representation of the l-th layer, and H (0) = X is the input feature. is the normalized adjacency matrix, A is the adjacency matrix, D is the node degree matrix, and D i,i = ∑ j A ij . is the learnable weight matrix of the l-th layer. σ is the activation function.

[0045] In addition, ST-GCN also needs to use graph convolution operations to aggregate the information of adjacent nodes and capture spatial dependencies. Specifically, the features of each node are weighted and averaged by the features of its adjacent nodes to generate a new feature representation.

[0046] The temporal convolution operation of ST-GCN is implemented through 1D convolution. Given a time series X of node features t ∈R N×T×F , where N is the number of nodes, T is the number of time steps, and F is the feature dimension. The calculation of the actual convolution layer is as follows: H k+t is the input feature matrix at time step k + t. is the learnable weight matrix of the temporal convolution layer. K is the size of the convolution kernel, that is, the number of time steps considered. is the output of the temporal convolution of the (l + 1)-th layer.

[0047] Among them, the core step of ST-GCN is the spatio-temporal graph convolution, which simultaneously processes spatial and temporal dependencies: is the spatial graph convolution part, which aggregates the features of the adjacency matrix. is the part of the temporal convolution, which considers the dependencies between time steps. and are the learnable weights of the temporal and spatial convolution layers.

[0048] ST-GCN aggregates spatio-temporal data through the joint operation of spatial convolution and temporal convolution, and then generates global spatio-temporal features. The calculation method is as follows: Here, X t is the node feature matrix at each time slice, is the normalized adjacency matrix, and H t is the output of the spatio-temporal features.

[0049] In the process of ST-GCN capturing temporal correlations, the training loss value at each step needs to be quantified by a loss function. Here, the mean squared error is used to measure the difference in the prediction results: is the actual traffic flow of the i-th node at time t. is the predicted traffic flow.

[0050] The training of the ST-GCN network includes updating the network parameters through the backpropagation algorithm. During the training process, the weight parameters in the model are updated by the gradient descent algorithm, and the Adam optimizer is used. The parameter update formula is:

[0051] θ t is the current model parameter, η is the learning rate, m t is the first moment estimate of the gradient, and v tis the second moment estimation of the matrix. ε is a constant to avoid division by zero.

[0052] 2. Transformer

[0053] At the input end of the Transformer, the input traffic flow data is represented as X = {x1, x2, …, x T}, where x t ∈R F is the traffic flow feature at time step t. T is the number of time steps, and F is the feature dimension, which is used for the prediction of long-term traffic conditions.

[0054] To process time series data, the Transformer needs to add positional encoding for each time step to manifest the sequential information. The formula is: and PE(t, 2m) and PE(t, 2m + 1) respectively represent the even and odd position encodings of the m-th dimension at time step t. Finally, the representation after adding the positional encoding to the input feature is: x′ t = x t + PE(t).

[0055] For the construction of the self-attention mechanism, first, the input x′ t needs to be mapped to query, key, and value vectors:

[0056] Q t = W Q x′ t , K t = W K x′ t , V t = W V x′ t

[0057] where W Q , W K , W V are trainable weight matrices.

[0058] The calculation process of self-attention is directly introduced: Q, K, and V are the query matrix, key matrix, and value matrix respectively. d k is the dimension of the key vector, and its function is to calculate the similarity between the query and the key. The similarity is normalized by softmax to obtain the attention weights.

[0059] To enable the algorithm to capture information in different subspaces, in the present invention, a multi-head attention design is also carried out for the Transformer. The formula is: MultiHead(Q, K, V) = Concat(head1, head2, …, head h )WO , the calculation method for each head:

[0060]

[0061] where h is the number of attention heads, and W O is the weight matrix of the output.

[0062] Build the encoder. Each layer of the encoder needs to calculate the self-attention layer and the feed-forward neural network. The calculation formula for the former is Attention output = MultiHead(Q, K, V), and then it also needs to go through layer normalization and residual connection: LayerNorm1(Attention output + x′ t ). For the latter: FFN(x′ t ) = ReLU(x′ t W1 + b1)W2 + b2 and its normalization: LayerNorm2(FFN(x′ t ) + LayerNorm1(Attention output + X′ t ). The encoder is stacked by L layers with the same structure, and the output is: Z = Encoder(X).

[0063] Build the decoder. Each layer of the decoder includes a self-attention mechanism, an encoder-decoder attention mechanism, and a feed-forward neural network: SelfAttention output = MultiHead(Q, K, V), CrossAttention output = MultiHead(Q, K′, V′), where K′, V′ are the queries and key values output by the encoder. The feed-forward neural network is the same as that of the encoder. Finally, each layer is normalized: DecoderOoutput t = LayerNorm3(FFN3(CrossAttention output ) + selfAttention output ).

[0064] The output of the last layer of the decoder is obtained through a linear transformation to get the traffic prediction as: where, is the predicted traffic, and W out is the weight coefficient of the output layer.

[0065] 3. For the traffic prediction problem with very complex traffic conditions, a single ST-GCN and Transformer often do not have good accuracy. For example, Figure 1As shown, in the present invention, hybrid attention (temporal, spatial, and channel attention) fusion will be introduced to optimize the prediction of complex traffic. Compared with existing prediction technologies, it can utilize both spatio-temporal relationships and long-sequence dependencies to improve prediction accuracy.

[0066] In the attention mixing stage, the input tensor is still: X ∈ R N×T×F , where the specific parameter explanations are the same as above. N, T, and F represent the number of traffic nodes, the number of historical time steps, and the feature dimension of each node, respectively.

[0067] 3.1. Construct the temporal attention mechanism

[0068] First, perform an average operation on X in the spatial dimension to obtain where, X i,t represents all the feature values of the i-th node at the t-th time step.

[0069] Next, calculate two process quantities, which are used to define the learnable linear transformation weights, namely the query vector and the key vector where is the projection matrix of the temporal attention, and d t represents the temporal attention dimension.

[0070] Subsequently, calculate the attention scores between time steps T tt′ represents the degree of attention of the t-th time step to the t'-th time step, and obtain the time step attention matrix T (T) .

[0071] Then, perform a tensor transformation on the time step attention matrix T (T) , that is, change the data dimension arrangement and then revert to the original input to obtain the N×T×F structure.

[0072] Finally, for the temporally weighted feature output X (T) = X · T (T) ∈R N×T×F , using this form to express that the matrix multiplication only acts on the time dimension.

[0073] After performing residual connection and layer normalization on the temporally weighted feature output X (T) , the model now has the ability to learn the dynamic dependencies between historical time steps and can identify which time steps have a strong impact on the current prediction.

[0074] 3.2. Construct the spatial attention mechanism, and the implementation method is similar to the temporal attention mechanism.

[0075] First, perform an average operation on X in the time dimension to obtain

[0076] Next, calculate two process quantities. The expression is a learnable spatial attention projection matrix, and d s represents the spatial attention dimension.

[0077] Subsequently, calculate the attention scores between nodes S ii′ represents the degree of attention of the i-th node to the i'-th node, and obtain the attention matrix S between nodes (S) .

[0078] Then, perform tensor deformation on the attention matrix S between nodes (S) That is, change the data dimension arrangement and then back to the original input to obtain the N×T×F structure.

[0079] Finally, for the spatially weighted feature output X (s) = S (S) · T (T) ∈ R N×T×F .

[0080] After performing residual connection and layer normalization on the spatially weighted feature output X (s) , the model now has the ability to learn the dynamic weights between nodes. Compared with the existing technology, this model can identify which nodes have strong correlations.

[0081] 3.3. Construct the channel attention mechanism

[0082] In this model, the channel attention mechanism is also used. Compared with the existing technology, the algorithm can weight each feature channel to highlight more important feature dimensions.

[0083] First, perform global pooling on the input X:

[0084] Then, perform non-linear mapping s = σ(W2·ReLU(W1·z)) ∈ R F , W1 ∈ R F×F / r , W2 ∈ R F / r×F are channel compression and channel restoration respectively. σ is the Sigmoid function, and the output range is between 0 and 1. s represents the channel attention weight vector.

[0085] Finally, perform channel re-weighting The expression means element-wise multiplication by channel.

[0086] After normalizing the channel attention weight vector, the corresponding feature output of channel attention is obtained.

[0087] 3.4. After constructing the three attention mechanisms, perform two-layer linear projection and feature concatenation on the feature outputs of the above three attention modules respectively, which are used as the output of the hybrid attention mechanism module, and feed the output of the hybrid attention mechanism module to the two multi-head attention modules in the Transformer model, so that the Transformer model can simultaneously focus on space and time during encoding, as Figure 2 shown. The existing Transformer model is optimized using the hybrid attention mechanism, making the Transformer model have stronger expressive ability and be more suitable for dealing with complex dynamic scenarios.

[0088] II. Traffic signal control process

[0089] Build a hierarchical critic-actor structure. The present invention introduces the hierarchical policy gradient method (A3C) to optimize the traditional model architecture. The high-level policy network is responsible for the global planning of the signal cycle at the regional level, and the low-level policy network focuses on the signal phase control of specific intersections. During the training process, a dynamic entropy coefficient adjustment mechanism is introduced to balance exploration and exploitation, improve the convergence stability and cross-scenario generalization ability. The overall system adopts a multi-threaded asynchronous training framework, which is suitable for dealing with complex traffic control problems in large-scale continuous state spaces. The policy loss function is defined as: L π =-E[log π (a|s)A(s,a)-β t Hπ((·|s))], where β t is the entropy adjustment coefficient that dynamically decays with the training progress.

[0090] In the traffic signal control link, it is necessary to construct a regional traffic state matrix as the input of this part of the model. Use v p to represent the total number of vehicles on the p-th lane, and use q p to represent the number of vehicles queuing for passage on the p-th lane. If an intersection has num lanes, the state vector of this road intersection can be defined as: r = [v1, v2,..., v num , q1, q2,... q num

[0091] If there are N intersections in this area and the state vector dimension of each intersection is d, then the regional traffic state matrix that the model needs to construct is

[0092] The actor structure adopts a two-level modeling method: the high-level actor takes the regional traffic state matrix S zone as the input and outputs the signal cycle vector of each intersection to guide the low-level policy target; the low-level actor takes the single intersection state vector s t =[q, v, φ, Δt​remain is used as the input to output the phase selection probability distribution, where the remaining time variable is Δt remain It is represented by sine coding as follows:

[0093]

[0094] The critic structure adopts a distributed value function estimation mechanism, including a global critic and a local critic. At the same time, a cross-level gradient compensation mechanism is set. When the local network error is backpropagated, the gradient information from the high-level network will be fused to improve the policy coordination.

[0095] The estimation method of the advantage function adopts a multi-time scale hybrid strategy, and its form is:

[0096] A hybrid = λA n-step +(1 - λ)A GAE

[0097] Among them, λ is the fusion coefficient, which is automatically adjusted by the meta-learning strategy during training to minimize the advantage variance.

[0098] The parameter definitions are as follows:

[0099] R t:t+n : The n-step discounted return starting from time step t; τ t+k : The immediate reward at the (t + k)-th step, which is represented by an asymmetric clipping function as:

[0100]

[0101] The discount factor γ adopts a dynamic adjustment form:

[0102] γ t = 0.99 × e -0.001t .

[0103] The training process of A3C includes the following steps: Initialize the hierarchical global network (including the high-level policy and the low-level policy), the local policy network, and multiple environment copies; In the multi-threaded environment with parallel execution, the worker threads simultaneously use the high-level policy to generate sub-goals and execute the low-level policy in the local environment; Introduce a cross-level sampling mechanism, the high-level network generates a sub-task goal g every K steps t , and the low-level policy makes a joint decision based on the state s t and the goal g t ; The update of the high-level policy network follows the following gradient calculation formula:

[0104]

[0105] During the policy update process, the overall loss function is composed as follows:

[0106]

[0107] Among them, α is the adaptive hierarchical weight, defined as:

[0108] is the high-level loss, and the calculation formula is:

[0109]

[0110] Weighted by the advantage function A zone (S, g), optimize the generation strategy π of the sub-goal g H .

[0111] represents the low-level loss, defined as:

[0112]

[0113] represents the policy gradient loss that maximizes the action probability under the advantage function A.

[0114] The update of the value network introduces a gradient truncation strategy, and the loss function is:

[0115]

[0116] During the implementation of traffic signal control, the low-level policy network jointly outputs the action probability distribution according to the state s and the high-level goal g:

[0117] π(a|s, g) = Softmax(f θ L(s)||W g (g))

[0118] Among them, W g is the high-level goal embedding matrix, which ensures information alignment and sharing among multiple levels in the policy network.

[0119] When performing traffic signal control, the action output by the policy network is based on the probability distribution in the current state, indicating the probability of selecting each phase (signal light state) in the current state. Let π θ (a t |s t ) represent the probability that the agent selects the action a t in the state s t , and the softmax function is used in the present invention to represent: Among them, θ is the parameter of the policy network, and W π and b t are the weights and biases of the network.

[0120] Performance Evaluation: Comparing the performance improvement of the A3C algorithm with traditional traffic signal control methods (such as fixed-time control, fixed-period control, etc.), the following formula is used to analyze traffic signal control:

[0121]

[0122] Among them, Metric represents the performance metrics used, such as traffic flow and average waiting time.

[0123] Performance evaluation metric for measuring the increase in traffic flow under the control of the A3C algorithm compared to the traditional method, i.e., the efficiency improvement:

[0124]

[0125] III. Cooperative Multi-AGV Path Planning Process

[0126] The signal light scheme optimized by A3C will affect the traffic state of the road network, and the AGV conducts path planning based on the updated traffic state; at the same time, the path decision of the vehicle will also act on the signal light optimization, making the two form a dynamic closed loop and jointly improving the overall efficiency of the traffic system.

[0127] When building the AGV control model system, first select the main traffic nodes of the city, and use the traffic state matrix at the input end to represent the real-time traffic state of the intersection. Each element of this matrix corresponds to the traffic characteristics of a specific lane, such as traffic flow, queue length, average vehicle speed, and signal light state, etc. These data are used as the input of the model to provide accurate basic information for signal control decision-making and path optimization.

[0128]

[0129] Among them, s a,b represents the b-th traffic state on the a-th lane.

[0130] For the action space of signal phase selection, taking the 4-phase topological structure as an example here, the action space of the traffic signal in the AGV model can be defined as A = {α′1, α′2, α′3, α′4}, where α′1 to α′4 can represent different signal phases, and need to be assigned according to the position requirements of specific road directions in the urban intersection.

[0131] Similar to the A3C model above, the purpose of designing the reward function for AGV is to drive the system towards the optimization goal, such as reducing vehicle waiting time, improving traffic efficiency, etc., and can be constructed as:

[0132] res v =-η·WaitingTime(s v )-β·QueueLength(s v) + γ·TaskCompletion(s v )。

[0133] Among them, WaitingTime(s v ) is the total waiting time of vehicles on the lane, QueueLength(s v ) is the number of queuing vehicles on the lane, TaskCompletion(s v ) is the number of tasks completed, and α, β, γ are weight coefficients.

[0134] The signal control strategy is adjusted through a reinforcement learning model, and the best signal light action is selected based on the traffic state. The agent selects an action a v according to the state s v , and there is: a v = π θ (s v ). Among them, π θ (s v ) represents the intermediate result of the policy network, that is, the probability of selecting different actions when the state s v is given.

[0135] Multi-AGV path planning. The core of the path planning of this algorithm is to minimize the path cost, and at the same time combine a single-objective combination planning model and a reward function mechanism to optimize the shortest path selection. Compared with the existing technology, it can dynamically optimize the path selection by combining real-time traffic data and historical experience, avoid congestion, and improve the traffic efficiency. Its advantageous weighted regression method not only minimizes the travel time, but also comprehensively considers factors such as signal light waiting and road risks.

[0136] This system uses a network graph abstracted into a data structure as the input end, and the weight of each edge of the network graph represents the path cost: Among them, u refers to the u-th edge of the network graph abstracted as the output end data structure, ω u is the weight of the u-th edge, and d u is the specific path distance or time consumption of the u-th edge.

[0137] The AGV algorithm of the present invention uses the Dijkstra algorithm for shortest path selection. The cost function of this algorithm is: f(i) = g(i) + h(i), where g(i) is the actual cost from the starting point to the current node i, and h(i) is the estimated cost from the current node i to the target node.

[0138] For the dynamic path planning problem, the AGV needs to update the path according to the real-time traffic conditions during the system operation. The optimization problem of real-time path planning can be expressed as: Among them, Cost(s v , a v) is the state traffic s v to the next state s v+1 cost, a v is the control action.

[0139] The optimization goal of the system is traffic congestion, waiting time, and maximizing the task completion rate. It can be expressed as:

[0140]

[0141] where α, β, and γ are weights, and M is the number of AGVs.

[0142] Furthermore, the model needs to quantify the impact of signals on the path. The state of traffic signals affects the path selection of AGVs. Considering the impact of the signal cycle, AGVs need to adjust the path according to the state of the traffic lights. Assume the current signal cycle is T cycle , and the AGV needs to complete the task within the cycle. The target task expression is:

[0143]

[0144] where N phases is the number of phases within the signal cycle.

[0145] Because traffic signals adjust their states according to real-time traffic conditions (such as traffic flow, vehicle waiting time, etc.). So the algorithm needs to adjust the signal state according to the waiting time:

[0146] SignalUpdate(v) = SignalingPolicy(s v , a v ).

[0147] where SignalingPolicy represents a policy function that selects the best signal according to the traffic state.

[0148] Subsequently, the model will select whether to adjust the path according to the dynamic signal state. The model will first evaluate the state of the traffic lights at the current intersection, that is, whether it is currently red or green, and the remaining time of the traffic lights. At this time, the model starts to judge, which can avoid additional time waste caused by waiting: NewPath(v) = PathAdjustment(s v , π θ ). The key factor PathAdjustment(s v , π θ ) represents the path adjustment function, and its role is to adjust the path according to the current state and control strategy.

[0149] Construction of a comprehensive reward function. The reward function is used to evaluate the effectiveness of the signal control strategy selected by the agent. The fundamental principle is to drive intelligent transportation informatization learning and needs to be defined according to indicators such as traffic flow, vehicle waiting time, and traffic efficiency.

[0150] Denote the queue length of the \(u\)-th lane at time \(v\). \(U\) is the number of lanes at the intersection. The goal of this function is to minimize the queue lengths of all lanes. The reward value is usually negative to penalize situations with long queues. Secondly, since vehicle delay is another important indicator to measure the traffic signal control effect. The smaller the vehicle delay, the more efficient the traffic signal control.

[0151] Define a reward function based on the delay time: where denotes the vehicle delay time of the \(u\)-th lane. The goal of this reward function is to minimize the vehicle delay time, thereby improving the traffic efficiency of the entire intersection.

[0152] Furthermore, construct a reward function based on the traffic volume: where denotes the traffic volume of the \(u\)-th lane at time \(v\). Through this reward function, the agent will be encouraged to adopt a signal control strategy that can increase the vehicle throughput. Finally, design a comprehensive hybrid reward function to balance the influence of various factors as: where \(\omega_1\), \(\omega_2\) and \(\omega_3\) are weight coefficients used to balance the influence of different indicators.

[0153] Based on the reward function \(R\) v Calculate the state-action value under the condition: where \(Q(s\) v ,a\) v ) reflects the long-term reward of choosing path \(l\) v under the current traffic state \(s\) v . \(\zeta\) is the discount factor that controls the influence degree of future rewards. Continue to obtain the advantage function: \(A(s\) v ,l\) v ) = \(Q(s\) v ,l\) v ) - \(VALUE(s\) v ), \(VALUE(s\) t ) is the state value function, indicating the average return of all possible paths in state \(s\) v . \(A(s\) v ,l\) v ) measures the quality of the current path selection relative to the average path.

[0154] Finally, obtain the optimized path selection strategy of the AGV: \(\pi(a\) v |s\) v) ∝ π(a v |s v ) exp{A(s v , a v )}. By adjusting the path selection strategy through weighted regression of advantages, the paths with greater advantages are more likely to be selected. Compared with the existing traffic road management technologies, this model generally has better adaptability under different traffic conditions.

[0155] AGV algorithm performance evaluation. The completion time of tasks is an important indicator to measure the performance of AGVs, defined as the average completion time of all tasks. The calculation formula for this indicator: T task,y is the completion time of the y-th task, where N represents the total number of intersection nodes.

[0156] Calculate the ratio of the path length to the shortest path:

[0157] Finally, the algorithm also needs to calculate the system congestion degree, which represents the degree of traffic congestion and requires calculating the average waiting time at all intersections: N is still the number of intersections.

[0158] Based on the same technical solution, the present invention also designs an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned traffic flow prediction and regulation method are implemented.

[0159] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned traffic flow prediction and regulation method are implemented.

[0160] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0161] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any variations or substitutions that can be understood and conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A traffic flow prediction method, characterized in that: include: S1: Obtain road traffic data and build a traffic flow dataset; S2: Build a traffic prediction network and use the traffic flow dataset for training to obtain a traffic prediction model; S3: Use the traffic prediction model to achieve real-time prediction of traffic flow.

2. The traffic flow prediction method according to claim 1, characterized in that: The traffic prediction network in step S2 includes a temporal attention mechanism module, a spatial attention mechanism module, a channel attention mechanism module, and a Transformer model. The steps of building the traffic prediction network include: S2.1: Based on the intersection topology of the target city, a graph structure is abstracted; in the graph structure, nodes represent intersections of urban traffic roads, and edges represent lanes connecting intersections; S2.2: Three-dimensional traffic flow data X∈R N×T×F They are used as inputs of the temporal attention mechanism module, the spatial attention mechanism module, the channel attention mechanism module, and the Transformer, respectively, where N is the number of nodes in the graph structure, T is the number of time steps, and F is the feature dimension; S2.3: The output of the temporal attention mechanism module is transformed through residual connection, layer normalization, and linear projection to obtain feature X time ; S2.4: The output of the spatial attention mechanism module is transformed through residual connection, layer normalization, and linear projection to obtain feature X space ; S2.5: The output of the channel attention mechanism module is normalized and linearly projected to obtain feature X channel ; S2.6: Feature X time , X space and X channel After feature concatenation, they are input into the two multi-head attention modules in the Transformer model respectively.

3. The traffic flow prediction method according to claim 2, characterized in that: In the temporal attention mechanism module, obtain the attention score matrix T between time steps (T) , and for T (T) Perform tensor deformation to N×T×F structure; T (T) The element in row t and column t′ T tt′ Indicates the degree of attention of the t-th time step to the t′-th time step, Q T (t) is the query vector at the tth time step, K T (t′) is the key vector at the t′th time step, K T (k) is the key vector for the kth time step.

4. The traffic flow prediction method according to claim 3, characterized in that: Q T (t), K T (t′) and K T (k) is obtained based on the average processing of X in the spatial dimension.

5. The traffic flow prediction method according to claim 2, characterized in that: In the spatial attention mechanism module, obtain the attention matrix S between nodes (S) , and for S (S) Perform tensor deformation to N×T×F structure; S (S) The element in row i and column i′ S ii′ Indicates the degree of attention of the i-th node to the i′th node, Q s (i) is the query vector of the i-th node, K s (i′) is the key vector of the i′th node, K s (c) is the key vector of the c-th node.

6. The traffic flow prediction method according to claim 5, characterized in that: Q s (i) K s (i′) and K s (c) Based on the result of averaging X in the time dimension.

7. The traffic flow prediction method according to claim 2, characterized in that: The channel attention mechanism module includes a cascaded global pooling module, a nonlinear mapping module, and a channel reweighting module.

8. The traffic flow prediction method according to claim 2, characterized in that: The positional encoding in Transformer includes: Among them, PE(t,2m) and PE(t,2m+1) represent the even and odd position encodings of the m-th dimension feature at the t-th time step, respectively.

9. A traffic signal control method, characterized in that: The control method comprises: A hierarchical critic-actor structure is constructed, in which: the actor structure adopts a two-level modeling method, the high-level actor takes the traffic state matrix as input and the signal cycle vector of each intersection as output; the bottom-level actor takes the single intersection state vector as input and the phase selection probability distribution as output; the critic structure adopts a distributed value function estimation mechanism, including global critics and local critics, and sets a cross-level gradient compensation mechanism; the advantage function estimation method adopts a multi-time scale hybrid strategy; Train the constructed hierarchical critic-actor structure, and use the trained hierarchical critic-actor structure to control traffic signals; The traffic state matrix is d is the dimension of the state vector of the intersection, r i =[v i,1 ,v i,2 ,…,v i,num ,q i,1 ,q i,2 ,…q i,nim ],v n,p represents the total number of vehicles on the pth lane of intersection i, q n,p Represents the number of vehicles waiting in line for passage on the p-th lane of intersection i, and the elements in the traffic state matrix are obtained based on the traffic flow prediction method as described in any one of claims 1 to 8.

10. A collaborative multi-AGV path planning, characterized in that: The path planning method is based on the traffic flow prediction result obtained by the traffic flow prediction method according to any one of claims 1 to 8, and the reward function is set as: Flow u (t) represents the vehicle traffic volume of lane u at the tth time step, QueueLength u (t) represents the queue length of lane u at the tth time step, Delay u (t) represents the vehicle delay time in lane u at the t-th time step.

Citation Information

Cited By

  • Intelligent road large model driven traffic signal optimization method and platform

    CN122050166A