Sparse self-attention fused roadside berth occupancy rate space-time prediction method

By integrating sparse self-attention, the spatial Transformer and time Informer models are used to solve the accuracy of roadside berth occupancy prediction, achieving efficient urban traffic management and convenient parking experience.

CN120199087AActive Publication Date: 2025-06-24BEIJING UNIV OF CHEM TECH +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510571777.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-24
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the occupancy rate of roadside berths, affecting urban traffic management and user parking experience.

Method used

The spatial and temporal prediction method of roadside berth occupancy with sparse self-attention is adopted. By obtaining historical parking data, building a weighted directed graph, and capturing the spatial and temporal dependencies using the spatial Transformer and temporal Informer models, accurate prediction of berth occupancy is achieved.

Benefits of technology

It realizes accurate prediction of roadside berth occupancy rate, improves the efficiency of urban traffic management, optimizes vehicle flow, reduces congestion, and provides citizens with a more convenient parking experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199087A_ABST
    Figure CN120199087A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse self-attention fused roadside berth occupancy rate space-time prediction method, and belongs to the technical field of traffic prediction, and the method comprises the steps: data collection and preprocessing: carrying out the statistics of the parking number in each time period according to historical parking data, calculating the berth occupancy rate of each berth region in each time period, and obtaining a berth occupancy rate data set; constructing an adjacent matrix of the map: obtaining a distance matrix among the parking areas by accessing a map API, and constructing the adjacent matrix through a Gaussian kernel; and constructing a parking space occupancy rate prediction model, and performing parking space occupancy rate prediction based on the parking space occupancy rate prediction model to obtain a roadside parking space occupancy rate prediction result. According to the method, the correlation between parking demand time and space is fully considered, and accurate prediction of the roadside parking space occupancy rate is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic prediction, and particularly to a spatio-temporal prediction method for roadside berth occupancy rate integrating sparse self-attention. Background Art

[0002] With the continuous increase of urban population and the growth of car ownership, parking management has become a common challenge in today's cities. Parking facilities include off-street parking lots, parking garages and roadside berths, all of which are important components of the traffic facility system. A roadside berth refers to a parking position set by the public security traffic management department using urban roads, including motor vehicle lanes, non-motor vehicle lanes, sidewalks and other parking spaces available for motor vehicles. Roadside berths are usually equipped with sensors and cameras, and the management background can obtain the occupancy situation of the berths in real time.

[0003] The prediction of roadside berth occupancy rate is crucial. On the one hand, it can provide parking information for users, guide users to available berths, reduce the time for users to search for berths, and improve the utilization rate of berth resources; more importantly, by understanding the occupancy situation of berths, the management department can better manage traffic flow, optimize vehicle flow, reduce congestion, and improve the overall traffic efficiency. Therefore, the prediction of roadside berth occupancy rate plays an important role in the efficient management of urban traffic, which is not only related to the parking experience of users, but also affects the operation efficiency of the overall urban traffic system. Summary of the Invention

[0004] The purpose of the present invention is to provide a spatio-temporal prediction method for roadside berth occupancy rate integrating sparse self-attention, which fully considers the correlation of parking demand in time and space and realizes the accurate prediction of roadside berth occupancy rate.

[0005] To achieve the above purpose, the present invention provides a spatio-temporal prediction method for roadside berth occupancy rate integrating sparse self-attention, and the steps include:

[0006] S1. Obtain historical parking data, preprocess the historical parking data, and according to the historical parking data, count the parking quantity in different time periods according to the entry and exit time of each vehicle in each roadside berth area, and calculate the berth occupancy rate of each berth area in different time periods according to the counted parking quantity and the total number of parking spaces in each berth area, so as to obtain a berth occupancy rate data set;

[0007] S2. Regard the urban berth area network as a weighted directed graph Construct the adjacency matrix of the graph;

[0008] S3. Build a parking berth occupancy rate prediction model, input the information obtained in step S1 and step S2, and perform parking berth occupancy rate prediction based on the parking berth occupancy rate prediction model to obtain the prediction result of the roadside berth occupancy rate.

[0009] Preferably, in step S2, the weighted directed graph where V is the set of vertices, E is the set of edges, and A is the adjacency matrix of the graph. Constructing the adjacency matrix of the graph includes:

[0010] S21. By accessing the map API, obtain the longitude and latitude of the roadside berth area, calculate the driving distance between each roadside berth area according to the longitude and latitude of the roadside berth area, and generate a distance matrix;

[0011] S22. According to the generated distance matrix, apply the Gaussian kernel function to each element of the distance matrix to convert the distance into similarity. The Gaussian kernel form is:

[0012]

[0013] where dis(x i →y j ) represents the distance between two points in the distance matrix, σ represents the bandwidth parameter of the Gaussian kernel, which is determined by the global standard deviation, and the Gaussian kernel makes points closer in distance have higher weights.

[0014] Preferably, the parking berth occupancy prediction model in step S3 includes:

[0015] Input layer: Use a convolutional layer to aggregate the input data and expand the data channel dimension;

[0016] Spatio-temporal layer: Includes a spatial Transformer and a temporal Informer. The spatio-temporal layer is composed of stacking spatio-temporal Transformer-Informer modules to capture dynamic spatio-temporal dependencies;

[0017] Spatial Transformer: Adopt a spatio-temporal position embedding layer, a fixed graph convolutional layer, and a dynamic self-attention mechanism layer to extract spatial information, and adopt a gating mechanism to fuse the spatial features obtained from the fixed graph convolutional layer and the dynamic self-attention mechanism layer;

[0018] Temporal Informer: Use a temporal Informer model based on a sparse self-attention mechanism to extract temporal information;

[0019] Prediction layer: Use two convolutional layers to make predictions according to the spatio-temporal features output by the last spatio-temporal module to obtain the prediction result of the roadside berth occupancy rate.

[0020] Preferably, the input of the convolutional layer in the input layer is with dimensions M, N, C, where M represents the height, N represents the width, represents the set of real numbers, C represents the number of channels, and after the convolution operation, is mapped to Indicates the number of output channels.

[0021] Preferably, the extracting of spatial information by using a spatiotemporal position embedding layer, a fixed graph convolution layer, and a dynamic self-attention mechanism layer, and fusing the spatial features obtained from the fixed graph convolution layer and the dynamic self-attention mechanism layer by using a gating mechanism specifically include:

[0022] The spatiotemporal position embedding layer is used to aggregate the input layer features. Perform spatiotemporal position embedding, where the spatial and temporal position embeddings are and Considering the connectivity and distance factors between nodes, the graph adjacency matrix is ​​initialized Flatten along the spatial axis and time axis, we get and and with Connect and get the embedded features, expressed as F c represents the convolutional layer;

[0023] Use a fixed graph convolutional layer to extract the adjacency matrix and Extract spatial features from the image The diagonal elements of the degree matrix D are defined as ii =∑ i A ij , i = 1, 2, ..., N, the symmetric normalized Laplace matrix L is defined as L = I n -D -1 / 2 AD -1 / 2 , scale the Laplacian matrix, and the scaled Laplacian matrix is where λ max represents the maximum eigenvalue of the Laplacian matrix, I n Represents the unit matrix, for the embedded feature Graph convolution using Chebyshev polynomial approximation to obtain node features The formula is:

[0024]

[0025] in, express The jth channel of ij,k represents the learned weight, T k is a Chebyshev polynomial of order k, where k represents the highest order of the Chebyshev polynomial, k = 1,…,K;

[0026] Use the dynamic self-attention mechanism layer to extract spatial information and embed the features of each time step Projection based on Train three subspaces for each node and The formula is:

[0027]

[0028] Wherein, respectively represent the weight matrices of Q S , K S , V S . The dynamic spatial dependence relationship between nodes is calculated through the dot product of Q S and K S . The spatial dependence is normalized using the Softmax function, controlling the range of the dot product, and the node features are updated through the attention weight matrix S . The formula is: S The formula is:

[0029]

[0030] A shared three-layer feedforward neural network and activation function are used to update the node features M S . The formula is:

[0031]

[0032] Wherein, represents the residual connection, W0, W1, and W2 represent the parameter matrices of the three layers, and U S represents the spatial features after passing through the three-layer neural network and activation function;

[0033] A gating mechanism is used to fuse the spatial information obtained from the fixed graph convolutional layer and the dynamic self-attention mechanism layer. The formula is:

[0034]

[0035] Wherein, f S and represent linear projections, respectively transforming and into one-dimensional vectors, represents the spatial features obtained from the fixed graph convolutional layer;

[0036] The output Y S is weighted by Y′ S and X S . The formula is:

[0037]

[0038] In the formula, the output As the input of the Time Informer.

[0039] Preferably, using the Time Informer model based on the sparse self-attention mechanism to extract time information includes:

[0040] Concatenate the input features with the time embedding D T to obtain G t Denote the 1×1 convolutional layer, which generates a dimensional vector for each node at each time step, X′ T represents the feature representation after time embedding;

[0041] Adopt the probabilistic sparse self-attention mechanism to model the time dependence, and construct three latent subspaces, namely and The formula is:

[0042]

[0043] where, represents the input sequence, represents the weight matrix;

[0044] The formula for the self-attention mechanism modeling is:

[0045]

[0046] In the formula, q i represents the i-th query vector, k j represents the j-th key vector, v j represents the j-th value vector, represents the calculation of the value vector v i related to the key vector k j under the condition of the given query vector q j ;

[0047] The probabilistic sparsity criterion measures the gap between p(k j |q i ) and the uniform distribution where L k represents the length of Q T , and the sparsity metric formula is:

[0048]

[0049] In the formula, K represents the key matrix, represents the embedding dimension, is the dot product of the query vector and the key vector, used to measure the similarity between them;

[0050] Calculate the filtered sparse matrix according to the sparsity metric formula According to the sparse matrix Calculate the time attention, and the formula is:

[0051]

[0052] Introduce a distillation mechanism to assign higher weights to the data features of important nodes in time attention, and the formula is:

[0053] Y T = MaxPool(ELU(Conv(S T )));

[0054] In the formula, MaxPool represents selecting the maximum value within a fixed receptive field and generating a new matrix with reduced dimensions, ELU represents the activation function, and Y T represents the output after the distillation mechanism.

[0055] Preferably, the spatio-temporal layer is composed of stacked spatio-temporal Transformer-Informer modules to capture dynamic spatio-temporal dependencies, including: integrating the spatial Transformer module and the time Informer module through a residual connection method to form a spatio-temporal module with spatio-temporal feature processing capabilities; subsequently, increasing the depth of the neural network by stacking spatio-temporal Transformer-Informer modules to form a spatio-temporal layer capable of efficiently capturing spatio-temporal dependency relationships.

[0056] Preferably, the process of the spatio-temporal layer capturing spatio-temporal dependency relationships includes:

[0057] In the l-th spatio-temporal module, the spatial Transformer module extracts spatial features from and the adjacency matrix represents the output of the (l - 1)-th spatio-temporal module, and the time Informer combines the spatial features and the input features The formula is:

[0058]

[0059] In the formula, represents the output of the time Informer module;

[0060] The spatio-temporal layer includes multiple spatio-temporal modules. After being processed by the spatio-temporal layer, the feature representation is

[0061] Preferably, the input of the last spatio-temporal module is the d-dimensional spatio-temporal features of the outputs of N nodes at the last time step. ST Perform multi-step prediction on the future T traffic conditions through N nodes to obtain the prediction result. The formula is as follows

[0062] Y = Conv(Conv(X ST ))

[0063] Therefore, the present invention adopts the above-mentioned spatio-temporal prediction method for roadside berth occupancy rate integrating sparse self-attention, fully considering the correlation of parking demand in time and space, realizing accurate prediction of roadside berth occupancy rate. By predicting the berth occupancy situation, the urban traffic management department can better conduct parking guidance and parking management, further improve the berth use efficiency, and provide a more convenient parking experience for citizens.

[0064] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0065] Figure 1 It is a method framework diagram of an embodiment of the present invention. Detailed Embodiments

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0067] Embodiment

[0068] Referring to Figure 1 , the present invention provides a spatio-temporal prediction method for roadside berth occupancy rate integrating sparse self-attention, and the steps include:

[0069] Obtain 70 roadside berths in a certain city, and collect a total of 51,585 parking data with a time span from August 1, 2019 to August 6, 2019.

[0070] S1. Obtain the historical parking data of 70 roadside berths in a certain city, preprocess the historical parking data, including extracting the in and out times of each vehicle and the berth location information, and clean the data to remove abnormal data, forming a dataset including the vehicle entry time into the berth (IN_TIME), the departure time from the berth (OUT_TIME), and the berth location (POSITION_NAME). An example of the data is shown in Table 1.

[0071] Table 1 Example of parking data

[0072]

[0073]

[0074] According to the in and out times of the vehicles, with a 5-minute time interval, and based on the in and out times of individual vehicles in each roadside berth area, count the parking quantities in different time periods. Calculate the berth occupancy rates of each berth area in different time periods according to the counted parking quantities and the total number of berths in each berth area, and obtain a dataset of berth occupancy rates.

[0075] S2. Regard the urban berth area network as a weighted directed graph Construct the adjacency matrix of the graph. Among them, the weighted directed graph where V is the set of vertices, E is the set of edges, and A is the adjacency matrix of the graph;

[0076] Constructing the adjacency matrix of the graph includes:

[0077] S21. By accessing the map API, obtain the longitude and latitude of the roadside berth areas, as shown in Table 2. Calculate the driving distances between each roadside berth area according to the longitude and latitude of the roadside berth areas, and generate a distance matrix, as shown in Table 3.

[0078] Table 2 Longitude and latitude of roadside berths

[0079]

[0080]

[0081] Table 3 Distance matrix

[0082]

[0083] S22. According to the generated distance matrix, apply the Gaussian kernel function to each element of the distance matrix to convert the distance into similarity. The form of the Gaussian kernel is:

[0084]

[0085] where, dis(x i →y j) represents the distance between two points in the distance matrix, and σ represents the bandwidth parameter of the Gaussian kernel, which is determined by the global standard deviation. The Gaussian kernel makes points closer to each other have higher weights.

[0086] S3. Construct a parking berth occupancy rate prediction model, input the information obtained in steps S1 and S2, and based on the parking berth occupancy rate prediction model, predict the parking berth occupancy rate to obtain the predicted result of the roadside berth occupancy rate. For example, input the historical occupancy rate data of 12 time steps in a total of 70 roadside berth areas during the period from 9:00 to 10:00 on August 6th, and the berth occupancy rate for the next 5 minutes (10:00 - 10:05) can be predicted. The result is shown in Table 4:

[0087] Table 4 Predicted Results of Berth Occupancy Rate

[0088] Berth area Occupancy rate Berth area Occupancy rate 1 75.95% 36 11.90% 2 56.63% 37 27.87% 3 14.78% 38 52.94% 4 21.11% 39 25.76% 5 29.34% 40 35.29% 6 17.88% 41 24.84% 7 66.42% 42 16.95% 8 19.28% 43 8.91% 9 34.80% 44 6.45% 10 12.86% 45 36.59% 11 18.37% 46 13.73% 12 8.85% 47 64.52% 13 37.18% 48 7.25% 14 34.09% 49 37.17% 15 27.27% 50 32.47% 16 55.83% 51 16.43% 17 9.64% 52 10.00% 18 9.80% 53 16.19% 19 17.14% 54 7.35% 20 13.94% 55 16.00% 21 25.45% 56 4.03% 22 18.95% 57 36.47% 23 60.87% 58 7.59% 24 37.72% 59 12.22% 25 10.29% 60 17.48% 26 32.35% 61 21.37% 27 13.33% 62 3.85% 28 22.50% 63 15.38% 29 8.45% 64 9.14% 30 52.94% 65 75.00% 31 13.24% 66 17.03% 32 37.36% 67 27.27% 33 15.79% 68 32.47% 34 18.18% 69 62.65% 35 10.53% 70 26.09%

[0089] The parking berth occupancy rate prediction model includes:

[0090] Input layer: Use a 1×1 convolutional layer to aggregate the input data and expand the data channel dimension. The 1×1 convolutional layer is a special type of convolutional operation in a convolutional neural network. It is usually used to change the number of channels of the feature map while keeping the spatial dimensions (height and width) of the feature map unchanged. The specific process of the convolutional layer is as follows:

[0091] The input of the 1×1 convolutional layer is with dimensions M, N, C, where M represents the height, N represents the width, and C represents the number of channels, represents the set of real numbers, that is, the depth of the feature map;

[0092] The convolutional kernel used in the 1×1 convolutional operation is a weight matrix of size (1, 1, C in , C out ), where: 1 represents the height of the convolutional kernel is 1, 1 represents the width of the convolutional kernel is 1, C in represents the number of input channels, and C out represents the number of output channels. Each weight in the convolutional kernel performs a convolutional operation with the corresponding position in the input tensor;

[0093] After the 1×1 convolutional operation, is mapped to that is, the number of output channels is

[0094] Spatio-temporal layer: It includes a spatial Transformer and a temporal Informer. The spatio-temporal layer is constructed by stacking spatio-temporal Transformer-Informer modules to capture dynamic spatio-temporal dependencies.

[0095] Spatial Transformer: It extracts spatial information by using a spatio-temporal position embedding layer, a fixed graph convolutional layer, and a dynamic self-attention mechanism layer, and uses a gating mechanism to fuse the spatial features obtained from the fixed graph convolutional layer and the dynamic self-attention mechanism layer, specifically including:

[0096] Use the spatio-temporal position embedding layer to perform spatio-temporal position embedding after the feature aggregation of the input layer where the spatial and temporal position embeddings are respectively and Considering the connectivity and distance factors between nodes, initialize with a graph adjacency matrix to model the spatial dependency relationship, where, The initialization of uses one-hot encoding. Tiled along the spatial axis and the temporal axis to obtain and and connect with to obtain the embedded feature, denoted as F c represents a 1×1 convolutional layer, which is used to convert its embedded feature into a -dimensional vector for each node at each time step.

[0097] Use the fixed graph convolutional layer to extract spatial features from the adjacency matrix and Define the diagonal element of the degree matrix D of the graph as D ii =∑ i A ij , i = 1, 2, …, N, and the symmetric normalized Laplacian matrix L is defined as L = I n - D -1 / 2 AD -1 / 2 , scale the Laplacian matrix, and the scaled Laplacian matrix is where λ max represents the largest eigenvalue of the Laplacian matrix, I n represents the identity matrix, and for the embedded feature Use the graph convolution approximated by Chebyshev polynomials to obtain the node feature The formula is:

[0098]

[0099] where, represents the j-th channel of, θ ij,k represents the learned weight, T k is the Chebyshev polynomial representing the k-th order, k represents the highest order of the Chebyshev polynomial, k = 1, …, K;

[0100] Use the dynamic self-attention mechanism layer to extract spatial information and embed the features of each time step Projection based on Train three subspaces for each node and The formula is:

[0101]

[0102] in, Respectively represent Q S , K S , V S The weight matrix is ​​obtained by Q S With K S Dynamic spatial dependencies between nodes in the dot product calculation The Softmax function is used to normalize the spatial dependency. Control the range of the dot product and make the node features Through the attention weight matrix S S To update, the formula is:

[0103]

[0104] A shared three-layer feedforward neural network and activation function are used. For node feature M S To update, the formula is:

[0105]

[0106] in, represents the residual connection, W0, W1, W2 represent the parameter matrix of the three layers, U S Represents the spatial features after three layers of neural network and activation function;

[0107] The gating mechanism is used to fuse the spatial information obtained from the fixed graph convolution layer and the dynamic self-attention mechanism layer. The formula is:

[0108]

[0109] in, f S and represents linear projection, respectively and Convert to a one-dimensional vector, Represents the spatial features obtained by the fixed graph convolution layer;

[0110] Through Y' S and X S Weighted with g to get the output Y S , the formula is:

[0111]

[0112] In the formula, the output is used as the input to the Time Informer.

[0113] Time Informer: The Time Informer model based on the sparse self-attention mechanism is used to extract time information and capture long-term dependencies. The Informer introduces an attention mechanism and a global feature capture mechanism, which can effectively capture the long-term dependencies in the sequence. Specifically, it includes:

[0114] The input feature is concatenated with the time embedding D T to obtain G t denotes a 1×1 convolutional layer, which generates a dimensional vector for each node at each time step. X′ T represents the feature representation after time embedding;

[0115] The probabilistic sparse self-attention mechanism is used to model the time dependence, and three potential subspaces are constructed, namely and The formula is:

[0116]

[0117] Among them, represents the input sequence, represents the weight matrix;

[0118] The traditional self-attention mechanism uses dot-product to calculate the time dependence. Based on the Transformer model, the Informer proposes ProbSparse Self-attention, where each key vector only processes a limited number of dominant query vectors, significantly reducing the number of query vectors to be processed and effectively reducing the time complexity and space complexity of the calculation. In long-term sequence prediction, Q T has sparsity, that is, there is redundant computational effort in the original self-attention calculation.

[0119] The initial self-attention mechanism modeling formula is:

[0120]

[0121] In the formula, q i represents the i-th query vector, k j represents the j-th key vector, v jrepresents the j-th value vector, represents that given the query vector q i under the condition, for the key vector k j associated value vector v j is calculated;

[0122] The probability sparsity criterion measures the p(k j |q i ) and the uniform distribution gap, where L k represents the length of Q T The greater the difference between the two distributions, the greater the KL divergence. The KL formula is:

[0123]

[0124] After removing the constant, the sparsity measure is as follows:

[0125]

[0126] However, the above formula has a large computational cost. To reduce the computational cost and complexity, the sparsity measure can be approximately simplified to the following formula:

[0127]

[0128] In the formula, K represents the key matrix, represents the embedding dimension, is the dot product of the query vector and the key vector, used to measure the similarity between them;

[0129] The filtered sparse matrix is calculated according to the sparsity measure formula According to the sparse matrix The temporal attention is calculated, and the formula is:

[0130]

[0131] The multi-head probability sparse self-attention mechanism will cause redundancy in feature mapping. Therefore, the Informer model introduces a distillation mechanism to assign higher weights to the data features of important nodes and is beneficial to extracting more key features in the next layer. The distillation mechanism uses one-dimensional convolution and max pooling to downsample the input temporal dimension, and the formula is as follows:

[0132] Y T = MaxPool(ELU(Conv(S T )));

[0133] In the formula, MaxPool represents selecting the maximum value within a fixed receptive field and generating a new matrix with reduced dimensions, ELU represents the activation function, Y TRepresents the output after the distillation mechanism.

[0134] The spatio-temporal layer is constructed by stacking spatio-temporal Transformer-Informer modules to capture dynamic spatio-temporal dependencies, including: integrating the spatial Transformer module and the temporal Informer module through residual connections to form a spatio-temporal module with spatio-temporal feature processing capabilities; subsequently, increasing the depth of the neural network by stacking spatio-temporal Transformer-Informer modules to form a spatio-temporal layer capable of efficiently capturing spatio-temporal dependency relationships. The specific steps are as follows:

[0135] In the l-th spatio-temporal module, the spatial Transformer module extracts spatial features from and the adjacency matrix and serves as the input to the temporal Informer. The input of the temporal Informer is composed of and merged, and the formula is:

[0136]

[0137] where the input of the l-th spatio-temporal module is the output of the (l - 1)-th spatio-temporal module, and

[0138] is the output of the temporal Informer module. The spatio-temporal layer is composed of multiple spatio-temporal modules, and these modules effectively capture global spatio-temporal features by increasing the depth of the neural network. After being processed by the spatio-temporal layer, the resulting feature representation is

[0139] Prediction layer: Use two 1×1 convolutional layers to make predictions based on the spatio-temporal features output by the last spatio-temporal module to obtain the prediction result of the roadside berth occupancy rate. ST d-dimensional spatio-temporal features through N nodes to make multi-step predictions of the future T traffic conditions to obtain the prediction result The formula is

[0140] Y = Conv(Conv(X ST ))

[0141] To verify the effectiveness of the method of the present invention, the method of the present invention was compared with 9 baseline models including HA, ARIMA, LSTM, GRU, Informer, T-GCN, STGCN, GraphWavenet, and STTN. The results are shown in Table 5. It can be seen from Table 5 that the method proposed by the present invention shows the highest prediction accuracy in the predictions of 5 min, 15 min, 30 min, and 60 min.

[0142] Table 5 Comparison results of the method of the present invention and the baseline models under different prediction durations

[0143]

[0144] Therefore, the present invention adopts the above-mentioned spatio-temporal prediction method for roadside berth occupancy rate integrating sparse self-attention, fully considering the correlation of parking demand in time and space, realizing the accurate prediction of roadside berth occupancy rate. By predicting the berth occupancy situation, the urban traffic management department can better conduct parking guidance and parking management, further improve the berth use efficiency, and provide a more convenient parking experience for citizens.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention, characterized in that the steps include: S1. Obtain historical parking data, pre-process the historical parking data, count the number of parking spaces in different time periods according to the entry and exit time of a single vehicle in each roadside parking area based on the historical parking data, calculate the parking space occupancy rate of each parking space area in different time periods based on the counted parking space and the total number of parking spaces in each parking space area, and obtain a parking space occupancy rate data set; S2. Treat the urban parking area network as a weighted directed graph Construct the adjacency matrix of the graph; S3. Construct a parking space occupancy prediction model, input the information obtained in step S1 and step S2, perform parking space occupancy prediction based on the parking space occupancy prediction model, and obtain a roadside parking space occupancy prediction result.

2. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 1 is characterized by: In step S2, the weighted directed graph Where V is the set of vertices, E is the set of edges, and A is the adjacency matrix of the graph. The construction of the adjacency matrix of the graph includes: S21. Accessing a map API to obtain the longitude and latitude of a roadside parking area, calculating the driving distance between each roadside parking area according to the longitude and latitude of the roadside parking area, and generating a distance matrix; S22. According to the generated distance matrix, a Gaussian kernel function is applied to each element of the distance matrix to convert the distance into similarity. The Gaussian kernel form is: Among them, dis(x i →y j ) represents the distance between two points in the distance matrix, σ represents the bandwidth parameter of the Gaussian kernel, which is determined by the global standard deviation. The Gaussian kernel gives higher weights to points that are closer.

3. The spatiotemporal prediction method of roadside parking space occupancy rate integrating sparse self-attention according to claim 2 is characterized in that: The parking space occupancy rate prediction model in step S3 includes: Input layer: Use the convolution layer to aggregate the input data and expand the data channel dimension; Spatiotemporal layer: includes spatial Transformer and temporal Informer. The spatiotemporal layer is constructed by stacking spatiotemporal Transformer-Informer modules to capture dynamic spatiotemporal dependencies. Spatial Transformer: It uses spatiotemporal position embedding layer, fixed graph convolution layer and dynamic self-attention mechanism layer to extract spatial information, and uses gating mechanism to fuse spatial features obtained from fixed graph convolution layer and dynamic self-attention mechanism layer. Time Informer: Extract time information using the time informer model based on sparse self-attention mechanism; Prediction layer: Use two convolutional layers to make predictions based on the spatiotemporal features output by the last spatiotemporal module to obtain the roadside parking space occupancy prediction results.

4. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 3 is characterized by: The convolutional layer input of the input layer is The dimensions are M, N, C, where M represents the height and N represents the width. represents a real number set, C represents the number of channels, and after the convolution operation, is mapped to Indicates the number of output channels.

5. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 4 is characterized in that: The method of extracting spatial information by using a spatiotemporal position embedding layer, a fixed graph convolution layer, and a dynamic self-attention mechanism layer, and fusing spatial features obtained from the fixed graph convolution layer and the dynamic self-attention mechanism layer by using a gating mechanism specifically includes: The spatiotemporal position embedding layer is used to aggregate the input layer features. Perform spatiotemporal position embedding, where the spatial and temporal position embeddings are and Considering the connectivity and distance factors between nodes, the graph adjacency matrix is ​​initialized Flatten along the spatial axis and time axis, we get and and with Connect and get the embedded features, expressed as F c represents the convolutional layer; Use a fixed graph convolutional layer to extract the adjacency matrix and Extract spatial features from the image The diagonal elements of the degree matrix D are defined as ii =∑ i A ij , i = 1, 2, ..., N, the symmetric normalized Laplace matrix L is defined as L = I n -D -1 / 2 AD -1 / 2 , scale the Laplace matrix, and the scaled Laplace matrix is where λ max represents the maximum eigenvalue of the Laplacian matrix, I n Represents the unit matrix, for the embedded feature Graph convolution using Chebyshev polynomial approximation to obtain node features The formula is: in, express The jth channel of ij,k represents the learned weight, T k is a Chebyshev polynomial of order k, where k represents the highest order of the Chebyshev polynomial, k = 1,…,K; Use the dynamic self-attention mechanism layer to extract spatial information and embed the features of each time step Projection based on Train three subspaces for each node and The formula is: in, Respectively represent Q S , K S , V S The weight matrix is ​​obtained by Q S With K S Dynamic spatial dependencies between nodes in the dot product calculation The Softmax function is used to normalize the spatial dependency. Control the range of the dot product and make the node features Through the attention weight matrix S S To update, the formula is: M S =S S V S ; A shared three-layer feedforward neural network and activation function are used. For node feature M S To update, the formula is: in, represents the residual connection, W0, W1, W2 represent the parameter matrix of the three layers, U S Represents the spatial features after three layers of neural network and activation function; The gating mechanism is used to fuse the spatial information obtained from the fixed graph convolution layer and the dynamic self-attention mechanism layer. The formula is: in, f S and Represents linear projection, respectively and Convert to a one-dimensional vector, Represents the spatial features obtained by the fixed graph convolution layer; Through Y' S and X S Weighted with g to get the output Y S , the formula is: In the formula, the output As the input of the time Informer.

6. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 5 is characterized in that: The temporal information extracted using the temporal Informer model based on the sparse self-attention mechanism includes: The input features Embedded with time T Splicing G t represents a 1×1 convolutional layer, which generates a dimensional vector, X′ T Represents the feature representation after time embedding; The probabilistic sparse self-attention mechanism is used to model the temporal dependency and construct three latent subspaces, namely and The formula is: in, represents the input sequence, represents the weight matrix; The modeling formula of the self-attention mechanism is: In the formula, q i represents the i-th query vector, k j represents the jth key vector, v j represents the jth value vector, Indicates that given a query vector q i Under the condition of j The associated value vector v j Calculate The probability sparsity criterion measures p(k j |q i ) and uniform distribution The gap, where L k Indicates Q T The length of , the sparsity measurement formula is: Where K represents the bond matrix, represents the embedding dimension, is the dot product of the query vector and the key vector, which is used to measure the similarity between them; The filtered sparse matrix is ​​calculated according to the sparsity measurement formula According to the sparse matrix The temporal attention is calculated as follows: The distillation mechanism is introduced to give higher weights to the data features of important nodes of temporal attention. The formula is: Y T =MaxPool(ELU(Conv(S T )))? In the formula, MaxPool represents selecting the maximum value within a fixed receptive field and generating a new matrix with reduced dimension, ELU represents the activation function, and Y T Represents the output after the distillation mechanism.

7. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 6 is characterized by: The method of forming a spatiotemporal layer by stacking spatiotemporal Transformer-Informer modules to capture dynamic spatiotemporal dependencies includes: integrating the spatial Transformer module with the temporal Informer module through residual connections to form a spatiotemporal module with spatiotemporal feature processing capabilities; then, increasing the depth of the neural network by stacking spatiotemporal Transformer-Informer modules to form a spatiotemporal layer that can capture spatiotemporal dependencies.

8. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 7 is characterized by: The process of capturing the spatiotemporal dependency relationship in the spatiotemporal layer includes: In the lth spatiotemporal module, the spatial Transformer module is and extract spatial features Y from the adjacency matrix l S , Represents the output of the l-1th spatiotemporal module, the temporal Informer merges the spatial feature Y l S and input features The formula is: Where Y l T Represents the output of the time Informer module; The spatiotemporal layer includes multiple spatiotemporal modules. After processing by the spatiotemporal layer, the feature representation is 9. The spatiotemporal prediction method for roadside parking space occupancy rate integrating sparse self-attention according to claim 8 is characterized in that: The prediction based on the spatiotemporal features output by the last spatiotemporal module includes: the input of the last spatiotemporal module is the output d of the N nodes at the last time step. ST Dimensional spatiotemporal characteristics Through N nodes, multi-step predictions are made for the future T traffic conditions to obtain the prediction results. The formula is: Y=Conv(Conv(X ST ))。

Citation Information

Patent Citations

  • Highway traffic flow prediction method based on Transform and graph attention network

    CN116092294A

  • Multi-dimensional roadside parking abnormal occupation identification method and device

    CN116720114A

  • Ship traffic flow long sequence space-time prediction method based on ST-Informer

    CN116821784A

  • Digraph-based regional traffic prediction method, apparatus and device, and medium

    CN117058884A

  • Traffic flow prediction method based on interactive dynamic graph convolution and probability sparse attention

    CN117290707A