Urban road traffic flow detection method and system

By performing clock synchronization and spatial grid alignment on pavement video streams on urban roads, combining deep learning models to extract the spatiotemporal characteristics of traffic elements and perform dynamic semantic mapping, the problem of low data coverage and processing efficiency in the existing traffic flow detection methods is solved, and high-precision and real-time traffic flow detection are achieved.

CN120148248AInactive Publication Date: 2025-06-13张俊忠
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510428871.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing urban road traffic flow detection methods have problems such as limited data coverage, low processing efficiency and insufficient real-time performance. Especially in complex and changeable urban environments, it is difficult to meet the requirements of data accuracy and timeliness.

Method used

By obtaining the road surface video stream and performing clock synchronization and alignment with the spatial grid, spatiotemporal synchronization data is generated. Then, the traffic elements in the video stream are extracted spatiotemporally based on the deep learning model to generate traffic vector data containing velocity, direction and position attributes. Finally, the traffic vector data is dynamically mapped to generate traffic status data associated with the geographic information data layer attributes.

Benefits of technology

Real-time conversion from unstructured video data to dynamic spatiotemporal vector data is realized, the accuracy and real-time nature of traffic flow detection are improved, and bottlenecks such as inefficient data conversion and semantic faults in traditional methods are broken.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148248A_ABST
    Figure CN120148248A_ABST
Patent Text Reader

Abstract

The invention provides an urban road traffic flow detection method and system, and the method comprises the steps: obtaining a pavement video stream, carrying out the clock synchronization and space grid alignment of the pavement video stream and geographic information data, and generating time-space synchronization data; based on the space-time synchronization data, space-time feature extraction is carried out on traffic elements in the road surface video stream through a deep learning model, traffic vector data are generated, and attributes of the traffic vector data comprise speed, direction and position; and carrying out dynamic semantic mapping processing on the traffic vector data to generate traffic state data, and associating the traffic state data with the layer attribute of the geographic information data. By adopting the method, the urban road traffic flow detection precision can be improved through real-time conversion from the unstructured video data to the dynamic space-time vector data, and a solid technical support is provided for optimization of the urban traffic flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent urban transportation, and particularly relates to a method and system for detecting urban road traffic flow. Background Art

[0002] The detection of urban road traffic flow is a core research area of intelligent transportation systems, which is directly related to the operation efficiency of cities, the travel experience of residents, and the scientific nature of traffic management decisions. Its importance is self-evident. With the acceleration of urbanization and the sharp increase in the vehicle ownership, the accurate acquisition and analysis of traffic flow data have become the key to improving the utilization rate of road resources and alleviating congestion. However, the current mainstream detection methods mostly rely on traditional sensors or single video monitoring, and there are problems such as limited coverage, low data processing efficiency, and insufficient real-time performance. These limitations make it difficult for traffic management departments to comprehensively grasp the dynamic changes of the road network. Especially in complex and changeable urban environments, the accuracy and timeliness of data often cannot meet the requirements. The defects of existing methods further highlight the technical bottlenecks in this field. On the one hand, when facing large-scale and multi-dimensional traffic scenarios, traditional detection means often become distorted due to incomplete data collection; on the other hand, although image-based monitoring systems can provide rich visual information, they are limited by the unstructured nature of the original data and are difficult to be efficiently converted into structured data that can be calculated and analyzed. This low conversion efficiency from massive images to useful information has become the main obstacle to the development of traffic flow detection technology. The core challenges focus on how to extract key traffic elements from monitoring images and endow them with computable attributes. Specifically, the dynamic characteristics (such as position, speed, direction) of entities such as vehicles and pedestrians are difficult to be accurately identified and quantified. At the same time, the storage and analysis efficiency of data are limited by the complexity of processing methods. In addition, the integration of traffic states from a global perspective is also a major problem. Especially when multi-source data needs to be docked with a geographic information system, the lack of a unified data expression form makes it extremely difficult to implement spatial queries and pattern recognition.

[0003] However, the current urban road traffic flow detection methods fail to achieve seamless integration with geographic information systems, which has become a key problem hindering the improvement of the accuracy of urban road traffic flow detection. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and system for detecting urban road traffic flow in view of the above technical problems, which can improve the accuracy of urban road traffic flow detection through the real-time conversion of unstructured video data into dynamic spatio-temporal vector data, and provide solid technical support for the optimization of urban traffic flow.

[0005] In the first aspect, the present application provides a method for detecting urban road traffic flow, including:

[0006] Obtain the road surface video stream, perform clock synchronization and spatial grid alignment processing on the road surface video stream and geographical information data to generate spatio-temporal synchronization data;

[0007] Based on the spatio-temporal synchronization data, use a deep learning model to extract spatio-temporal features of traffic elements in the road surface video stream to generate traffic vector data, and the attributes of the traffic vector data include speed, direction, and position;

[0008] Perform dynamic semantic mapping processing on the traffic vector data to generate traffic state data, and the traffic state data is associated with the layer attributes of the geographical information data.

[0009] In a possible embodiment, performing clock synchronization and spatial grid alignment processing on the road surface video stream and geographical information data to generate spatio-temporal synchronization data includes:

[0010] Perform clock synchronization on the camera corresponding to the road surface video stream and the server corresponding to the geographical information data through the Precision Time Protocol to generate a synchronization signal with aligned timestamps;

[0011] Based on the synchronization signal, perform spatio-temporal slice sampling on consecutive frames of the road surface video stream to generate a spatio-temporal slice sequence;

[0012] According to the spatial coordinate system of the geographical information data, perform grid mapping processing on the spatio-temporal slice sequence to generate spatio-temporal synchronization data aligned with the geographical space.

[0013] In a possible embodiment, based on the synchronization signal, performing spatio-temporal slice sampling on consecutive frames of the road surface video stream to generate a spatio-temporal slice sequence includes:

[0014] Perform sliding time window sampling on the road surface video stream, and the window length covers consecutive multiple frames of images to generate a time-continuous frame sequence;

[0015] Perform spatial grid division on each frame of the time-continuous frame sequence to generate multi-dimensional spatio-temporal slices;

[0016] Encode the multi-dimensional spatio-temporal slices into four-dimensional tensor data including time depth, spatial dimension, and number of channels to generate a spatio-temporal slice sequence.

[0017] In a possible embodiment, based on the spatio-temporal synchronization data, using a deep learning model to extract spatio-temporal features of traffic elements in the road surface video stream to generate traffic vector data includes:

[0018] Use a neural network to extract spatio-temporal features of the spatio-temporal synchronization data to generate a high-dimensional spatio-temporal feature tensor;

[0019] Perform a linear transformation on the high-dimensional spatio-temporal feature tensor through a weight matrix obtained by pre-training to generate initial traffic vector data;

[0020] Perform trajectory verification and outlier removal on the initial traffic vector data based on motion continuity constraints to generate traffic vector data.

[0021] In a possible embodiment, perform spatio-temporal feature extraction on the spatio-temporal synchronization data through a neural network to generate a high-dimensional spatio-temporal feature tensor, including:

[0022] Perform 3D convolution operations on the spatio-temporal synchronization data to extract spatio-temporal local features;

[0023] Fuse spatio-temporal local features at different levels through residual connections to generate a fused feature;

[0024] Perform channel attention weighting processing on the fused feature to generate a high-dimensional spatio-temporal feature tensor.

[0025] In a possible embodiment, perform dynamic semantic mapping processing on the traffic vector data to generate traffic state data, including:

[0026] Construct dynamic semantic mapping rules according to the layer attribute fields of the geographic information data;

[0027] Based on the dynamic semantic mapping rules, associate and map speed, direction, and position with the lane speed limit, topological relationship, and coordinate nodes of the geographic information data respectively to generate intermediate data;

[0028] Perform coordinate transformation error calibration processing on the intermediate data to generate calibrated traffic state data.

[0029] In a possible embodiment, based on the dynamic semantic mapping rules, associate and map speed, direction, and position with the lane speed limit, topological relationship, and coordinate nodes of the geographic information data respectively to generate intermediate data, including:

[0030] Calculate the normalized ratio of the speed to the lane speed limit to generate a speed compliance index, and write the speed compliance index into the intermediate data;

[0031] Calculate the motion direction angle based on the direction, and perform vector angle matching between the motion direction angle and the lane topological relationship of the geographic information data to generate a lane attribution identifier, and write the lane attribution identifier into the intermediate data;

[0032] Perform the nearest neighbor association between the position and the road network nodes of the geographic information data to generate a spatial positioning label, and write the spatial positioning label into the intermediate data.

[0033] In a second aspect, the present application also provides an urban road traffic flow detection system, including:

[0034] A data preprocessing module, configured to obtain a road surface video stream, perform clock synchronization and spatial grid alignment processing on the road surface video stream and geographical information data, and generate spatio-temporal synchronization data;

[0035] A feature extraction module, configured to perform spatio-temporal feature extraction on traffic elements in the road surface video stream based on the spatio-temporal synchronization data through a deep learning model, and generate traffic vector data, where the attributes of the traffic vector data include speed, direction, and position;

[0036] A traffic detection module, configured to perform dynamic semantic mapping processing on the traffic vector data to generate traffic state data, and the traffic state data is associated with the layer attributes of the geographical information data.

[0037] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above-mentioned urban road traffic flow detection method is implemented.

[0038] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned urban road traffic flow detection method is implemented.

[0039] The above-mentioned urban road traffic flow detection method and system, by obtaining a road surface video stream and performing clock synchronization and spatial grid alignment processing on it to generate spatio-temporal synchronization data; based on the spatio-temporal synchronization data, using a deep learning model to perform spatio-temporal feature extraction on traffic elements in the road surface video stream to generate traffic vector data including speed, direction, and position attributes; by performing dynamic semantic mapping processing on the traffic vector data, generating traffic state data associated with the layer attributes of the geographical information data. The above technical solution realizes the real-time conversion from unstructured video data to dynamic spatio-temporal vector data, can break through the bottlenecks such as low data conversion efficiency and semantic discontinuity in traditional methods, improve the traffic flow detection accuracy, and provide high-confidence data support for traffic optimization. Description of the Drawings

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0041] Figure 1 It is a flowchart of an urban road traffic flow detection method provided by an embodiment of the present invention;

[0042] Figure 2Schematic diagram of a structure of an urban road traffic flow detection system provided by an embodiment of the present invention. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] First, a brief introduction is made to the nouns involved in the embodiments of the present application.

[0045] Geographic information data refers to a collection of various information related to geographical locations. It has important value in fields such as urban planning, traffic management, and resource allocation. In an urban environment, geographic information data usually includes the detailed layout of the urban road network, the attributes of different road sections (such as the number of lanes, road type, speed limit information, etc.), the locations and control logics of traffic lights, and the distribution of various public facilities (such as bus stops, parking lots, gas stations, etc.). These data not only provide a basic geographical framework for urban traffic flow detection, but also can be combined with real-time traffic data to achieve accurate analysis and visual presentation of traffic states. By associating traffic vector data with the layer attributes of geographic information data, the dynamic changes of traffic flow can be more intuitively displayed, providing strong support for the optimized management of urban traffic.

[0046] Semantic mapping is a process of converting low-level data information into high-level semantic concepts. By establishing a mapping relationship between data and semantics, machines can understand the actual meaning expressed by the data. In the field of urban traffic, semantic mapping can convert traffic vector data (such as the speed, direction, and location of vehicles) into more semantic traffic state information, such as "congested", "smooth", "accident", etc. This conversion not only facilitates traffic managers to intuitively understand the current traffic conditions, but also provides a more accurate basis for formulating traffic optimization strategies, thus realizing intelligent traffic management from data-driven to semantic-driven.

[0047] 3D convolution operation is an extended operation in a convolutional neural network (CNN). It adds a time dimension or a depth dimension on the basis of traditional two-dimensional convolution, so as to be able to process data with a three-dimensional structure. In three-dimensional space, 3D convolution extracts local features of the input data by sliding the convolution kernel in space and time (or depth). This operation is particularly suitable for processing video data (where the time dimension is crucial) or three-dimensional images (such as voxel data in medical imaging), and can capture complex patterns and correlations in space and time of the data. Through 3D convolution, the network can learn richer feature representations, thereby improving the recognition and understanding capabilities for dynamic scenes or three-dimensional structures.

[0048] According to the above-mentioned glossary, the implementation environment of a method for detecting urban road traffic flow provided by an embodiment of the present application is described. Schematically, the implementation environment includes: an image sensor, a terminal, and a processor. Among them, the processor is signal-connected to the image sensor and the storage device through a network; the image sensor can be a high-definition camera, a low-light camera, an infrared camera, etc.; the processor includes, but is not limited to, a central processing unit, a multi-core processor, an artificial intelligence chip, etc., which are not limited here.

[0049] Combined with the above-mentioned glossary and implementation environment, the application scenarios of the embodiments of the present application are described. The method for detecting urban road traffic flow provided in the embodiments of the present application can be applied to the following scenarios, including but not limited to:

[0050] In urban traffic management, the above technical solution can monitor the traffic flow of urban main roads and key intersections in real time. Through in-depth learning analysis of the video stream, congested areas can be quickly identified, and this information can be fed back to the traffic management department in real time. The management department can adjust the signal light duration in a timely manner based on the traffic status data, guide vehicle diversion, thereby effectively alleviating congestion and optimizing traffic flow.

[0051] In the smart city traffic platform, the city-level traffic management platform integrates the detection data of multiple intersections. Through the overlay analysis of the spatio-temporal vector data (such as vehicle trajectories, congestion indexes) generated by this technology and the road network of geographic information data, frequently-occurring congestion nodes and bottleneck sections can be identified. For example, if it is found that the left-turn vehicles are severely congested around a business district during the weekend evening rush hour, tidal lanes can be dynamically added or the signal light linkage plan can be optimized. Long-term data can also be used to evaluate the effect of road reconstruction and provide a quantitative basis for urban planning.

[0052] In the scenario of autonomous driving vehicle-road collaboration, this technology converts real-time traffic flow data (such as vehicle positions, speeds) into a standardized vector format and broadcasts it to autonomous driving vehicles through a 5G network. The vehicle combines high-precision maps to predict the traffic conflict points at the upcoming intersections and adjusts the vehicle speed or lane-changing strategy in advance. For example, when it is detected that the lateral approaching vehicle speed at an intersection is too fast, an avoidance instruction can be sent to the autonomous driving vehicle to improve the traffic safety under complex road conditions.

[0053] Schematically, the method for detecting urban road traffic flow provided by the embodiments of the present application can also be applied to other application scenarios. Only examples are given here, and the specific application scenarios are not limited.

[0054] In one embodiment, as Figure 1As shown, a method for detecting urban road traffic flow is provided. In this embodiment, the method is exemplified by being applied to a terminal. It can be understood that the method can also be applied to a server and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The method includes:

[0055] Step 101: Obtain a road surface video stream, perform clock synchronization and spatial grid alignment processing on the road surface video stream and geographic information data, and generate spatio-temporal synchronization data.

[0056] Specifically, obtain a road surface video stream through a camera installed on an urban road, and perform clock synchronization and spatial grid alignment processing on the obtained video stream and geographic information data to ensure that the time and spatial information of the video data are consistent with the geographic information data, thereby generating spatio-temporal synchronization data, eliminating the spatio-temporal reference difference between the video data and the geographic data, solving the trajectory drift problem caused by time asynchronization or coordinate offset in the traditional method, and providing a high-precision spatio-temporal reference for subsequent analysis.

[0057] Step 102: Based on the spatio-temporal synchronization data, extract spatio-temporal features of traffic elements in the road surface video stream through a deep learning model, and generate traffic vector data. The attributes of the traffic vector data include speed, direction, and position.

[0058] Specifically, use a deep learning model to extract spatio-temporal features of traffic elements in the road surface video stream. The model can automatically identify and extract key attributes such as the speed, direction, and position of traffic elements, and generate structured traffic vector data. Among them, the deep learning model is trained through a large amount of data and has a powerful feature extraction ability, which can adapt to complex traffic scenarios and improve the accuracy of traffic flow detection.

[0059] Step 103: Perform dynamic semantic mapping processing on the traffic vector data to generate traffic state data, and the traffic state data is associated with the layer attributes of the geographic information data.

[0060] Specifically, perform dynamic semantic mapping processing on the extracted traffic vector data. By associating the traffic vector data with the layer attributes of the geographic information data, intuitive traffic state data is generated to improve the intelligent level of traffic flow detection.

[0061] The above-mentioned urban road traffic flow detection method obtains the road surface video stream, performs clock synchronization and spatial grid alignment processing on it to generate spatio-temporal synchronization data; based on the spatio-temporal synchronization data, uses a deep learning model to extract spatio-temporal features of traffic elements in the road surface video stream to generate traffic vector data including speed, direction, and position attributes; through dynamic semantic mapping processing of the traffic vector data, traffic state data associated with the attributes of the geographic information data layer is generated. The above technical solution realizes the real-time conversion from unstructured video data to dynamic spatio-temporal vector data, can break through the bottlenecks such as low data conversion efficiency and semantic discontinuity in traditional methods, improve the traffic flow detection accuracy, and provide high-confidence data support for traffic optimization.

[0062] In a possible embodiment, performing clock synchronization and spatial grid alignment processing on the road surface video stream and the geographic information data to generate spatio-temporal synchronization data may include:

[0063] Step 201, perform clock synchronization on the camera corresponding to the road surface video stream and the server corresponding to the geographic information data through the Precision Time Protocol to generate a synchronization signal with aligned timestamps.

[0064] Exemplarily, synchronization messages are transmitted between the camera and the server through an Ethernet or fiber optic communication link, the network transmission delay is calculated and compensated. After synchronization is completed, a synchronization signal with aligned timestamps is generated as the unified time reference for subsequent data processing to eliminate the time deviation between the camera and the geographic information data server, solve the problem of trajectory misalignment caused by clock asynchronization, and provide a high-precision time alignment basis for multi-source data fusion.

[0065] Step 202, perform spatio-temporal slice sampling on consecutive frames of the road surface video stream based on the synchronization signal to generate a spatio-temporal slice sequence.

[0066] Specifically, a sliding time window is used to segment and intercept the video stream to generate a spatio-temporal slice sequence. Each spatio-temporal slice contains consecutive video frames in the time dimension and the complete road area image in the spatial dimension. Through spatio-temporal slice sampling, the dynamic motion information of traffic elements (such as vehicle acceleration, steering trend) can be retained, while reducing the data redundancy of single-frame processing, improving the subsequent feature extraction efficiency, and adapting to the dynamic change requirements of urban traffic scenarios.

[0067] Step 203, perform grid mapping processing on the spatio-temporal slice sequence according to the spatial coordinate system of the geographic information data to generate spatio-temporal synchronization data aligned with the geographic space.

[0068] Specifically, the pixel coordinates of each spatio-temporal slice are converted into geographical coordinates (such as WGS-84, CGCS2000 or local coordinate system) through an affine transformation matrix, and the geographical space grid cells are divided according to a preset resolution to ensure that each grid corresponds to a fixed geographical scale, generating spatio-temporal synchronous data aligned in geographical space. Each data unit contains a timestamp, geographical coordinates and the original video frame information, so as to solve the problem of inconsistent spatial scales between video data and geographical data, eliminate the vehicle positioning error caused by coordinate system offset, and provide fusion data with consistent spatial reference for traffic flow detection.

[0069] In a possible embodiment, spatio-temporal slice sampling is performed on the road surface video stream based on a synchronization signal to generate a spatio-temporal slice sequence, including:

[0070] Step 301, perform sliding time window sampling on the road surface video stream, where the window length covers multiple consecutive frames of images, generating a temporally continuous frame sequence.

[0071] Specifically, by setting a sliding window that covers multiple consecutive frames of images and moving the window frame by frame to generate a temporally continuous frame sequence, the length of the sliding time window can be adjusted according to actual needs. For example, different lengths of sliding time windows are set during peak hours and off-peak hours to ensure that the sampled data can fully reflect the dynamic changes of traffic elements (such as vehicles and pedestrians), retain the motion correlation of traffic elements in the time dimension, avoid information loss caused by single-frame analysis, and provide dynamically coherent input data for subsequent spatio-temporal feature extraction.

[0072] Step 302, perform spatial grid division on each frame of the temporally continuous frame sequence to generate multi-dimensional spatio-temporal slices.

[0073] Specifically, each frame of image is subdivided into grid cells, and each grid cell corresponds to a fixed area of the actual road space. By superimposing the grid division results of all frames, multi-dimensional spatio-temporal slices are generated. Each slice contains continuous frames in the time dimension and grid-like image data in the spatial dimension. Spatial grid division enhances the local feature extraction ability, reduces redundant calculations (such as ignoring vehicle-free areas), and at the same time provides a structured data carrier for spatio-temporal feature fusion, improving the processing efficiency.

[0074] Step 303, encode the multi-dimensional spatio-temporal slices into four-dimensional tensor data including time depth, spatial dimension and number of channels, generating a spatio-temporal slice sequence.

[0075] Exemplarily, the time depth (e.g., 5 frames), spatial dimension (e.g., a grid of 32×32 pixels), and number of RGB channels of each spatio-temporal slice can be integrated into a four-dimensional tensor, whose dimensions are represented as (time depth × height × width × channels). By tensor encoding to unify the data format and generate a standardized spatio-temporal slice sequence, the above method encapsulates spatio-temporal information in tensor form, which can support efficient storage and parallel computing (such as GPU acceleration), reduce the complexity of subsequent model processing, and at the same time ensure the integrity of the data spatio-temporal structure and avoid feature distortion.

[0076] In a possible embodiment, based on spatio-temporal synchronized data, spatio-temporal feature extraction is performed on traffic elements in a road surface video stream through a deep learning model to generate traffic vector data, which may include:

[0077] Step 401, perform spatio-temporal feature extraction on the spatio-temporal synchronized data through a neural network to generate a high-dimensional spatio-temporal feature tensor.

[0078] Specifically, performing spatio-temporal feature extraction on the spatio-temporal synchronized data through a neural network can generate a high-dimensional spatio-temporal feature tensor. The powerful learning ability of the neural network enables it to automatically identify the complex features of traffic elements and extract key information such as speed, direction, and position, providing a rich feature basis for subsequent traffic flow detection.

[0079] Step 402, perform a linear transformation on the high-dimensional spatio-temporal feature tensor through a weight matrix obtained by pre-training to generate initial traffic vector data.

[0080] Specifically, flatten the feature tensor into a vector in the spatial dimension and perform matrix multiplication with the weight matrix to generate initial traffic vector data. Each vector data represents a set of attributes of a traffic element. Optionally, the process of obtaining the pre-trained weight matrix is as follows: On a large-scale traffic video dataset containing labeled traffic elements (such as vehicle position, speed, direction), train a 3D convolutional neural network so that it can predict target attributes from spatio-temporal data; optimize the network parameters through the backpropagation algorithm, and fix the weight matrix of the fully connected layer as the pre-trained weight. This matrix encapsulates the mapping relationship from spatio-temporal features to vector attributes and can be directly transferred to a new scenario for linear transformation. Using the pre-trained weight to achieve an efficient mapping from features to vectors can avoid manually designing feature engineering, improve the generalization ability of the algorithm, and at the same time reduce the computational complexity.

[0081] Step 403, perform trajectory verification and outlier removal on the initial traffic vector data based on the motion continuity constraint to generate traffic vector data.

[0082] Specifically, the position of traffic elements at the next moment is predicted through Kalman filtering or a kinematic model. If the deviation between the initial vector data and the predicted value exceeds the threshold, it is determined as an abnormal point and removed, and the vector data that conforms to the motion law is retained to generate an optimized traffic vector data set. Through the above method, abnormal trajectory points caused by factors such as sensor noise and occlusion can be eliminated, the detection accuracy and data reliability can be improved, and the accuracy and credibility of subsequent analysis results can be ensured.

[0083] In a possible embodiment, spatio-temporal feature extraction is performed on spatio-temporal synchronous data through a neural network to generate a high-dimensional spatio-temporal feature tensor, which may include:

[0084] Step 501, perform 3D convolution operation on spatio-temporal synchronous data to extract spatio-temporal local features.

[0085] Specifically, a 3D convolution kernel is used to slide on the spatio-temporal synchronous data, and local spatio-temporal features are extracted through convolution operations. The convolution kernel continuously scans multiple frames of images in the time dimension to capture the motion patterns of traffic elements (such as vehicle acceleration and pedestrian movement trajectories), and at the same time identifies the position and contour of the target in the space dimension. By synchronously fusing time and space information through 3D convolution, the dynamic behavior characteristics of traffic elements (such as lane changing and sudden braking) can be effectively captured, overcoming the limitation that traditional 2D convolution only focuses on spatial features and improving the accuracy of motion state detection.

[0086] Step 502, fuse spatio-temporal local features at different levels through residual connections to generate a fused feature.

[0087] Specifically, cross-layer skip connections are constructed in the neural network to splice channels or add elements of low-dimensional features (such as edges and textures) in the shallow network and high-dimensional features (such as motion trends and global context) in the deep network to generate a fused feature. For example, the 64-channel feature output by the first layer of 3D convolution is connected to the 256-channel feature output by the third layer through a residual block to form a multi-scale feature representation. By using the residual structure, the problem of gradient disappearance is avoided, the adaptability of the model to complex traffic scenarios (such as occlusion and illumination changes) is enhanced, and at the same time, fine-grained motion details are retained, improving the robustness of feature representation.

[0088] Step 503, perform channel attention weighting processing on the fused feature to generate a high-dimensional spatio-temporal feature tensor.

[0089] Specifically, the feature response intensity of each channel is statistically calculated through global average pooling, and a fully connected layer is used to generate a channel weight vector to enhance high-response channels and suppress low-response channels. For example, after applying attention weights to the fused features of 256 channels, a high-dimensional spatio-temporal feature tensor with the same output channel dimension but significantly enhanced key features (such as fast-moving vehicles) is obtained. Through the adaptive channel weighting mechanism, features strongly related to traffic flow detection (such as speed mutation regions) are focused on, background noise interference is suppressed, and the generated feature tensor is made more discriminative to support the high-precision decoding of subsequent vector data.

[0090] In a possible embodiment, performing dynamic semantic mapping processing on traffic vector data to generate traffic state data may include:

[0091] Step 601: Construct dynamic semantic mapping rules according to the layer attribute fields of geographic information data.

[0092] Specifically, dynamic semantic mapping rules are constructed according to the layer attribute fields of geographic information data (such as lane speed limit values, lane topological connection relationships, road network node coordinates). Exemplarily, by parsing the attribute table structure of geographic information data, the numerical range of the lane speed limit field, the connection directionality of the lane topological relationship, and the geographic coordinate information of road nodes are extracted to establish speed-limit association rules, direction-topology matching rules, and position-node mapping rules. For example, the speed mapping rule is defined as "real-time speed / lane speed limit", the direction matching rule is "the included angle threshold between the moving direction angle and the lane centerline vector ≤ 15°", and the position mapping rule is "the Euclidean distance between the coordinate and the nearest road node ≤ 0.5 meters". By dynamically adapting to different geographic information data formats (such as Shapefile, GeoJSON), seamless docking of multi-source GIS platforms is supported, and the inefficiency problem of manual rule configuration is avoided.

[0093] Step 602: Based on the dynamic semantic mapping rules, associate and map speed, direction, and position with the lane speed limit, topological relationship, and coordinate nodes of geographic information data respectively to generate intermediate data.

[0094] Specifically, by matching the traffic vector data with the specific attributes of geographic information data, the traffic state information is mapped into specific geographical space positions and traffic rule constraints, providing preliminary data support for the semantic expression of traffic states.

[0095] Step 603: Perform coordinate transformation error calibration processing on the intermediate data to generate calibrated traffic state data.

[0096] Specifically, based on the control point coordinates (such as road intersection marker points) in the geographic information data, calculate the position mapping error of the intermediate data, and fit the error compensation matrix through affine transformation or least squares method to perform translation, rotation, and scaling correction on the position coordinates, so as to eliminate the cumulative errors (such as projection deformation, device calibration deviation) in the coordinate system conversion process, ensure the geographical spatial consistency between the traffic state data and the GIS layer, and provide high-confidence input for traffic control decisions.

[0097] In a possible embodiment, based on the dynamic semantic mapping rules, the speed, direction, and position are respectively associated and mapped with the lane speed limit, topological relationship, and coordinate nodes of the geographic information data to generate intermediate data, which may include:

[0098] Step 701, calculate the normalized ratio of the speed to the lane speed limit to generate a speed compliance index, and write the speed compliance index into the intermediate data.

[0099] Specifically, calculate the compliance coefficient of each vehicle through the formula "speed compliance index = real-time speed / lane speed limit". When the coefficient is greater than 1, it is marked as an overspeed state. Write the calculation result into the "speed compliance" field of the intermediate data, and append the timestamp and lane number to automatically generate a standardized speed evaluation index, identify overspeed vehicles in real time, and provide data support for dynamic speed limit regulation and violation evidence collection.

[0100] Step 702, calculate the movement direction angle based on the direction, and perform vector angle matching between the movement direction angle and the lane topological relationship of the geographic information data to generate a lane attribution identifier, and write the lane attribution identifier into the intermediate data.

[0101] Specifically, calculate the vehicle movement direction angle based on the direction attribute in the traffic vector data, calculate the direction vector through the displacement difference of two consecutive frame position coordinates, and convert it into an angle value. Perform vector angle matching between this direction angle and the lane topological relationship (such as the lane centerline vector direction) of the geographic information data. For example, if the cosine value of the angle > 0.95 (corresponding angle difference ≤ 18°), it is determined that the vehicle is in the current lane, generate a lane attribution identifier, and write the identification result into the "lane number" field of the intermediate data. Accurately determine the lane where the vehicle is located through vector angle matching, solve the problem of lane drift misjudgment in traditional video detection, and support lane-level traffic flow statistics and illegal lane-changing detection.

[0102] Step 703, perform the nearest neighbor association between the position and the road network nodes of the geographic information data to generate a spatial positioning label, and write the spatial positioning label into the intermediate data.

[0103] Specifically, based on the k-d tree index, the network node closest to the vehicle position (such as the intersection center point or the road segment end point) can be quickly retrieved, the Euclidean distance is calculated and the node ID is recorded. If the distance is less than the preset threshold, a spatial positioning label is generated to mark the road segment or intersection where the vehicle is located, and the node ID, distance value, and road segment number are written into the "spatial positioning" field of the intermediate data to achieve centimeter-level geospatial positioning, support the accurate mapping of vehicle trajectories and road network topologies, and provide a high-precision data basis for applications such as path guidance and accident positioning.

[0104] In summary, the urban road traffic flow detection method provided by the embodiments of the present application generates spatio-temporal synchronization data with unified spatio-temporal benchmarks by performing hardware-level clock synchronization and spatial grid alignment processing on the road surface video stream and geographic information data, thereby solving the spatio-temporal misalignment problem of multi-source data; based on a deep learning model, spatio-temporal features are extracted from the spatio-temporal synchronization data, a 3D convolutional neural network is used to capture the dynamic motion patterns of traffic elements, and structured traffic vector data including speed, direction, and position attributes is generated through weight matrix mapping; further, in combination with the layer attributes of geographic information data, a dynamic semantic mapping rule is constructed to associate and map the dynamic attributes in the traffic vector data with lane speed limits, topological relationships, and road nodes, and spatial errors are eliminated through coordinate calibration to generate traffic state data with both geographic semantic labels and spatio-temporal dynamic characteristics. The above technical methods can improve the accuracy and real-time performance of traffic flow detection, provide accurate data support for signal light optimization, congestion relief, and road network planning, and promote the upgrade of urban traffic management from experience-driven to data-intelligence-driven.

[0105] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0106] Based on the same inventive concept, an embodiment of the present application further provides an urban road traffic flow detection system for implementing the urban road traffic flow detection method involved above. The implementation solutions provided by this system for solving problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the urban road traffic flow detection system provided below can refer to the limitations on the urban road traffic flow detection method in the above text, and will not be elaborated here.

[0107] In an exemplary embodiment, as Figure 2 shown, an urban road traffic flow detection system 10 is provided, including:

[0108] A data preprocessing module 11, configured to obtain a road surface video stream, perform clock synchronization and spatial grid alignment processing on the road surface video stream and geographical information data, and generate spatio-temporal synchronization data.

[0109] A feature extraction module 12, configured to perform spatio-temporal feature extraction on traffic elements in the road surface video stream based on the spatio-temporal synchronization data through a deep learning model, and generate traffic vector data. The attributes of the traffic vector data include speed, direction, and position.

[0110] A traffic detection module 13, configured to perform dynamic semantic mapping processing on the traffic vector data to generate traffic state data, and the traffic state data is associated with the layer attributes of the geographical information data.

[0111] In a possible embodiment, the data preprocessing module 11 may include:

[0112] A clock synchronization unit 111, configured to perform clock synchronization on the camera corresponding to the road surface video stream and the server corresponding to the geographical information data through a precise time protocol, and generate a synchronization signal with aligned timestamps.

[0113] A spatio-temporal slicing unit 112, configured to perform spatio-temporal slicing sampling on consecutive frames of the road surface video stream based on the synchronization signal, and generate a spatio-temporal slice sequence.

[0114] A data mapping unit 113, configured to perform grid mapping processing on the spatio-temporal slice sequence according to the spatial coordinate system of the geographical information data, and generate spatio-temporal synchronization data aligned with the geographical space.

[0115] In a possible embodiment, the spatio-temporal slicing unit 112 may include:

[0116] A sliding time window sampling sub-unit 1121, configured to perform sliding time window sampling on the road surface video stream, where the window length covers consecutive multiple frames of images, and generate a frame sequence with continuous time.

[0117] The spatial grid division sub-unit 1122 is used to perform spatial grid division on each frame image in a time-continuous frame sequence to generate multi-dimensional spatio-temporal slices.

[0118] The data encoding sub-unit 1123 is used to encode the multi-dimensional spatio-temporal slices into four-dimensional tensor data including time depth, spatial dimension, and number of channels, and generate a spatio-temporal slice sequence.

[0119] In a possible embodiment, the feature extraction module 12 may include:

[0120] The high-dimensional feature extraction unit 121 is used to perform spatio-temporal feature extraction on spatio-temporal synchronization data through a neural network to generate a high-dimensional spatio-temporal feature tensor.

[0121] The high-dimensional feature processing unit 122 is used to perform a linear transformation on the high-dimensional spatio-temporal feature tensor through a weight matrix obtained by pre-training to generate initial traffic vector data.

[0122] The data cleaning unit 123 is used to perform trajectory verification and outlier removal on the initial traffic vector data based on motion continuity constraints to generate traffic vector data.

[0123] In a possible embodiment, the high-dimensional feature extraction unit 121 may include:

[0124] The convolution operation sub-unit 1211 is used to perform 3D convolution operation on spatio-temporal synchronization data to extract spatio-temporal local features.

[0125] The residual connection sub-unit 1212 is used to fuse spatio-temporal local features at different levels through residual connection to generate a fused feature.

[0126] The fused feature processing sub-unit 1213 is used to perform channel attention weighting processing on the fused feature to generate a high-dimensional spatio-temporal feature tensor.

[0127] In a possible embodiment, the traffic detection module 13 may include:

[0128] The rule definition unit 131 is used to construct dynamic semantic mapping rules according to the layer attribute fields of geographic information data.

[0129] The association mapping unit 132 is used to perform association mapping on speed, direction, and position with lane speed limit, topological relationship, and coordinate nodes of geographic information data respectively based on the dynamic semantic mapping rules to generate intermediate data.

[0130] The data calibration unit 133 is used to perform coordinate transformation error calibration processing on the intermediate data to generate calibrated traffic state data.

[0131] In a possible embodiment, the association mapping unit 132 may include:

[0132] A speed compliance subunit 1321 is configured to calculate a normalized ratio of the speed to the lane speed limit, generate a speed compliance metric, and write the speed compliance metric into intermediate data.

[0133] A lane attribution subunit 1322 is configured to calculate a movement direction angle based on the direction, perform a vector angle matching between the movement direction angle and the lane topology relationship of the geographic information data, generate a lane attribution identifier, and write the lane attribution identifier into intermediate data.

[0134] A spatial positioning subunit 1323 is configured to perform a nearest neighbor association between the position and the road network nodes of the geographic information data, generate a spatial positioning label, and write the spatial positioning label into intermediate data.

[0135] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a method for detecting urban road traffic flow as described above are implemented.

[0136] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0137] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0138] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A method for detecting urban road traffic flow, characterized in that: The method comprises: Acquire a road video stream, perform clock synchronization and spatial grid alignment processing on the road video stream and geographic information data, and generate time-space synchronization data; Based on the spatiotemporal synchronization data, a deep learning model is used to extract spatiotemporal features of traffic elements in the road surface video stream to generate traffic vector data, wherein the attributes of the traffic vector data include speed, direction and position; Dynamic semantic mapping is performed on the traffic vector data to generate traffic status data, and the traffic status data is associated with the layer attributes of the geographic information data.

2. The method according to claim 1, characterized in that The performing clock synchronization and spatial grid alignment processing on the road surface video stream and the geographic information data to generate time-space synchronization data includes: The camera corresponding to the road surface video stream and the server corresponding to the geographic information data are synchronized by using a precision time protocol to generate a synchronization signal with aligned timestamps; Based on the synchronization signal, the road surface video stream is subjected to spatiotemporal slicing sampling of continuous frames to generate a spatiotemporal slicing sequence; According to the spatial coordinate system of the geographic information data, the time-space slice sequence is subjected to grid mapping processing to generate the time-space synchronization data aligned with the geographic space.

3. The method according to claim 2, characterized in that The performing spatiotemporal slicing sampling of continuous frames of the road surface video stream based on the synchronization signal to generate a spatiotemporal slicing sequence includes: Sliding time window sampling is performed on the road surface video stream, where the window length covers multiple consecutive frame images, to generate a time-continuous frame sequence; Performing spatial grid division on each frame image in the time-continuous frame sequence to generate multi-dimensional space-time slices; The multi-dimensional space-time slices are encoded into four-dimensional tensor data including time depth, space dimension and number of channels to generate the space-time slice sequence.

4. The method according to claim 1, characterized in that The extracting spatiotemporal features of traffic elements in the road surface video stream based on the spatiotemporal synchronization data through a deep learning model to generate traffic vector data includes: Extracting spatiotemporal features from the spatiotemporal synchronization data through a neural network to generate a high-dimensional spatiotemporal feature tensor; Performing a linear transformation on the high-dimensional spatiotemporal feature tensor through a weight matrix obtained through pre-training to generate initial traffic vector data; The initial traffic vector data is subjected to trajectory verification and outlier elimination based on motion continuity constraints to generate the traffic vector data.

5. The method according to claim 4, characterized in that The step of extracting spatiotemporal features from the spatiotemporal synchronization data by using a neural network to generate a high-dimensional spatiotemporal feature tensor includes: Performing a 3D convolution operation on the spatiotemporal synchronization data to extract spatiotemporal local features; The spatiotemporal local features at different levels are fused through residual connections to generate fused features; The fused features are subjected to channel attention weighted processing to generate the high-dimensional spatiotemporal feature tensor.

6. The method according to claim 1, characterized in that The step of performing dynamic semantic mapping processing on the traffic vector data to generate traffic status data includes: Constructing dynamic semantic mapping rules according to the layer attribute fields of the geographic information data; Based on the dynamic semantic mapping rule, the speed, the direction, and the position are respectively associated and mapped with the lane speed limit, the topological relationship, and the coordinate node of the geographic information data to generate intermediate data; The intermediate data is subjected to a coordinate conversion error calibration process to generate the calibrated traffic status data.

7. The method according to claim 6, characterized in that Based on the dynamic semantic mapping rule, the speed, the direction, and the position are respectively associated and mapped with the lane speed limit, the topological relationship, and the coordinate node of the geographic information data to generate intermediate data, including: Calculate a normalized ratio of the speed to the lane speed limit to generate a speed compliance index, and write the speed compliance index into the intermediate data; Calculate the moving direction angle based on the direction, match the moving direction angle with the lane topology relationship of the geographic information data by vector angle, generate a lane ownership mark, and write the lane ownership mark into the intermediate data; The location is associated with the road network node of the geographic information data by nearest neighbor, a spatial positioning tag is generated, and the spatial positioning tag is written into the intermediate data.

8. A city road traffic flow detection system, characterized in that: The system comprises: A data preprocessing module is used to obtain a road surface video stream, perform clock synchronization and spatial grid alignment processing on the road surface video stream and geographic information data, and generate time-space synchronization data; A feature extraction module, configured to extract spatiotemporal features of traffic elements in the road surface video stream based on the spatiotemporal synchronization data through a deep learning model to generate traffic vector data, wherein the attributes of the traffic vector data include speed, direction and position; The traffic detection module is used to perform dynamic semantic mapping processing on the traffic vector data to generate traffic status data, and the traffic status data is associated with the layer attributes of the geographic information data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Big data rapid processing operation method based on smart city

    CN120434362A