A vision measurement based traffic perception system and method

CN121725438BActive Publication Date: 2026-09-22CHINA AUTOMOTIVE INTELLIGENT TECHNOLOGY (TIANJIN) CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610226689.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-26
Publication Date
2026-09-22
Estimated Expiration
2046-02-26

AI Technical Summary

Technical Problem

整个过程不仅耗时耗力,响应滞后,且易因人为疲劳或判断差异导致处理结果不一致甚至误判

Benefits of technology

本申请提供了一种基于视觉测量的交通感知方法和系统,通过在路口的多个摄像头各自采集道路图像,逐步转换到实际环境中,云端对映射得到的同一交通元素的轮廓进行再处理,得到交通元素的运动信息;并对所述运动信息进行融合,得到融合运动信息;根据融合运动信息对道路交通情况进行感知。相较于传统监控摄像系统,本发明实现了智能化、自主定位、抗失效、多终端协同的替代性部署方式,适用于智慧交通感知系统构建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725438B_ABST
    Figure CN121725438B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of vehicle safety, in particular to a traffic perception system and method based on visual measurement. The method comprises the following steps: a plurality of cameras arranged at an intersection respectively acquire road images, and the poses of the cameras are obtained; each camera adopts a detection and segmentation integrated algorithm based on a light DQN network to synchronously perform traffic element segmentation and identification on the road images, so that the types and contours of the traffic elements are obtained; the contours of the traffic elements are mapped to an actual environment according to the poses of the cameras; the contours of the same traffic elements obtained by mapping are reprocessed in the cloud, so that the motion information of the traffic elements is obtained; the motion information is fused, so that fused motion information is obtained; and the road traffic condition is perceived according to the fused motion information. The application realizes detection, ranging and positioning of traffic elements, and improves the real-time performance, intelligence and redundancy of road monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent sensing and traffic safety technology, and more specifically, to a traffic sensing system and method based on visual measurement. Background Technology

[0002] In traditional road monitoring systems, ordinary cameras only have video capture capabilities. When emergencies such as traffic accidents, illegal lane changes, or pedestrians jaywalking occur, relevant management departments must rely on manual operation to retrieve historical video data from each camera individually. Professionals then visually observe and analyze the footage to determine the incident and assign responsibility. This entire process is not only time-consuming and labor-intensive, with delayed response times, but also prone to inconsistencies and even misjudgments due to human fatigue or differing judgments. Furthermore, the lack of coordination between cameras prevents automatic vehicle tracking and hinders intelligent traffic management.

[0003] Moreover, when identifying targets from images captured by cameras, there are technical problems such as low recognition efficiency and inability to accurately identify the target's location and size.

[0004] In view of the above, this application is hereby submitted. Summary of the Invention

[0005] The purpose of this application is to provide a traffic perception system and method based on visual measurement to detect, measure distances and locate traffic elements (such as vehicles, pedestrians, etc.), thereby improving the real-time performance, intelligence and redundancy of road monitoring.

[0006] To achieve the above objectives, this application adopts the following technical solution: In a first aspect, this application provides a traffic perception method based on visual measurement, comprising: Multiple cameras deployed at the intersection each capture road images and obtain the camera pose; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and outline of the traffic elements. The contours of traffic elements are mapped onto the actual environment based on the pose of each camera. The cloud platform reprocesses the outline of the same traffic element obtained from the mapping to obtain the motion information of the traffic element; and fuses the motion information to obtain fused motion information. The road traffic situation is perceived based on the fused motion information; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element, including: Each camera processes the road image into grayscale to obtain a grayscale image, on which the edges of traffic elements are displayed; Select an endpoint from the edge; The sub-image of the local window around the endpoint is input into the lightweight DQN network to obtain the target action direction corresponding to the maximum Q value output by the lightweight DQN network; the lightweight DQN network has a one-to-one correspondence with the type of traffic element; Move a predetermined distance from the endpoint toward the target action direction to hit a new endpoint; Returns the operation of feeding sub-images of local windows around the endpoints into the lightweight DQN network until the stopping condition is met; Connecting multiple endpoints yields a closed-loop path, and taking the convex hull of the closed-loop path yields a polygonal frame.

[0007] Secondly, this application provides a traffic perception system based on vision measurement, comprising: Multiple cameras are deployed at the intersection, and these cameras are connected to the cloud. Multiple cameras each capture road images and obtain the camera's pose; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element; and maps the contour of the traffic elements to the actual environment based on the pose of each camera. The cloud platform reprocesses the outline of the same traffic element obtained from the mapping to obtain the motion information of the traffic element; and fuses the motion information to obtain fused motion information. The road traffic situation is perceived based on the fused motion information; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element, including: Each camera processes the road image into grayscale to obtain a grayscale image, on which the edges of traffic elements are displayed; Select an endpoint from the edge; The sub-image of the local window around the endpoint is input into the lightweight DQN network to obtain the target action direction corresponding to the maximum Q value output by the lightweight DQN network; the lightweight DQN network has a one-to-one correspondence with the type of traffic element; Move a predetermined distance from the endpoint toward the target action direction to hit a new endpoint; Returns the operation of feeding sub-images of local windows around the endpoints into the lightweight DQN network until the stopping condition is met; Connecting multiple endpoints yields a closed-loop path, and taking the convex hull of the closed-loop path yields a polygonal frame.

[0008] Compared with the prior art, the beneficial effects of this application are as follows: This application provides a traffic perception method and system based on visual measurement. It involves acquiring road images from multiple cameras at an intersection, progressively converting them to the actual environment, and then reprocessing the contours of the same traffic element in the cloud to obtain motion information. This motion information is then fused to obtain fused motion information, which is used to perceive road traffic conditions. Compared to traditional surveillance camera systems, this invention provides an intelligent, autonomous, fault-resistant, and multi-terminal collaborative alternative deployment method, suitable for building intelligent traffic perception systems.

[0009] Furthermore, this application employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of traffic elements. By integrating detection and segmentation, efficiency is improved. Moreover, the final result is a contour that conforms to the shape of the traffic element itself, which participates in subsequent mapping and perception of the road traffic environment, thereby improving the accuracy of element position and size detection, as well as the accuracy of traffic perception. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a traffic perception method based on visual measurement provided in an embodiment of this application; Figure 2 This is a schematic diagram of endpoint movement provided in an embodiment of this application; Figure 3 This is a schematic diagram of the polygonal frame provided in an embodiment of this application; Figure 4 This is a schematic diagram of a common detection frame provided in the embodiments of this application. Detailed Implementation

[0012] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0013] The present application will be further described in detail below with reference to the embodiments.

[0014] Figure 1 This is a flowchart illustrating a traffic perception method based on visual measurement provided in an embodiment of this application. See also... Figure 1 The method provided in this application is applicable to situations where multiple cameras installed at intersections are used for traffic element perception and intelligent traffic management. The method provided in this embodiment includes: S110: Multiple cameras deployed at the intersection collect road images and obtain the camera pose.

[0015] Multiple cameras are deployed at the intersection to form a comprehensive monitoring network. Each camera includes an industrial-grade high-frame-rate monocular camera, an inertial navigation system (IMU), and a positioning module (GPS / BeiDou / GNSS). The camera installation locations ensure coverage of all key areas of the intersection, including pedestrian crossings, stop lines, and turning lanes. The camera sampling frequency is 25-30fps, and the resolution is no less than 1920×1080 to adapt to different lighting conditions.

[0016] The cameras acquire real-time pose information, including position coordinates (longitude, latitude, altitude) and attitude angles (pitch, roll, yaw), through a built-in inertial navigation system. All cameras have a unified time synchronization to ensure data consistency.

[0017] S120. Each camera uses an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of the traffic elements.

[0018] Each camera is equipped with an embedded edge computing board, integrating a deep learning-based object detection algorithm. Optionally, the object detection algorithm uses models such as YOLOv5 or Faster R-CNN to perform real-time analysis of each frame of road image, detecting traffic element types including vehicles, pedestrians, and non-motorized vehicles. The output of the object detection algorithm includes: 1) bounding box coordinates (x_min, y_min, x_max, y_max); 2) traffic element category (car, truck, pedestrian, bicycle, etc.); 3) confidence score. Furthermore, the bounding boxes of traffic elements are marked on the road image with special colored borders to highlight the location of the traffic elements.

[0019] Optionally, the object detection algorithm is an integrated detection and segmentation algorithm based on a lightweight DQN network (Deep Q-network). First, the lightweight DQN network simultaneously segments and identifies the edges of traffic elements from the road image. Then, the edges are processed to obtain the contours of the traffic elements, specifically represented by polygons. This will be described in detail in the following embodiments.

[0020] S130: Map the outline of traffic elements to the actual environment based on the pose of each camera.

[0021] Step 1: Based on the camera's internal and external parameters and pose, construct the first transformation matrix between the image coordinate system and the ground coordinate system; based on the first transformation matrix, map the contour corner points (i.e., the corner points of polygons) of traffic elements in the image to the ground coordinate system.

[0022] ; Where K is the intrinsic parameter of the camera, u and v are the image coordinates, and X is the image coordinate. c Y c Z c These are three-dimensional coordinates in the camera coordinate system.

[0023] ; Where R is the pose of the camera coordinate system relative to the ground coordinate system, t is the position of the camera in the ground coordinate system, and X is the position of the camera in the ground coordinate system. g Y g Z g These are coordinates in a ground coordinate system.

[0024] Step 2: Based on the second transformation matrix between the ground coordinate system and the station center coordinate system, map the outline corner points of the traffic elements to the station center coordinate system.

[0025] The station center coordinate system is selected based on the East-North-Sky coordinate system.

[0026] ; in, It is an offset matrix. It is the rotation matrix from the ground coordinate system to the station center coordinate system. , and These are coordinates in the station-centered coordinate system.

[0027] Step 3: Based on the transformation function between the station center coordinate system and the WGS-84 coordinate system, map the contour corner points of the traffic elements to the WGS-84 coordinate system.

[0028] ; in, , , These are the WGS-84 coordinates corresponding to the origin of the station-centered coordinate system. It is the transformation function from the station-centered coordinate system to the WGS-84 coordinate system. , , These are coordinates in the WGS-84 coordinate system.

[0029] S140. The cloud reprocesses the contour of the same traffic element obtained by mapping to obtain the motion information of the traffic element; and fuses the motion information to obtain fused motion information.

[0030] First, based on the contour and image frame rate, calculate the position and size of traffic elements, the spacing between adjacent traffic elements, and the instantaneous velocity vector. Figure 3 This is a schematic diagram of the outline provided in the embodiments of this application. Taking a vehicle as an example, firstly, the two parallel or nearly parallel edges that are furthest apart in the polygon are selected as the head and tail of the vehicle. Then, the distance between the head and tail is calculated as the vehicle length; the distance perpendicular to the vehicle length direction is the vehicle width; the intersection of the vehicle length and width is the vehicle position (i.e., the centroid position). The distance between the centroid position of the previous traffic element (e.g., a vehicle) and the centroid position of the next traffic element (e.g., a vehicle) is the spacing. Assuming that the displacement of the same traffic element along a certain direction in two adjacent road images is Ln, and the time interval between the two images is Tn, then the magnitude of the instantaneous velocity vector along a certain direction is Ln / Tn, and the vector direction points to the position of the traffic element in the next image.

[0031] Optionally, considering that in traditional static camera monitoring systems, the camera's pose is usually assumed to be constant, the motion of elements detected in road images is entirely caused by the element's own motion. However, when the camera itself is also in motion (e.g., a camera swaying due to strong winds or vibrations), the motion of elements observed in the image is actually a mixture of the element's actual motion and the camera's own motion. To improve the accuracy of element motion detection, relative motion compensation is introduced. High-frequency, real-time motion data provided by the IMU is used as a "reference benchmark" to remove the camera's own motion component from the observed mixed motion, thereby extracting the true absolute motion of the traffic element relative to the fixed ground. For example, the IMU data is converted to obtain the camera's motion direction and velocity in the WGS-84 coordinate system. The instantaneous velocity vector and magnitude of the traffic element are added to the instantaneous velocity vector and magnitude of the camera to correct the instantaneous velocity vector and magnitude of the traffic element.

[0032] Due to calibration errors, pixel noise, and different viewing angles, each camera's observations of the instantaneous velocity vector sum of the same traffic element (such as a vehicle) and its distance from adjacent traffic elements are uncertain. This uncertainty can be described by a probability distribution. The purpose of fusion is to calculate the "fused motion information" that maximizes the probability based on the uncertainties of each observation.

[0033] First, construct a covariance matrix based on the errors of each camera. Camera errors include intrinsic parameter errors, extrinsic parameter errors, and detection errors, which can be obtained through extensive prior testing. For example, construct a covariance matrix by taking the variances of the estimated errors in the x-direction, y-direction, and z-direction, and the covariances of the estimated errors in any two directions.

[0034] Then, the instantaneous velocity vectors of the same traffic element are fused using the covariance matrix to obtain the fused instantaneous velocity vector; the distances between the same traffic element and adjacent traffic elements are fused using the covariance matrix to obtain the fused distances.

[0035] In addition to the fusion algorithm based on the covariance matrix mentioned above, weighted summation, averaging and other methods can also be used to fuse motion information of the same traffic element from multiple cameras.

[0036] S150. Perceive road traffic conditions based on fused motion information.

[0037] The type (e.g., vehicle embedding vector), location, size, fused instantaneous velocity vector, and fused spacing of traffic elements are input into a large model to obtain the behavior of traffic elements, risk hotspots, and traffic congestion trends; optionally, the large model is a Transformer model or a graph neural network model.

[0038] Traffic element behavior identification includes: vehicle lane changes, turning, sudden braking, stopping, accident warnings, and other behavioral patterns; risk heat zones include: potential conflict points, generating real-time risk heat maps. Traffic congestion trends include: no congestion and congestion levels.

[0039] The following specific example illustrates in detail the processing procedure for large models.

[0040] Obtain the type (e.g., vehicle embedding vector), location, size, fused instantaneous velocity vector, and fused spacing of traffic elements detected in the road images of the 30 frames prior to the current time, and input this data into the Transformer model.

[0041] The Transformer model employed in this application is a specially designed encoder-decoder architecture specifically designed for analyzing the spatiotemporal dynamics of traffic elements. The model receives a serialized traffic data stream as input, with each time step containing the type of traffic element (e.g., vehicle embedding vector), location, size, fused instantaneous velocity vector, and fused spacing. At the input embedding layer, this diverse and heterogeneous data is first transformed into a unified high-dimensional feature vector representation. To preserve the sequential information of the time series, the model also introduces sine and cosine position coding, assigning a unique location identifier to each time step.

[0042] The input embedding layer is followed by an encoder consisting of six identical stacked layers. At the heart of each encoder layer is a multi-head self-attention mechanism, which captures the complex dependencies between traffic elements across multiple time steps from multiple subspaces in parallel. Specifically, the encoder computes the query, key, and value vectors for each time step, using attention weights to determine which states from historical moments are most important for the current moment's behavioral analysis. For example, a sudden braking action might be closely related to a deceleration trend in the preceding seconds; the attention mechanism automatically captures this long-range dependency. The attention output is then non-linearly transformed by a feedforward neural network. Each sublayer is equipped with residual connections and layer normalization to stabilize the training process and ensure effective gradient propagation.

[0043] The decoder also consists of six layers. The first layer is a masked multi-head self-attention layer, ensuring that information from previous time steps can only be accessed when predicting the current output, preventing the leakage of future information. The second layer is the crucial video encoder-decoder attention layer, where the decoder focuses on the most relevant spatiotemporal features in the encoder output. For example, when determining whether a vehicle is changing lanes, the decoder pays special attention to the state information of adjacent lanes and surrounding vehicles provided in the encoder. The final layer is still a feedforward network, used to further refine the features.

[0044] Following the decoder are multi-task output heads, including a behavior recognition output head, a risk heatmap output head, and a traffic congestion trend output head. The behavior recognition output head extracts salient features from the sequence using global max pooling, then classifies them using a multilayer perceptron to output the probability distribution of vehicle behaviors such as lane changes, turns, and sudden braking, accurately identifying complex driving intentions. The risk heatmap output head reconstructs the decoder's output into a spatial feature map, using a convolutional neural network to extract local spatial patterns and generate a high-resolution real-time risk heatmap, accurately locating potential danger zones such as intersection conflict points and pedestrian crossings in high-risk areas. The traffic congestion trend output head captures the overall macroscopic state of traffic flow through global average pooling, outputting multi-level assessment results ranging from no congestion to severe congestion.

[0045] The entire model employs an end-to-end multi-task learning strategy, which not only achieves parameter sharing and improves computational efficiency but also promotes knowledge transfer between different tasks. For example, the identified sudden braking behavior directly affects the risk level of the corresponding area in the risk heatmap. During inference, the model can process the trajectory sequences of multiple traffic elements in parallel, ultimately forming a comprehensive and in-depth situational awareness of the traffic scene.

[0046] The large model's output supports multi-format data encapsulation, including structured JSON, binary protocol, and ROS2 topic broadcasting. Output data can be integrated into intelligent transportation platforms, vehicle-road cooperative RSU systems, traffic situation platforms, etc. The large model's output is provided to users as daily, weekly, or real-time alerts, and supports chart / map displays.

[0047] The application scenarios of the implementation methods provided in this application include: 1) target measurement (including size, position, distance, speed, etc.) and accident detection of traffic networks; 2) multi-camera collaborative traffic situation monitoring platform; 3) replacing traditional cameras to realize edge intelligence upgrade and transformation.

[0048] It's important to note that before performing traffic perception, the large-scale model needs to be pre-trained using training samples to equip it with road traffic perception capabilities. For example, training samples include the type, location, size, fused instantaneous velocity vector, and fused spacing of traffic elements identified from multiple road images, along with labels indicating the behavior, risk hotspots, and traffic congestion trends of real-world traffic elements. By iterating through the parameters of the large-scale model, its predictions are made to approximate the labeled values. The specific training process will not be detailed here.

[0049] Optionally, each camera periodically reports its health status to the cloud; if any camera malfunctions or a new camera is added, distributed fusion perception is performed based on the road images collected by all the current cameras.

[0050] For example, each camera periodically (e.g., every 30 seconds) reports its health status to the cloud, including CPU / GPU load, memory usage, network connectivity, and sensor operating status. When the cloud detects a camera malfunction or adds a new camera, it automatically reconfigures the fusion weights to ensure the robustness of the perception network. Observational data from the malfunctioning camera is excluded, and newly added cameras participate in the fusion calculation after calibration. If a camera malfunctions, other working cameras can take over its computational functions. For instance, a malfunctioning camera with normal recording capabilities can send road images and its own pose to nearby cameras, which then perform target detection, location mapping, and upload the images to the cloud.

[0051] This application provides a traffic perception method and system based on visual measurement. It involves acquiring road images from multiple cameras at an intersection, progressively converting them to the actual environment, and then reprocessing the contours of the same traffic element in the cloud to obtain motion information. This motion information is then fused to obtain fused motion information, which is used to perceive road traffic conditions. Compared to traditional surveillance camera systems, this invention provides an intelligent, autonomous, fault-resistant, and multi-terminal collaborative alternative deployment method, suitable for building intelligent traffic perception systems.

[0052] Furthermore, this application employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of traffic elements. By integrating detection and segmentation, efficiency is improved. Moreover, the final result is a contour that conforms to the shape of the traffic element itself, which participates in subsequent mapping and perception of the road traffic environment, thereby improving the accuracy of element position and size detection, as well as the accuracy of traffic perception.

[0053] Optionally, each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in the road image, obtaining the type and contour of the traffic elements, including the following operations: Step 1: Each camera performs grayscale processing on the road image to obtain a grayscale image, on which the edges of traffic elements are displayed.

[0054] Optionally, the road image is processed using the Sobel operator to obtain initial edges. Simultaneously, morphological operations (such as erosion) are used to extract the edges of traffic elements. Then, the edges obtained using the Sobel operator are combined with the edges extracted through morphological operations to obtain the final edges. A grayscale image is output, where white pixels represent detected edges and black pixels represent the background. (See [link to relevant documentation]). Figure 2 As shown.

[0055] Step 2: Select an endpoint from the edge.

[0056] The first endpoint selected from the edge is used as the starting endpoint. These starting endpoints are randomly scattered in areas that may contain traffic elements (referring to areas on the edge), excluding locations such as green belts and rivers. For different areas, different starting endpoints can be scattered in different areas, and subsequent steps can be executed in parallel, thereby achieving parallel segmentation and recognition.

[0057] Step 3: Input the sub-image of the local window around the endpoint into the lightweight DQN network to obtain the target action direction corresponding to the maximum Q value output by the lightweight DQN network; the lightweight DQN network has a one-to-one correspondence with the type of traffic element. If it is necessary to detect vehicles and pedestrians in the road image, a lightweight DQN network for vehicle detection and a lightweight DQN network for pedestrian detection are needed to simultaneously segment and identify traffic elements in the road image.

[0058] Step 4: Starting from the endpoint, move a set length in the direction of the target action to hit the new endpoint. Figure 2 (The red dot in the image).

[0059] The lightweight DQN network consists of an input layer, convolutional layers, fully connected layers, and an output layer. The input layer preprocesses the input sub-image to obtain an appropriate resolution. The convolutional layer contains three kernels, used to extract low-dimensional, mid-dimensional, and high-dimensional features from the preprocessed image, respectively. These low-dimensional, mid-dimensional, and high-dimensional features are input to the fully connected layer (with 256 neurons), which integrates the multi-dimensional features into a global representation. The output layer has nine neurons, corresponding to eight action directions (up, down, left, right, top-left, bottom-left, top-right, bottom-right). Figure 2(The yellow arrow in the diagram) and one stop action. Each neuron in the output layer outputs the Q-value, intermediate vector, and confidence score for each action.

[0060] Define a local window (without specific size restrictions, e.g., 20 pixels * 20 pixels) centered on the endpoint. This local window should be able to cover a section of the edge of the traffic element. Input the sub-image within this local window into a lightweight DQN network. After forward propagation through the lightweight DQN network, obtain the Q-value of each action, and select the action with the largest Q-value.

[0061] For example, if you select the upper right action, you will move to the upper right several pixels (e.g., 3 to 5 pixels) from that endpoint, and the pixel at the new position will be the new endpoint.

[0062] Delineate a new local window of a square centered on the new endpoint. Figure 2 (using the blue square window in the image), the sub-images within the new local window are input into the lightweight DQN network to obtain the Q-value for each action, and the action with the largest Q-value is selected. This process is repeated to obtain multiple endpoints.

[0063] Before using a lightweight DQN network to select action directions, the network needs to be trained: Multiple road image samples are collected, and ground truth edges (the true edges of traffic elements) are labeled on these samples. A reward rule is constructed, including: increasing the first reward value (e.g., 1.0) if a new edge is hit; increasing the second reward value (e.g., 0.1) if a historical edge is hit; subtracting the third reward value (e.g., 0.2) if a non-edge (i.e., inside a traffic element) is hit; subtracting the fourth reward value (e.g., 0.02) if multiple visited windows are hit consecutively (indicating reverse endpoint movement); increasing the fifth reward value (e.g., 5.0) if the path is closed; and subtracting the sixth reward value (e.g., 0.01) after each step as the cost for each step. The lightweight DQN network is then trained based on multiple road image samples to update its parameters and maximize the reward value. For example, multiple road image samples can be placed in an experience replay pool. Road image samples are randomly selected from the experience replay pool, processed into grayscale, and initial endpoints are scattered along the initial edges. Sub-images are cropped with the endpoints as the center and input into a lightweight DQN network to obtain the predicted target action direction. New endpoints are obtained based on the predicted target action direction. At this point, the reward value is calculated according to the reward rule.

[0064] The reward rules here can suppress path jitter and empty runs away from the edge.

[0065] Step 5: If the stopping condition is not met, return to step 3.

[0066] The stopping conditions include: the path is closed (i.e., the action with the maximum Q value is the stopping action); or the distance between the starting endpoint and the current endpoint is less than a set value (e.g., 3 to 5 pixels), indicating that the path is close to being closed. Figure 2 The green path is shown in part.

[0067] Step 6: If the stopping condition is met, connect multiple endpoints to obtain a closed-loop path, and take the convex hull of the closed-loop path to obtain a polygonal frame.

[0068] For paths that are not completely closed but are close to closed, the shortest path bridging method (connecting the first and last endpoints with straight lines) is used to obtain the closed loop path.

[0069] See Figure 3 The traffic elements are represented by images, and the purple lines represent closed-loop paths, which perfectly depict the vehicle's shape. The next step is to find the smallest red convex polygon that encloses all closed-loop paths. Convex hull algorithms include, but are not limited to, Graham's scan method and Jarvis's stepping method.

[0070] Optionally, if the stopping condition is met, the closed-loop path needs to be evaluated for performance metrics before proceeding with convex hull calculation. Specifically, metrics are calculated for the predicted edge path relative to the real edge. These metrics include: coverage, edge hit rate, loop closure rate, and path redundancy rate. Coverage is the proportion of pixels in the intersection of the predicted closed-loop path and the real edge to the number of pixels in the real edge. Edge hit rate is the proportion of pixels in the predicted closed-loop path that belong to the real edge. Loop closure rate is the proportion of samples that ultimately achieve loop closure in a batch of training samples. Path redundancy rate is the proportion of the predicted closed-loop path length to the length of the real edge. If the metrics do not meet the requirements, such as coverage, edge hit rate, or loop closure rate being below a threshold, or path redundancy rate being above a threshold, the starting endpoint is reselected, and the process returns to step 3.

[0071] This embodiment has the following technical effects: 1. It allows for free allocation of computing power, replacing large-scale neural networks on the market with a large number of lightweight DQN networks. It controls the number of random starting endpoints to allocate computing power according to the actual situation. For example, if a region is more important, more starting endpoints can be allocated.

[0072] 2. The detection frame in the prior art is rectangular, see [reference]. Figure 4 Furthermore, the edge paths in this application can be regressed into polygons, which are more suitable for vehicles with abnormal postures. These polygons can better represent the spatial position of the target in subsequent processing, resulting in more accurate position, size, instantaneous velocity vectors, and spacing.

[0073] 3. The detection and segmentation integrated algorithm based on the lightweight DQN network can perform traffic element segmentation and recognition in parallel, greatly improving recognition efficiency.

[0074] This application also provides a traffic perception system based on visual measurement, including: multiple cameras deployed at intersections, wherein the multiple cameras are connected to the cloud.

[0075] Multiple cameras each capture road images and obtain the camera's pose; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element; and maps the contour of the traffic elements to the actual environment based on the pose of each camera. The cloud platform reprocesses the outline of the same traffic element obtained from the mapping to obtain the motion information of the traffic element; and fuses the motion information to obtain fused motion information. The road traffic situation is perceived based on the fused motion information.

[0076] Each camera includes: an embedded edge computing board; an industrial-grade high frame rate monocular camera; an inertial navigation system and positioning module; and a data upload module.

[0077] The traffic perception system based on vision measurement provided in this embodiment can execute the traffic perception method based on vision measurement provided in any embodiment and has the corresponding technical effects, which will not be elaborated here.

[0078] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means, such as coaxial cable, optical fiber, digital subscriber line (DSL), or wireless means, such as infrared, wireless, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium, etc. It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium, in other words, it can be a non-transient storage medium.

[0079] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0080] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A traffic perception method based on visual measurement, characterized in that, include: Multiple cameras deployed at the intersection each capture road images and obtain the camera pose; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and outline of the traffic elements. The contours of traffic elements are mapped onto the actual environment based on the pose of each camera. The cloud platform reprocesses the outline of the same traffic element obtained from the mapping to obtain the motion information of the traffic element; The motion information is then fused to obtain fused motion information. The fusion of the motion information to obtain fused motion information includes: constructing a covariance matrix based on the errors of each camera; and using the covariance matrix, fusing the instantaneous velocity vector of the same traffic element with the distance between adjacent traffic elements to obtain fused motion information. The road traffic situation is perceived based on the fused motion information, including: inputting the type, location, size, fused instantaneous velocity vector, and fused spacing of traffic elements into a large model to obtain the behavior of traffic elements, risk hotspots, and traffic congestion trends; the behavior of traffic elements includes identification of: vehicle lane changing, turning, sudden braking, stopping, and accident warning behavior patterns; risk hotspots include: potential conflict points, generating a real-time risk heat map; traffic congestion trends include: no congestion and congestion level; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element, including: Each camera processes the road image into grayscale to obtain a grayscale image, on which the edges of traffic elements are displayed; the starting endpoints are randomly scattered in the regions on the edges, and different starting endpoints are scattered in different regions, and subsequent steps are executed in parallel. Select an endpoint from the edge; The sub-image of the local window around the endpoint is input into the lightweight DQN network to obtain the target action direction corresponding to the maximum Q value output by the lightweight DQN network; the lightweight DQN network has a one-to-one correspondence with the type of traffic element; Move a predetermined distance from the endpoint toward the target action direction to hit a new endpoint; Returns the operation of feeding sub-images of local windows around the endpoints into the lightweight DQN network until the stopping condition is met; Connecting multiple endpoints yields a closed-loop path, and taking the convex hull of the closed-loop path yields a polygonal frame.

2. The method according to claim 1, characterized in that, Before feeding the sub-images of local windows around the endpoints into the lightweight DQN network, the following is also included: Collect multiple road image samples and annotate the true edges on the road image samples; Construct reward rules, which include: increasing the first reward value if a new edge is hit, increasing the second reward value if a historical edge is hit, subtracting the third reward value if a non-edge is hit, subtracting the fourth reward value if multiple visited windows are hit consecutively, increasing the fifth reward value if the path is closed, and subtracting the sixth reward value after each step. The lightweight DQN network is trained based on the multiple road image samples to update the parameters of the lightweight DQN network and maximize the reward value. The lightweight DQN network includes an input layer, a convolutional layer, a fully connected layer, and an output layer.

3. The method according to claim 2, characterized in that, The stopping conditions include: the path is closed, or the distance between the starting endpoint and the current endpoint is less than a set value; After the stopping condition is met, it also includes: The predicted edge paths are used to calculate metrics relative to the actual edges, including: coverage, edge hit rate, loop closure rate, and path redundancy rate. If the specified metrics do not meet the requirements, a new starting endpoint is selected.

4. The method according to claim 3, characterized in that, Based on the pose of each camera, the contours of traffic elements are mapped onto the real environment, including: Based on the camera's internal and external parameters and pose, construct the first transformation matrix between the image coordinate system and the ground coordinate system; Based on the first transformation matrix, the contour corner points of traffic elements in the image are mapped to the ground coordinate system; Based on the second transformation matrix between the ground coordinate system and the station center coordinate system, the outline corner points of the traffic elements are mapped to the station center coordinate system; Based on the transformation function between the station center coordinate system and the WGS-84 coordinate system, the contour corner points of traffic elements are mapped to the WGS-84 coordinate system.

5. The method according to claim 1, characterized in that, The large model is either a Transformer model or a graph neural network model.

6. A traffic perception system based on visual measurement, characterized in that, include: Multiple cameras are deployed at the intersection, and these cameras are connected to the cloud. Multiple cameras each capture road images and obtain the camera's pose; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element; and maps the contour of the traffic elements to the actual environment based on the pose of each camera. The cloud platform reprocesses the outline of the same traffic element obtained from the mapping to obtain the motion information of the traffic element; The motion information is then fused to obtain fused motion information. The fusion of the motion information to obtain fused motion information includes: constructing a covariance matrix based on the errors of each camera; and using the covariance matrix, fusing the instantaneous velocity vector of the same traffic element with the distance between adjacent traffic elements to obtain fused motion information. The road traffic situation is perceived based on the fused motion information, including: inputting the type, location, size, fused instantaneous velocity vector, and fused spacing of traffic elements into a large model to obtain the behavior of traffic elements, risk hotspots, and traffic congestion trends; the behavior of traffic elements includes identification of: vehicle lane changing, turning, sudden braking, stopping, and accident warning behavior patterns; risk hotspots include: potential conflict points, generating a real-time risk heat map; traffic congestion trends include: no congestion and congestion level; Each camera employs an integrated detection and segmentation algorithm based on a lightweight DQN network to simultaneously segment and identify traffic elements in road images, obtaining the type and contour of each traffic element. The process includes: each camera performing grayscale processing on the road image to obtain a grayscale image displaying the edges of traffic elements; starting endpoints randomly scattered across regions on the edges, with different starting endpoints scattered across different regions, and subsequent steps executed in parallel; selecting an endpoint from the edge; inputting the sub-images of the local windows surrounding the endpoint into the lightweight DQN network to obtain the target action direction corresponding to the maximum Q value output by the lightweight DQN network; the lightweight DQN network having a one-to-one correspondence with the type of traffic element; moving a set length from the endpoint towards the target action direction to hit a new endpoint; returning to the operation of inputting the sub-images of the local windows surrounding the endpoint into the lightweight DQN network until a stopping condition is met; connecting multiple endpoints to obtain a closed-loop path, and taking the convex hull of the closed-loop path to obtain a polygonal bounding box.

7. The system according to claim 6, characterized in that, Each camera includes: an embedded edge computing board; an industrial-grade high frame rate monocular camera; an inertial navigation system and positioning module; and a data upload module.

Citation Information

Patent Citations

  • Left ventricular intimal image segmentation method and system based on deep reinforcement learning

    CN117036689A

  • Image scene understanding method and system applied to Internet of Vehicles road condition analysis

    CN120726584A

  • Method and device for determining end-to-end perception decision regulation and control architecture

    CN120726596A

  • Traffic monitoring method and system for vehicle tracking and early warning

    CN120766534A