Traffic target monitoring and tracking method based on monocular vision daytime scene reconstruction

Through the traffic target monitoring and tracking system based on monocular vision, combined with multiple modules and algorithms, the shortcomings of radar and cameras in traffic monitoring in existing technologies are solved, and accurate monitoring and real-time tracking of vehicle positions are achieved, which is suitable for various complex road environments.

CN119600550BActive Publication Date: 2025-09-30XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411593504.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-09-30
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

Existing radar and camera detection technologies have difficulty achieving comprehensive and accurate vehicle target detection and tracking in traffic monitoring. Radar cannot obtain vehicle appearance features and is expensive. Camera performance degrades in bad weather, and monocular cameras cannot provide vehicle location information.

Method used

It adopts a traffic target monitoring and tracking system based on monocular vision, combined with a monocular camera, inertial navigation module, lightweight processor, radio frequency communication module and Beidou positioning module. Through frame difference detection, contour tracking and region growing algorithm, it realizes vehicle feature extraction and real-time trajectory tracking, and uses solar energy storage module for power supply.

Benefits of technology

It achieves accurate monitoring and real-time tracking of vehicle positions in different road environments. It is suitable for scenarios without power supply or rapid deployment. It has high flexibility and low operating costs and is suitable for real-time monitoring of urban roads and intercity highways.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600550B_ABST
    Figure CN119600550B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic target monitoring and tracking system and method based on monocular vision daytime scene reconstruction, including a monocular camera, an inertial navigation module, a lightweight processor, a radio frequency communication module, a solar energy storage module and a Beidou positioning module; the monocular camera is used to continuously capture multiple frames of images; the inertial navigation module uses a built-in high-precision gyroscope to measure the three-dimensional posture of the monocular camera when taking pictures, and calculates the camera's external parameter matrix; the lightweight processor is responsible for processing the multiple frames of images collected by the camera, performing image preprocessing, vehicle detection and trajectory calculation, and executing the core algorithm; the radio frequency communication module transmits the detected traffic flow information back to the traffic monitoring center; the Beidou positioning module provides geographic location information, so that the traffic flow information can be synchronized with the geographic location. Without relying on high-power consumption equipment, the present invention transmits vehicle information back to the monitoring center through the radio frequency communication module, and is suitable for daily traffic monitoring and signal control optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of traffic monitoring, and in particular relates to a traffic target monitoring and tracking system and method based on monocular vision daytime scene reconstruction. Background Art

[0002] In the field of traffic monitoring, vehicle target detection is one of the key technologies in intelligent transportation systems. Traffic monitoring equipment mainly relies on radar and cameras, but current applications still suffer from problems such as instability, high energy consumption, and inaccuracy.

[0003] In previous work, radar technology (including millimeter-wave radar, lidar, and pulse radar) has excelled in detecting vehicle distance and speed due to its long detection range and ability to provide highly accurate target location and velocity information. Radar's electromagnetic waves can penetrate rain, snow, and fog, enabling stable monitoring in adverse weather conditions (such as rain, snow, and fog), and are unaffected by changes in lighting. However, radar struggles to acquire vehicle appearance features, such as vehicle type and license plate information. Therefore, while radar performs well in detecting a vehicle's physical location, it cannot provide specific information about its identity. Furthermore, radar is susceptible to multipath effects, such as ground reflections, making it prone to generating false targets. Furthermore, because detection primarily relies on the Doppler effect, radar performs poorly for detecting slow-moving or stationary vehicles. Radar's ability to resolve close-range vehicles is limited, making it particularly prone to target loss when dealing with dense traffic or small vehicles. This leads to inaccurate traffic flow information and increases the probability of missed detections, potentially posing a significant risk. Furthermore, radar involves high-frequency signal processing, resulting in high hardware costs that restrict its large-scale deployment.

[0004] Camera vision technology, on the other hand, uses image processing and object recognition techniques to capture a vehicle's exterior features (such as vehicle model, color, and license plate number) and track the vehicle through continuous image frames. However, monocular cameras cannot directly provide vehicle position and speed information, typically requiring multi-camera collaboration or additional calibration to achieve 3D reconstruction and positioning of the vehicle. Furthermore, camera performance degrades significantly in rainy, snowy, and low-light environments.

[0005] Lanzhou Jiaotong University published a method for vehicle identification and speed measurement based on video images in its paper, "Single-lens Video Vehicle Speed ​​Detection Method Based on Close-Range Photogrammetry." However, the paper's shortcomings are that, while the method is flexible and low-cost, its reliance on a single-lens camera can lead to fluctuations in accuracy under varying conditions, and the improved collinearity equations fail to fully account for complex road environments and dynamic changes.

[0006] Zhongyuan University of Technology disclosed a technology for measuring vehicle speed using multi-view cameras in its paper "Vehicle Video Speed ​​Measurement Method Based on Multi-View Cameras".

[0007] The disadvantage of this method is that although the multi-camera system improves detection accuracy and avoids the limitations of a single camera, it increases equipment cost and layout complexity. In particular, in large-scale deployments, additional resources may be required to maintain the stability and accuracy of the system.

[0008] Therefore, it is difficult for existing radar and camera detection technologies to achieve comprehensive and accurate traffic monitoring. In order to address the above problems, a lightweight traffic target detection and tracking system based on monocular vision scene reconstruction was developed. Summary of the Invention

[0009] To overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a traffic target monitoring and tracking system and method based on monocular vision daytime scene reconstruction. The system is based on a monocular camera, automatically detects lane lines during the day, and establishes a mapping relationship between pixel points and actual positions based on prior knowledge of road design. Through frame difference detection, contour tracking and region growing algorithms, vehicle feature information is extracted in a lightweight manner, and real-time vehicle trajectory tracking is achieved. Without relying on high-power equipment, vehicle information is transmitted back to the monitoring center via a radio frequency communication module, making it suitable for daily traffic monitoring and signal control optimization.

[0010] In order to achieve the above object, the technical solution adopted by the present invention is:

[0011] In one aspect of the present invention, a traffic target monitoring and tracking system based on monocular vision daytime scene reconstruction is provided, comprising a monocular camera, an inertial navigation module, a lightweight processor, a radio frequency communication module, a solar energy storage module, and a Beidou positioning module;

[0012] The monocular camera is installed on the side of the road and shoots traffic at an oblique angle to continuously capture multiple frames of images, providing raw data for vehicle detection and trajectory analysis.

[0013] The inertial navigation module uses a built-in high-precision gyroscope to measure the three-dimensional posture of the monocular camera when taking pictures and calculate the camera's external parameter matrix; providing accurate position information for subsequent data processing;

[0014] The lightweight processor is responsible for processing multiple frames of images captured by the camera, performing image preprocessing, vehicle detection and trajectory calculation, and executing core algorithms to ensure real-time response of the system;

[0015] The Beidou positioning module provides geographic location information, enabling vehicle trajectory information to be synchronized with the geographic location; thereby enhancing the system's adaptability in different road environments and ensuring seamless integration with other traffic management systems, enabling accurate monitoring of vehicle locations;

[0016] The radio frequency communication module transmits the detected vehicle trajectory information and vehicle positioning information back to the traffic monitoring center to ensure that the traffic flow information can be transmitted in real time, realizing real-time monitoring and management of traffic conditions;

[0017] The monocular camera, lightweight processor and radio frequency communication module are powered by a solar energy storage module.

[0018] Another aspect of the present invention provides a method for monitoring and tracking a traffic target, which is implemented using the traffic target monitoring and tracking system based on monocular vision daytime scene reconstruction of the present invention, and includes the following steps:

[0019] Step 1: Lane line detection;

[0020] A monocular camera captures road images in real time, a lightweight processor is used to detect lane lines, and the coordinates of the four vertices of the lane lines are extracted, thereby providing a reference for the transformation matrix from pixel coordinates to world coordinates.

[0021] Step 2: spatial coordinate system mapping;

[0022] After lane detection is complete, the lightweight processor uses the camera's three-axis angles provided by the inertial navigation module to assist in establishing a conversion equation. Based on standard lane design parameters, the detected lane points are used as known quantities to solve the conversion equation, thereby establishing a transformation matrix to convert pixel coordinates in the image into actual world coordinates.

[0023] Step 3: frame difference detection;

[0024] A monocular camera captures multiple consecutive frames of images to generate a static background image. Then, a differential operation is performed between the consecutive frames captured by the monocular camera and the static background image, and potential target areas are identified based on a variable threshold method.

[0025] Step 4: Contour tracing;

[0026] Based on the identified potential target area, seed pixels are selected and contour tracking is performed using the worm-follower method (tracking algorithm) to ensure the integrity and continuity of the vehicle area boundary. Contour tracking relies on the results of frame difference detection to further lock the target area.

[0027] Step 5: Target position and texture feature estimation;

[0028] The minimum coordinate value of the target area (i.e., the lower left pixel) is used as the center of mass of the vehicle, and the world coordinates of the vehicle are calculated through the transformation matrix. At the same time, the size, color and texture information of the vehicle are recognized through the camera image.

[0029] Step 6: Vehicle target tracking and result output;

[0030] The system combines the vehicle's location, color and size characteristics to continuously track and associate the same vehicle, generating a complete trajectory and related information, thus forming a closed loop. Finally, the system uses the radio frequency module to transmit the tracking data and Beidou positioning data back to the control center.

[0031] The step 1 is specifically as follows:

[0032] (1) Grayscale conversion and noise removal

[0033] First, the lightweight processor converts the RGB image input by the camera into a grayscale image. The converted grayscale image G(x p ,y p ) is binarized by the dynamic threshold method of the local mean, and the dynamic threshold T(x p ,y p ) is calculated as:

[0034] T(x p ,y p )=M(x p ,y p )*factor

[0035] Among them, M(x p ,y p ) is the local mean calculated by the mean filter, and factor is the scaling factor used to adjust the threshold sensitivity. The result of the binarization process generates a white area mask, which is used to identify possible lane line areas. Next, a morphological erosion operation is performed to remove noise and isolated white pixels to generate a binary mask.

[0036] (2) Hough transform and line segment detection

[0037] On the generated binary mask, Hough transform is used to detect lane segments, and the parameterization of the line is expressed in polar coordinate form:

[0038] ρ=x p cosθ+y p sinθ

[0039] Among them, ρ is the vertical distance from the line to the origin of the image, θ is the inclination angle of the line, and then the local maximum is found in the Hough space to detect the straight line segment l in the image k , thereby calibrating the position of the lane line.

[0040] Each detected straight line segment l k Calculate the direction vector v of each line segment defined by its endpoints (x1, y1) and (x2, y2) k :

[0041]

[0042] By calculating the direction vector, the system determines the direction and position of each line segment;

[0043] (3) Line segment clustering to form complete lane lines

[0044] After completing the line segment detection, the line segments with similar directions are merged through the clustering algorithm to generate a complete lane line. For each two line segments li and lj, by calculating their angle θ ij Determine whether they belong to the same cluster:

[0045]

[0046] If the angle between the two line segments is less than the set threshold θ threshold , they are considered to have similar directions and belong to the same cluster. For each line segment in a cluster, the minimum distance d between them is calculated. min To determine whether the line segments are close enough:

[0047]

[0048] If the minimum distance d min Less than the set distance threshold d threshold , then the two line segments are merged into the same cluster. In this way, multiple line segments with similar directions and positions are aggregated into a complete lane line;

[0049] Use the Bresenham algorithm to extend the line segment, starting from the starting point of the line segment and gradually extending the line segment according to the change of pixel position;

[0050] (4) Distinguishing between solid lanes and dashed lanes

[0051] Accurately distinguish solid lanes from dashed lanes based on cluster generation, cluster number determination, and pixel distribution trends.

[0052] The specific steps are as follows:

[0053] A. Cluster Generation and Quantity Determination

[0054] First, the lightweight processor calculates the Euclidean distance d between adjacent pixels. ij , to determine whether they have spatial continuity, for two pixel points p i =(x i,y i ) and p j =(x j ,y j ), if the distance between them satisfies the following conditions:

[0055]

[0056] They are considered to belong to the same cluster. In addition, according to the direction vector v of the line segment i and v j To determine the direction similarity, the direction vector is calculated by the coordinates of the starting and ending points of the line segment:

[0057]

[0058] If the angle θ between the direction vectors of the two line segments ij The following similarity conditions are met:

[0059]

[0060] Then these two line segments have directional similarity and belong to the same cluster;

[0061] Based on these rules, the system generates multiple clusters and then calculates the number of clusters N. clusters The number of clusters is used as the basis for preliminary judgment of lane line type. If the number of clusters exceeds the set threshold N dashed ,Lane lines are likely to be dashed lines, since dashed lines are usually segmented into multiple discontinuous clusters;

[0062] B. Monotonically decreasing detection of pixel distribution

[0063] For dashed lanes, sort the pixels in the cluster by the ordinate y and calculate the number of pixels n in each interval i (y), if the following roughly decreasing relationship is satisfied:

[0064] n i (y1)≥n i (y2)≥n i (y3)≥…≥n i (y k ), y1<y2<…<y k

[0065] This indicates that the pixel distribution conforms to the characteristics of the dotted lane; C. Comprehensive judgment

[0066] Combine the number of clusters and the pixel distribution trend to make a final judgment. If the number of clusters N clusters Exceeding the threshold N dashed , and the distribution of pixel points roughly shows a decreasing trend, then the lane line is marked as a dotted line;

[0067] Otherwise, the system marks the lane line as a solid line. The final judgment rules are as follows:

[0068]

[0069] By comprehensively analyzing cluster generation, cluster quantity, and pixel distribution trends, the system can more accurately distinguish between solid lanes and dashed lanes, ensuring detection accuracy.

[0070] The step 2 is specifically as follows:

[0071] (1) Create an image grid

[0072] Set the starting point and ending point of each dashed lane segment in the image coordinate system to be (x start ,y start ) and (x end ,y end ), for each dotted line segment, the system draws a horizontal line segment perpendicular to the direction of the line segment, and the system calculates the direction vector v of the dotted line segment dashed , whose direction vector is determined by the coordinates of the starting and ending points of the line segment:

[0073]

[0074] Based on this direction vector, calculate the direction vector v perpendicular to the dotted line segment ⊥ , the formula of this vector is:

[0075] v ⊥ =(-v dashed (2), v dashed (1)

[0076] Through the starting point and end point of each dotted line segment, draw multiple horizontal line segments in the direction perpendicular to the dotted line segment, and define the starting point and end point of the horizontal line segment as:

[0077] L horizontal =(x start ,y start )+t·v ⊥ ,in

[0078] At the same time, the system uses the detected solid lanes as longitudinal dividing lines, which are represented in the image as a set of continuous pixel points (x line ,y line ), which are regarded as longitudinal lines; the lane line detection results are used to construct the longitudinal boundaries of the grid. Through these longitudinal lines and the transverse lines generated by the dotted line segments, the system can divide the entire road area into multiple regular grids;

[0079] The generation of grid points is represented as a set of intersection points G, which consists of the intersection points of lane lines and transverse lines:

[0080] G={(x i ,y i )|L vertical ∩L horizontal}

[0081] Through the grid points, the system divides the road area into several grids, providing a reference framework for subsequent spatial mapping and vehicle detection;

[0082] (2) Establishing a mapping relationship

[0083] After the grid is generated, spatial mapping is performed using known road design standards. Road design standards include lane width, dashed lane segment length, and dashed line spacing parameters to estimate the precise location of longitudinal and transverse lines in real space.

[0084] (3) Full road mapping

[0085] After generating the preliminary mapping relationship of the local grid, the mapping of the local grid is extended to the entire road surface area.

[0086] The step 3 is specifically as follows:

[0087] First, a camera captures multiple frames of images to generate a static background image, which represents the road background without vehicles.

[0088] Then, the current image of each frame is compared with the background image, and the difference between the two is calculated to detect possible vehicles;

[0089] When the change value of the image exceeds the set threshold, the areas that may contain vehicles are identified;

[0090] In the daytime detection mode, firstly, multiple frames of images I1, I2, ..., I n Perform median operation to generate a target-free background image B(x, y), the formula is:

[0091] B(x,y)=Median(I1(x,y),I2(x,y),...,I n (x, y)

[0092] The background image B(x, y) represents the road background without vehicle targets. Then, the system obtains the current frame image I current (x, y), and select the sampling point P(x i ,y i), where i = 1, 2, ..., N represents the number of sampling points. The interval between these sampling points is smaller than the size of the vehicle to ensure that each vehicle target area contains at least one sampling point. The frame difference detection is performed by calculating the difference between only the selected sampling points. The calculation of the frame difference detection is as follows:

[0093] D(x i ,y i )=|I current (x i ,y i )-B(x i ,y i )|foreachi

[0094] When the difference D(x i ,y i ) exceeds a certain threshold T d When the system identifies the target sampling point P(x i ,y i ) may be within the vehicle's target area:

[0095] D(x i ,y i )>T d .

[0096] The step 4 is specifically as follows:

[0097] Identify the target sampling point P(x i ,y i ), the edge points in the target pixel block are selected as initial seeds, and the target contour is detected by applying the contour tracking algorithm of the insect follower;

[0098] Determine whether a pixel point (x, y) belongs to the target edge based on the following conditions:

[0099] |ΔI(x, y)|=|I current (x, y)-B(x, y)|>T d and Make |ΔI(x′, y′)|≤T d

[0100] Among them, T d is the difference threshold, N(x, y) represents the neighborhood of the pixel (x, y). The first condition ensures that the pixel is significantly different from the background, and the second condition ensures that at least one pixel in the neighborhood is slightly different from the background, thus confirming that the pixel is on the edge;

[0101] After successfully drawing the outline of the vehicle, the region growing method R(x, y) is used to classify the target by color. The region growing starts from the initial seed point S(x0, y0) of the outline and detects the color value of the adjacent pixels (x′, y′). If the color difference of the pixels meets the condition:

[0102] |I(x′,y′)-I(x0,y0)|<T c

[0103] The pixel is classified into the same category and the area continues to expand.

[0104] The step 5 is specifically as follows:

[0105] The chassis position of each vehicle target is determined by the bottom pixel point of its chassis (x min ,y min ) is determined by first identifying the pixel position, and then converting the chassis pixel from the image coordinate system to the actual space coordinate system using the following formula:

[0106]

[0107] After completing vehicle target detection, the system further extracts all pixel points of the vehicle target and estimates its color and size characteristics.

[0108] The specific processing steps are as follows:

[0109] First, all pixels within the identified vehicle outline are extracted. These pixels constitute the complete two-dimensional outline of the vehicle. The pixel set of the vehicle target is expressed as:

[0110]

[0111] Where n represents the total number of pixels of the vehicle target;

[0112] Next, the system estimates the color characteristics of the vehicle target. By counting all the pixels of each vehicle, the system calculates its average color value C vehicle , each pixel (x i ,y i ) is composed of three channel values ​​(R i , G i , B i ) is composed of the average color value C of the vehicle vehicle The calculation is as follows:

[0113]

[0114] Among them, R i , G i , B iRepresents pixels (x i ,y i )’s red, green, and blue channel values. Through this formula, the system can obtain the overall color feature C of the vehicle vehicle ;

[0115] Finally, the system estimates the size characteristics of the vehicle target by analyzing its boundary points. The width W of the vehicle is vehicle and height H vehicle The width and height are determined by the coordinates of the leftmost, rightmost, topmost and bottommost pixels respectively. The calculation formulas for width and height are as follows:

[0116] W vehicle =x max -x min

[0117] H vehicle =y max -y min

[0118] Among them, x max and x min are the horizontal coordinates of the rightmost and leftmost boundary points of the vehicle target contour, y max and y min are the ordinates of the upper and lower boundary points, respectively, so as to accurately estimate the two-dimensional outline size of the vehicle.

[0119] The step 6 is specifically as follows:

[0120] A. Feature Matching and Association

[0121] The system uses the least squares method to fit the vehicle's motion trajectory in the previous frames and calculates the predicted position of the vehicle in the current frame. Then, the system will calculate the actual measurement value of the current frame The vehicle is compared with the predicted value and associated based on the difference between the measured value and the predicted value. The matching criteria for vehicle association include matching of color, size and position. The matching criteria are as follows:

[0122]

[0123] Among them, ∈ C ,∈ W ,∈ H are the matching tolerances for color, width, and height, respectively; ΔD is the position difference threshold between the measured value and the predicted value; if all the above conditions are met, it is determined to be the detection result of the same target;

[0124] B. Target status update

[0125] The system updates the vehicle's trajectory. The status update includes trajectory smoothing filtering, velocity estimation, and position prediction. The detected trajectory is smoothed using a Kalman filter or a mean filter. The smooth update formula for the vehicle's position is as follows:

[0126]

[0127] Among them, α is the smoothing coefficient that controls the filter strength, which is calculated by calculating the position change ΔX between two consecutive frames of the vehicle. vehicle , ΔY vehicle As well as the time interval Δt, the speed of the vehicle can be expressed as follows:

[0128]

[0129] The system performs least squares fitting based on the vehicle positions of the previous frames to predict the vehicle position at the next moment t+1. The prediction model is:

[0130] X vehicle =At+b

[0131] where X vehicle is the historical position of the vehicle, t is the time series, A and b are fitting coefficients;

[0132] C. Trajectory Generation and Update

[0133] Through feature matching, target state update and position prediction, the system generates and updates the vehicle's complete trajectory T vehicle ,The vehicle trajectory includes the smoothed position of the vehicle in each frame, the ,speed estimation and related feature information.

[0134] Beneficial effects of the present invention:

[0135] The present invention uses a multi-view camera system to conduct experimental verification on different traffic scenes to ensure that the system can accurately locate vehicle trajectories.

[0136] This paper simulates actual traffic monitoring scenarios. In the experimental environment, the multi-view camera configuration can effectively evaluate the system's initialization capabilities in complex traffic environments and ensure the stability and accuracy of vehicle detection.

[0137] The system of the present invention can successfully establish a mapping relationship between pixel coordinates and actual space coordinates under various shooting angles and camera installation heights.

[0138] Whether shooting from a high altitude, left or right, or at a lower roadside angle, the present invention can complete accurate mapping matrix initialization in four-lane and six-lane test scenarios.

[0139] Whether the camera is deployed above the overpass or on the roadside, the system achieves extremely high accuracy in vehicle detection and tracking. In each case, the vehicle's pixel coordinates are accurately mapped to real-world coordinates, enabling precise detection and tracking of vehicle targets in a variety of complex road environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0140] Figure 1 It is a schematic diagram of the process algorithm of the present invention.

[0141] Figure 2 It is a schematic diagram of the situation when the present invention is working.

[0142] Figure 3 Schematic diagram of lane line mapping point calculation in the present invention.

[0143] Figure 4 It is a schematic diagram of converting image pixels into spatial positions according to the present invention.

[0144] Figure 5 Schematic diagram of vehicle trajectory obtained by the present invention.

[0145] Figure 6 It is a schematic diagram of the video acquisition module of the present invention. DETAILED DESCRIPTION

[0146] The present invention will be described in further detail below with reference to the accompanying drawings.

[0147] This traffic target monitoring and tracking system, based on monocular daytime scene reconstruction, uses a monocular camera to automatically detect lane markings during the day and establishes a mapping relationship between pixels and actual positions based on prior knowledge of road design. Through frame difference detection, contour tracking, and region growing algorithms, it efficiently extracts vehicle feature information and enables real-time vehicle trajectory tracking. This solution, without relying on high-power devices, transmits vehicle information back to a monitoring center via a wireless communication module, making it suitable for daily traffic monitoring and signal control optimization.

[0148] The present invention is suitable for real-time monitoring in traffic scenarios such as urban roads and intercity highways, especially those without power supply or requiring rapid deployment. With solar power and wireless communication, the system has high flexibility and low operating costs.

[0149] This system is designed to provide lightweight, low-power vehicle detection and data return functions. Figure 6 As shown in the figure, the hardware structure includes the following main parts:

[0150] (1) Monocular camera: The camera is the main data input device of the system and is used to collect video data. It is installed on the side of the road and uses an oblique angle to shoot traffic. It can continuously capture multiple frames of images and provide raw data for vehicle detection and trajectory analysis.

[0151] (2) Inertial navigation module: Through the built-in high-precision gyroscope, the three-dimensional posture of the camera when taking pictures can be measured, which is used to calculate the camera's external parameter matrix.

[0152] (3) Lightweight processor: The system uses a single-chip microcomputer as a lightweight processor to locally process the video data collected by the camera. The processor is responsible for image preprocessing, vehicle detection and trajectory calculation, and executes core algorithms to ensure real-time response.

[0153] (4) Radio Frequency Communication Module: In order to transmit the detected traffic flow information back to the traffic monitoring center, the system is equipped with a radio frequency communication module. This module supports long-distance wireless communication, ensuring that traffic flow information can be transmitted in real time, thus achieving real-time monitoring and management of traffic conditions.

[0154] (5) Solar Energy Storage Module: To achieve self-power, the system is equipped with solar panels and energy storage modules. The power required to operate the camera, processor, and communication module is provided by solar energy. This design reduces dependence on external power sources and is particularly suitable for traffic monitoring in remote or off-grid areas.

[0155] (6) Beidou Positioning Module: This module provides high-precision geographic location information, allowing traffic flow information to be synchronized to a specific geographic location. The introduction of the Beidou Positioning Module improves the system's adaptability to different road environments, ensuring that the system can seamlessly integrate with other traffic management systems and achieve accurate monitoring of traffic flow locations.

[0156] In summary, the system hardware uses modular design, combined with lightweight processors, solar power supply and radio frequency communication, to achieve real-time monitoring and data feedback of vehicles, and is suitable for different road environments.

[0157] The implementation process of the present invention is as follows Figure 1-Figure 5 The specific steps are as follows:

[0158] Step 1: Lane Detection

[0159] First, a monocular camera is used to capture real-time images of the road, an image processing algorithm is used to detect lane lines on the road, and edge detection technology is used to extract the positions of lane lines on the road. The lane line detection results will provide a reference benchmark for subsequent vehicle positioning, ensuring that the system can accurately locate the relative position of the vehicle based on the lane line information.

[0160] (1) Grayscale conversion and noise removal

[0161] First, the lightweight processor converts the RGB image input by the camera into a grayscale image to simplify processing and reduce computational complexity. The converted grayscale image G(x p ,y p ) is binarized by the dynamic threshold method of the local mean; the dynamic threshold T(x p ,y p ) is calculated as:

[0162] T(x p ,y p )=M(x p ,y p )*factor

[0163] Among them, M(x p ,y p ) is the local mean calculated by the mean filter, and factor is the scaling factor used to adjust the threshold sensitivity. The binarization results in a white area mask, which is used to identify possible lane marking areas. Next, a morphological erosion operation is performed to remove noise and isolated white pixels, ensuring a clean and coherent binary mask. The erosion operation processes the image using a 4×4 structuring element to remove isolated noise, thereby improving the accuracy of subsequent detection.

[0164] (2) Hough transform and line segment detection

[0165] On the generated binary mask, Hough transform is used to detect lane segments. Hough transform is an algorithm for detecting straight line features in an image. It represents the parameterization of the line in polar coordinate form:

[0166] ρ=x p cosθ+y p sinθ

[0167] Where ρ is the perpendicular distance from the line to the image origin, and θ is the line's inclination angle. We then detect straight line segments in the image by searching for local maxima in Hough space (these peaks correspond to parameterized representations of lines in the image). Since lane lines in road images are unique straight line segments, we calibrate the positions of lane lines by finding the most likely positions of the lines in the image.

[0168] The relationship between the binary mask image and the straight line segment is that the white areas in the binary image represent possible lane line positions. Through the Hough transform, these positions can be converted into polar coordinates to extract the features of the lane lines.

[0169] Furthermore, the image processing module sorts line segments by length, prioritizing those that are longer and more likely to represent lane markings. This is because long line segments more reliably reflect actual lane markings, improving the accuracy of subsequent processing. Therefore, during detection, only the longest line segments with the highest likelihood of representing lane markings are retained, ensuring accurate and effective detection.

[0170] Each detected straight line segment l k Defined by its endpoints (x1, y1) and (x2, y2). The system calculates the direction vector v for each line segment k :

[0171]

[0172] By calculating the direction vector, the system determines the direction and position of each line segment.

[0173] (3) Line segment clustering to form complete lane lines

[0174] After completing the line segment detection, the line segments with similar directions are merged through the clustering algorithm to generate a complete lane line. The goal of clustering is to group the detected line segments according to direction and position so that adjacent line segments with similar directions can be combined into a complete lane line. For each two line segments li and lj, by calculating their angle θ ij Determine whether they belong to the same cluster:

[0175]

[0176] If the angle between the two line segments is less than the set threshold θ threshold , they are considered to have similar directions and belong to the same cluster. For each line segment in a cluster, the minimum distance d between them is calculated. min To determine whether the line segments are close enough:

[0177]

[0178] If the minimum distance d min Less than the set distance threshold d threshold , then the two line segments can be merged into the same cluster. In this way, the system clustering module aggregates multiple line segments with similar directions and close positions into a complete lane line.

[0179] The system uses the Bresenham algorithm to extend line segments to cover more possible pixels and ensure the continuity of longer line segments. Specifically, the Bresenham algorithm approximates straight lines by efficiently selecting appropriate pixels on a two-dimensional grid. The image processing module starts from the starting point of the line segment and gradually extends it based on the changes in pixel positions, ensuring the continuity and directional consistency of the line segment throughout the process, thereby improving the accuracy and effectiveness of image processing.

[0180] (4) Distinguishing between solid lanes and dashed lanes

[0181] In order to accurately distinguish between solid lanes and dashed lanes, the system performs a comprehensive analysis based on cluster generation, cluster number determination, and pixel distribution trends.

[0182] The specific steps are as follows:

[0183] A. Cluster Generation and Quantity Determination

[0184] Cluster generation is based on the spatial continuity and directional similarity of pixels. First, the lightweight processor calculates the Euclidean distance d between adjacent pixels. ij , to determine whether they have spatial continuity. For two pixel points p i =(x i ,y i ) and p j =(x j ,y j ), if the distance between them satisfies the following conditions:

[0185]

[0186] They are considered to belong to the same cluster. In addition, the system also calculates the direction vector v of the line segment. i and v j To determine the direction similarity. The direction vector is calculated by the coordinates of the starting and ending points of the line segment:

[0187]

[0188] If the angle θ between the direction vectors of the two line segments ij The following similarity conditions are met:

[0189]

[0190] Then these two line segments have similar directions and belong to the same cluster. Based on these rules, the system generates multiple clusters. Then, the system calculates the number of clusters N clusters The number of clusters is used as the basis for preliminary judgment of lane line type. If the number of clusters exceeds the set threshold N dashed, the lane lines are likely to be dashed lines, since dashed lines are usually segmented into multiple discontinuous clusters.

[0191] B. Monotonically decreasing detection of pixel distribution

[0192] In addition to the number of clusters, the system also checks whether the distribution of pixels within each cluster follows a monotonically decreasing trend. For dashed lanes, the number of pixels within a cluster generally decreases as the distance from the camera increases. Therefore, the system sorts the pixels within the cluster by the ordinate y and calculates the number of pixels n in each interval. i (y). If the following approximate decreasing relationship is satisfied:

[0193] n i (y1)≥n i (y2)≥n i (y3)≥…≥n i (y k ), y1<y2<…<y k

[0194] This indicates that the pixel distribution conforms to the characteristics of a dashed lane. This decreasing relationship is not strictly required, but should generally show a trend in which the number of pixels decreases with increasing distance. For solid lanes, due to possible occlusion or other influencing factors, the pixel distribution of the cluster will generally not conform to this decreasing pattern.

[0195] C. Comprehensive judgment

[0196] The system makes a final judgment based on the number of clusters and pixel distribution trends. clusters Exceeding the threshold N dashed , and the distribution of pixels shows a roughly decreasing trend, the lane line is marked as a dotted line; otherwise, the system marks the lane line as a solid line. The final judgment rules are as follows:

[0197]

[0198] By comprehensively analyzing cluster generation, cluster quantity, and pixel distribution trends, the system can more accurately distinguish between solid lanes and dashed lanes, ensuring detection accuracy.

[0199] Step 2: Spatial coordinate system mapping

[0200] After lane detection is complete, the lightweight processor converts the image pixels captured by the camera into a physical space coordinate system. This process uses geometric transformation models (such as perspective transformation) based on the camera's internal and external parameters to map the 2D coordinates in the image to actual 3D coordinates.

[0201] (1) Create an image grid

[0202] Dashed lanes are composed of multiple line segments, each with a start and end point. To construct the image grid, the lightweight processor first generates multiple transverse segmentation lines based on the detected start and end points of the dashed lane segments. These transverse segmentation lines intersect with the lane's longitudinal solid and dashed lines, dividing the road area into a regular grid structure.

[0203] Specifically, assume that the starting point and ending point of each dashed lane segment are (x start ,y start ) and (x end ,y end ). For each dotted line segment, the system draws a horizontal line segment perpendicular to the direction of the line segment. In order to ensure the accuracy of the horizontal line segment, the system calculates the direction vector V of the dotted line segment. dashed , whose direction vector is determined by the coordinates of the starting and ending points of the line segment:

[0204]

[0205] Based on this direction vector, calculate the direction vector v perpendicular to the dotted line segment ⊥ , the formula of this vector is:

[0206] v ⊥ =(-v dashed (2), V dashed (1)

[0207] Through the starting point and ending point of each dotted line segment, draw multiple horizontal line segments in the direction perpendicular to the dotted line segment. The starting point and ending point of the horizontal line segment are defined as:

[0208] L horizontal =(x start ,y start )+t·v ⊥ ,in

[0209] At the same time, the system uses the detected solid lanes as longitudinal dividing lines. These lane lines are represented in the image as a set of continuous pixel points (x line ,y line ) are considered longitudinal lines. The lane line detection results are used to construct the longitudinal boundaries of the grid. Using these longitudinal lines and the transverse lines generated by the dashed segments, the system is able to divide the entire road area into multiple regular grids.

[0210] The generation of grid points is represented as a set of intersection points G, which consists of the intersection points of lane lines and transverse lines:

[0211] G={(x i ,y i )|L vertical ∩Lhorizontal}

[0212] Through the grid points, the system divides the road area into several grids, providing a reference framework for subsequent spatial mapping and vehicle detection.

[0213] The demonstration road section contains two solid lane lines and one dotted lane line. In the dotted lane, the system detects four dotted line segments. Based on these detection results, the system generates three longitudinal lines L vertical = {L1, L2, L3} (two solid lines and one dashed lane line). For each dashed line segment, the system calculates the lane line based on its starting point and ending point (x start ,y start ) and (x end ,y end ) generates 8 horizontal lines L perpendicular to the dotted line segment direction horizontal , that is, draw a horizontal line through the starting point and the end point of each dotted line segment. Through the intersection of the vertical line and the horizontal line, the system generates a total of 24 landmark points G = {(x i ,y i )}, these landmarks are marked by the longitudinal line L vertical and transverse line L horizontal Finally, these landmark points generate 21 landmark quadrilaterals, with every four adjacent landmark points forming a grid unit. These grid units provide a precise reference frame for the subsequent mapping of pixel coordinates to spatial coordinates.

[0214] (2) Establishing a mapping relationship

[0215] After the image mesh is generated, it is spatially mapped using known road design criteria. These criteria include parameters such as lane width, length of dashed lane segments, and spacing between dashed lines, which are used to estimate the precise location of longitudinal and transverse lines in real space.

[0216] For example, on roads with speeds below 60 km / h, dashed lane segments are 2 meters long, 15 cm wide, and spaced 4 meters apart. Based on these known road parameters, the system derives the relative spatial position of each intersection in the image grid, thereby deriving a preliminary mapping relationship for the local grid.

[0217] (3) Full road mapping

[0218] After generating the preliminary mapping relationship of the local grid, the mapping of the local grid is extended to the entire road surface area.

[0219] Specifically, the area within the solid lane line is first extracted as the main detection road surface area. By identifying the road boundary defined by the solid lane line, the system can accurately define the road surface mapping range. Using the bilinear interpolation method, the system generates a large number of mapping pairs between pixel coordinates and actual space coordinates. Assuming that the pixel coordinates of the four vertices of the grid are (x1, y1), (x2, y2), (x3, y3), (x4, y4), the corresponding actual space coordinates are (X1, Y1), (X2, Y2), (X3, Y3), (X4, Y4). For any pixel point (x, y) in the grid, its actual space coordinates (X, Y) can be calculated using the following bilinear interpolation formula:

[0220]

[0221] These mapping pairs will serve as the basis for the subsequent solution of the perspective transformation matrix H to achieve accurate mapping of the entire road surface.

[0222] Step 3: Frame Difference Detection

[0223] First, a camera captures multiple frames of images to generate a static background image, which represents the road background without vehicles.

[0224] Next, the system compares each frame's current image with the background image, calculating the difference between the two to detect possible vehicles. When the image variance exceeds a set threshold, areas likely to contain vehicles are identified. This step reduces the amount of image data to be processed, effectively reducing computational effort and enabling rapid localization of potential vehicle targets.

[0225] In the daytime detection mode, firstly, multiple frames of images I1, I2, ..., I n Perform median operation to generate a target-free background image B(x, y), the formula is:

[0226] B(x,y)=Median(I1(x,y),I2(x,y),...,I n (x, y)

[0227] The background image B(x, y) represents the road background without vehicle targets. Then, the system obtains the current frame image I current (x, y), and select the sampling point P(x i ,y i ), where i = 1, 2, ..., N represents the number of sampling points. The interval between these sampling points is smaller than the size of the vehicle to ensure that each vehicle target area contains at least one sampling point. Frame difference detection reduces the amount of calculation by only calculating the differences between the selected sampling points, achieving lightweight processing. The calculation of frame difference detection is as follows:

[0228] D(x i ,y i )=|I current (x i ,y i )-B(x i ,y i )|foreachi

[0229] When the difference D(x i ,y i ) exceeds a certain threshold T d When the system identifies the sampling point P(x i ,y i ) may be within the vehicle's target area:

[0230] D(x i ,y i )>T d

[0231] At this stage, the system only considers the target sampling point P(x i ,y i ), thereby significantly reducing the amount of calculation, achieving coarse positioning and optimizing system performance.

[0232] Step 4: Contour Tracing

[0233] After identifying these target sampling points, edge points within the target pixel block are selected as initial seeds, and the target contour is detected using a worm-following contour tracking algorithm. The core principle of contour tracking is to iteratively detect the contour of the region until a closed loop is formed. Specifically, whether a pixel point (x, y) belongs to the target edge is determined based on the following conditions:

[0234] |ΔI(x, y)|=|I current (x, y)-B(x, y)|>T d and Make |ΔI(x′, y′)|≤T d

[0235] Among them, T d is the difference threshold, N(x, y) represents the neighborhood of the pixel (x, y). The first condition ensures that the pixel is significantly different from the background, and the second condition ensures that at least one pixel in the neighborhood is slightly different from the background, thus confirming that the pixel is on the edge.

[0236] Through these two combined criteria, the system can effectively extract the edge contour of the vehicle. The contour tracing process iteratively tracks adjacent pixels, gradually depicting the complete boundary and ensuring the vehicle's contour is continuous.

[0237] This method can not only effectively identify a single vehicle, but also handle contour extraction when multiple vehicles overlap or are occluded.

[0238] After successfully drawing the outline of the vehicle, the region growing method R(x, y) is used to classify the target by color. The region growing starts from the initial seed point S(x0, y0) of the outline and detects the color value of the adjacent pixels (x′, y′). If the color difference of the pixels meets the condition:

[0239] |I(x′,y′)-I(x0,y0)|<T c

[0240] The pixel is classified into the same category and the area continues to expand.

[0241] Region growing is very effective in dealing with situations where vehicles are close in position but have different colors.

[0242] Step 5: Target position and texture feature estimation

[0243] After using the contour tracking algorithm to obtain the complete vehicle area, in order to understand the actual displacement of the vehicle in space, it is necessary to convert the image coordinate system to the actual space coordinate system. According to the pixel point set defined by the contour tracking algorithm, since the establishment of the spatial mapping relationship is based on the road lane line, we can know that among all the pixel points, the actual space coordinate corresponding to the pixel in the lower left corner of the coordinate system is closest to the road surface and has the smallest distortion. It can be expressed in the pixel point set as (x min ,y min ).

[0244] The chassis position of each vehicle target is determined by the bottom pixel point of its chassis (x min ,y min ). First, identify the pixel position, and then convert the chassis pixel from the image coordinate system to the actual space coordinate system using the following formula:

[0245]

[0246] After completing vehicle target detection, the system further extracts all the pixels of the vehicle target and estimates its color and size characteristics. The specific processing steps are as follows:

[0247] First, all pixels within the identified vehicle outline are extracted. These pixels constitute the complete two-dimensional outline of the vehicle. The pixel set of the vehicle target can be expressed as:

[0248]

[0249] Among them, n represents the total number of pixels of the vehicle target. Through this set, the system completely describes the outline of the vehicle;

[0250] Next, the system estimates the color characteristics of the vehicle target. By counting all the pixels of each vehicle, the system calculates its average color value C vchicle Each pixel (x i ,y i ) is composed of three channel values ​​(R i , G i , B i ), so the average color value of the vehicle is C vehicle The calculation is as follows:

[0251]

[0252] Among them, R i , G i , B i Represents pixels (x i ,y i ) The red, green, and blue channel values ​​of the vehicle. Through this formula, the system can obtain the overall color characteristics C of the vehicle vehicle , providing an important basis for subsequent vehicle classification and identification.

[0253] Finally, the system estimates the size characteristics of the vehicle by analyzing its boundary points. The width of the vehicle W vehicle and height H vehicle The width and height are determined by the coordinates of the leftmost, rightmost, topmost and bottommost pixels respectively. The calculation formulas for width and height are as follows:

[0254] W vehicle =x max -x min

[0255] H vehicle =y max -y min

[0256] Among them, x max and x min are the horizontal coordinates of the rightmost and leftmost boundary points of the vehicle target contour, y max and y min are the vertical coordinates of the upper and lower boundary points respectively. Through these formulas, the system can accurately estimate the two-dimensional contour size of the vehicle.

[0257] Through the above steps, the system can accurately extract the color and size characteristics of the vehicle, providing accurate basic information for further vehicle identification and tracking.

[0258] Step 6: Vehicle target tracking and result output

[0259] A. Feature Matching and Association

[0260] In order to track and associate the same vehicle in multiple frames of vehicle detection, the system combines the vehicle's location information, color features, and size features. Specifically, the system uses the least squares method to fit the vehicle's motion trajectory in the previous frames and calculates the vehicle's predicted position in the current frame. Then, the system will calculate the actual measurement value of the current frame The measured value is compared with the predicted value and the association is performed based on the difference between the measured value and the predicted value. The matching criteria for vehicle association include matching of color, size and position. The matching criteria are as follows:

[0261]

[0262] Among them, ∈ C ,∈ W ,∈ H are the matching tolerances for color, width, and height, respectively, and ΔD is the position difference threshold between the measured value and the predicted value. If all of the above conditions are met, it is determined to be the detection result of the same target. Through this association mechanism, the system can track the movement trajectory of the same vehicle in multiple consecutive frames of images. B. Target Status Update

[0263] To further improve the accuracy and stability of vehicle tracking, the system updates the vehicle's trajectory. This state update includes trajectory smoothing filtering, velocity estimation, and position prediction. Due to detection errors and noise, directly using the measured position may lead to unstable vehicle trajectories. Therefore, the system uses a Kalman filter or mean filter to smooth the detected trajectory to ensure a smoother vehicle trajectory. The formula for smoothing the vehicle position is as follows:

[0264]

[0265] Where α is the smoothing coefficient that controls the filter strength. Based on the smoothing result of the vehicle trajectory, the system can estimate the vehicle's speed. Specifically, by calculating the position change ΔX between two consecutive frames vehicle , ΔY vehicle As well as the time interval Δt, the speed of the vehicle can be expressed as follows:

[0266]

[0267] Speed ​​estimation provides an important basis for subsequent traffic analysis and can also help determine the behavior characteristics of vehicles. In addition, the system uses the least squares fitting method to predict the vehicle's position at the next moment t+1 by using the vehicle position of the previous frames. The prediction model is:

[0268] X vehicle =At+b

[0269] where X vehicle is the historical position of the vehicle, t is the time series, A and b are fitting coefficients. Through position prediction, the system can further improve the accuracy of vehicle tracking.

[0270] C. Trajectory Generation and Update

[0271] Through feature matching, target state update and position prediction, the system generates and updates the vehicle's complete trajectory T vehicle The vehicle trajectory includes the smoothed position, velocity estimation, and related feature information of the vehicle in each frame, such as color C vehicle and size (W vehicle , H vehicle This output provides an accurate data foundation for vehicle behavior analysis, traffic monitoring, and control. This method not only enables the system to track the dynamic position of vehicles in real time but also provides multi-dimensional information, providing solid support for vehicle detection, tracking, and prediction in intelligent transportation systems.

[0272] The device developed by the present invention can achieve real-time detection and tracking of traffic targets through an efficient algorithm without adding additional hardware equipment, and is suitable for vehicle detection under different lighting conditions during the day.

[0273] Application of the present invention:

[0274] 1) Applied to temporary vehicle detection and speed monitoring in construction areas.

[0275] Due to the complex conditions on construction roads, traditional speed measurement equipment requires power grids and complex installation procedures. However, the lightweight monocular vision inspection system of this invention can be quickly deployed in construction areas without relying on a power grid, enabling scene construction and vehicle detection based on lane markings. This system monitors the speed of vehicles passing through the construction area in real time and provides detailed vehicle information, such as license plates and vehicle models, to ensure safety within the construction area. Furthermore, because the equipment is easily movable and disassembled, it is suitable for frequent deployment and removal within a short period of time, flexibly responding to the needs of different construction phases.

[0276] 2) Intelligent traffic guidance and early warning system for construction areas.

[0277] The system not only monitors vehicle speed but also uses cameras to capture detailed information about vehicles, such as vehicle type and volume, and uses lane information to guide vehicles through the road in an orderly manner. When the system detects speeding, illegal driving, or unusual behavior, it issues real-time warnings, prompting drivers to slow down or change routes, thereby improving safety and traffic efficiency in construction zones.

[0278] 3) Environmental assessment and traffic data analysis of the construction area.

[0279] During construction, the system continuously collects data such as vehicle speed, volume, and route. Based on this data, it analyzes traffic conditions in the construction area, assessing traffic safety and pressure points. This long-term data collection allows construction managers to optimize construction plans, rationalize traffic guidance measures, and improve construction efficiency and safety.

Claims

1. A traffic target monitoring and tracking method, characterized in that: The following steps are included: Step 1: Lane line detection; A monocular camera captures road images in real time, a lightweight processor is used to detect lane lines, and the coordinates of the four vertices of the lane lines are extracted, thereby providing a reference for the transformation matrix from pixel coordinates to world coordinates. Step 2: spatial coordinate system mapping; After lane detection is complete, the lightweight processor uses the camera's three-axis angles provided by the inertial navigation module to assist in establishing a conversion equation. Based on standard lane design parameters, the detected lane points are used as known quantities to solve the conversion equation, thereby establishing a transformation matrix to convert pixel coordinates in the image into actual world coordinates. Step 3: frame difference detection; A monocular camera captures multiple consecutive frames of images to generate a static background image. Then, a differential operation is performed between the consecutive frames captured by the monocular camera and the static background image, and potential target areas are identified based on a variable threshold method. Step 4: Contour tracing; Based on the identified potential target area, seed pixels are selected and contour tracking is performed using the worm-following method (tracking algorithm) to ensure the integrity and continuity of the vehicle area boundary; Contour tracking relies on the results of frame difference detection and can further lock the target area; Step 5: Target position and texture feature estimation; The minimum coordinate value of the target area is used as the center of mass of the vehicle, and the world coordinates of the vehicle are calculated through the transformation matrix. At the same time, the size, color and texture information of the vehicle are recognized through the camera image; Step 6: Vehicle target tracking and result output; The system continuously tracks and associates the same vehicle based on its location, color, and size, generating a complete trajectory and related information, thus forming a closed loop. The system then uses a radio frequency module to transmit tracking data and Beidou positioning data back to the control center. The method is implemented through a traffic target monitoring and tracking system based on monocular vision daytime scene reconstruction. The system includes a monocular camera, an inertial navigation module, a lightweight processor, a radio frequency communication module, a solar energy storage module, and a Beidou positioning module. The monocular camera is installed on the side of the road and shoots traffic at an oblique angle to continuously capture multiple frames of images, providing raw data for vehicle detection and trajectory analysis. The inertial navigation module uses a built-in high-precision gyroscope to measure the three-dimensional posture of the monocular camera when taking pictures, calculates the camera's external parameter matrix, and provides accurate position information for subsequent data processing; The lightweight processor is responsible for processing multiple frames of images captured by the monocular camera, performing image preprocessing, vehicle detection and trajectory calculation, and executing core algorithms to ensure real-time response of the system; The Beidou positioning module provides geographic location information so that vehicle trajectory information can be synchronized with the geographic location; The radio frequency communication module transmits the detected vehicle trajectory information and vehicle positioning information back to the traffic monitoring center to ensure that the traffic flow information can be transmitted in real time, realizing real-time monitoring and management of traffic conditions; The monocular camera, lightweight processor and radio frequency communication module are powered by a solar energy storage module.

2. A traffic target monitoring and tracking method according to claim 1, characterized in that: The step 1 is specifically as follows: (1) Grayscale conversion and noise removal First, the lightweight processor converts the RGB image input by the camera into a grayscale image. The converted grayscale image G(x p ,y p ) is binarized by the dynamic threshold method of the local mean, and the dynamic threshold T(x p ,y p ) is calculated as: T(x p ,y p )=M(x p ,y p )*factor Among them, M(x p ,y p ) is the local mean calculated by the mean filter, and factor is the scaling factor used to adjust the threshold sensitivity. The result of the binarization process generates a white area mask, which is used to identify possible lane line areas. Next, a morphological erosion operation is performed to remove noise and isolated white pixels to generate a binary mask. (2) Hough transform and line segment detection On the generated binary mask, Hough transform is used to detect lane segments, and the parameterization of the line is expressed in polar coordinate form: p=x p cosθ+y p sinth Among them, ρ is the vertical distance from the line to the origin of the image, θ is the inclination angle of the line, and then the local maximum is found in the Hough space to detect the straight line segment l in the image k , thereby calibrating the position of the lane line; Each detected straight line segment l k Calculate the direction vector v of each line segment defined by its endpoints (x1, y1) and (x2, y2) k : By calculating the direction vector, the system determines the direction and position of each line segment; (3) Line segment clustering to form complete lane lines After completing the line segment detection, the line segments with similar directions are merged through the clustering algorithm to generate a complete lane line. For each two line segments li and lj, by calculating their angle θ ij Determine whether they belong to the same cluster: If the angle between the two line segments is less than the set threshold θ threshold , they are considered to have similar directions and belong to the same cluster. For each line segment in a cluster, the minimum distance d between them is calculated. min To determine whether the line segments are close enough: If the minimum distance d min Less than the set distance threshold d threshold , then the two line segments are merged into the same cluster. In this way, multiple line segments with similar directions and positions are aggregated into a complete lane line; Use the Bresenham algorithm to extend the line segment, starting from the starting point of the line segment and gradually extending the line segment according to the change of pixel position; (4) Distinguishing between solid lanes and dashed lanes Accurately distinguish solid lanes from dashed lanes based on cluster generation, cluster number determination, and pixel distribution trends.

3. A traffic target monitoring and tracking method according to claim 2, characterized in that: The specific steps of step (4) are as follows: A. Cluster Generation and Quantity Determination First, the lightweight processor calculates the Euclidean distance d between adjacent pixels. ij , to determine whether they have spatial continuity, for two pixel points p i =(x i ,y i ) and p j =(x j ,y j ), if the distance between them satisfies the following conditions: They are considered to belong to the same cluster. In addition, according to the direction vector v of the line segment i and v j To determine the direction similarity, the direction vector is calculated by the start and end coordinates of the line segment: If the angle θ between the direction vectors of the two line segments ij The following similarity conditions are met: Then these two line segments have directional similarity and belong to the same cluster; Based on these rules, the system generates multiple clusters and then calculates the number of clusters N. clusters The number of clusters is used as the basis for preliminary judgment of lane line type. If the number of clusters exceeds the set threshold N dashed ,Lane lines are likely to be dashed lines, since dashed lines are usually segmented into multiple discontinuous clusters; B. Monotonically decreasing detection of pixel distribution For dashed lanes, sort the pixels in the cluster by the ordinate y and calculate the number of pixels n in each interval i (y), if the following roughly decreasing relationship is satisfied: <h2 style=";text-align:left;direction:ltr">n<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (y1)≥n<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (y2)≥n<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (y3)≥…≥n<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (y<h2 style=";text-align:left;direction:ltr"> k <h2 style=";text-align:left;direction:ltr"> ),y1 <y2<…<y <h2 style=";text-align:left;direction:ltr"> k This indicates that the pixel distribution conforms to the characteristics of the dotted lane; C. Comprehensive judgment Combine the number of clusters and the pixel distribution trend to make a final judgment. If the number of clusters N clusters Exceeding the threshold N dashed , and the distribution of pixel points roughly shows a decreasing trend, then the lane line is marked as a dotted line; Otherwise, the system marks the lane line as a solid line. The final judgment rules are as follows: By comprehensively analyzing cluster generation, cluster quantity, and pixel distribution trends, the system can more accurately distinguish between solid lanes and dashed lanes, ensuring detection accuracy.

4. A traffic target monitoring and tracking method according to claim 3, characterized in that: The step 2 is specifically as follows: (1) Create an image grid Set the starting point and ending point of each dashed lane segment in the image coordinate system to be (x start ,y start ) and (x end ,y end ), for each dotted line segment, the system draws a horizontal line segment perpendicular to the direction of the line segment, and the system calculates the direction vector v of the dotted line segment dashed , whose direction vector is determined by the coordinates of the starting and ending points of the line segment: Based on this direction vector, calculate the direction vector v perpendicular to the dotted line segment ⊥ , the formula of this vector is: v ⊥ =(-v dashed (2),v dashed (1)) Through the starting point and end point of each dotted line segment, draw multiple horizontal line segments in the direction perpendicular to the dotted line segment, and define the starting point and end point of the horizontal line segment as: L horizontal =(x start ,y start )+t·v ⊥ ,in At the same time, the system uses the detected solid lanes as longitudinal dividing lines, which are represented in the image as a set of continuous pixel points (x line ,y line ), which are regarded as longitudinal lines; the lane line detection results are used to construct the longitudinal boundaries of the grid. Through these longitudinal lines and the transverse lines generated by the dotted line segments, the system can divide the entire road area into multiple regular grids; The generation of grid points is represented as a set of intersection points G, which consists of the intersection points of lane lines and transverse lines: G={(x i ,y i )∣L vertical ∩L horizontal } Through the grid points, the system divides the road area into several grids, providing a reference framework for subsequent spatial mapping and vehicle detection; (2) Establishing a mapping relationship After the grid is generated, spatial mapping is performed using known road design standards. Road design standards include lane width, dashed lane segment length, and dashed line spacing parameters to estimate the precise location of longitudinal and transverse lines in real space. (3) Full road mapping After generating the preliminary mapping relationship of the local grid, the mapping of the local grid is extended to the entire road surface area.

5. A traffic target monitoring and tracking method according to claim 4, characterized in that: The step 3 is specifically as follows: First, a camera captures multiple frames of images to generate a static background image, which represents the road background without vehicles. Then, the current image of each frame is compared with the background image, and the difference between the two is calculated to detect possible vehicles; When the change value of the image exceeds the set threshold, the areas that may contain vehicles are identified; In the daytime detection mode, firstly, multiple frames of images I1, I2, ..., I n Perform median operation to generate a target-free background image B(x, y), the formula is: B(x,y)=Median(I1(x,y),I2(x,y),…,I n (x,y)) The background image B(x, y) represents the road background without vehicle targets. Then, the system obtains the current frame image I current (x, y), and select the sampling point P(x i ,y i ), where i = 1, 2, ..., N represents the number of sampling points. The interval between these sampling points is smaller than the size of the vehicle to ensure that each vehicle target area contains at least one sampling point. The frame difference detection is performed by calculating the difference between the selected sampling points. The calculation of the frame difference detection is as follows: D(x i ,y i )=|I current (x i ,y i )-B(x i ,y i )|foreachi When the difference D(x i ,y i ) exceeds a certain threshold T d When the system identifies the target sampling point P(x i ,y i ) may be within the vehicle's target area: D(x i ,y i )>T d 。 6. A traffic target monitoring and tracking method according to claim 5, characterized in that: The step 4 is specifically as follows: Identify the target sampling point P(x i ,y i ), the edge points in the target pixel block are selected as initial seeds, and the target contour is detected by applying the contour tracking algorithm of the insect follower; Determine whether a pixel point (x, y) belongs to the target edge based on the following conditions: |ΔI(x, y)| = |I current (x, y) - B(x, y)| > T d and such that |ΔI(x′, y′)| ≤ T d Among them, T d is the difference threshold, N(x, y) represents the neighborhood of the pixel (x, y). The first condition ensures that the pixel is significantly different from the background, and the second condition ensures that at least one pixel in the neighborhood is slightly different from the background, thus confirming that the pixel is on the edge; After successfully drawing the outline of the vehicle, the region growing method R(x, y) is used to classify the target by color. The region growing starts from the initial seed point S(x0, y0) of the outline and detects the color value of the neighboring pixels (x′, y′). If the color difference of the pixels meets the conditions: |I(x′,y′)-I(x0,y0)|<T c The pixel is classified into the same category and the area continues to expand.

7. A traffic target monitoring and tracking method according to claim 6, characterized in that: The step 5 is specifically as follows: The chassis position of each vehicle target is determined by the bottom pixel point of its chassis (x min ,y min ) is determined by first identifying the pixel position, and then converting the chassis pixel from the image coordinate system to the actual space coordinate system using the following formula: After completing vehicle target detection, the system further extracts all pixel points of the vehicle target and estimates its color and size characteristics.

8. A traffic target monitoring and tracking method according to claim 7, characterized in that: The specific processing steps are as follows: First, all the pixels within the identified vehicle outline are extracted. These pixels constitute the complete two-dimensional outline of the vehicle. The pixel set of the vehicle target is expressed as: Where n represents the total number of pixels of the vehicle target; Next, the system estimates the color characteristics of the vehicle target. By counting all the pixels of each vehicle, the system calculates its average color value C vehicle , each pixel (x i ,y i ) is composed of three channel values ​​(R i , G i , B i ) is composed of the average color value C of the vehicle vehicle The calculation is as follows: Among them, R i , G i , B i Represents pixels (x i ,y i )’s red, green, and blue channel values. Through this formula, the system can obtain the overall color feature C of the vehicle vehicle ; Finally, the system estimates the size characteristics of the vehicle target by analyzing its boundary points. The width W of the vehicle is vehicle and height H vehicle The width and height are determined by the coordinates of the leftmost, rightmost, topmost and bottommost pixels respectively. The calculation formulas for width and height are as follows: W vehicle =x max -x min H vehicle =y max -y min Among them, x max and x min are the horizontal coordinates of the rightmost and leftmost boundary points of the vehicle target contour, y max and y min are the ordinates of the upper and lower boundary points, respectively, so as to accurately estimate the two-dimensional outline size of the vehicle.

9. A traffic target monitoring and tracking method according to claim 8, characterized in that: The step 6 is specifically as follows: A. Feature Matching and Association The system uses the least squares method to fit the vehicle's motion trajectory in the previous frames and calculates the predicted position of the vehicle in the current frame. Then, the system will calculate the actual measurement value of the current frame The vehicle is compared with the predicted value and associated based on the difference between the measured value and the predicted value. The matching criteria for vehicle association include matching of color, size and position. The matching criteria are as follows: Among them, ∈ C ,∈ W ,∈ H are the matching tolerances for color, width, and height, respectively; ΔD is the position difference threshold between the measured value and the predicted value; if all the above conditions are met, it is determined to be the detection result of the same target; B. Target status update The system updates the vehicle's trajectory. The status update includes trajectory smoothing filtering, velocity estimation, and position prediction. The detected trajectory is smoothed using a Kalman filter or a mean filter. The smooth update formula for the vehicle's position is as follows: Among them, α is the smoothing coefficient that controls the filter strength, which is calculated by calculating the position change ΔX between two consecutive frames of the vehicle. vehicle , ΔY vehicle As well as the time interval Δt, the speed of the vehicle can be expressed as follows: The system performs least squares fitting based on the vehicle positions of the previous frames to predict the vehicle position at the next moment t+1. The prediction model is: X vehicle =At+b where X vehicle is the historical position of the vehicle, t is the time series, A and b are fitting coefficients; C. Trajectory Generation and Update Through feature matching, target state update and position prediction, the system generates and updates the vehicle's complete trajectory T vehicle ,The vehicle trajectory includes the smoothed position of the vehicle in each frame, the ,speed estimation and related feature information.

Citation Information

Patent Citations

  • Monocular camera outdoor three-dimensional reconstruction method based on full convolutional neural network

    CN110060331A

  • Map making method and device thereof, electronic equipment and computer readable storage medium

    CN111105695A