A traffic object positioning method based on roadside camera and latitude and longitude registration

Through the roadside camera combining image background information entropy and perspective transformation, the latitude and longitude problems of traffic object positioning are solved, accurate positioning of traffic object and vehicle status judgment are achieved, and vehicle perception ability and safety are improved.

CN116935336BActive Publication Date: 2025-08-15CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310862427.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-08-15
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

The existing roadside visual object detection methods cannot accurately obtain the latitude and longitude information of traffic objects, and it is difficult for the vehicle to judge the location and status of the surrounding traffic objects, resulting in limited vehicle perception range and increased safety hazards.

Method used

The latitude and longitude of traffic objects are calculated by using an adaptive threshold SIFT and FLANN based on the image background information entropy of roadside cameras, combining the latitude and longitude registration after image viewing angle transformation and the motion trajectory estimation method based on vehicle template matching.

Benefits of technology

It realizes real-time acquisition of traffic objects' location information on the roadside, expands the vehicle's perception range, provides data support for collision warning, and supports vehicle driving trajectory tracking and traffic flow prediction, improving the accuracy and safety of traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935336B_ABST
    Figure CN116935336B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for locating traffic objects based on roadside cameras and longitude and latitude registration, and belongs to the field of intelligent transportation. The present invention is divided into the following three parts: in response to the problem that the shaking of the roadside camera mounting rod causes the target pixel offset in the camera, a region of interest positioning method combining adaptive threshold SIFT and FLANN based on image background information entropy is designed and proposed to match image feature points to locate the position of the lane line ROI (region of interest) object in the image; in response to the problem that the longitude and latitude of the traffic object cannot be obtained after camera target detection in the roadside perception environment, a longitude and latitude registration method after image perspective transformation is designed and proposed to calculate the longitude and latitude of the perceived traffic object; in response to the problem that the roadside camera cannot obtain the precise position of the front of the vehicle on the road section through target detection, a front position estimation method based on vehicle template matching and motion trajectory is designed and proposed to obtain the longitude and latitude of the perceived vehicle front.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation and relates to a traffic object positioning method based on roadside camera and longitude and latitude registration. Background Art

[0002] Roadside image perception and object positioning are crucial components of intelligent traffic object identification, positioning, tracking, and vehicle-infrastructure collaborative perception. To achieve truly high-level autonomous driving and assisted safety, relying solely on vehicle-side perception is easily limited by road conditions and sensor performance, preventing the vehicle from obtaining information about a wide range of traffic participants, which poses a safety hazard. Therefore, implementing roadside perception and vehicle-infrastructure collaboration is essential. Roadside perception is a crucial component of vehicle-infrastructure collaborative application development. By deploying sensors on the roadside and transmitting collected road surface information to the vehicle via V2X communication, the vehicle possesses beyond-line-of-sight perception capabilities. As the primary method for roadside perception, visual object detection uses cameras to perceive information about the road's surroundings, including traffic objects, traffic signs, and road markings. By determining and estimating the vehicle's position, direction of travel, and other status, it enables cloud-based monitoring and management of traffic objects and events, while also helping the vehicle make accurate and safe decisions.

[0003] However, existing roadside visual target detection solutions are mostly based on deep learning technology, which cannot accurately obtain the latitude and longitude position information of traffic objects. Although the target detection method based on millimeter wave and visual fusion can detect the latitude and longitude information of traffic objects, it is mainly effective for moving detection targets. Even if the detected traffic object information is sent to the moving vehicle, it is difficult for the vehicle to determine the position and status of the surrounding traffic objects. The present invention proposes a positioning method based on roadside camera and latitude and longitude registration, which is used to obtain the position information of traffic objects on the roadside in real time. On the one hand, it expands the perception range of moving vehicles and provides data support for vehicle collision warning; on the other hand, it provides a data source for analysis of vehicle cross-domain trajectory tracking, traffic flow prediction, and urban development planning. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide a traffic object positioning method based on roadside camera and longitude and latitude registration.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A traffic object positioning method based on roadside camera and latitude and longitude registration, which includes a region of interest positioning method based on image background information entropy using adaptive threshold SIFT and FLANN, a latitude and longitude registration method after image perspective transformation, and a vehicle head position estimation method based on vehicle-type template matching and motion trajectory;

[0007] The region of interest (ROI) localization method, which combines adaptive threshold SIFT and FLANN based on image background information entropy, grayscales roadside images, extracts feature points using adaptive threshold SIFT, and then uses FLANN feature point matching to locate the lane markings on both sides of the road in the image. This method determines the lane ROI and addresses the issue of lane ROI offset caused by shaking of the roadside camera mounting poles due to typhoons and rainstorms.

[0008] Based on the latitude and longitude registration method after image perspective transformation, the lane line ROI in the image is transformed into a bird's-eye view. Then, the electronic fence grid is divided and the longitude and latitude coordinates are aligned. The longitude and latitude coordinates corresponding to any pixel point in the lane line are calculated. The vehicle traffic objects are classified, the class template is matched, and the direction angle of the motion trajectory is generated. The longitude and latitude of the vehicle head are calculated through the aligned electronic fence grid.

[0009] A vehicle head position estimation method based on vehicle-type template matching and motion trajectory extracts non-lane line ROIs to generate a mask overlaid on the original image for traffic object detection and tracking. A cascaded two-level network based on MobileNet-V2 is designed to segment the tracked vehicles into different types. A body information mapping table is established for each type to obtain the detected vehicle body length. Based on the body length and the midpoint of the lower edge of the detection frame, the heading angle and driving direction information are generated according to the vehicle tracking trajectory under the electronic fence grid of the bird's-eye view to calculate the latitude and longitude of the vehicle head.

[0010] Optionally, the region of interest positioning method based on the adaptive threshold SIFT and FLANN combined with image background information entropy is specifically as follows:

[0011] S101: Collect images from roadside cameras and convert them into grayscale;

[0012] S102: Determine whether the current frame is the first frame image. If it is the first frame image, manually select the lane line area on both sides of the road; if it is not the first frame image, extract the lane line ROI at the previous moment;

[0013] S103: Construct a Gaussian scale space pyramid for the intercepted lane line area by convolving the Gaussian function G(x, y, σ) with the image I(x, y), that is:

[0014]

[0015] symbol Indicates the convolution of two functions, G(x, y, σ) is a two-dimensional Gaussian function, σ represents the scale parameter, and the two-dimensional Gaussian function is expressed as follows:

[0016]

[0017] S104: Construct the DOG operator by subtracting images in two adjacent scale spaces to obtain an approximation of the Gaussian Lappass method LOG:

[0018] D(x,y,σ)=L(x,y,kσ)-L(x,y,σ)

[0019] S105: After obtaining the DOG space, each pixel is compared with the 26 points around it. If the point is an extreme point, it is defined as a candidate feature point;

[0020] S106: Determine whether the current frame of the image acquired after grayscale conversion is the first frame image. If it is the first frame image, manually select the non-lane ROI selection area; if it is not the first frame image, extract the non-lane ROI selection area at the previous moment;

[0021] S107: Calculate the grayscale histogram and probability distribution of the non-lane ROI at the current moment based on the non-lane ROI region selection at the previous moment:

[0022]

[0023] Where H and W represent the height and width of the non-lane ROI, respectively, I(x, y) represents the pixel value of the image at position (x, y), and δ(I=i) represents the indicator function, which is defined as follows:

[0024]

[0025] The grayscale distribution function p is expressed as:

[0026]

[0027] Among them, p i Indicates the probability of the pixel with gray value i appearing, n i Represents the number of pixels with grayscale value i, N represents the total number of pixels in the image; normalize the grayscale level i:

[0028]

[0029] Where L represents the number of gray levels;

[0030] S108: Calculate the one-dimensional information entropy H of the non-ROI area of the image at the current moment:

[0031]

[0032] S109: Design an adaptive contrast threshold function to eliminate the candidate feature points in S105:

[0033] thresh=thresh0*e αH

[0034] Where thresh0 is the initial threshold, and α is the adjustment parameter. The larger α is, the greater the influence of information entropy on the threshold. Roadside images under different lighting conditions, such as sunny days, cloudy days, rainy days, and nighttime, are taken. The grayscale entropy and the density of feature points under different thresholds of the SIFT algorithm are calculated to determine the optimal adjustment parameter of α.

[0035] S110: Calculate each candidate feature point Compare with thresh in S109, if it is less than the threshold, remove the candidate feature point, otherwise retain it;

[0036]

[0037] S111: Estimate the principal curvature of each candidate feature point using the curvature screening method in the SIFT algorithm. If the principal curvature is lower than the threshold, the feature point is retained; otherwise, it is eliminated.

[0038] S112: Use the gradient histogram weights in the feature point area to assign to each key point, so that the operator has rotation invariance, and calculate the amplitude and angle of each feature point

[0039]

[0040] θ(x,y)=αtan2((L(x,y+1)-L(x,y-1)) / (L(x+1,y)-I(x-1,y)))

[0041] Among them, L is the Gaussian smoothed image closest to the scale of the feature point;

[0042] S113: Generate feature point descriptors based on the gradient magnitude and direction within a 16×16 neighborhood window of each feature point. Each feature point is composed of 16 seed points, each with 8 vectors, forming a 128-dimensional SIFT feature vector.

[0043] S114: Generate a SIFT description operator for the entire image based on the adaptive threshold SIFT algorithm;

[0044] S115: Using the SIFT descriptor operator of the lane lines and the entire neighborhood image, perform feature point matching using the fast nearest neighbor algorithm FLANN to locate the positions of the lane lines ROI on both sides of the road in the image;

[0045] S116: Determine the road region of interest according to the position of the lane line ROI in the image.

[0046] Optionally, the latitude and longitude registration method after the image perspective transformation is specifically as follows:

[0047] S201: cropping the road ROI according to the located position of the road ROI in the image;

[0048] S202: Select four pixels in the road area of interest, and select four corresponding pixels at corresponding positions in the aerial image or the Google Earth bird's-eye view image;

[0049] S203: Perform perspective transformation. Assume that the coordinates of the four pixel points in the road area of interest are P1(x1, y1), P2(x2, y2), P3(x3, y3), and P4(x4, y4). Assume that the coordinates of the four corresponding pixel points in the corresponding positions in the bird's-eye view image are Q1(u1, v1), Q2(u2, v2), Q3(u3, v3), and Q4(u4, v4).

[0050]

[0051] Among them, [a ij ] represents the perspective transformation matrix, by constructing matrix A and vector b:

[0052]

[0053] Solve the linear equation system A[a 11 a 21 a 31 a 12 a 22 a 32 a 13 a 23 ] T =b, get the value of the perspective transformation matrix;

[0054] S204: The bird's-eye view of the road area of interest is rotated and adjusted to make the road vertical; let the image before rotation be I(x, y) and the image after rotation be I'(x', y'), then

[0055] x′=(xw / 2)cosθ-(yh / 2)sinθ+w / 2

[0056] y′=(xw / 2)sinθ+(yh / 2)cosθ+h / 2

[0057] Where w and h represent the width and height of the image respectively, and θ represents the rotation angle;

[0058] S205: Performing electronic fence grid division in the bird's-eye view of the road area of interest, dividing the lanes with a horizontal and vertical spacing of 50 pixels;

[0059] S206: synthesizing a video stream for the gridded road area of interest, and collecting the longitude and latitude of each grid point;

[0060] S207: Calculate the longitude and latitude coordinates of the test point. Assume that the pixel coordinates of any test point P in the bird's-eye view of the road area of interest are I(x, y), the grid spacing is spacing, and the grid row and column col of the test point are row and col, respectively.

[0061]

[0062]

[0063] S208: Take the grid point P0 (row + 1, col) as the origin O, P0P1 (row + 1, col + 1) as the X axis, and P0P2 (row, col) as the Y axis to establish a grid coordinate system. The coordinates of point P1 are (1, 0), the coordinates of point P2 are (0, 1), and the grid coordinates of the point P to be measured are P (x g ,y g ) is expressed as

[0064] x g =x mod spacing

[0065] y g =1-(y mod spacing)

[0066] Among them, mod is the remainder operation;

[0067] S209: The grid coordinate system is transformed into the equal longitude and latitude projection coordinate system O′X′Y′ by rotation, where O′X′ is the equal latitude and O′Y′ is the equal longitude. The rotation angle is θ, which is obtained by calibration with Google Earth. The coordinates of P2 in the coordinate system O′X′Y′ are (sinθ, cosθ), and the coordinates of the point to be measured P in O′X′Y′ are (x g ′,y g ')for;

[0068]

[0069] S210: Linearize the equal longitude and latitude lines in a grid. According to S206, let the longitude and latitude coordinates of point P2 be (lon2, lat2), and the longitude and latitude coordinates of point O be (lon0, lat0). Then the longitude and latitude coordinates (lon, lat) of the point P to be measured are:

[0070]

[0071]

[0072] Wherein, θ≠kπ / 2, k=0, 1, 2, ...

[0073] Optionally, the vehicle head position estimation method based on vehicle template matching and motion trajectory is specifically as follows:

[0074] S301: Add a non-lane line area mask to the image collected from the roadside;

[0075] S302: Traffic object detection and tracking based on the YOLO+deepsort series network;

[0076] S303: Determine whether the traffic object is a vehicle. If not, jump to S307.

[0077] S304: If the traffic object is a vehicle, determine the direction of the vehicle's movement based on the vehicle's trajectory between the two frames tracked by deepsort. If it is determined to be oncoming, jump to S307;

[0078] S305: If a vehicle is determined to be a destination, the vehicle in the image is captured according to the detection box. A two-level cascaded MobileNet-V2 network is designed and trained to segment the vehicles. A lightweight MobileNet-V2 network is designed and trained based on a public vehicle dataset to classify the identified vehicles into four categories: small cars, compact cars, mid-sized cars, and large cars.

[0079] S306: Establishing a body length matching relationship for the subdivided vehicle types, such as a small car body length of 4m, a compact car body length of 4.8m, a mid-size car body length of 5.5m, and a large car body length of 8m. After body length matching, assign body length information to the vehicles in the second-level classification.

[0080] S307: Calculate the longitude and latitude of the traffic object. If the detected traffic object is a pedestrian, the midpoint of the lower edge of the pedestrian detection frame is mapped to the position of the pedestrian in the bird's-eye view. The longitude and latitude of the pedestrian's location are obtained by using the latitude and longitude registration method after image perspective transformation. If the detected traffic object is a vehicle and its traveling direction is the incoming direction, the midpoint of the lower edge of the vehicle detection frame is mapped to the position of the vehicle's front end in the bird's-eye view. Similarly, the longitude and latitude of the vehicle's front end are obtained. If the detected traffic object is a vehicle and its traveling direction is the outgoing direction, the midpoint of the lower edge of the vehicle detection frame is mapped to the position of the vehicle's rear end in the bird's-eye view. Similarly, the longitude and latitude of the vehicle's front end are obtained.

[0081] S308: For a vehicle whose driving direction is detected as the direction to go, the vehicle heading angle information is obtained based on the trajectory of the vehicle's rear position in the bird's-eye view. The longitude and latitude of the vehicle's rear are set as P end (lon1, lat1), the vehicle body length is m, the heading angle is α, the rotation angle of the grid coordinate system OXY and the longitude and latitude projection coordinate system O′X′Y′ is θ, then the heading angle of the vehicle in the O′X′Y′ system is γ=α+θ; then the longitude and latitude of the vehicle head Pstart (lon2, lat2) is solved as follows:

[0082]

[0083]

[0084] Among them, Arc≈6371.393*1000 is the average distance from the center of the earth to all points on the earth's surface, in meters.

[0085] The beneficial effects of the present invention include: a region of interest location method combining adaptive threshold SIFT and FLANN based on image entropy, a latitude and longitude registration method after image perspective transformation, and a vehicle head position estimation method based on vehicle-type template matching and motion trajectory. The roadside image is grayscaled and then subjected to adaptive threshold SIFT and FLANN methods to quickly locate the positions of lane lines on both sides of the road in the image. By performing electronic fence grid division and longitude and latitude registration on the image road ROI after perspective transformation, the longitude and latitude coordinates corresponding to any pixel within the lane line can be calculated. Simultaneously, vehicle traffic objects are subdivided into categories, class template matching is performed, and motion trajectory is used to generate direction angles. The longitude and latitude of the vehicle head can be calculated using the registered electronic fence grid.

[0086] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0088] Figure 1 It is a structural block diagram of the method of the present invention;

[0089] Figure 2 This is a flow chart of the method for locating the road region of interest by combining SIFT and FLANN with an adaptive threshold of image background information entropy according to the present invention;

[0090] Figure 3 A location map of the region of interest for locating roads in the present invention;

[0091] Figure 4 This is a flow chart of the latitude and longitude registration method based on image perspective transformation according to the present invention;

[0092] Figure 5 A diagram showing the grid division of an electronic fence under a bird's-eye view of an area of interest on a road according to the present invention;

[0093] Figure 6 This is a diagram of the grid coordinate system converted to the equal longitude and latitude projection coordinate system of the present invention;

[0094] Figure 7 This is a flow chart of the vehicle head position estimation method based on vehicle template matching and motion trajectory of the present invention;

[0095] Figure 8 Schematic diagram of the calculation of the vehicle head position estimation method of the present invention. DETAILED DESCRIPTION

[0096] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0097] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0098] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0099] The present invention proposes a method for locating traffic objects based on roadside camera and latitude and longitude registration, comprising:

[0100] The method for locating the region of interest based on the adaptive threshold SIFT and FLANN combined with the image background information entropy, the latitude and longitude registration method after image perspective transformation, and the vehicle head position estimation method based on vehicle template matching and motion trajectory are shown in the following table. Figure 1 As shown in the figure, the region of interest positioning method based on the combination of adaptive threshold SIFT and FLANN based on image background information entropy can quickly locate the position of lane lines on both sides of the road in the image after grayscale conversion of the image collected from the road side, adaptive threshold SIFT extraction of feature points, and FLANN feature point matching to determine the lane area of interest, and solve the problem of lane area of interest offset caused by shaking of roadside camera mounting poles due to typhoons and rainstorms; the latitude and longitude registration method based on image perspective transformation transforms the image lane line ROI into a bird's-eye view, performs electronic fence grid division and latitude and longitude registration, and can calculate the latitude and longitude coordinates corresponding to any pixel point in the lane line; for vehicle traffic Objects are classified, class templates are matched, and direction angles are generated based on motion trajectories. The longitude and latitude of the vehicle head can be calculated through the registered electronic fence grid. The vehicle head position estimation method based on vehicle class template matching and motion trajectory is used to extract non-lane line ROIs to generate masks and overlay them on the original image for traffic object detection and tracking. A cascaded two-level network based on MobileNet-V2 is designed to segment the tracked vehicles. A body information mapping table is established for each vehicle model to obtain the detected vehicle body length. Based on the body length and the midpoint of the lower edge of the detection frame, the heading angle and driving direction information are generated according to the vehicle tracking trajectory under the electronic fence grid in a bird's-eye view to solve the longitude and latitude of the vehicle head.

[0101] In this embodiment, the region of interest positioning method based on the adaptive threshold SIFT and FLANN combined with the image background information entropy is shown in Figure 2 As shown, the process is as follows:

[0102] S101: Collect images from roadside cameras and convert them into grayscale;

[0103] S102: Determine whether the current frame is the first frame image. If it is the first frame image, manually select the lane line area on both sides of the road; if it is not the first frame image, extract the lane line ROI at the previous moment;

[0104] S103: Construct a Gaussian scale space pyramid for the intercepted lane line area by convolving the Gaussian function G(x, y, σ) with the image I(x, y), that is:

[0105]

[0106] symbol Indicates the convolution of two functions, G(x, y, σ) is a two-dimensional Gaussian function, σ represents the scale parameter, and the two-dimensional Gaussian function is expressed as follows:

[0107]

[0108] S104: Construct the DOG operator by subtracting images in two adjacent scale spaces to obtain an approximation of the Gaussian Lappass method LOG:

[0109] D(x,y,σ)=L(x,y,kσ)-L(x,y,σ)

[0110] S105: After obtaining the DOG space, each pixel is compared with the 26 points around it. If the point is an extreme point, it is defined as a candidate feature point;

[0111] S106: Determine whether the current frame of the image acquired after grayscale conversion is the first frame image. If it is the first frame image, manually select the non-lane ROI selection area; if it is not the first frame image, extract the non-lane ROI selection area at the previous moment;

[0112] S107: Calculate the grayscale histogram and probability distribution of the non-lane ROI at the current moment based on the non-lane ROI region selection at the previous moment:

[0113]

[0114] Where H and W represent the height and width of the non-lane ROI, respectively, I(x, y) represents the pixel value of the image at position (x, y), and δ(I = i) represents the indicator function, which is defined as follows

[0115]

[0116] The grayscale distribution function p is expressed as

[0117]

[0118] Among them, p i Indicates the probability of the pixel with gray value i appearing, n i Represents the number of pixels with grayscale value i, and N represents the total number of pixels in the image. Normalize the grayscale level i

[0119]

[0120] Where L represents the number of gray levels;

[0121] S108: Calculate the one-dimensional information entropy H of the non-ROI area of the image at the current moment:

[0122]

[0123] S109: Design an adaptive contrast threshold function to eliminate candidate feature points in S105

[0124] thresh=thresh0*e αH

[0125] Where thresh0 is the initial threshold, and α is the adjustment parameter. The larger α is, the greater the influence of information entropy on the threshold. Roadside images under different lighting conditions, such as sunny days, cloudy days, rainy days, and nighttime, are taken. The grayscale entropy and the density of feature points under different thresholds of the SIFT algorithm are calculated to determine the optimal adjustment parameter of α.

[0126] S110: Calculate each candidate feature point Compare with thresh in S109, if it is less than the threshold, remove the candidate feature point, otherwise retain it;

[0127]

[0128] S111: Estimate the principal curvature of each candidate feature point using the curvature screening method in the SIFT algorithm. If the principal curvature is lower than the threshold, the feature point is retained; otherwise, it is eliminated.

[0129] S112: Use the gradient histogram weights in the feature point area to assign to each key point, so that the operator has rotation invariance, and calculate the amplitude and angle of each feature point

[0130]

[0131] θ(x,y)=αtan2((L(x,y+1)-L(x,y-1)) / (L(x+1,y)-I(x-1,y)))

[0132] Among them, L is the Gaussian smoothed image closest to the scale of the feature point;

[0133] S113: Generate feature point descriptors based on the gradient magnitude and direction within a 16×16 neighborhood window of each feature point. Each feature point is composed of 16 seed points, each with 8 vectors, forming a 128-dimensional SIFT feature vector.

[0134] S114: The above steps S101-S113 describe the process of extracting lane line neighborhood features based on the adaptive threshold SIFT algorithm. The SIFT descriptor of the entire image is generated based on the same method.

[0135] S115: Using the SIFT descriptor operator of the lane lines and the entire neighborhood image, perform feature point matching using the FLANN (Fast Nearest Neighbor) algorithm to locate the positions of the lane lines ROI on both sides of the road in the image;

[0136] S116: Determine the road region of interest based on the position of the lane line ROI in the image, see Figure 3 shown.

[0137] In summary, the region of interest localization method combining adaptive threshold SIFT and FLANN based on image background information entropy can locate the position of the road region of interest in the camera before and after the pole shakes, and can adapt to roadside environments under different lighting environments such as sunny days, cloudy days, rainy days, and nights, and has strong robustness.

[0138] In this embodiment, the latitude and longitude registration method based on the image perspective transformation is shown in Figure 4 As shown, the process:

[0139] S201: cropping the road ROI according to the position of the road ROI located in the image by the above method;

[0140] S202: Select four pixels from the road area of interest. Select four corresponding pixels from the corresponding positions in the aerial image or the Google Earth bird's-eye view image. The selection principles are: (1) the pixel features are as obvious as possible, so that their positions can be well determined in both the area of interest and the bird's-eye view image; (2) the area surrounded by the four pixels covers the entire road area of interest as much as possible.

[0141] S203: Perform perspective transformation. Assume that the coordinates of the four pixel points in the road area of interest are P1(x1, y1), P2(x2, y2), P3(x3, y3), and P4(x4, y4). Assume that the coordinates of the four corresponding pixel points in the corresponding positions in the bird's-eye view image are Q1(u1, v1), Q2(u2, v2), Q3(u3, v3), and Q4(u4, v4).

[0142]

[0143] Among them, [a ij ] represents the perspective transformation matrix, by constructing matrix A and vector b

[0144]

[0145] Solve the linear equation system A[a 11 a 21 a 31 a 12 a 22 a 32 a 13 a 23 ] T =b, the value of the perspective transformation matrix can be obtained;

[0146] S204: The bird's-eye view of the road area of interest is rotated and adjusted to make the road vertical; let the image before rotation be I(x, y) and the image after rotation be I'(x', y'), then

[0147] x′=(xw / 2)cosθ-(yh / 2)sinθ+w / 2

[0148] y′=(xw / 2)sinθ+(yh / 2)cosθ+h / 2

[0149] Where w and h represent the width and height of the image respectively, and θ represents the rotation angle;

[0150] S205: Perform geo-fence grid division under the bird's-eye view of the road area of interest, see Figure 5 As shown, in this embodiment, the lanes are divided with a spacing of 50 pixels horizontally and vertically;

[0151] S206: synthesizing a video stream for the gridded road area of interest, and collecting the longitude and latitude of each grid point;

[0152] S207: Calculate the longitude and latitude coordinates of the test point. Assume that the pixel coordinates of any test point P in the bird's-eye view of the road area of interest are I(x, y), the grid spacing is spacing, and the grid row and column col of the test point are row and col, respectively.

[0153]

[0154]

[0155] S208: Take the grid point P0 (row + 1, col) as the origin O, P0P1 (row + 1, col + 1) as the X axis, and P0P2 (row, col) as the Y axis to establish a grid coordinate system. The coordinates of point P1 are (1, 0), the coordinates of point P2 are (0, 1), and the grid coordinates of the point P to be measured are P (x g ,y g ) is expressed as

[0156] x g =x mod spacing

[0157] y g =1-(y mod spacing)

[0158] Among them, mod is the remainder operation;

[0159] S209: Transform the grid coordinate system into the iso-longitude and latitude projection coordinate system O′X′Y′ by rotation, where O′X′ is the iso-latitude and O′Y′ is the iso-longitude. Set the rotation angle to θ (which can be obtained by calibration on Google Earth). Figure 6 As shown, the coordinates of P2 in the coordinate system O′X′Y′ are (sinθ, cosθ), and the coordinates of the point to be measured P in O′X′Y′ are (x g ′,y g ')for;

[0160]

[0161] S210: Linearize the equal longitude and latitude lines in a grid. According to S206, let the longitude and latitude coordinates of point P2 be (lon2, lat2), and the longitude and latitude coordinates of point O be (lon0, lat0). Then the longitude and latitude coordinates (lon, lat) of the point P to be measured are:

[0162]

[0163]

[0164] Where θ≠kπ / 2, k=0, 1, 2...

[0165] In summary, the latitude and longitude registration method based on image perspective transformation can calculate the latitude and longitude coordinates of any point in the road area of interest, and realize the absolute position positioning of any point in the lane line under the roadside camera.

[0166] In this embodiment, the vehicle head position estimation method based on vehicle template matching and motion trajectory is shown in Figure 7 As shown, the process is as follows:

[0167] S301: Add a non-lane line area mask to the image collected from the roadside;

[0168] S302: Traffic object detection and tracking based on the YOLO+deepsort series network;

[0169] S303: Determine whether the traffic object is a vehicle. If not, jump to S307.

[0170] S304: If the traffic object is a vehicle, determine the vehicle's direction of movement based on the vehicle's trajectory between the two frames tracked by deepsort. If it is determined to be oncoming, jump to S307;

[0171] S305: If a vehicle is determined to be a destination, the vehicle in the image is captured according to the detection box. A two-level cascaded MobileNet-V2 network is designed and trained to segment the vehicles. A lightweight MobileNet-V2 network is designed and trained based on a public vehicle dataset to classify the identified vehicles into four categories: small cars, compact cars, mid-sized cars, and large cars.

[0172] S306: Establishing a body length matching relationship for the subdivided vehicle types, such as a small car body length of 4m, a compact car body length of 4.8m, a mid-size car body length of 5.5m, and a large car body length of 8m. After body length matching, assign body length information to the vehicles in the second-level classification.

[0173] S307: Calculate the longitude and latitude of the traffic object. If the detected traffic object is a pedestrian, the midpoint of the lower edge of the pedestrian detection frame is mapped to the position of the pedestrian in the bird's-eye view. The longitude and latitude of the pedestrian's location can be obtained by using the latitude and longitude registration method after the image perspective is transformed. If the detected traffic object is a vehicle and its traveling direction is the incoming direction, the midpoint of the lower edge of the vehicle detection frame is mapped to the position of the vehicle's front end in the bird's-eye view. Similarly, the longitude and latitude of the vehicle's front end can be obtained. If the detected traffic object is a vehicle and its traveling direction is the outgoing direction, the midpoint of the lower edge of the vehicle detection frame is mapped to the position of the vehicle's rear end in the bird's-eye view. Similarly, the longitude and latitude of the vehicle's front end can be obtained.

[0174] S308: For a vehicle whose driving direction is detected as the direction to go, the vehicle heading angle information can be obtained based on the trajectory of the vehicle's rear position in the bird's-eye view. Let the longitude and latitude of the vehicle's rear be P end (lon1, lat1), the vehicle body length is m, the heading angle is α, and the rotation angle of the grid coordinate system OXY and the longitude and latitude projection coordinate system O′X′Y′ is θ, see Figure 8 As shown, the heading angle of the vehicle in the O′X′Y′ system is γ=α+θ; then the latitude and longitude of the vehicle head is P start (lon2, lat2) is solved as follows

[0175]

[0176]

[0177] Among them, Arc≈6371.393*1000 is the average distance from the center of the earth to all points on the earth's surface, in meters.

[0178] In summary, the vehicle head position estimation method based on vehicle template matching and motion trajectory can realize the absolute position positioning of traffic objects in the lane and the position estimation of the vehicle head.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for locating traffic objects based on roadside camera and latitude and longitude registration, characterized by: The method includes a region of interest localization method based on an adaptive threshold SIFT and FLANN method based on image background information entropy, a latitude and longitude registration method after image perspective transformation, and a vehicle head position estimation method based on vehicle-type template matching and motion trajectory. The region of interest (ROI) localization method, which combines adaptive threshold SIFT and FLANN based on image background information entropy, grayscales roadside images, extracts feature points using adaptive threshold SIFT, and then uses FLANN feature point matching to locate the lane markings on both sides of the road in the image. This method helps address the issue of lane ROI offset caused by shaking of the roadside camera mounting poles due to typhoons and heavy rains. Based on the latitude and longitude registration method after image perspective transformation, the lane line ROI in the image is transformed into a bird's-eye view. Then, the electronic fence grid is divided and the longitude and latitude coordinates are aligned. The longitude and latitude coordinates corresponding to any pixel point in the lane line are calculated. The vehicle traffic objects are classified, the class template is matched, and the direction angle of the motion trajectory is generated. The longitude and latitude of the vehicle head are calculated through the aligned electronic fence grid. A vehicle head position estimation method based on vehicle-type template matching and motion trajectory extracts non-lane line ROIs to generate a mask overlaid on the original image for traffic object detection and tracking. A cascaded two-level network based on MobileNet-V2 is designed to segment the tracked vehicles into different types. A body information mapping table is established for each type to obtain the detected vehicle body length. Based on the body length and the midpoint of the lower edge of the detection frame, the heading angle and driving direction information are generated according to the vehicle tracking trajectory under the electronic fence grid of the bird's-eye view to calculate the latitude and longitude of the vehicle head.

2. The method for locating traffic objects based on roadside camera and latitude and longitude registration according to claim 1, characterized in that: The method for locating the region of interest based on the adaptive threshold SIFT and FLANN combined with image background information entropy is specifically as follows: S101: Collect images from roadside cameras and convert them into grayscale; S102: Determine whether the current frame is the first frame image. If it is the first frame image, manually select the lane line area on both sides of the road; if it is not the first frame image, extract the lane line ROI at the previous moment; S103: Construct a Gaussian scale space pyramid for the intercepted lane line area by convolving the Gaussian function G(x, y, σ) with the image I(x, y), that is: symbol Indicates the convolution of two functions, G(x, y, σ) is a two-dimensional Gaussian function, σ represents the scale parameter, and the two-dimensional Gaussian function is expressed as follows: S104: Construct the DOG operator by subtracting images in two adjacent scale spaces to obtain an approximation of the Gaussian Lappass method LOG: D(x,y,σ)=L(x,y,kσ)-L(x,y,σ) S105: After obtaining the DOG space, each pixel is compared with the 26 points around it. If the point is an extreme point, it is defined as a candidate feature point; S106: Determine whether the current frame of the image acquired after grayscale conversion is the first frame image. If it is the first frame image, manually select the non-lane ROI selection area; if it is not the first frame image, extract the non-lane ROI selection area at the previous moment; S107: Calculate the grayscale histogram and probability distribution of the non-lane ROI at the current moment based on the non-lane ROI region selection at the previous moment: Where H and W represent the height and width of the non-lane ROI, respectively, I(x, y) represents the pixel value of the image at position (x, y), and δ(I=i) represents the indicator function, which is defined as follows: The grayscale distribution function p is expressed as: Among them, p i Indicates the probability of the pixel with gray value i appearing, n i Represents the number of pixels with grayscale value i, N represents the total number of pixels in the image; normalize the grayscale level i: Where L represents the number of gray levels; S108: Calculate the one-dimensional information entropy H of the non-ROI area of the image at the current moment: S109: Design an adaptive contrast threshold function to eliminate the candidate feature points in S105: thresh=thresh0*e αH Where thresh0 is the initial threshold, and α is the adjustment parameter. The larger α is, the greater the influence of information entropy on the threshold. Roadside images under different lighting conditions, such as sunny days, cloudy days, rainy days, and nighttime, are taken. The grayscale entropy and the density of feature points of the SIFT algorithm under different thresholds are calculated to determine the optimal adjustment parameter of α. S110: Calculate each candidate feature point Compare with thresh in S109, if it is less than the threshold, remove the candidate feature point, otherwise retain it; S111: Estimate the principal curvature of each candidate feature point using the curvature screening method in the SIFT algorithm. If the principal curvature is lower than the threshold, the feature point is retained; otherwise, it is eliminated. S112: Use the gradient histogram weights in the feature point area to assign to each key point, so that the operator has rotation invariance, and calculate the amplitude and angle of each feature point θ(x,y)=αtan2((L(x,y+1)-L(x,y-1)) / (L(x+1,y)-I(x-1,y))) Among them, L is the Gaussian smoothed image closest to the scale of the feature point; S113: Generate feature point descriptors based on the gradient magnitude and direction within a 16×16 neighborhood window of each feature point. Each feature point is composed of 16 seed points, each with 8 vectors, forming a 128-dimensional SIFT feature vector. S114: Generate a SIFT description operator for the entire image based on the adaptive threshold SIFT algorithm; S115: Using the SIFT descriptor operator of the lane lines and the entire neighborhood image, perform feature point matching using the fast nearest neighbor algorithm FLANN to locate the positions of the lane lines ROI on both sides of the road in the image; S116: Determine the road region of interest according to the position of the lane line ROI in the image.

3. The method for locating traffic objects based on roadside camera and latitude and longitude registration according to claim 2, characterized in that: The latitude and longitude registration method after the image perspective transformation is specifically as follows: S201: cropping the road ROI according to the located position of the road ROI in the image; S202: Select four pixels in the road area of interest, and select four corresponding pixels at corresponding positions in the aerial image or the Google Earth bird's-eye view image; S203: Perform perspective transformation. Assume that the coordinates of the four pixel points in the road area of interest are P1(x1, y1), P2(x2, y2), P3(x3, y3), and P4(x4, y4). Assume that the coordinates of the four corresponding pixel points in the corresponding positions in the bird's-eye view image are Q1(u1, v1), Q2(u2, v2), Q3(u3, v3), and Q4(u4, v4). Among them, [a ij ] represents the perspective transformation matrix, by constructing matrix A and vector b: Solve the linear equation system A[a 11 a 21 a 31 a 12 a 22 a 32 a 13 a 23 ] T =b, get the value of the perspective transformation matrix; S204: The bird's-eye view of the road area of interest is rotated and adjusted to make the road vertical; let the image before rotation be I(x, y) and the image after rotation be I'(x', y'), then x′=(xw / 2)cosθ-(yh / 2)sinθ+w / 2 y′=(xw / 2)sinθ+(yh / 2)COsθ+h / 2 Where w and h represent the width and height of the image respectively, and θ represents the rotation angle; S205: Performing electronic fence grid division in the bird's-eye view of the road area of interest, dividing the lanes with a horizontal and vertical spacing of 50 pixels; S206: synthesizing a video stream for the gridded road area of interest, and collecting the longitude and latitude of each grid point; S207: Calculate the longitude and latitude coordinates of the test point. Assume that the pixel coordinates of any test point P in the bird's-eye view of the road area of interest are I(x, y), the grid spacing is spacing, and the grid row and column col of the test point are row and col, respectively. S208: Take the grid point P0 (row + 1, col) as the origin O, P0P1 (row + 1, col + 1) as the X axis, and P0P2 (row, col) as the Y axis to establish a grid coordinate system. The coordinates of point P1 are (1, 0), the coordinates of point P2 are (0, 1), and the grid coordinates of the point P to be measured are P (x g ,y g ) is expressed as x g =x mod spacing y g =1-(y mod spacing) Among them, mod is the remainder operation; S209: The grid coordinate system is transformed into the equal longitude and latitude projection coordinate system O′X′Y′ by rotation, where O′X′ is the equal latitude and O′Y′ is the equal longitude. The rotation angle is θ, which is obtained by calibration with Google Earth. The coordinates of P2 in the coordinate system O′X′Y′ are (sinθ, cosθ), and the coordinates of the point to be measured P in O′X′Y′ are (x g ′,y g ')for; S210: Linearize the equal longitude and latitude lines in a grid. According to S206, let the longitude and latitude coordinates of point P2 be (lon2, lat2), and the longitude and latitude coordinates of point O be (lon0, lat0). Then the longitude and latitude coordinates (lon, lat) of the point P to be measured are: Wherein, θ≠kπ / 2, k=0, 1, 2… 4. The method for locating traffic objects based on roadside camera and latitude and longitude registration according to claim 3, characterized in that: The vehicle head position estimation method based on vehicle template matching and motion trajectory is specifically as follows: S301: Add a non-lane line area mask to the image collected from the roadside; S302: Traffic object detection and tracking based on the YOLO+deepsort series network; S303: Determine whether the traffic object is a vehicle. If not, jump to S307. S304: If the traffic object is a vehicle, determine the vehicle's direction of movement based on the vehicle's trajectory between the two frames tracked by deepsort. If it is determined to be oncoming, jump to S307; S305: If a vehicle is determined to be a destination, the vehicle in the image is captured according to the detection box. A two-level cascaded MobileNet-V2 network is designed and trained to segment the vehicles. A lightweight MobileNet-V2 network is designed and trained based on a public vehicle dataset to classify the identified vehicles into four categories: small cars, compact cars, mid-sized cars, and large cars. S306: Establishing a body length matching relationship for the subdivided vehicle types, such as a small car body length of 4m, a compact car body length of 4.8m, a mid-size car body length of 5.5m, and a large car body length of 8m. After body length matching, assign body length information to the vehicles in the second-level classification. S307: Calculate the longitude and latitude of the traffic object. If the detected traffic object is a pedestrian, the midpoint of the lower edge of the pedestrian detection frame is mapped to the position of the pedestrian in the bird's-eye view. The longitude and latitude of the pedestrian's location are obtained by using the latitude and longitude registration method after image perspective transformation. If the detected traffic object is a vehicle and its traveling direction is the incoming direction, the midpoint of the lower edge of the vehicle detection frame is mapped to the position of the vehicle's front end in the bird's-eye view. Similarly, the longitude and latitude of the vehicle's front end are obtained. If the detected traffic object is a vehicle and its traveling direction is the outgoing direction, the midpoint of the lower edge of the vehicle detection frame is mapped to the position of the vehicle's rear end in the bird's-eye view. Similarly, the longitude and latitude of the vehicle's front end are obtained. S308: For a vehicle whose driving direction is detected as the direction to go, the vehicle heading angle information is obtained based on the trajectory of the vehicle's rear position in the bird's-eye view. The longitude and latitude of the vehicle's rear are set as P end (lon1, lat1), the vehicle body length is m, the heading angle is α, the rotation angle of the grid coordinate system OXY and the longitude and latitude projection coordinate system O′X′Y′ is θ, then the heading angle of the vehicle in the O′X′Y′ system is γ=α+θ; then the longitude and latitude of the vehicle head P start (lon2, lat2) is solved as follows: Among them, Arc≈6371.393*1000 is the average distance from the center of the earth to all points on the earth's surface, in meters.

Citation Information

Patent Citations

  • Image registration method combining target detection and semantic segmentation

    CN110097584A

  • Vehicle deviation alarm method based on lane line gradient image adaptive threshold segmentation

    CN110298216A