Vehicle real position resolving method for monocular camera monitoring scene

By processing the video stream of the monocular camera monitoring device, constructing the calibration matrix and feature point solution, the problem that the monocular camera cannot accurately solve the vehicle's three-dimensional coordinates is solved, and high-precision vehicle three-dimensional position solution and detection are achieved.

CN120823540APending Publication Date: 2025-10-21ANHUI KELI INFORMATION IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510869681.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing monocular camera monitoring equipment cannot accurately calculate the three-dimensional coordinates of a vehicle in the real world, resulting in low detection accuracy. Traditional methods also increase hardware costs and system complexity.

Method used

By acquiring the real-time video stream of the monocular surveillance camera, constructing the camera calibration matrix, performing vehicle target detection and tracking, extracting feature points, calculating the three-dimensional motion displacement vector, and solving the three-dimensional coordinates of the vehicle in the real world.

Benefits of technology

It achieves high-precision three-dimensional position solution of vehicles, reduces hardware costs, solves the scale distortion problem of traditional methods, and improves detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823540A_ABST
    Figure CN120823540A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle real position resolving method for a monocular camera monitoring scene, which relates to the technical field of traffic monitoring, and comprises the following steps: acquiring real-time video stream information of a monocular camera, and decoding to generate a continuous frame image sequence; constructing a camera calibration matrix; target detection and tracking are carried out on each frame of image, a unique identity ID is distributed, and first position information of a vehicle target is extracted; based on the first position information, cutting a vehicle target local image, and extracting all feature point sets in the local image as second position information; matching the second position information of the same vehicle ID in the adjacent frames, and calculating a three-dimensional motion displacement vector of the vehicle target in the real world coordinate system; and according to the three-dimensional motion displacement vector and the camera calibration matrix, calculating the three-dimensional coordinates of the vehicle feature points in the real world coordinate system, and fitting to generate the three-dimensional position and size information of the vehicle target. And the real three-dimensional position is calculated through the motion displacement of the continuous frame feature points of the vehicle in the video stream of the monocular camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic monitoring, and in particular to a method for calculating the real position of a vehicle in a monocular camera monitoring scene. Background Art

[0002] In the field of traffic monitoring, accurate detection of traffic parameters such as vehicle spacing and real-time vehicle speed, as well as traffic events such as crossing the line and changing lanes, requires obtaining parameters related to physical properties such as the distance and size of the target in the real world, in order to accurately calculate the real-world three-dimensional positional relationship between vehicles and lanes, and between vehicles. Existing surveillance camera equipment can detect motor vehicles in terms of target detection and recognition, but it can only provide the two-dimensional coordinate values ​​of the pixel area formed by the projection of the vehicle in the two-dimensional camera image, and cannot directly provide the three-dimensional coordinate values ​​of the vehicle in the real-world coordinate system. When some manufacturers associate the detection target mapping with the real scene, they directly equate the two-dimensional pixel area in the camera image with its position in the real-world three-dimensional coordinate system. However, if the two-dimensional pixel coordinate values ​​are directly used to restore the real-world three-dimensional coordinate values ​​corresponding to the object through the camera intrinsic and extrinsic parameters obtained by camera calibration, the calculated three-dimensional coordinate position of the vehicle will be much larger than its real position in the world coordinate system, resulting in low accuracy and precision. For example, in a real electric police monitoring scenario, when the vehicle is in front of the stop line, a large part of the pixel area occupied by the vehicle body in the camera image often exceeds the stop line. The integrated radar and vision device transmits and receives millimeter waves or lasers, measuring the vehicle's true position based on their reflection from the vehicle's main body. While this method can accurately detect the vehicle's true position, it suffers from high hardware costs, difficulty in implementation, and a short service life. In addition, some fields use two sets of cameras to build a binocular camera system, or use 3D structured light cameras, etc. to detect corners based on the optical feature points of the image, and then perform 3D reconstruction based on the corners to calculate the true 3D position of the target object. However, such solutions using dedicated cameras not only increase hardware costs but also significantly increase system complexity, leading to problems such as low robustness, poor real-time performance, high scene dependence, and high construction costs. It can be seen that there is an urgent need for a method for calculating the real position of a vehicle in a monocular camera monitoring scene to solve the problems in the existing technology of using only traditional monocular monitoring camera equipment, which cannot construct the three-dimensional scale of objects, cannot calculate the actual position size of the vehicle, or has low data accuracy. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides a method for calculating the real position of a vehicle in a scene monitored by a monocular camera, comprising the following steps: S1. Obtain the real-time video stream information of the monocular surveillance camera and decode it to generate a continuous frame image sequence; S2. Construct the camera calibration matrix of the monocular surveillance camera, including the intrinsic parameter matrix , extrinsic rotation matrix and translation matrix ; S3, detecting and tracking a vehicle target for each frame of image, assigning a unique ID to each vehicle target, and extracting the first position information of the vehicle target in the image pixel coordinate system; S4. Based on the first position information, crop the vehicle target partial image, and extract all feature point sets in the partial image as second position information; S5. Match the second position information of the same vehicle ID in adjacent frames and calculate the three-dimensional motion displacement vector of the vehicle target in the real-world coordinate system; S6. Calculate the three-dimensional coordinates of the vehicle feature points in the real-world coordinate system based on the three-dimensional motion displacement vector and the camera calibration matrix, and generate the three-dimensional position and size information of the vehicle target by fitting.

[0004] Furthermore, the calculation method of the three-dimensional motion displacement vector in S5 includes: Selecting a matching feature point with the largest sum of vertical coordinate pixel values ​​in the second position information as a reference feature point; Substitute the reference feature points into the camera imaging model and set the initial depth coordinates , calculating the initial three-dimensional coordinates of the reference feature point in the real-world coordinate system; Generate a three-dimensional motion displacement vector based on the difference in the initial three-dimensional coordinates of the reference feature points of the vehicle target in adjacent frames .

[0005] Furthermore, the S6 includes: The three-dimensional motion displacement vector Substituting the camera imaging model into the real world coordinate system, and calculating the three-dimensional coordinates of the non-reference feature points; The three-dimensional coordinates of all feature points are integrated, and the three-dimensional outline of the vehicle target is restored through spatial fitting, and the position information, size information and motion state parameters of the vehicle target in the real-world coordinate system are output.

[0006] Furthermore, the three-dimensional motion displacement vector Substituting into the camera imaging model, the following relationship is satisfied:

[0007] in, 、 They are the three-dimensional motion displacement vectors of the vehicle target in the real world exist Axis and The displacement component in the axial direction.

[0008] Furthermore, the feature point extraction method in S4 includes: Converting the local image into a grayscale image; Traverse the pixel points of the grayscale image and determine the pixel points on the circle centered at the point. Whether the absolute value of the grayscale difference between the pixel point and the center point exceeds the threshold ; If the threshold is exceeded If the percentage of points is greater than the preset ratio, it is determined to be a feature point; The final feature point set is filtered by non-maximum suppression.

[0009] Furthermore, the first position information in S3 is expressed as a rectangular frame in a two-dimensional coordinate system, specifically: The coordinates of the diagonal vertices of the rectangular frame in the image pixel coordinate system; or, The center coordinates of the rectangular frame in the image pixel coordinate system and the width and height of the rectangular frame.

[0010] Optionally, the camera calibration method in S2 includes any one of a traditional camera calibration method, a camera self-calibration algorithm, or a Zhang Zhengyou calibration method.

[0011] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention calculates the true three-dimensional position by the motion displacement of the feature points of the vehicle in consecutive frames in the monocular camera video stream, reducing hardware costs. It also infers the depth information based on the consistency constraint of the vehicle motion, solving the scale distortion problem of traditional monocular solutions that directly map the two-dimensional detection frame into three-dimensional space. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0013] Figure 1 This is a flowchart of the overall process disclosed in the embodiment of the present invention; Figure 2 This is a schematic diagram of imaging of a monocular surveillance camera disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0014] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0015] The present invention aims to provide a method for calculating the true position of a vehicle in a monocular camera surveillance scenario. This method enables high-precision three-dimensional positioning by reusing existing surveillance cameras without the need for expensive equipment such as radar and binocular cameras. It also solves the scale distortion problem of traditional monocular solutions that directly map a two-dimensional detection frame into three-dimensional space.

[0016] See also Figure 1 The present invention provides a method for calculating the real position of a vehicle in a monocular camera monitoring scene, which mainly includes the following steps: S1. Obtain the real-time video stream information of the monocular surveillance camera and decode it to generate a continuous frame image sequence.

[0017] First, configure the network video stream information of the road monitoring camera.

[0018] The network video stream information of the road monocular surveillance camera includes at least the communication protocol and communication address of the network video stream, and may also include identity authentication information such as user name and password for access control of the monocular surveillance camera; it may also include encoding format information for local decoding application after the method obtains the real-time network video stream.

[0019] In this embodiment, the communication protocol for the real-time network video stream from the monocular road surveillance camera can use international and domestic standard protocols, including but not limited to RTP (Real-time Transport Protocol), RTCP (Real-time Transport Control Protocol), RTSP (Real-time Streaming Protocol), RTMP (Real-time Messaging Protocol), and GB28181. The communication address includes information such as the path name and port number for the network video stream.

[0020] According to the configured road monocular surveillance camera network video stream information, a long connection is established with the road monocular surveillance camera; then, frame-by-frame data of the real-time video stream is continuously obtained, frame-by-frame image information is decoded, and the frame-by-frame images are arranged and numbered in sequence.

[0021] S2. Construct the camera calibration matrix of the monocular surveillance camera, including the intrinsic parameter matrix , extrinsic rotation matrix and translation matrix .

[0022] After obtaining the RGB image of the camera monitoring screen, it is necessary to calibrate the monocular monitoring camera to obtain the intrinsic parameter matrix K, extrinsic parameter rotation matrix R and translation matrix T of the monocular monitoring camera.

[0023] In image measurement and machine vision applications, a geometric model of camera imaging must be established to determine the relationship between the three-dimensional geometric position of a point on a spatial object's surface and its corresponding point in the image. These geometric model parameters are known as camera parameters. Under most conditions, these parameters (intrinsic parameters, extrinsic parameters, and distortion parameters) must be determined through experimentation and calculation. This process of determining these parameters is called camera calibration. Camera parameter calibration is a critical step, as the accuracy of the calibration results and the stability of the algorithm directly impact the accuracy of the calculated real-world coordinate position information for any point in the camera's surveillance image.

[0024] Optionally, the camera calibration method includes any one of a traditional camera calibration method, a camera self-calibration algorithm, or a Zhang Zhengyou calibration method.

[0025] Alternatively, select an object with known dimensions in the real scene as a calibration reference, such as a lane area or building, with parallel or orthogonal information. If no suitable calibration reference exists in the real scene, you can also prepare a calibration reference and temporarily place it on the road in the monitoring scene for camera calibration.

[0026] By establishing the correspondence between the points on the calibration object with known real-world coordinates and their pixel coordinates in the image, a certain algorithm is used to calculate the internal and external parameters of the camera imaging model:

[0027]

[0028]

[0029]

[0030] in,( ) is a point in space The coordinates in the world coordinate system. for Corresponding points in the pixel coordinate system of the surveillance camera imaging screen 's coordinates.

[0031] is the rotation matrix, The translation vector is the external parameter of the camera. It is the transformation relationship between the coordinate axis rotation angle and the origin offset vector established based on factors such as the installation point and angle when converting from the world coordinate system to the camera coordinate system. axis, axis, The axis rotates by an angle of 、 、 .

[0032] is the scaling factor ( is not 0), is the effective focal length (the distance from the camera optical center to the image plane), is the homogeneous coordinate of the spatial point in the camera coordinate system, and is the homogeneous coordinate of the imaging point in the image coordinate system.

[0033] It is the camera's intrinsic parameter, which is the inherent parameter determined when the camera device is produced. , For Camera axis, The normalized focal length on the axis, in pixels. , and Represents a pixel on the image plane Axis and The physical size of the axis direction is calculated by the ratio of the sensor size and the camera resolution. In general, the ratio of the height and width transformation of the object when the camera is imaging is different, and this ratio is obtained based on the focal length. Therefore, it is necessary to calculate the camera separately. Direction and The focal length of the direction.

[0034] Through camera calibration, the camera intrinsic parameters in the above formula can be obtained , the rotation matrix , the translation vector Other camera parameters.

[0035] S3. Detect and track vehicle targets for each frame of image, assign a unique ID to each vehicle target, and extract the first position information of the vehicle target in the image pixel coordinate system.

[0036] S31. Load image target detection model The pre-trained deep learning target detection neural network model is loaded into the memory, and a calling interface is provided for inputting the original image information data. The deep learning target detection neural network model can be, but is not limited to, a two-stage image target detection algorithm based on candidate regions, such as R-CNN and Fast R-CNN; it can also be a one-stage image target detection algorithm such as SSD, YOLOv5, YOLOv8, and YOLOv10 that directly generates the category probability and position coordinate values ​​of the object. Generally speaking, the two-stage algorithm has an advantage in the accuracy of the detection results. The one-stage detection algorithm has an advantage in speed and is more suitable for front-end monitoring scenarios where hardware performance is limited but real-time detection is required.

[0037] The deep learning target detection neural network model of this embodiment identifies specific target objects including but not limited to small cars, trucks, large buses, special vehicles and other types of motor vehicles that meet the requirements of road traffic management based on the needs of actual road application scenarios.

[0038] Step S32: Image detection to obtain target information By inputting a frame of raw RGB image data from the camera's surveillance footage into the loaded neural network model interface, the target information contained in the frame's raw image information can be determined based on the output of the neural network model. Each target information in the detection result includes two parts: the target's category and location coordinates.

[0039] In this embodiment, the first position information is the position coordinate information of the target detection result frame in the pixel coordinate system of the monocular surveillance camera image. The first position information represents the original position information of the target detection frame in the RGB image. Its mathematical representation is a rectangular frame in a two-dimensional coordinate system, which requires at least two sets of coordinate points to determine. Optionally, the coordinates of the diagonal vertices of the rectangular frame in the image pixel coordinate system, or the coordinates of the center of the rectangular frame in the image pixel coordinate system, as well as the width and height of the rectangular frame.

[0040] The first position information can be the pixel coordinates of the two corner points of the target detection rectangle on any diagonal line in the monocular camera coordinate system. ,in, It can represent the horizontal pixel coordinate value of the upper left corner of the target detection rectangle. Indicates the vertical pixel coordinate value of the upper left corner; the corresponding The first position information is the horizontal pixel coordinate value and the vertical pixel coordinate value of the lower right corner of the target detection rectangle. It is composed of the upper left corner and lower right corner of the rectangular frame.

[0041] Similarly, the first location information You can also choose to use the lower left corner and upper right corner of the rectangular box. Then, the horizontal pixel coordinate values ​​and vertical pixel coordinate values ​​of the lower left corner and upper right corner of the target detection rectangle can be represented respectively.

[0042] The first position information can also be the pixel coordinate values ​​of the center point of the target detection rectangular frame and the width and height of the rectangular frame in the monocular camera coordinate system. ,in, Indicates the horizontal pixel coordinate value of the center point of the target detection rectangle. Indicates the vertical pixel coordinate value of the center point; the corresponding They represent the pixel values ​​occupied by the width and height of the target detection rectangle respectively.

[0043] and The equation relationship between them is:

[0044] Step S33: Target tracking and assigning vehicle ID When detecting a series of continuous RGB images, the same motor vehicle target may be photographed dozens or even hundreds of times, but there is no direct correlation between the results of adjacent image frames. At this time, it is impossible to obtain the first position information of the vehicle in different image frames. association.

[0045] In view of this, based on the image detection results, Kalman filtering and Hungarian matching algorithm are used on a series of continuous RGB image detection results. According to the continuity of vehicle driving, target feature matching between upper and lower different video image frames is achieved, target tracking is performed, and a unique "identity card" is issued to each vehicle target to ensure that the output results are consistent with the vehicle target situation in the real scene.

[0046] In this embodiment, the target tracking technology solution provided is a plug-in solution. The tracking algorithm used for target tracking can be, but is not limited to, open-source plug-and-play target detection algorithms such as SORT, DeepSort, ByteTrack, OC-Sort, Deep OC-Sort, BoT-SORT, and StrongSORT. Alternatively, based on actual scenarios, target characteristics, and other application requirements, parameters of open-source tracking algorithms can be adjusted and optimized, and custom processing logic can be added to obtain improved tracking algorithms with better results.

[0047] The current The detection results of the frame RGB image are consistent with the previous The detection results of the frame image detection are predicted and matched, and the unique identity ID of the vehicle target is obtained based on the original detection information.

[0048] Vehicle ID Build a dataset , the first position information corresponding to the vehicle ID in a series of continuous image frame detection results , added to the dataset Data association is performed in the dataset, and a set of vehicle target information datasets is formed based on vehicle trajectories:

[0049]

[0050] S4. Based on the first position information, crop the vehicle target partial image, and extract all feature point sets in the partial image as second position information.

[0051] In this embodiment, the method for extracting feature points includes converting the local image into a grayscale image; traversing the pixel points of the grayscale image, and determining the pixel points on the circle centered at the point. Whether the absolute value of the grayscale difference between the pixel point and the center point exceeds the threshold If the threshold is exceeded If the number of points is greater than the preset ratio, it is determined to be a feature point; the final feature point set is filtered through non-maximum suppression.

[0052] Specifically, based on the detection results of the RGB image, the vehicle ID " "The first location information The corresponding area is cropped from the original RGB image of the current frame to obtain the minimum circumscribed rectangle local RGB image containing all the feature information of the vehicle .

[0053] First, the local RGB image Calculate the grayscale value to obtain a simplified grayscale image .

[0054] Then the local grayscale image Extract feature points, which are composed of key points and descriptors. Key points are feature points in grayscale images. The position in the keypoint, and sometimes also the direction and size. A descriptor is a vector that describes the information of the pixels around the keypoint. The descriptor information between feature keypoints with similar appearance is similar.

[0055] Calculate candidate points The formula for determining whether a point is a feature point is as follows:

[0056] in, The candidate point is the center of the circle, A circle with a radius of . yes Any point on the circumference, is the pixel gray value of the point, The candidate point for the center of the circle The pixel gray value. It is the threshold of the difference between the grayscale values ​​of two pixels when the difference is obvious. are all selected points If N is greater than a given threshold, which is generally set to 0.75, then the candidate point is considered is a feature point, otherwise, the candidate point Not a feature point.

[0057] In this embodiment, the pixel point is the center of the circle, Generates a circle for a circle of radius In the circle Select the top as the starting point and select 16 evenly distributed pixels in a clockwise direction. Based on the selected 16 pixels, get the pixel gray value of each pixel ~ . Then 16 pixels Pixel grayscale value ~ Pixels Pixel grayscale value Compare. Determine the pixel The pixel gray value is the same as the pixel to be compared Is the absolute value of the difference between the pixel gray values ​​greater than the gray difference threshold? ,in, Indicates the preset pixel gray value change threshold. If yes, the result value is 1, otherwise the result value is 0. And the 16 result values ​​are accumulated and summed up to get N. Further determine whether the value of N is not less than 12 (16*0.75=12). If so, determine the pixel point is a feature point, otherwise, determine the pixel point Not a feature point. Loop through the grayscale image Each pixel in , thus obtaining a local RGB image All the characteristic points.

[0058] In a further solution of this embodiment, the local RGB image The operation of calculating candidate feature points for each pixel in the image is performed. When all the original feature corner points are initially obtained, the original feature corner points are very likely to be densely distributed and clustered. It is necessary to further discard the feature corner points that are too close together through maximum value suppression and only retain the core feature key points:

[0059] in, The original feature points is the center of the circle, A circle with a radius of . yes Any point on the circumference, is the pixel gray value of the point, The original feature point of the circle center Pixel grayscale value of the original feature point The value range is based on the current pixel point As the center, All original feature points within the neighborhood of the radius. Represents the original feature points The optimal feature point in the neighborhood, the rest of the pixels need to be discarded.

[0060] Generally, local RGB images Finally, more than a dozen optimal feature points can be calculated. "The first location information The corresponding local RGB image The set of all feature points extracted is used as the second position information . The second location information Also added to the dataset middle.

[0061]

[0062]

[0063]

[0064] S5. Match the second position information of the same vehicle ID in adjacent frames and calculate the three-dimensional motion displacement vector of the vehicle target in the real-world coordinate system.

[0065] When the camera is monocular, it can only obtain the two-dimensional pixel coordinates of the object when it is imaged, and cannot directly obtain the absolute scale information of the object. Therefore, when mapping it back to the real world scene, the system cannot determine the actual size and true position of the target. In this case, it is necessary to calculate the target motion based on two sets of two-dimensional pixel coordinates.

[0066] In this embodiment, the calculation method of the three-dimensional motion displacement vector includes selecting the matching feature point with the largest sum of vertical coordinate pixel values ​​in the second position information as the reference feature point; substituting the reference feature point into the camera imaging model, setting the initial depth coordinate , calculate the initial three-dimensional coordinates of the reference feature points in the real world coordinate system; generate the three-dimensional motion displacement vector based on the difference in the initial three-dimensional coordinates of the reference feature points of the vehicle target in adjacent frames .

[0067] Specifically, in the above steps of this embodiment, the current The vehicle in the frame RGB image is different from the previous The same vehicle target in the frame image establishes a unique ID for matching and association. Frame image continuously detects and tracks the vehicle ID ", get the first location information . Calculate and obtain the first position information The corresponding local RGB image The second position information of all feature points in the set . From a pre-built dataset Get the vehicle ID in the query "In the previous Second position information in the frame image . Set the second position information of the feature point With the second location information Perform feature point matching. Specifically, the second position information Each feature point in With the second location information All feature points in Calculate the descriptor distance in sequence. Then from the second position information Select the feature point with the smallest descriptor distance as the matching point. Traverse the second position information All feature points in the second position information are found in turn The matching points in .

[0068] As the vehicle target moves, it may be blocked by other objects, causing some feature points to disappear or reappear. Feature points that cannot be successfully matched are treated as redundant information and will be recovered later.

[0069] Assume that a continuously moving object in the real world is regarded as a point mass : At the moment , the object moves to point The world coordinates are:

[0070] At this point, there are:

[0071] in, Yes The imaging point in the camera image is . is the camera calibration matrix parameter of the monocular surveillance camera obtained in the above step S2.

[0072] Similarly, we know that at time , the point moves to point The world coordinates of , Yes The imaging point in the camera image is .

[0073]

[0074] in, is the target's three-dimensional motion change in the real world exist The displacement component in the axial dimension, is The displacement component in the axial dimension.

[0075] Calculate the second position information and second location information The sum of the vertical coordinate pixel values ​​of each set of feature points that match in the , and take the set of feature points with the largest sum of vertical coordinate pixel values ​​as the reference feature points Feature points The coordinate value of can be directly picked up from the pixel coordinate system of the RGB image. Substitute into the above formulas respectively, and let , we can calculate the corresponding three-dimensional point coordinates in the real world . The target's three-dimensional motion change in the real world can be obtained :

[0076] S6. Calculate the three-dimensional coordinates of the vehicle feature points in the real-world coordinate system based on the three-dimensional motion displacement vector and the camera calibration matrix, and fit the three-dimensional position and size information of the vehicle target.

[0077] The three-dimensional motion displacement vector Substitute the camera imaging model to solve the three-dimensional coordinates of the non-reference feature points in the real-world coordinate system; integrate the three-dimensional coordinates of all feature points, restore the three-dimensional outline of the vehicle target through spatial fitting, and output the position information, size information and motion state parameters of the vehicle target in the real-world coordinate system.

[0078] According to the overall consistency of the object, the second position information can be known and second location information Each set of matching feature points in the 3D coordinates of the real-world coordinate system always maintains the same relative distance. Then there is a relationship between any set of feature points:

[0079] in, is the target's three-dimensional motion change in the real world exist The displacement component in the axial dimension, is The displacement component in the axial direction. According to this formula, any set of known feature point coordinates can be used to , calculated .

[0080] Repeat the above calculation formula for all feature points to further obtain the target second position information and second location information The three-dimensional coordinate value of each set of feature points matched in the real world coordinate system is used as the third position information .

[0081]

[0082] Third-party location information Each set of 3D coordinate points in the dataset is integrated based on their spatial relationships, and internal coordinate points are filtered. Through data fitting, the target's specific 3D spatial information, such as its shape, size, and dimensions, is transformed from a scattered distribution to a more regular pattern that can be expressed using mathematical functions. This results in the vehicle's 3D position coordinates, accurately capturing the vehicle's 3D information in real-world coordinates.

[0083] In a further solution of this embodiment, based on the three-dimensional information of the vehicle in the real world coordinates, the distance from the target to the camera installation position in the real world and the physical property related parameters such as the length and width of the vehicle body can be directly obtained.

[0084] Based on the detection results and lane areas of the RGB image of the same frame, the lane to which each vehicle belongs and its position information in the lane can be analyzed and determined. Traffic parameters such as the headway, body distance, vehicle density, and space occupancy of adjacent vehicles in the same lane at the current moment can also be calculated.

[0085] Based on the lane a vehicle belongs to and its position within that lane, it can not only determine whether the vehicle is moving normally or parked, but also analyze and judge whether the vehicle has engaged in traffic incidents such as crossing a lane or changing lanes across a lane. Furthermore, the time difference between adjacent frames can be calculated based on the time interval between the frames. The ratio of the vehicle's displacement vector to the time difference can be used to calculate the real-time speed of each vehicle.

[0086] See also Figure 2 The imaging diagram of a monocular camera monitoring is shown in FIG. Indicates the road, Indicates the camera mounting rod. Represents the camera imaging surface. At time , the vehicle moves to point , then there are feature points Feature points The imaging points in the camera image are .at this time, The three-dimensional coordinates of are unknown. The two-dimensional pixel coordinate values ​​can be directly obtained from the camera image.

[0087] It should be noted that if we directly use the imaging point and camera calibration parameters, and take , and perform calculations. Calculate the restored three-dimensional coordinate value (let ) will correctly correspond to the points in the real world coordinate system .but Calculate the restored three-dimensional coordinate value (let ) The corresponding point in the real world coordinate system will be incorrectly mapped to in the position.

[0088] At the moment , the vehicle moves to point , then there are feature points Feature points The imaging points in the camera image are .

[0089] Similarly, if we use the imaging point directly and camera calibration parameters. Calculate the restored three-dimensional coordinate value (let ) will correctly correspond to the points in the real world coordinate system .but Calculate the restored three-dimensional coordinate value (let ) The corresponding point in the real world coordinate system will be incorrectly mapped to in the position.

[0090] pass Solve and restore the correct three-dimensional points in the real-world coordinate system .from arrive The horizontal displacement between The vehicle from time At the time The amount of motion displacement . Then according to the vehicle from time At the time The amount of motion displacement , substitute into the formula, now, The imaging point can be obtained by solving The corresponding three-dimensional point in the real-world coordinate system .

[0091] This method leverages prior information about the vehicle's three-dimensional spatial position and overall motion consistency to extract the motion displacement of ground points, accurately estimating the target's overall motion scale. This method accurately calculates and restores the vehicle's true three-dimensional position for cameras installed at various heights and angles. It also fits the target's characteristic points, transforming their scattered distribution into a more regular pattern that can be expressed using mathematical functions, reducing system complexity and computing power.

[0092] The present invention calculates the real-world three-dimensional positional relationship between vehicles and lanes, and between vehicles and vehicles more accurately by calculating the distance and size of the target in the real world. It also supports further analysis and judgment to obtain the lane to which each vehicle belongs and its position information in the lane, and calculates traffic parameters such as the head distance, body distance, vehicle density, and space occupancy of adjacent vehicles in the same lane at the current moment. It also supports not only solving and analyzing whether the vehicle is driving or parking normally based on the lane to which the vehicle belongs and its position information in the lane, but also analyzing and judging whether the vehicle has traffic incidents such as crossing the line or changing lanes across the line. On this basis, the time difference can be further obtained based on the generation time interval between adjacent frames. The real-time speed of each vehicle can be calculated by the ratio of the vehicle displacement vector and the time difference.

[0093] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for calculating the real position of a vehicle in a monocular camera monitoring scene, characterized in that: The following steps are involved: S1. Obtain the real-time video stream information of the monocular surveillance camera and decode it to generate a continuous frame image sequence; S2. Construct the camera calibration matrix of the monocular surveillance camera, including the intrinsic parameter matrix , extrinsic rotation matrix and translation matrix ; S3, detecting and tracking a vehicle target for each frame of image, assigning a unique ID to each vehicle target, and extracting the first position information of the vehicle target in the image pixel coordinate system; S4. Based on the first position information, crop the vehicle target partial image, and extract all feature point sets in the partial image as second position information; S5. Match the second position information of the same vehicle ID in adjacent frames and calculate the three-dimensional motion displacement vector of the vehicle target in the real-world coordinate system; S6. Calculate the three-dimensional coordinates of the vehicle feature points in the real-world coordinate system based on the three-dimensional motion displacement vector and the camera calibration matrix, and generate the three-dimensional position and size information of the vehicle target by fitting.

2. The method for calculating the real position of a vehicle in a monocular camera monitoring scene according to claim 1, characterized in that: The calculation method of the three-dimensional motion displacement vector in S5 includes: Selecting a matching feature point with the largest sum of vertical coordinate pixel values ​​in the second position information as a reference feature point; Substitute the reference feature points into the camera imaging model and set the initial depth coordinates , calculating the initial three-dimensional coordinates of the reference feature point in the real-world coordinate system; Generate a three-dimensional motion displacement vector based on the difference in the initial three-dimensional coordinates of the reference feature points of the vehicle target in adjacent frames .

3. The method for calculating the vehicle's true position in a monocular camera monitoring scene according to claim 2, characterized in that: The S6 includes: The three-dimensional motion displacement vector Substituting the camera imaging model into the real world coordinate system, and calculating the three-dimensional coordinates of the non-reference feature points; The three-dimensional coordinates of all feature points are integrated, and the three-dimensional outline of the vehicle target is restored through spatial fitting, and the position information, size information and motion state parameters of the vehicle target in the real-world coordinate system are output.

4. The method for calculating the real position of a vehicle in a monocular camera monitoring scene according to claim 3 is characterized in that: The three-dimensional motion displacement vector Substituting into the camera imaging model, the following relationship is satisfied: in, 、 They are the three-dimensional motion displacement vectors of the vehicle target in the real world exist Axis and The displacement component in the axial direction.

5. The method for calculating the vehicle's true position in a monocular camera monitoring scene according to claim 1, characterized in that: The method for extracting feature points in S4 includes: Converting the local image into a grayscale image; Traverse the pixel points of the grayscale image and determine the pixel points on the circle centered at the point. Whether the absolute value of the grayscale difference between the pixel point and the center point exceeds the threshold ; If the threshold is exceeded If the percentage of points is greater than the preset ratio, it is determined to be a feature point; The final feature point set is filtered by non-maximum suppression.

6. The method for calculating the vehicle's true position in a monocular camera monitoring scene according to claim 1, characterized in that: The first position information in S3 is expressed as a rectangular frame in a two-dimensional coordinate system, specifically: The coordinates of the diagonal vertices of the rectangular frame in the image pixel coordinate system; or, The center coordinates of the rectangular frame in the image pixel coordinate system and the width and height of the rectangular frame.

7. The method for calculating the real position of a vehicle in a monocular camera monitoring scene according to claim 1, characterized in that: The camera calibration method in S2 includes any one of a traditional camera calibration method, a camera self-calibration algorithm, or a Zhang Zhengyou calibration method.

Citation Information

Patent Citations

  • Method for realizing SLAM positioning based on monocular vision and related device

    CN111928842A

  • Method and device for estimating GPS coordinates of multiple target objects and tracking target objects on basis of camera image information about unmanned aerial vehicle

    WO2024096691A1