Traffic violation judgment system and method based on monocular camera

Through perspective distortion correction and dynamic calibration of monocular cameras, the problem of difficulty in accurately calculating absolute distances at the camera terminals at traffic intersections is solved, and low-cost and efficient traffic violation detection and recording are achieved.

CN120472682APending Publication Date: 2025-08-12BEIJING VOCATIONAL COLLEGE OF TRANSPORTATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510748695.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and at low cost to calculate the absolute distance between pedestrians or vehicles and the reference position at traffic intersection camera terminals, resulting in difficulty in judging violations, especially in large field of view and complex scenarios.

Method used

A monocular camera is used to combine perspective distortion correction and dynamic calibration methods. By installing a monocular camera at a traffic intersection, the imaging field of view is obtained and perspective distortion correction is performed. The absolute distance between the target and the camera is calculated using the dynamic marker positioning parameters of the static marker.

Benefits of technology

It realizes accurate measurement of the absolute distance of multiple targets without increasing hardware costs and complexity, supports real-time traffic violation judgment and recording, and reduces the burden of manual review, and is suitable for violation detection in existing monitoring blind spots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472682A_ABST
    Figure CN120472682A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of traffic violation positioning, and provides a traffic violation judgment method based on a monocular camera in order to solve the problems of complex calculation and high cost of an absolute distance between a traffic intersection target and a reference position in the prior art, and the method comprises the following steps: determining an imaging visual field of the monocular camera by taking a mounting position of the monocular camera as a reference point; determining a first image of the monocular camera based on the imaging field of view of the monocular camera; setting a basic hypothesis according to a proportional relation between an imaging view field of the monocular camera and each edge in the first image; dynamically calibrating the monocular camera based on the static marker to obtain calibration parameters; based on the calibration parameters, positioning a target in the view field of the monocular camera, and determining an absolute distance between the target and the monocular camera; and judging whether the target violates rules or not based on the obtained absolute distance between the target and the monocular camera. According to the invention, the absolute distance between the plurality of cameras and the target can be measured while the visual information is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic violation positioning, and in particular to a traffic violation judgment system and method based on a monocular camera. Background Art

[0002] Traffic violations generally refer to violations committed by drivers, pedestrians, passengers, and any entities or individuals involved in road traffic. Common violations include running red lights, crossing solid lines, and driving outside of lanes. Traffic violation location tracking uses technology (such as electronic monitoring, satellite positioning, and image recognition) to track and record violations on the road in real time or after the fact, thereby ensuring road safety and order.

[0003] Traffic lights and cameras are typically deployed at intersections, and cameras are also deployed at non-intersection locations on highways. Traffic lights guide pedestrians and vehicles, while cameras record the scene 24 hours a day. Typically, the camera terminal includes a low-cost chip for acquiring, processing, and uploading video signals. The video recorded by the cameras serves two main purposes: 1. It facilitates backend staff to spot-check pedestrian or vehicle violations; 2. It facilitates scene reconstruction when a traffic accident occurs at an intersection. For the first purpose, the traditional approach relies on backend staff spot-checking, which is labor-intensive and inefficient, and often misses violations. A more advanced approach now involves deploying AI analysis algorithms in the backend to automatically identify violations, which are then manually analyzed for targeted results. However, due to the large number of intersections and cameras, uploading all video to backend servers for real-time analysis would place significant strain on bandwidth and computing power. Therefore, through the chip and algorithm of the camera terminal, the camera terminal can judge in real time whether there are any violations and what kind of violations there are in the video signal, and then upload it to the background. This can improve the violation investigation and punishment rate, reduce the workload of the background staff, and reduce the pressure on the background server.

[0004] Currently, determining whether a traffic violation has occurred and the type of violation requires at least two conditions: 1) identification of pedestrians and vehicles within the field of view; and 2) accurate calculation of the distance between the pedestrian or vehicle and a reference position. For the first condition, various methods have been developed, such as the open-source YOLO series of algorithms, which can perform pedestrian and vehicle recognition within camera terminals at low cost, with mature algorithms and good results.

[0005] For the second condition, the commonly used distance measurement methods at traffic intersections include the following:

[0006] First, LiDAR. LiDAR has a transmitting aperture and a receiving aperture for radar waves. It can use triangulation and Time of Flight (ToF) to measure absolute distance. The accuracy of the former decreases quadratically with increasing distance, while the latter uses the extremely rapid motion of photons and the time it takes to calculate distance. Because photons move so quickly, the timer must be very accurate. Within its normal operating range, LiDAR has very high ranging accuracy, but its hardware cost is high, and the number of spatial points it can scan is small. The longer the distance, the lower the resolution. Traffic intersections typically have a large field of view, which means that LiDAR can capture a small number of target pedestrians or vehicles simultaneously. Furthermore, it can typically only obtain longitudinal distance (i.e., the distance between the target vehicle or pedestrian and the LiDAR transmitter). This longitudinal distance information is difficult to combine with camera video and traffic sign information to determine the presence and type of traffic violations. Therefore, LiDAR is typically used to determine speeding violations and is not suitable for violations such as illegal lane changes or driving across lanes.

[0007] Second, structured light cameras use structured light or Time of Flight (ToF) technology to measure the absolute distance of objects. While these cameras offer relatively accurate results, they are expensive and require high computing power. Furthermore, the information captured by structured light cameras can generally only be used to calculate distance and is difficult to use for pedestrian or vehicle recognition.

[0008] Third, ultrasonic devices use the emission and reception of sound waves to measure absolute distance, but cannot simultaneously obtain other environmental information, such as visual information.

[0009] Fourth, binocular cameras. Two cameras are rigidly connected, and the line connecting them is called the baseline, which has a fixed length. For a fixed point P in space, its projections on the imaging surfaces of the two cameras are denoted as P1 and P2. According to the principle of camera imaging, P-P1-P2 forms a triangle, and the pixel distances between P1 and the edge of each image are generally different. For example, suppose the binocular cameras are placed horizontally, and the pixel distance between P1 and the left edge of the image of the first camera is a1, and the pixel distance between P2 and the left edge of the image of the second camera is a2. Generally speaking, a1 is not equal to a2 (when the line connecting P and the midpoint of the baseline is perpendicular to the baseline, they are equal). Since the absolute length of the baseline is known, the absolute distance from point P to the camera can be calculated through certain mathematical calculations. The absolute distance obtained by binocular cameras is relatively accurate, but the upper limit of the distance it can measure is limited by the baseline length. In addition, the hardware of binocular cameras is relatively complex and costly, and there are certain blind spots in the field of view.

[0010] Fifth, consider monocular cameras. Generally, monocular cameras need to be in motion to calculate the distance to a target object using motion difference, and this distance is relative, not absolute. Consider a moving monocular camera. At a certain moment, consider spatial points A and B, located at different distances along its principal optical axis. Their projections on the camera's imaging plane are the same spatial point O, which is also the optical center of the camera's imaging plane. A-B-O is a straight line perpendicular to the camera's imaging plane at point O. At this point, it is impossible to distinguish between points A and B. Next, the monocular camera moves forward a certain distance, and its optical center moves to a new spatial point, O'. A-B-O' is no longer a straight line, but instead forms a triangle. The projections of A' and B on the camera's imaging plane no longer overlap. The pixel difference between the pixels of A' and B' and the optical center can be used to determine the distance between points A and B and the camera. The larger the pixel difference, the closer the distance. However, only the relative distance between A' and B' can be determined; their absolute distance from the camera cannot be determined. This indicates that monocular cameras suffer from scale uncertainty when measuring distance.

[0011] In summary, the current methods for calculating the distance between pedestrians or vehicles and a reference position have problems such as high computational cost, complex hardware, or uncertain measurement scale. Summary of the Invention

[0012] In view of the shortcomings of the above-mentioned prior art, the present invention proposes a traffic violation judgment system and method based on a monocular camera, which is used to measure the absolute distance between the camera on the road and multiple target persons and vehicles within the field of view.

[0013] In a first aspect, the present invention discloses a method for determining traffic violations based on a monocular camera, the method comprising:

[0014] Install a monocular camera at the traffic intersection to be inspected, and determine the projection point and projection distance of the monocular camera;

[0015] A rectangular area is obtained in the ground shooting field of view of the monocular camera as the imaging field of view, and at least one target to be measured is located in the imaging field of view;

[0016] Performing perspective distortion correction on an image of an imaging field of view captured by a camera to obtain a first corrected image, and setting a basic assumption for the first corrected image;

[0017] Determine marker points within the imaging field of view, and determine a first distance between each marker point and the projection point, and a second distance between two adjacent marker points;

[0018] According to the basic assumption, dynamically calibrate the pose parameters of the camera based on the pixel coordinates of the first corrected image and the first distance and the second distance;

[0019] Identify the imaging target in the first corrected image, obtain the pixel coordinates of the imaging target, and then determine the absolute distance between the target to be measured and the monocular camera in combination with the pose parameters;

[0020] Based on the obtained absolute distance between the target and the monocular camera, the target violation category and the violation confidence corresponding to the target violation category are determined.

[0021] In a second aspect, the present invention discloses a traffic violation judgment system based on a monocular camera, the system comprising:

[0022] The determination module is used to install a monocular camera at the traffic intersection to be detected and determine the projection point and projection distance of the monocular camera.

[0023] The imaging field of view determination module is used to obtain a rectangular area in the ground shooting field of view of the monocular camera as the imaging field of view, and at least one target to be measured is located in the imaging field of view.

[0024] The acquisition module is used to perform perspective distortion correction on the image of the imaging field of view area captured by the camera to obtain a first corrected image, and set a basic hypothesis for the first corrected image.

[0025] A distance calculation module determines the marker points within the imaging field of view, and determines a first distance between each marker point and the projection point, and a second distance between two adjacent marker points;

[0026] A parameter calibration module dynamically calibrates the camera's pose parameters based on the pixel coordinates of the first calibrated image, the first distance, and the second distance according to a basic assumption;

[0027] The absolute distance determination module identifies the imaging target in the first corrected image, obtains the pixel coordinates of the imaging target, and then determines the absolute distance between the target to be measured and the monocular camera in combination with the posture parameters;

[0028] The traffic violation positioning module is used to determine the target violation category and the violation confidence corresponding to the target violation category based on the absolute distance between the target and the monocular camera.

[0029] The traffic violation positioning module is used to determine whether the target has violated the traffic rules based on the absolute distance between the target and the monocular camera.

[0030] Therefore, the present invention adopts the above-mentioned traffic violation location method and device based on measuring absolute distance with a monocular camera, which has the following beneficial effects:

[0031] The present invention can use a monocular camera to measure absolute distances in multiple directions in specific scenarios without adding additional components, structural complexity, and hardware costs; while the present invention only uses a monocular camera to obtain visual information, it can also measure the absolute distance between the camera on the road and multiple target persons and vehicles within the field of view, thereby making real-time judgments and records of traffic violations, avoiding manual review of massive video recordings and consuming a lot of manpower. In addition, the real-time judgment result data can also be used as auxiliary evidence in legal disputes to prove the scientific nature of the traffic violation judgment process. In addition, the present invention is applicable to existing monitoring blind spots (such as intersections without traffic lights), and the absolute distance between the target and the dynamic baseline can be measured by this method to detect illegal lane changes, road occupation, and other behaviors.

[0032] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a flow chart of the traffic violation location method proposed by the present invention using a monocular camera to measure absolute distance.

[0034] Figure 2 This is a schematic diagram of a scene device in the prior art.

[0035] Figure 3 Schematic diagram of the monocular camera's imaging field of view.

[0036] Figure 4 Schematic diagram of the pixel image captured by a monocular camera.

[0037] Figure 5 Schematic diagram of the pixel image captured by the monocular camera after perspective distortion correction.

[0038] Figure 6 The original image of the Go board captured by the monocular camera.

[0039] Figure 7 This is the result image after perspective distortion correction based on the four outer corner points of the Go board.

[0040] Figure 8 The image is the result of perspective distortion correction based on the four corners of the Go board where the pieces can be placed.

[0041] Figure 9 Schematic diagram of the imaging field of view during dynamic calibration of a monocular camera.

[0042] Figure 10 Schematic diagram of the monocular camera's imaging field of view when calculating the absolute distance between the target and the monocular camera.

[0043] Figure 11Schematic diagram of the imaging field of view after perspective distortion correction when calculating the absolute distance between the target and the monocular camera. DETAILED DESCRIPTION

[0044] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0045] Example 1

[0046] The monocular camera of the present invention is installed on a scene device, and the scene device can be a column for fixing the camera, a beam on the column and related fasteners, such as Figure 2 Before using the positioning method proposed in the present invention to measure absolute distance, the monocular camera and the scene device are assembled to ensure that the monocular camera is stably mounted on the scene device. The monocular camera is then deployed on top of the traffic light pole at the intersection using the scene device. Furthermore, the monocular camera can be a standard black and white camera or a color camera.

[0047] Reference Figure 1 , Figure 1 This embodiment provides a method for determining traffic violations based on a monocular camera, which may include:

[0048] Step 1: Install a monocular camera at the traffic intersection to be detected and determine the projection point and projection distance of the monocular camera;

[0049] As described above, the monocular camera is mounted on the scene device, and the position of the scene device is used as the reference point. In the embodiment of the present application, it is assumed that the road surface is a horizontal surface. Figure 3 As shown, the monocular camera is fixed at point O. The projection point S of the fixed point O on the road is the monocular camera's projection point on the road. At this time, the camera is in its initial position. Point O is the center point of the scene device. During the subsequent rotation of the monocular camera, the scene device remains stationary relative to the ground, while the monocular camera can rotate relative to the scene device.

[0050] like Figure 3 As shown, draw a perpendicular line from the fixed point O to the horizontal plane (i.e. the road surface), and the perpendicular point obtained is the projection point S of point O on the horizontal plane.

[0051] Step 2: Obtain a rectangular area in the ground shooting field of view of the monocular camera as the imaging field of view, and at least one target to be measured is located in the imaging field of view;

[0052] In traffic violation photography, monocular cameras are usually installed at higher positions such as traffic pillars, and the target of the photography is on the road below it, so the main optical axis of the camera is facing downward and straight ahead. Assuming that the main optical axis intersects the horizontal plane at point P, point P is the bottom point of the image. According to the imaging principle of the camera, when the angle between the main optical axis OP of the camera and the horizontal plane is an acute angle, the camera's shooting field of view is a trapezoid, that is, Figure 3 The trapezoid A1B1CD in the figure is A1B1 / / CD, and the length of A1B1 is greater than the length of CD. The specific size of the trapezoid A1B1CD and the positions of the four vertices can be determined according to the specific model and parameters of the monocular camera.

[0053] Connect the projection point S and point P, and extend the line SP until it intersects the lower base A1B1 of the trapezoid A1B1CD. Figure 3 It can be seen that SP intersects with the upper base CD of trapezoid A1B1CD at T, and the extension line of SP intersects with the lower base A1B1 at M. Points T and M are the first and second intersection points respectively; at the same time, point T is the midpoint of the upper base CD, and point M is the midpoint of the lower base A1B1.

[0054] Draw parallel lines CB and DA through the two endpoints C and D of the upper base CD, where TM is the line connecting the first intersection point T and the second intersection point M; Figure 3 It can be seen that CD = AB, meaning that point M is the midpoint of AB. Also, according to the projection relationship, points B and A are the feet of perpendiculars CB and DA, respectively, on the lower base A1B1. The rectangular area ABCD enclosed by AB, CD, AD, and BC serves as the imaging field of view, with TM being the midline. The direction from point A to point B is the horizontal direction of the image, and the direction from point D to point A is the vertical direction of the image.

[0055] It should be noted that, for ease of observation, Figure 3 The A1B1CD in the figure is illustrated as a trapezoid with thickness, but in fact A1B1CD as a working surface has no thickness. The following A1B1CD has no thickness.

[0056] The following briefly introduces the causes of perspective distortion in monocular camera imaging. The corresponding image captured by the monocular camera should be a Figure 4 The vertices of the rectangle A′1B′1C′D′ shown are respectively Figure 3 The vertices A1, B1, C, and D of the first trapezoid in the image correspond to each other. The size of the rectangle A'1B'1C'D' is w×h (w is width, h is height, and the unit is pixel), where the length of A'1B'1 is w pixels, the length of C'D' is also w pixels, the length of A'1D' is h pixels, and the length of B'1C' is h pixels. Connecting the diagonals A'1C' and B'1D', we can determine the nadir point P at Figure 4At the same time, according to the positions of points M and T on A1B1 and CD (at the midpoint), points M' and T' can be determined in the rectangle A'1B'1C'D'. T' and M' are the midpoints of A'1B'1 and C'D' respectively. Correspondingly, according to Figure 3 The positional relationship between midpoint A, point B and points A1, B1 can be found in Figure 4 Determine the corresponding points A', B', and connect A', B', C', and D' to get the trapezoid A'B'C'D'. The trapezoid A'B'C'D' is the image of the imaging field of view captured by the monocular camera. Obviously, Figure 3-Figure 4 Although the actual length of AB is equal to the actual length of CD ( Figure 3 ), but the pixel length of A′B′ is smaller than the pixel length of C′D′ ( Figure 4 This is due to the perspective principle of camera imaging and has nothing to do with camera manufacturing errors. The specific reason is that the length of A1B1 is greater than the length of CD. Therefore, when imaging, the compression factor between the actual length of A1B1 and the pixel length of A′1B′1 is greater than the compression factor between the actual length of CD and the pixel length of C′D′, which leads to this perspective distortion.

[0057] Step 3: Perform perspective distortion correction on the image of the imaging field of view area captured by the camera to obtain a first corrected image, and set a basic assumption for the first corrected image; wherein the basic assumption is: the image after perspective distortion correction has the same horizontal pixel ratio as the horizontal length ratio of the real plane and the same vertical pixel ratio as the vertical length ratio of the real plane.

[0058] For the perspective distortion caused by the perspective principle mentioned above, existing perspective distortion correction methods can be used to restore the image to its original proportions. These methods can include feature point-based homography, geometric transformation, and camera calibration using a calibration plate. These methods are all known in the art, and their principles will not be elaborated here.

[0059] After perspective distortion correction, the Figure 4 The image of the imaging field area with perspective distortion shown in FIG is restored to Figure 5 The rectangle shown is the first corrected image A′B′C′D′. Because AB = CD and AB / / CD, A′B′ = C′D′. It's important to note that the number and spacing of photosensitive elements in the horizontal and vertical directions of a camera differ, causing the ratio of the horizontal true length to the pixel length of the corresponding line segment to differ from the ratio of the vertical true length to the pixel length. For example, if AB = AD, then generally speaking, A′B′ ≠ A′D′.

[0060] according to Figure 5 , the image points A', B', C', D' in the first corrected image A'B'C'D' correspond to the four vertices A, B, C, D of the imaging field of view ABCD, respectively, such as the image point A' corresponds to the vertex A, and at the same time, the pixel midline T'M' of A'B'C'D' corresponds to the field of view midline TM of ABCD, such as Figure 3 As shown, the endpoints T and M of the visual field midline TM are located on the broadside CD and the broadside AB respectively, and the distance between the broadside CD and the projection point S is smaller than the distance between the broadside AB and the projection point S.

[0061] Based on the above analysis, the following assumptions are made:

[0062] Assumption 1: The line OS connecting the monocular camera's reference point O and its projection point S is perpendicular to the horizontal plane, with S as the vertical point. The height of OS is a known and fixed value, and the value of OS is denoted by l. During camera operation, l does not change and can be obtained manually.

[0063] Assumption 2: The width of the imaging field of view ABCD of the monocular camera is fixed, and the distance between the projection point S and the imaging field of view ABCD and the distance between the projection point S and the image base point P are fixed.

[0064] When the installation position, installation angle and height (l) of the monocular camera from the working surface are fixed and known values, the width AB, CD and length TM of the imaging field of view ABCD are fixed, and the distance ST between the projection point S and the imaging field of view SBCD, as well as the distance SP between the projection point S and the image bottom point P are fixed. In principle, ST, SP, TM, AB, CD can be calculated based on manual measurement, and then it is assumed that they will not change in the future.

[0065] Assumption 3: The monocular camera used has been distortion corrected.

[0066] Due to various distortions generated during lens design, manufacturing, and assembly, images or videos captured by the camera may be distorted. Generally, the images captured by the camera can be corrected after leaving the factory to obtain distortion correction parameters to reduce or eliminate distortion, thereby improving image quality and visual effects. The monocular cameras mentioned in the embodiments of this application are assumed to have undergone distortion correction.

[0067] Assumption 4: All images captured by the monocular camera have been corrected for perspective distortion.

[0068] When a camera takes a picture, the image and video it captures will experience perspective distortion due to the principle of perspective. Perspective distortion is inherent; even if the camera's design, manufacturing, and assembly are flawless, it will still be present during capture. Perspective distortion can be corrected. In the examples of this application, it is assumed that all images used have been processed to eliminate perspective distortion.

[0069] Assumption 5: For an image corrected for perspective distortion, the horizontal pixel ratio is the same as the horizontal length ratio of the true plane, and the vertical pixel ratio is the same as the vertical length ratio of the true plane.

[0070] Assumption 5 is used as the basic assumption for subsequent absolute distance calculations. Under this assumption, the horizontal length ratio of the pixels in the perspective-corrected image is the same as the horizontal length ratio of the actual horizontal ground, and the vertical length ratio of the pixels is the same as the vertical length ratio of the actual horizontal ground.

[0071] In order to verify the rationality of Assumption 5, a verification case is also given in this embodiment for demonstration.

[0072] Verification Case

[0073] In this embodiment, a Go chessboard is used to conduct an experiment to verify the above hypothesis 5. A monocular camera is used to shoot the Go chessboard. The captured chessboard image is as follows: Figure 6 As shown in the figure, each small grid on the Go board is a square with a side length of 2 cm. The horizontal length (from A to T) of the entire board is 41.1 cm (hereinafter referred to as the board width), and the vertical length (from 01 to 19) is 38.6 cm (hereinafter referred to as the board height). The aspect ratio is 1.065 calculated by width / height (41.1÷38.6). The resolution of the board image is 6530×4561. Figure 6 As can be seen from the original chessboard image, this image has undergone perspective distortion. The following uses a perspective distortion correction algorithm (such as the homography matrix method) to handle perspective distortion.

[0074] Based on the four corner points of the chessboard (the four vertices of the outermost quadrilateral black frame), the perspective distortion correction algorithm is used to correct the perspective distortion. The correction result is as follows: Figure 7 As shown, the resolution of the image after perspective distortion correction is 6449×4593.

[0075] Next, use Figure 6 The four corners of the chessboard shown in the figure can be used to place chess pieces. The perspective distortion correction is performed and the correction result is as follows: Figure 8 The image resolution of the image after perspective distortion correction is 6049×6283.

[0076] against Figure 7 , record the pixel coordinates of each placement point on the four sides of the chessboard, the results are shown in Table 1.

[0077]

[0078]

[0079] against Figure 8 , record the pixel coordinates of each placement point on the four sides of the chessboard, and the results are shown in Table 2.

[0080]

[0081]

[0082] Combined with the contents of Table 1, the width and height of each small grid on the chessboard are counted in turn. The results are shown in Table 3:

[0083] Table 3

[0084]

[0085] Combined with the contents of Table 2, the width and height of each small grid on the chessboard are counted in turn. The results are shown in Table 4:

[0086] Table 4

[0087]

[0088]

[0089] The results in Tables 3 and 4 show that in the horizontal direction (width direction) of the perspective-corrected image, the pixel width of each grid is essentially the same, with a maximum coefficient of variation of 0.0060, less than 1%. This is attributed to accidental factors such as the slightly offset coordinates recorded due to the overly thick chessboard lines. Similarly, in the vertical direction (height direction) of the perspective-corrected image, the pixel width of each grid is essentially the same, with a maximum coefficient of variation of 0.0104, or 1.04%. This is attributed to accidental factors such as the slightly offset coordinates recorded due to the overly thick chessboard lines. Considering that each small grid on the chessboard is a 2-cm square, the following conclusions can be drawn:

[0090] 1) After perspective distortion correction, the pixel ratio of the horizontal line segments in the image is the same as the actual length ratio of the corresponding horizontal line segments in the real plane.

[0091] 2) After the perspective distortion is corrected, the pixel ratio of the longitudinal line segment of the image is the same as the actual length ratio of the corresponding longitudinal line segment in the real plane.

[0092] For example, in Figure 5 In the perspective distortion corrected image A′B′C′D′ shown in the figure, the pixel ratio of the horizontal line segments A′M′ to A′B′ is The ratio of the lengths of the transverse segments AM and AB corresponding to A′M′ and A′B′ respectively in the true horizontal plane projection same.

[0093] In addition to perspective distortion, the images captured by the camera also have distortions such as barrel distortion and radial distortion. These distortions are not related to perspective distortion. Generally speaking, the effects of distortions such as barrel distortion and radial distortion on the camera itself can be reduced or eliminated through common calibration methods. In this embodiment, the monocular camera can be dynamically calibrated using a calibration method to eliminate the camera's own distortion.

[0094] Step 4: Determine the marker points within the imaging field of view, and determine the first distance between each marker point and the projection point, as well as the second distance between two adjacent marker points;

[0095] As described in Assumption 2, in principle, the lengths of ST, SP, TM, AB, and CD are fixed. However, the orientations of different roads and intersections vary significantly. To ensure the camera's proper alignment, the camera's pitch and left / right angles must be adjusted during installation. Therefore, cameras are typically equipped with two-dimensional mechanical or electric rotation mechanisms, one for each. Because these two rotation mechanisms are adjustable, they may deflect slightly during operation due to factors such as wind and rain. Because the camera's field of view is relatively long and wide, even slight deflection can result in significant deviations from the distant field of view. To achieve this, it is necessary to use fixed, static ground markers as reference points to dynamically calibrate the camera's position. This involves dynamically measuring ST, TM, AB, and CD. Once measured, they can be used as known values in subsequent calculations. These static markers can be fixed, immovable objects such as pillars and ground markings.

[0096] Optionally, step 4 includes the following sub-steps:

[0097] Step 401: Within the imaging field of view ABCD of the monocular camera, three static markers are used as the first calibration point E, the second calibration point F, and the third calibration point G respectively;

[0098] like Figure 9 As shown, when performing dynamic calibration, three obvious static markers that are not easily changed are selected as the first calibration point E, the second calibration point F, and the third calibration point G within the imaging field of view ABCD of the monocular camera located on the ground.

[0099] Step 402: Connect the projection point S with the first calibration point E, the second calibration point F, and the third calibration point G to obtain the lengths e, f, and g of SE, SF, and SG, respectively;

[0100] Wherein, SE, SF, and SG represent the line SE between the projection point S and the first calibration point E, the line SF between the projection point S and the second calibration point F, and the line SF between the projection point S and the third calibration point E, respectively. By manually measuring the lengths of SE, SF, and SG, the first distances e, f, and g are obtained, respectively, and SE=e, SF=f, and SG=g; e, f, and g are taken as the first distances;

[0101] Step 403: sequentially connect the first calibration point E, the second calibration point F, and the third calibration point G to obtain the lengths l1, l2, and l3 of EF, FG, and GE, respectively;

[0102] EF, FG and GE respectively represent the line EF between the first calibration point E and the second calibration point F, the line FG between the second calibration point F and the third calibration point G, and the line GE between the third calibration point G and the first calibration point E; similar to the first distance, the second distances l1, l2 and l3 can be obtained by manual measurement, and EF=l1, FG=l2, GE=l3.

[0103] Step 5: According to the basic assumption, dynamically calibrate the camera's pose parameters based on the pixel coordinates of the first corrected image and the first distance and the second distance;

[0104] Optionally, step 5 includes the following sub-steps:

[0105] Step 501: Draw a perpendicular line through the first calibration point E, the second calibration point F, and the third calibration point G to the extension line SM of the visual field midline TM of the imaging visual field ABCD to obtain the first perpendicular foot E0, the second perpendicular foot F0, and the third perpendicular foot G0;

[0106] Step 502: Determine the true lengths of the first line segment SE0, the second line segment SF0, the third line segment GG0, and the fourth line segment E0F0 respectively according to the first distance and the second distance;

[0107] The specific calculation process is as follows:

[0108] Extend the line GG0 and the line SE between the third calibration point G and the third perpendicular foot G0 so that GG0 and SE intersect at point H2. Figure 9 shown.

[0109] exist Figure 9In triangle SEF, first calculate the first angle η between lines SE and SF using the law of cosines based on the values of e, f, and l1. Next, in triangle SFG, calculate the second angle μ between lines SF and SG using the law of cosines based on the values of f, g, and l2. Finally, using η and μ as known parameters, calculate the third angle θ between lines SM and SF, which is ∠FSM.

[0110] Combine Figure 9 , the calculation process of the third angle θ is as follows:

[0111] SG0=SG·cos(μ-θ)=gcos(μ-θ)

[0112] G0G2=SG0·tan(θ+η)=gcos(μ-θ)tan(θ+η)

[0113] GG0=SG·sin(μ-θ)=gsin(μ-θ)

[0114] GG2=GG0+G0G2=gsin(μ-θ)+gcos(μ-θ)tan(θ+η)

[0115] SE0=SE·cos(θ+η)=ecos(θ+η)

[0116]

[0117] because so

[0118]

[0119] In △EGG2, ∠EG2G=π / 2-(θ+η), and from the cosine theorem we know that:

[0120]

[0121] In equation (1), there is only one unknown variable, θ. Therefore, θ can be obtained by solving the equation. At this point, η, μ, and θ can all be used as known parameters for subsequent calculations.

[0122] According to θ, η and e, the first line segment SE0 between the projection point S and the foot E0 of the first calibration point E can be calculated by the following formula:

[0123] SE0=SE·cos(θ+η)=ecos(θ+η);

[0124] According to θ and f, the second line segment SF0 between the projection point S and the foot of the perpendicular F0 of the second calibration point F can be calculated by the following formula:

[0125] SF0=SF·cosθ=tfcosθ

[0126] The third line segment GG0 between the projection point S and the foot of the perpendicular G0 of the third calibration point G can be calculated according to θ, μ and g by the following formula:

[0127] GG0=SG·sin(μ-θ)=gsin(μ-θ)

[0128] Based on SE0 and SF0, the fourth line segment E0F0 between the foot of the perpendicular E0 of the first calibration point E and the foot of the perpendicular F0 of the second calibration point f can be obtained:

[0129] E0F0=SE0-SF0=ecos(θ+η)-fcosθ.

[0130] Step 503: Determine the pixel coordinates of each point in the first corrected image A′B′C′D′, and obtain the pixel lengths of the second pixel segment T′M′, the third pixel segment G′G′0, the fourth pixel segment E′0F0′, and the fifth pixel segment E′0T′, as well as the width of the wide side A′B′ of the image in pixels.

[0131] Since the first corrected image A'B'C'D' corresponds to the imaging field of view ABCD, the image points corresponding to T, M, E, F, G, E0, F0, and G0 in A'B'C'D' are T', M', E', F', G', E'0, F0', and G'0. F0'E'0 represents the pixel distance between points E'0 and F0' in the image, and the same applies to the other representations.

[0132] For the convenience of description, the pixel coordinates of D′, A′, B′, and C′ are first set to (0, 0), (0, h), (w, h), and (w, 0), respectively, where w and h are the width and height pixels of the image, respectively.

[0133] Since E, F, and G are landmark objects on the ground, they can be identified by existing image recognition algorithms. The YOLO series image recognition algorithm can be used for recognition. Finally, the pixel coordinates of E′, F′, and G′ corresponding to the first calibration point E, the second calibration point F, and the third calibration point G in A′B′C′D′ are (i e ,j e )、(i f ,j f )、(i g ,j g ).

[0134] Next, since T' and M' are the midpoints of the upper and lower edges of the image, the coordinates of T' and M' corresponding to the first intersection T and the second intersection M in A'B'C'D' are determined to be (w / 2,0) and (w / 2,h) respectively. Since G'0 is on the central axis of the image and G'G'0 / / A'B', the coordinates of G'0 are (w / 2,j g ).

[0135] From the above, we can see that the pixel length of the second pixel line segment T′M′ is h, and the pixel length of G′G′0 is i g -w / 2, the pixel length of E′0F0′ is j e -j f , the pixel length of E′0T′ is j e , the width of A′B′ is w pixels.

[0136] Step 504: Determine the values of each pose parameter based on the basic assumptions; the pose parameters include the visual field centerline TM, the first boundary line segment ST, and the broadside AB of the imaging visual field;

[0137] from Figure 9 From the perspective of , the first boundary line segment ST represents the line segment between the projection point S and the endpoint M. The combination of ST and TM can be used as the distance between the projection point S and the imaging field of view ABCD.

[0138] According to Assumption 5, in the vertical direction of the image, the actual length ratio of the fourth line segment E0F0 to the visual field center line TM is the same as the pixel length ratio of the fourth pixel line segment E′0F0′ to the second pixel line segment T′M′ in the first corrected image. Then:

[0139]

[0140] The true length of the visual field centerline TM is:

[0141]

[0142] Similarly, according to the basic assumption, in the vertical direction of the image, there are:

[0143]

[0144] It can be obtained that the true length of the fifth line segment E0T between the first perpendicular foot E0 and the first intersection point T is:

[0145]

[0146] According to the first line segment SE0 and the fifth line segment E0T, the actual length of the first boundary line segment ST is:

[0147]

[0148] Similarly, based on the basic assumption, in the horizontal direction of the image, the actual length ratio of the third line segment GG0 to the wide side AB of the imaging field of view is equal to the pixel length ratio of the third pixel line segment G′G′0 to the wide side A′B′ of the image in the first corrected image, then:

[0149]

[0150] It can be obtained that the actual width of the broadside AB of the imaging field of view is:

[0151]

[0152] At this point, through dynamic calibration, the calibration parameters AB, ST, and TM can be obtained, which can be used for subsequent violation positioning of the monocular camera.

[0153] Step 6: Identify the imaging target in the first corrected image, obtain the pixel coordinates of the imaging target, and then determine the absolute distance between the target to be measured and the monocular camera in combination with the posture parameters.

[0154] like Figure 10 As shown, the contact point between the target to be measured and the ground in the imaging field of view of the monocular camera (rectangle ABCD) is recorded as point Q, and the projection point of Q on SM is recorded as Q0. The target to be measured can be a pedestrian or a vehicle. Figure 10 In the equation, the absolute distance between the target and the monocular camera is the distance SQ between point Q and the projection point S of the monocular camera.

[0155] The image captured by the monocular camera is corrected for perspective distortion, as shown in the following example: Figure 11 As shown, the image point Q in the corresponding image captured in the imaging field is Q′. First, the pixel coordinate position Q′(i, j) of the target in the image is obtained by the image recognition algorithm in the prior art, and the pixel coordinates of Q′0 are (w / 2, j). The image recognition algorithm can be a YOLO series algorithm, which is generally embedded in the camera. The YOLO series algorithm is an efficient target detection algorithm. Its core idea is to convert the target detection task into a single regression problem, and directly predict the position and category of the bounding box through a single convolutional neural network (CNN). The pixel coordinates of the image points Q′ and Q′0 can be obtained according to the YOLO series algorithm.

[0156] According to assumption 5, in the horizontal direction of the image, then:

[0157]

[0158] According to the calibration parameters AB, substituting equation (4) into the above equation, we can get:

[0159]

[0160] Similarly, according to the basic assumptions, there is the following equation:

[0161]

[0162] According to the calibration parameter TM, substituting formula (2) into the above formula, we can get:

[0163]

[0164] Then according to the calibration parameters

[0165]

[0166] We can get:

[0167]

[0168] So far, QQ0 and SQ0 have been calculated, and combined with Figure 10 ,Since the triangle SQQ0 is a right triangle, the Pythagorean theorem can be used to calculate SQ, that is, the absolute distance between the target and the monocular camera.

[0169] Step 7: Based on the obtained absolute distance between the target and the monocular camera, determine the target violation category and the violation confidence corresponding to the target violation category.

[0170] Specifically, the target violation category refers to the judgment result of whether the target has committed a traffic violation. The target violation category can be represented by a number, such as 1 for violation and 0 for no violation. The violation confidence level indicates the credibility of the target violation category.

[0171] When determining the target violation category, the obtained absolute distance can be compared with the preset threshold range. If the absolute distance is within the threshold range, it means that no traffic violation has occurred. Otherwise, it means that a traffic violation has occurred, and the corresponding confidence level is given.

[0172] During subsequent manual review, the confidence level can be used as an auxiliary indicator to help screen cases that require manual review (e.g., cases with a confidence level of 90% or above are automatically marked).

[0173] In practice, if there is a two-way lane or a one-way three-lane lane in the field of view in front of the camera, and the target to be measured is detected in the field of view to be pressed against the solid line of the lane, then the absolute distance between the target to be measured (i.e., the wheel of the vehicle) and the camera can be calculated by the absolute distance calculation method of the present application. At the same time, the solid line of the lane at the traffic intersection is a static marker, so the distance between the solid line and the projection point of the camera can be obtained by manual measurement, and considering that there are errors in the instrument measurement, a line-pressing threshold value can be preset according to the measured value, and the degree of violation can be judged by the size relationship between the calculated absolute distance and the line-pressing threshold value. Specifically, when the absolute distance is greater than the line-pressing threshold value, the absolute value of the difference between the absolute distance and the line-pressing threshold value is calculated. If the absolute value (i.e., the part exceeding the threshold value) is less than ε (ε is a smaller value), a warning is issued. Otherwise, it is determined that the target to be measured has a traffic violation and a fine is imposed.

[0174] Example 2

[0175] This application utilizes the above-mentioned traffic violation judgment method based on a monocular camera to design a traffic violation judgment system based on a monocular camera. The system may include:

[0176] The determination module is used to install a monocular camera at the traffic intersection to be detected and determine the projection point and projection distance of the monocular camera.

[0177] The imaging field of view determination module is used to obtain a rectangular area in the ground shooting field of view of the monocular camera as the imaging field of view, and at least one target to be measured is located in the imaging field of view.

[0178] The acquisition module is used to perform perspective distortion correction on the image of the imaging field of view area captured by the camera to obtain a first corrected image, and set a basic hypothesis for the first corrected image.

[0179] A distance calculation module determines the marker points within the imaging field of view, and determines a first distance between each marker point and the projection point, and a second distance between two adjacent marker points;

[0180] A parameter calibration module dynamically calibrates the camera's pose parameters based on the pixel coordinates of the first calibrated image, the first distance, and the second distance according to a basic assumption;

[0181] The absolute distance determination module identifies the imaging target in the first corrected image, obtains the pixel coordinates of the imaging target, and then determines the absolute distance between the target to be measured and the monocular camera in combination with the posture parameters;

[0182] The traffic violation positioning module is used to determine the target violation category and the violation confidence corresponding to the target violation category based on the absolute distance between the target and the monocular camera.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A traffic violation judgment method based on a monocular camera, characterized in that: The method includes: Determine the projection point and projection distance of the monocular camera; A rectangular area is obtained in the ground shooting field of view as the imaging field of view, and at least one target to be measured is located in the imaging field of view; Performing perspective distortion correction on the captured image of the imaging field of view to obtain a first corrected image, and setting a basic assumption for the first corrected image; Determine marker points within the imaging field of view, and determine a first distance between each marker point and the projection point, and a second distance between two adjacent marker points; According to the basic assumption, dynamically calibrate the pose parameters of the camera based on the pixel coordinates of the first corrected image and the first distance and the second distance; Identify the imaging target in the first corrected image, obtain the pixel coordinates of the imaging target, and then determine the absolute distance between the target to be measured and the monocular camera in combination with the pose parameters; Based on the obtained absolute distance between the target and the monocular camera, the target violation category and the violation confidence corresponding to the target violation category are determined.

2. A traffic violation judgment method based on a monocular camera as claimed in claim 1, characterized in that: The image points A′, B′, C′, and D′ in the first corrected image A′B′C′D′ correspond to the vertices A, B, C, and D of the imaging field of view ABCD. At the same time, the pixel midline T′M′ of A′B′C′D′ corresponds to the field of view midline TM of ABCD, and the end points T and M of the field of view midline TM are located on the wide side CD and the wide side AB, respectively, and the distance between the wide side CD and the projection point S is smaller than the distance between the wide side AB and the projection point S.

3. The method for determining traffic violations based on a monocular camera as claimed in claim 1, wherein: The basic assumption is that, for an image corrected for perspective distortion, the horizontal pixel ratio is the same as the horizontal length ratio of the real plane, and the vertical pixel ratio is the same as the vertical length ratio of the real plane.

4. A traffic violation judgment method based on a monocular camera as claimed in claim 2, characterized in that: Selecting marker points within the imaging field of view, and determining a first distance between each marker point and the projection point, and a second distance between two adjacent marker points, including: In the imaging field of view ABCD, three static markers are used as the first calibration point E, the second calibration point F and the third calibration point G respectively; Connect the projection point S with E, F, and G respectively to obtain the lengths e, f, and g of SE, SF, and SG respectively; Connect E, F, and G in sequence to obtain the lengths l1, l2, and l3 of EF, FG, and GE respectively; Among them, e, f and g are the first distances, and l1, l2 and l3 are the second distances.

5. A traffic violation judgment method based on a monocular camera as claimed in claim 4, characterized in that: According to the basic assumption, the camera's pose parameters are dynamically calibrated based on the pixel coordinates of the first calibrated image and the first and second distances, including: Draw a perpendicular line through the first calibration point E, the second calibration point F, and the third calibration point G to the extension line SM of the visual field midline TM of the imaging visual field ABCD, and obtain the first perpendicular foot E0, the second perpendicular foot F0, and the third perpendicular foot G0; Determine the true lengths of the first line segment SE0, the second line segment SF0, the third line segment GG0, and the fourth line segment E0F0 according to the first distance and the second distance respectively; Determine the pixel coordinates of each point in the first corrected image A′B′C′D′, and obtain the pixel lengths of the second pixel segment T′M′, the third pixel segment G′G′0, the fourth pixel segment E′0F0′, and the fifth pixel segment E′0T′, as well as the width of the wide side A′B′ of the image in pixels; The center line TM of the field of view, the first boundary line segment ST, and the broadside AB of the imaging field of view are used as the pose parameters of the camera; where ST represents the line segment between the projection point S and the endpoint M; According to the basic assumptions, the values of each posture parameter are determined respectively.

6. A traffic violation judgment method based on a monocular camera as claimed in claim 5, characterized in that: Determining the true lengths of the first line segment SE0, the second line segment SF0, the third line segment GG0, and the fourth line segment E0F0 according to the first distance and the second distance respectively includes: SE0=ecos(θ+η) SF0=fcosθ GG0=gsin(μ-θ) E0F0=ecos(θ+η)-fcosθ Where θ is the angle between SE and SM; η is the angle between SE and SF; μ is the angle between SF and SG.

7. A traffic violation judgment method based on a monocular camera as claimed in claim 6, characterized in that: Determining the pixel coordinates of each point in the first corrected image A′B′C′D′, respectively obtaining the pixel lengths of the second pixel segment T′M′, the third pixel segment G′G′0, the fourth pixel segment E′0F0′, and the fifth pixel segment E′0T′, and the width in pixels of the wide side A′B′ of the image, includes: Define the pixel coordinates of D′, A′, B′, and C′ as (0,0), (0,h), (w,h), and (w,0), respectively, where w and h are the width and height of the image, respectively. The coordinates of the image points E′, F′, and G′ corresponding to the first calibration reference point E, the second calibration reference point F, and the third calibration reference point G in A′B′C′D′ are defined as (i e ,i e )、(i f ,j f )、(i g ,j g ); Determine the pixel length of T′M′ as h and the pixel length of G′G′0 as i g -w / 2, the pixel length of E′0F′0 is j e -j f , the pixel length of E′0T′ is j e , the width of A′B′ is w pixels.

8. A traffic violation judgment method based on a monocular camera as claimed in claim 7, characterized in that: According to the basic assumptions, the values of the posture parameters are determined respectively, including: 1) Since the horizontal pixel ratio is the same as the horizontal length ratio of the real plane, we have: The true length of the visual field centerline TM is: Similarly, Then the true length of the fifth line segment E0T is: Based on the true length of the first line segment SE0, we can get: 2) Since the vertical pixel ratio is the same as the vertical length ratio of the real plane, we have: The actual width of the broadside AB of the imaging field of view is:

9. A traffic violation judgment method based on a monocular camera as claimed in claim 8, characterized in that: Identify the imaging target in the first corrected image, obtain the pixel coordinates of the imaging target, and then determine the absolute distance between the target to be measured and the monocular camera in combination with the pose parameters, including: Determine the projection point Q0 of the road contact point Q of the target to be measured on the visual field center line TM, and obtain the right triangle SQQ0 based on the projection points S, Q and Q0; In the first corrected image A′B′C′D′, the pixel coordinates of the imaging target Q′ are obtained as (i, j), and the pixel coordinates of Q′0 are obtained as (w / 2, j). The pixel length of the first pixel side Q′Q′0 and the pixel length of the sixth pixel line segment T′Q′0 are determined. According to the basic assumption, the real lengths of the first side QQ0 and the second side SQ0 are determined based on the real widths of the visual field centerline TM, the first line segment ST, and the broad side AB of the imaging visual field, respectively; Based on QQ0 and SQ0, determine the absolute distance SQ between the target to be measured and the monocular camera.

10. A traffic violation judgment system based on a monocular camera, characterized in that: The system includes: The determination module is used to install a monocular camera at the traffic intersection to be detected and determine the projection point and projection distance of the monocular camera. The imaging field of view determination module is used to obtain a rectangular area in the ground shooting field of view of the monocular camera as the imaging field of view, and at least one target to be measured is located in the imaging field of view. The acquisition module is used to perform perspective distortion correction on the image of the imaging field of view area captured by the camera to obtain a first corrected image, and set a basic hypothesis for the first corrected image. A distance calculation module determines the marker points within the imaging field of view, and determines a first distance between each marker point and the projection point, and a second distance between two adjacent marker points; A parameter calibration module dynamically calibrates the camera's pose parameters based on the pixel coordinates of the first calibrated image, the first distance, and the second distance according to a basic assumption; The absolute distance determination module identifies the imaging target in the first corrected image, obtains the pixel coordinates of the imaging target, and then determines the absolute distance between the target to be measured and the monocular camera in combination with the posture parameters; The traffic violation positioning module is used to determine the target violation category and the violation confidence corresponding to the target violation category based on the absolute distance between the target and the monocular camera.