Road surface pit depth calculation method and system based on monocular camera image

Through the road pit depth calculation method based on a monocular camera, the vertical coordinate of the point of extinguishing point and 3D information combined with the camera pitch angle and imaging plane distance, the problem of difficulty in obtaining pit depth information in the prior art is solved, and efficient and accurate depth calculation and position detection are achieved, reducing cost and complexity.

CN119941861AActive Publication Date: 2025-05-06SHANGHAI TONGLU CLOUD TRANSPORTATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510429255.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing pavement pit groove detection method based on monocular cameras has problems in feature extraction, depth information acquisition, equipment cost, calibration complexity, generalization capability, real-timeness and computing resource requirements, and it is especially difficult to accurately calculate the depth information of pit grooves.

Method used

A method for calculating the depth of the road surface pit groove based on a monocular camera image is proposed. By obtaining the vertical coordinate of the point of the image to be detected and the 3D information of the pit groove is calculated based on the pitch angle of the camera and the imaging plane distance. The method includes steps such as image acquisition, point-detection, 3D information extraction and depth calculation.

Benefits of technology

It realizes the rapid calculation of pit depth information while detecting pit slot locations, which reduces implementation costs and calculation costs, improves operating efficiency, simplifies the calibration process, and improves the anti-interference ability and the accuracy of depth calculations of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941861A_ABST
    Figure CN119941861A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of road engineering, municipal engineering and highway engineering, and relates to a pavement pit depth calculation method and system based on a monocular camera image, and the method comprises the steps: obtaining a to-be-detected image, and extracting the vanishing point ordinate of the to-be-detected image, pit 3D information and pit height; according to the vanishing point ordinate, the actual distance from the central point of the upper top surface of the pit to the camera and the pixel distance of the pit depth in the horizontal view are obtained by combining the pit 3D information and the pit height; and the actual depth of the pit slot is obtained by combining the actual distance and the pixel distance with the distance from the imaging plane of the monocular camera to the camera. While the position of the pit slot is intelligently detected, the depth information of the pit slot can be quickly calculated, and the method has the characteristics of low implementation cost, low calculation cost, high operation efficiency and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of road engineering, municipal engineering and highway engineering, and in particular to a method and system for calculating the depth of road potholes based on monocular camera images. Background Art

[0002] With the acceleration of urbanization, road maintenance and management are becoming more and more important. Potholes are one of the common road damage phenomena, which not only increases the wear and tear cost of vehicles, but also affects driving safety. Traditional pothole detection methods usually rely on manual inspections, which is not only inefficient but also easily affected by subjective factors, resulting in inaccurate detection results. In recent years, the development of computer vision technology has provided new solutions for automated detection. By using a monocular camera to collect road images and combining it with deep learning target detection technology, automatic recognition of potholes can be achieved.

[0003] Among the existing pit detection methods based on monocular cameras, there are mainly the following technical routes:

[0004] Methods based on traditional image processing: These methods usually use edge detection, morphological operations and other means to extract pit features. However, these methods are less effective in pit recognition under complex backgrounds and cannot directly obtain pit depth information:

[0005] Instability of feature extraction: Traditional image processing methods rely on manually designed features, such as edge detection and threshold segmentation. In actual applications, road conditions are complex and changeable, with shadows, reflections, and texture differences. These factors interfere with the accurate extraction of pothole edges, resulting in a significant decrease in the accuracy of feature extraction in complex scenarios, and prone to false detection and missed detection. For example, on a road with direct sunlight, the edge of a pothole may be blurred due to reflections, making it difficult for edge detection-based methods to accurately locate the pothole boundary.

[0006] Limitations of depth information acquisition: This type of method can only process at the two-dimensional image level and cannot directly obtain the depth information of the pothole. However, the depth of the pothole is crucial for road maintenance decisions. Potholes of different depths require different maintenance measures and priorities. Lack of depth information will make maintenance work less targeted and timely.

[0007] Methods based on binocular or multi-camera detection: Although binocular or multi-camera cameras can obtain depth information through parallax calculation, the equipment cost is high and requires a complex calibration process. In addition, the installation and maintenance of multi-camera systems in practical applications are also difficult:

[0008] High equipment cost: Binocular or multi-camera systems need to be equipped with multiple cameras, which not only increases the hardware procurement cost, but also requires the purchase of additional brackets for mounting the cameras and devices to ensure that the cameras work synchronously. The overall equipment cost is high, which is unaffordable for some road maintenance projects with limited budgets.

[0009] Complex calibration process: To ensure the accuracy of the relative position relationship between multiple cameras, complex calibration operations are required. The calibration process involves many parameter adjustments, and any slight deviation may cause large errors in the depth calculation results. In addition, if the equipment moves during use due to vehicle vibration or other reasons, it must be recalibrated, which undoubtedly increases the difficulty and cost of system maintenance.

[0010] Difficulty in installation and maintenance: The installation of a multi-camera system requires careful planning of the position and angle of each camera to ensure the consistency of its coverage and viewing angle and to achieve accurate depth information collection. However, in the actual installation environment, it is not easy to meet such requirements. Once a camera fails or its performance degrades, the entire system may need to be readjusted and optimized, which is complex and time-consuming.

[0011] Methods based on deep learning: In recent years, deep learning has made significant progress in the field of image recognition. By training a neural network model, efficient detection of potholes on the road surface can be achieved. However, most existing methods can only provide the location information of the potholes on a two-dimensional image, and it is difficult to directly calculate the depth data of the potholes:

[0012] Lack of depth information: Currently, most deep learning-based pothole detection methods focus on target detection tasks in two-dimensional images. Although they perform well in identifying pothole locations, they cannot provide pothole depth data. This makes it difficult to determine reasonable maintenance measures and priorities based on pothole depth during road maintenance, which may lead to delayed maintenance work, increased maintenance costs and driving risks. For example, for deeper potholes, if they are not discovered and repaired in time, they may cause greater damage to the vehicle and even cause traffic accidents.

[0013] In summary, the existing pit detection methods based on monocular cameras have problems in feature extraction, depth information acquisition, equipment cost, calibration complexity, generalization ability, real-time performance, and computing resource requirements. In particular, due to the timeliness requirements for maintenance of different pit depths, the depth information of the pit becomes indispensable during actual maintenance. Therefore, developing a method that can both efficiently and accurately detect the pit position and accurately calculate the pit depth has become a technical problem that needs to be solved urgently. Summary of the invention

[0014] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to propose a method and system for calculating the pothole depth of a road surface based on a monocular camera image, which can quickly calculate the pothole depth information while intelligently detecting the pothole position, and the present invention has the characteristics of low implementation cost, low calculation cost, and high operating efficiency.

[0015] To achieve the above object, the present invention provides the following solutions:

[0016] A method for calculating the depth of road potholes based on a monocular camera image, comprising:

[0017] Acquire an image to be detected, and extract the vertical coordinate of the vanishing point of the image to be detected, as well as the 3D information of the pit and the height of the pit;

[0018] Acquire, according to the vertical coordinate of the vanishing point, the actual distance from the center point of the top surface of the pit to the camera and the pixel distance of the pit depth in the horizontal view in combination with the pit 3D information and the pit height;

[0019] The actual depth of the pit is obtained by using the actual distance and the pixel distance in combination with the distance from the imaging plane of the monocular camera to the camera.

[0020] Optionally, extracting the vertical coordinate of the vanishing point of the image to be detected includes:

[0021] The image to be detected is input into an image vanishing point detection model to obtain the vertical coordinate of the vanishing point; the image vanishing point detection model is trained using a first training set, and the first training set includes: images marked with vanishing point information.

[0022] Optionally, extracting the pit 3D information and the pit height comprises:

[0023] The image to be detected is input into a pit 3D detection model to obtain the pit 3D information and the pit height; the pit 3D detection model is trained by using a second training set, and the second training set includes: image data containing pit 3D information; the Depth Block module in the pit 3D detection model is used to extract target depth feature information, and the target depth information representation is added to the output layer.

[0024] Optionally, obtaining the actual distance from the center point of the top surface of the pit to the camera includes:

[0025] Calculate the vertical coordinate of the vanishing point to obtain the current pitch angle, obtain the mapping matrix from the image to the overhead image according to the current pitch angle combined with the constant matrix of the image to the overhead image, and further use the vertical coordinate of the vanishing point to obtain the distance from the bottom edge of the image to the camera;

[0026] Based on the 3D information of the pit, the coordinates of the center point of the top surface of the pit in the image are obtained, and the coordinates of the center of the bottom surface of the pit in the image are obtained through the height of the pit;

[0027] Calculating the mapping matrix from the image to the top-view image and the coordinates of the center point of the upper top surface of the pit in the image, obtaining the coordinates of the center point of the upper top surface of the pit in the image in the top-view image, calculating the coordinates of the center point of the upper top surface of the pit in the image in the top-view image, and obtaining the actual distance from the center point of the upper top surface of the pit to the lower edge of the original image;

[0028] The actual distance from the center point of the top surface of the pit to the lower edge of the original image is added to the distance from the lower edge of the image to the camera to obtain the actual distance from the center point of the top surface of the pit to the camera.

[0029] Optionally, acquiring a constant matrix from the image to the overhead view image includes:

[0030] Calibrate the monocular camera, determine the camera pitch angle, and obtain a rotation matrix based on the camera pitch angle; the rotation matrix includes: a first rotation matrix and a second rotation matrix:

[0031]

[0032] in, is the rotation matrix of the y-axis rotation angle relative to the world coordinate system, is the pitch angle of the camera;

[0033] Acquire an original image through a calibrated monocular camera, perform calibration calculation on the original image, and acquire a mapping matrix from the original image to the overhead view image;

[0034] The first rotation matrix is ​​combined with the mapping matrix from the original image to the overhead image to obtain the constant matrix from the image to the overhead image:

[0035]

[0036] in, is the mapping matrix from image to top-view image, is the first rotation moment, is the constant matrix from image to overhead view image.

[0037] Optionally, obtaining the actual distance from the center point of the top surface of the pit to the lower edge of the original image includes:

[0038]

[0039] in, is the height of the top view, is the ordinate of point P in the top view, is the actual distance from point P to the lower edge of the top view.

[0040] Optionally, obtaining the pixel distance of the pit depth in the horizontal view includes:

[0041] acquiring a mapping matrix from the image to the horizontal view according to the current pitch angle combined with a constant matrix from the image to the horizontal view, and acquiring coordinates of the center point of the upper top surface and the center coordinates of the lower bottom surface in the horizontal view based on the coordinates of the center point of the upper top surface of the pit in the image and the center coordinates of the lower bottom surface of the pit in the image combined with the mapping matrix from the image to the horizontal view;

[0042] Based on the coordinates of the center point of the upper top surface in the horizontal view and the coordinates of the center point of the lower bottom surface in the horizontal view, the pixel distance of the pit depth in the horizontal view is acquired.

[0043] Optionally, acquiring a constant matrix of the image to a horizontal view includes:

[0044] Perform a checkerboard calibration calculation on the original image to obtain a mapping matrix from the original image to the horizontal view, and use the second rotation matrix in combination with the mapping matrix from the original image to the horizontal view to obtain a constant matrix from the image to the horizontal view:

[0045]

[0046] in, is the mapping matrix from image to horizontal view, is the second rotation matrix, A constant matrix for the image to horizontal view.

[0047] Optionally, obtaining the actual depth of the pit comprises:

[0048]

[0049] in, is the actual depth of the pit, is the pixel distance of the ground pit depth in the horizontal view, is the distance from the imaging plane to the monocular camera position, is the position from the pit to the camera.

[0050] To achieve the above object, the present invention also provides a road pothole depth calculation system based on a monocular camera image, comprising:

[0051] An image acquisition module, used for acquiring an image to be detected;

[0052] The cloud storage calculation module extracts the vertical coordinate of the vanishing point of the image to be detected, the 3D information of the pit, and the height of the pit, and obtains the actual distance from the center point of the top surface of the pit to the camera and the pixel distance of the pit depth in the horizontal view according to the vertical coordinate of the vanishing point combined with the 3D information of the pit and the height of the pit, and obtains the actual depth of the pit by combining the actual distance and the pixel distance with the distance from the imaging plane of the monocular camera to the camera.

[0053] The beneficial effects of the present invention are:

[0054] The present invention can be used in conjunction with the intelligent road disease detection system to output the pothole height information while detecting the pothole position. The system can obtain the pothole depth through simple calibration calculation, thus completing the gap of the pothole detection based on the monocular camera image, which lacks the depth information.

[0055] In terms of execution efficiency, compared with the inefficiency and high cost of traditional pit depth measurement, the present invention can basically achieve no manual operation, reduce measurement costs and improve efficiency.

[0056] Compared with the calculation of pothole depth based on binocular or multi-objective methods, the present invention adopts a monocular camera as the acquisition device, and realizes accurate measurement of the pothole depth through camera calibration and pixel distance conversion. It not only reduces the hardware cost, but also simplifies the calibration process, making the system easier to install and maintain.

[0057] Compared with simply relying on deep learning target detection to realize road pothole detection, the present invention can not only intelligently detect the location of the pothole, but also obtain the depth information of the pothole, providing essential data support for the intelligent maintenance of the pothole.

[0058] The method of detecting image vanishing points and correcting the image mapping matrix in real time introduced in the present invention greatly improves the anti-interference capability of the overall system and the accuracy of pit depth calculation.

[0059] The deep learning 3D pit detection algorithm based on monocular camera images used in the present invention optimizes the traditional deep learning target detection algorithm, introduces the depth information feature map, optimizes the network structure, and enriches the data samples, so that the deep learning 3D pit detection algorithm used in the present invention can output both the position information of the pit and the depth information of the pit, and the algorithm has strong generalization ability and high accuracy.

[0060] In summary, the present invention can quickly calculate the pit depth information while intelligently detecting the pit position, and the present invention has the characteristics of low implementation cost, low calculation cost, high operation efficiency, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0062] Figure 1 This is a schematic diagram of the installation of an edge acquisition device according to an embodiment of the present invention;

[0063] Figure 2 A schematic diagram of a camera pitch angle according to an embodiment of the present invention;

[0064] Figure 3 This is a schematic diagram of imaging with a monocular camera according to an embodiment of the present invention;

[0065] Figure 4 The calibration original image for obtaining the top view image mapping matrix according to the embodiment of the present invention;

[0066] Figure 5 A schematic diagram of 3D marking / detection of road potholes according to an embodiment of the present invention;

[0067] Figure 6 A schematic diagram of a deep learning 3D object detection algorithm flow in an embodiment of the present invention;

[0068] Figure 7 This is a flow chart of the equipment installation and calibration phase of an embodiment of the present invention;

[0069] Figure 8 The present invention is a flowchart of a method for calculating the depth of road potholes based on monocular camera images according to an embodiment of the present invention. DETAILED DESCRIPTION

[0070] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0071] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0072] Explanation of the idea of ​​dynamically correcting the camera's original image to the top-view image mapping matrix and the original image to the horizontal view mapping matrix: It is known that the mapping matrix of the original image collected by the camera is converted to the top-view image and the camera's pitch angle (Camera coordinate system and world coordinate system The angle of the y-axis), roll angle (Camera coordinate system and world coordinate system The angle of the x-axis), heading angle (Camera coordinate system and world coordinate system For example, when the roll angle When the pitch angle changes, the image captured by the camera will appear to tilt up and down in the vertical direction. When the heading angle changes, the relative position of the object in the image in the vertical direction will change. When the heading angle changes, it determines the left and right rotation of the camera shooting direction. When taking a panoramic photo, by changing the heading angle You can get scenes in different directions.

[0073] The transformation matrix formula from the camera original image to the top view image is as follows:

[0074]

[0075]

[0076] in, is the mapping matrix from the camera original image to the top view image, is the rotation matrix of the camera, is the translation matrix, Camera intrinsic matrix; Camera intrinsic matrix The general form is as follows:

[0077]

[0078] in, and They are images The focal lengths in the a-axis and y-axis directions, and It is the coordinate value of the principal point in the image coordinate system (the pixel coordinate corresponding to the intersection of the optical axis and the image plane).

[0079] The general form of the translation vector T is as follows: ;

[0080] The rotation matrix R is a 3x3 matrix, which can be composed of the rotation angles of the camera coordinate system relative to the three coordinate axes of the world coordinate system. Multiply them together to get, as shown in formula (2).

[0081] is the rotation matrix of the camera coordinate system relative to the three coordinate axes of the world coordinate system. Let the camera’s pitch angle, roll angle, and heading angle be hour:

[0082]

[0083]

[0084]

[0085] It should be pointed out that, as far as the present invention is concerned, the purpose of obtaining the top view image involved is only to accurately calculate the actual ground distance between two ground points in the top view image. In view of this, the process has no direct relationship with the position selection of the origin coordinates of the XY plane (ground) of the world coordinate system. Therefore, the top view mapping matrix M of the present invention can ignore the heading angle (i.e., the angle between the camera coordinate system and the z-axis in the world coordinate system) when constructing it. When the roll angle and installation position of the camera remain constant, the roll angle matrix, translation matrix, and intrinsic parameter matrix of the camera coordinate system are also fixed accordingly. At this time, the mapping matrix from the original image to the top view image can be regarded as a function related only to the pitch angle, as shown below:

[0086]

[0087] Where A is a 3x3 constant matrix.

[0088] Similarly, because the camera is affected by the pitch angle and roll angle when shooting the original image, the image obtained is not a horizontal perspective. The present invention can transform the coordinate system of the original image into an ideal coordinate system equivalent to the camera being in a horizontal state through a series of coordinate transformation operations. In this horizontal view coordinate system, there is a significant feature, that is, the same straight line perpendicular to the ground has the same horizontal coordinate in the horizontal view. Figure 2 As shown, during the camera installation phase, when the camera's pitch angle is the pitch angle value When the camera coordinate system rotates from the initial state to the top view, the rotation angle along the y-axis is ( ), the rotation matrix at this time is ; When rotating to a horizontal viewing angle, the rotation angle along the y-axis is , the rotation matrix at this time is Based on the above principles and related formulas, it can be deduced that when the camera pitch angle is When , the mapping matrix from the original image to the horizontal view can be expressed as follows:

[0089] (8)

[0090] When the camera installation position and the camera roll angle remain unchanged, formula (7) and formula (8) are only related to the change of the camera's pitch angle. The mapping matrix from the camera's original image to the top view image and the mapping matrix from the camera's original image to the horizontal view can be updated in real time by monitoring the change of the camera's pitch angle.

[0091] An explanation of the relationship between the vanishing point of the original image and the pitch angle of the camera, and the relationship between the vanishing point of the original image and the distance from the bottom edge of the original image to the camera.

[0092] It is known that the value of the vanishing point ordinate of the original image has a linear relationship with the pitch angle of the camera and also has a linear relationship with the distance from the lower edge of the original image to the camera. Therefore, after installing the camera, fixing the installation position and the camera roll angle, the camera pitch angle can be manually adjusted to collect the original images of the camera at different pitch angles, and the actual distance D1 from the lower edge of the original image to the camera at each pitch angle is measured. Then, the vanishing point of the original image is detected using a deep learning image vanishing point detection algorithm to obtain the vanishing point ordinate of each original image; and then a set of data Data1 consisting of the camera pitch angle and the vanishing point ordinate of the original image and a set of data Data2 consisting of the vanishing point ordinate of the original image and the distance from the lower edge of the original image to the camera is obtained. The linear formula from the vanishing point ordinate of the original image to the camera pitch angle is obtained by linear fitting of Data1, as shown in Formula 9. The linear formula from the vanishing point ordinate of the original image to the distance from the lower edge of the original image to the camera is obtained by linear fitting of Data2, as shown in Formula 10.

[0093]

[0094]

[0095] Wherein, x is the vertical coordinate value of the vanishing point of the original image.

[0096] During real-time acquisition, the vanishing point of the original image is detected by the deep learning image vanishing point detection algorithm, and the real-time pitch angle of the camera is calculated according to formula (9), and the new mapping matrix can be calculated by formula (7). The distance from the bottom edge of the original image to the camera can also be calculated by formula (10).

[0097] Explanation of the relationship between the height of an object in a horizontal view and the pixel distance in a horizontal view:

[0098] The imaging principle of a monocular camera can be compared to pinhole imaging, such as Figure 3 As shown, the ratio of the actual height RH of the object to the actual distance Da from the object to the camera is equal to the ratio of the pixel distance of the object in the imaging to the distance Db from the imaging plane to the camera position. The formula is as follows:

[0099]

[0100]

[0101]

[0102] Where Db is the distance from the imaging plane to the monocular camera position, PH is the pixel distance of the object in the horizontal view, RH is the actual height of the object, Da is the actual distance from the object to the camera, Dc is the actual distance from the object to the lower edge of the original image, and Dl is the actual distance from the lower edge of the original image to the camera.

[0103] In the calibration stage, the actual height RH of the object and the horizontal distance Da from the object to the camera can be obtained through on-site measurement. The pixel distance PH of the object in the horizontal view can be calculated by the ground point pixel coordinates and vertex pixel coordinates of the object in the horizontal view. The unknown parameter Db is the parameter that this method needs to obtain through calibration calculation. Therefore, during calibration, Db is calculated by collecting the values ​​of Da, PH, and RH. Usually, multiple groups of RH, Da, and PH values ​​are obtained, and multiple Db are calculated. The average value of multiple Db is taken as the calculation parameter of this method, and the calculation formula of the object height in the horizontal view can be obtained:

[0104]

[0105] From formula (14), it can be seen that when the pixel distance PH of the ground pit depth in the horizontal view, the position Da of the pit to the camera, and the distance Db from the imaging plane to the monocular camera position calculated during calibration are known, the actual depth RH of the pit can be obtained.

[0106] Introduction to deep learning image vanishing point detection algorithm:

[0107] The deep learning image vanishing point detection algorithm is an important technical means used in the present invention to obtain key information such as the camera pitch angle and the distance from the bottom edge of the original image to the camera. The algorithm is based on the principle of deep learning and is trained on a large number of images with annotated vanishing point information, so that the model can automatically identify the location of the vanishing point in the image.

[0108] During the training process, the algorithm learns the characteristic patterns of vanishing points in different scenes and camera postures. For example, for images in road scenes, the algorithm can identify the points where the extension direction of the road, the edges of buildings, etc. converge in the image, i.e., the vanishing points. The positions of these vanishing points are closely related to the posture of the camera (especially the pitch angle), and also have a certain geometric relationship with the distance from the bottom edge of the image to the camera.

[0109] The algorithm is based on the convolutional neural network (CNN) architecture, which takes advantage of its powerful feature extraction capabilities. The input image is processed step by step through a combination of multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layer is used to extract local features of the image, the pooling layer reduces the data dimension, and the fully connected layer maps the extracted features to the predicted output of the vanishing point position.

[0110] In order to improve the accuracy and generalization ability of the algorithm, the training data covers images under various road conditions, weather conditions and time. This enables the algorithm to adapt to image changes in different environments and accurately detect the location of the vanishing point, thereby providing a reliable basis for the subsequent calculation of the camera pitch angle and the distance from the bottom edge of the image to the camera.

[0111] Deep learning 3D pothole detection algorithm: The deep learning 3D pothole detection algorithm is one of the core algorithms for calculating the pothole depth in the present invention, which aims to extract the 3D information of the pothole from the 2D image. Figure 5 As shown, it includes the quadrilateral frame information of the top surface of the pit and the pixel height of the pit depth in the image.

[0112] The traditional deep learning 2D target detection algorithm of monocular camera images mainly relies on deep learning neural networks to extract target features from the input RGB images. However, when directly using the 2D target detection algorithm to learn the feature expression of the target from the original image, it is difficult to learn the depth information features of the target. In order to improve the accuracy of the prediction of the pit depth information, the deep learning pit 3D detection algorithm used in the present invention (such as Figure 6 Compared with traditional deep learning algorithms, the following optimizations are available:

[0113] 1. Increase input information: By using a mature depth estimation algorithm to extract the depth information of the entire image, it is used as an independent feature map and concatenated with the original RGB image into a 4D input. This provides the deep learning network with more direct and accurate full-image depth information, thereby reducing the difficulty of learning the target depth information.

[0114] 2. Optimize network structure: Compared with the traditional 2D target detection algorithm, the deep learning 3D pit detection algorithm used in the present invention adds a network structure (Depth Block) for fitting target depth information, making it easier to extract target depth feature information, thereby improving the learning ability of target depth information.

[0115] 3. Enhanced output layer information: The depth information representation of the target is added to the output layer of the algorithm, so that the deep learning 3D pit detection algorithm outputs the depth information of the target while outputting the target polygon position information, thereby achieving the goal of 3D pit detection.

[0116] When labeling training data, Figure 5As shown in the figure, the deep learning 3D pit detection algorithm used in the present invention uses quadrilaterals to mark the top surface of the pit, which can more accurately describe the shape and position information of the pit in the image. At the same time, the pixel height of the pit depth in the image is clearly marked, so that the algorithm can learn the relationship between the pit depth and the image features, thereby accurately outputting the 3D information of the pit during the detection process.

[0117] During training, the present invention uses a large amount of manually annotated image data containing 3D information of potholes. These data are collected from actual scenes with different road conditions (such as highways, urban roads, rural roads, etc.), different weather (sunny, rainy, cloudy, etc.) and different times (day, night, etc.) to ensure that the algorithm has good generalization ability.

[0118] During inference, the inference process of the deep learning pit 3D detection algorithm used in the present invention is similar to the inference process of the traditional deep learning target detection network. The feature map of the image and depth information is first extracted through the convolutional neural network. Then, the region proposal network (RPN) is used to generate candidate regions that may contain pits. These candidate regions are further classified and regressed to determine the exact position of the pit (quadrilateral box information) and the pixel height of the pit depth in the image.

[0119] Based on the above information, this embodiment discloses a method for calculating the depth of a road pothole based on a monocular camera image, including: obtaining an image to be detected, extracting the vertical coordinate of the vanishing point of the image to be detected, as well as the 3D information of the pothole and the height of the pothole; obtaining the actual distance from the center point of the top surface of the pothole to the camera and the pixel distance of the pothole depth in the horizontal view according to the vertical coordinate of the vanishing point combined with the 3D information of the pothole and the height of the pothole; and obtaining the actual depth of the pothole by combining the actual distance and the pixel distance with the distance from the imaging plane of the monocular camera to the camera.

[0120] Furthermore, obtaining the actual distance from the center point of the upper top surface of the pit to the camera includes: calculating the vertical coordinate of the vanishing point, obtaining the current pitch angle, obtaining the mapping matrix from the image to the overhead image according to the current pitch angle combined with the constant matrix from the image to the overhead image, and further using the vertical coordinate of the vanishing point to obtain the distance from the lower edge of the image to the camera; based on the 3D information of the pit, obtaining the coordinates of the center point of the upper top surface of the pit in the image, and obtaining the center coordinates of the lower bottom surface of the pit in the image through the pit height; calculating the mapping matrix from the image to the overhead image and the coordinates of the center point of the upper top surface of the pit in the image, obtaining the coordinates of the center point of the upper top surface of the pit in the image in the overhead image, and calculating the coordinates of the center point of the upper top surface of the pit in the image in the overhead image, and obtaining the actual distance from the center point of the upper top surface of the pit to the lower edge of the original image; adding the actual distance from the center point of the upper top surface of the pit to the lower edge of the original image to the distance from the lower edge of the image to the camera, to obtain the actual distance from the center point of the upper top surface of the pit to the camera.

[0121] Obtaining the pixel distance of the pit depth in the horizontal view includes: obtaining a mapping matrix from the image to the horizontal view according to the current pitch angle combined with a constant matrix from the image to the horizontal view, obtaining the coordinates of the center point of the upper top surface and the center coordinates of the lower bottom surface in the horizontal view based on the coordinates of the center point of the pit top surface in the image and the coordinates of the center of the lower bottom surface in the image combined with the mapping matrix from the image to the horizontal view; obtaining the pixel distance of the pit depth in the horizontal view based on the coordinates of the center point of the upper top surface in the horizontal view and the coordinates of the center coordinates of the lower bottom surface in the horizontal view.

[0122] Specifically: installation and calibration of equipment, such as Figure 7 shown.

[0123] S101: Install the equipment, install the monocular camera, edge control machine, and positioning module on the collection vehicle.

[0124] S102: Fitting to obtain the relationship between the vertical coordinate of the vanishing point of the original image and the camera pitch angle , and the relationship between the vertical coordinate of the vanishing point of the original image and the distance from the bottom edge of the original image to the camera The specific method is as follows.

[0125] After fixing the camera installation position and camera roll angle, adjust the pitch angle of the monocular camera, collect multiple sets of original images at different pitch angles, and measure the distance from the bottom edge of each original image to the camera. Use the deep learning image vanishing point detection algorithm to detect the vertical coordinate of the vanishing point of each original image. Get a set of data Data1 consisting of the camera pitch angle and the image vanishing point vertical coordinate, and fit the relationship between the original image vanishing point vertical coordinate and the pitch angle , get a set of data Data2 consisting of the vertical coordinates of the vanishing point of the original image and the distance from the bottom edge of the original image to the camera, and get the relationship between the vertical coordinates of the vanishing point of the original image and the distance from the bottom edge of the original image to the camera by fitting .

[0126] S103: After adjusting the appropriate pitch angle, fix the camera pitch angle. Measure the camera pitch angle at this time , use formula (5) to calculate the rotation matrix at this time and .

[0127] S104: Obtain the mapping matrix M from the original image to the overhead view image through calibration calculation.

[0128] The original image to overhead view image mapping matrix M can be obtained by using the coordinate information of known points in two pixel spaces (the original image pixel space and the overhead view image pixel space); the specific method is as follows.

[0129] By laying a chessboard of known size on the road surface, the lower edge of the chessboard coincides with the lower edge of the original image, and the original image of the monocular camera is collected at this time. Figure 4As shown, the pixel coordinates of the four corner points of the chessboard in the original image are [(xtl, ytl), (xtr, ytr), (xbl, RawImageH), (xbr, RawImageH)], where (xtl, ytl) represents the coordinates of the upper left corner of the chessboard in the original image; (xtr, ytr) represents the coordinates of the upper right corner of the chessboard in the original image; (xbl, RawImageH) represents the coordinates of the lower left corner of the chessboard in the original image, and RawImageH is the height of the original image (the lower edge of the chessboard is flush with the lower edge of the imaging plane); (xbr, RawImageH) represents the coordinates of the lower left corner of the chessboard in the original image. The coordinates of the lower right corner of the chessboard; since the chessboard does not deform in the top view image, the length and width of the chessboard are set to 2 meters, and the distance between each two pixels in the top view image represents 1 mm, thus obtaining the coordinates of the four corners of the chessboard in the top view image [(Xtl, IMGH-2000), (Xtl+2000, IMGH-2000), (Xtl, IMGH), (Xtl+2000, IMGH)], Xtl represents the pixel coordinates of the left border of the chessboard in the top view image, which can be set as needed. Generally, the chessboard is placed in the center of the top view image, that is, Xtl=IMGW / 2-1000. Among them, IMGW is the width of the top view image; IMGH is the height of the top view image (generally set to the effective acquisition distance of each image). Use the getPerspectiveTransform method in opencv, input the coordinates of the corner points of the chessboard in the original image and the coordinates of the corner points of the chessboard in the top view image, and you can quickly calculate the mapping matrix M. Since the mapping matrix M assumes that the distance between every two pixels in the top view image represents 1 mm, the height of the top view image is IMGH, and the bottom edge of the top view coincides with the bottom edge of the original image. Therefore, it can be seen that the actual distance Dc from the point P (x, y) in the top view image to the bottom edge of the top view satisfies the following formula:

[0130]

[0131] Among them, IMGH is the height of the top view, y is the vertical coordinate of point P in the top view, and Dc is the actual distance from point P to the lower edge of the top view.

[0132] S105: Use the rotation matrix calculated by S103 The mapping matrix M obtained in S104 is calculated according to formula (7) to obtain a constant matrix A from the original image to the overhead view image.

[0133] S106: Obtain the mapping matrix from the original image to the horizontal view by chessboard calibration calculation .

[0134] The specific method is as follows: a chessboard of known size is placed vertically on the road surface, and the original image at this time is collected; the pixel coordinates of the four corner points of the chessboard placed vertically on the ground in the original image are [(xtl1, ytl1), (xtr1, ytr1), (xbl1, ybl1), (xbr11, ybr11)], where (xtl1, ytl1) represents the coordinates of the upper left corner of the chessboard placed vertically on the ground in the original image; (xtr1, ytr1) represents the coordinates of the upper right corner of the chessboard in the original image; (xbl1, ybl1) represents the coordinates of the lower left corner of the chessboard in the original image; (xbr1, ybr1) represents the coordinates of the lower right corner of the chessboard in the original image; since the chessboard does not deform in the horizontal view , let the coordinates of the upper left corner of the chessboard in the horizontal view be consistent with the coordinates of the upper left corner on the original image, let the width of the chessboard in the horizontal view be consistent with the width on the original image, and you can get the coordinates of the four corner points of the chessboard in the horizontal view [(xtl1, ytl1), (xtr1, ytl1), (xtl1, ybl1), (xtr1, ybl1)], where (xtl1, ytl1) represents the coordinates of the upper left corner of the chessboard in the horizontal view; (xtr1, ytl1) represents the coordinates of the upper right corner of the chessboard in the horizontal view; (xtl1, ybl1) represents the coordinates of the lower left corner of the chessboard in the horizontal view; (xtr1, ybl1) represents the coordinates of the lower right corner of the chessboard in the horizontal view. Use the getPerspectiveTransform method in opencv, input the coordinate information of the chessboard perpendicular to the ground in the original image and the coordinate information in the horizontal view, and you can calculate the mapping matrix from the original image to the horizontal view. .

[0135] S107: Use the rotation matrix obtained in S103 , and the mapping matrix obtained by S106 According to formula (8), the constant matrix from the original image to the horizontal view is calculated .

[0136] S108: Calculate the distance Db from the imaging plane of the monocular camera to the camera through calibration.

[0137] The specific steps are as follows: place a calibration pole of known height vertically at different horizontal distances from the camera to collect the original image; obtain multiple sets of grounding point coordinates P1 (x1, y1) and vertex coordinates P2 (x2, y2) of the calibration pole in the original image; use the mapping matrix Mt from the original image to the horizontal view to calculate the coordinates PP1 and PP2 of P1 and P2 in the horizontal view; according to PP1 and PP2, the pixel distance PH of the calibration pole in the horizontal view can be calculated; use a ruler to measure the actual distance Da from the grounding point of the calibration pole to the camera; given the actual height RH of the calibration pole, the required parameter Db can be calculated according to formula (12); multiple sets of Db can be calculated and Db can be obtained by taking the average value to reduce the measurement error.

[0138] The second stage is the collection and operation stage, such as Figure 8 shown.

[0139] S201: Run the collection vehicle to collect original road images in real time, match the positioning information and upload it to the cloud server.

[0140] S201: After receiving the information, the cloud service uses a deep learning image vanishing point detection algorithm to detect the vertical coordinate ynew of the vanishing point of the original image.

[0141] S203: Calculate the camera pitch angle when the original image is captured using ynew and formula (9) .

[0142] S204: Use S203 to get the camera pitch angle and S105 to obtain the constant matrix A and the formula At this time, the mapping matrix Mnew from the original image to the overhead image.

[0143] S205: Use S203 to get the camera pitch angle The constant matrix is ​​calculated by S106 And formula (8) calculates the mapping matrix from the original image to the horizontal view at this time .

[0144] S206: Calculate the distance Dlnew from the lower edge of the original image to the camera when the original image is captured using ynew obtained in S201 and formula (10).

[0145] S207: Detect the 3D information of the pit in the original image (including the information of the quadrilateral box [x1, y1, x2, y2, x3, y3, x4, y5] and the pixel height h) by using the deep learning pit 3D detection algorithm.

[0146] S208: According to the position information box[x1,y1,x2,y2,x3,y3,x4,y5] of the top surface of the pit in the original image detected by S207, the center point coordinates Ptc (xc,yc) of the top surface of the pit in the original image are calculated, and according to the pit pixel height h output by the deep learning pit 3D detection algorithm in S207, the center coordinates Pbc (xc,yc-h) of the lower bottom surface of the pit in the original image are calculated.

[0147] S209: According to S204, the mapping matrix Mnew from the original image to the top view image and the coordinates Ptc (xc, yc) of the center point of the top surface of the pit in the original image are calculated, and the cv2.perspectiveTransform method in opencv is used to calculate the coordinates PFtc (xfc, yfc) of the center point of the top surface of the pit in the original image in the top view image. Using PFtc (xfc, yfc) and formula (15), the actual distance Dcnew from the center point of the top surface of the pit to the lower edge of the original image is calculated; the distance Dlnew from the lower edge of the original image to the camera obtained by S206 is added to Dcnew to obtain the actual distance Danew from the center point of the top surface of the pit to the camera.

[0148] S210: Based on the top center point coordinates Ptc (xc, yc) and bottom center coordinates Pbc (xc, yc-h) of the pit in the original image obtained in S208 and the mapping matrix from the original image to the horizontal view obtained in S205 The cv2.perspectiveTransform method in opencv is used to calculate the coordinates PTtc (xttc, yttc) and PTbc (xtbc, ytbc) of Ptc and Pbc in the horizontal view of the original image. The pixel distance PHnew of the pit depth in the horizontal view is calculated according to the coordinates PTtc (xtt, ytt) and PTbc (xtb, ytb) in the horizontal view.

[0149] S211: Based on the distance Danew from the center point of the top surface of the pit to the camera obtained in S209 and the pixel distance PHnew of the pit depth in the horizontal view obtained in S210, and Db obtained in the calibration stage S108, the actual depth RHnew of the pit is calculated using formula (14).

[0150] The present embodiment also provides a road pothole depth calculation system based on a monocular camera image, including: an image acquisition module, used to obtain an image to be detected; a cloud storage calculation module, which extracts the vertical coordinate of the vanishing point of the image to be detected and the 3D information of the pothole and the height of the pothole, and obtains the actual distance from the center point of the top surface of the pothole to the camera and the pixel distance of the pothole depth in the horizontal view according to the vertical coordinate of the vanishing point combined with the 3D information of the pothole and the height of the pothole, and obtains the actual depth of the pothole by combining the actual distance and the pixel distance with the distance from the imaging plane of the monocular camera to the camera.

[0151] Specifically: The present invention discloses a road pothole depth calculation system based on monocular camera images, the system architecture mainly consists of two parts: a vehicle-mounted acquisition device and a cloud storage and calculation module. Among them, the specific structure of the vehicle-mounted acquisition device is as follows: Figure 1 As shown in the figure, it includes key components such as a monocular camera for image acquisition, a GPS / Beidou positioning module with precise positioning function, and an edge central industrial computer. This vehicle-mounted acquisition equipment is cleverly installed on the acquisition vehicle. Its core task is to collect the original image of the road surface, quickly match it with the corresponding precise positioning information, and upload this data to the cloud module in real time.

[0152] The cloud storage computing module mainly includes several important sub-modules: the first is a storage module specifically used to store the original images uploaded by edge devices to ensure that massive image data can be properly preserved; the second is a database module responsible for storing original image information, such as acquisition time, positioning information and other structured data, to provide a data basis for subsequent analysis; the third is the image vanishing point detection module, which can detect the vanishing point position of the original image in real time and provide key parameters for subsequent calculations; the fourth is the pit position and depth information detection module, which uses a deep learning 3D target detection algorithm to detect the pit position information and depth information in the image; the fifth is the pit depth calculation module, which accurately calculates the pit depth based on the data collected and analyzed in the early stage.

[0153] The specific workflow of the entire system is mainly divided into two key stages: the first is the equipment installation and calibration stage, during which the on-board acquisition equipment must be carefully installed and accurately calibrated; the second is the real-time acquisition operation stage, during which the on-board acquisition equipment and the cloud storage and computing module work closely together. The former continuously collects and uploads data, while the latter quickly processes and analyzes it, ultimately achieving accurate calculation of the depth of potholes on the road.

[0154] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for calculating the depth of road potholes based on monocular camera images, characterized in that: include: Acquire an image to be detected, and extract the vertical coordinate of the vanishing point of the image to be detected, as well as the 3D information of the pit and the height of the pit; Acquire, according to the vertical coordinate of the vanishing point, the actual distance from the center point of the top surface of the pit to the camera and the pixel distance of the pit depth in the horizontal view in combination with the pit 3D information and the pit height; The actual depth of the pit is obtained by using the actual distance and the pixel distance in combination with the distance from the imaging plane of the monocular camera to the camera.

2. The method for calculating road pothole depth based on monocular camera images according to claim 1, characterized in that: Extracting the vertical coordinate of the vanishing point of the image to be detected includes: The image to be detected is input into an image vanishing point detection model to obtain the vertical coordinate of the vanishing point; the image vanishing point detection model is trained using a first training set, and the first training set includes: images marked with vanishing point information.

3. The method for calculating road pothole depth based on monocular camera images according to claim 1, characterized in that: Extracting the pit 3D information and the pit height includes: The image to be detected is input into a pit 3D detection model to obtain the pit 3D information and the pit height; the pit 3D detection model is trained by using a second training set, and the second training set includes: image data containing pit 3D information; the Depth Block module in the pit 3D detection model is used to extract target depth feature information, and the target depth information representation is added to the output layer.

4. The method for calculating road pothole depth based on monocular camera images according to claim 1, characterized in that: Obtaining the actual distance from the center point of the top surface of the pit to the camera includes: Calculate the vertical coordinate of the vanishing point to obtain the current pitch angle, obtain the mapping matrix from the image to the overhead image according to the current pitch angle combined with the constant matrix of the image to the overhead image, and further use the vertical coordinate of the vanishing point to obtain the distance from the bottom edge of the image to the camera; Based on the 3D information of the pit, the coordinates of the center point of the top surface of the pit in the image are obtained, and the coordinates of the center of the bottom surface of the pit in the image are obtained through the height of the pit; Calculating the mapping matrix from the image to the top-view image and the coordinates of the center point of the upper top surface of the pit in the image, obtaining the coordinates of the center point of the upper top surface of the pit in the image in the top-view image, calculating the coordinates of the center point of the upper top surface of the pit in the image in the top-view image, and obtaining the actual distance from the center point of the upper top surface of the pit to the lower edge of the original image; The actual distance from the center point of the top surface of the pit to the lower edge of the original image is added to the distance from the lower edge of the image to the camera to obtain the actual distance from the center point of the top surface of the pit to the camera.

5. The method for calculating road pothole depth based on monocular camera images according to claim 4, characterized in that: The constant matrix for obtaining the image to the top view image includes: Calibrate the monocular camera, determine the camera pitch angle, and obtain a rotation matrix based on the camera pitch angle; the rotation matrix includes: a first rotation matrix and a second rotation matrix: in, is the rotation matrix of the y-axis rotation angle relative to the world coordinate system, is the pitch angle of the camera; Acquire an original image through a calibrated monocular camera, perform calibration calculation on the original image, and acquire a mapping matrix from the original image to the overhead view image; The first rotation matrix is ​​combined with the mapping matrix from the original image to the overhead image to obtain the constant matrix from the image to the overhead image: in, is the mapping matrix from image to top-view image, is the first rotation moment, is the constant matrix from image to overhead view image.

6. The method for calculating road pothole depth based on monocular camera images according to claim 4, characterized in that: Obtaining the actual distance from the center point of the top surface of the pit to the lower edge of the original image includes: in, is the height of the top view, is the ordinate of point P in the top view, is the actual distance from point P to the lower edge of the top view.

7. The method for calculating road pothole depth based on monocular camera images according to claim 4, characterized in that: Obtaining the pixel distance of the pit depth in the horizontal view includes: acquiring a mapping matrix from the image to the horizontal view according to the current pitch angle combined with a constant matrix from the image to the horizontal view, and acquiring coordinates of the center point of the upper top surface and the center coordinates of the lower bottom surface in the horizontal view based on the coordinates of the center point of the upper top surface of the pit in the image and the center coordinates of the lower bottom surface of the pit in the image combined with the mapping matrix from the image to the horizontal view; Based on the coordinates of the center point of the upper top surface in the horizontal view and the coordinates of the center point of the lower bottom surface in the horizontal view, the pixel distance of the pit depth in the horizontal view is acquired.

8. The method for calculating road pothole depth based on monocular camera images according to claim 7, characterized in that: The constant matrix for getting the image into a horizontal view includes: Perform a checkerboard calibration calculation on the original image to obtain a mapping matrix from the original image to the horizontal view, and use the second rotation matrix in combination with the mapping matrix from the original image to the horizontal view to obtain a constant matrix from the image to the horizontal view: in, is the mapping matrix from image to horizontal view, is the second rotation matrix, A constant matrix for the image to horizontal view.

9. The method for calculating road pothole depth based on monocular camera images according to claim 1, characterized in that: Obtaining the actual depth of the pit comprises: in, is the actual depth of the pit, is the pixel distance of the ground pit depth in the horizontal view, is the distance from the imaging plane to the monocular camera position, is the position from the pit to the camera.

10. A road pothole depth calculation system based on monocular camera images, characterized in that: include: An image acquisition module, used for acquiring an image to be detected; The cloud storage calculation module extracts the vertical coordinate of the vanishing point of the image to be detected, the 3D information of the pit, and the height of the pit, and obtains the actual distance from the center point of the top surface of the pit to the camera and the pixel distance of the pit depth in the horizontal view according to the vertical coordinate of the vanishing point combined with the 3D information of the pit and the height of the pit, and obtains the actual depth of the pit by combining the actual distance and the pixel distance with the distance from the imaging plane of the monocular camera to the camera.

Citation Information

Patent Citations

  • Ramp monocular distance measurement method and device based on vanishing point and target grounding point

    CN115597550A

  • Pit hole detection method, electronic equipment and storage medium

    CN116331245A

  • Monocular camera pitch angle determination and depth estimation method, device and equipment

    CN117557616A

  • Driving decision system and method based on pit recognition

    CN117864169A

  • Road facility height calculation method based on monocular camera image

    CN118918169A