An expressway construction inspection optimization method based on artificial intelligence
By deploying depth cameras and mobile inspection vehicles at highway construction sites, and combining boundary recognition models and 3D matching algorithms, the accuracy of foundation pit boundary identification and measurement was solved, enabling intelligent construction and refined assessment of safety risks at the construction site.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LIAOCHENG TRANSPORTATION DEV CO LTD
- Filing Date
- 2025-09-18
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are insufficient for accurately identifying and measuring the boundaries of foundation pits in highway construction under complex environments. Traditional methods cannot effectively identify the spatial geometric features of foundation pits, leading to inaccurate construction safety risk assessments.
A fixed-position depth camera combined with a boundary recognition model is used to identify the pixels at the foundation pit boundary. The model is then combined with a LiDAR and a binocular camera on a mobile inspection vehicle to perform 3D point cloud matching. A stereo matching algorithm for image-point cloud fusion is constructed, and discrete grid areas are divided for safety risk assessment.
It significantly improved the accuracy and precision of foundation pit boundary identification, realized intelligent inspection of construction sites, reduced false detection and missed detection rates, and enhanced the refinement of construction safety risk assessment.
Smart Images

Figure CN121169876B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of construction inspection, and more particularly to image recognition technology, specifically an artificial intelligence-based optimization method for highway construction inspection. Background Technology
[0002] Highway construction and maintenance are crucial for ensuring long-term road safety and improving traffic efficiency. With the continuous growth of traffic volume and increasing road usage frequency, construction operations have gradually shifted from traditional manual inspections to intelligent and information-based processes. However, excavation pits, as a common construction structure, pose significant safety risks during construction and maintenance. Firstly, excavation pits are typically formed through excavation, and their boundaries and shapes are often complex and variable. If not monitored and accurately marked in real time, they can easily lead to construction machinery entering the pit or workers falling in, causing serious personal injury and equipment damage. Secondly, if excavation pits are not identified in a timely manner or their parameters are inaccurately assessed, vehicles traveling near the construction area may be affected by pit collapses or insufficient road surface support, potentially causing traffic accidents and posing risks to road safety and operational efficiency.
[0003] Excavation pits vary not only in depth and width, but their area and spatial distribution also dynamically adjust with the progress of construction. Therefore, traditional methods relying on manual inspections and experience-based judgment are insufficient to guarantee comprehensive and accurate information about the pits, resulting in monitoring delays and inadequate identification accuracy. In actual construction scenarios, environmental factors such as nighttime operations, complex terrain, and rainwater accumulation further complicate the identification of pit boundaries and risk assessment. Thus, achieving rapid identification and high-precision measurement of excavation pit areas has become a key technological bottleneck restricting construction safety management.
[0004] Existing research largely focuses on highway inspection methods based on artificial intelligence and big data. Patent CN117726324B, "A Highway Traffic Construction Inspection Method and System Based on Data Recognition," proposes an inspection approach combining neural networks and predictive models. This method acquires future maintenance values, predicts vehicle transport data, predicts temperature data, and predicts precipitation data, and imports this data into a constructed neural network model to obtain the future trend of maintenance values. Furthermore, by estimating the time it takes for maintenance values to reach maintenance thresholds, it achieves prediction and early warning of road maintenance time, thereby reducing the time lag between road damage and repair, and improving the foresight and timeliness of road maintenance. Although the aforementioned patent has made progress in road damage prediction and maintenance time early warning, it still has the following technical challenges: the patent focuses on time-dimensional prediction and cannot quantitatively measure the spatial geometric features of the foundation pit, such as its boundary, depth, and width; it uses predicted maintenance time for early warning but fails to combine actual site topography information and structural risks, and does not perform targeted intelligent analysis of the most dangerous foundation pits during construction.
[0005] To address this issue, this invention proposes an AI-based optimization method for highway construction inspection. By introducing AI technology to automate the calculation and measurement of foundation pits at construction sites, intelligent inspection is achieved, improving the efficiency and accuracy of inspections, reducing the cost of manual inspections, and enhancing the overall management level of the construction process. Summary of the Invention
[0006] This invention provides an AI-based optimization method for highway construction inspection. Traditional construction inspections often rely on manual visual inspection or single camera equipment, which are easily affected by changes in lighting, dust obstruction, and weak textures, leading to inaccurate detection of pit boundaries. The proposed method, in step S1, uses a fixed-position depth camera combined with a boundary recognition model to automatically extract and map pit boundary pixels, ensuring accurate identification of pit areas in complex environments and reducing false positives and false negatives. Step S3 uses a binocular camera to calculate parallax and incorporates the 3D point cloud of a mobile inspection vehicle as geometric constraints to construct a stereo matching algorithm that fuses images and point clouds, significantly improving the accuracy and robustness of pit depth and width estimation. Traditional construction safety assessments are often based on single-point measurements or empirical judgments, lacking refined spatial distribution analysis and easily overlooking local high-risk areas. In step S4, this invention divides the construction site into discrete grid areas and, combined with comprehensive information such as pit area, depth, and width, quantitatively calculates the safety risk value of each discrete grid, achieving spatialization and visualization of risk and effectively identifying high-risk areas.
[0007] To achieve the above objectives, this invention provides an artificial intelligence-based optimization method for highway construction inspection, comprising the following steps:
[0008] S1: Use a fixed-position depth camera to acquire construction ground images of the highway construction site, use a boundary recognition model to identify the foundation pit boundary pixels in the construction ground images, map the foundation pit boundary pixels to the highway construction site, and take the area enclosed by the mapped foundation pit boundary pixels as the foundation pit.
[0009] S2: Control the lidar scanner deployed on the mobile inspection vehicle to collect auxiliary point cloud data of the foundation pit, and control the binocular camera deployed on the mobile inspection vehicle to collect the left and right views of the foundation pit.
[0010] S3: Based on the left and right views, the disparity information of the foundation pit is calculated, and combined with the auxiliary point cloud data, a stereo matching algorithm is used to estimate the depth and width of the foundation pit to obtain the depth and width of the foundation pit;
[0011] S4: Divide the highway construction site into discrete grid areas, generate safety risk values for the discrete grids based on the foundation pit information within the discrete grids, and conduct secondary inspections on discrete grids whose safety risk values exceed the allowable risk. The foundation pit information includes the area, depth, and width of the foundation pit.
[0012] As a further improvement of the present invention:
[0013] Furthermore, a boundary recognition model is used to identify the boundary pixels of the foundation pit in the construction ground image, including:
[0014] Multiple fixed-position depth cameras are deployed at the highway construction site to collect images of the construction ground, and the collected images are sent to the boundary recognition model.
[0015] The boundary recognition model includes an input layer, a contrast enhancement layer, and a boundary recognition layer. The input layer receives the construction ground image and performs grayscale processing to obtain a grayscale image of the construction ground. The contrast enhancement layer calculates the contrast of the grayscale image of the construction ground. For grayscale images of construction ground with low contrast, a histogram equalization method is used to enhance the contrast. For grayscale images of construction ground that are overexposed or underexposed, gamma correction is performed to obtain a contrast-enhanced image of the construction ground. The boundary recognition layer is used to identify the foundation pit boundary pixels in the contrast-enhanced image of the construction ground, obtain the foundation pit boundary pixels in the contrast-enhanced image of the construction ground, and extract the pixel coordinates of the foundation pit boundary pixels.
[0016] The boundary recognition layer adopts an improved YOLOv8 model structure. The improvement is achieved by introducing a multi-scale feature fusion structure into the backbone network and a spatial attention mechanism into the detection head.
[0017] A boundary recognition model is used to identify the boundaries of the construction ground image, and the pixel coordinates of the pixels identified as the foundation pit boundary are obtained. The pixel coordinates of the foundation pit boundary pixels are then used to map the foundation pit boundary pixels to the highway construction site.
[0018] Furthermore, the step of mapping the foundation pit boundary pixels to the highway construction site using the pixel coordinates of the foundation pit boundary pixels, and defining the area enclosed by the mapped positions of the foundation pit boundary pixels as the foundation pit, includes:
[0019] Extract the depth value at the pixel coordinates of the pixels at the edge of the pit, and combine the depth value to map the pixel coordinates to the camera coordinate system based on the intrinsic parameter matrix of the depth camera, so as to obtain the three-dimensional camera coordinates of the pixel coordinates of the pixels at the edge of the pit in the camera coordinate system.
[0020] The 3D camera coordinates are transformed to a unified highway coordinate system using the extrinsic rotation matrix and extrinsic translation vector of the depth camera, thus obtaining the 3D highway coordinates of the 3D camera coordinates in the highway coordinate system.
[0021] Connect the three-dimensional highway coordinates in the highway coordinate system to form a closed polygon. The closed polygon is the area enclosed by the pixel mapping position of the pit boundary.
[0022] Furthermore, the system controls the lidar scanner deployed on the mobile inspection vehicle to collect auxiliary point cloud data of the foundation pit, and controls the binocular camera deployed on the mobile inspection vehicle to collect left and right views of the foundation pit, including:
[0023] At highway construction sites, mobile inspection vehicles are deployed in appropriate locations around the foundation pits to ensure that the mobile inspection vehicles can cover each foundation pit along the inspection route.
[0024] The mobile inspection vehicle is equipped with a lidar scanner and a binocular camera. The binocular camera includes a left-view camera and a right-view camera, which are used to acquire the left view and the right view, respectively.
[0025] The mobile inspection vehicle is controlled to move along the perimeter of the foundation pit, so that the laser radar scanner can perform a 360° rotation scan of the foundation pit at a preset scanning frequency, and collect three-dimensional point clouds of the foundation pit and its perimeter, forming a three-dimensional point cloud set as auxiliary point cloud data of the foundation pit. The three-dimensional point cloud is in the form of three-dimensional coordinates.
[0026] The mobile inspection vehicle is controlled to simultaneously activate the left and right view cameras to capture images, obtaining a panoramic view of the foundation pit from both the left and right sides.
[0027] Set the parallax search range to ;
[0028] Calculate the disparity values of any pixel coordinates in the grayscale left view under different disparities, and use them as the disparity information of the foundation pit:
[0029] ;
[0030] in, This represents the disparity value of pixel coordinate u in the grayscale left view under disparity q. This represents the set of pixel coordinates of the grayscale left view. This represents the left view in grayscale centered at pixel coordinate u. Pixel region, v represents pixel region Any pixel coordinate in the array, This represents the grayscale value at pixel coordinate v in the grayscale left view. This represents the pixel coordinates after the pixel coordinate v has been translated q pixels horizontally. Represents the pixel coordinates in the grayscale right view The grayscale value at that location;
[0031] Obtain auxiliary point cloud data of the foundation pit.
[0032] Furthermore, by combining auxiliary point cloud data, a stereo matching algorithm is used to estimate the depth and width of the foundation pit, obtaining the depth and width of the foundation pit, including:
[0033] The point cloud consistency matching cost between any pixel coordinate in the grayscale left view and the auxiliary point cloud data is calculated. The point cloud consistency matching cost is the minimum coordinate difference between the pixel coordinate and the three-dimensional point cloud projection result of all three-dimensional point clouds in the auxiliary point cloud data. The three-dimensional point cloud projection result is the projection of the three-dimensional point cloud onto the pixel coordinate system where the pixel coordinate is located in the grayscale left view.
[0034] The point cloud consistency matching cost between pixel coordinates in the grayscale left view and auxiliary point cloud data, as well as the disparity value of pixel coordinates under different disparities, are added together and used as the initial matching cost of pixel coordinates in the grayscale left view under different disparities.
[0035] The initial matching cost of pixel coordinates in the grayscale left view under different disparities is accumulated in different directions, and the cost accumulation results in multiple directions are summed as the path accumulation cost of pixel coordinates under different disparities.
[0036] The disparity that minimizes the path accumulation cost of pixel coordinates in the grayscale left view is selected. Based on the selected disparity, camera baseline, and focal length, the disparity depth of pixel coordinates in the grayscale left view is generated. The disparity depth is used as the depth value of pixel coordinates in the grayscale left view. The pixel coordinates in the grayscale left view are mapped to the highway coordinate system using the method described in step S1, as described in step S1. The mapped three-dimensional highway coordinates are obtained. The coordinate value of the mapped three-dimensional highway coordinates on the Z-axis is used as the depth of pixel coordinates in the grayscale left view at the pit location.
[0037] Select the mapped 3D highway coordinates of all pixel coordinates in the grayscale left view, and project the mapped 3D highway coordinates onto the plane coordinate system in the highway coordinate system to obtain the projected coordinates of each mapped 3D working coordinate. Calculate the Euclidean distance between any two projected coordinates and select the largest Euclidean distance as the width of the pit.
[0038] The maximum value of the depth of all pixel coordinates at the pit location in the grayscale left view is selected as the depth of the pit.
[0039] Further, step S4 includes:
[0040] Project the highway construction site onto a plane coordinate system within the highway coordinate system;
[0041] The projected highway construction site is divided into several square discrete grids;
[0042] The safety risk value of the discrete grid is generated based on the foundation pit information within the discrete grid.
[0043] Furthermore, the formula for calculating the security risk value of the discrete mesh is as follows:
[0044] ;
[0045] Where R represents the safety risk value of the discrete grid, Len represents the side length of the discrete grid, and H represents the number of foundation pits within the discrete grid. Let represent the area of the h-th pit within the discrete grid. This represents the depth of the h-th pit within the discrete grid. This represents the width of the h-th pit within the discrete grid. This indicates the maximum depth of all foundation pits at the highway construction site. The areas are respectively ,depth and width The normalized value, This represents the volume percentage factor of the h-th foundation pit;
[0046] All represent effect coefficients. This represents the nonlinear amplification factor.
[0047] Compared with existing technologies, this invention proposes an artificial intelligence-based optimization method for highway construction inspection, which has the following beneficial effects:
[0048] First, this invention improves the YOLOv8 model to accurately identify the boundary pixels of the foundation pit in the contrast-enhanced image of the construction site, significantly improving the automation and accuracy of construction site monitoring. Specifically, by introducing a multi-scale feature fusion structure into the backbone network, the model can simultaneously capture the overall outline of the foundation pit at low resolution and the information of slender boundary lines and auxiliary markers at medium to high resolution, thus effectively solving the problem that traditional single-scale feature extraction is difficult to accurately identify complex edges. The introduction of the spatial attention mechanism enables the model to produce a higher response to key boundaries and weak texture areas in the fused feature map, improving robustness to changes in illumination, shadow occlusion, and noise interference.
[0049] Meanwhile, this invention tightly couples the discrete grid of the highway construction site with the information of the foundation pit, significantly improving the precision of construction risk assessment. Specifically, the formula introduces normalization and nonlinear penalty terms on the basis of traditional area, depth, and width indicators. This makes the safety risk value amplified for foundation pits with excessive depth or width exceeding the limit, which is more in line with the actual propagation law of construction safety hazards. By coupling the discrete grid scale with area proportion factors and volume proportion factors, it can reflect the degree of damage of the foundation pit to the overall stability of the local foundation, avoiding the shortcomings of simple area or depth evaluation methods in missing large-scale hazards. In this way, not only can the comparable calculation of foundation pits of different sizes be achieved, but also the discrete grids that exceed the allowable risk can be highlighted, thereby providing a scientific basis and accurate positioning for subsequent secondary inspections, effectively improving the intelligence and safety assurance level of highway construction site inspections. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating an artificial intelligence-based optimization method for highway construction inspection, as provided in an embodiment of the present invention.
[0051] Figure 2 This is a flowchart of the foundation pit boundary pixel recognition process provided in an embodiment of the present invention. Detailed Implementation
[0052] The realization of the objectives, functional characteristics, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0053] This invention provides an AI-based optimization method for highway construction inspection. The executing entity of this AI-based highway construction inspection optimization method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this invention: a server, a terminal, etc. In other words, the AI-based highway construction inspection optimization method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0054] Reference Figure 1 as well as Figure 2 Embodiment 1 of the present invention is as follows:
[0055] An AI-based optimization method for highway construction inspection includes the following steps:
[0056] S1: Use a fixed-position depth camera to acquire construction ground images of the highway construction site, use a boundary recognition model to identify the foundation pit boundary pixels in the construction ground images, map the foundation pit boundary pixels to the highway construction site, and take the area enclosed by the mapped foundation pit boundary pixel position as the foundation pit.
[0057] A boundary recognition model is used to identify the boundary pixels of the foundation pit in the construction ground image, including:
[0058] Multiple fixed-position depth cameras are deployed at the highway construction site to collect images of the construction ground, and the collected images are sent to the boundary recognition model.
[0059] The boundary recognition model includes an input layer, a contrast enhancement layer, and a boundary recognition layer. The input layer receives the construction ground image and performs grayscale processing to obtain a grayscale image of the construction ground. The contrast enhancement layer calculates the contrast of the grayscale image of the construction ground. For grayscale images of construction ground with low contrast, a histogram equalization method is used to enhance the contrast. For grayscale images of construction ground that are overexposed or underexposed, gamma correction is performed to obtain a contrast-enhanced image of the construction ground. The boundary recognition layer is used to identify the foundation pit boundary pixels in the contrast-enhanced image of the construction ground, obtain the foundation pit boundary pixels in the contrast-enhanced image of the construction ground, and extract the pixel coordinates of the foundation pit boundary pixels.
[0060] As an embodiment of the present invention, the histogram equalization method is the CLAHE (Contrast Limited Adaptive Histogram Equalization) method; the contrast calculation method of the grayscale image of the construction ground is to calculate the standard deviation and mean of the pixel grayscale values of the grayscale image of the construction ground. If the standard deviation is lower than the empirically set preset standard deviation threshold (e.g., 25), the grayscale image of the construction ground is low contrast; if the mean is higher than the empirically set preset mean threshold (e.g., 40), the grayscale image of the construction ground is overexposed; if the mean is lower than the empirically set preset mean threshold, the grayscale image of the construction ground is underexposed.
[0061] Specifically, for overexposed grayscale images of construction ground, a gamma correction factor greater than 1 is set, and for underexposed grayscale images of construction ground, a gamma correction factor less than 1 is set.
[0062] It should be noted that the contrast enhancement layer can accurately determine whether the image is low in contrast, overexposed, or underexposed by calculating the pixel mean and standard deviation of the grayscale image of the construction ground. For different situations, it uses CLAHE histogram equalization or gamma correction methods to process the image. The gamma coefficient is set to be greater than 1 for overexposed images and less than 1 for underexposed images, thereby effectively enhancing the texture and details of the construction ground and improving the visibility of the pit edge.
[0063] The boundary recognition layer adopts an improved YOLOv8 model structure. The improvements are as follows: a multi-scale feature fusion structure is introduced into the backbone network to enhance the ability to capture the slender structure of the foundation pit edge; a spatial attention mechanism is introduced into the detection head to improve the localization accuracy of boundary points in weak texture areas; and a dataset of foundation pit edges and auxiliary markers in the construction scenario is added during the training phase to improve the robustness of the boundary recognition model under complex lighting and noise conditions.
[0064] As a preferred embodiment of the present invention, refer to Figure 2 As shown, the process for identifying the pit boundary pixels in the boundary recognition layer is as follows:
[0065] S101: The backbone network in the boundary recognition layer performs multi-scale feature extraction on the contrast-enhanced image of the construction ground, capturing image features at low, medium, and high resolutions respectively. The image features at low resolution are used to capture the overall outline of the foundation pit, while the image features at medium and high resolutions are used to capture slender boundary lines and details of auxiliary markers.
[0066] S102: It fuses image features at different scales and enhances the response capability to key boundaries and weakly textured regions through a spatial attention mechanism;
[0067] S103: Use the detection head to generate a spatial attention map on the fused feature map to highlight the possible pit boundary pixel regions, and use the detection head to generate a boundary probability map on the fused feature map, wherein the boundary probability map consists of the probability that each pixel is a boundary.
[0068] S104: Perform weighted fusion of the spatial attention map and the boundary probability map to obtain the boundary weighted probability map;
[0069] S105: Extract the probability of each pixel being a pit boundary pixel from the boundary weighted probability map, and mark pixels with a probability higher than a preset probability threshold (such as 0.6) as pit boundary pixels;
[0070] S106: Perform connected component analysis and boundary smoothing on all foundation pit boundary pixels in the contrast-enhanced image of the construction ground to obtain the final foundation pit boundary pixel recognition result.
[0071] As an embodiment of the present invention, the training process for the model structure parameters of the boundary recognition layer is as follows:
[0072] The construction site was photographed by a fixed camera to obtain high-resolution images from different angles and under different lighting conditions. The high-resolution images collected covered the entire foundation pit area and surrounding auxiliary markers (such as guardrails, warning signs, etc.). The real labels of each pixel in the high-resolution images were manually labeled to form a training set. A real label of 1 indicates that the pixel is a foundation pit boundary pixel, and a real label of 0 indicates that the pixel is not a foundation pit boundary pixel.
[0073] The training loss function of the boundary recognition layer is constructed based on the training set. The model structure parameters in the boundary recognition layer are trained using Adam or SGD optimizer. The model structure parameters are updated by backpropagation based on the training loss function results.
[0074] The training loss function of the boundary recognition layer is: :
[0075] ;
[0076] in, This represents the probability that the i-th pixel in the training set, predicted by the boundary recognition layer, is a pit boundary pixel. This represents the true label of the i-th pixel in the training set. , This indicates that the i-th pixel in the training set is the pit boundary pixel. Let S represent the number of pixels in the training set, indicating that the i-th pixel is not a pit boundary pixel. The smaller the value of the training loss function, the higher the overlap between the predicted probability and the true label, and the higher the recognition accuracy of the pit boundary pixels.
[0077] It should be noted that this invention improves the YOLOv8 model to achieve accurate identification of foundation pit boundary pixels in contrast-enhanced construction site images, significantly improving the automation and accuracy of construction site monitoring. Specifically, by introducing a multi-scale feature fusion structure into the backbone network, the model can simultaneously capture the overall foundation pit outline at low resolution and the slender boundary lines and auxiliary marker information at medium to high resolution, effectively solving the problem that traditional single-scale feature extraction cannot accurately identify complex edges. The introduction of a spatial attention mechanism enables the model to produce a higher response to key boundaries and weakly textured areas in the fused feature map, improving robustness to changes in illumination, shadow occlusion, and noise interference. By generating a spatial attention map and a boundary probability map through the detection head and then weighted and fused, the potential foundation pit boundary pixel areas can be accurately highlighted. By setting a probability threshold to filter high-confidence boundary points, false detections and false negatives are reduced. Connectivity analysis and boundary smoothing further enhance the continuity and realism of the boundaries, providing reliable input for subsequent 3D reconstruction, depth measurement, and slope stability assessment. The boundary recognition model not only improves the recognition accuracy and reliability of foundation pit boundary pixels, but also reduces the reliance on manual inspection. It can quickly and accurately obtain foundation pit boundary information in complex construction environments, providing a solid data foundation for construction anomaly detection, construction safety early warning and management decision-making, thereby significantly improving the intelligence level of highway construction monitoring and construction safety assurance capabilities.
[0078] A boundary recognition model is used to identify the boundaries of the construction ground image, and the pixel coordinates of the pixels identified as the foundation pit boundary are obtained. The pixel coordinates of the foundation pit boundary pixels are then used to map the foundation pit boundary pixels to the highway construction site.
[0079] The process of mapping the boundary pixels of the foundation pit to the highway construction site using the pixel coordinates of the foundation pit boundary pixels, and defining the area enclosed by the mapped boundary pixel positions as the foundation pit, includes:
[0080] Extract the depth value at the pixel coordinates of the pixels at the foundation pit boundary (directly acquired by the depth camera). Combine the depth value with the intrinsic parameter matrix of the depth camera to map the pixel coordinates to the camera coordinate system, and obtain the three-dimensional camera coordinates of the pixel coordinates of the foundation pit boundary pixels in the camera coordinate system.
[0081] Specifically, pixel coordinates The mapping formula to the camera coordinate system is:
[0082] ;
[0083] ;
[0084] in, Represents pixel coordinates 3D camera coordinates mapped to the camera coordinate system. Represents pixel coordinates The depth value, Let T represent the intrinsic parameter matrix of the depth camera, and let T denote the transpose. Representing the intrinsic parameter matrix The inverse matrix; specifically, This indicates the horizontal focal length of the depth camera. This indicates the vertical focal length of the depth camera. Indicates the principal point coordinates of the depth camera;
[0085] The 3D camera coordinates are transformed to a unified highway coordinate system by using the extrinsic rotation matrix and extrinsic translation vector of the depth camera, thus obtaining the 3D highway coordinates of the 3D camera coordinates in the highway coordinate system. Specifically, the transformation to a unified highway coordinate system enables the pit boundary pixels captured by depth cameras deployed at different fixed locations to be mapped to a unified highway coordinate system.
[0086] 3D camera coordinates The transformation formula for converting to a unified highway coordinate system is as follows:
[0087] ;
[0088] in, Represents 3D camera coordinates Transform to a unified highway coordinate system using three-dimensional highway coordinates. This represents the 3x3 extrinsic rotation matrix of the depth camera. This represents the 3x1 extrinsic translation vector of the depth camera.
[0089] Connect the three-dimensional highway coordinates in the highway coordinate system to form a closed polygon. The closed polygon is the area enclosed by the pixel mapping position of the pit boundary.
[0090] Optionally, the three-dimensional road coordinates can be connected into a closed polygon using a distance-based connection method, a connected domain closure method, or a curve fitting method. In the distance-based connection method, if the Euclidean distance between the three-dimensional road coordinates is less than a preset distance threshold (e.g., 0.4 meters), then the two three-dimensional road coordinates are connected.
[0091] S2: Control the lidar scanner deployed on the mobile inspection vehicle to collect auxiliary point cloud data of the foundation pit, and control the binocular camera deployed on the mobile inspection vehicle to collect the left and right views of the foundation pit.
[0092] Control the lidar scanner deployed on the mobile inspection vehicle to collect auxiliary point cloud data of the foundation pit, and control the binocular camera deployed on the mobile inspection vehicle to collect left and right views of the foundation pit, including:
[0093] At highway construction sites, mobile inspection vehicles are deployed in appropriate locations around the foundation pits to ensure that the mobile inspection vehicles can cover each foundation pit along the inspection route.
[0094] The mobile inspection vehicle is equipped with a lidar scanner and a binocular camera. The binocular camera includes a left-view camera and a right-view camera, which are used to acquire the left and right views, respectively. The horizontal spacing of the binocular camera has been calibrated to ensure that image pairs that can be used for stereo matching can be acquired. The lidar scanner and the binocular camera are time-synchronized to ensure the consistency of data acquired at the same time in space and time.
[0095] The mobile inspection vehicle is controlled to move along the perimeter of the foundation pit, so that the laser radar scanner can perform a 360° rotation scan of the foundation pit at a preset scanning frequency (such as 10 Hz) to collect three-dimensional point clouds of the foundation pit and its perimeter, forming a three-dimensional point cloud set as auxiliary point cloud data of the foundation pit. The three-dimensional point cloud is in the form of three-dimensional coordinates.
[0096] Specifically, the collected 3D point cloud is mapped to the highway coordinate system, the 3D point cloud is replaced with the mapped coordinates, and a 3D point cloud set is constructed.
[0097] The mobile inspection vehicle is controlled to simultaneously activate the left and right view cameras to capture images, obtaining a panoramic view of the foundation pit from both the left and right sides.
[0098] S3: Based on the left and right views, the disparity information of the foundation pit is calculated. Combined with auxiliary point cloud data, a stereo matching algorithm is used to estimate the depth and width of the foundation pit, thus obtaining the depth and width of the foundation pit.
[0099] The disparity information of the foundation pit is calculated based on the left and right views, including:
[0100] The left and right views are converted to grayscale to obtain grayscale left and right views;
[0101] Set the parallax search range to Optionally, Q can be set to 5 (in pixels).
[0102] ;
[0103] in, This represents the disparity value of pixel coordinate u in the grayscale left view under disparity q. This represents the set of pixel coordinates of the grayscale left view. This represents the left view in grayscale centered at pixel coordinate u. Pixel region, v represents pixel region Any pixel coordinate in the array, This represents the grayscale value at pixel coordinate v in the grayscale left view. This represents the pixel coordinates after the pixel coordinate v has been translated q pixels horizontally. Represents the pixel coordinates in the grayscale right view The grayscale value at that location.
[0104] Combining auxiliary point cloud data, a stereo matching algorithm is used to estimate the depth and width of the foundation pit, obtaining the depth and width of the foundation pit, including:
[0105] The point cloud consistency matching cost between any pixel coordinate in the grayscale left view and the auxiliary point cloud data is calculated. The point cloud consistency matching cost is the minimum coordinate difference between the pixel coordinate and the three-dimensional point cloud projection result of all three-dimensional point clouds in the auxiliary point cloud data. The larger the coordinate difference, the higher the point cloud consistency matching cost. The three-dimensional point cloud projection result is the projection of the three-dimensional point cloud onto the pixel coordinate system where the pixel coordinate is located in the grayscale left view.
[0106] The point cloud consistency matching cost between pixel coordinates in the grayscale left view and auxiliary point cloud data, as well as the disparity value of pixel coordinates under different disparities, are added together and used as the initial matching cost of pixel coordinates in the grayscale left view under different disparities.
[0107] Specifically, the initial matching cost of pixel coordinate u in the grayscale left view under disparity q is: :
[0108] ;
[0109] in, This represents the point cloud consistency matching cost for pixel coordinate u in the grayscale left view. This indicates the control weight, which is set based on experience. It is 0.5;
[0110] The initial matching cost of pixel coordinates in the grayscale left view under different disparities is accumulated in different directions, and the cost accumulation results in multiple directions are summed as the path accumulation cost of pixel coordinates under different disparities.
[0111] Specifically, the cumulative path cost calculation method for pixel coordinate u in the grayscale left view under disparity q is as follows:
[0112] ;
[0113] ;
[0114] in, This represents the pixel coordinate u in the grayscale left view, along with the disparity q and direction. The cumulative result of the cost, in which direction The value range includes 8 directions: up, down, left, right, and the 4 diagonal directions. Represents pixel coordinate u along the direction The pixel coordinates after translation by one pixel. Represents the pixel coordinates in the grayscale left view In terms of parallax q and direction The cumulative result of the cost, Represents the pixel coordinates in the grayscale left view In parallax q+1 and direction The cumulative result of the cost, Represents the pixel coordinates in the grayscale left view In terms of parallax q-1 and direction The cumulative result of the cost;
[0115] This represents the cumulative path cost of pixel coordinate u in the grayscale left view under disparity q;
[0116] Represents the parallax variation penalty constant, set It is 2;
[0117] Indicates selection , , The minimum value in;
[0118] The disparity that minimizes the path accumulation cost of pixel coordinates in the grayscale left view is selected. Based on the selected disparity, camera baseline, and focal length, the disparity depth of pixel coordinates in the grayscale left view is generated. The disparity depth is used as the depth value of pixel coordinates in the grayscale left view. The pixel coordinates in the grayscale left view are mapped to the highway coordinate system using the method described in step S1, as described in step S1. The mapped three-dimensional highway coordinates are obtained. The coordinate value of the mapped three-dimensional highway coordinates on the Z-axis is used as the depth of pixel coordinates in the grayscale left view at the pit location.
[0119] Specifically, the formula for generating the disparity depth is:
[0120] ;
[0121] Where deep represents the parallax depth, B represents the camera baseline, and f represents the camera's focal length. Indicates the selected parallax;
[0122] The mapped 3D highway coordinates of all pixel coordinates in the grayscale left view are selected, and the mapped 3D highway coordinates are projected onto the plane coordinate system within the highway coordinate system to obtain the projected coordinates of each mapped 3D working coordinate. The Euclidean distance between any two projected coordinates is calculated, and the largest Euclidean distance is selected as the width of the pit. Specifically, the highway coordinate system is a 3D coordinate system. The two horizontal axes in the highway coordinate system are extracted, and the resulting plane coordinate system is the plane coordinate system within the highway coordinate system.
[0123] The maximum value of the depth of all pixel coordinates at the pit location in the grayscale left view is selected as the depth of the pit.
[0124] It should be noted that this invention, by introducing a point cloud consistency matching cost between the grayscale left view pixel coordinates and the auxiliary point cloud data, can effectively compensate for the shortcomings of matching based solely on pixel grayscale or disparity. Traditional disparity matching is prone to mismatches in construction scenarios with sparse textures, varying lighting, or noise. Point cloud projection provides additional three-dimensional geometric constraints, making the matching cost calculation more reliable, thereby improving the accuracy of the initial matching cost.
[0125] In the path cost accumulation stage, this invention employs a multi-directional (up and down, left and right, and four diagonals) cost accumulation method, combined with a disparity change penalty mechanism, effectively suppressing continuous disparity jumps. Compared to traditional single-directional cost accumulation, this method can reduce erroneous disparity distributions caused by noise or occlusion. In the depth generation stage, a precise geometric model is introduced based on the camera baseline and focal length, ensuring the consistency of the physical meaning from disparity estimation to depth values. Furthermore, the conversion from pixels to actual 3D space is completed by mapping to the highway coordinate system, thereby ensuring that the depth of the foundation pit in the construction scenario can intuitively reflect the actual terrain depression. Specifically, the method proposed in this invention not only improves the automation level of foundation pit identification and measurement, but also provides reliable basic data support for subsequent safety risk value calculation and secondary inspection decision-making, thereby improving the intelligence and safety of highway construction site inspections.
[0126] S4: Divide the highway construction site into discrete grid areas, generate safety risk values for the discrete grids based on the foundation pit information within the discrete grids, and conduct secondary inspections on discrete grids whose safety risk values exceed the allowable risk. The foundation pit information includes the area, depth, and width of the foundation pit.
[0127] The highway construction site is projected onto a plane coordinate system within the highway coordinate system; specifically, the highway coordinate system is a three-dimensional coordinate system, and the two horizontal axes in the highway coordinate system are extracted to form a plane coordinate system within the highway coordinate system.
[0128] The projected highway construction site is divided into several square discrete grids;
[0129] Optionally, the side length Len of the discrete grid of the square satisfies ,in The range of side lengths for the discrete grid of the set square;
[0130] The side length of the discrete grid can be adjusted arbitrarily within the range of side length, so that the discrete grid can cover the projected highway construction site. For the foundation pit that crosses the discrete grid, the side length of the discrete grid is adjusted so that the same foundation pit is completely contained in a single discrete grid, ensuring that each foundation pit belongs to only one discrete grid.
[0131] The safety risk value of the discrete grid is generated based on the foundation pit information within the discrete grid.
[0132] Optionally, allowable risk can be set based on engineering specifications, expert experience, or the quantile of safety risk values of all historical discrete grids. For the method of setting allowable risk based on the quantile of safety risk values of all historical discrete grids, the safety risk values of all discrete grids during the same period of highway construction in history are obtained and sorted in descending order. The safety risk value with a quantile of 10% is selected as the allowable risk. The safety risk value with a quantile of 10% is the safety risk value that is only lower than the obtained 10% safety risk value.
[0133] The formula for calculating the security risk value of the discrete grid is:
[0134] ;
[0135] Where R represents the safety risk value of the discrete grid, Len represents the side length of the discrete grid, and H represents the number of foundation pits within the discrete grid. Let represent the area of the h-th pit within the discrete grid. This represents the depth of the h-th pit within the discrete grid. This represents the width of the h-th pit within the discrete grid. This indicates the maximum depth of all foundation pits at the highway construction site. The areas are respectively ,depth and width The normalized value, The volume ratio factor of the h-th foundation pit is indicated; optionally, the area of the foundation pit in the discrete network is calculated as follows: the projected coordinates of the mapped three-dimensional working coordinates are marked in the plane coordinate system within the highway coordinate system, the marked area is taken as the foundation pit plane in the plane coordinate system, and the area of the smallest circumscribed polygon surrounding the foundation pit plane is calculated in the plane coordinate system.
[0136] Specifically, This reflects the overall degree of damage caused by the h-th pit to the discrete grid surface. A comprehensive safety risk assessment was conducted by combining the depth and width of the h-th foundation pit. The introduction of a nonlinear amplification mechanism exponentially increases the risk of deep pits. It reflects the overall weakening effect of the h-th foundation pit on the local foundation stability and is suitable for identifying small, high-risk foundation pits that are not large in area but are very deep.
[0137] All represent effect coefficients. Indicates the nonlinear amplification factor; set based on experience. The values are set to 0.3, 0.4, and 0.3 respectively. It is 1.2.
[0138] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0139] It should be noted that the sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0141] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. An optimization method for highway construction inspection based on artificial intelligence, characterized in that, The method includes: S1: Use a fixed-position depth camera to acquire construction ground images of the highway construction site, use a boundary recognition model to identify the foundation pit boundary pixels in the construction ground images, map the foundation pit boundary pixels to the highway construction site, and take the area enclosed by the mapped foundation pit boundary pixels as the foundation pit. S2: Control the lidar scanner deployed on the mobile inspection vehicle to collect auxiliary point cloud data of the foundation pit, and control the binocular camera deployed on the mobile inspection vehicle to collect the left and right views of the foundation pit. S3: Based on the left and right views, the disparity information of the foundation pit is calculated, and combined with the auxiliary point cloud data, a stereo matching algorithm is used to estimate the depth and width of the foundation pit to obtain the depth and width of the foundation pit; S4: Divide the highway construction site into discrete grid areas, generate safety risk values for the discrete grids based on the foundation pit information within the discrete grids, and conduct secondary inspections on discrete grids whose safety risk values exceed the allowable risk. The foundation pit information includes the area, depth, and width of the foundation pit. The formula for calculating the security risk value of the discrete grid is: ; Where R represents the safety risk value of the discrete grid, Len represents the side length of the discrete grid, and H represents the number of foundation pits within the discrete grid. Let represent the area of the h-th pit within the discrete grid. This represents the depth of the h-th pit within the discrete grid. This represents the width of the h-th pit within the discrete grid. This indicates the maximum depth of all foundation pits at the highway construction site. The areas are respectively ,depth and width The normalized value, This represents the volume percentage factor of the h-th foundation pit; All represent effect coefficients. This represents the nonlinear amplification factor.
2. The method for optimizing highway construction inspection based on artificial intelligence as described in claim 1, characterized in that, A boundary recognition model is used to identify the boundary pixels of the foundation pit in the construction ground image, including: Multiple fixed-position depth cameras are deployed at the highway construction site to collect images of the construction ground, and the collected images are sent to the boundary recognition model. The boundary recognition model includes an input layer, a contrast enhancement layer, and a boundary recognition layer. The input layer receives the construction ground image and performs grayscale processing to obtain a grayscale image of the construction ground. The contrast enhancement layer calculates the contrast of the grayscale image of the construction ground. For grayscale images of construction ground with low contrast, a histogram equalization method is used to enhance the contrast. For grayscale images of construction ground that are overexposed or underexposed, gamma correction is performed to obtain a contrast-enhanced image of the construction ground. The boundary recognition layer is used to identify the foundation pit boundary pixels in the contrast-enhanced image of the construction ground, obtain the foundation pit boundary pixels in the contrast-enhanced image of the construction ground, and extract the pixel coordinates of the foundation pit boundary pixels. The boundary recognition layer adopts an improved YOLOv8 model structure. The improvement is achieved by introducing a multi-scale feature fusion structure into the backbone network and a spatial attention mechanism into the detection head. A boundary recognition model is used to identify the boundaries of the construction ground image, and the pixel coordinates of the pixels identified as the foundation pit boundary are obtained. The pixel coordinates of the foundation pit boundary pixels are then used to map the foundation pit boundary pixels to the highway construction site.
3. The method for optimizing highway construction inspection based on artificial intelligence as described in claim 2, characterized in that, The process of mapping the boundary pixels of the foundation pit to the highway construction site using the pixel coordinates of the foundation pit boundary pixels, and defining the area enclosed by the mapped boundary pixel positions as the foundation pit, includes: Extract the depth value at the pixel coordinates of the pixels at the edge of the pit, and combine the depth value to map the pixel coordinates to the camera coordinate system based on the intrinsic parameter matrix of the depth camera, so as to obtain the three-dimensional camera coordinates of the pixel coordinates of the pixels at the edge of the pit in the camera coordinate system. The 3D camera coordinates are transformed to a unified highway coordinate system using the extrinsic rotation matrix and extrinsic translation vector of the depth camera, thus obtaining the 3D highway coordinates of the 3D camera coordinates in the highway coordinate system. Connect the three-dimensional highway coordinates in the highway coordinate system to form a closed polygon. The closed polygon is the area enclosed by the pixel mapping position of the pit boundary.
4. The method for optimizing highway construction inspection based on artificial intelligence as described in claim 1, characterized in that, Control the lidar scanner deployed on the mobile inspection vehicle to collect auxiliary point cloud data of the foundation pit, and control the binocular camera deployed on the mobile inspection vehicle to collect left and right views of the foundation pit, including: At highway construction sites, mobile inspection vehicles are deployed in appropriate locations around the foundation pits to ensure that the mobile inspection vehicles can cover each foundation pit along the inspection route. The mobile inspection vehicle is equipped with a lidar scanner and a binocular camera. The binocular camera includes a left-view camera and a right-view camera, which are used to acquire the left view and the right view, respectively. The mobile inspection vehicle is controlled to move along the perimeter of the foundation pit, so that the laser radar scanner can perform a 360° rotation scan of the foundation pit at a preset scanning frequency, and collect three-dimensional point clouds of the foundation pit and its perimeter, forming a three-dimensional point cloud set as auxiliary point cloud data of the foundation pit. The three-dimensional point cloud is in the form of three-dimensional coordinates. The mobile inspection vehicle is controlled to simultaneously activate the left and right view cameras to capture images, obtaining a panoramic view of the foundation pit from both the left and right sides.
5. The method for optimizing highway construction inspection based on artificial intelligence as described in claim 4, characterized in that, The disparity information of the foundation pit is calculated based on the left and right views, including: The left and right views are converted to grayscale to obtain grayscale left and right views; Set the parallax search range to ; Calculate the disparity values of any pixel coordinates in the grayscale left view under different disparities, and use them as the disparity information of the foundation pit: ; in, This represents the disparity value of pixel coordinate u in the grayscale left view under disparity q. This represents the set of pixel coordinates of the grayscale left view. This represents the left view in grayscale centered at pixel coordinate u. Pixel region, v represents pixel region Any pixel coordinate in the array, This represents the grayscale value at pixel coordinate v in the grayscale left view. This represents the pixel coordinates after the pixel coordinate v has been translated q pixels horizontally. Represents the pixel coordinates in the grayscale right view The grayscale value at that location; Obtain auxiliary point cloud data of the foundation pit.
6. The method for optimizing highway construction inspection based on artificial intelligence as described in claim 5, characterized in that, Combining auxiliary point cloud data, a stereo matching algorithm is used to estimate the depth and width of the foundation pit, obtaining the depth and width of the foundation pit, including: The point cloud consistency matching cost between any pixel coordinate in the grayscale left view and the auxiliary point cloud data is calculated. The point cloud consistency matching cost is the minimum coordinate difference between the pixel coordinate and the three-dimensional point cloud projection result of all three-dimensional point clouds in the auxiliary point cloud data. The three-dimensional point cloud projection result is the projection of the three-dimensional point cloud onto the pixel coordinate system where the pixel coordinate is located in the grayscale left view. The point cloud consistency matching cost between pixel coordinates in the grayscale left view and auxiliary point cloud data, as well as the disparity value of pixel coordinates under different disparities, are added together and used as the initial matching cost of pixel coordinates in the grayscale left view under different disparities. The initial matching cost of pixel coordinates in the grayscale left view under different disparities is accumulated in different directions, and the cost accumulation results in multiple directions are summed as the path accumulation cost of pixel coordinates under different disparities. The disparity that minimizes the path accumulation cost of pixel coordinates in the grayscale left view is selected. Based on the selected disparity, camera baseline and focal length, the disparity depth of pixel coordinates in the grayscale left view is generated. The disparity depth is used as the depth value of pixel coordinates in the grayscale left view. The pixels are mapped to the highway construction site using step S1. The pixel coordinates in the grayscale left view are mapped to the highway coordinate system to obtain the mapped three-dimensional highway coordinates. The coordinate value of the mapped three-dimensional highway coordinates on the Z-axis is used as the depth of pixel coordinates in the grayscale left view at the pit location. Select the mapped 3D highway coordinates of all pixel coordinates in the grayscale left view, and project the mapped 3D highway coordinates onto the plane coordinate system in the highway coordinate system to obtain the projected coordinates of each mapped 3D working coordinate. Calculate the Euclidean distance between any two projected coordinates and select the largest Euclidean distance as the width of the pit. The maximum value of the depth of all pixel coordinates at the pit location in the grayscale left view is selected as the depth of the pit.
7. The method for optimizing highway construction inspection based on artificial intelligence as described in claim 1, characterized in that, Step S4 includes: Project the highway construction site onto a plane coordinate system within the highway coordinate system; The projected highway construction site is divided into several square discrete grids; The safety risk value of the discrete grid is generated based on the foundation pit information within the discrete grid.