Concrete structure apparent defect identification method and device based on inspection robot
By using inspection robots to collect concrete structure images in real time and using the YOLOv5 model to identify defects, the problems of low efficiency and poor safety of manual identification are solved, and efficient and accurate defect identification is achieved, especially improving robustness in complex environments.
Patent Information
- Application Number
- CN202510553142.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-04-29
AI Technical Summary
In the existing technology, manual identification of apparent defects in concrete structures is time-consuming and labor-intensive, easily affected by subjective factors, and poses safety risks in high-altitude or dangerous environments.
A method for identifying surface defects of concrete structures based on an inspection robot is adopted. A depth camera is used to collect images in real time, and the defects are identified through the YOLOv5 defect recognition model. LiDAR and light sensors are combined to optimize image quality, and the Transformer module and CBAM module are used to improve recognition accuracy.
It improves the recognition efficiency and accuracy of surface defects in concrete structures, reduces the safety risks of manual identification, and enhances the robustness in complex environments, especially the recognition ability of complex lighting and texture images.
Smart Images

Figure CN120673220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of concrete structure detection, and in particular to a method and device for identifying surface defects of concrete structures based on an inspection robot. Background Art
[0002] With the acceleration of urbanization and the continuous expansion of infrastructure construction, concrete, as one of the most important materials in the construction and engineering fields, its quality and durability are directly related to public safety and economic benefits. However, during use, concrete structures will inevitably develop surface defects such as cracks, spalling, holes and leakage due to environmental factors (such as temperature changes, humidity, chemical corrosion), loads or construction quality problems. These defects are not only early signs of structural deterioration, but may also further develop into serious safety hazards, such as bridge collapse or tunnel seepage. Therefore, the timely identification of surface defects in concrete structures has become a core demand in the field of engineering maintenance and health monitoring.
[0003] Currently, defect identification on the surface of concrete structures is typically performed manually. However, this manual identification method is time-consuming and labor-intensive, and is subject to subjective factors and work experience, resulting in missed detections and misjudgments. Furthermore, manual operation in high-altitude, narrow, or dangerous environments poses safety risks. Summary of the Invention
[0004] The present invention provides a method and device for identifying apparent defects of concrete structures based on an inspection robot, which are mainly capable of improving the recognition efficiency and recognition accuracy of apparent defects of concrete structures and ensuring the safety of workers.
[0005] According to a first aspect of the present invention, a method for identifying surface defects of concrete structures based on an inspection robot is provided, comprising:
[0006] In response to an apparent defect recognition instruction for a target concrete structure, a depth camera device carried by the inspection robot is used to collect an apparent image of the target concrete structure in real time;
[0007] Performing image quality enhancement processing on the apparent image to obtain the processed apparent image;
[0008] Obtain a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, wherein the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a concrete structure apparent defect recognition dataset;
[0009] The processed appearance image is input into the defect recognition model, and global image features and context image features are extracted from the appearance image through each Transformer module in the backbone network. The global image features and context image features extracted by each Transformer module are merged through the SPPF module to obtain the merged image features corresponding to each Transformer module. The key features of the merged image features are enhanced through the channel attention layer and the spatial attention layer of each CBAM module in the neck network, and the enhanced merged image features are feature fused to obtain the fused image features corresponding to each CBAM module. The fused image features corresponding to each CBAM module are used to perform defect recognition through the head network to obtain the defect recognition results output by each prediction head module, wherein the defect recognition results are reflected on the appearance image in the form of a labeled box with a defect type label.
[0010] Optionally, the collecting of the surface image of the target concrete structure in real time by using a depth camera device carried by the inspection robot includes:
[0011] Using multi-source sensors to perceive in real time regional environmental data of the area where the target concrete structure is located, and to obtain in real time illumination information of the area where the target concrete structure is located, surface reflection characteristic information of the target concrete structure, and shooting requirement information;
[0012] fusing the regional environmental data sensed by the multi-source sensors to obtain fused environmental data;
[0013] Based on the fused environmental data, a three-dimensional point cloud map of the area where the target concrete structure is located is constructed in real time;
[0014] Based on the three-dimensional point cloud map, dynamically generate a reachability topology map of the inspection robot, and based on the lighting information, the surface reflection feature information, and the shooting requirement information, determine candidate shooting viewpoints, and determine multiple candidate inspection paths where the candidate shooting viewpoints are located in the reachability topology map;
[0015] Based on the terrain data of the area where each candidate inspection path is located, determining the terrain slope and land use type of each candidate inspection path, as well as the slope resistance corresponding to the terrain slope and the land resistance corresponding to the land use type;
[0016] Determining a relative importance ratio of the terrain slope and the land use type to the inspection resistance, constructing a judgment matrix of the terrain slope and the land use type to the inspection resistance based on the relative importance ratio, and determining resistance weight coefficients of the terrain slope and the land use type respectively based on the judgment matrix;
[0017] Based on the resistance weight coefficient, the slope resistance and the land resistance are weighted and summed to obtain a comprehensive recommendation coefficient corresponding to each candidate inspection path, and based on the comprehensive recommendation coefficient, the inspection path of the inspection robot is determined in each candidate inspection path;
[0018] The inspection robot is controlled to travel along the inspection path, and a depth camera device carried by the inspection robot is used to collect the surface image of the target concrete structure in real time during the traveling process.
[0019] Optionally, the inspection robot is also equipped with a laser radar;
[0020] The performing image quality enhancement processing on the apparent image to obtain the processed apparent image includes:
[0021] Using the laser radar to scan the distance between the inspection robot and the target concrete structure in real time, and generating high-density point cloud data based on the real-time scanned distance;
[0022] Determining depth data between an image pixel point of the apparent image and the inspection robot, wherein the depth data includes a distance, a viewing angle, and a height difference between the image pixel point and the inspection robot;
[0023] Time-synchronizing the high-density point cloud data with the depth data, and fusing the synchronized high-density point cloud data with the synchronized depth data to obtain fused data;
[0024] Determining a relative position relationship between the laser radar and the depth camera device, and generating a transformation matrix based on the relative position relationship;
[0025] Based on the transformation matrix, the fused data is converted into a preset global coordinate system corresponding to the inspection robot to obtain the converted fused data, and based on the converted fused data, the apparent image is geometrically corrected to obtain the processed apparent image, wherein the geometric correction includes at least one of radial distortion correction, tangential distortion correction, perspective alignment correction, and height matching correction.
[0026] Optionally, the inspection robot is further equipped with a light sensor;
[0027] The method of collecting the surface image of the target concrete structure in real time by using the depth camera device carried by the inspection robot includes:
[0028] Using the light sensor to monitor the light intensity around the target concrete structure in real time, and determining the sensor exposure parameters of the depth camera device based on the light intensity, wherein the sensor exposure parameters include exposure time, gain, and infrared projector intensity;
[0029] Based on the sensor exposure parameters, the parameters of the depth camera device are adjusted in real time, and the depth camera device with the adjusted parameters is used to capture the apparent image of the target concrete structure.
[0030] Optionally, before obtaining the defect recognition model based on YOLOv5, the method further includes:
[0031] Construct an initial defect recognition model based on YOLOv5 and obtain a concrete structure surface defect recognition dataset, wherein the concrete structure surface defect recognition dataset contains sample surface images of sample concrete structures with different defect type labels under various lighting conditions;
[0032] Performing operations such as rotation, flipping, and brightness adjustment on the sample appearance images to obtain an expanded concrete structure appearance defect recognition dataset, and dividing the expanded concrete structure appearance defect recognition dataset into training data and test data;
[0033] The initial defect recognition model is trained using the training data, and the trained initial defect recognition model is tested using the test data, and the trained initial defect recognition model that meets the test conditions is determined as the defect recognition model.
[0034] Optionally, after performing defect recognition on the fused image features corresponding to each CBAM module through the head network to obtain the defect recognition results output by each prediction head module, the method further includes:
[0035] extracting defect geometric features in the target concrete structure based on the defect identification result;
[0036] Based on the defect geometric characteristics, the target concrete structure is classified into defect levels using a preset level classification configuration table to obtain the defect level of the target concrete structure;
[0037] determining, based on the three-dimensional point cloud data of the target concrete structure, defect locations corresponding to defects in the target concrete structure, correlating defect geometric features, defect levels, and defect locations corresponding to the defects in the target concrete structure, and generating a defect distribution map based on the correlated data;
[0038] The defect distribution map is uploaded to a cloud database for distributed storage, and the defect distribution map is transmitted to a user terminal for visual display of a preset dimension.
[0039] Optionally, after classifying the target concrete structure by using a preset grade classification configuration table to obtain the defect grade of the target concrete structure, the method further includes:
[0040] Determine whether the defect level is greater than a preset level threshold; if so, generate an alarm message and trigger an alarm mechanism; and send the alarm message to a user terminal via a preset communication tool interface so that an operation and maintenance personnel of the user terminal can perform maintenance on the corresponding target concrete structure based on the alarm message;
[0041] Log data is generated based on the alarm information of the target concrete structure, and the log data is stored in the cloud database.
[0042] According to a second aspect of the present invention, there is provided a device for identifying surface defects of concrete structures based on an inspection robot, comprising:
[0043] An image acquisition unit is configured to respond to an apparent defect recognition instruction of a target concrete structure and acquire an apparent image of the target concrete structure in real time using a depth camera device carried by the inspection robot;
[0044] An image processing unit, configured to perform image quality enhancement processing on the apparent image to obtain the processed apparent image;
[0045] an acquisition unit, configured to acquire a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, wherein the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a concrete structure apparent defect recognition dataset;
[0046] A defect recognition unit is used to input the processed appearance image into the defect recognition model, perform global image feature extraction and context image feature extraction on the appearance image through each Transformer module in the backbone network, merge the global image features and context image features extracted by each Transformer module through the SPPF module to obtain merged image features corresponding to each Transformer module, perform key feature enhancement on the merged image features through the channel attention layer and spatial attention layer of each CBAM module in the neck network, perform feature fusion on the enhanced merged image features to obtain fused image features corresponding to each CBAM module, perform defect recognition on the fused image features corresponding to each CBAM module through the head network, and obtain the defect recognition result output by each prediction head module, wherein the defect recognition result is reflected on the appearance image in the form of a labeled box with a defect type label.
[0047] According to a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for identifying surface defects of concrete structures based on an inspection robot is implemented.
[0048] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for identifying surface defects of concrete structures based on an inspection robot when executing the program.
[0049] According to the present invention, a method and device for identifying surface defects of concrete structures based on an inspection robot are provided. Compared with the current method of manually identifying defects in concrete structures, the present invention uses an inspection robot to collect surface images of the target concrete structure in real time and uses a defect recognition model based on YOLOv5 to identify defects in the surface images. This can improve the efficiency and accuracy of identifying surface defects in concrete structures and avoid the safety issues caused by manual inspections. At the same time, the embodiments of the present invention use a customized model architecture, including a backbone network consisting of a Transformer module and an SPPF module, a neck network consisting of a CBAM module, and a head network consisting of multiple prediction head modules. Training is completed using a dataset of surface images of concrete structures with different illumination intensities and image textures. This achieves defect recognition in complex environments for various images, particularly images with complex illumination intensities and complex textures, and improves the model's robustness to different lighting conditions and viewing angles. Considering the slender morphology of cracks, the model introduces a Transformer module to improve crack detection. The Transformer module's strong ability to capture long-range dependencies enables the model to fully learn contextual and global information about the crack area, thereby improving defect detection accuracy. In addition, the network also integrates CBAM (Convolutional Block Attention Module) in the feature fusion part, which can realize real-time detection of cracks. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0051] Figure 1 A flow chart of a method for identifying surface defects of concrete structures based on an inspection robot according to an embodiment of the present invention is shown;
[0052] Figure 2 A schematic diagram of the network structure of a defect recognition model based on YOLOv5 provided in an embodiment of the present invention is shown;
[0053] Figure 3 A schematic diagram of the structure of a C3TranS module provided in an embodiment of the present invention is shown;
[0054] Figure 4 A schematic diagram of the structure of a CBAM module provided by an embodiment of the present invention is shown;
[0055] Figure 5 A flow chart of another method for identifying surface defects of concrete structures based on an inspection robot provided by an embodiment of the present invention is shown;
[0056] Figure 6 A schematic structural diagram of a concrete structure surface defect recognition device based on an inspection robot provided by an embodiment of the present invention is shown;
[0057] Figure 7 A schematic structural diagram of another device for identifying surface defects of concrete structures based on an inspection robot according to an embodiment of the present invention is shown;
[0058] Figure 8 A schematic diagram of the physical structure of a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0059] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.
[0060] At present, the manual defect identification of concrete structure surface is time-consuming and labor-intensive. It is also affected by the subjective factors and work experience of the staff, which may lead to missed detection and misjudgment.
[0061] In order to solve the above problems, the embodiment of the present invention provides a method for identifying surface defects of concrete structures based on an inspection robot, such as Figure 1 As shown, the method includes:
[0062] 101. In response to an apparent defect recognition instruction for a target concrete structure, a depth camera device carried by an inspection robot is used to collect an apparent image of the target concrete structure in real time.
[0063] Among them, the surface defects include: honeycombs, holes, cracks, rough surfaces, exposed tendons, etc.; the depth camera device can be an image acquisition sensor device such as a depth camera.
[0064] In an embodiment of the present invention, a high-precision positioning system for the inspection robot can be constructed using the combined navigation technology of the Global Navigation Satellite System (GNSS) and the Inertial Navigation System (INS). In environments where the GPS signal is weak or absent, the INS provides high-frequency dynamic measurements, and the positioning data is optimized in combination with the Kalman filter algorithm to ensure stable navigation of the robot in complex construction site environments. The inspection robot is equipped with a lidar, a screen, an ultrasonic sensor, and a camera (such as the Intel RealSense D435i) to perceive the surrounding environment in real time, build a two-dimensional or three-dimensional map of the construction site, and plan the optimal inspection path for the inspection robot. For example, a SLAM (Simultaneous Localization and Mapping) algorithm is used to construct a two-dimensional or three-dimensional map of the construction site environment where the target concrete structure is located in real time, and simultaneously determine the position of the inspection robot itself. Based on the inertial measurement unit data, the SLAM positioning accuracy is optimized and the cumulative error in the lidar point cloud matching is reduced; based on the real-time map and sensor data, a dynamic path planning algorithm is used to achieve obstacle avoidance and generate the optimal inspection path. The robot then follows the optimal inspection route, capturing images of the target concrete structure's surface during its journey. The inspection robot is equipped with a depth camera that integrates a high-resolution RGB sensor (e.g., resolution up to 1920×1080 pixels, 30 frames per second), a pair of stereo depth sensors, and an infrared projector. The camera boasts a field of view (FOV) of 87°×58° (horizontal×vertical) and a global shutter function, making it suitable for image acquisition in dynamic scenes. The depth camera, working in conjunction with a navigation module based on SLAM technology through the inspection robot's control system, follows a pre-planned optimal inspection route. This route planning considers the three-dimensional spatial distribution of the target concrete structure and the distribution of obstacles on the construction site. The robot can travel along the optimal inspection route at a constant speed (e.g., 0.5 m / s), while the camera's sensors capture images in real time, producing high-quality images of the concrete structure's surface. To ensure data integrity, the system sets a minimum 60% overlap between adjacent images. Subsequent image stitching techniques generate a continuous panoramic view of the concrete surface. The acquisition targets include key components such as the target concrete structure's walls, beam bottoms, and columns, making it suitable for a variety of complex construction sites, including bridges, tunnels, and high-rise buildings. Furthermore, the depth camera's built-in inertial measurement unit (IMU, Bosch BMI055, 6 degrees of freedom) provides real-time motion data, further assisting with image stability and positioning accuracy in dynamic environments. This improves the acquisition quality of surface images and, in turn, the accuracy of identifying surface defects in the target concrete structure.
[0065] In another embodiment of the present invention, in order to improve the image acquisition quality and avoid the influence of complex lighting, it is necessary to control the exposure parameters of the depth camera device when acquiring the apparent image. Based on this, the method includes: using the light sensor to monitor the light intensity around the target concrete structure in real time, and determining the sensor exposure parameters of the depth camera device based on the light intensity, wherein the sensor exposure parameters include exposure time, gain, and infrared projector intensity; based on the sensor exposure parameters, adjusting the parameters of the depth camera device in real time, and using the depth camera device after parameter adjustment to acquire the apparent image of the target concrete structure.
[0066] Specifically, to adapt to the variable lighting conditions in the construction site environment where the target concrete structure is located (such as low light in a tunnel or strong reflections outdoors), the inspection robot in an embodiment of the present invention is equipped with an ambient light sensor. The light sensor monitors the light intensity of the surrounding environment in real time and transmits the monitoring data to the control module of the depth camera device. The control module dynamically adjusts the exposure parameters of the RGB sensor of the depth camera device according to the light intensity, including exposure time (T), gain (Gain), and the intensity of the infrared projector. For example, in a dim environment with a light intensity of less than 50 lux (such as inside a tunnel), the system extends the exposure time to 1 / 15 second, increases the gain to a medium level, and increases the output power of the infrared projector (the default power is approximately 150 mW) to improve the low-light performance of the depth camera device and ensure uniform brightness of the RGB image; in strong light conditions (such as direct sunlight with an intensity of >50,000 lux), the exposure time is shortened to 1 / 1000 second, the gain is reduced, and the infrared projector is turned off to avoid overexposure. This enables the capture of high-quality surface images, which undergo real-time preprocessing (such as gamma correction or histogram equalization) to further optimize contrast and detail clarity. The depth camera's global shutter design, combined with an adaptive exposure mechanism, enables high-quality image output even in day-night conditions and in highly reflective environments, providing a reliable data foundation for subsequent defect detection using the YOLOv5 model.
[0067] 102. Perform image quality enhancement processing on the apparent image to obtain a processed apparent image.
[0068] Among them, the inspection robot is also equipped with a laser radar. For the embodiment of the present invention, after the apparent image of the target concrete structure is collected, in order to further improve the image collection quality, it is necessary to perform quality enhancement processing on the apparent image. Based on this, step 102 specifically includes: using the laser radar to scan the distance between the inspection robot and the target concrete structure in real time, and generating high-density point cloud data based on the real-time scanning distance; determining the depth data between the image pixel points of the apparent image and the inspection robot, wherein the depth data includes the distance, viewing angle, and height difference between the image pixel points and the inspection robot; synchronizing the high-density point cloud data with the depth data, and The synchronized high-density point cloud data is fused with the synchronized depth data to obtain fused data; the relative posture relationship between the laser radar and the depth camera device is determined, and a transformation matrix is generated based on the relative posture relationship; based on the transformation matrix, the fused data is converted into a preset global coordinate system corresponding to the inspection robot to obtain the converted fused data, and based on the converted fused data, the apparent image is geometrically corrected to obtain the processed apparent image, wherein the geometric correction includes at least one of radial distortion correction, tangential distortion correction, perspective alignment correction, and height matching correction.
[0069] Specifically, a lidar (LiDAR) system acquires precise distance information between the inspection robot and the target concrete structure, generating high-density point cloud data. LiDAR calculates distance by emitting a laser beam and measuring its return time (ToF) or phase difference, using multi-line scanning (e.g., 16- or 32-line) for high-density data acquisition. A depth camera (e.g., an RGB-D camera) uses structured light, ToF, or binocular vision to provide depth information for each image pixel. This depth data not only includes distance but also requires calculating the spatial position of the pixel relative to the inspection robot (viewing angle, height difference). Pixel depth values are obtained through infrared structured light encoding and decoding. The camera's intrinsic parameter matrix is used to convert pixel coordinates into three-dimensional coordinates in the camera coordinate system. The angle with the inspection robot's central axis is calculated to obtain the viewing angle. Combined with the inspection robot's attitude angle provided by the IMU (Inertial Measurement Unit), the pixel height is projected onto the global coordinate system to obtain the height difference. A trigger signal is then used to synchronize the exposure times of the LiDAR and camera, and software-level synchronization is achieved through timestamp matching. Then, geometric features such as normal vectors and curvature are extracted from the high-density point cloud data, and visual features are extracted from the depth data of the apparent image. Correspondence is established based on the similarity of feature descriptors. The RANSAC algorithm is used to eliminate false matches, and finally a probabilistic fusion framework is used to fuse the high-density point cloud data and depth data according to the weighted feature confidence. Furthermore, the spatial position relationship between the lidar and the depth camera device (rotation matrix R and translation vector t) is determined through joint calibration, and the transformation matrix T is constructed to achieve data unification in different coordinate systems. The initial position of the inspection robot is taken as the origin, the X-axis is along the forward direction, and the Z-axis is vertically upward. The fused data is converted from the sensor coordinate system to the preset global coordinate system using the transformation matrix T (wherein the preset global coordinate system is set according to actual needs). Based on the converted fused data, the Brown-Conrady model is then used to correct radial and tangential lens distortion through polynomial fitting. Based on essential matrix decomposition, the homography matrix H is calculated to achieve epipolar alignment of the images, i.e., perspective alignment correction. A digital elevation model (DEM) is constructed, and pixel heights are normalized through bilinear interpolation to achieve height matching correction. By geometrically correcting the apparent image, the embodiments of the present invention can improve image quality, thereby increasing the accuracy of the model's processing of the apparent image.
[0070] 103. Obtain a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a dataset of concrete structure surface images with different lighting intensities and image textures.
[0071] in, Figure 2 The structure of the defect recognition model is shown in the figure. Backbone is the backbone network, Neck is the neck network, and Head is the head network. Backbone includes Conv module (convolution to feature module), C3 layer (convolution layer), Transformer module, and SPPF (Spatial Pyramid Pooling Fast) module; Neck includes Concat (Concatenation Module, splicing) module, Upsample (Upsample Module, upsampling) module, Conv module, C3 layer, CBAM (Convolutional Block Attention Module, combining channel attention and spatial attention) module; Head includes multiple Conv (convolution operation) modules. Figure 3 The schematic diagram of the C3TranS (CSP, Cross Stage Partial Network, Bottleneck, Transformers, a hybrid architecture combining local feature extraction and global information capture) module is shown, including the CBS (Convolution, Batch Normalization, ReLU, a combination of convolution layer, batch normalization layer and ReLU activation function) module, Transformer-block (processing input sequence and generating output sequence) module, Conv module, Concat module, and the working process of the Transformer module is: Embedded Patches (one-dimensional sequence representation generated after the apparent image is divided into blocks and embedded), LayerNom (normalization operation on the image after Embedded Patches), Mutil-Head Attention (parallel calculation of attention distribution on the normalized image), Dropout (prevention of model overfitting operation), MLP (Multilayer Perceptron, multi-layer perceptron, used to learn the nonlinear relationship between input and output). Figure 3 Where Q, K, and V are query vector, key vector, and value vector respectively. Figure 4 The CBAM module structure is shown, including the channel attention module.
[0072] For an embodiment of the present invention, in order to improve the defect recognition accuracy of a defect recognition model based on YOLOv5, it is first necessary to train and construct a defect recognition model. Based on this, the method includes: constructing an initial defect recognition model based on YOLOv5, and obtaining a concrete structure apparent defect recognition dataset, wherein the concrete structure apparent defect recognition dataset contains sample apparent images of sample concrete structures with different defect type labels under multiple lighting conditions; performing operations such as rotation, flipping, and brightness adjustment on the sample apparent images to obtain an expanded concrete structure apparent defect recognition dataset, and dividing the expanded concrete structure apparent defect recognition dataset into training data and test data; using the training data to train the initial defect recognition model, and using the test data to test the trained initial defect recognition model, and determining the trained initial defect recognition model that meets the test conditions as the defect recognition model.
[0073] Specifically, first, an initial object detection model based on YOLOv5 is constructed. The model structure can be as described above. Next, a dataset for identifying surface defects in concrete structures is obtained. This dataset contains images of concrete structures with various defect labels, different lighting conditions, and different textures, as well as images of concrete structures without defects. This dataset includes image files and corresponding annotated labels. The dataset is then divided into training and test data using a random or pre-set strategy. The initial object detection model is trained using the training data. During training, metrics such as loss and mean average prediction (MAP) are monitored to evaluate model performance. Training parameters such as the learning rate, optimizer, and regularization are adjusted as needed to optimize the training effect. The trained initial object detection model is then tested on the test data to evaluate its performance on unseen data. Metrics such as mAP, precision, and recall are calculated and recorded on the test set. If the model performance does not meet the requirements, the training phase can be returned to for further iterations or adjustments until the model performance meets the requirements. Therefore, the embodiment of the present invention completes the training of the model through the concrete structure surface image data set with different lighting intensities and image textures, so that the trained defect recognition model can identify defects in images under different lighting and textures, avoiding the influence of complex lighting and textures, thereby improving the defect recognition accuracy.
[0074] 104. The processed appearance image is input into the defect recognition model, and global image features and context image features are extracted from the appearance image through each Transformer module in the backbone network. The global image features and context image features extracted by each Transformer module are merged through the SPPF module to obtain the merged image features corresponding to each Transformer module. The key features of the merged image features are enhanced through the channel attention layer and the spatial attention layer of each CBAM module in the neck network, and the enhanced merged image features are feature fused to obtain the fused image features corresponding to each CBAM module. The fused image features corresponding to each CBAM module are used to perform defect recognition through the head network to obtain the defect recognition results output by each prediction head module, wherein the defect recognition results are reflected on the appearance image in the form of labeled boxes with defect type labels.
[0075] For the embodiment of the present invention, considering that concrete cracks usually extend over a large span, a Transformer module is introduced into the feature extraction network. First, the surface image is input into the backbone network, and a feature map with translation invariance is obtained through the convolution module in the backbone network. Then, the Transformer is used to obtain the global connection between the pixels of the feature map. Combining the advantages of the two, the feature expression ability of the crack target under a complex background is enhanced and the global information extraction ability of the feature map is strengthened. The C3 module in the YOLOv5 backbone network contains 3 convolutional layers, which uses a bottleneck structure and a 1x1 convolution layer to realize image feature extraction. The Transformer is introduced into the C3 module to construct the C3Trans structure. Each Transformer module consists of multiple sub-layers, and the main functions are realized by multi-head attention and a multi-layer perceptron composed of fully connected layers. The sub-layers are connected using a residual structure. Compared with the original convolution, the multi-head attention of the Transformer module can calculate the correlation between all features of the entire feature map to obtain the global information of the feature map and sufficient context information. Therefore, in complex scenarios, the C3Trans module has better feature extraction capabilities for the target, and uses the SPPF module at the end of the feature extraction network to ensure that the context information and global information of cracks of different scales are merged without losing any feature information, and obtain the merged image features corresponding to each Transformer module. The output of the SPPF module is input to the neck network. In the feature fusion part, in order to eliminate redundant feature information unrelated to the cracks, more attention resources are used for the target area that needs to be focused on. At the same time, CBAM is a lightweight module that can be integrated into the C architecture and can be trained in an end-to-end manner. Therefore, the CBAM attention mechanism is added to the feature fusion part of the YOLOv5 network, such as Figure 4As shown. For a given feature map, CBAM sequentially infers attention maps along two independent dimensions of channel and space, and collaboratively learns the key local detail information in the image. C The corresponding elements of the matrix are multiplied to obtain the feature map F' which can effectively reflect the key channel information of the feature. On the basis of channel weighting, the spatial feature information is adaptively weighted using the serial spatial attention mechanism, and F' is used as the input of the spatial attention module and combined with the spatial weight coefficient M S The corresponding elements of the matrix are multiplied to obtain the feature map F″ containing channel position information and spatial position information. The process is shown in the following formula:
[0076]
[0077] in, represents the multiplication of the corresponding elements of the two matrices. Furthermore, the feature map F″ (the enhanced merged image features) is fused. Finally, the prediction head modules in the head network are used to identify defects based on the corresponding fused features, resulting in a labeled box with the defect type on the apparent image. This allows the defect type and location on the apparent image to be identified.
[0078] According to the present invention, a method for identifying surface defects of concrete structures based on an inspection robot is provided. Compared with the current method of manually identifying defects in concrete structures, the present invention uses an inspection robot to collect surface images of the target concrete structure in real time and uses a defect recognition model based on YOLOv5 to identify defects in the surface images. This method can improve the efficiency and accuracy of identifying surface defects of concrete structures and avoid the safety issues caused by manual inspections. At the same time, the embodiment of the present invention adopts a special model architecture, including a backbone network consisting of a Transformer module and an SPPF module, a neck network consisting of a CBAM module, and a head network consisting of multiple prediction head modules. It is trained using a dataset of surface images of concrete structures with different illumination intensities and image textures. This method realizes defect recognition in complex environments for various images, especially images with complex illumination intensities and complex textures, and improves the model's robustness to different lighting conditions and viewing angles. Considering the slender morphology of cracks, the model introduces a Transformer module to improve crack detection. The Transformer module's strong ability to capture long-range dependencies enables the model to fully learn contextual and global information about the crack area, thereby improving defect detection accuracy. In addition, the network also integrates CBAM (Convolutional Block Attention Module) in the feature fusion part, which can realize real-time detection of cracks.
[0079] Furthermore, in order to better illustrate the above process of classifying data, as a refinement and extension of the above embodiment, the embodiment of the present invention provides another method for identifying surface defects of concrete structures based on an inspection robot, such as Figure 5 As shown, the method includes:
[0080] 201. In response to an apparent defect recognition instruction for a target concrete structure, use a depth camera device carried by an inspection robot to collect an apparent image of the target concrete structure in real time.
[0081] The inspection robot is equipped with a depth camera, laser radar, a screen, and various sensors (such as ultrasonic sensors). The depth camera captures images of the concrete structure, while the laser radar and various sensors plan the inspection path. The screen displays inspection results.
[0082] For the embodiment of the present invention, in order to accurately collect the surface image of the concrete structure, it is first necessary to reasonably plan the inspection path for the inspection robot. Based on this, the method includes: using multi-source sensors to perceive the regional environmental data of the area where the target concrete structure is located in real time, and obtaining the lighting information of the area where the target concrete structure is located, the surface reflection characteristic information of the target concrete structure, and the shooting requirement information in real time; fusing the regional environmental data perceived by the multi-source sensors to obtain fused environmental data; based on the fused environmental data, constructing a three-dimensional point cloud map of the area where the target concrete structure is located in real time; based on the three-dimensional point cloud map, dynamically generating an accessibility topology map of the inspection robot, and based on the lighting information, the surface reflection characteristic information, and the shooting requirement information, determining candidate shooting viewpoints, and determining multiple candidate inspection paths where the candidate shooting viewpoints are located in the accessibility topology map; based on the terrain data of the area where each candidate inspection path is located, According to the above, the terrain slope and land use type of each candidate inspection path, as well as the slope resistance corresponding to the terrain slope and the land resistance corresponding to the land use type are determined respectively; the relative importance ratio of the terrain slope and the land use type to the inspection resistance is determined, and based on the relative importance ratio, a judgment matrix of the terrain slope and the land use type to the inspection resistance is constructed, and based on the judgment matrix, the resistance weight coefficients of the terrain slope and the land use type are determined respectively; based on the resistance weight coefficient, the slope resistance and the land resistance are weightedly summed to obtain the comprehensive recommendation coefficient corresponding to each candidate inspection path, and based on the comprehensive recommendation coefficient, the inspection path of the inspection robot is determined in each candidate inspection path; the inspection robot is controlled to travel along the inspection path, and the surface image of the target concrete structure is collected in real time by using the depth camera device carried by the inspection robot during the driving process.
[0083] Among them, multi-source sensors can include optical sensors, acoustic sensors, visual sensors, etc.; regional environmental data includes road condition information, obstacle information, three-dimensional size information of the area where the target concrete is located, etc.: lighting information refers to the light intensity around the target concrete structure; surface reflection characteristic information refers to the physical properties of the concrete surface when it is illuminated (such as reflected light intensity, reflected light direction, spectral reflection characteristics, etc.); shooting demand information refers to information such as the shooting location and shooting angle of the target concrete structure; land use types include: industrial land, residential land, mountainous and hilly areas, coastal land, agricultural land, etc.
[0084] Specifically, the regional environmental data collected by multi-source sensors is first normalized, processed for outliers, and aligned for features. The processed regional environmental data corresponding to each sensor is then fused, such as through weighted fusion, to obtain fused environmental data. Based on the fused environmental data, i.e., the three-dimensional coordinate information of the region, the iterative nearest point algorithm is used to align consecutive frames, and then a three-dimensional point cloud map of the region where the target concrete structure is located is determined. Furthermore, based on the three-dimensional point cloud map, multiple shooting nodes are determined. If a collision-free path exists between any two shooting nodes, edges are connected, with edge weights reflecting path length or energy consumption. Thus, an accessibility topology map can be generated based on each node and the edges between them. Then, candidate shooting viewpoints that meet the lighting intensity requirements, the concrete structure surface light reflection intensity requirements, and the shooting requirements are found in the accessibility topology map. The paths where the candidate shooting viewpoints are located are determined as candidate inspection paths. The terrain slope of each candidate inspection route is then calculated using digital elevation model data, and the land use type is determined. The corresponding slope resistance is then determined based on the terrain slope, and the corresponding land resistance is determined based on the land use type. For example, the greater the slope, the greater the corresponding resistance, such as clay resistance being greater than sand resistance. The relative importance ratio of terrain slope to land use type is determined using multiple expert ratings or historical data (e.g., slope:land = 3:1). For example, 5-10 domain experts are organized to conduct independent ratings. Based on their experience, the experts determine the impact of terrain slope and land use type on inspection resistance. If the experts believe that the impact of terrain slope on resistance is significantly greater than that of land use type (scale 5), the ratio of land use type to terrain slope is 1 / 5. Based on the expert evaluation results, a judgment matrix is constructed, as shown below:
[0085]
[0086] Among them, a 12 is the relative importance ratio of terrain slope and land use type to inspection resistance, a 21 =1 / a 12. Further, the maximum eigenvalue of the judgment matrix and the eigenvector corresponding to the maximum eigenvalue are calculated, and the eigenvector is normalized to obtain a weight vector, that is, the resistance weight coefficient of the terrain slope and the land use type is obtained. Then, the slope resistance and the land resistance are weighted and summed according to the resistance weight coefficient to obtain the comprehensive recommendation coefficient corresponding to each candidate inspection path, and finally the path corresponding to the minimum comprehensive recommendation coefficient is selected as the inspection path of the inspection robot. Then, the inspection robot is controlled to move according to the inspection path, and the depth camera device is controlled to collect the surface image of the target concrete structure during the movement. The embodiment of the present invention determines the optimal inspection path by comprehensively considering various information such as path resistance, shooting requirements, and lighting information. It can reduce the movement time and energy consumption of the inspection robot by optimizing path planning, thereby improving shooting efficiency; comprehensively consider lighting information to ensure shooting under optimal lighting conditions to improve image quality; and reduce equipment wear and failure by avoiding complex terrain and harsh environment, thereby extending equipment life.
[0087] 202. Perform image quality enhancement processing on the apparent image to obtain the processed apparent image.
[0088] Specifically, the inspection robot is additionally equipped with a laser radar (LiDAR). For example, the operating frequency of the LiDAR is not less than 10Hz, the ranging accuracy is ±2cm, and the scanning range covers 360° horizontal angle and ±15° vertical angle. The LiDAR scans the distance between the robot and the concrete structure in real time, generates high-density point cloud data, and synchronizes it with the apparent image captured by the depth camera (for example, the synchronization error can be set to <10ms). The depth sensor of the depth camera itself can provide a depth map with a resolution of up to 1280×720, with a measurement range from 0.11m to 10m. The distance (D), viewing angle deviation (θ) and height difference (H) between the pixel point of each frame image and the robot are calculated through a stereo vision algorithm. The high-density point cloud data of the LiDAR is fused with the depth data of the depth camera and mapped to the global map coordinate system of the inspection robot through a coordinate transformation algorithm (such as a homogeneous transformation matrix) to ensure the geometric consistency of the image data with the concrete structure. For example, when an inspection robot inspects the bottom of a concrete beam, if the depth camera measures a depth of 1.5 meters and the lidar calibration angle θ = 3°, the system can correct image distortion and accurately mark the spatial location of the defect. This multi-sensor fusion method effectively eliminates image distortion caused by varying perspectives or distances, providing high-precision input for subsequent defect identification.
[0089] 203. Obtain a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a dataset of concrete structure surface images with different lighting intensities and image textures.
[0090] 204. The processed appearance image is input into the defect recognition model, and global image features and context image features are extracted from the appearance image through each Transformer module in the backbone network. The global image features and context image features extracted by each Transformer module are merged through the SPPF module to obtain the merged image features corresponding to each Transformer module. The key features of the merged image features are enhanced through the channel attention layer and spatial attention layer of each CBAM module in the neck network, and the enhanced merged image features are feature fused to obtain the fused image features corresponding to each CBAM module. The fused image features corresponding to each CBAM module are used to perform defect recognition through the head network to obtain the defect recognition results output by each prediction head module, wherein the defect recognition results are reflected on the appearance image in the form of labeled boxes with defect type labels.
[0091] Specifically, the original YOLOv5 network was modified to obtain a defect detection model. The surface image was input into the defect detection model. The Transformer module in the backbone network extracted global and contextual information from the surface image, generating global and contextual features. The SPFF module then merged these features and fed them into the CBAM in the neck network for feature fusion. Finally, the fused features were fed into the head network. At the output, multiple prediction heads of different scales were used to perform bounding box regression and classification on cracks, adapting to cracks of varying sizes. For example, the output feature maps of the prediction heads were sized 80x80x18, 40x40x18, and 20x20x18. The channel dimension contained the predicted bounding box information: the predicted category, bounding box coordinates, and confidence value. Furthermore, the prediction heads performed non-maximum suppression based on the confidence values to produce the final prediction. The network can choose between loss functions such as regression loss and cross-entropy loss for classification.
[0092] 205. Based on the defect identification results, the defect geometric features are extracted in the target concrete structure.
[0093] Based on the geometric characteristics of the defects, the target concrete structure is classified into defect levels using a preset level classification configuration table to obtain the defect level of the target concrete structure.
[0094] Among them, the preset grading configuration table stores various defect combination sizes and their corresponding defect grades. In the embodiment of the present invention, after the defect is identified in the target concrete structure, the geometric features of the defect, such as length (L), width (W), and depth (D), are measured using an edge detection algorithm. For example, the crack width W is converted into the actual size by combining pixel-level edge spacing with lidar ranging data. Then, the defect grade corresponding to the geometric feature is determined according to the preset grading configuration table. For example, the grading criteria are as follows:
[0095] Level 1 defect (excellent): no obvious cracks or only very fine cracks (crack width less than 0.1mm);
[0096] Second level defect (good): a small number of fine cracks (crack width between 0.1mm and 0.3mm);
[0097] Level 3 defect (general): There are obvious cracks (the width of the crack is between 0.3mm and 0.5mm);
[0098] Level 4 defect (poor): There are many wide cracks (crack width greater than 0.5mm);
[0099] Level 5 defect (dangerous): There are a large number of wide cracks or cross cracks (crack width greater than 1mm).
[0100] Furthermore, based on the three-dimensional point cloud data of the target concrete structure, the defect position corresponding to the defect in the target concrete structure is determined, the defect geometric features, defect levels, and defect positions corresponding to the defects in the target concrete structure are associated, and a defect distribution map is generated based on the associated data; the defect distribution map is uploaded to the cloud database for distributed storage, and the defect distribution map is transmitted to the user terminal for visualization of preset dimensions. Among them, the preset dimensions can be set according to actual needs. Specifically, the defect recognition results and defect levels are associated with the positioning data generated by the inspection robot, and the geometric features and severity levels of each defect are bound to its real-time spatial coordinates (x, y, z). The positioning data can be derived from the SLAM map, fusing the IMU (Inertial Measurement Unit) information of the lidar point cloud and the depth camera device to ensure the correspondence between the defect position and the global environment of the construction site. The system integrates these data into a defect distribution map and displays it in a two-dimensional or three-dimensional visual form. The two-dimensional map uses a planar projection to mark the location of defects and uses color coding to indicate severity (for example, green for good, yellow for fair, and red for dangerous). The three-dimensional map combines point cloud data to display the spatial distribution of defects on concrete components, supporting perspective rotation and zooming for detailed inspection. The generated defect distribution map is uploaded to a cloud database via an efficient communication module (such as a 5G network). The database uses a distributed storage architecture (such as MongoDB or PostgreSQL) to support efficient management and query of large-scale data. Each defect record contains the following fields: defect type, classification result, geometric features, spatial coordinates, acquisition time (for example, accurate to the second, such as 2025-03-1714:35:22), image evidence (JPEG format), etc. Cloud storage supports multi-user access. Construction managers can view the map in real time through a web interface or mobile application to analyze the defect distribution pattern. The maintenance team can compare defect change trends based on historical data and formulate repair plans.
[0101] In another embodiment of the present invention, the inspection robot is equipped with an efficient communication module (such as 5G or Wi-Fi) to transmit the defect identification results to the remote monitoring terminal in real time. The remote terminal includes a portable device of the construction management personnel (such as a smartphone, tablet computer) and a server of the monitoring center. The terminal receives data through a dedicated application or a web interface. The application uses a real-time streaming protocol to communicate with the robot and supports low-latency updates of the data stream. The received defect distribution map is displayed in a visual form, and users can view the specific defect location and detailed information through interactive operations (such as zooming and rotating). The system supports multi-user simultaneous access, and the monitoring center can process data streams from multiple inspection robots at the same time, which is convenient for centralized management of large-scale construction sites.
[0102] In another embodiment of the present invention, after the surface defects of the target concrete are graded, in order to ensure the rapid maintenance of high-grade defects, it is also necessary to trigger an alarm mechanism. Based on this, the method includes: determining whether the defect level is greater than a preset level threshold; if so, generating an alarm message and triggering an alarm mechanism; sending the alarm message to the user terminal through a preset communication tool interface, so that the operation and maintenance personnel of the user terminal can perform maintenance on the corresponding target concrete structure based on the alarm information; generating log data based on the alarm information of the target concrete structure, and storing the log data in the cloud database.
[0103] The preset level threshold is set based on actual needs. Specifically, when a severe defect is identified (e.g., a crack width W > 1mm, meaning the defect level exceeds the preset level threshold), the system immediately triggers an alarm mechanism. Alarm mechanisms include: Local alarm: For example, the robot's built-in buzzer emits an intermittent alarm, and the status indicator switches to flashing red, alerting nearby construction workers. Remote notification: High-priority alarm information is sent to a remote terminal via the communication module. The message format includes the defect type, level, location, timestamp, and image evidence. Alarm information is sent to construction workers' mobile devices via push notifications, accompanied by vibration and an audible tone, ensuring timely notification. Logging: All alarm events are automatically recorded in a cloud database. Fields include the alarm trigger time, defect details, and handling status (initially "unhandled") for subsequent tracking and accountability. After receiving the notification, construction workers can view defect details through their terminals and take appropriate measures, such as suspending work, conducting on-site inspections, or arranging repairs. The system supports two-way communication, allowing managers to send commands to the robot through the terminal (e.g., pausing inspections or rescanning designated areas). This ensures timely handling of severe defects.
[0104] According to another method for identifying surface defects of concrete structures based on an inspection robot, compared to the current method of manually identifying defects on the surface of concrete structures, the present invention uses an inspection robot to collect surface images of the target concrete structure in real time and uses a defect recognition model based on YOLOv5 to identify defects on the surface images. This method can improve the recognition efficiency and accuracy of surface defects of concrete structures and avoid the safety issues caused by manual inspections. At the same time, the embodiment of the present invention adopts a special model architecture, including a backbone network consisting of a Transformer module and an SPPF module, a neck network consisting of a CBAM module, and a head network consisting of multiple prediction head modules. It is trained using a dataset of surface images of concrete structures with different light intensities and image textures. This method realizes defect recognition in complex environments for various images, especially images with complex light intensities and complex textures, and improves the model's robustness to different lighting conditions and viewing angles. Considering the slender morphology of cracks, the model introduces a Transformer module to improve crack detection. The Transformer module's strong ability to capture long-range dependencies enables the model to fully learn contextual and global information about the crack area, thereby improving defect detection accuracy. In addition, the network also integrates CBAM (Convolutional Block Attention Module) in the feature fusion part, which can realize real-time detection of cracks.
[0105] Further, as Figure 1 In a specific implementation, an embodiment of the present invention provides a concrete structure surface defect recognition device based on an inspection robot, such as Figure 6 As shown, the device includes: an image acquisition unit 31, an image processing unit 32, an acquisition unit 33, and a defect recognition unit 34.
[0106] The image acquisition unit 31 may be configured to respond to an apparent defect recognition instruction for a target concrete structure and utilize a depth camera device carried by the inspection robot to acquire an apparent image of the target concrete structure in real time.
[0107] The image processing unit 32 may be configured to perform image quality enhancement processing on the apparent image to obtain the processed apparent image.
[0108] The acquisition unit 33 can be used to obtain a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition. The backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales. The defect recognition model is pre-trained based on a concrete structure apparent defect recognition dataset.
[0109] The defect recognition unit 34 can be used to input the processed appearance image into the defect recognition model, perform global image feature extraction and context image feature extraction on the appearance image through each Transformer module in the backbone network, and merge the global image features and context image features extracted by each Transformer module through the SPPF module to obtain the merged image features corresponding to each Transformer module, perform key feature enhancement on the merged image features through the channel attention layer and spatial attention layer of each CBAM module in the neck network, and perform feature fusion on the enhanced merged image features to obtain the fused image features corresponding to each CBAM module, perform defect recognition on the fused image features corresponding to each CBAM module through the head network to obtain the defect recognition result output by each prediction head module, wherein the defect recognition result is reflected on the appearance image in the form of a labeled box with a defect type label.
[0110] In specific application scenarios, in order to collect the surface image of the target concrete structure, such as Figure 7 As shown, the image acquisition unit 31 includes an acquisition module 311 , a fusion module 312 , a construction module 313 , a first determination module 314 , and an acquisition module 315 .
[0111] The acquisition module 311 can be used to use multi-source sensors to perceive the regional environmental data of the area where the target concrete structure is located in real time, and to obtain the lighting information of the area where the target concrete structure is located, the surface reflection characteristic information of the target concrete structure, and the shooting requirement information in real time.
[0112] The fusion module 312 may be configured to perform fusion processing on the regional environment data sensed by the multi-source sensors to obtain fused environment data.
[0113] The construction module 313 can be used to construct a three-dimensional point cloud map of the area where the target concrete structure is located in real time based on the fused environmental data.
[0114] The first determination module 314 can be used to dynamically generate a reachability topology map of the inspection robot based on the three-dimensional point cloud map, and determine candidate shooting viewpoints based on the lighting information, the surface reflection feature information, and the shooting requirement information, and determine multiple candidate inspection paths where the candidate shooting viewpoints are located in the reachability topology map.
[0115] The first determination module 314 can also be used to determine the terrain slope and land use type of each candidate inspection path, as well as the slope resistance corresponding to the terrain slope and the land resistance corresponding to the land use type based on the terrain data of the area where each candidate inspection path is located.
[0116] The first determination module 314 can also be used to determine the relative importance ratio of the terrain slope and the land use type to the inspection resistance, construct a judgment matrix of the terrain slope and the land use type to the inspection resistance based on the relative importance ratio, and determine the resistance weight coefficients of the terrain slope and the land use type respectively based on the judgment matrix.
[0117] The first determination module 314 can also be used to perform weighted summation of the slope resistance and the land resistance based on the resistance weight coefficient to obtain a comprehensive recommendation coefficient corresponding to each candidate inspection path, and determine the inspection path of the inspection robot in each candidate inspection path based on the comprehensive recommendation coefficient.
[0118] The acquisition module 315 may be used to control the inspection robot to travel along the inspection path, and utilize the depth camera device carried by the inspection robot during the driving process to acquire the surface image of the target concrete structure in real time.
[0119] In a specific application scenario, the inspection robot is also equipped with a laser radar; in order to enhance the image quality of the apparent image, the image processing unit 32 includes a scanning module 321, a second determination module 322, a synchronization module 323, a generation module 324, and a correction module 325.
[0120] The scanning module 321 can be used to use the laser radar to scan the distance between the inspection robot and the target concrete structure in real time, and generate high-density point cloud data based on the real-time scanned distance.
[0121] The second determination module 322 can be used to determine the depth data between the image pixel point of the apparent image and the inspection robot, wherein the depth data includes the distance, viewing angle, and height difference between the image pixel point and the inspection robot.
[0122] The synchronization module 323 may be used to perform time synchronization on the high-density point cloud data and the depth data, and to perform data fusion on the synchronized high-density point cloud data and the synchronized depth data to obtain fused data.
[0123] The generation module 324 can be used to determine the relative posture relationship between the laser radar and the depth camera device, and generate a transformation matrix based on the relative posture relationship.
[0124] The correction module 325 can be used to convert the fused data into a preset global coordinate system corresponding to the inspection robot based on the transformation matrix to obtain the converted fused data, and perform geometric correction on the apparent image based on the converted fused data to obtain the processed apparent image, wherein the geometric correction includes at least one of radial distortion correction, tangential distortion correction, perspective alignment correction, and height matching correction.
[0125] In a specific application scenario, the inspection robot is also equipped with a light sensor; in order to collect the surface image of the target concrete structure, the first determination module 314 can also be used to use the light sensor to monitor the light intensity around the target concrete structure in real time, and determine the sensor exposure parameters of the depth camera device based on the light intensity, wherein the sensor exposure parameters include exposure time, gain, and infrared projector intensity.
[0126] The acquisition module 315 may also be configured to adjust parameters of the depth camera device in real time based on the sensor exposure parameters, and use the depth camera device after parameter adjustment to acquire the apparent image of the target concrete structure.
[0127] In a specific application scenario, in order to construct a defect recognition model, the device further includes: a construction unit 35 .
[0128] The construction unit 35 can be used to construct an initial defect recognition model based on YOLOv5 and obtain a concrete structure apparent defect recognition dataset, wherein the concrete structure apparent defect recognition dataset includes sample apparent images of sample concrete structures with different defect type labels under multiple lighting conditions; the sample apparent images are rotated, flipped, brightness adjusted, and other operations are performed on the sample apparent images to obtain an expanded concrete structure apparent defect recognition dataset, and the expanded concrete structure apparent defect recognition dataset is divided into training data and test data; the initial defect recognition model is trained using the training data, and the trained initial defect recognition model is tested using the test data, and the trained initial defect recognition model that meets the test conditions is determined as the defect recognition model.
[0129] In a specific application scenario, in order to post-process the defect identification result, the device further includes a post-processing unit 36.
[0130] The post-processing unit 36 can be used to extract defect geometric features in the target concrete structure based on the defect identification result; based on the defect geometric features, classify the target concrete structure into defect levels using a preset level classification configuration table to obtain the defect level of the target concrete structure; based on the three-dimensional point cloud data of the target concrete structure, determine the defect position corresponding to the defect in the target concrete structure, associate the defect geometric features, defect level, and defect position corresponding to the defect in the target concrete structure, and generate a defect distribution map based on the associated data; upload the defect distribution map to a cloud database for distributed storage, and transmit the defect distribution map to a user terminal for visual display of preset dimensions.
[0131] In a specific application scenario, in order to issue an alarm for serious defects, the post-processing unit 36 can also be used to determine whether the defect level is greater than a preset level threshold. If it is, an alarm message is generated and an alarm mechanism is triggered. The alarm message is sent to the user terminal through a preset communication tool interface so that the operation and maintenance personnel of the user terminal can repair the corresponding target concrete structure based on the alarm information; log data is generated based on the alarm information of the target concrete structure, and the log data is stored in the cloud database.
[0132] It should be noted that for other corresponding descriptions of the functional modules involved in the apparatus for identifying surface defects of concrete structures based on inspection robots provided in the embodiment of the present invention, reference can be made to Figure 1 The corresponding description of the method shown will not be repeated here.
[0133] Based on the above Figure 1The method shown, accordingly, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the following steps when executed by a processor: in response to an apparent defect recognition instruction of a target concrete structure, using a depth camera device carried by an inspection robot to collect an apparent image of the target concrete structure in real time; performing image quality enhancement processing on the apparent image to obtain the processed apparent image; obtaining a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a dataset of concrete structure apparent images with different light intensities and image textures; the processed The appearance image is input into the defect recognition model, and global image features and context image features are extracted from the appearance image through each Transformer module in the backbone network. The global image features and context image features extracted by each Transformer module are merged through the SPPF module to obtain the merged image features corresponding to each Transformer module. The key features of the merged image features are enhanced through the channel attention layer and the spatial attention layer of each CBAM module in the neck network, and the enhanced merged image features are feature fused to obtain the fused image features corresponding to each CBAM module. The fused image features corresponding to each CBAM module are used to perform defect recognition through the head network to obtain the defect recognition results output by each prediction head module, wherein the defect recognition results are reflected on the appearance image in the form of labeled boxes with defect type labels.
[0134] Based on the above Figure 1 The method shown and Figure 6 The embodiment of the device shown in the figure, the embodiment of the present invention also provides a physical structure diagram of a computer device, such as Figure 8As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor, wherein the memory 42 and the processor 41 are both arranged on a bus 43; when the processor 41 executes the program, the following steps are implemented: in response to an apparent defect recognition instruction of a target concrete structure, an apparent image of the target concrete structure is collected in real time by using a depth camera device carried on an inspection robot; image quality enhancement processing is performed on the apparent image to obtain the processed apparent image; a defect recognition model based on YOLOv5 is obtained, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-determined based on the apparent defects of concrete structures with different light intensities and image textures. The image data set completes training; the processed appearance image is input into the defect recognition model, and global image feature extraction and context image feature extraction are performed on the appearance image through each Transformer module in the backbone network, and the global image features and context image features extracted by each Transformer module are merged through the SPPF module to obtain the merged image features corresponding to each Transformer module, and the key features of the merged image features are enhanced through the channel attention layer and spatial attention layer of each CBAM module in the neck network, and the enhanced merged image features are feature fused to obtain the fused image features corresponding to each CBAM module, and the fused image features corresponding to each CBAM module are used to perform defect recognition on the head network to obtain the defect recognition result output by each prediction head module, wherein the defect recognition result is reflected on the appearance image in the form of a labeled box with a defect type label.
[0135] Through the technical solution of the present invention, the present invention uses an inspection robot to collect the surface images of the target concrete structure in real time, and uses a defect recognition model based on YOLOv5 to identify defects in the surface images. This can improve the recognition efficiency and accuracy of surface defects in concrete structures and avoid safety issues caused by manual inspections. At the same time, the embodiment of the present invention adopts a special model architecture, including a backbone network consisting of a Transformer module and an SPPF module, a neck network consisting of a CBAM module, and a head network consisting of multiple prediction head modules. It uses a dataset of concrete structure surface images with different light intensities and image textures to complete training. This achieves defect recognition for various images in complex environments, especially complex light intensity images and complex texture images, and improves the model's robustness to different lighting conditions and viewing angles. Considering the slender morphology of cracks, the model introduces a Transformer module to improve crack detection. The Transformer module's strong long-range dependency capture capability enables the model to fully learn the contextual information and global information of the crack area, thereby improving defect detection accuracy. In addition, the network also integrates CBAM (Convolutional Block Attention Module) in the feature fusion part, which can realize real-time detection of cracks.
[0136] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0137] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for identifying surface defects of concrete structures based on an inspection robot, characterized in that: include: In response to an apparent defect recognition instruction for a target concrete structure, a depth camera device carried by the inspection robot is used to collect an apparent image of the target concrete structure in real time; Performing image quality enhancement processing on the apparent image to obtain the processed apparent image; Obtain a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, wherein the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a dataset of concrete structure surface images with different illumination intensities and image textures; The processed appearance image is input into the defect recognition model, and global image features and context image features are extracted from the appearance image through each Transformer module in the backbone network. The global image features and context image features extracted by each Transformer module are merged through the SPPF module to obtain the merged image features corresponding to each Transformer module. The key features of the merged image features are enhanced through the channel attention layer and the spatial attention layer of each CBAM module in the neck network, and the enhanced merged image features are feature fused to obtain the fused image features corresponding to each CBAM module. The fused image features corresponding to each CBAM module are used to perform defect recognition through the head network to obtain the defect recognition results output by each prediction head module, wherein the defect recognition results are reflected on the appearance image in the form of a labeled box with a defect type label.
2. The method according to claim 1, characterized in that The method of collecting the surface image of the target concrete structure in real time by using the depth camera device carried by the inspection robot includes: Using multi-source sensors to perceive in real time regional environmental data of the area where the target concrete structure is located, and to obtain in real time illumination information of the area where the target concrete structure is located, surface reflection characteristic information of the target concrete structure, and shooting requirement information; fusing the regional environmental data sensed by the multi-source sensors to obtain fused environmental data; Based on the fused environmental data, a three-dimensional point cloud map of the area where the target concrete structure is located is constructed in real time; Based on the three-dimensional point cloud map, dynamically generate a reachability topology map of the inspection robot, and based on the lighting information, the surface reflection feature information, and the shooting requirement information, determine candidate shooting viewpoints, and determine multiple candidate inspection paths where the candidate shooting viewpoints are located in the reachability topology map; Based on the terrain data of the area where each candidate inspection path is located, determining the terrain slope and land use type of each candidate inspection path, as well as the slope resistance corresponding to the terrain slope and the land resistance corresponding to the land use type; Determining a relative importance ratio of the terrain slope and the land use type to the inspection resistance, constructing a judgment matrix of the terrain slope and the land use type to the inspection resistance based on the relative importance ratio, and determining resistance weight coefficients of the terrain slope and the land use type respectively based on the judgment matrix; Based on the resistance weight coefficient, the slope resistance and the land resistance are weighted and summed to obtain a comprehensive recommendation coefficient corresponding to each candidate inspection path, and based on the comprehensive recommendation coefficient, the inspection path of the inspection robot is determined in each candidate inspection path; The inspection robot is controlled to travel along the inspection path, and a depth camera device carried by the inspection robot is used to collect the surface image of the target concrete structure in real time during the traveling process.
3. The method according to claim 1, characterized in that The inspection robot is also equipped with a laser radar; The performing image quality enhancement processing on the apparent image to obtain the processed apparent image includes: Using the laser radar to scan the distance between the inspection robot and the target concrete structure in real time, and generating high-density point cloud data based on the real-time scanned distance; Determining depth data between an image pixel point of the apparent image and the inspection robot, wherein the depth data includes a distance, a viewing angle, and a height difference between the image pixel point and the inspection robot; Time-synchronizing the high-density point cloud data with the depth data, and fusing the synchronized high-density point cloud data with the synchronized depth data to obtain fused data; Determining a relative position relationship between the laser radar and the depth camera device, and generating a transformation matrix based on the relative position relationship; Based on the transformation matrix, the fused data is converted into a preset global coordinate system corresponding to the inspection robot to obtain the converted fused data, and based on the converted fused data, the apparent image is geometrically corrected to obtain the processed apparent image, wherein the geometric correction includes at least one of radial distortion correction, tangential distortion correction, perspective alignment correction, and height matching correction.
4. The method according to claim 1, wherein The inspection robot is also equipped with a light sensor; The method of collecting the surface image of the target concrete structure in real time by using the depth camera device carried by the inspection robot includes: Using the light sensor to monitor the light intensity around the target concrete structure in real time, and determining the sensor exposure parameters of the depth camera device based on the light intensity, wherein the sensor exposure parameters include exposure time, gain, and infrared projector intensity; Based on the sensor exposure parameters, the parameters of the depth camera device are adjusted in real time, and the depth camera device with the adjusted parameters is used to capture the apparent image of the target concrete structure.
5. The method according to claim 1, wherein Before obtaining the defect recognition model based on YOLOv5, the method further includes: Construct an initial defect recognition model based on YOLOv5 and obtain a concrete structure surface defect recognition dataset, wherein the concrete structure surface defect recognition dataset contains sample surface images of sample concrete structures with different defect type labels under various lighting conditions; Performing operations such as rotation, flipping, and brightness adjustment on the sample appearance images to obtain an expanded concrete structure appearance defect recognition dataset, and dividing the expanded concrete structure appearance defect recognition dataset into training data and test data; The initial defect recognition model is trained using the training data, and the trained initial defect recognition model is tested using the test data, and the trained initial defect recognition model that meets the test conditions is determined as the defect recognition model.
6. The method according to claim 1, characterized in that After performing defect recognition on the fused image features corresponding to each CBAM module through the head network to obtain a defect recognition result output by each prediction head module, the method further includes: extracting defect geometric features in the target concrete structure based on the defect identification result; Based on the defect geometric characteristics, the target concrete structure is classified into defect levels using a preset level classification configuration table to obtain the defect level of the target concrete structure; determining, based on the three-dimensional point cloud data of the target concrete structure, defect locations corresponding to defects in the target concrete structure, correlating defect geometric features, defect levels, and defect locations corresponding to the defects in the target concrete structure, and generating a defect distribution map based on the correlated data; The defect distribution map is uploaded to a cloud database for distributed storage, and the defect distribution map is transmitted to a user terminal for visual display of a preset dimension.
7. The method according to claim 6, characterized in that After classifying the target concrete structure by using the preset grade classification configuration table to obtain the defect grade of the target concrete structure, the method further includes: Determine whether the defect level is greater than a preset level threshold; if so, generate an alarm message and trigger an alarm mechanism; and send the alarm message to a user terminal via a preset communication tool interface so that an operation and maintenance personnel of the user terminal can perform maintenance on the corresponding target concrete structure based on the alarm message; Log data is generated based on the alarm information of the target concrete structure, and the log data is stored in the cloud database.
8. A device for identifying surface defects of concrete structures based on an inspection robot, characterized in that: include: An image acquisition unit is configured to respond to an apparent defect recognition instruction of a target concrete structure and acquire an apparent image of the target concrete structure in real time using a depth camera device carried by the inspection robot; An image processing unit, configured to perform image quality enhancement processing on the apparent image to obtain the processed apparent image; an acquisition unit, configured to acquire a defect recognition model based on YOLOv5, wherein the defect recognition model includes a backbone network for extracting image features, a neck network for fusing image features, and a head network for defect recognition, wherein the backbone network includes multiple Transformer modules and SPPF modules, the neck network includes multiple CBAM modules, and the head network includes multiple prediction head modules of different scales, wherein the defect recognition model is pre-trained based on a concrete structure apparent defect recognition dataset; A defect recognition unit is used to input the processed appearance image into the defect recognition model, perform global image feature extraction and context image feature extraction on the appearance image through each Transformer module in the backbone network, merge the global image features and context image features extracted by each Transformer module through the SPPF module to obtain merged image features corresponding to each Transformer module, perform key feature enhancement on the merged image features through the channel attention layer and spatial attention layer of each CBAM module in the neck network, perform feature fusion on the enhanced merged image features to obtain fused image features corresponding to each CBAM module, perform defect recognition on the fused image features corresponding to each CBAM module through the head network, and obtain the defect recognition result output by each prediction head module, wherein the defect recognition result is reflected on the appearance image in the form of a labeled box with a defect type label.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Power transmission line inspection image detection method based on deep convolutional neural network
CN117541535A
Method and device for robot to automatically inspect medical equipment
CN117637136A
Inspection robot intelligent scheduling management method based on multi-path inspection data
CN118211741A
Inspection route optimization method and device, storage medium, equipment and product
CN118295447A
Overhead line inspection method and device and storage medium
CN119576011A
Cited By
Subway maintenance workshop semantic map construction method based on improved YOLOv5
CN116740710A
A subway maintenance workshop semantic map construction method based on improved YOLOv5
CN116740710B
Photovoltaic module EL image intelligent defect identification method and system based on deep learning, storage medium and electronic terminal
CN121147231A
Concrete construction defect detection and repair system based on image recognition
CN121639616A
Image Recognition-Based Concrete Construction Defect Detection and Repair System
CN121639616B