Road surface disease real-time inspection method and system based on unmanned aerial vehicle
By combining the improved YOLOv5 and Vision Transformer algorithms with 3D reconstruction technology, the problems of low efficiency and insufficient accuracy in road damage detection were solved, real-time transmission and processing of damage information was achieved, and road maintenance efficiency was improved.
Patent Information
- Application Number
- CN202511157423.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-19
AI Technical Summary
The existing road defect detection technology has low efficiency, insufficient detection accuracy and cannot provide real-time feedback. It is difficult to adapt to complex lighting conditions and various types of defects. The three-dimensional reconstruction accuracy and efficiency are insufficient, and real-time inspections cannot be achieved.
The improved YOLOv5 algorithm and Vision Transformer algorithm are used to detect and classify disease targets, combined with a three-dimensional reconstruction algorithm to assess the severity of the disease, and the disease information is transmitted to the ground control center in real time via drones.
It improves the efficiency and accuracy of road damage inspections, realizes real-time transmission and processing of damage information, and supports rapid response and reasonable maintenance plans.
Smart Images

Figure CN120655646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road inspection, and in particular to a method and system for real-time inspection of road surface defects based on an unmanned aerial vehicle (UAV). Background Art
[0002] Currently, road defect detection primarily relies on manual inspections or on-board equipment. Manual inspections are inefficient, susceptible to subjective factors, and pose safety risks. While on-board equipment can improve efficiency, it is costly and less adaptable to complex road conditions. Furthermore, traditional methods struggle to detect and classify defects in real time, failing to provide accurate and timely information, leading to delays in repair decisions.
[0003] While existing drone-based road inspection methods can improve inspection efficiency, they still have shortcomings in image processing and disease identification. For example, existing methods use relatively limited technical means in image preprocessing, disease feature extraction, and classification and identification, and are unable to effectively handle complex lighting conditions, diverse terrain, and multiple disease types.
[0004] In recent years, deep learning technology has made significant progress in image recognition and processing. However, applying deep learning to road defect detection still faces numerous challenges. For example, how to achieve efficient image processing and defect identification within the limited computing resources of drones, and how to ensure the accuracy and real-time nature of detection results, remain pressing issues.
[0005] 3D reconstruction technology has important applications in road damage assessment, but existing techniques have limitations in terms of reconstruction accuracy and efficiency. For example, traditional 3D reconstruction methods are prone to errors when processing complex damage morphologies and are computationally complex, making them difficult to meet the needs of real-time inspections. Summary of the Invention
[0006] In order to solve the technical problems in the above background, the present invention aims to solve the problems of low inspection efficiency, insufficient detection accuracy and inability to provide real-time feedback in the prior art.
[0007] To achieve the above objectives, the present invention provides a real-time inspection method for road surface defects based on drones, comprising the following steps:
[0008] Use drones to collect and pre-process images of the road surface;
[0009] Use the improved YOLOv5 algorithm to detect disease targets in preprocessed images;
[0010] An improved convolutional neural network is used to extract features of detected disease targets;
[0011] The improved Vision Transformer algorithm is used to classify and identify the extracted disease features to determine the type of disease;
[0012] After the disease type is determined, the severity of the disease is assessed through a 3D reconstruction algorithm and a disease report is generated.
[0013] Preferably, the step of capturing an image of the road surface includes: monitoring the lighting conditions in real time by a light intensity sensor, and adjusting the exposure parameters of the camera using an intelligent exposure algorithm, wherein the exposure parameters of the camera are adjusted according to the following formula:
[0014] ,
[0015] Where E represents the exposure parameter of the camera, L represents the light intensity detected by the light intensity sensor, and a and b are determined coefficients used to adjust the exposure parameter according to different lighting conditions.
[0016] Preferably, the method for preprocessing the collected image includes: identifying and removing blurred images, overexposed images and underexposed images by calculating the gradient information and brightness distribution of the image; using a denoising algorithm based on wavelet transform to perform multi-scale denoising on the image; calculating the histogram of the image, generating a cumulative distribution function, and adjusting the pixel value of the image according to the cumulative distribution function; applying an image super-resolution reconstruction algorithm based on deep learning to generate an adversarial network architecture.
[0017] Preferably, when using the improved YOLOv5 algorithm to detect disease targets, the channel attention weight of the feature map is calculated by using the channel attention mechanism:
[0018] ,
[0019] in, Represents the channel attention weight of the feature map F, where F is the feature map, is the activation function, MLP is the multi-layer perceptron, avgpool and maxpool are global average pooling and global maximum pooling respectively;
[0020] Introduce the spatial attention mechanism and calculate the spatial attention weight of the feature map:
[0021] ,
[0022] in, Represents the spatial attention weight of the feature map F, conv is the convolution operation, and concat is the concatenation operation of the feature map;
[0023] Apply channel attention and spatial attention weights to the feature map to obtain the enhanced feature map:
[0024] ,
[0025] in, Represents the feature map after feature enhancement.
[0026] Preferably, the extracted damage features are classified by a hierarchical clustering algorithm, and the damage types include: cracks, potholes, looseness, rutting, oil flooding, spalling and settlement;
[0027] The meta-learning algorithm is used to dynamically learn disease characteristics. The specific formula is:
[0028] ,
[0029] in, represents the updated model parameters, represents the initial model parameters, is the adjustment parameter, Represents the model parameters The gradient, represents the task loss function, Represents a model.
[0030] Preferably, the method for assessing the severity of the disease by a three-dimensional reconstruction algorithm comprises:
[0031] ,
[0032] in, Represents a point in three-dimensional space, x is the horizontal coordinate on the image plane, y is the vertical coordinate on the image plane, z is the depth information of the point, f is the focal length of the camera, Depth map midpoint The depth value of .
[0033] Preferably, the method for generating the disease report includes: using wireless communication technology to transmit the detected disease information to the ground control center in real time; the ground control center receives and decodes the transmitted disease information, and processes the received disease information in real time, and the processing includes screening, classification and sorting; the disease report includes the disease type, location, size, severity, image fragments of the diseased area cropped from the original image, the unique number of the diseased area and the detection timestamp; the generated disease report is stored in the database of the ground control center and backed up.
[0034] The present invention also provides a real-time inspection system for road surface diseases based on drones, the system being used to implement the above method, comprising: an acquisition module, a detection module, an extraction module, an identification module and a generation module;
[0035] The acquisition module is used to use a drone to collect images of the road surface and perform preprocessing;
[0036] The detection module is used to detect disease targets on the preprocessed image using the improved YOLOv5 algorithm;
[0037] The extraction module is used to extract features of detected disease targets using an improved convolutional neural network;
[0038] The recognition module is used to classify and identify the extracted disease features using an improved Vision Transformer algorithm to determine the type of disease;
[0039] The generation module is used to evaluate the severity of the disease and generate a disease report through a three-dimensional reconstruction algorithm after the disease type is determined.
[0040] Preferably, the workflow for capturing images of the road surface includes: monitoring the lighting conditions in real time through a light intensity sensor, and adjusting the exposure parameters of the camera using an intelligent exposure algorithm, wherein the exposure parameters of the camera are adjusted according to the following formula:
[0041] ,
[0042] Where E represents the exposure parameter of the camera, L represents the light intensity detected by the light intensity sensor, and a and b are determined coefficients used to adjust the exposure parameter according to different lighting conditions.
[0043] Preferably, the process of preprocessing the image includes: identifying and removing blurred images, overexposed images and underexposed images by calculating the gradient information and brightness distribution of the image; using a denoising algorithm based on wavelet transform to perform multi-scale denoising on the image; calculating the histogram of the image, generating a cumulative distribution function, and adjusting the pixel value of the image according to the cumulative distribution function; applying an image super-resolution reconstruction algorithm based on deep learning to generate an adversarial network architecture.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention collects images through automatic flight of drones, combined with efficient image processing and disease recognition algorithms, significantly improving the efficiency of road disease inspections, covering large areas in a short time, and reducing labor and time costs.
[0046] Through the improved YOLOv5 algorithm and Vision Transformer algorithm, this invention can accurately detect and identify various types of diseases, improve the accuracy of disease detection, and the application of multi-scale feature extraction and deep learning technology makes the disease characteristics more obvious, further improving the accuracy of detection.
[0047] The present invention realizes the real-time transmission and processing of disease information. The ground control center can obtain disease information and generate disease reports in a timely manner, so that the road maintenance department can quickly respond to diseases and take repair measures in time, thereby improving the efficiency and response speed of road maintenance.
[0048] The present invention uses 3D reconstruction technology to not only measure the size of the disease, but also reconstruct the 3D morphology of the disease, comprehensively assess the severity of the disease, provide road maintenance departments with more comprehensive disease information, and help formulate more reasonable maintenance plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;
[0051] Figure 2 Schematic diagram of the convolution principle and process of an embodiment of the present invention; (a) is a schematic diagram of the convolution principle; (b) is a schematic diagram of the convolution process;
[0052] Figure 3 Schematic diagram of the purpose and principle of pooling according to an embodiment of the present invention; wherein (a) is a schematic diagram of the pooling principle; and (b) is a schematic diagram of the purpose of pooling. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] Example 1:
[0056] like Figure 1 FIG. 1 is a flow chart of the method of this embodiment, and the steps include:
[0057] S1. Use drones to collect and preprocess images of the road surface.
[0058] A drone equipped with a camera and positioning system captures images of the road surface along a pre-set route. Before takeoff, the drone's camera is calibrated, and image quality is tested using multi-point light intensity detection and color correction algorithms to reduce image distortion. The altitude and speed of the pre-set route are adjusted based on the road's topography and environmental conditions. The positioning system's terrain matching algorithm and environmental perception capabilities optimize the drone's flight path to achieve complete image acquisition.
[0059] During image acquisition, the light intensity sensor monitors the lighting conditions in real time, and the intelligent exposure algorithm is used to adjust the camera's exposure parameters to improve image clarity. The camera's exposure parameters are adjusted according to the following formula:
[0060] ,
[0061] Where E represents the exposure parameter of the camera, L represents the light intensity detected by the light intensity sensor, and a and b are determined coefficients used to adjust the exposure parameter according to different lighting conditions.
[0062] After image acquisition is complete, the collected images are screened and blurred, overexposed, and underexposed images are identified and removed by calculating image gradient information and brightness distribution. A wavelet-based denoising algorithm is used to perform multi-scale denoising on the images. An adaptive threshold selection algorithm preserves edge information while removing noise interference. The image histogram is calculated to generate a cumulative distribution function (CDF). The pixel values are then adjusted based on the CDF to enhance image contrast. A deep learning-based image super-resolution reconstruction algorithm, using a generative adversarial network architecture, improves image resolution and enhances detail.
[0063] S2. Use the improved YOLOv5 algorithm to detect disease targets in the preprocessed image.
[0064] Perform background segmentation on the preprocessed image, use the Gaussian mixture model to model the background pixels, and calculate the matching degree between each pixel and the background model. The specific formula is:
[0065] ,
[0066] in, is the matching degree of the background model of pixel point p, where p represents a pixel in the image and exp represents the exponential operation of the natural exponential function e. is the intensity value of pixel p, and are the mean and standard deviation of the background model.
[0067] Based on the matching degree, the pixels are classified into background and foreground areas, and the foreground areas are extracted as possible disease areas. The improved YOLOv5 algorithm is used to detect diseased targets, and the channel attention mechanism is introduced to calculate the channel attention weight of the feature map. The specific formula is:
[0068] ,
[0069] in, Represents the channel attention weight of the feature map F, where F is the feature map, is the activation function, MLP is a multi-layer perceptron, avgpool and maxpool are global average pooling and global maximum pooling respectively.
[0070] The spatial attention mechanism is introduced to calculate the spatial attention weight of the feature map. The specific formula is:
[0071] ,
[0072] in, Represents the spatial attention weight of the feature map F, conv is the convolution operation, and concat is the concatenation operation of the feature map.
[0073] Apply channel attention and spatial attention weights to the feature map to enhance the feature representation of key areas. The specific formula is:
[0074] ,
[0075] in, Represents the feature map after feature enhancement.
[0076] The improved YOLOv5 algorithm is used to perform target detection on the enhanced feature map, and the location, range and confidence of the diseased target are output.
[0077] S3. Use the improved convolutional neural network to extract features of the detected disease targets.
[0078] Convolution operations are performed on the feature maps of the diseased targets using convolution kernels of different scales to extract multi-scale features. These multi-scale feature maps are fused to generate a composite feature map. Combined with the Transformer architecture, the self-attention mechanism and multi-head attention module are used to analyze the features and extract long-range dependencies. The extracted features are then normalized using batch normalization and layer normalization techniques to adjust the mean and variance of the feature map, optimize the feature distribution, and achieve feature consistency.
[0079] The above improved convolutional neural network consists of convolutional layers and pooling layers. The specific structure includes:
[0080] like Figure 2 (a) Convolution refers to the use of a group of filters to move across an image to collect information from small areas. Here, the filters are called convolution kernels, the number of filters is called the convolution kernel number, and the size of the filters is called the convolution kernel size. The filters move in two directions, horizontally and vertically, and the interval between their movements is called the strides. The original size includes three parameters: length, width, and height. The height refers to 3 for color images and 1 for black and white images. After filtering and merging, higher-level information is obtained. Here, "filtering and merging" refers to passing through a convolution layer, and "obtaining higher-level information" refers to the output information being continuously compressed in length and width while increasing in height (height refers to the number of channels). For example, Figure 2 (b) Gain a deeper understanding of the image information.
[0081] like Figure 3 (a) Pooling refers to the unconscious information loss caused by each convolution. To solve this problem, the image length and width are not compressed during the convolution process, only the height is increased. A pooling layer is added after each convolution layer to filter the data and compress the length and width. The pooling layer ensures the integrity of the convolution layer information and the effective filtering of the input information, reducing the burden of network calculations and effectively improving the accuracy of the model. Figure 3 (b) The important parameters in the pooling layer are the pooling kernel size (k size )m×n means that each area with length m and width n is pooled once; the pooling types include maximum pooling, average pooling, etc. For example, maximum pooling refers to taking the maximum grayscale value of each pooled area as the representation; strides refers to the interval size of the pooling kernel function movement. Another parameter of the convolution and pooling layer is padding, which refers to the processing method of the image boundary. If padding is selected as SAME, it means that 0 is added to the insufficient part so that the input and output of the convolution layer can maintain the same size. Figure 3 (b) is a pooling layer with a pooling kernel size of 2×2 and a stride of 2×2 (both horizontally and vertically are 2). The pooling method is maximum pooling. You can see that the 4×4 matrix on the left becomes a 2×2 matrix after pooling, and each value after pooling is the maximum value of the corresponding pooling area in the original matrix.
[0082] S4. Use the improved Vision Transformer algorithm to classify and identify the extracted disease features and determine the type of disease.
[0083] The extracted damage features are classified and the types of damage are identified by a hierarchical clustering algorithm. The damage types include cracks, potholes, looseness, rutting, oil flooding, spalling and settlement.
[0084] Initialize the parameters of the meta-learning algorithm and use the meta-learning algorithm to dynamically learn the disease characteristics and adapt to the disease type. The specific formula is:
[0085] ,
[0086] in, represents the updated model parameters, represents the initial model parameters, is the adjustment parameter, Represents the model parameters The gradient, represents the task loss function, represents the model, and its parameters are .
[0087] Use the dynamic feature extraction module to adjust the feature extraction process according to the distribution of disease characteristics. The specific formula is:
[0088] ,
[0089] in, It represents the feature vector after dynamic feature extraction. Feature Extractor represents the feature extraction module, which is a neural network module used to extract useful features from input features. T represents the input feature vector. Represents the updated model parameters.
[0090] Use a dynamic weight adjustment mechanism to dynamically adjust the weights based on the distribution of disease characteristics. The specific formula is:
[0091] ,
[0092] in, represents the updated model parameters, represents the initial model parameters, is the adjustment parameter, Indicates the adjustment amount of the parameter.
[0093] Adaptive classifier adjustment algorithm is used to adjust the classifier weight according to the distribution of disease characteristics. The specific formula is:
[0094] ,
[0095] ,
[0096] in, represents the weight of the i-th disease, exp is the exponential operation of the natural exponential function e, Is an adjustment parameter used to control the influence of confidence on weight. Indicates disease characteristics The confidence score of Represents the extracted disease features, n represents the total number of disease categories, j is the index variable used to traverse all disease categories, and Classifier is the classifier function used to classify disease features.
[0097] S5. After the disease type is determined, the severity of the disease is evaluated through a 3D reconstruction algorithm and a disease report is generated.
[0098] A deep learning-based 3D reconstruction algorithm is applied to generate a depth map of the diseased area from the image, wherein the depth map represents the depth information of each pixel.
[0099] Generate point cloud data of the diseased area based on the depth map. Use the depth information in the depth map and combine it with the camera parameters to calculate the three-dimensional coordinates of each pixel. The specific formula is:
[0100] ,
[0101] in, Represents a point in three-dimensional space, x is the horizontal coordinate on the image plane, y is the vertical coordinate on the image plane, z is the depth information of the point, f is the focal length of the camera, Depth map midpoint The depth value of .
[0102] The point cloud data is converted into a voxel grid, and the Marching Cubes algorithm is used to extract isosurfaces from the voxel grid to generate a 3D model of the damage. The reconstructed 3D model is analyzed, and the geometric characteristics of the damaged area are measured using volume calculation and shape analysis algorithms to assess the severity of the damage. These geometric characteristics include volume, aspect ratio, and curvature. The severity of the damage is classified as mild, moderate, or severe.
[0103] Finally, wireless communication technology is used to transmit the detected disease information in real time to the ground control center. The ground control center receives and decodes the transmitted disease information and processes the received disease information in real time, including screening, classification, and sorting.
[0104] Based on the processing results, a damage report is generated, which includes the damage type, location, size, severity, an image segment of the damaged area cropped from the original image, a unique number of the damaged area, and a detection timestamp; the generated damage report is stored in the database of the ground control center and backed up.
[0105] Example 2:
[0106] This embodiment also provides a real-time inspection system for road surface defects based on a drone, comprising: a collection module, a detection module, an extraction module, a recognition module, and a generation module.
[0107] The following will describe in detail how the present invention solves technical problems in real life in conjunction with this embodiment.
[0108] First, the acquisition module uses drones to collect images of the road surface and perform preprocessing.
[0109] A drone equipped with a camera and positioning system captures images of the road surface along a pre-set route. Before takeoff, the drone's camera is calibrated, and image quality is tested using multi-point light intensity detection and color correction algorithms to reduce image distortion. The altitude and speed of the pre-set route are adjusted based on the road's topography and environmental conditions. The positioning system's terrain matching algorithm and environmental perception capabilities optimize the drone's flight path to achieve complete image acquisition.
[0110] During image acquisition, the light intensity sensor monitors the lighting conditions in real time, and the intelligent exposure algorithm is used to adjust the camera's exposure parameters to improve image clarity. The camera's exposure parameters are adjusted according to the following formula:
[0111] ,
[0112] Where E represents the exposure parameter of the camera, L represents the light intensity detected by the light intensity sensor, and a and b are determined coefficients used to adjust the exposure parameter according to different lighting conditions.
[0113] After image acquisition is complete, the collected images are screened and blurred, overexposed, and underexposed images are identified and removed by calculating image gradient information and brightness distribution. A wavelet-based denoising algorithm is used to perform multi-scale denoising on the images. An adaptive threshold selection algorithm preserves edge information while removing noise interference. The image histogram is calculated to generate a cumulative distribution function (CDF). The pixel values are then adjusted based on the CDF to enhance image contrast. A deep learning-based image super-resolution reconstruction algorithm, using a generative adversarial network architecture, improves image resolution and enhances detail.
[0114] The detection module uses the improved YOLOv5 algorithm to detect disease targets in the preprocessed images.
[0115] Perform background segmentation on the preprocessed image, use the Gaussian mixture model to model the background pixels, and calculate the matching degree between each pixel and the background model. The specific formula is:
[0116] ,
[0117] in, is the matching degree of the background model of pixel point p, where p represents a pixel in the image and exp represents the exponential operation of the natural exponential function e. is the intensity value of pixel p, and are the mean and standard deviation of the background model.
[0118] Based on the matching degree, the pixels are classified into background and foreground areas, and the foreground areas are extracted as possible disease areas. The improved YOLOv5 algorithm is used to detect diseased targets, and the channel attention mechanism is introduced to calculate the channel attention weight of the feature map. The specific formula is:
[0119] ,
[0120] in, Represents the channel attention weight of the feature map F, where F is the feature map, is the activation function, MLP is a multi-layer perceptron, avgpool and maxpool are global average pooling and global maximum pooling respectively.
[0121] The spatial attention mechanism is introduced to calculate the spatial attention weight of the feature map. The specific formula is:
[0122] ,
[0123] in, Represents the spatial attention weight of the feature map F, conv is the convolution operation, and concat is the concatenation operation of the feature map.
[0124] Apply channel attention and spatial attention weights to the feature map to enhance the feature representation of key areas. The specific formula is:
[0125] ,
[0126] in, Represents the feature map after feature enhancement.
[0127] The improved YOLOv5 algorithm is used to perform target detection on the enhanced feature map, and the location, range and confidence of the diseased target are output.
[0128] The extraction module then uses an improved convolutional neural network to extract features of the detected disease targets.
[0129] The improved convolutional neural network consists of a convolutional layer and a pooling layer, and its specific structure includes:
[0130] like Figure 2(a) Convolution refers to the use of a group of filters to move across an image to collect information from small areas. Here, the filters are called convolution kernels, the number of filters is called the convolution kernel number, and the size of the filters is called the convolution kernel size. The filters move in two directions, horizontally and vertically, and the interval between their movements is called the strides. The original size includes three parameters: length, width, and height. The height refers to 3 for color images and 1 for black and white images. After filtering and merging, higher-level information is obtained. Here, "filtering and merging" refers to passing through a convolution layer, and "obtaining higher-level information" refers to the output information being continuously compressed in length and width while increasing in height (height refers to the number of channels). For example, Figure 2 (b) Gain a deeper understanding of the image information.
[0131] like Figure 3 (a) Pooling refers to the unconscious information loss caused by each convolution. To solve this problem, the image length and width are not compressed during the convolution process, only the height is increased. A pooling layer is added after each convolution layer to filter the data and compress the length and width. The pooling layer ensures the integrity of the convolution layer information and the effective filtering of the input information, reducing the burden of network calculations and effectively improving the accuracy of the model. Figure 3 (b) The important parameters in the pooling layer are the pooling kernel size (k size ) m×n means that each area with length m and width n is pooled once; the pooling types include maximum pooling, average pooling, etc. For example, maximum pooling refers to taking the maximum grayscale value of each pooled area as the representation; strides refers to the interval size of the pooling kernel function movement. Another parameter of the convolution and pooling layer is padding, which refers to the processing method of the image boundary. If padding is selected as SAME, it means that 0 is added to the insufficient part so that the input and output of the convolution layer can maintain the same size. Figure 3 (b) is a pooling layer with a pooling kernel size of 2×2 and a stride of 2×2 (both horizontally and vertically are 2). The pooling method is maximum pooling. You can see that the 4×4 matrix on the left becomes a 2×2 matrix after pooling, and each value after pooling is the maximum value of the corresponding pooling area in the original matrix.
[0132] Convolution operations are performed on the feature maps of the diseased targets using convolution kernels of different scales to extract multi-scale features. These multi-scale feature maps are fused to generate a composite feature map. Combined with the Transformer architecture, the self-attention mechanism and multi-head attention module are used to analyze the features and extract long-range dependencies. The extracted features are then normalized using batch normalization and layer normalization techniques to adjust the mean and variance of the feature map, optimize the feature distribution, and achieve feature consistency.
[0133] The recognition module uses the improved Vision Transformer algorithm to classify and identify the extracted disease features and determine the type of disease.
[0134] The extracted damage features are classified and the types of damage are identified by a hierarchical clustering algorithm, wherein the types of damage include cracks, potholes, looseness, rutting, oil flooding, spalling and settlement;
[0135] Initialize the parameters of the meta-learning algorithm and use the meta-learning algorithm to dynamically learn the disease characteristics and adapt to the disease type. The specific formula is:
[0136] ,
[0137] in, represents the updated model parameters, represents the initial model parameters, is the adjustment parameter, Represents the model parameters The gradient, represents the task loss function, represents the model, and its parameters are .
[0138] Use the dynamic feature extraction module to adjust the feature extraction process according to the distribution of disease characteristics. The specific formula is:
[0139] ,
[0140] in, It represents the feature vector after dynamic feature extraction. Feature Extractor represents the feature extraction module, which is a neural network module used to extract useful features from input features. T represents the input feature vector. Represents the updated model parameters.
[0141] Use a dynamic weight adjustment mechanism to dynamically adjust the weights based on the distribution of disease characteristics. The specific formula is:
[0142] ,
[0143] in, represents the updated model parameters, represents the initial model parameters, is the adjustment parameter, Indicates the adjustment amount of the parameter.
[0144] Adaptive classifier adjustment algorithm is used to adjust the classifier weight according to the distribution of disease characteristics. The specific formula is:
[0145] ,
[0146] ,
[0147] in, represents the weight of the i-th disease, exp is the exponential operation of the natural exponential function e, Is an adjustment parameter used to control the influence of confidence on weight. Indicates disease characteristics The confidence score of Represents the extracted disease features, n represents the total number of disease categories, j is the index variable used to traverse all disease categories, and Classifier is the classifier function used to classify disease features.
[0148] After the disease type is determined, the final generation module uses a three-dimensional reconstruction algorithm to evaluate the severity of the disease and generate a disease report.
[0149] Applying a deep learning-based 3D reconstruction algorithm to generate a depth map of the diseased area from the image, wherein the depth map represents the depth information of each pixel;
[0150] Generate point cloud data of the diseased area based on the depth map. Use the depth information in the depth map and combine it with the camera parameters to calculate the three-dimensional coordinates of each pixel. The specific formula is:
[0151] ,
[0152] in, Represents a point in three-dimensional space, x is the horizontal coordinate on the image plane, y is the vertical coordinate on the image plane, z is the depth information of the point, f is the focal length of the camera, Depth map midpoint The depth value of .
[0153] The point cloud data is converted into a voxel grid, and the Marching Cubes algorithm is used to extract isosurfaces from the voxel grid to generate a 3D model of the damage. The reconstructed 3D model is analyzed, and the geometric characteristics of the damaged area are measured using volume calculation and shape analysis algorithms to assess the severity of the damage. These geometric characteristics include volume, aspect ratio, and curvature. The severity of the damage is classified as mild, moderate, or severe.
[0154] Finally, wireless communication technology is used to transmit the detected disease information in real time to the ground control center. The ground control center receives and decodes the transmitted disease information and processes the received disease information in real time, including screening, classification, and sorting.
[0155] Based on the processing results, a damage report is generated, which includes the damage type, location, size, severity, an image segment of the damaged area cropped from the original image, a unique number of the damaged area, and a detection timestamp; the generated damage report is stored in the database of the ground control center and backed up.
[0156] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A real-time inspection method for road surface defects based on drones, characterized by the following steps: include: Use drones to collect and pre-process images of the road surface; Use the improved YOLOv5 algorithm to detect disease targets in preprocessed images; An improved convolutional neural network is used to extract features of detected disease targets; The improved Vision Transformer algorithm is used to classify and identify the extracted disease features to determine the type of disease; After the disease type is determined, the severity of the disease is assessed through a 3D reconstruction algorithm and a disease report is generated.
2. The real-time inspection method for road surface defects based on drones according to claim 1 is characterized in that: The steps of capturing images of the road surface include: monitoring the lighting conditions in real time using a light intensity sensor, and adjusting the exposure parameters of the camera using an intelligent exposure algorithm. The exposure parameters of the camera are adjusted according to the following formula: , Where E represents the exposure parameter of the camera, L represents the light intensity detected by the light intensity sensor, and a and b are determined coefficients used to adjust the exposure parameter according to different lighting conditions.
3. The real-time inspection method for road surface defects based on drones according to claim 1 is characterized in that: The methods for preprocessing the collected images include: identifying and removing blurred images, overexposed images, and underexposed images by calculating the gradient information and brightness distribution of the image; using a denoising algorithm based on wavelet transform to perform multi-scale denoising on the image; calculating the histogram of the image, generating a cumulative distribution function, and adjusting the pixel value of the image according to the cumulative distribution function; applying an image super-resolution reconstruction algorithm based on deep learning to generate an adversarial network architecture.
4. The real-time inspection method for road surface defects based on drones according to claim 1 is characterized in that: When using the improved YOLOv5 algorithm to detect disease targets, the channel attention weight of the feature map is calculated through the channel attention mechanism: , in, Represents the channel attention weight of the feature map F, where F is the feature map, is the activation function, MLP is the multi-layer perceptron, avgpool and maxpool are global average pooling and global maximum pooling respectively; Introduce the spatial attention mechanism and calculate the spatial attention weight of the feature map: , in, Represents the spatial attention weight of the feature map F, conv is the convolution operation, and concat is the concatenation operation of the feature map; Apply channel attention and spatial attention weights to the feature map to obtain the enhanced feature map: , in, Represents the feature map after feature enhancement.
5. The real-time inspection method for road surface defects based on drones according to claim 1 is characterized in that: The extracted damage features are classified by a hierarchical clustering algorithm, and the damage types include: cracks, potholes, looseness, rutting, oil flooding, spalling and settlement; The meta-learning algorithm is used to dynamically learn disease characteristics. The specific formula is: , in, represents the updated model parameters, represents the initial model parameters, is the adjustment parameter, Represents the model parameters The gradient, represents the task loss function, Represents a model.
6. The real-time inspection method for road surface defects based on drones according to claim 1 is characterized in that: Methods for assessing disease severity through 3D reconstruction algorithms include: , in, Represents a point in three-dimensional space, x is the horizontal coordinate on the image plane, y is the vertical coordinate on the image plane, z is the depth information of the point, f is the focal length of the camera, Depth map midpoint The depth value of .
7. The real-time inspection method for road surface defects based on drones according to claim 1 is characterized in that: The method for generating the disease report includes: using wireless communication technology to transmit the detected disease information to a ground control center in real time; the ground control center receives and decodes the transmitted disease information, and processes the received disease information in real time, and the processing includes screening, classification and sorting; the disease report includes the disease type, location, size, severity, an image segment of the diseased area cropped from the original image, a unique number of the diseased area and a detection timestamp; the generated disease report is stored in a database of the ground control center and backed up.
8. A real-time inspection system for road surface defects based on drones, the system being used to implement the method according to any one of claims 1 to 7, characterized in that: include: Acquisition module, detection module, extraction module, recognition module and generation module; The acquisition module is used to use a drone to collect images of the road surface and perform preprocessing; The detection module is used to detect disease targets on the preprocessed image using the improved YOLOv5 algorithm; The extraction module is used to extract features of detected disease targets using an improved convolutional neural network; The recognition module is used to classify and identify the extracted disease features using an improved Vision Transformer algorithm to determine the type of disease; The generation module is used to evaluate the severity of the disease and generate a disease report through a three-dimensional reconstruction algorithm after the disease type is determined.
9. The real-time road surface disease inspection system based on drone according to claim 8 is characterized in that: The workflow for capturing images of the road surface includes real-time monitoring of lighting conditions using a light intensity sensor and adjusting the camera's exposure parameters using an intelligent exposure algorithm. The camera's exposure parameters are adjusted according to the following formula: , Where E represents the exposure parameter of the camera, L represents the light intensity detected by the light intensity sensor, and a and b are determined coefficients used to adjust the exposure parameter according to different lighting conditions.
10. The real-time road surface disease inspection system based on drone according to claim 8 is characterized in that: The image preprocessing process includes: identifying and removing blurred images, overexposed images, and underexposed images by calculating the image's gradient information and brightness distribution; performing multi-scale denoising on the image using a wavelet transform-based denoising algorithm; calculating the image's histogram, generating a cumulative distribution function, and adjusting the image's pixel values based on the cumulative distribution function; and applying a deep learning-based image super-resolution reconstruction algorithm to generate an adversarial network architecture.
Citation Information
Patent Citations
Method and equipment for detecting internal diseases of highway pavement structure
CN118396996A
Road defect identification method and system, storage medium, equipment and program
CN118865382A
Multi-modal small sample image classification method and system based on multi-scale dynamic feature fusion
CN119559435A
Pavement structure disease recognition model for poor interlayer bonding and training method
CN119963969A
Airport pavement underground structure disease automatic detection method based on deep learning
WO2022147969A1
Cited By
Highway traffic management method and system based on highway disease dynamic early warning
CN122090619A