Road disease image recognition and positioning detection method and system based on deep learning

Through improved lightweight neural network and multi-task loss function optimization, combined with super-resolution positioning module, the problems of low efficiency, insufficient accuracy and deployment difficulties of road disease detection in the prior art are solved, and efficient and low-cost subpixel-level disease detection is achieved.

CN120339981APending Publication Date: 2025-07-18INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353787.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing road disease detection methods rely on low efficiency, high cost and are susceptible to subjective factors. Traditional image processing methods are sensitive to light changes and complex backgrounds, have poor generalization capabilities, deep learning models have low detection accuracy in small targets, large parameters, and are difficult to deploy to edge devices, and insufficient positioning accuracy cannot meet road maintenance needs.

Method used

The improved lightweight neural network MobileNetV3 is adopted in conjunction with the CBAM attention module, and the optimization of data augmentation and multi-task loss function, combined with the super-resolution positioning module to realize sub-pixel-level detection on edge devices, including data acquisition, data augmentation, model construction, multi-task training and deployment optimization.

Benefits of technology

It realizes efficient detection and subpixel-level positioning of small target diseases, meeting the millimeter-level accuracy requirements for road maintenance, the model runs in real time on edge devices, with fast inference speed and low memory footprint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339981A_ABST
    Figure CN120339981A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and deep learning, in particular to a road disease image recognition and positioning detection method and system based on deep learning, and the method comprises the following steps: data collection, data enhancement, model construction, multi-task training, and deployment optimization. The method has the beneficial effects that multi-scale features are extracted through an improved lightweight network architecture (MobileNetV3 in combination with a CBAM attention module), and the detection capability of small target diseases (such as cracks and net cracks) is enhanced; meanwhile, a super-resolution positioning module (SRGAN) is introduced, up-sampling is carried out on a low-resolution feature map by four times, sub-pixel-level coordinate regression is achieved, and the millimeter-level positioning requirement of road maintenance is met. The technical scheme comprises data enhancement, multi-task loss function optimization and edge device deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and deep learning, and specifically provides a method and system for road disease image recognition and location detection based on deep learning. Background Technique

[0002] The existing road disease detection methods have the following technical bottlenecks: Manual inspection: relying on manpower, with low efficiency, high cost, and being easily affected by subjective factors. Traditional image processing: based on threshold segmentation or edge detection, sensitive to light changes and complex backgrounds, and with poor generalization ability.

[0003] The existing deep learning models, such as Faster R-CNN or YOLO, have the following problems in road disease detection:

[0004] Low detection accuracy for small targets (such as slender cracks);

[0005] Large number of model parameters, making it difficult to deploy to edge devices;

[0006] Insufficient location accuracy, unable to meet the millimeter-level requirements of road maintenance. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for road disease image recognition and location detection based on deep learning to solve the problems raised in the above background technique.

[0008] To achieve the above purpose, the present invention provides the following technical solution: A method for road disease image recognition and location detection based on deep learning, including the following steps:

[0009] Data collection: Collect road images through an in-vehicle camera with a resolution of not less than 1920×1080 pixels;

[0010] Data augmentation: Perform random rotation, brightness adjustment, and noise injection on the collected images to simulate rainy days or shadow environments and improve the robustness of the model;

[0011] Model construction: Design an improved lightweight neural network. The backbone network adopts the depthwise separable convolution structure of MobileNetV3 to reduce the number of parameters. The detection head integrates channel and spatial attention mechanisms to enhance the small target feature extraction ability. The location branch combines a super-resolution module to achieve sub-pixel coordinate regression;

[0012] Multi-task training: Use the classification loss function FocalLoss to solve the problem of class imbalance, and combine the location loss functions SmoothL1Loss and IoULoss to improve the accuracy of the bounding box;

[0013] Deployment Optimization: Through model quantization and TensorRT acceleration technology, achieve an inference speed ≥ 30 FPS and a memory occupancy < 500 MB on the NVIDIA Jetson Xavier edge device.

[0014] Preferably, the data augmentation strategy specifically includes:

[0015] Randomly rotate the image within the range of ±15°;

[0016] Adjust the brightness to 50% - 150% of the original image;

[0017] Inject Gaussian noise or simulate raindrop texture, and the noise intensity is dynamically adjusted according to the preset environmental parameters.

[0018] Preferably, the improved lightweight neural network further includes: generating channel attention weights through global average pooling and max pooling in the channel dimension, generating spatial attention weights through convolution in the spatial dimension, dynamically weighting the feature map to enhance the response in the crack area; upsampling the low-resolution feature map by 4 times through SRGAN, outputting the coordinates of the disease center point, and the positioning accuracy reaches ±1 mm.

[0019] Preferably, the multi-task loss function is achieved through weighted fusion, and the specific form is:

[0020] Classification loss: Adopt Focal Loss, and the adjustment factor α ∈ [0.25, 0.75] to balance positive and negative samples;

[0021] Localization loss: Combine SmoothL1Loss and IoULoss, and the weight coefficients are λ1 = 0.5 and λ2 = 0.5 respectively. The total loss function is: L = λ12L Focal +λ22(L SmoothL1 +L IoU ).

[0022] Preferably, when the model is deployed, it meets the following performance requirements:

[0023] When the input image resolution is 1920×1080, the inference speed ≥ 30 FPS;

[0024] Support real-time detection of 6 types of road diseases including cracks, potholes, ruts, looseness, bleeding, and subsidence;

[0025] After optimization by TensorRT, the model memory occupancy is compressed to < 500 MB, and the positioning accuracy error is maintained ≤ 2 mm.

[0026] A system for a road disease image recognition and localization detection method based on deep learning, including:

[0027] Data acquisition module: Configure the in-vehicle camera to collect road images with a resolution of ≥1920×1080;

[0028] Data augmentation module: Perform random rotation, brightness adjustment, and noise injection on the images to enhance the generalization of the model;

[0029] Lightweight neural network model: Use MobileNetV3 as the backbone network, combine the detection head with channel and spatial attention mechanisms, and a super-resolution localization branch to achieve sub-pixel coordinate regression;

[0030] Multi-task loss calculation unit: Combine FocalLoss and SmoothL1Loss+IoULoss to dynamically balance class imbalance and bounding box accuracy;

[0031] Edge deployment unit: Run in real-time on the NVIDIA Jetson Xavier device through FP16 quantization and TensorRT acceleration, with an inference speed of ≥30 FPS and a memory occupancy of <500MB.

[0032] Preferably, the data augmentation module further includes:

[0033] Environment simulation unit: Simulate complex lighting conditions by superimposing Gaussian noise or raindrop textures, and the noise intensity is dynamically adjusted according to preset environmental parameters;

[0034] Geometric transformation unit: Support random rotation, translation, and scaling of images to expand sample diversity.

[0035] Preferably, the lightweight neural network model includes:

[0036] CBAM module: Generate attention weights in the channel dimension through global average pooling and max pooling, and generate spatial weights through convolution in the spatial dimension to dynamically enhance the feature response in the crack area;

[0037] Super-resolution localization branch: Upsample the low-resolution feature map by 4 times through the SRGAN generator, output the coordinates of the disease center point, with an accuracy of ±1mm.

[0038] Preferably, the specific implementation method of the multi-task loss calculation unit is:

[0039] Classification loss: Use FocalLoss, and set the adjustment factor α∈[0.25,0.75] to balance the positive and negative sample ratios;

[0040] Localization loss: Combine SmoothL1Loss and IoULoss, and perform weighted fusion through the weight coefficients λ1 = 0.5 and λ2 = 0.5. The total loss function is: L = λ1×2L Focal +λ2×2(L SmoothL1 +LIoU )。

[0041] Preferably, the edge deployment unit meets the following performance indicators:

[0042] When the input image resolution is 1920×31080, the end-to-end inference latency ≤ 33 ms;

[0043] It supports real-time detection of 6 types of diseases, including cracks, potholes, ruts, looseness, bleeding, and subsidence. The memory occupancy for single-frame processing < 500 MB;

[0044] After optimization by TensorRT, the model parameters are compressed to FP16 precision, and the positioning accuracy error ≤ 2 mm.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] The method and system for road disease image recognition and localization detection based on deep learning proposed by the present invention extract multi-scale features through an improved lightweight network architecture (MobileNetV3 combined with CBAM attention module), enhancing the detection ability of small target diseases (such as cracks and net cracks); at the same time, a super-resolution localization module (SRGAN) is introduced to upsample the low-resolution feature map by 4 times to achieve sub-pixel coordinate regression (accuracy ±1 mm), meeting the millimeter-level positioning requirements for road maintenance. The technical solutions include data augmentation, optimization of multi-task loss functions (Focal Loss + Smooth L1 Loss + IoU Loss), and edge device deployment (model quantization and TensorRT acceleration). BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the method flow of the present invention;

[0048] Figure 2 is the sub-pixel level positioning error distribution diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] In order to clearly and completely describe the purpose, technical solutions of the present invention, and make the advantages more clearly understood, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are some, but not all, embodiments of the present invention, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0050] Embodiment 1, the present invention provides a technical solution: A method for road disease image recognition and localization detection based on deep learning, including the following steps:

[0051] Data collection: Collect road images through in-vehicle cameras with a resolution of not less than 1920×1080 pixels.

[0052] Data augmentation: Randomly rotate, adjust the brightness, and inject noise into the collected images to simulate rainy or shadow environments and improve the robustness of the model; the specific data augmentation strategy includes:

[0053] Randomly rotate the image within the range of ±15°;

[0054] Adjust the brightness to 50%-150% of the original image;

[0055] Inject Gaussian noise or simulate raindrop textures, and the noise intensity is dynamically adjusted according to preset environmental parameters.

[0056] Model construction: Design an improved lightweight neural network. The backbone network adopts the MobileNetV3 depthwise separable convolution structure to reduce the number of parameters. The detection head integrates the channel and spatial attention mechanism (CBAM module) to enhance the ability to extract small target features. The localization branch combines the super-resolution module (SRGAN) to achieve sub-pixel coordinate regression; the improved lightweight neural network further includes:

[0057] CBAM module: Generate channel attention weights through global average pooling and max pooling in the channel dimension, and generate spatial attention weights through convolution in the spatial dimension, and dynamically weight the feature map to enhance the response in the crack area;

[0058] Super-resolution localization module: Upsample the low-resolution feature map 4 times through SRGAN and output the coordinates of the disease center point with a positioning accuracy of ±1mm.

[0059] Multi-task training: Use the classification loss function Focal Loss to solve the class imbalance problem, and combine the localization loss functions Smooth L1 Loss and IoU Loss to improve the accuracy of the bounding box; the multi-task loss function is achieved through weighted fusion, and the specific form is:

[0060] Classification loss: Use Focal Loss, and the adjustment factor α∈[0.25,0.75] to balance positive and negative samples;

[0061] Localization loss: Combine Smooth L1 Loss and IoU Loss, and the weight coefficients are λ1 = 0.5 and λ2 = 0.5 respectively. The total loss function is: L = λ12L Focal +λ22(L SmoothL1 +L IoU )

[0062] Deployment Optimization: Through model quantization (FP16) and TensorRT acceleration technology, the inference speed ≥ 30FPS is achieved on the NVIDIA Jetson Xavier edge device, and the memory occupancy < 500MB. The following performance requirements are met during model deployment:

[0063] When the input image resolution is 1920×1080, the inference speed ≥ 30FPS;

[0064] It supports real-time detection of 6 types of diseases, including cracks, potholes, ruts, looseness, bleeding, and settlement;

[0065] After optimization by TensorRT, the model memory occupancy is compressed to < 500MB, and the positioning accuracy error is maintained at ≤ 2mm.

[0066] Performance Comparison:

[0067]

[0068] Example 2, based on Example 1, a system for road disease image recognition and positioning detection method based on deep learning is proposed, including:

[0069] Data Acquisition Module: Configure an in-vehicle camera to collect road images with a resolution ≥ 1920×1080;

[0070] Data Augmentation Module: Perform random rotation, brightness adjustment, and noise injection on the images to enhance the model's generalization ability; further include:

[0071] Environmental Simulation Unit: Simulate complex lighting conditions by superimposing Gaussian noise or raindrop textures, and the noise intensity is dynamically adjusted according to preset environmental parameters;

[0072] Geometric Transformation Unit: Support random rotation, translation, and scaling of images to expand sample diversity.

[0073] Lightweight Neural Network Model: Using MobileNetV3 as the backbone network, combined with a detection head with channel and spatial attention mechanisms, and a super-resolution positioning branch to achieve sub-pixel coordinate regression; include:

[0074] CBAM Module: Generate attention weights in the channel dimension through global average pooling and max pooling, and generate spatial weights through convolution in the spatial dimension to dynamically enhance the feature response in the crack area;

[0075] Super-Resolution Positioning Branch: Upsample the low-resolution feature map by 4 times through the SRGAN generator, and output the coordinates of the disease center point with an accuracy of ±1mm.

[0076] Multi-task Loss Calculation Unit: Combining Focal Loss with Smooth L1 Loss + IoU Loss to dynamically balance class imbalance and bounding box accuracy; the specific implementation method is as follows:

[0077] Classification Loss: Adopt Focal Loss, and set the adjustment factor α ∈ [0.25, 0.75] to balance the positive and negative sample ratios;

[0078] Localization Loss: Combine Smooth L1 Loss and IoU Loss, and perform weighted fusion through the weight coefficients λ1 = 0.5 and λ2 = 0.5. The total loss function is: L = λ12L Focal +λ22(L SmoothL1 +L IoU ).

[0079] Edge Deployment Unit: Through FP16 quantization and TensorRT acceleration, it runs in real time on the NVIDIA Jetson Xavier device, with an inference speed ≥ 30 FPS and a memory occupancy < 500 MB. The edge deployment unit meets the following performance indicators:

[0080] When the input image resolution is 1920×1080, the end-to-end inference latency ≤ 33 ms;

[0081] It supports real-time detection of 6 types of diseases, including cracks, potholes, ruts, looseness, bleeding, and subsidence, with a memory occupancy < 500 MB for single-frame processing;

[0082] After optimization by TensorRT, the model parameters are compressed to FP16 precision, and the localization accuracy error ≤ 2 mm.

[0083] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for road disease image recognition and location detection based on deep learning, characterized in that: It includes the following steps: Data acquisition: Collect road images through an in-vehicle camera with a resolution of no less than 1920×1080 pixels; Data augmentation: Perform random rotation, brightness adjustment, and noise injection on the collected images to simulate rainy or shadow environments and improve the robustness of the model; Model construction: Design an improved lightweight neural network. The backbone network adopts the depthwise separable convolution structure of MobileNetV3 to reduce the number of parameters. The detection head integrates channel and spatial attention mechanisms to enhance the ability to extract small object features. The localization branch combines a super-resolution module to achieve sub-pixel coordinate regression; Multi-task training: Use the classification loss function FocalLoss to solve the problem of class imbalance, and combine the localization loss functions SmoothL1Loss and IoULoss to improve the accuracy of the bounding box; Deployment optimization: Through model quantization and TensorRT acceleration technology, achieve an inference speed ≥ 30FPS and a memory occupancy < 500MB on the NVIDIA Jetson Xavier edge device.

2. The method for road disease image recognition and positioning detection based on deep learning according to claim 1, wherein: The data augmentation strategy specifically includes: Randomly rotate the image within the range of ±15°; Adjust the brightness to 50%-150% of the original image; Inject Gaussian noise or simulate raindrop textures, and the noise intensity is dynamically adjusted according to the preset environmental parameters.

3. The method for road disease image recognition and location detection based on deep learning according to claim 2, wherein: The improved lightweight neural network further includes: Generating channel attention weights through global average pooling and max pooling in the channel dimension, generating spatial attention weights through convolution in the spatial dimension, and dynamically weighting the feature map to enhance the response in the crack area; Upsampling the low-resolution feature map by 4 times through SRGAN and outputting the coordinates of the disease center point with a positioning accuracy of ±1mm.

4. The method for road disease image recognition and location detection based on deep learning according to claim 3, wherein: The multi-task loss function is achieved through weighted fusion, and the specific form is: Classification loss: Adopt FocalLoss, and the adjustment factor α ∈ [0.25, 0.75] to balance positive and negative samples; Localization loss: Combining SmoothL1Loss and IoULoss with weight coefficients λ1 = 0.5 and λ2 = 0.5 respectively, the total loss function is: L = λ1·L Focal + λ2·(L SmoothL1 + L IoU ).

5. The method for road disease image recognition and location detection based on deep learning according to claim 4, characterized in that: When the model is deployed, it meets the following performance requirements: When the input image resolution is 1920×1080, the inference speed ≥ 30FPS; Support real-time detection of 6 types of diseases including cracks, potholes, ruts, looseness, bleeding, and subsidence; After optimization by TensorRT, the model memory occupancy is compressed to < 500MB, and the positioning accuracy error is maintained ≤ 2mm.

6. A system for the method of road disease image recognition and positioning detection based on deep learning according to claim 5, characterized in that: It includes: Data acquisition module: Configure an in-vehicle camera to collect road images with a resolution ≥ 1920×1080; Data augmentation module: Perform random rotation, brightness adjustment, and noise injection on the image to enhance the generalization of the model; Lightweight neural network model: Use MobileNetV3 as the backbone network, a detection head combined with channel and spatial attention mechanisms, and a super-resolution localization branch to achieve sub-pixel coordinate regression; Multi-task loss calculation unit: Combine FocalLoss and SmoothL1Loss + IoULoss to dynamically balance class imbalance and bounding box accuracy; Edge deployment unit: Run in real time on the NVIDIA Jetson Xavier device through FP16 quantization and TensorRT acceleration, with an inference speed ≥ 30FPS and a memory occupancy < 500MB.

7. A system according to claim 6, characterized in that: The data augmentation module further includes: Environment simulation unit: Simulates complex lighting conditions by superimposing Gaussian noise or raindrop textures, and the noise intensity is dynamically adjusted according to preset environment parameters; Geometric transformation unit: Supports random rotation, translation, and scaling of images to expand sample diversity.

8. A system according to claim 6, characterized in that: The lightweight neural network model includes: CBAM module: Generates attention weights through global average pooling and max pooling in the channel dimension, and generates spatial weights through convolution in the spatial dimension to dynamically enhance the feature response in the crack area; Super-resolution localization branch: Upsamples the low-resolution feature map by 4 times through the SRGAN generator, outputs the coordinates of the disease center point, and the accuracy reaches ±1mm.

9. A system according to claim 6, wherein: The specific implementation method of the multi-task loss calculation unit is: Classification loss: Adopts FocalLoss, and sets the adjustment factor α∈[0.25,0.75] to balance the proportion of positive and negative samples; Localization loss: Combining SmoothL1Loss and IoULoss, weighted fusion is performed with weight coefficients λ1 = 0.5 and λ2 = 0.

5. The total loss function is: L = λ1·L Focal + λ2·(L SmoothL1 + L IoU ).

10. A system according to claim 6, characterized in that: The edge deployment unit meets the following performance indicators: When the input image resolution is 1920×1080, the end-to-end inference latency ≤ 33ms; Supports real-time detection of 6 types of diseases including cracks, potholes, ruts, looseness, bleeding, and settlement, and the memory occupancy for single-frame processing < 500MB; After optimization by TensorRT, the model parameters are compressed to FP16 precision, and the positioning accuracy error ≤ 2mm.