Highway disease detection method capable of accurately positioning

By using an instance segmentation model that combines high-definition binocular cameras and GPS, the problem of insufficient recognition accuracy and adaptability in highway defect detection has been solved. Pixel-level recognition and continuous tracking of defects have been achieved, improving detection accuracy and adaptability and reducing road maintenance costs.

CN120877121APending Publication Date: 2025-10-31GUANGDONG ZHIDIAN HI-TECH CO LTD

Patent Information

Application Number
CN202511141553.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies for detecting highway defects suffer from low accuracy and poor adaptability, making it difficult to detect multiple defects simultaneously, particularly for minor cracks, small potholes, and repaired pavements.

Method used

A high-definition binocular camera combined with GPS is used for real-time positioning to build an instance segmentation model. The CBAM attention mechanism and P2 detection head are introduced, and disease tracking is performed by combining Kalman filtering and optical flow. The model is deployed using an edge machine, and the detection accuracy is optimized by CIoU bounding box regression strategy and BCE loss function.

Benefits of technology

It achieves pixel-level identification and continuous tracking of highway defects, accurately locates the defect positions, improves detection accuracy and adaptability, and reduces road maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877121A_ABST
    Figure CN120877121A_ABST
Patent Text Reader

Abstract

The invention provides an expressway disease detection method capable of accurately positioning. The expressway disease detection method comprises the following steps: S1, image acquisition and data annotation; s2, constructing an instance segmentation model; s3, extracting a multi-scale feature map and fusing multi-scale features; s4, outputting a category label, a bounding box coordinate and a confidence coefficient score of the detection box; s5, performing multi-target trajectory tracking; s6, eliminating global motion influence; s7, establishing an association relationship between the prediction trajectory and the current detection frame; s8, dynamically maintaining a track list; and S9, outputting bounding box coordinates and ID tags in each target continuous frame, and storing the bounding box coordinates and the ID tags as structured data. According to the method, a high-performance target detection framework, an advanced instance segmentation tracking technology and a GPS positioning technology are combined, pixel-level identification and continuous tracking of diseases can be realized while high detection precision is ensured, and the positions of the diseases can be accurately positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method for detecting highway defects that can be accurately located. Background Technology

[0002] With economic development and the continuous increase in traffic flow, prolonged use and the effects of natural environmental factors, such as repeated crushing by heavy vehicles, fatigue of road materials caused by temperature changes, and water damage, can lead to various defects on highway pavements, including cracks, potholes, and ruts. These defects not only affect driving safety and comfort but also shorten the service life of the road and increase maintenance costs.

[0003] While manual inspection offers a certain level of accuracy, it suffers from low efficiency, high labor intensity, and strong subjectivity, making it difficult to meet the high-frequency inspection needs of large-scale road networks. Traditional automated inspection methods, although more efficient, generally suffer from low accuracy, poor adaptability, and an inability to accurately locate defects. Furthermore, they often focus on identifying single types of defects (such as cracks), making it difficult to simultaneously detect multiple defects. This is particularly problematic for highways, where the road surface often contains fine cracks, small potholes, and repaired surfaces, where high accuracy is crucial. Summary of the Invention

[0004] In view of the problems existing in the prior art, the present invention proposes a method for detecting highway defects that can accurately locate them.

[0005] To achieve the above objectives, the technical solution adopted by this invention is: a method for accurately locating highway defects, comprising:

[0006] S1. Image Acquisition and Data Labeling: Real-time video is acquired using a high-definition binocular camera installed on the acquisition vehicle. The acquired real-time video is preprocessed, and GPS is used for real-time positioning. Polygon annotation is performed on the target area using image labeling tools to construct a training sample set with pixel-level accuracy.

[0007] S2. Constructing an instance segmentation model: A general algorithm model is used as the basic architecture, and the CBAM attention mechanism is introduced and a P2 detection head is added for multiple iterations of training. At the same time, an instance segmentation module is integrated to achieve pixel-level recognition of road defects. The finally trained instance segmentation model can simultaneously recognize multiple common defect types. The instance segmentation model includes: an input layer, a backbone network, a neck network, and a head network.

[0008] S3. Use the edge machine to obtain the training sample set in S1 via the RTSP protocol through the network cable, and deploy the instance segmentation model on the edge machine. Extract multi-scale feature maps through the backbone network in the instance segmentation model, and fuse the multi-scale feature maps with the neck network.

[0009] S4. The head network of the instance segmentation model outputs the category label, bounding box coordinates, and confidence score for each detection box;

[0010] S5. Use Kalman filtering to predict the position and size of each detection box, predict the potential location of the disease in the current frame, and generate a predicted trajectory. At the same time, determine whether the target trajectory is the same target trajectory based on the confidence score. The higher the confidence score, the higher the weight of the target trajectory being the same trajectory. Perform multi-target trajectory tracking.

[0011] S6. For situations involving camera movement or large-scale scene displacement, the global motion effect is eliminated using optical flow, where the global motion effect includes target tracking drift, loss, or incorrect matching.

[0012] S7. Establish the association between the predicted trajectory and the current detection box to ensure that the target maintains a unique ID in consecutive frames, avoid ID switching, and avoid incorrect matching due to position deviation.

[0013] S8. Based on target detection and trajectory prediction, the system dynamically maintains the trajectory list to adapt to the appearance, disappearance and changes in motion state of the target;

[0014] S9. Finally, output the bounding box coordinates, category label, and ID of each target in consecutive frames and store them as structured data.

[0015] Furthermore, during the training phase of the instance segmentation model in S2, a bounding box regression strategy based on CIoU was adopted, the formula of which is as follows:

[0016]

[0017] Where IoU represents the intersection-over-union ratio, which measures the degree of overlap between two bounding boxes. CIoU is the object detection bounding box regression loss function, d represents the Euclidean distance between the center points of the predicted box and the ground truth box, c represents the diagonal length of the smallest closure region that can simultaneously cover the predicted box and the ground truth box, v represents the aspect ratio consistency measure, and α represents the weighting factor used to balance the influence of intersection-over-union ratio and aspect ratio consistency.

[0018] The parameters of the detector head are updated by combining the SGD optimizer and the BCE loss function, and the formula is as follows:

[0019] BCE(p,y)=-ilog(p)-(1-y)log(1-p);

[0020] Where p represents the probability value output by the model, and y represents the actual label.

[0021] Based on the above, this invention, while maintaining the accuracy of bounding box prediction, effectively improves the model's ability to locate irregular targets in complex scenes by introducing a consistency metric v for aspect ratio and the Euclidean distance d between the center points of the predicted and ground truth boxes. Especially in the task of identifying slender defects such as road cracks, CIoU exhibits stronger scale adaptability and convergence stability compared to traditional IoU or GIoU, helping to enhance the model's sensitivity to edge details and thus improving overall detection performance.

[0022] Furthermore, the operation method for real-time GPS positioning in S1 is as follows:

[0023] The first valid location point (lon0, lat0) is obtained from GPS and used as the initial reference point. The timestamp t0 at this time is recorded. New GPS data is obtained at regular intervals, including:

[0024] latitude and longitude i ,lat i ), heading i speed i timestamp t i ;

[0025] For image acquisition frames between two GPS update points, dead reckoning is used to estimate the position at the intermediate time:

[0026] Calculate the time interval d t =t i -t i-1 ;

[0027] Based on the speed of the previous frame i-1 and heading i-1 The distance traveled was calculated as follows:

[0028] d = speed i-1 ×d t ;

[0029] After converting the heading angle to radians, calculate the eastward and northward displacements:

[0030] dx = d·cos(θ);

[0031] dy = d·sin(θ);

[0032] Update the current vehicle's local coordinates or calculate them in reverse to obtain new latitude and longitude.

[0033] When a defect is detected in a frame of an image, the estimated location of the defect between two GPS points is found based on the timestamp of the frame, and its latitude and longitude information is added to the defect detection result to generate a defect report with geographic tags.

[0034] Furthermore, the backbone network in S2 includes:

[0035] CSPDarknet53: Employs a cross-stage partial connection structure to reduce computation and improve feature representation capabilities;

[0036] Multi-level convolutional module: including first Conv layer, second Conv layer, third Conv layer, fourth Conv layer and fifth Conv layer;

[0037] The first C3k2 module includes the first C3k2, the second C3k2, the third C3k2, and the fourth C3k2;

[0038] SPPF module: Introduces spatial pyramid pooling, which integrates contextual information through pooling windows of different scales to enhance the model's robustness to changes in the target scale;

[0039] The first Conv layer, second Conv layer, first C3k2 layer, third Conv layer, second C3k2 layer, fourth Conv layer, third C3k2 layer, fifth Conv layer, and fourth C3k2 layer are arranged sequentially. The first Conv layer performs initial downsampling on the image data to extract shallow texture and outputs the target feature map P1. Then, the second Conv layer performs downsampling to extract edges / texture and outputs the target feature map P2. Next, the first C3k2 layer stacks two layers of lightweight residual blocks to enhance the representation, and the third Conv layer performs downsampling again. The network then obtains a small to medium target feature map P3. Next, the network depth is deepened through the second C3k2 layer to improve semantic understanding. The fourth Conv layer downsamples the medium to large targets to obtain a medium target feature map P4. Then, the third C3k2 layer enhances the residual path and maintains gradient propagation. After the image data is processed by the fifth Conv layer, the maximum target feature map P5 is obtained. The fourth C3k2 layer enhances the deep features and inputs them into the SPPF module. The SPPF module then inputs the features into the C2PSA module to enhance the expression of key regions, and then inputs them into the neck network.

[0040] Furthermore, the neck network in S2 adopts a feature pyramid / path aggregation structure, which achieves bidirectional feature fusion and cross-scale connectivity through top-down feature pyramids and bottom-up path aggregation.

[0041] Multiple concatenation module and sampling module: The multiple concatenation module fuses feature maps from different levels and restores resolution through upsampling to improve the detection capability of the instance segmentation model; the multiple concatenation module includes a first concat, a second concat, a third concat, and a fourth concat, and the sampling module includes a first upsample, a second upsample, a first downsample, and a second downsample;

[0042] The second C3k2 module includes the fifth, sixth, seventh, and eighth C3k2 modules;

[0043] The first Upsample, first Concat, fifth C3k2, second Upsample, second Concat, sixth C3k2, first Downsample, third Concat, seventh C3k2, second Downsample, fourth Concat, and eighth C3k2 are set sequentially. The first Upsample performs upsampling to enhance spatial resolution. Then, the first Concat fuses the P4 output from the backbone network. The fifth C3k2 compresses the feature map channels and fuses semantics. Then, the second Upsample further performs upsampling for small target detection. The second Concat fuses the P3 from the backbone network. The sixth C3k2 obtains the final P3. Then, the first Downsample constructs the target detection output path, and the third Concat fuses the P3 from the backbone network. The seventh C3k2 obtains the final P4. Then, the second Downsample constructs the large target detection output path to prepare for large target detection. The fourth Concat fuses the P5 from the backbone network. Finally, the eighth C3k2 obtains the final P5.

[0044] Furthermore, the head network in S2 includes:

[0045] Classification branch: Outputs the target category;

[0046] Regression branch: Predicts bounding box coordinates (x, y, w, h);

[0047] Mask branch: Generates pixel-level segmentation masks to achieve precise contour division of each target;

[0048] Multi-scale output: Outputs detection results through feature maps at multiple levels, adapting to targets of different sizes.

[0049] Furthermore, the training sample set in S1 is divided into a training subset, a validation subset, and a test subset according to a set ratio, which are used for parameter learning of the deep learning model, monitoring of the training process, and final performance evaluation, respectively, to ensure that the model has good recognition ability and stability.

[0050] Furthermore, in S2, when inferring the instance segmentation model, the video frames are scaled to 640×640 and normalized before being input into the instance segmentation model. In the post-processing stage, overlapping boxes are removed using non-maximum suppression (NMS) and the segmentation boundaries are smoothed using morphological operations.

[0051] Meanwhile, this invention provides:

[0052] A server includes a processor and a memory, the memory storing at least one program that is loaded and executed by the processor to implement the aforementioned method for accurately locating highway defects.

[0053] A computer-readable storage medium storing at least one program, which is loaded and executed by a processor to implement the above-described method for accurately locating highway defects.

[0054] In summary, the beneficial technical effects of this invention are as follows: This invention combines a high-performance target detection framework, advanced instance segmentation and tracking technology, and GPS positioning technology, which can achieve pixel-level identification and continuous tracking of defects while ensuring high detection accuracy, and can accurately locate the location of defects. This allows for accurate and timely repairs before the defects on highways become more serious, eliminating road safety hazards and reducing road maintenance costs. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating the overall workflow of the present invention;

[0056] Figure 2 This is a network structure diagram of the instance segmentation model algorithm of the present invention. Detailed Implementation

[0057] like Figures 1 to 2 As shown, a method for accurately locating highway defects includes:

[0058] S1. Image Acquisition and Data Labeling: Real-time video is acquired using a high-definition binocular camera mounted on a data acquisition vehicle. The acquired video undergoes preprocessing, including noise reduction and contrast enhancement to minimize interference from invalid data in subsequent algorithms. Simultaneously, GPS is used for real-time positioning, providing latitude, longitude, heading angle, speed, and timestamp information. Furthermore, polygonal annotations are applied to target areas using image labeling tools to construct a training sample set with pixel-level accuracy. This training sample set is divided into a training subset, a validation subset, and a test subset in a 7:1:2 ratio, used for parameter learning of the deep learning model, monitoring of the training process, and final performance evaluation, respectively, ensuring the model possesses good recognition capabilities and stability. In addition, the acquired video content also includes data on weather conditions, light intensity, road type, and tunnel type.

[0059] S2. Constructing an Instance Segmentation Model: A general algorithm model is used as the basic architecture, and multiple iterations of training are performed to form an instance segmentation model. The general algorithm model uses YOLOv11 as the original architecture and RDD2022 is used to train pre-trained weights. The instance segmentation model is then structurally optimized based on this, mainly by introducing the CBAM attention mechanism to enhance the attention capability of disease areas and adding a P2 detection head to improve the detection accuracy of small diseases. The CBAM attention mechanism is located in the backbone and neck networks of the instance segmentation model, embedded in the first C3k2 module, the second C3k2 module, and C2PSA. The CBAM attention mechanism does not directly participate in detection but indirectly enhances feature representation to make subsequent detection more accurate, thereby enhancing the expressive power of the feature map and improving detection accuracy. The P2 detection head is extracted from the backbone network, enters the neck network for necessary channel alignment, fusion, and enhancement, and then serves as the fourth output branch in the Detect layer of the head network. Finally, it generates bounding boxes and classification results for small targets, thereby increasing the detection capability for extremely small targets and improving the model's accuracy in small target scenes. These improvements make the model more adaptable to the needs of identifying defects such as cracks and potholes in highway scenarios. An instance segmentation module is also integrated to obtain pixel-level contours of defects for subsequent size calculations and to provide better road maintenance suggestions. The finally trained instance segmentation model can simultaneously identify multiple common defect types, including but not limited to: transverse cracks, longitudinal cracks, alligator cracks, ruts, and potholes, thereby achieving refined defect area segmentation. The instance segmentation model includes:

[0060] Input layer: This layer receives the input image data and outputs images in a uniform size of [B, 3, 640, 640], which serves as the basis for model processing. Here, B is the batch size, 3 is the number of channels, and 640×640 represents the width and height. The input layer simply standardizes the image format, providing a uniform input size for the backbone network; it does not perform feature extraction.

[0061] Backbone network: Extracts multi-level features from images and progressively enhances semantic information;

[0062] The backbone network consists of the following parts:

[0063] (1)CSPDarknet53: It adopts a cross-stage partial connection structure to reduce the amount of computation and improve the feature representation ability;

[0064] (2) Multi-level convolutional module: including the first Conv layer, the second Conv layer, the third Conv layer, the fourth Conv layer and the fifth Conv layer;

[0065] (3) First C3k2 module: including first C3k2, second C3k2, third C3k2 and fourth C3k2;

[0066] (4) SPPF module: Introduces spatial pyramid pooling, which integrates contextual information through pooling windows of different scales to enhance the robustness of the model to changes in the target scale;

[0067] The first Conv layer, second Conv layer, first C3k2 layer, third Conv layer, second C3k2 layer, fourth Conv layer, third C3k2 layer, fifth Conv layer, and fourth C3k2 layer are set sequentially. The first Conv layer performs initial downsampling on the image data to extract shallow texture. Then, the second Conv layer performs downsampling to extract edges / texture. Next, the first C3k2 layer stacks two lightweight residual blocks to enhance the representation. The third Conv layer performs downsampling again to obtain small to medium target feature maps P3. Then, the second C3k2 layer deepens the network depth to improve semantic understanding. The fourth Conv layer performs downsampling on medium to large targets to obtain medium target feature maps P4. Then, the third C3k2 layer enhances the residual path and maintains gradient propagation. After the image data is processed by the fifth Conv layer, the maximum target feature map P5 is obtained. The fourth C3k2 layer enhances the deep features and inputs them into the SPPF module. The SPPF module then inputs the features into the C2PSA module to enhance the expression of key regions, and then inputs them into the neck network.

[0068] Neck network: Employs a feature pyramid / path aggregation structure, achieving bidirectional feature fusion and cross-scale connectivity through top-down feature pyramids and bottom-up path aggregation; helps the model understand targets of different sizes, forming three final output feature maps: P3, P4, and P5, which are used for small, medium, and large targets, respectively;

[0069] The neck network consists of the following parts:

[0070] (1) Multiple stitching module and sampling module: The multiple stitching module fuses feature maps of different levels and restores resolution through upsampling to improve the detection capability of the instance segmentation model; the multiple stitching module includes first Concat, second Concat, third Concat and fourth Concat, and the sampling module includes first Upsample, second Upsample, first Downsample and second Downsample;

[0071] (2) Second C3k2 module: including fifth C3k2, sixth C3k2, seventh C3k2 and eighth C3k2; wherein the function of the fifth C3k2, sixth C3k2, seventh C3k2 and eighth C3k2 is to enhance the feature extraction capability, so that the model can capture more details and contextual information at different scales.

[0072] The first Upsample, first Concat, fifth C3k2, second Upsample, second Concat, sixth C3k2, first Downsample, third Concat, seventh C3k2, second Downsample, fourth Concat, and eighth C3k2 are set sequentially. The first Upsample performs upsampling to enhance spatial resolution. Then, the first Concat fuses the P4 output from the backbone network. The fifth C3k2 compresses the feature map channels and fuses semantics. Then, the second Upsample further performs upsampling for small target detection. The second Concat fuses the P3 from the backbone network. The sixth C3k2 obtains the final P3. Then, the first Downsample constructs the target detection output path, and the third Concat fuses the P3 from the backbone network. The seventh C3k2 obtains the final P4. Then, the second Downsample constructs the large target detection output path to prepare for large target detection. The fourth Concat fuses the P5 from the backbone network. Finally, the eighth C3k2 obtains the final P5.

[0073] Head network: Performs object detection, classification, and instance segmentation tasks;

[0074] The head network consists of the following parts:

[0075] (1) Classification branch: Output the target category;

[0076] (2) Regression branch: Predict the bounding box coordinates (x,y,w,h);

[0077] (3) Mask branch: Generate pixel-level segmentation mask to achieve accurate contour division of each target;

[0078] (4) Multi-scale output: The detection results P3, P4 and P5 are output through feature maps of multiple levels to adapt to targets of different sizes.

[0079] S3. Use the edge machine to obtain the training sample set in S1 via the RTSP protocol through the network cable, and deploy the instance segmentation model on the edge machine. Extract multi-scale feature maps through the backbone network in the instance segmentation model, and fuse the multi-scale feature maps with the neck network.

[0080] S4. The head network of the instance segmentation model outputs the category label, bounding box coordinates, and confidence score for each detection box. The category label is the name of the identified disease, and the confidence score reflects the probability of the target existing in the detection box. It is used to filter high-reliability results and is also used for weight calculation of trajectory association. The higher the confidence score, the higher the weight of the target trajectory being the same trajectory; ensuring that high-confidence targets are matched first.

[0081] S5. The positions of the defects (cracks, potholes) change little in consecutive video frames, but the detection boxes may jitter due to changes in the camera's perspective caused by vehicle movement. Kalman filtering is used to predict the position and size of each detected defect box, predicting the potential location of the defect in the current frame. This smooths out position changes and reduces detection errors. Simultaneously, a confidence score is used to determine whether the target trajectories are the same, and finally, a predicted trajectory is generated. The confidence score ranges from 0 to 1; in this embodiment, the confidence threshold is set to 0.5. When the confidence score between two detected target trajectories reaches 0.5, they are considered to be the same target trajectory.

[0082] S6. For situations involving camera movement or large-scale scene displacement, the global motion effect is eliminated using optical flow.

[0083] S7. Establish the association between the predicted trajectory and the current detection box. The steps include:

[0084] S71. Calculate the IoU value between each trajectory prediction box and the detection box as the basis for position matching. At the same time, use Reid to perform similarity calculation on the appearance features (texture and shape feature vectors) of the disease. For occluded scenes, introduce Transformer temporal feature aggregation and use Referring Cross Attention to associate inter-frame queries to restore trajectory consistency, ensure that the target maintains a unique ID in consecutive frames, and avoid ID switching.

[0085] S72. The Hungarian algorithm is used for matching, pairwise matching the detection result of the current frame with the predicted trajectory of the previous frame. IoU is prioritized to measure spatial overlap; when the IoU is less than a set value, the two are considered mismatched, thus filtering out unreasonable matches. IoU is combined with appearance similarity to avoid erroneous matches caused by positional deviations.

[0086] S8. Based on target detection and trajectory prediction, the system dynamically maintains the trajectory list to adapt to the appearance, disappearance and motion state changes of the target; the system creates a trajectory based on the confidence of the detection box, initializes the trajectory parameters, assigns the center position and velocity information of the detection box to the state vector of the Kalman filter, establishes a state space model, uses the Kalman filter to predict the trajectory position in subsequent frames, and updates the state of the Kalman filter in conjunction with the detection box to achieve smooth tracking.

[0087] S9. Finally, output the bounding box coordinates, category label, and ID of each target in consecutive frames, and store them as structured data. The ID label is used to determine whether the target tracking is the same identified target. The target damage types are divided into pavement damage, cracks, deformation, and structural types, which can be further subdivided into pavement peeling, transverse cracks, longitudinal cracks, network cracks, potholes, swells, subsidence, bleeding, frost heave, etc.

[0088] In object detection using the instance segmentation model, each result is assigned an ID value. S7 is used to determine if the detected objects are the same target. If they are the same target, they are assigned the same ID value. If they are not the same target, a new ID value is assigned. If no image with the same ID is detected in the subsequent 10 images, the ID is released for reuse later, thus avoiding redundant calculations.

[0089] During the training phase of the instance segmentation model in S2, a bounding box regression strategy based on CIoU was adopted, and its formula is as follows:

[0090]

[0091] Where IoU represents the intersection-over-union ratio, which measures the degree of overlap between two bounding boxes. CIoU is the object detection bounding box regression loss function, d represents the Euclidean distance between the center points of the predicted box and the ground truth box, c represents the diagonal length of the smallest closure region that can simultaneously cover the predicted box and the ground truth box, v represents the aspect ratio consistency measure, and α represents the weighting factor used to balance the influence of intersection-over-union ratio and aspect ratio consistency.

[0092] The parameters of the detector head are updated by combining the SGD optimizer and the BCE loss function, and the formula is as follows:

[0093] BCE(p,y)=-ilog(p)-(1-y)log(1-p);

[0094] Where p represents the probability value output by the model, and y represents the actual label.

[0095] Based on the above, this invention, while maintaining the accuracy of bounding box prediction, effectively improves the model's ability to locate irregular targets in complex scenes by introducing a consistency metric v for aspect ratio and the Euclidean distance d between the center points of the predicted and ground truth boxes. Especially in the task of identifying slender defects such as road cracks, CIoU exhibits stronger scale adaptability and convergence stability compared to traditional IoU or GIoU, helping to enhance the model's sensitivity to edge details and thus improving overall detection performance.

[0096] The operation method for real-time GPS positioning in S1 is as follows:

[0097] The first valid location point (lon0, lat0) is obtained from GPS and used as the initial reference point. The timestamp t0 at this time is recorded. New GPS data is obtained at regular intervals, including:

[0098] latitude and longitude i ,lat i ), heading i speed i timestamp t i ;

[0099] For image acquisition frames between two GPS update points, dead reckoning is used to estimate the position at the intermediate time:

[0100] Calculate the time interval d t =t i -t i-1 ;

[0101] Based on the speed of the previous frame i-1 and heading i-1 The distance traveled was calculated as follows:

[0102] d = speed i-1 ×d t ;

[0103] After converting the heading angle to radians, calculate the eastward and northward displacements:

[0104] dx = d·cos(θ);

[0105] dy = d·sin(θ);

[0106] Update the current vehicle's local coordinates or calculate them in reverse to obtain new latitude and longitude.

[0107] When a defect is detected in a frame of an image, the estimated location of the defect between two GPS points is found based on the timestamp of the frame, and its latitude and longitude information is added to the defect detection result to generate a defect report with geographic tags.

[0108] In S2, when inferring the instance segmentation model, video frames are scaled to 640×640 and normalized before being input into the instance segmentation model. In the post-processing stage, overlapping boxes are removed using non-maximum suppression (NMS) and segmentation boundaries are smoothed using morphological operations.

[0109] In step S2, within the instance segmentation model, the specific changes of the image in the input layer, backbone network, neck network, and head network are as follows:

[0110] The input layer standardizes the image size to [B, 3, 640, 640], where B is the batch size, 3 is the number of channels (RGB), and 640×640 is the width and height. This only standardizes the image format; no feature extraction is performed. It provides a uniform input size for the backbone network.

[0111] 1. The backbone network is responsible for extracting multi-scale semantic features of the image in advance, performing 5 downsampling operations, and extracting 3 main output layers.

[0112] First Conv(64,3,2): Initial downsampling, extracting shallow texture, outputting P1;

[0113] The second Conv(128,3,2): downsamples to extract edges / textures, outputting P2;

[0114] First C3k2(256,2): Stack two layers of lightweight residual blocks to enhance characterization;

[0115] The third Conv(256,3,2): downsampling to extract mid-scale target features;

[0116] The second C3k2(512,2): increases network depth and improves semantic understanding;

[0117] Fourth Conv(512,3,2): Downsampling for medium to large-sized targets;

[0118] The third C3k2(512,2,shortcut=True): a strong residual path that preserves gradient propagation;

[0119] Fifth Conv(1024,3,2): Deepest feature, focusing on large targets;

[0120] Fourth C3k2(1024,2): Deep feature enhancement;

[0121] SPPF module (5): Multi-scale pooling to expand the receptive field;

[0122] C2PSA module (1024,2): Self-attention mechanism to enhance the expression of key regions.

[0123] Second: The neck network, feature fusion and scale unification, helps the model understand targets of different sizes and forms three layers of output feature maps (P3, P4, P5) for small, medium and large targets respectively;

[0124] First Upsample: Upsampling enhances spatial resolution;

[0125] First Concat: Merges the P4 output (layer 6) in the backbone network;

[0126] Fifth C3k2: Compress feature map channels and fuse semantics;

[0127] Second Upsample: Further upsampling for small target detection;

[0128] Second Concat: Merges the P3 output (layer 4) in the backbone network;

[0129] Sixth C3k2: Obtain the final P3 (small target detection output);

[0130] First Downsample: Used to construct the target detection output path;

[0131] Third Concat: Connects to the output of C3k2 (layer 13);

[0132] Seventh C3k2: Obtain P4 (target detection output);

[0133] Second Downsample (512): Prepare for large target detection;

[0134] Fourth Concat: Integrate P5 (layer 10) in the backbone network;

[0135] Eighth C3k2: Obtain the final P5 (large target detection output);

[0136] Three: The head network performs prediction output;

[0137] 16→P3(80×80): Detect small targets;

[0138] 19→P4(40×40): Target under detection;

[0139] 22→P5(20×20): Detect large targets.

[0140] The precise positioning provided by this invention plays a crucial role in the subsequent inspection and repair of roads. Without the positioning effect, the detected results would have no practical value.

[0141] The application scenario of this invention is unique: the defects on highways differ from those on municipal roads and rural roads. Highway defects mainly consist of minor cracks, small potholes, and repaired surfaces. The main requirement is to repair the defects before they become more severe, demanding higher precision in the model's detection of minute details and minimizing misidentification of repaired surfaces. In contrast, potholes on ordinary municipal roads are larger, and the precision required for the model's detection of minute details is not as high. The technical solution of this invention addresses these issues by adding specific measures, such as including repaired road conditions in the dataset. Furthermore, we have enhanced the detection of tiny targets smaller than 32×32 pixels by adding a CBAM attention detection mechanism, a P2 probe, and multi-scale feature fusion.

[0142] Meanwhile, this invention provides:

[0143] A server includes a processor and a memory, the memory storing at least one program that is loaded and executed by the processor to implement the aforementioned method for accurately locating highway defects.

[0144] A computer-readable storage medium storing at least one program, which is loaded and executed by a processor to implement the above-described method for accurately locating highway defects.

[0145] The above description is merely the optimal embodiment of the present invention and is not intended to limit the present invention. Any modifications or substitutions made by those skilled in the art without departing from the essence and scope of protection of the present invention should also be within the scope of protection of the present invention.

Claims

1. A method for accurately locating highway defects, characterized in that, Includes the following steps: S1. Image Acquisition and Data Labeling: Real-time video is acquired using a high-definition binocular camera installed on the acquisition vehicle. The acquired real-time video is preprocessed, and GPS is used for real-time positioning. Polygon annotation is performed on the target area using image labeling tools to construct a training sample set with pixel-level accuracy. S2. Constructing an instance segmentation model: A general algorithm model is used as the basic architecture, and the CBAM attention mechanism is introduced and a P2 detection head is added for multiple iterations of training. At the same time, an instance segmentation module is integrated to achieve pixel-level recognition of road defects. The finally trained instance segmentation model can simultaneously recognize multiple common defect types. The instance segmentation model includes: an input layer, a backbone network, a neck network, and a head network. S3. Use the edge machine to obtain the training sample set in S1 via the RTSP protocol through the network cable, and deploy the instance segmentation model on the edge machine. Extract multi-scale feature maps through the backbone network in the instance segmentation model, and fuse the multi-scale feature maps with the neck network. S4. The head network of the instance segmentation model outputs the category label, bounding box coordinates, and confidence score for each detection box; the category label of the detection box is the highway defect type of the detected target. S5. Use Kalman filtering to predict the position and size of each detection box, predict the potential location of the disease in the current frame, and generate a predicted trajectory. At the same time, determine whether the target trajectory is the same target trajectory based on the confidence score. The higher the confidence score, the higher the weight of the target trajectory being the same trajectory. Perform multi-target trajectory tracking. S6. For situations involving camera movement or large-scale scene displacement, the global motion effect is eliminated using optical flow, where the global motion effect includes target tracking drift, loss, or incorrect matching. S7. Establish the association between the predicted trajectory and the current detection box to ensure that the target maintains a unique ID in consecutive frames, avoid ID switching, and avoid incorrect matching due to position deviation. S8. Based on target detection and trajectory prediction, the system dynamically maintains the trajectory list to adapt to the appearance, disappearance and changes in motion state of the target; S9. Finally, output the bounding box coordinates, category label, and ID of each target in consecutive frames and store them as structured data.

2. The method for accurately locating highway defects according to claim 1, characterized in that: During the training phase of the instance segmentation model in S2, a bounding box regression strategy based on CIoU was adopted, and its formula is as follows: Where IoU represents the intersection-over-union ratio, which measures the degree of overlap between two bounding boxes. CIoU is the object detection bounding box regression loss function, d represents the Euclidean distance between the center points of the predicted box and the ground truth box, c represents the diagonal length of the smallest closure region that can simultaneously cover the predicted box and the ground truth box, v represents the aspect ratio consistency measure, and α represents the weighting factor used to balance the influence of intersection-over-union ratio and aspect ratio consistency. The parameters of the detector head are updated by combining the SGD optimizer and the BCE loss function, and the formula is as follows: BCE(p,y)=-ilog(p)-(1-y)log(1-p); Where p represents the probability value output by the model, and y represents the actual label.

3. The method for accurately locating highway defects according to claim 1, characterized in that, The operation method for real-time GPS positioning in S1 is as follows: The first valid location point (lon0, lat0) is obtained from GPS and used as the initial reference point. The timestamp t0 at this time is recorded. New GPS data is obtained at regular intervals, including: latitude and longitude i ,lat i ), heading i speed i timestamp t i ; For image acquisition frames between two GPS update points, dead reckoning is used to estimate the position at the intermediate time: Calculate the time interval d t =t i -t i-1 ; Based on the speed of the previous frame i-1 and heading i-1 The distance traveled was calculated as follows: d=speed i-1 ×d t ; After converting the heading angle to radians, calculate the eastward and northward displacements: dx = d·cos(θ); dy = d·sin(θ); Update the current vehicle's local coordinates or calculate them in reverse to obtain new latitude and longitude. When a defect is detected in a frame of an image, the estimated location of the defect between two GPS points is found based on the timestamp of the frame, and its latitude and longitude information is added to the defect detection result to generate a defect report with geographic tags.

4. The method for accurately locating highway defects according to claim 1, characterized in that, The backbone network in S2 includes: Multi-level convolutional module: including first Conv layer, second Conv layer, third Conv layer, fourth Conv layer and fifth Conv layer; The first C3k2 module includes the first C3k2, the second C3k2, the third C3k2, and the fourth C3k2; SPPF module: Introduces spatial pyramid pooling, which integrates contextual information through pooling windows of different scales to enhance the model's robustness to changes in the target scale; The first Conv layer, second Conv layer, first C3k2 layer, third Conv layer, second C3k2 layer, fourth Conv layer, third C3k2 layer, fifth Conv layer, and fourth C3k2 layer are set sequentially. The first Conv layer performs initial downsampling on the image data to extract shallow texture. Then, the second Conv layer performs downsampling to extract edges / texture. Next, the first C3k2 layer stacks two lightweight residual blocks to enhance the representation. The third Conv layer performs downsampling again to obtain small to medium target feature maps P3. Then, the second C3k2 layer deepens the network depth to improve semantic understanding. The fourth Conv layer performs downsampling on medium to large targets to obtain medium target feature maps P4. Then, the third C3k2 layer enhances the residual path and maintains gradient propagation. After the image data is processed by the fifth Conv layer, the maximum target feature map P5 is obtained. The fourth C3k2 layer enhances the deep features and inputs them into the SPPF module. The SPPF module then inputs the features into the C2PSA module to enhance the expression of key regions, and then inputs them into the neck network.

5. The method for accurately locating highway defects according to claim 1, characterized in that, In S2, the neck network adopts a feature pyramid / path aggregation structure, which achieves bidirectional feature fusion and cross-scale connectivity through top-down feature pyramids and bottom-up path aggregation. The neck network includes a multiple splicing module, a sampling module, and a second C3k2 module; wherein: Multiple concatenation module and sampling module: The multiple concatenation module fuses feature maps from different levels and restores resolution through upsampling to improve the detection capability of the instance segmentation model; the multiple concatenation module includes a first concat, a second concat, a third concat, and a fourth concat, and the sampling module includes a first upsample, a second upsample, a first downsample, and a second downsample; The second C3k2 module includes the fifth, sixth, seventh, and eighth C3k2 modules; The first Upsample, first Concat, fifth C3k2, second Upsample, second Concat, sixth C3k2, first Downsample, third Concat, seventh C3k2, second Downsample, fourth Concat, and eighth C3k2 are set sequentially. The first Upsample performs upsampling to enhance spatial resolution. Then, the first Concat fuses the P4 output from the backbone network. The fifth C3k2 compresses the feature map channels and fuses semantics. Then, the second Upsample further performs upsampling for small target detection. The second Concat fuses the P3 from the backbone network. The sixth C3k2 obtains the final P3. Then, the first Downsample constructs the target detection output path, and the third Concat fuses the P3 from the backbone network. The seventh C3k2 obtains the final P4. Then, the second Downsample constructs the large target detection output path to prepare for large target detection. The fourth Concat fuses the P5 from the backbone network. Finally, the eighth C3k2 obtains the final P5.

6. The method for accurately locating highway defects according to claim 1, characterized in that, The head network in S2 includes: Classification branch: Outputs the target category; Regression branch: Predicts bounding box coordinates (x, y, w, h); Mask branch: Generates pixel-level segmentation masks to achieve precise contour division of each target; Multi-scale output: Outputs detection results through feature maps at multiple levels, adapting to targets of different sizes.

7. The method for accurately locating highway defects according to claim 1, characterized in that: The training sample set in S1 is divided into a training subset, a validation subset, and a test subset according to a set ratio. These subsets are used for parameter learning of the deep learning model, monitoring of the training process, and final performance evaluation, respectively, to ensure that the model has good recognition ability and stability.

8. The method for accurately locating highway defects according to claim 1, characterized in that: In S2, when inferring the instance segmentation model, video frames are scaled to 640×640 and normalized before being input into the instance segmentation model. In the post-processing stage, overlapping boxes are removed using NMS technology, and segmentation boundaries are smoothed through morphological operations.

9. A server, the server comprising a processor and a memory, characterized in that: The memory stores at least one program, which is loaded and executed by the processor to implement the highway defect detection method capable of precise positioning as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing at least one program, characterized in that: The program is loaded and executed by a processor to implement the highway defect detection method capable of precise location as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Road disease identification and tracking method

    CN116485843A

  • Road maintenance intelligent disease identification system based on Beidou positioning

    CN118334628A

  • Intelligent pavement disease detection method and system based on instance segmentation

    CN118864443A

  • Road disease tracking method in complex interference road scene

    CN119206175A

  • Road disease detection method based on scale difference

    CN119206176A

Cited By

  • Road surface disease automatic detection method and system based on deep learning

    CN121708425A

  • A Deep Learning-Based Automated Detection Method and System for Road Surface Defects

    CN121708425B

  • Pavement disease identification method, system and equipment based on semi-supervised learning, and medium

    CN122116155A

  • Road disease identification method, system, device and medium based on semi-supervised learning

    CN122116155B