FPN-fused R-FCN road disease image recognition method and system
By integrating the feature pyramid network and the regional full convolution network, and extracting and fusing multi-scale feature maps, the identification limitations of traditional R-FCN in multi-scale and complex backgrounds are solved, efficient and automated road disease detection is achieved, and detection accuracy and efficiency are significantly improved.
Patent Information
- Application Number
- CN202510075458.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional R-FCNs have limitations in road disease image recognition, especially when dealing with small-scale features in multiple scales and complex contexts, they are prone to missed or missed detection.
The method of fusion feature pyramid network (FPN) and regional full convolutional network (R-FCN) is adopted to extract multi-scale feature maps through convolutional neural network and feature pyramid network, and feature fusion is performed through top-down and horizontal connection methods to generate multi-scale feature maps.
It significantly improves the accuracy and efficiency of road disease detection, can more accurately identify diseases of different scales, reduce missed and missed detection, improves the robustness and automation of detection, and reduces road maintenance costs.
Smart Images

Figure CN119942288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to road quality detection, and in particular to an R-FCN road damage image recognition method and system integrated with FPN. Background Art
[0002] With the continuous acceleration of urbanization, road infrastructure plays a vital role in transportation. However, road damage problems, such as cracks, potholes, and subsidence, pose a serious threat to the service life of roads and driving safety. Traditional road damage detection methods mainly rely on manual inspections, which are not only time-consuming and labor-intensive, but also have problems such as strong subjectivity and low detection efficiency. Therefore, it is of great practical significance to develop an efficient and accurate road damage image recognition method.
[0003] Radar scanning technology has been widely used in the field of road inspection because of its advantages of not being restricted by lighting conditions, strong penetration ability, and high detection accuracy. By performing high-resolution scanning of the road surface with radar equipment, accurate road surface morphology and disease data can be obtained. However, the complexity and noise interference of radar scanning images make it still challenging to directly identify diseased areas from radar images. It is usually necessary to combine advanced image processing and machine learning algorithms to further improve detection accuracy and efficiency.
[0004] With the advancement of deep learning technology, convolutional neural networks (CNNs) and target detection algorithms have made remarkable achievements in the field of image recognition. As an efficient target detection algorithm, the regional fully convolutional network (R-FCN) improves recognition accuracy while maintaining high detection speed through a fully convolutional architecture. The application of R-FCN in road disease detection makes it possible to automatically identify road surface cracks, potholes, deformations and other diseases.
[0005] However, the traditional R-FCN mainly relies on single-scale feature extraction, which has certain limitations when dealing with multi-scale and complex background disease detection. In particular, for small-scale features in road disease images, such as small cracks and tiny potholes, the traditional R-FCN is prone to missed detection or false detection.
[0006] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0007] In view of this, the present invention provides a road damage image recognition method and system integrating FPN with R-FCN to solve the above-mentioned problems.
[0008] In order to solve the above problems, the specific technical solutions adopted by the present invention are as follows:
[0009] According to one aspect of the present invention, a road damage image recognition method of R-FCN fused with FPN is provided, and the recognition method comprises the following steps:
[0010] S1. Collect road images, obtain road images under different disease conditions, preprocess the road images, and construct a road image dataset based on the preprocessing results;
[0011] S2. Extract the road image dataset based on the convolutional neural network and feature pyramid network to generate feature maps at different levels; and fuse the feature maps through the top-down path method and the lateral connection method to obtain multi-scale feature maps;
[0012] S3. Divide the multi-scale feature map through the region proposal network to generate candidate regions and obtain the initial disease candidate frame; detect and process the initial disease candidate frame based on the non-maximum suppression algorithm to obtain the disease candidate frame;
[0013] S4. Use the R-FCN model to perform target detection on the disease candidate box, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; output the disease recognition result based on the R-FCN model, optimize the recognition result through post-processing, and obtain the final recognition result.
[0014] Preferably, extracting the road image dataset based on the convolutional neural network and the feature pyramid network to generate feature maps at different levels; and fusing the feature maps through the top-down path method and the lateral connection method to obtain the multi-scale feature map includes the following steps:
[0015] S21, inputting the road image dataset into a convolutional neural network, and using the convolutional layer of the convolutional neural network to extract features of each image in the road image dataset to obtain feature maps of different depths;
[0016] S22, sorting the feature maps according to their depths, constructing a feature pyramid network according to the sorting results, and taking the deepest feature map as the starting point, upsampling the feature map based on a top-down path;
[0017] S23, horizontally connect the up-sampled feature map with the feature map of the previous layer, and fuse the feature maps using convolution operations to obtain multi-scale feature maps at different levels.
[0018] Preferably, the multi-scale feature map is divided and processed by a region proposal network to generate candidate regions to obtain an initial defect candidate frame; the initial defect candidate frame is detected and redundantly processed based on a non-maximum suppression algorithm to obtain the defect candidate frame, including the following steps:
[0019] S31, input the multi-scale feature map into the region proposal network, and use the Anchor mechanism to generate candidate regions on each multi-scale feature map;
[0020] S32, calculating the disease classification probability and performing bounding box regression for each generated candidate region, and optimizing the region proposal network using a multi-task loss function based on the output of the region proposal network and the true annotation box;
[0021] S33, sorting all candidate regions according to the confidence of disease classification, and screening out candidate regions that meet the confidence threshold to obtain an initial disease candidate frame;
[0022] S34, performing a non-maximum suppression algorithm on the initial defect candidate frame, and obtaining a defect candidate frame by detecting and filtering redundant initial defect candidate frames.
[0023] Preferably, the expression of the multi-task loss function is:
[0024] ;
[0025] In the formula, L(p i ,t i ) represents the multi-task loss function;
[0026] L cls represents the classification function;
[0027] L reg represents the bounding box regression function;
[0028] p i represents the disease classification probability of the i-th candidate area;
[0029] represents the true label;
[0030] t i represents the translation and scaling parameters of the regression prediction;
[0031] Parameters representing the true bounding box;
[0032] N cls and N reg Represents the number of terms used to normalize the classification function and the bounding box regression function respectively;
[0033] γ represents the weight.
[0034] Preferably, performing a non-maximum suppression algorithm on the initial defect candidate frame, and obtaining the defect candidate frame by detecting and filtering redundant initial defect candidate frames comprises the following steps:
[0035] S341, selecting a preliminary candidate region with the highest confidence in the initial disease candidate frame, and calculating an intersection-over-union ratio between the preliminary candidate region with the highest confidence and the remaining preliminary candidate regions;
[0036] S342, removing the preliminary candidate regions whose intersection-over-union ratio is greater than a preset intersection-over-union ratio threshold, and updating the preliminary candidate region set;
[0037] S343, selecting the next preliminary candidate region with the highest confidence from the updated preliminary candidate region set, and calculating the intersection-over-union ratio between the next preliminary candidate region with the highest confidence and the remaining preliminary candidate regions;
[0038] S344, repeating steps S342 and S343 until the preliminary candidate region set reaches a preset stop condition, and taking the remaining preliminary candidate regions as disease candidate frames.
[0039] Preferably, the R-FCN model is used to perform target detection on the disease candidate box, a position-sensitive score map is generated, and the final category of the disease is determined through a voting mechanism; and the disease recognition result is output based on the R-FCN model, and the recognition result is optimized through post-processing to obtain the final recognition result, which includes the following steps:
[0040] S41. For each disease candidate box, the disease candidate box is divided into several sub-regions using the R-FCN model, and a position sensitive score map is calculated through a convolution operation;
[0041] S42, mapping the disease candidate box to a grid of a preset size through a region of interest pooling operation, and for the score map of each sub-region, calculating the final category score of the disease candidate box through a summary voting mechanism to obtain a category prediction result of the disease candidate box;
[0042] S43, performing post-processing optimization on the category prediction results of all disease candidate frames to obtain the final disease recognition result.
[0043] Preferably, the calculation formula for calculating the position sensitive score through the convolution operation is:
[0044] ;
[0045] In the formula, represents the output of the i-th position-sensitive score map;
[0046] f represents the convolution operation;
[0047] X represents the disease candidate frame;
[0048] Represents the convolution parameters for generating the i,j-th score map;
[0049] cls represents the disease classification label.
[0050] Preferably, the disease candidate box is mapped to a grid of a preset size through a region of interest pooling operation, and for the score map of each sub-region, the final category score of the disease candidate box is calculated through a summary voting mechanism, and the category prediction result of the disease candidate box is obtained, which includes the following steps:
[0051] S421, mapping the features of the disease candidate frame to a grid of a preset size, and performing a maximum pooling operation on each sub-region of the disease candidate frame;
[0052] S422, for each sub-region of the disease candidate box, extracting a corresponding position sensitive score from the position sensitive score map;
[0053] S423. According to the position-sensitive score of each sub-region, the final category score of the disease candidate box is calculated through a summary voting mechanism.
[0054] Preferably, the calculation formula for calculating the final category score of the disease candidate box through the aggregate voting mechanism is:
[0055] ;
[0056] In the formula, Score cls Represents the final category score of the disease candidate box;
[0057] w i,j represents the weight of the i, j-th sub-region;
[0058] represents the output of the i-th position-sensitive score map;
[0059] k represents the dimension of sub-region division.
[0060] According to another aspect of the present invention, a road damage image recognition system integrating FPN and R-FCN is provided, the recognition system comprising: a data construction module, a multi-scale feature map analysis module, a damage candidate frame analysis module and a recognition result analysis module;
[0061] A data construction module is used to collect road surface images, obtain road images in different disease states, preprocess the road images, and construct a road image dataset based on the preprocessing results;
[0062] The multi-scale feature map analysis module is used to extract the road image data set based on the convolutional neural network and the feature pyramid network to generate feature maps at different levels; and to fuse the feature maps through the top-down path method and the lateral connection method to obtain the multi-scale feature map;
[0063] The defect candidate frame analysis module is used to divide the multi-scale feature map through the region proposal network, generate candidate regions, and obtain the initial defect candidate frame; the initial defect candidate frame is detected and redundantly processed based on the non-maximum suppression algorithm to obtain the defect candidate frame;
[0064] The recognition result analysis module is used to perform target detection on the disease candidate box using the R-FCN model, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; and output the disease recognition result based on the R-FCN model, optimize the recognition result through post-processing, and obtain the final recognition result.
[0065] The beneficial effects of the present invention are:
[0066] 1. The present invention realizes efficient and automated detection of multi-scale road defects by integrating feature pyramid network and regional full convolution network. Compared with traditional methods, it has stronger multi-scale feature extraction capability, significantly improves detection accuracy and efficiency, and has high robustness, especially in complex environments. Through denoising, contrast enhancement and data augmentation, it can cope with interference from different environments, realize accurate disease classification and positioning, reduce missed detection and false detection, and effectively improve the recognition accuracy of diseases of different scales and reduce missed detection rate through multi-level feature extraction. At the same time, it is highly automated, reduces manual intervention, reduces road maintenance costs, and has wide applicability and scalability.
[0067] 2. When facing complex road scenes, noise interference or severe weather conditions, the R-FCN integrated with FPN can still maintain a high recognition accuracy and is suitable for different types of road disease detection tasks. It can identify road internal diseases faster and more accurately, improve the efficiency and accuracy of road monitoring, and achieve a better balance when processing multi-scale target detection. It not only significantly improves the detection accuracy, but also maintains a high computational efficiency, which is helpful for road safety and maintenance.
[0068] 3. The present invention realizes efficient feature sharing through the fully convolutional structure of R-FCN, so that all candidate boxes can perform convolution calculations on the same feature map, significantly reducing redundant calculations and improving detection speed. Without sacrificing detection accuracy, the model has real-time processing capabilities and is suitable for large-scale road disease detection tasks.
[0069] 4. The present invention adopts adaptive denoising, contrast enhancement and data augmentation technology, which can enhance the contrast of the diseased area and improve the model's anti-interference ability to diverse noises. It still maintains high recognition rate and stability when dealing with complex environments with different lighting, shadows and weather changes.
[0070] 5. The present invention realizes full process automation, from the acquisition and preprocessing of radar scanning images to the identification and classification of defects, without the need for human intervention. It can accurately classify and locate defects and distinguish various types of defects (such as cracks, spalling, potholes, etc.), thereby improving the accuracy and reliability of detection and effectively reducing the burden of manual inspections.
[0071] 6. The architecture of the present invention is highly flexible and can adapt to different types of radar equipment and road scenarios. It is not only suitable for disease detection on highways and urban roads, but can also be extended to detection tasks of other infrastructure such as airport runways and bridges. At the same time, the algorithm structure of the method is easy to expand and can be combined with other deep learning models or detection frameworks to further improve the detection effect. Its high efficiency, accuracy and low cost make it have broad application prospects in smart city construction and infrastructure management, and can provide technical support and guarantee for related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0073] Figure 1 is a flow chart of a method for recognizing road damage images using R-FCN and FPN according to an embodiment of the present invention;
[0074] Figure 2 It is a principle block diagram of the R-FCN road damage image recognition system integrated with FPN according to an embodiment of the present invention.
[0075] In the figure:
[0076] 1. Data construction module; 2. Multi-scale feature map analysis module; 3. Disease candidate box analysis module; 4. Recognition result analysis module. DETAILED DESCRIPTION
[0077] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0078] According to an embodiment of the present invention, a road damage image recognition method and system using R-FCN integrated with FPN are provided.
[0079] The present invention is further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, according to one embodiment of the present invention, a road damage image recognition method integrating FPN with R-FCN is provided, and the recognition method comprises the following steps:
[0080] S1. Collect road images, obtain road images under different disease conditions, preprocess the road images, and construct a road image dataset based on the preprocessing results;
[0081] It should be noted that the images of road internal defects collected by radar scanning equipment should at least include the following types: cracks, settlements, voids, cavities, fissures or broken images. The collected data will serve as input for subsequent recognition algorithms.
[0082] Among them, radar equipment can penetrate the surface layer and capture subtle disease information inside the road. The collection of radar image data includes the following common disease types:
[0083] Cracks: Includes fine cracks and deep cracks that appear on the road surface.
[0084] Settlement: Deformation of a local area of a road due to subsidence of the roadbed.
[0085] Void: Holes or missing materials in the road's substructure.
[0086] Void: An underground cavity created by external forces or geological factors, causing the road surface to collapse.
[0087] Crushed: The surface appears to be broken or peeling.
[0088] Specifically, the road images scanned by radar are preprocessed, including denoising, contrast enhancement, and data augmentation, to improve the robustness of the model and obtain a data set for training.
[0089] Denoising: The median filter method is used to process the noise in the radar image. The median filter can effectively remove the noise, retain the edge details of the diseased area, and reduce the interference of environmental noise on the image.
[0090] Contrast enhancement: The contrast of radar images is enhanced using the adaptive histogram equalization (CLAHE) method; CLAHE can enhance the grayscale contrast locally and highlight the features of the diseased area, thereby improving the detection effect.
[0091] Data augmentation: In order to improve the generalization ability of the model, the data set is expanded in a variety of ways, including image flipping, rotation, scaling, cropping, adding noise, adjusting brightness and contrast, etc. These augmentation operations can simulate the characteristics of road damage under different environmental conditions and improve the model's ability to recognize a variety of damages.
[0092] S2. Extract the road image dataset based on the convolutional neural network and feature pyramid network to generate feature maps at different levels; and fuse the feature maps through the top-down path method and the lateral connection method to obtain multi-scale feature maps;
[0093] It should be noted that the core algorithm model of FPN is based on the concept of multi-scale feature representation. In traditional convolutional neural networks, as the layers go deeper, the spatial resolution of the feature map gradually decreases, but the receptive field increases, making the network have more abstract features at high levels, while the low-level feature maps maintain a high resolution but lack abstract expression. FPN fuses these different levels of features through top-down paths and lateral connections to generate feature maps with multi-scale features.
[0094] It is usually combined with a deep convolutional neural network (ResNet) and uses residual connections to enable deeper networks to maintain gradient propagation during feature extraction, solving the gradient vanishing problem that is prone to occur in deep networks. The residual block in ResNet is described by the formula:
[0095] y=F(x,{W i})+x;
[0096] Among them, y represents output, x represents input, F(x,{W i}) represents the transformation after passing through several convolutional layers, W i Represents the weights of the convolutional layer.
[0097] As a preferred implementation, the road image dataset is extracted based on a convolutional neural network and a feature pyramid network to generate feature maps at different levels; and the feature maps are fused through a top-down path and a lateral connection method to obtain a multi-scale feature map, including the following steps:
[0098] S21, inputting the road image dataset into a convolutional neural network, and using the convolutional layer of the convolutional neural network to extract features of each image in the road image dataset to obtain feature maps of different depths;
[0099] S22, sorting the feature maps according to their depths, constructing a feature pyramid network according to the sorting results, and taking the deepest feature map as the starting point, upsampling the feature map based on a top-down path;
[0100] S23, horizontally connect the up-sampled feature map with the feature map of the previous layer, and fuse the feature maps using convolution operations to obtain multi-scale feature maps at different levels.
[0101] It should be noted that the feature fusion in FPN adopts a top-down path and lateral connections. The top-down path gradually upsamples the feature maps of the deeper layers of the network, and the lateral connection fuses the upsampled feature maps of each layer with the corresponding shallow feature maps to generate multi-scale features.
[0102] P i =Conv(C i )+UpSample(P i+1 );
[0103] Where P i represents the feature map generated by the i-th layer, C i Represents the original convolutional feature map, Conv(C i ) indicates that C i Perform 1x1 convolution operation, UpSample(P i+1 ) indicates upsampling the feature map of the previous layer.
[0104] For example, the following are the steps for the Feature Pyramid Network (FPN) to generate a multi-level feature map:
[0105] (1) Backbone network (ResNet) feature extraction: The input image passes through several convolutional layers of the ResNet backbone network to generate feature maps of different depths. Assuming that ResNet has five stages, the output feature maps are denoted as C1, C2, C3, C4, and C5 respectively. The spatial resolution of these feature maps decreases layer by layer, but the abstract expression increases layer by layer. The feature maps of the shallower layers have higher resolution and more detailed information, while the feature maps of the deeper layers have stronger semantic information.
[0106] (2) Top-down path: Starting from the deepest feature map C5, the corresponding multi-scale feature map P5 is generated and used as the starting point of the top-down path. Then, P5 is up-sampled to increase its resolution to the same size as C4.
[0107] (3) Horizontal connection: The upsampled P5 is horizontally connected to the shallow feature map C4, and 1x1 convolution is used to reduce the number of channels so that the number of channels of the two feature maps matches. Then they are added to obtain a new feature map P4. This process is repeated, P4 is upsampled again and horizontally connected to C3 to generate P3, and so on, until the feature maps P1 of all levels are generated.
[0108] (4) Output multi-scale feature maps: The final P5, P4, P3, P2 and P1 are feature maps of different scales, which represent features from coarse-grained to fine-grained. These feature maps can be input into subsequent detection or classification networks to identify targets of different scales.
[0109] Among them, multi-scale feature extraction: The FPN structure generates detailed low-level features from shallow feature maps for detecting small defects (such as small cracks), and extracts higher-level semantic information from deep feature maps for detecting large defects (such as subsidence, potholes, etc.). FPN builds a feature pyramid from the bottom up and fuses deep features with shallow features through upsampling to ensure that defects of different scales can be accurately detected.
[0110] S3. Divide the multi-scale feature map through the region proposal network to generate candidate regions and obtain the initial disease candidate frame; detect and process the initial disease candidate frame based on the non-maximum suppression algorithm to obtain the disease candidate frame;
[0111] It should be noted that the candidate regions are generated based on the multi-scale feature map generated by RPN (Region Proposal Network), the candidate frames of the disease are obtained, and the redundant candidate frames are removed by the NMS (Non-Maximum Suppression) algorithm to retain the most representative disease regions. RPN ensures the efficiency and accuracy of disease detection.
[0112] Among them, RPN is a lightweight convolutional neural network used to generate candidate regions (Region Proposals) in target detection tasks. Compared with traditional sliding window and selective search algorithms, RPN can learn to generate candidate boxes directly from feature maps, which makes the detection process more efficient. RPN does not directly output Anchor as the final candidate region, but performs bounding box regression on the generated Anchor. This regression model adjusts the size and position of the Anchor by learning the displacement and scaling between the Anchor and the ground truth, so that it can locate the target more accurately. RPN uses an end-to-end training method to closely integrate the feature extraction and candidate box generation steps, significantly improving the speed and accuracy of detection.
[0113] Specifically, the formula for bounding box regression is as follows:
[0114] t x =(xx a ) / w a ;
[0115] t y =(yy a ) / h a ;
[0116] t w =log(w / w a );
[0117] t h =log(h / h a );
[0118] Where, t x , t y , t w , t h represents the predicted translation and scaling parameters, x, y, w, h represent the center coordinates and width and height of the Ground Truth bounding box, x a ,y a , w a ,h a Indicates the center coordinates and width and height of the Anchor.
[0119] As a preferred implementation, the multi-scale feature map is divided and processed by a region proposal network to generate candidate regions and obtain an initial defect candidate frame; the initial defect candidate frame is detected and redundantly processed based on a non-maximum suppression algorithm to obtain the defect candidate frame, including the following steps:
[0120] S31, input the multi-scale feature map into the region proposal network, and use the Anchor mechanism to generate candidate regions on each multi-scale feature map;
[0121] It should be noted that RPN uses the Anchor mechanism to generate candidate regions. On the input feature map, RPN uses a sliding window (such as a 3×3 convolution kernel) to generate anchors, and each sliding window position generates multiple anchors of different sizes and ratios. These anchors represent candidate regions on the feature map, covering different target sizes and aspect ratios.
[0122] Anchor size: There are several preset anchors of different sizes (such as 128×128, 256×256, 512×512), as well as multiple aspect ratios (such as 1:1, 1:2, 2:1).
[0123] Number of candidate boxes generated: For each feature map position, the number of anchors generated by RPN is equal to the number of combinations of different scales and aspect ratios.
[0124] S32, calculating the disease classification probability and performing bounding box regression for each generated candidate region, and optimizing the region proposal network using a multi-task loss function based on the output of the region proposal network and the true annotation box;
[0125] It should be noted that the goal of RPN is to classify and regress Ancho, so the loss function contains two parts: classification loss and bounding box regression loss. Generally, a multi-task loss function is used to optimize the network. The classification loss is combined with the bounding box regression loss to update the model parameters through back propagation. In order to improve the convergence and stability of the model, the cosine annealing learning rate is used to gradually reduce the learning rate to ensure that the model can converge smoothly in the later stage of training and avoid overfitting.
[0126] As a preferred implementation, the expression of the multi-task loss function is:
[0127] ;
[0128] In the formula, L(p i ,t i ) represents the multi-task loss function; L cls Represents the classification function, using cross entropy loss to calculate whether each Anchor contains the target; L reg represents the bounding box regression function, usually using Smooth L1Loss, which is used to optimize the bounding box position of the Anchor; p i represents the disease classification probability of the i-th candidate area; represents the true label (positive and negative sample label); t i represents the translation and scaling parameters of the regression prediction; Represents the parameters of the true bounding box; N cls and N reg They represent the number of terms used to normalize the classification function and the bounding box regression function respectively; γ represents the weight used to balance the impact of classification loss and regression loss.
[0129] S33, sorting all candidate regions according to the confidence of disease classification, and screening out candidate regions that meet the confidence threshold to obtain an initial disease candidate frame;
[0130] S34, performing a non-maximum suppression algorithm on the initial defect candidate frame, and obtaining a defect candidate frame by detecting and filtering redundant initial defect candidate frames.
[0131] As a preferred implementation, performing a non-maximum suppression algorithm on the initial defect candidate frame, and obtaining the defect candidate frame by detecting and filtering redundant initial defect candidate frames includes the following steps:
[0132] S341, selecting a preliminary candidate region with the highest confidence in the initial disease candidate frame, and calculating an intersection-over-union ratio between the preliminary candidate region with the highest confidence and the remaining preliminary candidate regions;
[0133] S342, removing the preliminary candidate regions whose intersection-over-union ratio is greater than a preset intersection-over-union ratio threshold, and updating the preliminary candidate region set;
[0134] S343, selecting the next preliminary candidate region with the highest confidence from the updated preliminary candidate region set, and calculating the intersection-over-union ratio between the next preliminary candidate region with the highest confidence and the remaining preliminary candidate regions;
[0135] S344, repeating steps S342 and S343 until the preliminary candidate region set reaches a preset stop condition, and taking the remaining preliminary candidate regions as disease candidate frames.
[0136] It should be noted that NMS (Non-Maximum Suppression): After RPN generates a large number of candidate boxes, the NMS algorithm is used to filter out redundant candidate boxes (preliminary candidate areas) and retain the most representative areas. NMS (Non-Maximum Suppression) is a post-processing algorithm used to filter redundant detection boxes in target detection tasks. The core idea is to retain the detection box with the highest score (confidence) while removing other boxes with a high overlap rate, thereby ensuring that each target is labeled with only one detection box. The algorithm decides whether to remove certain detection boxes by calculating the overlap between detection boxes (usually through IoU, Intersection over Union). The core of NMS is to calculate the overlap between detection boxes, which is usually measured by the intersection over union (IoU). The formula for IoU is:
[0137] ;
[0138] In the formula, A and B are two candidate boxes;
[0139] Area of Intersection represents the intersection area of two candidate boxes;
[0140] Area of Union represents the union area of two candidate boxes, and the calculation formula is:
[0141] Area of Union(A,B)=Area(A)+Area(B)-Area of Intersection(A,B);
[0142] The execution steps of NMS are as follows:
[0143] S1. Sort all candidate boxes according to classification confidence.
[0144] S2: Select the candidate box with the highest confidence as the detection result, and remove other candidate boxes whose overlapping area is greater than a certain threshold (such as 0.5).
[0145] S3. Repeat this process until the candidate box list is empty.
[0146] S4. Use the R-FCN model to perform target detection on the disease candidate box, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; output the disease recognition result based on the R-FCN model, optimize the recognition result through post-processing, and obtain the final recognition result.
[0147] It should be noted that R-FCN generates a position-sensitive score map for disease classification of each candidate region. The score map maps the candidate box region to the feature map and classifies and votes based on the features in each region to determine the final category of the disease.
[0148] Through the fully convolutional structure, R-FCN achieves feature sharing, thereby improving the computational efficiency of the model. Each candidate box is detected on the shared feature map, which greatly reduces the computational overhead.
[0149] The core innovation of the R-FCN model is to divide each candidate region (Ro) into multiple sub-regions and generate different score maps for each sub-region. In this way, each sub-region has its corresponding "position-sensitive" score, thus avoiding a large number of fully connected layer operations in traditional object detection networks. The position-sensitive score map is closely related to the spatial position of the image and is responsible for predicting the category score of each sub-region. The specific generation process is as follows:
[0150] R-FCN generates different category prediction responses for each sub-region of the candidate region by generating multiple position-sensitive score maps. Each pixel value of the score map represents the prediction score of a certain category at that position, and these local scores are aggregated for classification voting. Assuming that the feature map of the input image is X, in the R-FCN model, the feature map generates score maps of multiple categories through a series of convolution operations. For each category, k×k position-sensitive score maps are generated.
[0151] As a preferred implementation, the R-FCN model is used to detect the target of the disease candidate box, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; and the disease recognition result is output based on the R-FCN model, and the recognition result is optimized through post-processing. The final recognition result includes the following steps:
[0152] S41. For each disease candidate box, the disease candidate box is divided into several sub-regions using the R-FCN model, and a position sensitive score map is calculated through a convolution operation;
[0153] As a preferred implementation, the calculation formula for calculating the position-sensitive score through the convolution operation is:
[0154] ;
[0155] In the formula, represents the output of the i-th position-sensitive score map; f represents the convolution operation; X represents the disease candidate box; represents the convolution kernel parameters for generating the i,j-th score map; cls represents the disease classification label.
[0156] S42, mapping the disease candidate box to a grid of a preset size through a region of interest pooling operation, and for the score map of each sub-region, calculating the final category score of the disease candidate box through a summary voting mechanism to obtain a category prediction result of the disease candidate box;
[0157] As a preferred implementation, the disease candidate box is mapped to a grid of a preset size through a region of interest pooling operation. For the score map of each sub-region, the final category score of the disease candidate box is calculated through a summary voting mechanism. The category prediction result of the disease candidate box is obtained, which includes the following steps:
[0158] S421, mapping the features of the disease candidate frame to a grid of a preset size, and performing a maximum pooling operation on each sub-region of the disease candidate frame;
[0159] S422, for each sub-region of the disease candidate box, extracting a corresponding position sensitive score from the position sensitive score map;
[0160] S423. According to the position-sensitive score of each sub-region, the final category score of the disease candidate box is calculated through a summary voting mechanism.
[0161] It should be noted that RoI Pooling (region of interest pooling operation) is combined with position-sensitive score map; RoIPooling maps the feature map of each candidate region to a fixed-size grid by downsampling, and the response of each position-sensitive score map corresponds to a specific sub-region of Ro. This mapping maintains spatial alignment:
[0162]
[0163] In the formula, represents the feature of the i-th position on the RoI, and R represents the position parameter of the candidate region (i.e., its bounding box in the original image).
[0164] The core idea of the voting mechanism is to aggregate the category prediction results of each sub-region and make the final classification decision based on the aggregated scores. In this way, the prediction error that may occur in a single location can be avoided and the robustness of the overall prediction can be improved. If the contributions of the sub-regions are uneven, a weighted voting mechanism can be used. That is, a weight w is assigned to each sub-region according to some strategy. i, the final category score can be expressed as:
[0165] ;
[0166] In the formula, Score cls represents the final category score of the disease candidate box; w i,j represents the weight of the ith and jth sub-regions, which can be dynamically adjusted based on the position, feature response strength, etc.; represents the output of the i-th position-sensitive score map; k represents the division dimension of the sub-region.
[0167] The final classification decision is made by comparing the summary scores of all categories. For each candidate region, the final prediction of category cls is:
[0168] ;
[0169] In the formula, Represents the final prediction result of category cls, and the category with the highest score is selected as the classification result of the candidate region.
[0170] S43, performing post-processing optimization on the category prediction results of all disease candidate frames to obtain the final disease recognition result.
[0171] It should be noted that in order to further improve the detection effect, the NMS algorithm is used to remove overlapping detection frames and filter out low-confidence detection results, with a threshold value of 0.5. Finally, a disease detection report is generated, outputting accurate disease location and classification information.
[0172] The core idea of NMS is: for multiple overlapping detection boxes, only the box with the highest confidence is retained, and other overlapping boxes are suppressed to ensure that each target is detected only once.
[0173] Overlapping detection box problem: In the object detection task, the detector generates multiple bounding boxes that may contain the object. These bounding boxes are highly overlapped in a certain area, but they may correspond to the same object. If the overlapping detection boxes are not processed, the same object may be detected repeatedly.
[0174] The purpose of NMS is to suppress overlapping low-confidence boxes by filtering out the boxes with the highest confidence. The algorithm decides whether to suppress a box based on the confidence score and overlapping area of each box. It determines the degree of overlap between two detection boxes by calculating the intersection over union (loU) of the overlapping area.
[0175] Suppression strategy: If the IoU of two boxes is greater than a certain threshold (for example, 0.5), the two boxes are considered to have a high degree of overlap. NMS will retain the boxes with higher confidence and suppress (delete) the boxes with lower confidence. For each category of detection boxes, all detection boxes are first sorted according to the confidence score. The sorted detection boxes are sequentially subjected to the following operations:
[0176] Filter low confidence detections: Before NMS processing, the detection results are usually initially filtered. Set a confidence threshold Tconf. If the confidence of a detection box is lower than this threshold, the box is directly deleted:
[0177] RemoveB i ,if s i <T conf ;
[0178] Among them, B i represents bounding boxes arranged from high to low confidence; s i represents the confidence score; T conf Represents the confidence threshold.
[0179] Specifically, in order to verify the effectiveness of the present invention, multiple experimental indicators were used for evaluation, including:
[0180] Mean Average Precision (mAP): used to evaluate the overall accuracy of the model in different disease detection tasks.
[0181] Recall rate: used to measure the detection rate of the model for diseased areas.
[0182] F1 score: combines precision and recall to evaluate the overall performance of the model, especially its performance in small-scale disease detection.
[0183] like Figure 2 As shown, according to another embodiment of the present invention, a road damage image recognition system integrating FPN and R-FCN is provided, the recognition system comprising: a data construction module 1, a multi-scale feature map analysis module 2, a damage candidate frame analysis module 3 and a recognition result analysis module 4;
[0184] Data construction module 1, used to collect road surface images, obtain road images in different disease states, preprocess the road images, and construct a road image data set based on the preprocessing results;
[0185] Multi-scale feature map analysis module 2 is used to extract the road image data set based on the convolutional neural network and the feature pyramid network to generate feature maps at different levels; and to fuse the feature maps through the top-down path method and the lateral connection method to obtain the multi-scale feature map;
[0186] The defect candidate frame analysis module 3 is used to divide the multi-scale feature map through the region proposal network, generate candidate regions, and obtain the initial defect candidate frame; detect and perform redundancy processing on the initial defect candidate frame based on the non-maximum suppression algorithm to obtain the defect candidate frame;
[0187] The recognition result analysis module 4 is used to perform target detection on the disease candidate box using the R-FCN model, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; and output the disease recognition result based on the R-FCN model, optimize the recognition result through post-processing, and obtain the final recognition result.
[0188] In summary, with the help of the above technical solution of the present invention, the present invention realizes efficient and automated detection of multi-scale road diseases by fusing feature pyramid network and regional full convolution network. Compared with traditional methods, it has stronger multi-scale feature extraction ability, significantly improves detection accuracy and efficiency, especially in complex environments, has high robustness, can cope with interference from different environments through denoising, contrast enhancement and data augmentation, realizes accurate disease classification and positioning, reduces missed detection and false detection, and effectively improves the recognition accuracy of diseases of different scales through multi-level feature extraction, reduces missed detection rate, and at the same time, is highly automated, reduces manual intervention, reduces road maintenance costs, and has wide applicability and scalability. In the face of complex road scenes, noise interference or severe weather conditions, the present invention can still maintain a high recognition accuracy by fusing FPN R-FCN, which is suitable for different types of road disease detection tasks, can identify road internal diseases faster and more accurately, improve the efficiency and accuracy of road monitoring, and can achieve a better balance when processing multi-scale target detection, which not only significantly improves detection accuracy, but also maintains a high computational efficiency, which is helpful for road safety and maintenance. The present invention realizes efficient feature sharing through the full convolution structure of R-FCN, so that all candidate boxes can perform convolution calculations on the same feature map, significantly reducing redundant calculations and improving detection speed. Without sacrificing detection accuracy, the model has real-time processing capabilities and is suitable for large-scale road disease detection tasks. The present invention adopts adaptive denoising, contrast enhancement and data augmentation technologies, which can enhance the contrast of the diseased area and improve the model's anti-interference ability to diversified noise. The system still maintains high recognition rate and stability when dealing with complex environments with different lighting, shadows and weather changes. The present invention realizes full process automation, from the acquisition and preprocessing of radar scanning images to disease identification and classification, without manual intervention, and can accurately classify and locate diseases, distinguish various types of diseases (such as cracks, peeling, potholes, etc.), improve the accuracy and reliability of detection, and effectively reduce the burden of manual inspections. The architecture of the present invention is highly flexible and can adapt to different types of radar equipment and road scenarios. It is not only suitable for disease detection on highways and urban roads, but can also be extended to detection tasks of other infrastructure such as airport runways and bridges. At the same time, the algorithm structure of the method is easy to expand and can be combined with other deep learning models or detection frameworks to further improve the detection effect. Its high efficiency, accuracy and low cost make it have broad application prospects in smart city construction and infrastructure management, and can provide technical support and guarantee for related fields.
[0189] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0190] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. The R-FCN road disease image recognition method fused with FPN is characterized by: The identification method comprises the following steps: S1. Collect road images, obtain road images under different disease conditions, preprocess the road images, and construct a road image dataset based on the preprocessing results; S2. Extract the road image dataset based on the convolutional neural network and feature pyramid network to generate feature maps at different levels; and fuse the feature maps through the top-down path method and the lateral connection method to obtain multi-scale feature maps; S3. Divide the multi-scale feature map through the region proposal network to generate candidate regions and obtain the initial disease candidate frame; detect and process the initial disease candidate frame based on the non-maximum suppression algorithm to obtain the disease candidate frame; S4. Use the R-FCN model to perform target detection on the disease candidate box, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; output the disease recognition result based on the R-FCN model, optimize the recognition result through post-processing, and obtain the final recognition result.
2. The R-FCN road damage image recognition method fused with FPN according to claim 1 is characterized in that: The method extracts the road image data set based on the convolutional neural network and the feature pyramid network to generate feature maps at different levels; and fuses the feature maps through a top-down path method and a lateral connection method to obtain a multi-scale feature map, including the following steps: S21, inputting the road image dataset into a convolutional neural network, and using the convolutional layer of the convolutional neural network to extract features of each image in the road image dataset to obtain feature maps of different depths; S22, sorting the feature maps according to their depths, constructing a feature pyramid network according to the sorting results, and taking the deepest feature map as the starting point, upsampling the feature map based on a top-down path; S23, horizontally connect the up-sampled feature map with the feature map of the previous layer, and fuse the feature maps using convolution operations to obtain multi-scale feature maps at different levels.
3. The R-FCN road damage image recognition method fused with FPN according to claim 1 is characterized in that: The multi-scale feature map is divided and processed by the region proposal network to generate candidate regions and obtain initial disease candidate frames; The initial defect candidate frame is detected and redundantly processed based on the non-maximum suppression algorithm, and the defect candidate frame is obtained by the following steps: S31, input the multi-scale feature map into the region proposal network, and use the Anchor mechanism to generate candidate regions on each multi-scale feature map; S32, calculating the disease classification probability and performing bounding box regression for each generated candidate region, and optimizing the region proposal network using a multi-task loss function based on the output of the region proposal network and the true annotation box; S33, sorting all candidate regions according to the confidence of disease classification, and screening out candidate regions that meet the confidence threshold to obtain an initial disease candidate frame; S34, performing a non-maximum suppression algorithm on the initial defect candidate frame, and obtaining a defect candidate frame by detecting and filtering redundant initial defect candidate frames.
4. The R-FCN road damage image recognition method fused with FPN according to claim 3 is characterized in that: The expression of the multi-task loss function is: ; In the formula, L(p i ,t i ) represents the multi-task loss function; L cls represents the classification function; L reg represents the bounding box regression function; p i represents the disease classification probability of the i-th candidate area; represents the true label; t i represents the translation and scaling parameters of the regression prediction; Parameters representing the true bounding box; N cls and N reg Represents the number of terms used to normalize the classification function and the bounding box regression function respectively; γ represents the weight.
5. The R-FCN road damage image recognition method fused with FPN according to claim 3 is characterized in that: The non-maximum suppression algorithm is executed on the initial defect candidate frame to obtain the defect candidate frame by detecting and filtering the redundant initial defect candidate frames, which includes the following steps: S341, selecting a preliminary candidate region with the highest confidence in the initial disease candidate frame, and calculating an intersection-over-union ratio between the preliminary candidate region with the highest confidence and the remaining preliminary candidate regions; S342, removing the preliminary candidate regions whose intersection-over-union ratio is greater than a preset intersection-over-union ratio threshold, and updating the preliminary candidate region set; S343, selecting the next preliminary candidate region with the highest confidence from the updated preliminary candidate region set, and calculating the intersection-over-union ratio between the next preliminary candidate region with the highest confidence and the remaining preliminary candidate regions; S344, repeating steps S342 and S343 until the preliminary candidate region set reaches a preset stop condition, and taking the remaining preliminary candidate regions as disease candidate frames.
6. The method for recognizing road damage images by using R-FCN and FPN as claimed in claim 1, characterized in that: The method of using the R-FCN model to perform target detection on the disease candidate box, generating a position-sensitive score map, and determining the final category of the disease through a voting mechanism; and outputting the disease recognition result based on the R-FCN model, and optimizing the recognition result through post-processing to obtain the final recognition result includes the following steps: S41. For each disease candidate box, the disease candidate box is divided into several sub-regions using the R-FCN model, and a position sensitive score map is calculated through a convolution operation; S42, mapping the disease candidate box to a grid of a preset size through a region of interest pooling operation, and for the score map of each sub-region, calculating the final category score of the disease candidate box through a summary voting mechanism to obtain a category prediction result of the disease candidate box; S43, performing post-processing optimization on the category prediction results of all disease candidate frames to obtain the final disease recognition result.
7. The method for recognizing road damage images by using R-FCN and FPN as claimed in claim 6, characterized in that: The calculation formula for calculating the position-sensitive score through the convolution operation is: ; In the formula, represents the output of the i-th position-sensitive score map; f represents the convolution operation; X represents the disease candidate frame; Represents the convolution parameters for generating the i,j-th score map; cls represents the disease classification label.
8. The method for recognizing road damage images by using R-FCN and FPN as claimed in claim 6, characterized in that: The process of mapping the disease candidate frame to a grid of a preset size through a pooling operation of the region of interest, and calculating the final category score of the disease candidate frame through a summary voting mechanism for the score map of each sub-region, and obtaining the category prediction result of the disease candidate frame includes the following steps: S421, mapping the features of the disease candidate frame to a grid of a preset size, and performing a maximum pooling operation on each sub-region of the disease candidate frame; S422, for each sub-region of the disease candidate box, extracting a corresponding position sensitive score from the position sensitive score map; S423. According to the position-sensitive score of each sub-region, the final category score of the disease candidate box is calculated through a summary voting mechanism.
9. The method for recognizing road damage images by using R-FCN and FPN as claimed in claim 8, characterized in that: The calculation formula for calculating the final category score of the disease candidate frame through the aggregate voting mechanism is: ; In the formula, Score cls Represents the final category score of the disease candidate box; w i,j represents the weight of the i, j-th sub-region; represents the output of the i-th position-sensitive score map; k represents the dimension of sub-region division.
10. A FPN-integrated R-FCN road damage image recognition system, used to implement the FPN-integrated R-FCN road damage image recognition method according to any one of claims 1 to 9, characterized in that: The recognition system includes: a data construction module, a multi-scale feature map analysis module, a disease candidate frame analysis module and a recognition result analysis module; The data construction module is used to collect road surface images, obtain road images in different disease states, preprocess the road images, and construct a road image data set based on the preprocessing results; The multi-scale feature map analysis module is used to extract the road image data set based on the convolutional neural network and the feature pyramid network to generate feature maps at different levels; and to fuse the feature maps through the top-down path method and the lateral connection method to obtain the multi-scale feature map; The defect candidate frame analysis module is used to divide the multi-scale feature map through the region proposal network, generate candidate regions, and obtain initial defect candidate frames; detect and perform redundancy processing on the initial defect candidate frames based on the non-maximum suppression algorithm to obtain defect candidate frames; The recognition result analysis module is used to perform target detection on the disease candidate box using the R-FCN model, generate a position-sensitive score map, and determine the final category of the disease through a voting mechanism; and output the disease recognition result based on the R-FCN model, optimize the recognition result through post-processing, and obtain the final recognition result.
Citation Information
Patent Citations
Target detection method based on a dense connection characteristic pyramid network
CN109614985A
Automatic detection method for polymorphic target in continuous two-dimensional image
CN110009628A
Highway pavement three-dimensional disease identification method based on improved Faster R-CNN
CN114445708A
CT image lesion recognition method based on improved training network
CN115019065A
Faster R-CNN-based high-speed rail overhead line system part identification and positioning method
CN116363105A