Pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism

By combining the road surface disease detection method with wavelet transformation and high-frequency enhancement mechanism, the road surface features are extracted using the wavelet three-branch module and the multi-head attention mechanism, the problem of insufficient detection accuracy and reliability in the existing technology is solved, and more efficient disease identification and positioning is achieved.

CN120279402APending Publication Date: 2025-07-08CHONGQING JIAOTONG UNIV

Patent Information

Application Number
CN202311488406.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing deep learning-based pavement disease detection methods are not targeted, resulting in poor detection accuracy and reliability, and the inability to effectively identify and locate pavement diseases.

Method used

Combining wavelet transformation and high-frequency enhancement mechanism, the road surface detail texture and scale information are extracted through discrete wavelet transformation, and the attention enhancement mechanism is used to pay attention to disease characteristics. The wavelet three-branch module and the multi-head attention mechanism are used for feature extraction and fusion.

Benefits of technology

It improves the accuracy and reliability of road surface disease detection, can more accurately identify and locate diseases, reduce mis-detection and missed detection, and enhances the targetedness and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279402A_ABST
    Figure CN120279402A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep learning and pavement disease detection, in particular to a pavement disease detection method combining wavelet transform and a high-frequency enhancement mechanism, which comprises the following steps: acquiring a to-be-detected target pavement image and inputting the image into a backbone network for feature extraction to generate a pavement disease feature map, wherein the backbone network extracts and generates a pavement disease feature map through a wavelet three-branch module and a multi-head attention mechanism; inputting the pavement disease feature map into a Neck network for multi-scale feature fusion to generate a fusion feature map sequence; and inputting the fused feature map sequence into a target detection head, and outputting a corresponding pavement disease detection result. According to the method, road surface features such as road surface detail textures and scales and frequency and scale information of the road surface features are extracted through discrete wavelet transform, and meanwhile, the model pays more attention to disease information features contained in the road surface features through an attention enhancement mechanism, so that the road surface diseases can be identified and positioned more accurately; therefore, the accuracy and reliability of pavement disease detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and pavement disease detection, and particularly to a pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism. Background Art

[0002] In recent years, the highway construction in China has developed rapidly. How to maintain highways is the current research focus, and the pavement disease detection method is one of the cores of highway maintenance. In the early stage, pavement diseases were mainly detected by manual methods. This method not only has a long detection period, high work intensity, low accuracy, strong subjectivity, but also needs to close traffic routes when necessary, bringing potential personal safety problems to detection personnel and maintenance personnel. Subsequently, with the development of computer software and hardware technologies, detecting pavement diseases through machine vision technology has become the mainstream. This technology first obtains pavement information, that is, pavement image information, through a high-speed camera, then uses image processing and other intelligent algorithms to process this information, and finally determines whether the pavement image contains diseases and the locations of the diseases.

[0003] To achieve accurate pavement disease detection, an important link is to extract pavement disease characteristics as the main basis for detection, and the discriminative power of the characteristics will directly affect the detection results. Currently, the methods for extracting pavement disease characteristics are mainly divided into three types: the first is to use traditional image processing methods combined with manually designed features to extract pavement disease characteristics; the second introduces traditional machine learning algorithms on the basis of the first; the third is to extract image features based on deep learning methods. The first type belongs to an early method, which is easily affected by environmental factors such as light, shadow, background, and texture. Especially when inspecting defects such as potholes, looseness, and cracks, the inspection results have poor reliability and poor generalization; compared with the first method, the second method has been improved, but still highly depends on manually extracting features, and the machine learning training process is cumbersome; compared with manually extracting features, the method based on deep learning does not require sufficient prior knowledge and complex features designed manually, and has better robustness to changes in light, camera angle, and scene, and is the current mainstream method.

[0004] Among them, "An Intelligent Detection Method for Road Surface Diseases Suitable for Complex Backgrounds" (Publication No. CN 115661032 A) and "A Method for Detecting Road Surface Diseases in UAV Images Based on YOLO v4" (Publication No. CN 116310785 A) both use the currently popular YOLO model and make some improvements for disease detection. For example, they both choose to add an SE module for channel enhancement. However, this enhancement is a general enhancement for model features and not a feature enhancement specifically for road surface diseases such as cracks. Its pertinence is not strong. "Road Surface Quality Detection Method, Device and Related Products" (Publication No. CN115035305A) also uses a YOLO-based model for detection. Different from the above two patents, this patent specifically improves the YOLO model in combination with road surface disease characteristics, and improves the receptive field of the YOLO feature extraction network to the receptive field size suitable for road surface disease detection. Although this method has been improved for road surface disease characteristics, the model used still lacks pertinence in feature extraction; "A Method for Perceiving Highway Asphalt Pavement Diseases Based on a Fusion Convolutional Neural Network" (Publication No. CN 115393587A) fuses the Faster-RCNN, YOLOv5s, and SSD models and makes adaptive improvements. Although it does not make targeted improvements for road surface diseases, it combines the disease segmentation model with the disease detection model, introduces pixel-level marking of pictures, and improves the segmentation accuracy and detection accuracy at the same time. However, this method is relatively complex as a whole, and the model training and disease annotation costs are relatively high.

[0005] In summary, most of the current road surface disease detection methods based on deep learning methods use general methods to improve the feature extraction ability of the network model, do not conduct targeted module design and improvement for the characteristics of road surface diseases, and cannot accurately identify and locate diseases, thus resulting in poor accuracy and reliability of road surface disease detection. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned existing technologies, the technical problem to be solved by the present invention is: how to provide a road surface disease detection method combining wavelet transform and high-frequency enhancement mechanism, extract road surface features such as road surface details, textures, scales, as well as the frequency and scale information of road surface features through discrete wavelet transform, and at the same time make the model pay more attention to the disease information features contained in the road surface features through the attention enhancement mechanism, so as to be able to more accurately identify and locate road surface diseases, thereby improving the accuracy and reliability of road surface disease detection.

[0007] To solve the above technical problems, the present invention adopts the following technical solutions:

[0008] A road surface disease detection method combining wavelet transform and high-frequency enhancement mechanism, comprising:

[0009] S1: Obtain the target road surface image to be detected;

[0010] S2: Input the target road surface image into the backbone network for feature extraction to generate the corresponding road surface disease feature map;

[0011] The processing steps of the backbone network are as follows:

[0012] S201: Use the target road surface image as the input of the backbone network;

[0013] S202: First stage: First, perform patch embedding on the target road surface image to generate the corresponding high-dimensional feature map; then process the high-dimensional feature map through the wavelet three-branch module to generate the attention feature map with enhanced high-frequency; subsequently, perform residual fusion on the attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding first-stage feature map;

[0014] The processing process of the wavelet three-branch module: First, perform discrete wavelet transform on the input high-dimensional feature map, then perform attention enhancement, fusion, inverse discrete wavelet transform, and linear mapping on the result of the discrete wavelet transform, and finally generate the attention feature map with enhanced high-frequency;

[0015] S203: Second stage: First, perform patch embedding and downsampling on the first-stage feature map to obtain the corresponding high-dimensional feature map; then process the high-dimensional feature map through the wavelet three-branch module to generate the attention feature map with enhanced high-frequency; subsequently, perform residual fusion on the attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding second-stage feature map;

[0016] S204: Third stage: First, perform patch embedding and downsampling on the second-stage feature map to obtain the corresponding high-dimensional feature map; then calculate the attention feature of the high-dimensional feature map through the multi-head attention mechanism module to obtain the corresponding multi-head attention feature map; subsequently, perform residual fusion on the multi-head attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding third-stage feature map;

[0017] S205: Fourth stage: First, perform patch embedding and downsampling on the feature map of the third stage to obtain the corresponding high-dimensional feature map; then calculate the attention features of this high-dimensional feature map through the multi-head attention mechanism module to obtain the corresponding multi-head attention feature map; subsequently, perform residual fusion on this multi-head attention feature map and the corresponding high-dimensional feature map to generate the corresponding fused feature map; finally, after calculating this fused feature map through the PVT2FFN module, perform residual fusion with this fused feature map again to generate the corresponding feature map of the fourth stage;

[0018] S206: Take the feature maps of the first stage to the fourth stage as the road surface disease feature maps of the backbone network;

[0019] S3: Input the road surface disease feature maps into the Neck network for multi-scale feature fusion to generate a sequence of fused feature maps;

[0020] S4: Input the sequence of fused feature maps into the object detection head to output the corresponding road surface disease detection results.

[0021] Preferably, execute the first stage three times in a loop, and take the feature map of the first stage generated in the last time as the final feature map of the first stage;

[0022] In step S203, execute the second stage four times in a loop, and take the feature map of the second stage generated in the last time as the final feature map of the second stage;

[0023] In step S204, execute the third stage six times in a loop, and take the feature map of the third stage generated in the last time as the final feature map of the third stage;

[0024] In step S205, execute the fourth stage three times in a loop, and take the feature map of the fourth stage generated in the last time as the final feature map of the fourth stage.

[0025] Preferably, patch embedding means that first, the input image or feature map is segmented into several image patches or feature patches through two-dimensional convolution operations, and then the image patches or feature patches are mapped to a high-dimensional space through two-dimensional convolution and dimension transformation methods to obtain the corresponding high-dimensional feature map.

[0026] Preferably, the processing steps of the wavelet three-branch module are as follows:

[0027] S211: Define the input of the wavelet three-branch module as F;

[0028] S212: Convert the input F into a 2D feature map F2, and compress the number of channels of the 2D feature map F2 through convolution operations to obtain the feature map F 2′ ;

[0029] S213: For the feature map F2′ Perform discrete wavelet transform and decomposition to obtain a low-frequency sub-band F Low and three high-frequency sub-bands F LH 、F HL 、F HH ;

[0030] S214: Perform a convolution operation on the low-frequency sub-band F Low to obtain a low-frequency feature F LOW′ ;

[0031] S215: Concatenate the three high-frequency sub-bands F LH 、F HL 、F HH to obtain a high-frequency sub-band feature map F high ;

[0032] S216: Process the high-frequency sub-band feature map F high through the WFHA module to perform weighted processing on the spatial dimension and enhance the detailed features of the target, suppressing the unnecessary detailed features, and obtaining a high-frequency feature F high′ ;

[0033] S217: Concatenate the high-frequency feature F high′ with the low-frequency feature F LOW′ to obtain a concatenated feature F high″ ;

[0034] S218: Perform gated convolution processing on the concatenated feature F high″ ; then concatenate the result of the gated convolution processing with the low-frequency feature F LOW′ ; finally, perform inverse discrete wavelet transform on the concatenated result to obtain a feature map F idwt containing local feature information;

[0035] S219: Use the multi-head attention mechanism to perform feature extraction on the input F, the high-frequency feature F high′ and the feature map F idwt containing local feature information to generate an attention feature map with high-frequency enhancement.

[0036] Preferably, the processing steps of the WFHA module are as follows:

[0037] S2161: Perform a fuzzy operation on the high-frequency sub-band feature map F high through a fine-grained - maximum membership function to generate a membership matrix R ∈ R C*H*W with a value range between [0, 1];

[0038] S2162: Perform fuzzy enhancement on the membership matrix R through an adaptive fine-grained fuzzy enhancement method. The specific steps are as follows:

[0039] 1) Represents the membership matrix R through m1 and m2:

[0040]

[0041] In the formula: H represents the height of F high ; W represents the width of F high ; m1 represents the degree of division on H; m2 represents the degree of division on W; r c represents the image block in the membership matrix R; s represents the s-th r c ; r c The c in represents the number of channels; represents the i-th element of the c-th channel, and one r c has a total of elements;

[0042] 2) Determine the adaptive fine-grained fuzzy enhancement operator g;

[0043] The formula description is:

[0044] g = [g1, g2,..., g c T ;

[0045]

[0046] 3) Calculate the similarity matrix f between the membership matrix R and the adaptive fine-grained fuzzy enhancement operator g through the Manhattan distance;

[0047] The formula description is:

[0048]

[0049] f(R, g) = [f 1 , f 2 ,... f c , f c ∈ R H*W ;

[0050] In the formula: f(R, g) represents the similarity matrix between R and g; f c represents the similarity between each channel in R and each channel in g; represents the image block in R; represents the image block in g; and The i in represents the i-th channel; s represents the s-th or

[0051] 4) Adjust and update the similarity matrix f through the Gaussian function to obtain the updated similarity matrix f1;

[0052] ​The formula is described as follows:

[0053]

[0054] Wherein: Gaussian(f) represents performing a Gaussian function processing on the similarity matrix f(R,g); σ is an adjustable hyperparameter;

[0055] S2163: Perform defuzzification operation on the updated similarity matrix f1 through the fine-grained - maximum membership inverse function to obtain the corresponding high-frequency feature F high′ .

[0056] Preferably, the processing formula of the gated convolution is as follows:

[0057] Gate = F high″ ⊙W g ;

[0058] Feature = F high″ ⊙W f ;

[0059]

[0060] Wherein: Output represents the result of the gated convolution processing; W g and W f represent two uncorrelated convolutional kernels; ⊙ represents the convolution operation; represents element-wise multiplication of matrix elements; represents the activation function; σ represents the Sigmoid activation function, making the output of the gate value Gate between 0 and 1.

[0061] Preferably, the processing steps of the multi-head attention mechanism are as follows:

[0062] S2191: Map F ∈ R B*N*C to Q through linear mapping;

[0063] S2192: Use convolution and linear mapping to map F high′ to K w and V w ;

[0064] S2193: Calculate the multi-head attention learning Attention through the following formula w :

[0065]

[0066] Wherein: represents the Key value and Value value after high-frequency enhancement of the jth one, and the aggregated output of each head j is the global information combining the input after high-frequency enhancement;

[0067] S2194: Concatenate all the global information of each head and the reconstructed feature map F containing local information, and then perform a linear transformation to form the output of the wavelet three-branch module: idwt Concatenate them, and then perform a linear transformation to form the output of the wavelet three-branch module:

[0068] WaveletsBranchBlock(F) = MultiHead w (FW q , F high W k , F high W v , F idwt );

[0069]

[0070] In the formula: WaveletBranchBlock(F) represents the output of the wavelet three-branch module, that is, the attention feature map; W o represents the linear transformation matrix; W q , W k , W v represent the linear transformation matrices, and through them, Q j , K j w , V j w .

[0071] Compared with the prior art, the pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism in the present invention has the following beneficial effects:

[0072] In the present invention, a discrete wavelet transform is performed on a feature map through a wavelet three-branch module (i.e., Wavelets Branch Block), and then the results of the discrete wavelet transform are subjected to attention enhancement, fusion, inverse discrete wavelet transform, and linear mapping, finally generating an attention feature map with enhanced high-frequency components. Firstly, through the discrete wavelet transform, road surface details such as texture and scale, as well as the frequency and scale information of road surface features, can be effectively extracted, providing more clues for subsequent road surface disease detection; secondly, through the discrete wavelet transform, road surface features can be represented at different scales, and the features at these scales contain information at different levels of the road surface. By fusing the features at different scales, the model can consider road surface information more comprehensively, improving the reliability of road surface disease detection; at the same time, the attention enhancement mechanism enables the model to pay more attention to features related to road surface diseases. By performing attention enhancement on the high-frequency subbands of the discrete wavelet transform, the model can pay more attention to those features containing disease information, thereby improving the accuracy of road surface disease detection; finally, through the inverse discrete wavelet transform, the features after wavelet transform can be restored to the road surface feature map at the original scale, enabling the model to understand the road surface condition from both global and local perspectives and providing more accurate clues for subsequent disease detection; finally, through linear mapping, the attention feature map is transformed into more discriminative features and feature alignment is achieved, providing a more accurate feature representation for subsequent road surface disease detection. In summary, the present invention processes the feature map through the wavelet three-branch module, enabling the model to more accurately identify and locate road surface diseases, thereby improving the accuracy and reliability of road surface disease detection.

[0073] In the present invention, feature extraction is performed on a feature map through a multi-head attention mechanism module. The multi-head attention mechanism can recalibrate the information in the feature map by weighting the features at different positions and different channels in the feature map, enabling features related to diseases to receive greater attention, thereby improving the accuracy of road surface disease detection; at the same time, the multi-head attention mechanism can effectively fuse various information in the feature map, enabling the model to consider road surface features (such as shape, texture, color, etc.) more comprehensively, improving the reliability of the model and reducing the possibility of false detection and missed detection; in addition, the multi-head attention mechanism can perform independent weighting processing on different channels in the feature map, enabling the model to pay more attention to those channel features related to diseases, improving the pertinence of road surface disease detection. In summary, the present invention processes the feature map through the multi-head attention mechanism module, enabling the road surface disease detection model to more accurately identify and locate diseases, thereby improving the accuracy and reliability of road surface disease detection.

[0074] Before the wavelet three-branch module and the multi-head attention mechanism module process the feature map, patch embedding is used to segment the large feature image into sub-images and perform downsampling, enabling more effective data to be provided for the wavelet three-branch module and the multi-head attention mechanism module, and reducing the image size while increasing the number of channels of the image, thereby improving the accuracy and efficiency of road disease detection.

[0075] The present invention divides the feature extraction of the backbone network into four stages. First, as the depth of the network increases, the size of the feature map will continuously decrease, and the degree of abstraction of the features will gradually increase. By dividing the feature extraction process into multiple stages, each stage can have an appropriate depth and the size of the feature map can be maintained suitable for processing inputs of different sizes. Second, since the number of channels of the feature map gradually increases in each stage, the network extracts features at different abstraction levels in different stages, thereby improving the expression ability of the model. At the same time, dividing the feature extraction process into four stages makes the entire network have symmetry and balance, and the increase in the depth and the number of channels in each stage is relatively balanced, which helps to improve the stability and training effect of the model. Finally, multiple feature extraction stages make the network design more scalable. By increasing the number of stages or adjusting the number of each module in each stage, the depth and expression ability of the network can be easily changed to adapt to different tasks or data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:

[0077] Figure 1 is the network structure diagram of the backbone network;

[0078] Figure 2 is the network structure diagram of the wavelet three-branch module (Wavelets Branch Block);

[0079] Figure 3 is the workflow diagram of the WFHA module;

[0080] Figure 4 is the workflow diagram of the fine-grained - maximum membership function;

[0081] Figure 5 is the graphical schematic diagram of the membership matrix instance;

[0082] Figure 6 is the graphical schematic diagram of the adaptive fine-grained fuzzy enhancement method;

[0083] Figure 7 is the workflow diagram of the adaptive fine-grained fuzzy enhancement method;

[0084] Figure 8 It is a flowchart of the gating convolution process. Specific implementation manners

[0085] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the drawings herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0086] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings. In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use. These are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly defined and limited, the terms "set", "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0087] The following will be further described in detail through specific implementation manners:

[0088] Embodiment:

[0089] In this embodiment, a pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism is disclosed.

[0090] The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism includes:

[0091] S1: Obtain the target pavement image to be detected;

[0092] S2: Input the target pavement image into the backbone network for feature extraction to generate the corresponding pavement disease feature map;

[0093] In this embodiment, the backbone network is designed with reference to the PVT v2 model and the Wave-ViT model.

[0094] Among them, Vision Transformer, namely the ViT model, first introduced the Transformer model into the field of image processing and achieved the best results at that time in image classification tasks. Although ViT is suitable for image classification, directly applying it to pixel-level dense task prediction, such as object detection, still has certain difficulties because the output feature map of the ViT model is single-scale and low-resolution. Secondly, for common input image sizes, the computational and memory costs of ViT are relatively high. To solve the above problems, a large number of improved ViT models have emerged, and the PVT model, that is, the Pyramid ViT model, is one of them. This model can be used as a substitute for backbone networks such as convolutional neural networks in many object detection tasks.

[0095] PVT introduces a pyramid structure into the Transformer architecture so that it can generate multi-scale feature maps for dense prediction tasks. The PVT model follows the ResNet design rules and is divided into 4 stages. In the shallow layer, a smaller number of output channels are used to concentrate the main computational resources in the middle stage. As the network deepens, the image size will become smaller and smaller, while the output channels will become larger and larger.

[0096] In addition, the present invention combines gated convolution and the WFHA (Wave-Fuzzy High frequency Attention) module into the Wavelets Block of Wave-ViT to form a new Wavelets Branch Block (subsequently also referred to as the wavelet three-branch module), and finally forms the Wave-Fuzzy-ViT model, thereby strengthening the model's mining and utilization of high-frequency features. Its structure is as Figure 1 shown.

[0097] The processing steps of the backbone network are as follows:

[0098] S201: Use the target pavement image (assuming the input size is 224x224) as the input of the backbone network;

[0099] S202: The first stage: First, perform patch embedding on the target road surface image (that is, convert the target road surface image from a 3-channel image with a resolution of 224x224 into a feature map with a length of 3136 and 64 channels), generating a corresponding high-dimensional feature map; then process this high-dimensional feature map through a wavelet three-branch module (i.e., Wavelets Branch Block) to generate an attention feature map with enhanced high-frequency; subsequently, perform residual fusion on this attention feature map and the corresponding high-dimensional feature map to generate a corresponding fused feature map; finally, after calculating this fused feature map through the PVT2FFN module, perform residual fusion with this fused feature map again to generate a corresponding first-stage feature map;

[0100] Execute the first stage three times in a loop, and use the first-stage feature map generated in the last time as the final first-stage feature map of the first stage;

[0101] Among them, the processing process of the wavelet three-branch module: First, perform discrete wavelet transform on the input high-dimensional feature map, then perform attention enhancement, fusion, inverse discrete wavelet transform, and linear mapping on the result of the discrete wavelet transform, and finally generate an attention feature map with enhanced high-frequency;

[0102] In this embodiment, the PVT2FFN module used is a convolutional feed-forward network proposed by the existing PVTv2 model, which is used to replace the Feed-Forward module in the ordinary Transformer Block. The Feed-Forward module is a module containing two fully connected layers, which is used to perform non-linear transformation on the output data of the Attention Block to further enhance the feature expression ability of the input data. However, PVT2FFN adds depth convolution and GELU activation function between the two fully connected layers of Feed-Forward, thereby introducing zero-padding position encoding in the Transformer Block and strengthening the non-linear transformation ability, and can extract more semantic features.

[0103] S203: The second stage: First, perform patch embedding and downsampling on the first-stage feature map (that is, change the feature map dimension to a length of 784 and 128 channels) to obtain a corresponding high-dimensional feature map; then process this high-dimensional feature map through a wavelet three-branch module to generate an attention feature map with enhanced high-frequency; subsequently, perform residual fusion on this attention feature map and the corresponding high-dimensional feature map to generate a corresponding fused feature map; finally, after calculating this fused feature map through the PVT2FFN module, perform residual fusion with this fused feature map again to generate a corresponding second-stage feature map;

[0104] Execute the second stage four times in a loop, and use the feature map generated in the last time of the second stage as the final feature map of the second stage;

[0105] S204: The third stage: First, perform patch embedding and downsampling on the feature map of the second stage (changing to a length of 196 and 320 channels) to obtain the corresponding high-dimensional feature map; then, calculate the attention features of the high-dimensional feature map through the multi-head attention mechanism module (MSA) to obtain the corresponding multi-head attention feature map; subsequently, perform residual fusion on the multi-head attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it again to generate the corresponding feature map of the third stage;

[0106] In this embodiment, the multi-head attention mechanism module (Multi-Head Attention, MSA) used is an important module in natural language processing, mainly used for feature extraction and selection of the input sequence. Its basic idea is to split the input sequence into multiple heads, each head independently performs attention calculation, and then the attention results of each head are combined through a concatenation operation to form a richer feature representation.

[0107] In the multi-head attention mechanism, each head is an independent attention calculation module, which can freely focus on different parts of the input sequence and obtain independent weight parameters through learning. These heads can be regarded as some small encoder-decoders, which can observe and analyze the input sequence from different angles. Each head can focus on different parts of the input sequence. For example, one head focuses on local information and another head focuses on global information, so that the input sequence can be understood from multiple angles.

[0108] The multi-head attention mechanism is used to extract the global feature information of the input data. Similar to the convolution operation, it is also used for feature extraction. Convolution is used to extract local features, and the multi-head attention mechanism is used to extract global features. For any head head j , the cosine similarity is used internally to calculate the similarity between Q j and K j , and is scaled by ; then the softmax function is used to convert the similarity obtained in the previous step into weights between 0 and 1 and multiply by V j to obtain the final output, which contains global features.

[0109] The formula description of the multi-head attention mechanism is:

[0110]

[0111] MSA = MultiHead(FW q , FW k , FW v );

[0112]

[0113] Execute the third stage six times in a loop, and use the feature map of the third stage generated in the last time as the final feature map of the third stage;

[0114] S205: Fourth stage: First, perform patch embedding and downsampling on the feature map of the third stage (becoming 49 in length and 448 channels) to obtain the corresponding high-dimensional feature map; then calculate the attention features of the high-dimensional feature map through the multi-head attention mechanism module to obtain the corresponding multi-head attention feature map; subsequently, perform residual fusion on the multi-head attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on the fusion feature map again to generate the corresponding feature map of the fourth stage;

[0115] Execute the fourth stage three times in a loop, and use the feature map of the fourth stage generated in the last time as the final feature map of the fourth stage.

[0116] S206: Use the feature maps of the first stage to the fourth stage as the road surface disease feature maps of the backbone network;

[0117] S3: Input the road surface disease feature maps into the Neck network for multi-scale feature fusion to generate a sequence of fusion feature maps;

[0118] In this embodiment, the Neck network is an existing Neck network, and the present invention does not make improvements to this part.

[0119] S4: Input the sequence of fusion feature maps into the target detection head to output the corresponding road surface disease detection results.

[0120] In this embodiment, the target detection head is an existing target detection head, and the present invention does not make improvements to this part. The target detection head can use the existing RetinaNet, which performs classification prediction and bounding box prediction based on a set of obtained feature maps, and finally marks the disease area and the disease category on the disease image. The road surface disease detection result in the present invention specifically refers to marking the disease area on the given road surface disease image and indicating the disease type.

[0121] In the present invention, the feature map is subjected to discrete wavelet transform through a wavelet three-branch module (i.e., Wavelets Branch Block), and then the results of the discrete wavelet transform are subjected to attention enhancement, fusion, inverse discrete wavelet transform, and linear mapping, finally generating an attention feature map with enhanced high-frequency components. Firstly, through discrete wavelet transform, road surface details such as texture and scale, as well as the frequency and scale information of road surface features, can be effectively extracted, providing more clues for subsequent road surface disease detection; secondly, through discrete wavelet transform, road surface features can be represented at different scales, and the features at these scales contain information at different levels of the road surface. By fusing the features at different scales, the model can consider the road surface information more comprehensively, improving the reliability of road surface disease detection; at the same time, the attention enhancement mechanism enables the model to pay more attention to the features related to road surface diseases. By enhancing the attention of the high-frequency sub-bands of the discrete wavelet transform, the model can focus more on the features containing disease information, thereby improving the accuracy of road surface disease detection; finally, through inverse discrete wavelet transform, the features after wavelet transform can be restored to the road surface feature map at the original scale, enabling the model to understand the road surface condition from both global and local perspectives and providing more accurate clues for subsequent disease detection; finally, through linear mapping, the attention feature map is transformed into more discriminative features and feature alignment is achieved, providing a more accurate feature representation for subsequent road surface disease detection. In summary, through the processing of the feature map by the wavelet three-branch module in the present invention, the model can more accurately identify and locate road surface diseases, thereby improving the accuracy and reliability of road surface disease detection.

[0122] In the present invention, feature extraction is performed on the feature map through a multi-head attention mechanism module. Among them, the multi-head attention mechanism can recalibrate the information in the feature map by weighting the features at different positions and different channels in the feature map, enabling the features related to diseases to receive greater attention, thereby improving the accuracy of road surface disease detection; at the same time, the multi-head attention mechanism can effectively fuse various information in the feature map, enabling the model to consider road surface features (such as shape, texture, color, etc.) more comprehensively, improving the reliability of the model and reducing the possibility of false detection and missed detection; in addition, the multi-head attention mechanism can perform independent weighting processing on different channels in the feature map, enabling the model to pay more attention to the channel features related to diseases and improving the pertinence of road surface disease detection. In summary, through the processing of the feature map by the multi-head attention mechanism module in the present invention, the road surface disease detection model can more accurately identify and locate diseases, thereby improving the accuracy and reliability of road surface disease detection.

[0123] Before the wavelet three-branch module and the multi-head attention mechanism module process the feature map, patch embedding is used to segment the large feature image into sub-images and perform downsampling, enabling more effective data to be provided for the wavelet three-branch module and the multi-head attention mechanism module, reducing the image size while increasing the number of channels of the image, thereby improving the accuracy and efficiency of road disease detection.

[0124] The feature extraction of the backbone network of the present invention is divided into four stages. First, as the depth of the network increases, the size of the feature map will continuously decrease, and the degree of abstraction of the features will gradually increase. By dividing the feature extraction process into multiple stages, each stage can have an appropriate depth and can keep the size of the feature map suitable for processing inputs of different sizes. Second, since the number of channels of the feature map gradually increases in each stage, the network extracts features at different levels of abstraction in different stages, thereby improving the expression ability of the model. At the same time, dividing the feature extraction process into four stages makes the entire network have symmetry and balance, and the increase in the depth and the number of channels in each stage is relatively balanced, which helps to improve the stability and training effect of the model. Finally, multiple feature extraction stages make the network design more scalable. By increasing the number of stages or adjusting the number of each module in each stage, the depth and expression ability of the network can be conveniently changed to adapt to different tasks or data sets.

[0125] I. Patch Embedding

[0126] Patch Embedding means that first, the input image or feature map is segmented into several image patches or feature patches through a two-dimensional convolution operation, and then the image patches or feature patches are mapped to a high-dimensional space through a two-dimensional convolution and dimension transformation method to obtain the corresponding high-dimensional feature map.

[0127] The backbone network of the present invention is a network model based on the Transformer structure. Patch embedding is a common method of the Transformer Block, and its purpose is to segment a large image into several sub-images and perform downsampling (image patches). Then, MSA (multi-head attention) calculation is performed on the basis of the sub-images (the Wavelets Branch Block proposed in the present invention is also an improvement based on the MSA method). Without patch embedding, subsequent calculations cannot be performed; and downsampling follows the design principle of existing classical models, that is, reducing the image size while increasing the number of channels of the image.

[0128] Assume that the image data input to the network is F I ∈R B*C*H*W, where B represents the number of input images in each batch, C represents the channel dimension of the input images, and H and W represent the height and width of the images respectively. After Patch Embedding, the output is F I′ ∈R B*N*C′ , where N = H' * W', the number of channels is increased through convolution operations, while the width and height are correspondingly reduced, i.e., C' > C, H' < H, W' < W.

[0129] II. Wavelets Branch Block

[0130] In the specific implementation process, F is obtained through patch embedding I′ and then input into the improved Wavelets BranchBlock for discrete wavelet transform to obtain high-frequency and low-frequency components. At this time, three branches can be obtained, namely the spatial domain branch containing the spatial domain feature F I′ , the high-frequency branch containing high-frequency features, and the low-frequency branch containing low-frequency features. The three branches are processed separately, and finally, after fusion, inverse discrete wavelet transform, and linear mapping, it is restored to the same dimension as F I′ for other processing. In the Wavelets Branch Block of the same stage, the input and output dimensions are always the same, and Patch Embedding is performed again for downsampling when entering the next stage.

[0131] Since different semantic information is contained in the spatial domain, high-frequency, and low-frequency components respectively, it is reasonable to process them separately to obtain better features. Combining Figure 2 as shown, the processing steps of the wavelets three-branch module are as follows:

[0132] S211: Define the input of the wavelets three-branch module as F ∈ R B*N*C (For the convenience of understanding, assume that there is only one image in each batch, i.e., B = 1);

[0133] S212: Convert the input F into a 2D feature map F2 ∈ R C*H*W , and compress the number of channels of the 2D feature map F2 through convolution operations (because discrete wavelet transform will generate 4 subbands while downsampling the image, which means expanding the dimension of the input features by 4 times. First, use convolution operations to compress its number of channels), to obtain the feature map

[0134] S213: Perform discrete wavelet transform (DWT) and decomposition on the feature map F 2′ to obtain a low-frequency subband i.e., and three high-frequency subbands

[0135] Since the discrete wavelet transform itself is a lossless downsampling, the receptive fields of the feature maps of the high-frequency branch and the low-frequency branch are enlarged.

[0136] S214: Perform a convolution operation on the low-frequency subband F Low to obtain low-frequency features

[0137] S215: Concatenate the three high-frequency subbands F LH 、F HL 、F HH to obtain the high-frequency subband feature map

[0138] S216: Process the high-frequency subband feature map F high through the WFHA (Wave-Fuzzy High frequency Attention) module to perform weighted processing on the spatial dimension, strengthen the detailed features of the target, suppress the unnecessary detailed features, and obtain high-frequency features

[0139] S217: Concatenate the high-frequency features F high′ with the low-frequency features F LOW′ to obtain the concatenated features

[0140] S218: Perform gated convolution processing on the concatenated features F high″ ; then concatenate the result of the gated convolution processing with the low-frequency features F LOW′ ; finally, perform the inverse discrete wavelet transform (IDWT) on the concatenated result to obtain the feature map F idwt ∈R C*H*W containing local feature information, and adjust the feature map to F idwt ∈R N*C , N = H * W;

[0141] Since the wavelet transform has no information loss, there is no loss of any information in the downsampling through the wavelet transform and then the upsampling through the inverse wavelet transform, and the convolution operation in the wavelet domain can extract better local features.

[0142] S219: Use the improved multi-head attention mechanism to perform feature extraction on the input F, the high-frequency features F high′ and the feature map F idwt containing local feature information to generate an attention feature map with high-frequency enhancement.

[0143] In the present invention, a discrete wavelet transform is performed on an input high-dimensional feature map through a wavelet three-branch module, and then the result of the discrete wavelet transform is subjected to attention enhancement, fusion, inverse discrete wavelet transform, and linear mapping, finally generating an attention feature map with enhanced high-frequency components. First, through the discrete wavelet transform, road surface details, textures, scales, and other road surface features, as well as the frequency and scale information of the road surface features, can be effectively extracted, providing more clues for subsequent road surface disease detection. Secondly, through the discrete wavelet transform, road surface features can be represented at different scales, and the features at these scales contain information at different levels of the road surface. By fusing the features at different scales, the model can consider the road surface information more comprehensively, improving the reliability of road surface disease detection. At the same time, the attention enhancement mechanism enables the model to pay more attention to the features related to road surface diseases. By performing attention enhancement on the high-frequency subbands of the discrete wavelet transform, the model can pay more attention to the features containing disease information, thereby improving the accuracy of road surface disease detection. Finally, through the inverse discrete wavelet transform, the features after wavelet transform can be restored to the road surface feature map at the original scale, enabling the model to understand the road surface condition from both global and local perspectives and providing more accurate clues for subsequent disease detection. Finally, through linear mapping, the attention feature map is transformed into more discriminative features and feature alignment is achieved, providing a more accurate feature representation for subsequent road surface disease detection. In summary, the present invention processes the feature map through a wavelet three-branch module, enabling the model to more accurately identify and locate road surface diseases, thereby improving the accuracy and reliability of road surface disease detection.

[0144] The wavelet three-branch module (Wavelets Branch Block) of the present invention mainly includes a wavelet transform, a WFHA module, and a gated convolution module, and the input is divided into three branches: the spatial domain, high frequency, and low frequency for processing. The WFHA module is a spatial attention mechanism that can enhance the disease features in disease images; the gated convolution is similar in function to the WFHA module, the difference being that the gated convolution combines both low-frequency and high-frequency features at the same time, allowing the model to pay more attention to the disease area to achieve the effect of enhancing the disease feature extraction ability.

[0145] III. WFHA Module

[0146] In the specific implementation process, wavelet transform has good multi-scale analysis ability. Through discrete wavelet transform (DWT), an image can be decomposed into a low-frequency sub-band LL and three high-frequency sub-bands HL, LH, and HH. The low-frequency sub-band contains most of the information of the original image, while the three high-frequency components retain the texture details of the object. Compared with normal road surfaces, most road surface diseases have obvious characteristics visually, and in the frequency domain, there will be an obvious jump from low frequency to high frequency. Aiming at the problem of road surface diseases, in order to enhance the feature intensity of the disease area and suppress the feature intensity of the non-disease area, inspired by image enhancement based on fuzzy theory, the present invention proposes a high-frequency attention mechanism (WFHA) module that combines wavelet high-frequency sub-bands and fuzzy enhancement. As Figure 3 shown, the WFHA module includes two sub-methods, namely the fine-grained - maximum membership function and the adaptive fine-grained fuzzy enhancement method. First, the present invention uses the fine-grained - maximum membership function to map the high-frequency sub-band to the fuzzy domain U between 0 and 1, performs fuzzy enhancement on U, and then transforms it back to the wavelet domain through the membership inverse transformation function.

[0147] Combined with Figure 3 shown, the processing steps of the WFHA module are as follows:

[0148] S2161: Perform a fuzzy operation on the high-frequency sub-band feature map F high through the fine-grained - maximum membership function to generate a membership matrix R ∈ R C*H*W with a value range between [0, 1];

[0149] In this embodiment, during the calculation process of the fine-grained - maximum membership function, F high is divided into multiple image blocks (some image blocks contain target features, and some do not). After each image block undergoes corresponding calculations, the elements within the image block become values within the range of 0 - 1. From the calculation process, R can be composed of image blocks.

[0150] Specifically, for the fine-grained - maximum membership function:

[0151] Define A as a fuzzy subset of the high-frequency sub-band on U, which represents the set of the high-frequency sub-band belonging to "target features" on U, then its complement A c is the set of non-"target features" on U. Theoretically, the high-frequency sub-band only contains all the high-frequency information of the original image, and low-frequency signals such as the background have been filtered out. Diseases such as cracks belong to high-frequency signals, and the high-frequency sub-band must contain target features. Therefore, the information of the high-frequency sub-band is directly divided into two categories: "target features" and "non-target features". During the model training process, based on the backpropagation mechanism, the membership function gradually increases the membership degree of "target features", and the enhancement method gradually enhances "target features" and suppresses "non-target features".

[0152] Combined Figure 4 As shown, the fine-grained - maximum membership function divides the high-frequency subband feature map F high band into different regions u i , and then extracts the maximum value Max of different regions i Based on this, the corresponding membership matrix R is obtained;

[0153] The formula description is as follows:

[0154]

[0155]

[0156]

[0157] In the formula: σ represents the ReLU activation function; Max represents the maximum value function; Max i ∈R C*1*1 ; represents the maximum value of the i-th region and the C-th channel; represents the element in the j-th channel, m-th row, and n-th column of u i , where i ∈ [1, n1n2].

[0158] The membership function of A is:

[0159]

[0160] In the formula: μ A (x) is the fine-grained - maximum membership function

[0161]

[0162] In the formula: μ A -1 (x) is the inverse fine-grained - maximum membership function;

[0163] A c The membership function of is:

[0164]

[0165]

[0166] In the formula: where i ≠ j. When n1 = 1 and n2 = 1, the maximum membership function degenerates into an ordinary maximum membership function;

[0167] The target feature fuzzy set A = {(x1, μ A (x1)), (x2, μ A (x2)), … (xn , μ A (x n ))}, n = HW; A c = 1 - A, then the corresponding membership matrix R ∈ R C*H*W ,

[0168] S2162: The membership matrix R is fuzzy enhanced by the adaptive fine-grained fuzzy enhancement method, and the specific steps are as follows:

[0169] 1) The membership matrix R is re-expressed by m1 and m2:

[0170]

[0171] It means that R is divided into m1 * m2 r c , from the perspective of the set, that is, m1m2 r c combined together is R.

[0172] In the formula: H represents the height of F high ; W represents the width of F high ; m1 represents the division degree on H; m2 represents the division degree on W; r c represents the image block in the membership matrix R; s represents the s-th r c ; r c in which c represents the number of channels; represents the i-th element of the c-th channel, and 1 r c has a total of elements;

[0173] Figure 5 What represents is a membership matrix R, which has only one channel, H = W = 4, m1 = m2 = 2, and R is divided into 4 regions, that is, there are 4 r c=1 , and different colors are different r. 1 - 4, 5 - 8, 9 - 12, 13 - 16 in the figure are the specific x, that is, different r c inside the elements, so there are a total of 4 elements. Figure 5 The membership matrix R in is expressed by the formula:

[0174]

[0175] r 1 = [1, 2, 3, 4]; r 2 = [5, 6, 7, 8]; r 3 = [9, 10, 11, 12]; r 4= [13, 14, 15, 16].

[0176] 2) Determine the adaptive fine-grained fuzzy enhancement operator Specifically, given the membership matrix R ∈ R C *H*W as the input, the corresponding adaptive fuzzy enhancement operator is g is a randomly generated learnable parameter matrix, g ∈ [0, 1].

[0177] The formula is described as:

[0178] g = [g1, g2, …, g c T ;

[0179]

[0180] 3) Calculate the similarity matrix f between the membership matrix R and the adaptive fine-grained fuzzy enhancement operator g through the Manhattan distance algorithm; the Manhattan distance algorithm is an algorithm for calculating the distance between two points, also known as the city block distance, which calculates the absolute distance between two points in an n-dimensional coordinate system.

[0181] The formula is described as:

[0182]

[0183] f(R, g) = [f 1 , f 2 , … f c , f c ∈ R H*W ;

[0184] In the formula: f(R, g) represents the similarity matrix between R and g; f c represents the similarity between each channel in R and each channel in g; represents the image block in R; represents the image block in g; and in i represents the i-th channel; s represents the s-th or

[0185] In this embodiment, the adaptive fine-grained fuzzy enhancement method is described in a graphical manner:

[0186] As Figure 6 (a) shows, the membership matrix R ∈ R C*H*W ​, assume C = 1, H = W = 4. At this time, R is a 4*4 matrix, that is, R contains 16 elements. Assume m1 = m2 = 2, that is, R is divided into 4 regions (image patches), and each region is represented by a different color. At the same time, there are 4 regions represented by i = 4, u1, u2, u3, u4. Correspondingly, g = 4, that is, g has 4 parameters, represented by A, B, C, and D, and the same color represents the same region.

[0187] The calculation process of the Manhattan distance algorithm is as Figure 6 (b) and Figure 7 shown. R can be regarded as a set or combination of image patches. At the same time, R can also be viewed from the perspective of its scale C*H*W. f c is the result after calculating the similarity between the image patches in R and the parameters in g channel by channel. f is the result after subtracting g from R. The values in f are the similarities, and f is f(R, g).

[0188] 4) Adjust and update the similarity matrix f through the Gaussian function to obtain the updated similarity matrix f1;

[0189] In this embodiment, the similarity matrix f is a matrix with a value range between 0 and 1. The closer the value in it is to 0, the higher the similarity represents, and the closer it is to 1, the lower the similarity represents. Use the Gaussian function with the axis of symmetry at x = 0 to adjust f. For the values in f close to 0, their values are close to 1 after being calculated by the Gaussian function; for the values in f close to 1, their values are close to 0 after being calculated by the Gaussian function. Calculating through the Gaussian function can realize the attention adjustment of f, strengthen the attention weight of the target features, and suppress the attention weight of the non-target features, where the suppression degree is adjusted by the hyperparameter σ.

[0190] The formula description is:

[0191]

[0192] In the formula: Gaussian(f) represents performing Gaussian function processing on the similarity matrix f(R, g); σ is an adjustable hyperparameter;

[0193] S2163: Perform defuzzification operation on the updated similarity matrix f1 through the fine-grained - maximum membership inverse function to obtain the corresponding high-frequency feature F high′ .

[0194] The present invention can amplify or maintain the similarity of small-distance element pairs and reduce the similarity of large-distance element pairs through the WFHA module to achieve attention adjustment. For each image patch containing target features The random number in g will be updated to be near the target eigenvalue, making its Manhattan distance small enough; for the image patches that do not contain the target feature, the random number will be as far away from the internal values of the image as possible, making its Manhattan distance large enough. After obtaining the updated similarity matrix, perform the defuzzification operation to return to the value range space of the high-frequency subband. It should be noted that the enhancement algorithm proposed in the present invention is linear. Therefore, the Gaussian function in the WFHA attention method is not only used as an attention weight adjustment function but also a non-linear activation function.

[0195] IV. Gated Convolution

[0196] In the specific implementation process, although the high-frequency subband contains a lot of high-frequency semantic information and features, it lacks the semantic information of the low-frequency subband. Therefore, it is necessary to fuse some low-frequency information in the high-frequency subband to accelerate the model convergence and improve the accuracy. Concatenate the low-frequency subband and the high-frequency subband together, use the high-frequency subband as the mask information of the low-frequency subband, and then use gated convolution for the fusion of semantic features.

[0197] Gated convolution can dynamically select low-frequency features for fusion according to high-frequency features, and the input and output data dimensions of gated convolution remain the same. As shown in Figure 8 , the processing formula of gated convolution is as follows:

[0198] Gate = F high″ ⊙W g ;

[0199] Feature = F high″ ⊙W f ;

[0200]

[0201] In the formula: Output represents the result of the gated convolution processing; W g and W f represent two uncorrelated convolution kernels; ⊙ represents the convolution operation; represents the element-wise multiplication of matrices; represents the activation function, such as ReLU, etc.; σ represents the Sigmoid activation function, making the output of the gate value Gate between 0 and 1.

[0202] V. Improved Multi-Head Attention Mechanism

[0203] In the specific implementation process, the processing steps of the improved multi-head attention mechanism are as follows:

[0204] S2191: Map F ∈ R B*N*C to

[0205] S2192: Use convolution and linear mapping for Fhigh′ Mapped to and

[0206] S2193: Calculate the multi - head attention learning Attention through the following formula w (where Attention w represents the attention calculation operation, which is used to model or extract the global features of the input data. Different from Attention, Attention w is an Attention operation based on wavelet transform):

[0207]

[0208] In the formula: K j w , V j w represent the Key value and Value value of the j - th high - frequency enhanced, and the aggregated output of each head j is the global information combining the high - frequency enhanced input;

[0209] S2194: Concatenate all the global information of each head and the reconstructed feature map F containing local information idwt , and then perform a linear transformation to form the output of the wavelet three - branch module:

[0210] WaveletsBranchBlock(F) = MultiHead w (FW q , F high W k , F high W v , F idwt );

[0211]

[0212] In the formula: WaveletBranchBlock(F) represents the output of the wavelet three - branch module, that is, the attention feature map; W o represents the linear transformation matrix; W q , W k , W v represent the linear transformation matrices, through which Q j , K j w , V j w are obtained respectively.

[0213] VI. Neck Network

[0214] In the specific implementation process, the Neck network selects the FPN network (Feature Pyramid Network).

[0215] The FPN network is a feature fusion method commonly used in object detection tasks. Using it as the Neck network of the road surface disease detection model can effectively fuse features at different levels and perform dimensionality reduction and dimensionality increase operations on the features.

[0216] Specifically, in the bottom-up path of the FPN network, through continuous convolution and pooling operations, the resolution of the input feature map is gradually reduced, and more and more abstract features are extracted. In the top-down path, the FPN network gradually increases the resolution of the feature map through upsampling operations, so that each feature map can contain more detailed information. Through the combination of these two paths, the FPN network can capture the information of low-level and high-level features simultaneously in a unified framework and effectively fuse them.

[0217] In the road surface disease detection task, the Neck network of the FPN network can connect the front end and the back end of the detector and further process and fuse the features extracted by the detector. Specifically, the Neck network can perform dimensionality reduction, dimensionality increase, and fusion operations on the features through a series of convolutional layers, pooling layers, upsampling layers, etc., so that the detection model can better utilize the feature information at different levels and improve the detection accuracy.

[0218] In summary, the present invention uses the FPN network as the Neck network of the road surface disease detection model, which can effectively improve the performance of object detection and has important application value.

[0219] VII. Object Detection Head

[0220] In the specific implementation process, the RetinaNet is selected as the object detection head.

[0221] RetinaNet is a new neural network architecture for object detection tasks. It effectively improves the accuracy of object detection by constructing multi-scale anchor boxes on the Feature Pyramid Network (FPN). In the road surface disease detection task, RetinaNet can be used as the object detection head.

[0222] Specifically, RetinaNet first extracts multi-scale features through a Feature Pyramid Network, and then constructs anchor boxes of different scales on these features. Each anchor box is matched with the real targets in the image to calculate the loss function and update the network parameters. In this way, RetinaNet can perform object detection on multiple scales and multiple feature pyramids, effectively improving the detection accuracy. In the road surface disease detection task, RetinaNet can be used as an object detection head. Specifically, a set of generated feature maps are used as the input of the network, and then feature extraction and object detection are performed through RetinaNet, so that information of different scales and different features can be processed in a unified framework, improving the accuracy and efficiency of object detection.

[0223] In summary, the present invention uses RetinaNet as the object detection head of the road surface disease detection model, which can effectively improve the performance of object detection and has important application value.

[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements to the technical solutions of the present invention, without departing from the purpose and scope of the technical solutions, should be covered by the scope of the claims of the present invention.

Claims

1. A road surface disease detection method combining wavelet transform and high-frequency enhancement mechanism, characterized in that, Including: S1: Obtain the target road surface image to be detected; S2: Input the target road surface image into the backbone network for feature extraction to generate the corresponding road surface disease feature map; The processing steps of the backbone network are as follows: S201: Use the target road surface image as the input of the backbone network; S202: First stage: First, perform patch embedding on the target road surface image to generate the corresponding high-dimensional feature map; then process the high-dimensional feature map through the wavelet three-branch module to generate the high-frequency enhanced attention feature map; subsequently, perform residual fusion on the attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding first-stage feature map; Among them, the processing process of the wavelet three-branch module: First, perform discrete wavelet transform on the input high-dimensional feature map, then perform attention enhancement, fusion, inverse discrete wavelet transform, and linear mapping on the result of the discrete wavelet transform, and finally generate the high-frequency enhanced attention feature map; S203: Second stage: First, perform patch embedding and downsampling on the first-stage feature map to obtain the corresponding high-dimensional feature map; then process the high-dimensional feature map through the wavelet three-branch module to generate the high-frequency enhanced attention feature map; subsequently, perform residual fusion on the attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding second-stage feature map; S204: Third stage: First, perform patch embedding and downsampling on the second-stage feature map to obtain the corresponding high-dimensional feature map; then perform attention feature calculation on the high-dimensional feature map through the multi-head attention mechanism module to obtain the corresponding multi-head attention feature map; Subsequently, perform residual fusion on the multi-head attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding third-stage feature map; S205: Fourth stage: First, perform patch embedding and downsampling on the third-stage feature map to obtain the corresponding high-dimensional feature map; then perform attention feature calculation on the high-dimensional feature map through the multi-head attention mechanism module to obtain the corresponding multi-head attention feature map; Subsequently, perform residual fusion on the multi-head attention feature map and the corresponding high-dimensional feature map to generate the corresponding fusion feature map; finally, after calculating the fusion feature map through the PVT2FFN module, perform residual fusion on it and the fusion feature map to generate the corresponding fourth-stage feature map; S206: Use the feature maps of the first stage to the fourth stage as the road surface disease feature map of the backbone network; S3: Input the road surface disease feature map into the Neck network for multi-scale feature fusion to generate a sequence of fusion feature maps; S4: Input the sequence of fusion feature maps into the target detection head to output the corresponding road surface disease detection result.

2. The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism as described in claim 1, wherein: In step S202, the first stage is executed three times in a loop, and the first stage feature map generated in the last time is used as the final first stage feature map of the first stage; In step S203, the second stage is executed four times in a loop, and the second stage feature map generated in the last time is used as the final second stage feature map of the second stage; In step S204, the third stage is executed six times in a loop, and the third stage feature map generated in the last time is used as the final third stage feature map of the third stage; In step S205, the fourth stage is executed three times in a loop, and the fourth stage feature map generated in the last time is used as the final fourth stage feature map of the fourth stage.

3. The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism according to claim 1, wherein: In step S2, patch embedding means that first, the input image or feature map is segmented into several image patches or feature patches through a two-dimensional convolution operation, and then the image patches or feature patches are mapped to a high-dimensional space through a two-dimensional convolution and dimension transformation method to obtain the corresponding high-dimensional feature map.

4. The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism according to claim 1, characterized in that: In step S2, the processing steps of the wavelet three-branch module are as follows: S211: Define the input of the wavelet three-branch module as F; S212: Convert the input F into a 2D feature map F2, and compress the number of channels of the 2D feature map F2 through a convolution operation to obtain a feature map F 2′ ; S213: Perform discrete wavelet transform and decomposition on the feature map F 2′ to obtain a low-frequency sub-band F Low and three high-frequency sub-bands F LH , F HL , F HH ; S214: Convolve the low-frequency sub-band F Low to obtain the low-frequency feature F LOw′ ; S215: Concatenate the three high-frequency sub-bands F LH 、F HL 、F HH to obtain the high-frequency sub-band feature map F high ; S216: Process the high-frequency subband feature map F through the WFHA module high to perform weighting on the spatial dimension, enhance the detailed features of the target, suppress the unwanted detailed features, and obtain the high-frequency feature F high′ ; S217: Concatenate the high-frequency feature F high′ with the low-frequency feature F LOW′ to obtain the concatenated feature F high″ ; S218: Perform gated convolution processing on the splicing feature F high″ ; then splice the result of the gated convolution processing with the low-frequency feature F LOW′ ; finally, perform inverse discrete wavelet transform on the spliced result to obtain the feature map F idwt ; S219: Use the multi-head attention mechanism to perform feature extraction on the input F, the high-frequency feature F high′ and the feature map F idwt containing local feature information, and generate an attention feature map with enhanced high-frequency components.

5. The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism according to claim 4, characterized in that: In step S216, the processing steps of the WFHA module are as follows: S2161: Perform a fuzzy operation on the high-frequency subband feature map F through a fine-grained - maximum membership function high to generate a membership matrix R ∈ R with a value range between [0, 1] C*H*W ; S2162: Perform fuzzy enhancement on the membership matrix R through an adaptive fine-grained fuzzy enhancement method. The specific steps are as follows: 1) Represent the membership matrix R through m1 and m2: Where: H represents the height of F high ; W represents the width of F high ; m1 represents the degree of division on H; m2 represents the degree of division on W; r c represents the image block in the membership matrix R; s represents the s-th r c ; r c where c represents the number of channels; represents the i-th element of the c-th channel, and one r c has a total of elements; 2) Determine the adaptive fine-grained fuzzy enhancement operator g; The formula description is: g = [g1, g2, …, g c T ;​ 3) Calculate the similarity matrix f between the membership matrix R and the adaptive fine-grained fuzzy enhancement operator g through the Manhattan distance; The formula description is: f(R, g) = [f 1 , f 2 , … f c , f c ∈ R H*W ; where: f(R, g) represents the similarity matrix of R and g; f c represents the similarity between each channel in R and each channel in g; represents the image patch in R; represents the image patch in g; and the i in represents the i-th channel; s represents the s-th or 4) Adjust and update the similarity matrix f through a Gaussian function to obtain the updated similarity matrix f1; The formula description is: In the formula: Gaussian(f) represents performing Gaussian function processing on the similarity matrix f(R, g); σ is an adjustable hyperparameter; S2163: Perform defuzzification operation on the updated similarity matrix f1 through the fine-grained - maximum membership degree inverse function to obtain the corresponding high-frequency feature F hig h ′ 。 6. The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism according to claim 4, characterized in that: In step S218, the processing formula of the gated convolution is as follows: Gate=F high″ ⊙W g ; Feature=F high″ ⊙W f ; Where: Output represents the result of the gated convolution process; W g and W f represent two uncorrelated convolutional kernels; ⊙ represents the convolution operation; represents element-wise multiplication of matrix elements; represents the activation function; σ represents the Sigmoid activation function, making the output of the gate value Gate between 0 and 1.

7. The pavement disease detection method combining wavelet transform and high-frequency enhancement mechanism according to claim 4, characterized in that: In step S219, the processing steps of the multi-head attention mechanism are as follows: S2191: Map F ∈ R to Q through linear mapping B*N*C ; S2192: Map F to K high′ and V w using convolution and linear mapping; w ​ S2193: Calculate the multi-head attention learning Attention through the following formula w : In the formula: represents the Key value and Value value after high-frequency enhancement for the j-th one, and the aggregated output of each head j is the global information that combines the input after high-frequency enhancement; S2194: Concatenate all the global information of each head and the reconstructed feature map F containing local information, and then perform a linear transformation to form the output of the wavelet three-branch module: idwt Then concatenate them and perform a linear transformation to form the output of the wavelet three-branch module: WaveletsBranchBlock(F) = MultiHead w (FW q ,F high W k ,F high W v ,F idwt ); Where: WaveletBranchBlock(F) represents the output of the wavelet three-branch module, i.e., the attention feature map; W o represents the linear transformation matrix; W q , W k , W v represent the linear transformation matrices, through which Q j ,

Citation Information

Patent Citations

  • Road surface quality detection method and device and related product

    CN115035305A

  • Fusion convolutional neural network-based highway asphalt pavement disease sensing method

    CN115393587A

  • Road surface disease intelligent detection method suitable for complex background

    CN115661032A

  • Unmanned aerial vehicle image pavement disease detection method based on YOLO v4

    CN116310785A

Cited By

  • Farm disease identification method and device based on multi-modal data fusion

    CN121685464A