Road image oriented motion blur removal and super-resolution reconstruction method and system
Patent Information
- Application Number
- CN202611160161.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-08-03
AI Technical Summary
道路图像中的模糊退化、裂缝边缘、高频纹理和空间细节之间具有较强关联性,仅通过简单串联去模糊模型和超分辨率模型,难以充分发挥二者之间的互补作用
在本发明中,利用运动模糊去除分支,通过编码器、瓶颈层和解码器逐级提取与恢复道路图像中的结构信息,超分辨率重建分支提取道路图像中的高频纹理信息和空间细节信息,为后续高分辨率图像重建提供细节特征;通过跨尺度双向特征交互模块对运动模糊去除分支和超分辨率重建分支中的特征进行双向交互,能够使超分辨率重建分支中的细节特征能够辅助运动模糊去除分支恢复边缘结构,运动模糊去除分支中的清晰结构特征能够辅助超分辨率重建分支重建高分辨率纹理,进而将共享浅层特征、交互后去模糊特征和交互后的超分辨率特征进行融合和重建,得到重建后道路图像。本发明方法实现道路图像清晰结构恢复与高分辨率细节重建的协同处理,从而提高道路裂缝、路面纹理和病害边缘的复原质量,为后续道路病害检测、分割和技术状况评价提供更加清晰、可靠的图像输入。
Smart Images

Figure CN122656920B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and in particular relates to a method and system for motion blur removal and super-resolution reconstruction of road images. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the continuous expansion of road infrastructure, automatic detection and intelligent maintenance of road defects have gradually become an important development direction in the field of road engineering.
[0004] In actual road image acquisition, the acquisition platform is usually in motion, making images susceptible to motion blur due to factors such as vehicle speed, camera shake, road surface bumps, exposure time, acquisition distance, and ambient lighting. Motion blur weakens the clarity of crack edges, road texture, and disease contours, making small cracks, shallow textures, and boundary areas less distinct, thus increasing the difficulty of subsequent detection, segmentation, and recognition tasks. Furthermore, road images may suffer from insufficient resolution during acquisition, transmission, storage, or long-distance shooting. Low-resolution images result in the loss of crack details, texture particles, and disease boundary information, especially for long, thin cracks, minor diseases, and local texture variations, whose effective features are easily weakened. Directly inputting low-resolution blurred images into subsequent disease recognition models may lead to inaccurate disease edge localization, missed detection of small cracks, and discontinuous segmentation results.
[0005] Most existing methods treat motion blur removal and super-resolution reconstruction as two independent tasks. However, there is a strong correlation between blur degradation, crack edges, high-frequency textures, and spatial details in road images. Simply concatenating the deblurring model and the super-resolution model is insufficient to fully leverage their complementary effects. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a method and system for motion blur removal and super-resolution reconstruction of road images, which realizes the coordinated processing of clear structure restoration and high-resolution detail reconstruction of road images, and improves the restoration quality of road cracks, pavement texture and defect edges.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for motion blur removal and super-resolution reconstruction of road images, comprising: Feature extraction is performed on the acquired road images to obtain shared shallow features; The shared shallow features are input into the motion blur removal branch, and the structural information in the road image is extracted and restored step by step through the encoder, bottleneck layer and decoder. The shared shallow features are input into the super-resolution reconstruction branch. Multiple residual SwinTransformer modules and crack frequency detail enhancement modules are used to extract high-frequency texture information and spatial detail information of the road image from the shared shallow features. The crack frequency detail enhancement module is based on wavelet decomposition and uses the low-frequency sub-band as structural guidance information to perform window cross attention calculation on the high-frequency sub-band, thereby enhancing the reconstruction capability of crack details and texture edges. By utilizing a cross-scale bidirectional feature interaction module, the correlation between features in the motion blur removal branch and the super-resolution reconstruction branch is calculated through window cross attention. The features of the motion blur removal branch and the super-resolution reconstruction branch are bidirectionally interacted to obtain the deblurred features and the super-resolution features after interaction. The shared shallow features, the deblurred features after interaction, and the super-resolution features after interaction are fused and reconstructed to obtain the reconstructed road image.
[0008] Secondly, the present invention provides a motion blur removal and super-resolution reconstruction system for road images, characterized in that it includes: The feature extraction unit is configured to: extract features from the acquired road image to obtain shared shallow features; The deblurring unit is configured to input shared shallow features into the motion blur removal branch, and extract and restore structural information in the road image step by step through the encoder, bottleneck layer and decoder. The super-resolution reconstruction unit is configured to: input shared shallow features into the super-resolution reconstruction branch, and use multiple residual SwinTransformer modules and crack frequency detail enhancement modules to extract high-frequency texture information and spatial detail information of the road image from the shared shallow features; wherein, the crack frequency detail enhancement module is based on wavelet decomposition, and uses low-frequency sub-band as structural guidance information to perform window cross attention calculation on high-frequency sub-band to enhance the reconstruction capability of crack details and texture edges; The interaction unit is configured to: utilize the cross-scale bidirectional feature interaction module to calculate the correlation between features between the motion blur removal branch and the super-resolution reconstruction branch through window cross attention, perform bidirectional interaction on the features of the motion blur removal branch and the super-resolution reconstruction branch, and obtain the deblurred features and the super-resolution features after interaction. The reconstruction unit is configured to fuse and reconstruct shared shallow features, interactive deblurred features, and interactive super-resolution features to obtain the reconstructed road image.
[0009] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0010] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0011] The above one or more technical solutions have the following beneficial effects: In this invention, a motion blur removal branch is used to extract and restore structural information from road images step by step through an encoder, bottleneck layer, and decoder. A super-resolution reconstruction branch extracts high-frequency texture and spatial detail information from the road image, providing detailed features for subsequent high-resolution image reconstruction. A cross-scale bidirectional feature interaction module enables bidirectional interaction between the features in the motion blur removal and super-resolution reconstruction branches. This allows the detailed features in the super-resolution reconstruction branch to assist the motion blur removal branch in restoring edge structures, and the clear structural features in the motion blur removal branch to assist the super-resolution reconstruction branch in reconstructing high-resolution textures. Finally, shared shallow features, interacted deblurred features, and interacted super-resolution features are fused and reconstructed to obtain the reconstructed road image. This invention achieves collaborative processing of clear structural restoration and high-resolution detail reconstruction of road images, thereby improving the restoration quality of road cracks, pavement textures, and defect edges, providing clearer and more reliable image input for subsequent road defect detection, segmentation, and technical condition evaluation.
[0012] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0013] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0014] Figure 1 This is a schematic diagram of the overall network structure in an embodiment of the present invention; Figure 2 This is a structural diagram of the Restormer module in an embodiment of the present invention; Figure 3 This is a structural diagram of the MDAKB module in an embodiment of the present invention; Figure 4 This is a structural diagram of the RSTB module in an embodiment of the present invention; Figure 5 This is a structural diagram of the CFDEB module in an embodiment of the present invention; Figure 6 This is a structural diagram of the BWCA module in an embodiment of the present invention; Figure 7 This is a structural diagram of the ACFF module in an embodiment of the present invention; Detailed Implementation It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0015] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0016] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0017] Example 1 This embodiment proposes a motion blur removal and super-resolution reconstruction method for road images. This method is primarily used to process low-resolution blurred road images generated during acquisition by road inspection vehicles, drones, or vehicle-mounted cameras. By constructing motion blur removal and super-resolution reconstruction branches, the model can simultaneously recover blurred edges, crack contours, road surface textures, and high-resolution details from the road image.
[0018] This embodiment discloses a method for motion blur removal and super-resolution reconstruction of road images, including: Feature extraction is performed on the acquired road images to obtain shared shallow features; The shared shallow features are input into the motion blur removal branch, and the structural information in the road image is extracted and restored step by step through the encoder, bottleneck layer and decoder. The shared shallow features are input into the super-resolution reconstruction branch. Multiple residual SwinTransformer modules and crack frequency detail enhancement modules are used to extract high-frequency texture information and spatial detail information of the road image from the shared shallow features. The crack frequency detail enhancement module is based on wavelet decomposition and uses the low-frequency sub-band as structural guidance information to perform window cross attention calculation on the high-frequency sub-band, thereby enhancing the reconstruction capability of crack details and texture edges. By utilizing a cross-scale bidirectional feature interaction module, the correlation between features in the motion blur removal branch and the super-resolution reconstruction branch is calculated through window cross attention. The features of the motion blur removal branch and the super-resolution reconstruction branch are bidirectionally interacted to obtain the deblurred features and the super-resolution features after interaction. The shared shallow features, the deblurred features after interaction, and the super-resolution features after interaction are fused and reconstructed to obtain the reconstructed road image.
[0019] like Figure 1 As shown, the proposed overall network consists of a shared shallow feature extraction module, a motion blur removal branch, a super-resolution reconstruction branch, a cross-scale bidirectional feature interaction module, an adaptive complementary feature fusion module, and an image reconstruction module. Its overall processing flow is as follows: low-resolution blurred road image → shared shallow feature extraction → parallel processing of the motion blur removal and super-resolution reconstruction branches → cross-scale bidirectional feature interaction → adaptive complementary feature fusion → high-resolution clear road image output.
[0020] The overall network input is a low-resolution blurred road image, denoted as I. BLR The overall network output is a high-resolution, clear road image, denoted as I. HR_pred .
[0021] First, input image I BLR The data is fed into the shared shallow feature extraction module, where shared shallow features F0 are extracted through two-dimensional convolution with a kernel size of 3×3.
[0022] F0=Conv 3×3 (I BLR ) Among them, I BLR This represents a low-resolution, blurred road image as input; Conv 3×3 F0 represents a 2D convolution operation with a kernel size of 3×3; F0 represents the shared shallow features extracted from the input road image. These shared shallow features contain the basic texture, edge, and color information of the road image and serve as input to both the motion blur removal branch and the super-resolution reconstruction branch.
[0023] Subsequently, the shared shallow feature F0 is input to the motion blur removal branch. This branch is mainly used to recover blurred edges and degraded textures in road images caused by vehicle motion, camera shake, road vibration, or long exposure times. The deblurring feature output by the motion blur removal branch is denoted as F. D .
[0024] F D =φ D (F0) Where, φ D Indicates the motion blur removal branch; F D This indicates the deblurring feature.
[0025] This motion blur removal branch employs an encoder, bottleneck layer, and decoder structure to recover road crack, pavement texture, and defect edge information through multi-scale feature extraction and progressive upsampling. The motion blur removal branch is primarily used to recover motion blur in road images caused by vehicle movement, camera shake, road vibration, or long exposure times. This branch takes shared shallow features F0 as input and progressively extracts and recovers structural information from the road image through the encoder, bottleneck layer, and decoder to obtain the deblurred features F. D .
[0026] The motion blur removal branch adopts an encoder-bottleneck-decoder structure. The encoder is used to progressively reduce the feature map spatial resolution and increase the number of channels to extract blur degradation features at different scales; the bottleneck layer is used to model deep structural information in the low-resolution feature space; the decoder is used to progressively restore the feature map resolution and, combined with the skip connection features in the encoder, restore detailed information such as road cracks, pavement texture, and defect edges.
[0027] The process of the motion blur removal branch for the shared shallow feature F0 is as follows: First, the shared shallow feature F0 is subjected to a two-dimensional convolution operation with a kernel size of 3×3 to extract the first feature F1. F1=Conv 3×3 (F0) Subsequently, the first feature F1 is input into the first-level deblurring coding module to further extract shallow texture, local edges and blurring degradation information in the road image, thus obtaining the first-level coding feature E1.
[0028] E1=MDAKB(Restormer(F1)) Here, Restormer represents the attention-based feature extraction unit, which is used to perform global information modeling and local texture enhancement on the first input feature F1; MDAKB represents the motion-aware orientation adaptive convolution module, which is used to adaptively enhance features of different scales and directions according to the degree and direction of motion blur in the road image; E1 represents the first-level deblurring encoding feature.
[0029] Subsequently, the first-level deblurred coding feature E1 is downsampled to obtain the first low-resolution feature X1.
[0030] X1=Down(E1) Here, Down indicates a downsampling operation.
[0031] After downsampling, the spatial resolution of the feature map decreases while the number of channels increases, enabling the network to extract blurred and degraded information over a larger receptive field.
[0032] Then, the first low-resolution feature X1 is input into the second-level deblurring coding module to obtain the second-level coding feature E2.
[0033] E2=MDAKB(Restormer(X1)) Next, the second-level encoded feature E2 is downsampled to obtain the second low-resolution feature X2.
[0034] X2=Down(E2) Then, the second low-resolution feature X2 is input into the third-level deblurring coding module to obtain the third-level coding feature E3.
[0035] E3=MDAKB(Restormer(X2)) In this embodiment, the third-level coding feature E3 can further interact with the previous super-resolution features of the super-resolution reconstruction branch across scales, so that the detailed information in the super-resolution reconstruction branch can participate in the deblurring feature recovery.
[0036] The feature after interaction is denoted as E3. / Subsequently, the interactive feature E3 will be... / Downsampling is performed and input into the bottleneck layer to obtain bottleneck feature B.
[0037] B=MDAKB(Restormer(Down(E3 / ))) Here, B represents the bottleneck characteristic.
[0038] The bottleneck layer is located at the lowest resolution of the network and is mainly used to extract deep semantic information and global structural information, thereby enhancing the model's ability to express the overall direction of road cracks, the continuity of road surface texture, and fuzzy degradation patterns.
[0039] After bottleneck layer feature extraction is completed, the bottleneck features enter the decoder section. First, the bottleneck feature B is upsampled and then interacted with the third-level encoded feature E3. / The features are then concatenated and then the third-level decoding feature D3 is obtained through feature compression and deblurring decoding modules.
[0040] D3=MDAKB(Restormer(Reduce(Concat(Up(B)),E3 / )))) Here, Up represents upsampling; Concat represents channel-level concatenation; Reduce represents channel compression; and D3 represents third-level decoding features. This is related to E3. / By using skip connection fusion, the decoder can supplement the structural information retained during the encoding stage and reduce the loss of details during upsampling.
[0041] Then, the third-level decoding feature D3 is upsampled and concatenated with the second-level encoding feature E2 to obtain the second-level decoding feature D2.
[0042] D2=MDAKB(Restormer(Reduce(Concat(Up(D3),E2)))) Here, D2 represents the second-level decoding feature.
[0043] This stage primarily focuses on restoring mid-scale texture and hazard edge information in road images.
[0044] Next, the second-level decoding feature D2 is upsampled and concatenated with the first-level encoding feature E1 to obtain the first-level decoding feature D1.
[0045] D1=MDAKB(Restormer(Reduce(Concat(Up(D2),E1)))) Where D1 represents the first-level decoding feature.
[0046] This stage gradually restores the image to the same spatial resolution as the input image, which is used to generate the final deblurred features.
[0047] Finally, the first-level decoded feature D1 is input into the deblurring feature output layer to obtain the output feature F of the motion blur removal branch. D .
[0048] F D =Conv 3×3 (D1) Among them, Conv 3×3 This represents a two-dimensional convolution operation with a kernel size of 3×3; F D This represents the deblurring feature output from the motion blur removal branch.
[0049] Among them, such as Figure 2 As shown, the Restormer module is mainly used to enhance the global correlation and local texture representation capabilities of road image features. During motion blur removal, road cracks, pavement textures, and defect edges are often affected by motion blur, leading to local structural discontinuities and weakened edge responses. Therefore, this invention introduces the Restormer module into the deblurring encoding module to perform attention modeling and feedforward feature enhancement on the input features. Its network structure diagram is shown below. Figure 2 As shown.
[0050] Let the input features of the Restormer module be... The Restormer module first processes the input features Normalization is performed, and then the data is input into the Multi-Head Transposed Attention Module (MDTA) to obtain attention-enhanced features.
[0051] = +MDTA(LN( )) in, LN represents the input features; LN represents the two-dimensional layer normalization operation; MDTA represents the multi-head transpose attention module. This indicates features enhanced by attention.
[0052] By using residual connections, correlation information between different channels and spatial locations can be introduced while preserving the original feature information.
[0053] In the MDTA module, input features First, a 1×1 convolution is used for channel mapping, and then a 3×3 depthwise convolution is used to introduce local spatial information, resulting in query feature Q, key feature K, and value feature V.
[0054] Q,K,V=Split(DWConv 3×3 (Conv 1×1 ( ))) Among them, Conv 1×1 This represents a 2D convolution with a kernel size of 1×1, used for channel transformation; DWConv 3×3 This indicates a depthwise convolution with a kernel size of 3×3, used to extract local neighborhood information; Split indicates that the convolution output is divided into three parts: Q, K, and V; Q represents the query feature; K represents the key feature; and V represents the value feature.
[0055] Then, Q and K are normalized, and the attention weights between them are calculated.
[0056] =Softmax((Normalize(Q)·Normalize(K) T )τ) in, τ represents the attention weight; τ represents the learnable scaling parameter; Normalize(·) represents the normalization operation.
[0057] This attention weight is used to represent the degree of correlation between different feature channels or between different feature responses.
[0058] Next, the attention weights Applying this to the value feature V yields the attention output feature F. attn .
[0059] F attn =Conv 1×1 ( ·V) Among them, F attn Represents the attention output features; Conv 1×1 This indicates the output projection convolution. Through this process, the Restormer module can enhance the correlation between crack edges, texture direction, and disease contours in road images, improving the model's ability to recover blurred structures.
[0060] Features after attention enhancement Continue inputting a gated depthwise convolutional feedforward network (GDFN) to further enhance local texture and nonlinear representation capabilities.
[0061] = +GDFN(LN( )) Wherein, GDFN represents a gated depthwise convolutional feedforward network; This indicates the output characteristics of the Restormer module.
[0062] GDFN further refines the attention-enhanced features through channel expansion, depthwise convolution, and gated activation.
[0063] In the GDFN module, input features First, the channels are expanded using 1×1 convolution, then local features are extracted using 3×3 depth convolution, and then the features are divided into two parts and fused using a gating method.
[0064] U,V=Split(DWConv 3×3 (Conv 1×1 (LN( )))) F ffn =Conv 1×1 (GELU(U)·V) Where U and V represent the two sets of features after partitioning; GELU represents the nonlinear activation function; F ffn This represents the feedforward enhancement feature. By element-wise multiplying GELU(U) with V, GDFN can filter out more effective texture and edge responses and suppress irrelevant background interference.
[0065] The Restormer module achieves global correlation modeling and local texture enhancement of road image features through a process of "normalization-multi-head transposed attention-residual connection-normalization-gated feedforward network-residual connection", providing a more stable feature foundation for the subsequent MDAKB module to perceive motion blur direction and blur degree.
[0066] Among them, such as Figure 3As shown, the MDAKB module, or Motion-Aware Direction Adaptive Convolutional Module, is mainly used to adaptively enhance features at different scales and in different directions based on the degree and direction of motion blur in road images. Since image blurring from road inspection vehicles, drones, or vehicle-mounted cameras typically exhibits directional and degree variations, using only a fixed convolutional kernel is insufficient to simultaneously adapt to slight blur, severe blur, and blur degradation in different directions. Therefore, this embodiment designs the MDAKB module to enhance features from both blur intensity and blur direction perspectives.
[0067] Let the input characteristics of the MDAKB module be... First, the MDAKB module uses the fuzzy intensity estimator BIE to process the input features. The analysis is performed to generate weights for convolutional branches at different scales.
[0068] W=Softmax(BIE( )) W=[w3,w5,w7] Here, BIE represents the blur intensity estimator, which consists of a 1×1 convolution, a 3×3 depthwise convolution, a GELU activation function, and another 1×1 convolution; Softmax represents the normalization function; w3, w5, and w7 represent the weights corresponding to the 3×3, 5×5, and 7×7 convolution branches, respectively. This process enables the model to adaptively select convolutional responses of different scales based on the degree of blur in local areas of the road image.
[0069] Subsequently, input features By passing through separable convolutional branches of depths of 3×3, 5×5 and 7×7 respectively, multi-scale features under different receptive fields are obtained.
[0070] F3=SepConv 3×3 ( ) F5=SepConv 5×5 ( ) F7=SepConv 7×7 ( ) Among them, SepConv 3×3 SepConv 5×5 and SepConv 7×7 The numbers 3×3, 5×5, and 7×7 represent depthwise separable convolutions with kernel sizes of 3×3, 5×5, and 7×7, respectively; F3, F5, and F7 represent the features extracted at different scales. Smaller kernels are better suited for extracting local edges and small cracks, while larger kernels are better suited for perceiving larger areas of blurred and degraded regions.
[0071] Based on the fuzzy intensity weights, features at different scales are weighted and fused to obtain the multi-scale fuzzy perception features F. scale .
[0072] F scale =w3·F3+w5·F5+w7·F7 Among them, F scale This represents multi-scale blurred perception features. Through this weighted fusion method, the model can dynamically adjust the receptive field of the convolution according to the blur level of different regions, thereby enhancing its adaptability to motion blur degradation of road images.
[0073] While acquiring multi-scale features, the MDAKB module also sets up directional convolution branches to extract blurred structures and texture responses in different directions. Input features Four types of directionally related features are obtained by performing horizontal convolution, vertical convolution, dilated convolution, and 1×1 convolution branches respectively.
[0074] F h =SepConv 1×7 ( ) F v =SepConv 7×1 ( ) F c =SepConv 3×3 ,d=2( ) F o =Conv 1×1 ( ) Among them, F h Indicates horizontal features; F v Indicates vertical characteristics; F c F represents the dilated convolution branch feature; o Represents channel mapping features; SepConv 1×7 This indicates a depthwise separable convolution with a kernel size of 1×7 in the horizontal direction; SepConv 7×1 This indicates a depthwise separable convolution with a kernel size of 7×1 in the vertical direction; SepConv 3×3 d=2 indicates a 3×3 hole-depth separable convolution with an inflation rate of 2; Conv 1×1 This represents a two-dimensional convolution with a kernel size of 1×1.
[0075] Then, the features of the four directional branches are concatenated along the channel dimension and input into the Directional Weight Generator (DWG) to obtain the adaptive weights of different directional branches.
[0076] =Softmax(DWG(Concat(F h ,F v ,F c ,F o ))) =[a h ,a v ,a c ,a o ] Where Concat represents channel dimension concatenation; DWG represents the orientation weight generator, which consists of sequentially processed 1×1 convolutions, GELU activation functions, and 1×1 convolutions; a h a v a c and a o These represent the weights for the horizontal, vertical, dilated convolutional, and 1×1 convolutional branches, respectively. This process enables the model to adaptively select more effective directional features based on the blurred directions and texture orientation in the road image.
[0077] Based on the direction weights, the features of different directions are weighted and fused to obtain the direction-aware feature F. dir .
[0078] F dir =a h ·F h +a v ·F v +a c ·F c +a o ·F o Among them, F dir This represents the orientation-aware features. Through this orientation-weighted process, the model can enhance edge and texture responses related to motion-blurred orientations, which helps to recover road cracks, pavement textures, and defect contours.
[0079] Next, the multi-scale fuzzy perception feature F scale With orientation perception feature F dir The features are concatenated along the channel dimension and then fused using a 1×1 convolution to obtain the comprehensive motion blur enhancement feature F. m .
[0080] F m =Conv 1×1 (Concat(F scale ,F dir )) Among them, F m This represents the enhanced motion blur feature after fusion. This feature simultaneously includes multi-scale responses under different blur levels and directional responses under different blur directions.
[0081] To avoid over-correction caused by direct superposition of enhanced features, the MDAKB module further introduces a gated residual fusion mechanism. Specifically, the input features are... With enhanced features F m The components are concatenated, and the gating weights G are generated using the Sigmoid function.
[0082] G = Sigmoid(Conv) 1×1 (Concat( ,F m ))) Where G represents the gating weight; Sigmoid represents the Sigmoid activation function. The gating weight is used to control the strength of the enhanced features at different spatial locations and channels.
[0083] Finally, the gated enhanced features are compared with the input features. Residual fusion is performed to obtain the output features of the MDAKB module. .
[0084] = +G·F m in, This indicates the output characteristics of the MDAKB module; Represents input features; G·F m This indicates the enhanced features after gating control.
[0085] Through residual connections, the MDAKB module can selectively enhance the orientation structure and multi-scale texture information related to motion blur while preserving the original feature information.
[0086] The MDAKB module uses a processing flow of "fuzz intensity estimation - multi-scale convolution enhancement - directional convolution response - directional weight generation - gated residual fusion" to enable the network to perform adaptive feature enhancement based on the degree and direction of fuzziness in different regions of the road image, thereby improving the ability of the motion blur removal branch to restore crack edges, road surface texture and disease contours.
[0087] Through the above structure, the motion blur removal branch can progressively restore the clear structure of the road image in a multi-scale feature space. On the one hand, the encoder and bottleneck layer can extract deep blur degradation features in the road image; on the other hand, the decoder and skip connections can preserve and restore detailed information such as road cracks, pavement texture, and defect edges.
[0088] In this embodiment, the shared shallow feature F0 is input into the super-resolution reconstruction branch. This super-resolution reconstruction branch is mainly used to extract high-frequency texture information and spatial detail information from the road image, and to provide detailed features for subsequent high-resolution image reconstruction. The super-resolution feature output by the super-resolution reconstruction branch is denoted as F. S .
[0089] F S =φ S (F0) Where, φ S Indicates the super-resolution reconstruction branch; F S This branch represents super-resolution features. It further models crack details, texture structure, and defect contours in road images through residual feature extraction, window self-attention, and frequency detail enhancement.
[0090] The super-resolution reconstruction branch is mainly used to extract high-frequency texture information and spatial detail information from shared shallow features of road images, providing more comprehensive detail features for subsequent high-resolution image reconstruction. Since cracks, textures, and defect edges in road images are typically small, continuous, and directional, simple upsampling can easily lead to blurred crack edges, lost texture details, and unclear defect boundaries. Therefore, this embodiment sets up a super-resolution reconstruction branch to perform deep modeling of the detailed structure of road images.
[0091] The specific operation process of the super-resolution reconstruction branch on the input shared shallow feature F0 is as follows: First, the shared shallow feature F0 undergoes a two-dimensional convolution operation with a kernel size of 3×3 to obtain the initial feature S0 of the super-resolution reconstruction branch.
[0092] S0=Conv 3×3 (F0) Where F0 represents shared shallow features; Conv 3×3 S0 represents a 2D convolution operation with a kernel size of 3×3; S0 represents the initial feature of the super-resolution reconstruction branch. This feature provides input for subsequent super-resolution feature reconstruction while preserving the shallow texture information of the original road image.
[0093] Subsequently, the initial feature S0 of the super-resolution reconstruction branch is sequentially input into multiple residual SwinTransformer modules and crack frequency detail enhancement modules to gradually extract local texture, long-distance dependencies and high-frequency detail information in the road image.
[0094] The initial feature extraction process of the super-resolution reconstruction branch can be represented as: S1=CFDEB(RSTB(RSTB(S0))) In this module, RSTB represents the Residual SwingTransformer module; CFDEB represents the Crack Frequency Detail Enhancement module; and S1 represents the Pre-process Super-resolution Features. This stage is mainly used to extract the basic texture, crack edges, and shallow high-frequency details from the road image.
[0095] In order for the super-resolution reconstruction branch to complement the motion blur removal branch, the initial super-resolution feature S1 and the third-level encoded feature E3 in the motion blur removal branch are used. / The first cross-scale bidirectional feature interaction is performed. Through this process, the super-resolution reconstruction branch can obtain sharp structural information from the motion blur removal branch, while the motion blur removal branch can also obtain detail enhancement information from the super-resolution reconstruction branch. The feature after the interaction is denoted as S1. / .
[0096] Super-resolution features S1 after the first interaction / Continuing with the intermediate super-resolution feature extraction stage, the process can be represented as follows: S2=CFDEB(RSTB(RSTB(S1 / ))) S2 represents the intermediate super-resolution features. This stage further enhances the ability to represent texture continuity, crack direction, and disease contours in road images.
[0097] Subsequently, the super-resolution features enter the mid-to-late stage of detail reconstruction, which can be represented as follows: S3=CFDEB(RSTB(S2)) S3 represents the mid-to-late stage super-resolution features. This stage is mainly used to further recover small cracks, high-frequency textures, and details of road damage boundaries in the road image.
[0098] Subsequently, the mid-to-late stage super-resolution feature S3 and the third-level decoding feature E3 in the motion blur removal branch. / A second cross-scale bidirectional feature interaction is performed, enabling the super-resolution reconstruction branch to further utilize the deblurred structural information, thereby improving the accuracy of high-resolution detail reconstruction. The features after the interaction are denoted as S3. / .
[0099] Subsequently, after interaction, feature S3 / The process of entering the later detailed reconstruction stage can be represented as follows: S4=RSTB(S3 / ) Finally, the later super-resolution feature S4 is input into the convolutional output layer to obtain the super-resolution reconstruction branch output feature F. S .
[0100] FS = Conv 3×3 (S4) Among them, F S S4 represents the output features of the super-resolution reconstruction branch; Conv represents the later super-resolution features. 3×3 This represents a two-dimensional convolution operation with a kernel size of 3×3.
[0101] In summary, the super-resolution reconstruction branch achieves deep reconstruction of crack details, pavement texture, and defect edges in road images through a process of "shallow feature mapping - RSTB long-distance dependency modeling - CFDEB high-frequency detail enhancement - cross-scale bidirectional interaction," providing effective feature support for the final output of high-resolution, clear road images.
[0102] like Figure 4 As shown, the RSTB module, or Residual Swing Transformer module, is mainly used for window self-attention modeling and local texture enhancement of road image features. Cracks, pavement textures, and defect edges in road images often have characteristics such as being long and thin, continuous, and having obvious local directionality. Ordinary convolution alone is insufficient to fully model structural relationships over a large area. Therefore, this invention introduces the RSTB module into the super-resolution reconstruction branch to enhance the model's ability to express local textures and long-distance dependencies in road images.
[0103] Let the input characteristics of the RSTB module be... The RSTB module consists of two SwingTransformer layers and a 3×3 convolutional layer, and uses residual connections to output features. The first SwingTransformer layer uses standard window self-attention, while the second SwingTransformer layer uses shifted window self-attention.
[0104] The overall process of the RSTB module can be represented as follows: =STL1( ) =STL2( ) = +Conv 3×3 ( ) in, STL1 represents the input characteristics of the RSTB module; STL2 represents the first SwingTransformer layer; STL2 represents the second SwingTransformer layer. This represents the output feature of the first SwingTransformer layer; This represents the output feature of the second SwingTransformer layer; Conv 3×3 This represents a two-dimensional convolution operation with a kernel size of 3×3; This represents the output features of the RSTB module. Through residual connections, the RSTB module can further enhance the texture details and structural expression of road images while preserving the input features.
[0105] In the first SwinTransformer layer, the input features are... First, the data is normalized, and then feature association modeling within a local window is performed using a standard window self-attention module. This process can be represented as follows: X 1a = +WSA(LN( )) Where LN represents layer normalization operation; WSA represents ordinary window self-attention module; X 1a This represents the features enhanced by the normal window self-attention module. The normal window self-attention module divides the feature map into multiple fixed-size local windows and calculates the correlation between pixels or feature locations within each window, thereby enhancing the spatial connection between crack edges, road textures, and disease contours within local areas.
[0106] Subsequently, the feature X after self-attention enhancement using a normal window 1a The input feedforward feature enhancement network enhances channel information and local nonlinear features, resulting in the output features of the first SwingTransformer layer. .
[0107] =X 1a +MLP(LN(X 1a )) Wherein, MLP represents a feedforward feature enhancement network; This represents the output features of the first SwingTransformer layer. This process further enhances the texture response and edge details in road images while suppressing irrelevant background interference.
[0108] In the second SwinTransformer layer, the input features Similarly, after normalization, feature association modeling is performed using a shifted window self-attention module. The process can be represented as follows: X 2a = +SWSA(LN( )) Where SWSA represents the shift window self-attention module; X 2aThis represents the features enhanced by shifted window self-attention. Shifted window self-attention changes the window partitioning position, allowing features that were originally in adjacent windows to enter the same window for information interaction, thus overcoming the limitation of ordinary window self-attention, which can only model within a single window.
[0109] Subsequently, the feature X after shift window self-attention enhancement 2a The input is fed forward feature enhancement network to obtain the output features of the second SwinTransformer layer. .
[0110] =X 2a +MLP(LN(X 2a )) in, This represents the output feature of the second SwingTransformer layer. This feature, based on the local modeling of a normal window, further integrates texture and structural information between adjacent windows, which is beneficial for recovering continuous cracks, cross-window textures, and details of disease boundaries.
[0111] The window self-attention calculation process can be represented as: Q,K,V=Linear(X w ) =Softmax((Q·K T ) / √d) F attn = ·V Among them, X w The window features are represented after partitioning; Q, K, and V represent query features, key features, and value features, respectively; K T d represents the transpose of K; d represents the feature dimension of a single attention head; F represents attention weights; attn This represents the output feature of window self-attention. For ordinary window self-attention, X... w Features derived from fixed window partitioning; for shifted window self-attention, X w The features are derived from the re-division after window shifting.
[0112] In summary, the RSTB module enhances the expressive power of road image features while maintaining computational efficiency through a process of "ordinary window self-attention—feedforward feature enhancement—shifted window self-attention—feedforward feature enhancement—3×3 convolution—residual connection". Ordinary window self-attention is used to extract crack edges and texture relationships within local regions, while shifted window self-attention is used to enable information exchange between adjacent windows. The combination of the two can improve the model's ability to reconstruct continuous cracks, pavement textures, and defect edges, providing effective support for the recovery of high-resolution details in the super-resolution reconstruction branch.
[0113] like Figure 5 As shown, the CFDEB module, or Crack Frequency Detail Enhancement Module, is mainly used to enhance high-frequency detail information in road images. Information such as small cracks, texture particles, and defect boundaries in road images is usually concentrated in high-frequency components, and ordinary convolutional or attention modules may not be able to adequately recover these details during reconstruction. Therefore, this invention introduces the CFDEB module into the super-resolution reconstruction branch, improving the model's ability to reconstruct crack details and texture edges through wavelet decomposition and high-frequency subband cross-attention enhancement.
[0114] Let the input features of the crack frequency detail enhancement module be... First, the CFDEB module processes the input features. Perform Haar wavelet decomposition to decompose it into one low-frequency subband and three high-frequency subbands.
[0115] LL,LH,HL,HH=DWT( ) Wherein, DWT represents the Haar wavelet decomposition operation; LL represents the low-frequency subband, which mainly contains the overall structure and brightness information of the road image; LH, HL and HH represent the high-frequency subband, which mainly contains the edge, texture, crack details and orientation change information of the road image.
[0116] Subsequently, the CFDEB module uses the low-frequency subband LL as structural guidance information to perform window cross-attention enhancement on the three high-frequency subbands LH, HL and HH respectively.
[0117] D LH =WCA(LH,LL) D HL =WCA(HL,LL) D HH =WCA(HH,LL) Where WCA represents window cross attention calculation; D LH D HL and D HHThese represent three high-frequency enhancement features obtained by guiding low-frequency structural information. Through this process, the model can utilize the overall structural information in the low-frequency subband to guide the detail recovery of the high-frequency subband, avoiding invalid textures or artifacts generated during the high-frequency enhancement process.
[0118] Then, the high-frequency enhancement features are superimposed back onto the corresponding high-frequency subbands to obtain the enhanced high-frequency subbands.
[0119] LH'=LH+αLH·D LH HL' = HL + αHL·D HL HH'=HH+αHH·D HH Where LH', HL', and HH' represent the enhanced high-frequency subbands; αLH, αHL, and αHH represent the learnable high-frequency enhancement coefficients, all initialized to 0.1, used to control the enhancement intensity in different high-frequency directions.
[0120] After completing the high-frequency subband enhancement, the low-frequency subband LL and the enhanced high-frequency subbands LH', HL', and HH' are subjected to Haar wavelet inverse transform to reconstruct the frequency enhancement feature Xrec.
[0121] X rec =IDWT(LL,LH',HL',HH') Where IDWT represents the inverse Haar wavelet transform; X rec This represents the reconstructed features after frequency detail enhancement.
[0122] Finally, the reconstructed feature X will be... rec After 1×1 convolution mapping, and with the input features The residuals are summed to obtain the output features of the CFDEB module. .
[0123] = +Conv 1×1 (X rec ) in, Indicates the output characteristics of the CFDEB module; Conv 1×1 This represents a 2D convolution operation with a kernel size of 1×1. Through residual connections, the CFDEB module can further supplement high-frequency textures and crack details while preserving the original feature information.
[0124] In summary, the CFDEB module enhances the crack edges, texture details, and disease contours in road images in the frequency domain through the method of "Haar wavelet decomposition - low-frequency structure guidance - high-frequency subband cross-attention enhancement - inverse wavelet transform - residual fusion", thereby improving the ability of the super-resolution reconstruction branch to restore high-resolution details.
[0125] To avoid the two tasks of motion blur removal and super-resolution reconstruction being independent of each other, this embodiment sets up a cross-scale bidirectional feature interaction module between the two branches. This cross-scale bidirectional feature interaction module is used to realize the information transfer between deblurring features and super-resolution features. On the one hand, the detailed features in the super-resolution reconstruction branch can help the motion blur removal branch recover edge structures; on the other hand, the sharp structural features in the motion blur removal branch can help the super-resolution reconstruction branch reconstruct high-resolution textures.
[0126] Since the feature scales of the motion blur removal branch and the super-resolution reconstruction branch may be different, the feature sizes of the two branches are aligned by upsampling or downsampling operations before bidirectional interaction, and then the correlation between the two branches is calculated by window cross attention.
[0127] The cross-scale bidirectional feature interaction process can be represented as: D F / = D F +α D ·WCA(D F ,S F ) S F / = S F +α S ·WCA(S F D F ) Among them, D F With S F This represents the features of the motion blur removal branch and the super-resolution reconstruction branch before interaction; D F / With S F / The features of the motion blur removal branch and the super-resolution reconstruction branch after interaction are represented; WCA represents window cross-attention calculation; α D and α S α represents the learnable feature interaction strength coefficient. D and α S The initial value is 0. Through this bidirectional interaction, the motion blur removal branch and the super-resolution reconstruction branch can form a complementary relationship at the feature level, rather than simply being linked one after the other.
[0128] In this embodiment, as Figure 6 As shown, the BWCA module, or cross-scale bidirectional feature interaction module, is mainly used to realize information interaction between the motion blur removal branch and the super-resolution reconstruction branch. The motion blur removal branch focuses more on restoring sharp structures in road images, such as crack edges, defect contours, and texture continuity; while the super-resolution reconstruction branch focuses more on restoring high-frequency details in road images, such as fine cracks, texture grains, and boundary details. If the two branches extract features completely independently, it is difficult for the deblurring features and super-resolution features to form an effective complementarity. Therefore, this embodiment designs a cross-scale bidirectional feature interaction module to enable bidirectional information transmission between the two branches at the feature level.
[0129] Let the input feature of the motion blur removal branch be D. F The super-resolution reconstruction branch input feature is S F Among them, D F The spatial dimensions are HD×WD, and the number of channels is CD; S F The spatial dimensions are HS×WS, and the number of channels is CS. Since the two branches may be at different scales in the network, channel mapping and spatial scale alignment of the features of the two branches are required before interaction.
[0130] First, information transfer is performed from the super-resolution reconstruction branch to the motion blur removal branch. This process uses the feature D of the motion blur removal branch. F As a query feature, the feature S of the branch is reconstructed in super-resolution. F As a key-value feature, it is used to enable the detailed information in the super-resolution reconstruction branch to assist the motion blur removal branch in restoring a clear structure.
[0131] First, remove the branch of feature D from the motion blur. F Perform a 1×1 convolution mapping to obtain the query feature Q. D .
[0132] Q D =Conv 1×1 (D F ) Among them, Q D This represents the query feature corresponding to the motion blur removal branch; Conv 1×1 This represents a 2D convolution operation with a kernel size of 1×1. This operation is used to map motion blur removal branch features to a unified attention computation dimension.
[0133] Then, the super-resolution branch feature S F The interpolation operation is used to adjust the feature D to match the deblurred branch feature. F For the same spatial dimensions, scale-aligned super-resolution features F are obtained.S_down .
[0134] F S_down =Resize(F S HD, WD) Where Resize represents the scaling operation; F S_down This represents a super-resolution feature with the same branch space size as the motion blur removal branch.
[0135] Next, the scale-aligned super-resolution features F S_down Perform a 1×1 convolution mapping to obtain key-value features (KVS).
[0136] KV S =Conv 1×1 (F S_down ) Among them, KV S This represents the key-value features generated by the super-resolution reconstruction branch.
[0137] Then, Q D and KV S The input window cross-attention module yields the supplementary feature ΔF of the super-resolution reconstruction branch to the motion blur removal branch. D .
[0138] ΔF D =WCA(Q D KV S ) Where WCA represents window cross attention calculation; ΔF D This represents the deblurring supplementary features obtained by super-resolution feature guidance. This process can transfer the detailed texture and high-frequency information in the super-resolution reconstruction branch to the motion blur removal branch, which helps to restore the edges and texture details of road cracks.
[0139] Then, the super-resolution reconstruction branch supplements the motion blur removal branch with the feature ΔF. D After a 1×1 convolution mapping back to the original channel dimension of the motion blur removal branch, and then using a learnable coefficient α... D Perform residual update to obtain the feature D of the interactive motion blur removal branch. F / .
[0140] D F / =D F +α D ·Conv 1×1 (ΔF D ) Among them, D F / Features representing the motion blur removal branch after interaction; α D This represents the interaction strength coefficient of the learnable motion blur removal branch. This coefficient is used to control the degree of influence of super-resolution features on the motion blur removal branch, avoiding excessive interference of external branch information with the original deblurring features.
[0141] Secondly, information transfer is performed from the motion blur removal branch to the super-resolution reconstruction branch. This process uses the features S of the super-resolution reconstruction branch. F As a query feature, feature D is used to remove branches using motion blur. F As a key feature, it is used to enable the clear structural information in the motion blur removal branch to assist the super-resolution reconstruction branch in high-resolution detail reconstruction.
[0142] First, the features S of the super-resolution reconstruction branch are analyzed. F Perform a 1×1 convolution mapping to obtain the query feature Q. S .
[0143] Q S =Conv 1×1 (S F ) Among them, Q S This indicates the query features corresponding to the super-resolution reconstruction branch.
[0144] Then, the feature D of the motion blur removal branch is used. F The feature S is adjusted to match the super-resolution reconstruction branch through interpolation. F For the same spatial dimensions, scale-aligned deblurred features F are obtained. D_up .
[0145] F D_up =Resize(D F ,HS,WS) Among them, F D_up This represents a deblurred feature that is consistent with the branch space size of the super-resolution reconstruction.
[0146] Next, regarding F D_up Perform a 1×1 convolution mapping to obtain the key-value features (KV). D .
[0147] KV D =Conv 1×1 (F D_up ) Among them, KV D This represents the key-value features generated by the motion blur removal branch.
[0148] Then, Q S and KV DThe input window cross-attention module yields the supplementary feature ΔF of the motion blur removal branch to the super-resolution reconstruction branch. S .
[0149] ΔF S =WCA(Q S KV D ) Where, ΔF S This represents the super-resolution supplementary features obtained by deblurring features. This process can transfer the sharp structure, edge contours, and texture continuity from the motion blur removal branch to the super-resolution reconstruction branch, which helps to improve the structural consistency in the high-resolution reconstruction process.
[0150] Finally, ΔF S After a 1×1 convolution, the original channel dimension of the super-resolution reconstruction branch is mapped back to the original channel dimension, and then processed by a learnable coefficient α. S Perform residual update to obtain the interactive super-resolution features F S / .
[0151] S F / =S F +α S ·Conv 1×1 (ΔF S ) Among them, S F / Features representing the super-resolution reconstruction branch after interaction; α S This represents the interaction strength coefficient of the learnable super-resolution reconstruction branches. This coefficient controls the degree of influence of deblurring features on the super-resolution reconstruction branches.
[0152] In the window cross-attention calculation process, query features and key-value features are first divided into multiple local windows, and then cross-branch attention relationships are calculated within each window. The calculation process can be represented as follows: Q=Linear(Q win ) K,V=Linear(KV win ) A = Softmax((Q·K) T ) / √d) F attn =A·V Among them, Q win Indicates the window characteristics after query branching; KV win This represents the window features after key-value branching; Q represents the query vector; K represents the key vector; V represents the value vector; K TK represents the transpose of K; d represents the feature dimension of a single attention head; A represents the window cross-attention weights; F attn This represents the output features of window cross attention.
[0153] In summary, the BWCA module achieves information complementarity between the motion blur removal branch and the super-resolution reconstruction branch through a process of "scale alignment—channel mapping—window cross-attention—bidirectional residual update." On one hand, high-frequency details in the super-resolution reconstruction branch can assist the motion blur removal branch in restoring crack edges and texture details; on the other hand, sharp structures in the motion blur removal branch can assist the super-resolution reconstruction branch in maintaining road defect boundaries and texture continuity. Through this module, this invention avoids the simple independent processing between deblurring and super-resolution tasks, enabling the restoration of sharp structures and the reconstruction of high-resolution details in road images to mutually promote each other.
[0154] In this embodiment, the final feature F output by the motion blur removal branch is... D The final feature F output by the super-resolution reconstruction branch S And the shallow feature F0 is input into the adaptive complementary feature fusion module to obtain the fused feature F. fusion .
[0155] F fusion =ACFF(F D ,F S ,F0) Where ACFF represents the adaptive complementary feature fusion module; F fusion This represents the restored features after fusion. The adaptive complementary feature fusion module adaptively adjusts the contribution weights of deblurring features and super-resolution features based on the degree of degradation in different image regions. For severely blurred regions, the model focuses more on deblurring features; for regions with richer texture and detail, the model focuses more on super-resolution features.
[0156] Finally, the fused features F fusion The input image reconstruction module outputs a high-resolution, clear road image through convolution and upsampling operations.
[0157] I HR_pred =R(F fusion ) Where R represents the image reconstruction module (which consists of two Resblocks, 3×3 convolutions, and PixelShuffle, 3×3 convolutions); I HR_pred This represents the final output, a high-resolution, clear road image.
[0158] like Figure 7As shown, the ACFF module, or Adaptive Complementary Feature Fusion Module, is mainly used to fuse motion blur removal branch features, super-resolution reconstruction branch features, and shared shallow features. After processing by the aforementioned motion blur removal branch and super-resolution reconstruction branch, the motion blur removal branch provides strong and clear structural information, the super-resolution reconstruction branch provides rich high-frequency detail information, and the shared shallow features still retain the basic texture, color, and local edge information of the input road image. To fully utilize the three types of features, this invention designs an adaptive complementary feature fusion module, enabling the model to adaptively select more effective restoration information based on the degradation features of different regions.
[0159] Motion blur removal branch output feature F D Super-resolution reconstruction branch output features F S The shared shallow features are F0. First, for F... D F S A 1×1 convolution mapping is performed with F0 to ensure that the three types of features maintain consistency in channel representation.
[0160] F D a =Conv 1×1 (F D ) F S a =Conv 1×1 (F S ) F0 a =Conv 1×1 (F0) Among them, F D a F represents the aligned, deblurred features. S a F0 represents the aligned super-resolution features. a Represents the aligned shallow features; Conv 1×1 This represents a two-dimensional convolution operation with a kernel size of 1×1. This channel alignment process reduces the representational differences between features in different branches, providing a unified feature space for subsequent feature fusion.
[0161] Subsequently, the aligned deblurred features F are calculated. D a With super-resolution features F S a The differences between the features are diff.
[0162] Diff=|F D a -F S a | Here, Diff represents the difference between the two branches; |·| represents the element-wise absolute value operation. This difference feature can reflect the response differences between the motion blur removal branch and the super-resolution reconstruction branch in different image regions. For example, in regions with severe motion blur, the deblurring feature may be more important; in regions with rich texture details, the super-resolution feature may be more important.
[0163] Next, F D a F S a The Diff signal is concatenated along the channel dimension and input into the weight generation network to obtain the adaptive fusion weights corresponding to the motion blur removal branch and the super-resolution reconstruction branch.
[0164] W=Softmax(φ w (Concat(F D a ,F S a ,Diff))) W=[w D ,w S ] Where Concat represents the channel-level concatenation operation; φ w The weight generation network consists of a 1×1 convolution, a 3×3 depthwise convolution, a GELU activation function, and another 1×1 convolution in sequence; Softmax represents the normalization function; W represents the fused weights; w D w represents the weights corresponding to the deblurred features. S This represents the weights corresponding to the super-resolution features. Through this weight generation process, the model can adaptively determine whether the current region needs more deblurring of structural information or super-resolution detail information based on the image degradation of different regions.
[0165] Then, the deblurred features and super-resolution features are weighted and fused according to adaptive weights to obtain the preliminary fused features F. w .
[0166] F w =w D ·F D a +w S ·F S a Among them, F w This represents the initial fusion feature. This feature integrates the sharp structural information provided by the motion blur removal branch and the high-frequency detail information provided by the super-resolution reconstruction branch.
[0167] To further supplement the shallow texture and edge information in the input image, this invention will initially fuse feature F. w Aligned shallow features F0 a The input features F are concatenated to obtain the fused input features. in .
[0168] F in =Concat(F w ,F 0a ) Among them, F in This represents the fused input features used for subsequent feature refinement and gating control. A shallow feature F0 is introduced. a It can help the model retain the basic texture, color distribution and local edge information in the original road image, and avoid detail shift or texture loss during deep feature reconstruction.
[0169] Subsequently, F in Input the feature refinement network to obtain refined supplementary features F r .
[0170] F r =φ r (F in ) Where, φ r The feature refinement network consists of a 3×3 depthwise convolution, a GELU activation function, and a 1×1 convolution in sequence; F r This represents the supplementary features generated jointly by the initial fusion features and shallow features. This refinement process is used to further integrate deblurred structural information, super-resolution detail information, and shallow texture information.
[0171] At the same time, F in Input the gating generator network to obtain the gating weights G.
[0172] G=Sigmoid(φ g (F in )) Where, φ g This represents a gated generative network, consisting of 1×1 convolutions and a sigmoid activation function in sequence; sigmoid represents the sigmoid activation function; G represents the gate weights. The gate weights are used to control the supplementary features F. r The intensity of the effect at different spatial locations and channels is determined to avoid over-correction caused by direct superposition of enhancement features.
[0173] Finally, the gated supplementary features are residually fused with the preliminary fused features to obtain the final fused feature F. fusion .
[0174] F fusion =Fw +G·F r Among them, F fusion G·F represents the output feature of the adaptive complementary feature fusion module. r This represents the refined and supplemented features after gating control. Through this residual fusion method, the model can further supplement shallow textures and edge details while retaining the initial fusion results of deblurring and super-resolution.
[0175] In summary, the ACFF module achieves effective fusion of deblurred features, super-resolution features, and shallow features through a process of "feature alignment—difference modeling—adaptive weight generation—dual-branch weighted fusion—shallow feature supplementation—gated residual refinement." This module can dynamically adjust the contribution ratio of clear structural information and high-frequency detail information according to the degradation characteristics of different regions of the road image, thereby improving the restoration quality of the final high-resolution clear road image.
[0176] Through the above process, this embodiment achieves the coordinated processing of motion blur removal and super-resolution reconstruction of road images. Compared with the traditional simple concatenation method of first deblurring and then super-resolution, this embodiment uses dual-branch parallel processing, cross-scale bidirectional interaction, and adaptive complementary fusion to enable the restoration of clear structures and the reconstruction of high-resolution details in road images to promote each other, thereby improving the restoration quality of road cracks, pavement textures, and defect edges.
[0177] It should be noted that the cross-scale bidirectional feature interaction module can be set in the intermediate feature layer of the motion blur removal branch and the super-resolution reconstruction branch. In this embodiment, it acts between the third-level coding feature and the early super-resolution feature, and between the third-level decoding feature and the mid-to-late super-resolution feature, respectively, so as to realize the complementary intermediate layer features of the two branches in the forward propagation process.
[0178] In practical engineering applications, road images can be acquired by road inspection vehicles, drones, vehicle-mounted camera equipment, or mobile inspection equipment. In this embodiment, to facilitate the explanation of the complete process of model training and image restoration, the Crack500 public road crack image dataset is selected as the example data source. The road crack images in this dataset can be used as samples of clear road images to construct a pairing relationship between low-resolution blurred road images, low-resolution clear road images, and high-resolution clear road images.
[0179] Step 1: Obtain road image data.
[0180] First, clear road crack image data is acquired. In real-world engineering scenarios, road images can be collected using road inspection vehicles, drones, or vehicle-mounted camera equipment. In this embodiment, road crack images from the Crack500 dataset are used as the original clear images. The original images contain road cracks, road surface texture, local edges, and background areas, which can be used to simulate motion blur and insufficient resolution issues that exist in actual road image acquisition.
[0181] Step 2: Preprocess the original road image.
[0182] The acquired road images undergo preprocessing, including image cropping, resizing, and pixel normalization. Specifically, the original clear road image is cropped or scaled to a clear road image of size 512×512, and this image is designated as the high-resolution clear road image, denoted as I. HR .
[0183] Step 3: Construct a low-resolution, clear road image.
[0184] For a high-resolution, clear road image of size 512×512 HR A 2x downsampling was performed to obtain a low-resolution, clear road image with a size of 256×256, denoted as I. LR .
[0185] This process can be represented as: I LR =Downsample(I HR ) Among them, I HR Represents a high-resolution, clear road image; Downsample indicates a downsampling operation; I LR This indicates a low-resolution, clear road image.
[0186] Step 4: Construct a low-resolution blurred road image.
[0187] A clear road image of 256×256 was obtained. LR Then, motion blur degradation processing is applied to it to obtain a low-resolution blurred road image with a size of 256×256, denoted as I. BLR Motion blur degradation can be achieved by setting motion blur kernels with different directions, lengths, and intensities to simulate image blur caused by factors such as road inspection vehicle movement, camera shake, drone flight disturbances, or road vibrations. This process can be represented as: I BLR =MotionBlur(I LR ) Wherein, MotionBlur represents the motion blur degradation operation; I BLRThis represents a low-resolution, blurred road image. This image serves as the input image for the network model in this embodiment.
[0188] Step 5: Establish training sample pairing relationships.
[0189] After the above processing, each original road image can form a set of paired samples: 256×256 low-resolution clear road image I LR Used to generate low-resolution blurred road images I BLR ; 256×256 low-resolution blurred road image I BLR , as model input; 512×512 high-resolution clear road image I HR , as the supervision target for the final super-resolution reconstruction.
[0190] This results in a "low-resolution blurred image I" BLR —High-resolution clear images I HR The relationship between the training input and the supervised target. Among them, the low-resolution, clear image I... LR Only as a means of constructing low-resolution blurred images I BLR Intermediate degraded images are not used as model output or as separate supervision targets.
[0191] Step six, network training process.
[0192] After completing the construction of the training samples, the low-resolution blurred road image I with a size of 256×256 is used. BLR As input to the model, a high-resolution, clear road image of size 512×512 is used. HR This serves as the supervised objective for the final super-resolution reconstruction stage. During training, the model completes motion blur feature recovery, super-resolution detail reconstruction, cross-scale bidirectional feature interaction, adaptive complementary feature fusion, and image reconstruction in one forward propagation, enabling the model to learn the end-to-end mapping relationship from low-resolution blurred road images to high-resolution clear road images.
[0193] Specifically, the model training loss function consists of image reconstruction loss and structural similarity loss, and its overall loss function is expressed as: L=λ1L rec +λ2L ssim Where L represents the total loss function for model training; L rec L represents the image reconstruction loss; ssimλ1 represents the structural similarity loss; λ2 represents the weight coefficients corresponding to the image reconstruction loss and structural similarity loss, respectively, which are used to adjust the contribution of different loss terms in the model training process. In this embodiment, λ1 is set to 0.9 and λ2 is set to 0.1.
[0194] Image reconstruction loss is used to constrain the model output image I HR_pred With high-resolution clear road images I HR The formula for calculating overall consistency at the pixel level is: L rec =||I HR_pred -I HR || 1 Among them, ||·|| 1 This represents the L1 norm. Through this loss term, the model can reduce the pixel deviation between the reconstructed image and the true sharp image, making the output image approach the resolution of the sharp road image in terms of overall brightness, color, and texture distribution.
[0195] Structural similarity loss is used to constrain the model output image I HR_pred With high-resolution clear road images I HR The formula for calculating the consistency in local structure, brightness, and contrast is as follows: L ssim =1-SSIM(I HR_pred ,I HR ) Here, SSIM represents the structural similarity calculation function. By introducing structural similarity loss, the model can pay more attention to the structural preservation ability of road crack contours, pavement texture structure, and defect boundary areas during training, avoiding the problem of structural blurring or local texture discontinuities in the reconstructed image caused by relying solely on pixel errors.
[0196] During training, the total loss function L is used as the optimization objective, and the network parameters are updated through the backpropagation algorithm to achieve the desired final output image I. HR_pred In terms of pixel accuracy and structural consistency, it gradually approaches the supervised target image I. HR Therefore, the model can achieve joint optimization of motion blur removal and super-resolution reconstruction through supervision of the final high-resolution clear image without setting intermediate deblurring supervision.
[0197] First, input image I BLR The shared shallow feature extraction module is input, and the shared shallow feature F0 is extracted through a two-dimensional convolution operation with a kernel size of 3×3: F0=Conv 3×3 (I BLR ) F0 serves as the input feature for both the motion blur removal branch and the super-resolution reconstruction branch.
[0198] In the motion blur removal branch, F0 is first convolved with a 3×3 layer to obtain feature F1, which is then input into the first-level deblurring coding module to obtain the first-level coded feature E1. F1=Conv 3×3 (F0) E1=MDAKB(Restormer(F1)) Among them, Restormer is used for attention modeling and local texture enhancement; MDAKB is used for adaptive enhancement of features based on the degree and direction of motion blur; E1 represents the first-level encoded features.
[0199] Subsequently, E1 is downsampled to obtain the first low-resolution feature X1, which is then input into the second-level deblurring coding module to obtain the second-level coded feature E2. X1=Down(E1) E2=MDAKB(Restormer(X1)) Then, E2 is further downsampled to obtain feature X2, which is then input into the third-level deblurring coding module to obtain the third-level coded feature E3: X2=Down(E2) E3=MDAKB(Restormer(X2)) Meanwhile, the super-resolution reconstruction branch also uses the shared shallow feature F0 as input. First, F0 is mapped by a 3×3 convolution to obtain the initial feature S0 of the super-resolution reconstruction branch: S0=Conv 3×3 (F0) Then, S0 passes through two RSTB modules and one CFDEB module in sequence to extract the previous super-resolution features S1: S1=CFDEB(RSTB(RSTB(S0))) Among them, RSTB is used to model local texture relationships and cross-window structural connections in road images through normal window self-attention and shifted window self-attention; CFDEB is used to recover crack details and texture edges through Haar wavelet decomposition and high-frequency subband enhancement.
[0200] When the motion blur removal branch obtains the third-level coding feature E3 / Furthermore, after the super-resolution reconstruction branch obtains the previous super-resolution features S1, the two are input into the first cross-scale bidirectional feature interaction module (BWCA) for information exchange. The first BWCA interaction includes two directions.
[0201] The first direction involves information transfer from the super-resolution reconstruction branch to the motion blur removal branch. This direction uses the deblurring encoded feature E3 as the query feature source and the super-resolution feature S1 as the key-value feature source. First, a 1×1 convolution mapping is performed on E3 to obtain the query feature Q. D1 : Q D1 =Conv 1×1 (E3 / ) Then, S1 is resized to the same spatial size as E3 to obtain the scale-aligned super-resolution feature S. 1_down : S 1_down =Resize(S1,H E3 W E3 ) Among them, H E3 and W E3 They represent E3 respectively / The characteristic height and width. Then, for S... 1_down Perform a 1×1 convolution mapping to obtain the joint key-value features (KV). S1 and KV S1 The channel dimension is divided into key features K S1 Sum characteristic V S1 : KV S1 =Conv 1×1 (S 1_down ) K S1 V S1 =Split(KV S1 ) Next, Q D1 K S1 and V S1 The input window cross-attention module calculates the correlation between deblurred features and super-resolution features within a local window: A D1 =Softmax((Q D1 ·K S1 T ) / √d) ΔE3=A D1 ·V S1 Among them, A D1 K represents the cross-attention weights of the super-resolution reconstruction branch and the motion blur removal branch. S1 T K represents S1 The transpose of ; d represents the feature dimension of a single attention head; ΔE3 represents the supplementary features provided by the super-resolution reconstruction branch to the motion blur removal branch.
[0202] Subsequently, ΔE3 is mapped back to the feature dimension of the motion blur removal branch through a 1×1 convolution, and then processed by a learnable coefficient α. D1 Control its intensity, and then combine it with the original deblurred coding feature E3. / The residuals are summed to obtain the updated deblurred coding feature E3. / .
[0203] E3 / =E3 / +α D1 ·Conv 1×1 (ΔE3) Through this interaction, high-frequency details and texture information in the super-resolution reconstruction branch can be supplemented into the motion blur removal branch to help restore the edges of road cracks and local texture details.
[0204] The second direction involves information transfer from the motion blur removal branch to the super-resolution reconstruction branch. This direction uses the super-resolution feature S1 as the query feature source and the updated deblurred coding feature E3. / As a source of key-value features, S1 is first subjected to a 1×1 convolution mapping to obtain the query feature Q. S1 : Q S1 =Conv 1×1 (S1) Then, E3 / The scale-aligned deblurred feature E is obtained by adjusting the size of S1 to the same spatial dimensions. 3_up : E 3_up =Resize(E3 / H S1 W S1 ) Among them, H S1 and W S1 Let E represent the feature height and width of S1, respectively. Then, for E... 3_up Perform a 1×1 convolution mapping to obtain the joint key-value features (KV). D1 and KV D1 The channel dimension is divided into key features K D1 Sum characteristic V D1 : KV D1 =Conv 1×1 (E 3_up ) K D1 V D1 =Split(KV D1 ) Next, QS1 K D1 and V D1 The input window cross-attention module yields the supplementary feature ΔS1 of the motion blur removal branch to the super-resolution reconstruction branch: AS1=Softmax((Q S1 ·K D1 T) / √d) ΔS1=AS1·V D1 Where AS1 represents the cross-attention weight between the motion blur removal branch and the super-resolution reconstruction branch; K D1 T K represents D1 The transpose of ; ΔS1 represents the supplementary features provided by the motion blur removal branch to the super-resolution reconstruction branch.
[0205] Finally, ΔS1 is mapped back to the feature dimension of the super-resolution reconstruction branch through a 1×1 convolution, and then processed by the learnable coefficient α. S1 By controlling its intensity, and then adding the residuals of the original super-resolution feature S1, the updated super-resolution feature S1 is obtained. / : S1 / =S1+α S1 ·Conv 1×1 (ΔS1) Through this interaction, the clear structure, crack contours, and texture continuity information in the motion blur removal branch can be transferred to the super-resolution reconstruction branch, thereby assisting in the subsequent high-resolution detail reconstruction.
[0206] After completing the first BWCA interaction, the motion blur removal branch continues with E3. / As subsequent encoding input, the super-resolution reconstruction branch continues with S1 / As input for mid-term detailed reconstruction.
[0207] Subsequently, the motion blur removal branch interacts with the E3. / Downsampling is performed and input into the bottleneck layer to obtain bottleneck feature B: B=MDAKB(Restormer(Down(E3 / ))) After the bottleneck layer completes the extraction of deep blur degradation features and global structural information, the motion blur removal branch enters the decoding stage. First, the bottleneck feature B is upsampled and compared with the feature E3 from the encoding stage. / After splicing and processing by the channel compression and decoding modules, the third-level decoding feature D3 is obtained: D3=MDAKB(Restormer(Reduce(Concat(Up(B)),E3 / )))) At the same time, the super-resolution reconstruction branch will use the features S1 after the first interaction. / In the intermediate detail reconstruction stage, feature extraction is further performed using two RSTB modules and one CFDEB module to obtain the intermediate super-resolution feature S2: S2=CFDEB(RSTB(RSTB(S1 / ))) Then, S2 continues to be input into the mid-to-late stage of detail reconstruction, and after passing through the RSTB module and CFDEB module, the mid-to-late stage super-resolution features S3 are obtained: S3=CFDEB(RSTB(S2)) After the motion blur removal branch obtains the third-level decoding feature D3 and the super-resolution reconstruction branch obtains the mid-to-late-stage super-resolution feature S3, the two are input into the second cross-scale bidirectional feature interaction module (BWCA) for further information exchange. The second BWCA interaction also includes two directions.
[0208] The first direction involves information transfer from the super-resolution reconstruction branch to the motion blur removal branch. This direction uses the decoded feature D3 as the source of the query feature and the mid-to-late-stage super-resolution feature S3 as the source of the key-value feature. First, a 1×1 convolution mapping is performed on D3 to obtain the query feature Q. D2 : Q D2 =Conv 1×1 (D3) Then, S3 is resized to the same spatial size as D3 to obtain the scale-aligned super-resolution feature S. 3_down : S 3_down =Resize(S3,H D3 W D3 ) Among them, H D3 and W D3 Let S represent the feature height and width of D3, respectively. Then, for S... 3_down Perform a 1×1 convolution mapping to obtain the joint key-value features (KV). S2 and KV S2 The channel dimension is divided into key features K S2 Sum characteristic V S2 : KV S2 =Conv 1×1 (S 3_down ) K S2 V S2 =Split(KV S2 ) Next, Q D2 K S2 and V S2 The input window cross-attention module calculates supplementary information from the super-resolution features to the deblurred decoding features: A D2 =Softmax((Q D2 ·K S2 T ) / √d) ΔD3=A D2 ·V S2 Among them, A D2 K represents the cross-attention weights of the super-resolution reconstruction branch and the motion blur removal branch in the second interaction; S2 T K represents S2 The transpose of ; ΔD3 represents the supplementary feature provided by the super-resolution reconstruction branch to the motion blur removal branch.
[0209] Subsequently, ΔD3 is mapped back to the feature dimension of the motion blur removal branch through a 1×1 convolution, and then processed by a learnable coefficient α. D2 By controlling its intensity, and then adding the residual of the original decoded feature D3, the updated decoded feature D3 is obtained. / : D3 / =D3+α D2 ·Conv 1×1 (ΔD3) The second direction involves information transfer from the motion blur removal branch to the super-resolution reconstruction branch. This direction uses super-resolution feature S3 as the source of query features and the updated decoded feature D3. / As a source of key-value features. First, for S3 Perform a 1×1 convolution mapping to obtain the query feature Q. S2 : Q S2 =Conv 1×1 (S3) Then, D3 / The scale-aligned deblurred decoding feature D is obtained by adjusting the size of the feature to match that of S3 using a resize operation. 3_up : D 3_up =Resize(D3 / H S3 W S3 ) Among them, H S3 and W S3 Let D represent the feature height and width of S3, respectively. Then, for D... 3_upPerform a 1×1 convolution mapping to obtain the joint key-value features (KV). D2 and KV D2 The channel dimension is divided into key features K D2 Sum characteristic V D2 : KV D2 =Conv 1×1 (D 3_up ) K D2 V D2 =Split(KV D2 ) Next, Q S2 K D2 and V D2 The input window cross-attention module yields the supplementary feature ΔS3 of the motion blur removal branch to the super-resolution reconstruction branch: AS2=Softmax((Q S2 ·K D2 T ) / √d) ΔS3=AS2·V D2 Where AS2 represents the cross-attention weight between the motion blur removal branch and the super-resolution reconstruction branch in the second interaction; K D2 T K represents D2 The transpose of ; ΔS3 represents the supplementary features provided by the motion blur removal branch to the super-resolution reconstruction branch.
[0210] Finally, ΔS3 is mapped back to the feature dimension of the super-resolution reconstruction branch through a 1×1 convolution, and then processed by the learnable coefficient α. S2 By controlling its intensity, and then adding the residuals of the original super-resolution feature S3, the updated super-resolution feature S3 is obtained. / : S3 / =S3+α S2 ·Conv 1×1 (ΔS3) After completing the second BWCA interaction, the motion blur removal branch continues with D3. / As the input for subsequent decoding, the super-resolution reconstruction branch continues with S3. / As input for later detail reconstruction, the motion blur removal branch and the super-resolution reconstruction branch achieve complementary intermediate layer features during the network's forward propagation through two BWCA interactions, rather than simply fusing them after both branches have completely finished.
[0211] Subsequently, the motion blur removal branch continues to complete the subsequent decoding process. This applies to the interactive D3. / Upsampling is performed and concatenated with the second-level encoded feature E2 to obtain the second-level decoded feature D2: D2=MDAKB(Restormer(Reduce(Concat(Up(D3 / ),E2)))) Then, D2 is upsampled and concatenated with the first-level encoded feature E1 to obtain the first-level decoded feature D1: D1=MDAKB(Restormer(Reduce(Concat(Up(D2),E1)))) Finally, input D1 into the deblurring feature output layer to obtain the motion blur removal branch output feature F. D : F D =Conv 3×3 (D1) At the same time, the super-resolution reconstruction branch will use the features S3 after the second interaction. / Inputting the late-stage detail reconstruction stage yields the late-stage super-resolution feature S4: S4=RSTB(S3 / ) Then, input S4 into the 3×3 convolutional output layer to obtain the super-resolution reconstruction branch output feature F. S : F S = Conv 3×3 (S4) In obtaining the output feature F of the motion blur removal branch D Super-resolution reconstruction branch output features F S After sharing the shallow feature F0, the three are input into the adaptive complementary feature fusion module ACFF. This module first processes F... D F S Channel alignment is performed with F0, then the difference between the deblurred features and the super-resolution features is calculated, and adaptive fusion weights are generated based on the difference information to obtain the preliminary fusion features F. w : F D a =Conv 1×1 (F D ) F S a =Conv 1×1 (F S ) F0 a =Conv 1×1 (F0) Diff=|F D a -FS a | W=Softmax(φ w (Concat(F D a ,F S a ,Diff))) W=[w D ,w S ] Fw=w D ·F D a +w S ·F S a Subsequently, the initial fusion feature F w Alignment feature F0 with shallow layers a The input features F are concatenated to obtain the fused input features. in : Fin = Concat(Fw, F0) a ) F in The refined and supplementary features F are obtained by inputting them into the feature refinement network and the gated generation network, respectively. r And gating weight G: Fr=φr(Fin) G=Sigmoid(φg(Fin)) Finally, the gated supplementary features are residually fused with the preliminary fused features to obtain the final fused feature F. fusion : F fusion =F w +G·F r Finally, the fused feature F fusion Input image reconstruction module. The image reconstruction module reconstructs the fused features into a high-resolution, clear road image of size 512×512 through residual block, convolution, and pixel rearrangement upsampling operations. HR_pred : I HR_pred =R(F fusion ) Throughout the training process, the model input is a low-resolution blurred road image I of 256×256. BLR The final output is a high-resolution, clear road image of 512×512. HR_pred During training, a corresponding 512×512 high-resolution clear road image I was used. HR As a supervised objective, the model learns an end-to-end mapping from low-resolution blurry road images to high-resolution clear road images.
[0212] Step 7: Model application process.
[0213] After model training, in practical applications, 256×256 low-resolution blurred road images captured by road detection vehicles, drones, or vehicle-mounted cameras are input into the model. The input image first passes through a shared shallow feature extraction module to obtain F0; then it enters a motion blur removal branch and a super-resolution reconstruction branch for parallel feature extraction, and cross-scale bidirectional feature interaction is performed twice through a BWCA module in the intermediate feature layer; subsequently, adaptive complementary feature fusion is performed through an ACFF module; finally, a clear, high-resolution 512×512 road image I is output by the image reconstruction module. HR_pred The output image can be used as input for subsequent road crack detection, defect identification, road image segmentation, or pavement technical condition evaluation.
[0214] Step 8: Model evaluation and comparative analysis.
[0215] To verify the image restoration effect of the model proposed in this embodiment, Bicubic, EDSR, Stripformer, Restormer, GFN, and JMD-SR were selected as comparison methods, and the restoration results of different methods were evaluated under the same test set conditions. Specifically, the low-resolution blurred road image IBLR was input into different comparison methods to obtain the corresponding high-resolution restored images, and the output results of each method were compared with the high-resolution clear road image IHR.
[0216] In this embodiment, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Indicator (SSIM) are used as model evaluation metrics. PSNR is used to evaluate the pixel reconstruction error between the model's output image and the real high-resolution clear image; a higher PSNR value indicates a smaller image reconstruction error. SSIM is used to evaluate the structural similarity between the model's output image and the real high-resolution clear image; a higher SSIM value indicates better image structure preservation. Table 1 shows a comparison of the image restoration effects of different methods.
[0217] Table 1: Comparison of Image Restoration Results of Different Methods
[0218] As shown in Table 1, the method of the present invention outperforms other comparative methods in both PSNR and SSIM indices, indicating that the method of the present invention can achieve better image restoration results. Compared with methods such as Bicubic, EDSR, Stripformer, Restormer, GFN, and JMD-SR, the method of the present invention can more effectively restore the edges of road cracks, pavement texture, and defect contours.
[0219] Example 2 The purpose of this embodiment is to provide a motion blur removal and super-resolution reconstruction system for road images, including: The feature extraction unit is configured to: extract features from the acquired road image to obtain shared shallow features; The deblurring unit is configured to input shared shallow features into the motion blur removal branch, and extract and restore structural information in the road image step by step through the encoder, bottleneck layer and decoder. The super-resolution reconstruction unit is configured to: input shared shallow features into the super-resolution reconstruction branch, and use multiple residual SwinTransformer modules and crack frequency detail enhancement modules to extract high-frequency texture information and spatial detail information of the road image from the shared shallow features; wherein, the crack frequency detail enhancement module is based on wavelet decomposition, and uses low-frequency sub-band as structural guidance information to perform window cross attention calculation on high-frequency sub-band to enhance the reconstruction capability of crack details and texture edges; The interaction unit is configured to: utilize the cross-scale bidirectional feature interaction module to calculate the correlation between features between the motion blur removal branch and the super-resolution reconstruction branch through window cross attention, perform bidirectional interaction on the features of the motion blur removal branch and the super-resolution reconstruction branch, and obtain the deblurred features and the super-resolution features after interaction. The reconstruction unit is configured to fuse and reconstruct shared shallow features, interactive deblurred features, and interactive super-resolution features to obtain the reconstructed road image.
[0220] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0221] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0222] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0223] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0224] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0225] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0226] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0227] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0228] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0229] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0230] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for motion blur removal and super-resolution reconstruction of road images, characterized in that, include: Feature extraction is performed on the acquired road images to obtain shared shallow features; The shared shallow features are input into the motion blur removal branch, and the structural information in the road image is extracted and restored step by step through the encoder, bottleneck layer and decoder. The shared shallow features are input into the super-resolution reconstruction branch. Multiple residual SwinTransformer modules and crack frequency detail enhancement modules are used to extract high-frequency texture information and spatial detail information of the road image from the shared shallow features. The crack frequency detail enhancement module is based on wavelet decomposition and uses the low-frequency sub-band as structural guidance information to perform window cross attention calculation on the high-frequency sub-band, thereby enhancing the reconstruction capability of crack details and texture edges. By utilizing a cross-scale bidirectional feature interaction module, the correlation between features in the motion blur removal branch and the super-resolution reconstruction branch is calculated through window cross attention. The features of the motion blur removal branch and the super-resolution reconstruction branch are bidirectionally interacted to obtain the deblurred features and the super-resolution features after interaction. The shared shallow features, the deblurred features after interaction, and the super-resolution features after interaction are fused and reconstructed to obtain the reconstructed road image.
2. The method for motion blur removal and super-resolution reconstruction of road images as described in claim 1, characterized in that, The shared shallow features are input into the motion blur removal branch, where structural information in the road image is extracted and restored step by step through the encoder, bottleneck layer, and decoder. Specifically: The first feature is obtained by extracting features from the shared shallow features through convolution operations. The encoder extracts the fuzzy degradation features at different scales of the first feature to obtain the first-level coded features, the second-level coded features, and the third-level coded features; The third-level coding features are bidirectionally interacted across scales with the early super-resolution features of the super-resolution reconstruction branch to obtain the interacted third-level coding features. The third-level encoded features after interaction are downsampled and input into the bottleneck layer to obtain the bottleneck features; The bottleneck features after upsampling and the third-level encoded features after interaction are concatenated by the decoder. The concatenated features are then compressed and decoded to obtain the third-level decoded features. The upsampled third-level decoded features and the second-level encoded features are concatenated, and the concatenated features are compressed and decoded to obtain the second-level decoded features. The upsampled second-level decoded features and the first-level encoded features are concatenated. The concatenated features are then compressed and decoded to obtain the first-level decoded features. A convolution operation is then performed on the first-level decoded features to obtain the deblurred features.
3. The method for motion blur removal and super-resolution reconstruction of road images as described in claim 2, characterized in that, By extracting fuzzy degradation features at different scales of the first feature using an encoder, we obtain the first-level encoded features, the second-level encoded features, and the third-level encoded features, specifically: The first feature is obtained by performing feature extraction through convolution operation on the first feature; The first feature is input into the first-level deblurring coding module to extract shallow texture, local edges and blurring degradation information in the road image, thus obtaining the first-level coding feature; The downsampled first-level coding features are input into the second-level deblurring coding module to obtain the second-level coding features; The downsampled second-level encoded features are input into the third-level deblurring encoding module to obtain the third-level encoded features. The first, second, and third-level deblurring encoding modules all include an attention-based feature extraction unit and a motion-aware orientation adaptive convolution module. The attention-based feature extraction unit performs global information modeling and local texture enhancement on the input features. Then, the motion-aware orientation adaptive convolution module adaptively enhances features of different scales and directions according to the degree and direction of motion blur in the road image.
4. The method for motion blur removal and super-resolution reconstruction of road images as described in claim 3, characterized in that, The input features are modeled globally and enhanced locally using an attention-based feature extraction unit, specifically as follows: The input features are enhanced by multi-head transpose attention to obtain attention-enhanced features. Attention-enhanced features are input into a gated deep convolutional feedforward network to further enhance local texture and nonlinear representation capabilities. The output of the gated deep convolutional feedforward network is then fused with the attention-enhanced features to obtain the output features of the attention-based feature extraction unit.
5. The method for motion blur removal and super-resolution reconstruction of road images as described in claim 3 or 4, characterized in that, Then, using a motion-aware orientation-adaptive convolution module, features at different scales and in different directions are adaptively enhanced based on the degree and direction of motion blur in the road image. Specifically: The output features of the attention-based feature extraction unit are analyzed using a fuzzy intensity estimator to generate weights for convolutional branches at different scales; The output features of the attention-based feature extraction unit are processed using separable convolutional branches of different depths to obtain multi-scale features under different receptive fields. By using the weights of convolutional branches at different scales, multi-scale features under different receptive fields are weighted and fused to obtain multi-scale fuzzy perception features. The output features of the attention-based feature extraction unit are used to extract fuzzy structures and texture responses in different directions, resulting in multiple directional related features. Multiple directionally related features are concatenated, and an adaptive weight for different directional branches is obtained using a directional weight generator. By using adaptive weights for branches in different directions, multiple directionally related features are weighted and fused to obtain directional-aware features; The multi-scale fuzzy perception features, direction perception features, and the output features of the attention-based feature extraction unit are fused to obtain the output features of the motion perception direction adaptive convolution module.
6. The method for motion blur removal and super-resolution reconstruction of road images as described in claim 1, characterized in that, The shared shallow features are input into the super-resolution reconstruction branch. Multiple residual SwinTransformer modules and a crack frequency detail enhancement module are used to extract high-frequency texture and spatial detail information from the shared shallow features. Specifically: The first residual SwinTransformer module and the crack frequency detail enhancement module are used to extract features from the shared shallow features to obtain the early super-resolution features; The first cross-scale bidirectional feature interaction is performed between the previous super-resolution features and the third-level encoded features in the motion blur removal branch to obtain the previous super-resolution features after interaction. The second residual SwinTransformer module and the crack frequency detail enhancement module are used to extract features from the early super-resolution features after interaction to obtain the mid-term super-resolution features. The mid-term super-resolution features are then reconstructed using the residual SwinTransformer module and the crack frequency detail enhancement module to obtain the mid-to-late-term super-resolution features. The mid-to-late stage super-resolution features are combined with the third-level decoding features in the motion blur removal branch to perform a second cross-scale bidirectional feature interaction, resulting in the mid-to-late stage super-resolution features after the interaction. The super-resolution features after interaction are reconstructed in the later stages using the residual SwinTransformer module to obtain the output features of the super-resolution reconstruction branch.
7. The method for motion blur removal and super-resolution reconstruction of road images as described in claim 1 or 6, characterized in that, The processing procedure of the crack frequency detail enhancement module for its input features is as follows: Wavelet decomposition is performed on the input features of the crack frequency detail enhancement module to obtain low-frequency sub-bands and multiple high-frequency sub-bands; Using low-frequency subbands as structural guidance information, window cross-attention enhancement is performed on multiple high-frequency subbands to obtain high-frequency enhanced features; The enhanced high-frequency features are superimposed on the corresponding high-frequency sub-bands to obtain the enhanced high-frequency sub-bands; The frequency enhancement features are reconstructed by performing inverse wavelet transform on the low-frequency subband and the enhanced high-frequency subband. The frequency enhancement features are convolved and mapped, and then added to the input feature residuals of the crack frequency detail enhancement module to obtain the output features of the crack frequency detail enhancement module.
8. A system for motion blur removal and super-resolution reconstruction of road images, characterized in that, include: The feature extraction unit is configured to: extract features from the acquired road image to obtain shared shallow features; The deblurring unit is configured to input shared shallow features into the motion blur removal branch, and extract and restore structural information in the road image step by step through the encoder, bottleneck layer and decoder. The super-resolution reconstruction unit is configured to: input shared shallow features into the super-resolution reconstruction branch, and use multiple residual SwinTransformer modules and crack frequency detail enhancement modules to extract high-frequency texture information and spatial detail information of the road image from the shared shallow features; wherein, the crack frequency detail enhancement module is based on wavelet decomposition, and uses low-frequency sub-band as structural guidance information to perform window cross attention calculation on high-frequency sub-band to enhance the reconstruction capability of crack details and texture edges; The interaction unit is configured to: utilize the cross-scale bidirectional feature interaction module to calculate the correlation between features between the motion blur removal branch and the super-resolution reconstruction branch through window cross attention, perform bidirectional interaction on the features of the motion blur removal branch and the super-resolution reconstruction branch, and obtain the deblurred features and the super-resolution features after interaction. The reconstruction unit is configured to fuse and reconstruct shared shallow features, interactive deblurred features, and interactive super-resolution features to obtain the reconstructed road image.
9. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-7.
Citation Information
Patent Citations
An image super-resolution and non-uniform blur removal method based on fusion network
CN109345449A
Vehicle-mounted dual-spectrum deblurring super-resolution multi-task processing method and system
CN119540061A