A super-resolution processing method and system for a drone to capture road surface images

Through the highly efficient KAN reference image super-resolution ESR-KAN network model, the problems of low resolution and low image processing efficiency of drone shooting road surfaces are solved, and the image detail clarity and quality are significantly improved, and are suitable for devices with resource limitations.

CN119648527BActive Publication Date: 2025-06-13EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411620786.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-06-13
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

When the prior art improves the resolution of road image shooting by drones, the equipment is costly, cumbersome and may introduce errors. Traditional image processing methods may introduce artifacts or lead to loss of feature information. Convolutional neural networks have problems such as loss of fine-grained feature information and blurred feature when processing wide-angle road images captured by drones.

Method used

Using the highly efficient KAN reference image super-resolution ESR-KAN network model, the KAN fusion module, reconstruction module and M serial KAN-channel attention, combined with global average pooling and continuous KAN layers, the importance weight of each channel is learned, the fine regulation of the feature map is realized, and the spatial details of the image are restored through the 3×3 convolution layer and the ReLU activation function.

Benefits of technology

It significantly improves the detail clarity and overall quality of the road image taken by the drone, reduces the amount of calculation parameters, improves the interpretability and sparseness of the model, and is suitable for equipment with resource-constrained, and has a wide range of application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648527B_ABST
    Figure CN119648527B_ABST
Patent Text Reader

Abstract

The present invention relates to a super-resolution processing method and system for a drone to capture road surface images. The method includes the following steps: capturing road surface targets through aerial photography by the drone to obtain road surface images; at the same time, using a high-pixel camera to capture the same targets as reference images, and the reference images of all targets form a reference image set; constructing an efficient KAN reference-image super-resolution ESR-KAN network model: the ESR-KAN network model includes a KAN fusion module, a reconstruction module, and M serial KAN-channel attentions. The reconstruction module is used to restore the spatial shape of the image, and the result after being processed by the reconstruction module is subjected to residual connection with the result of the reference image after upsampling processing to obtain the final high-resolution image Y. In the super-resolution processing of the drone capturing road surface images, the present invention not only improves the image quality, enhances the usability of the image and the accuracy of analysis, but also provides a new solution for image processing of the drone in various application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image super-resolution processing, and particularly to a method and system for super-resolution processing of road surface images captured by an unmanned aerial vehicle (UAV). Background Art

[0002] When a UAV performs tasks such as aerial reconnaissance, terrain mapping, and traffic monitoring, the quality of the images captured by the camera it carries, especially the resolution, is crucial for obtaining accurate information. During the process of the UAV inspecting and detecting the road surface, the image resolution determines the ability of the UAV camera to capture road surface details. High-resolution images can provide more detailed information, which helps in accurately identifying, tracking, and analyzing road surface defect targets.

[0003] To improve the resolution of road surface images captured by a UAV, existing methods mainly use technologies such as high-resolution cameras and multi-frame image synthesis. For example, some studies have suggested using high-resolution single-lens reflex cameras with more than 20 million pixels and selecting the maximum resolution mode during shooting to reduce noise and blur in road surface images. However, directly using a high-resolution camera requires a significant increase in equipment costs. Multi-frame image synthesis and other technologies require an increase in the number of photos taken, and the synthesis takes a lot of time, is cumbersome to operate, and may introduce errors.

[0004] In recent years, image super-resolution technology has been widely applied to the field of improving image resolution. Through deep learning models, high-resolution images are reconstructed from low-resolution images, thereby restoring more detailed features (wide-angle images, macroscopic cracks visible to the naked eye on the road surface, shooting methods). However, these methods have some defects and deficiencies. Traditional image processing methods, such as filtering and sharpening, may introduce artifacts or result in the loss of important feature information. In addition, due to the low resolution and small targets of road surface images captured by a UAV, the convolutional neural network structure is affected by the loss of fine-grained feature information and feature blurring, resulting in a significant reduction in performance. The super-resolution methods for some microscopic images have insufficient recognition accuracy for such wide-angle road surface images captured by UAV aerial photography.

[0005] Today, with the rapid development of artificial intelligence technology, image super-resolution technology has become one of the key areas for improving image quality. Through deep learning models, this technology can effectively reconstruct high-resolution images from low-resolution images, thereby restoring rich detail features. The processing technology based on image super-resolution provides a new solution to address the limitations of traditional image processing methods. Compared with traditional methods, these advanced image processing technologies can improve the resolution of images captured by drones with higher efficiency and better quality, gradually becoming a new trend to replace traditional image processing means. Through these innovative technologies, when drones perform road surface shooting tasks, they will be able to obtain clearer and more accurate image data, greatly enhancing the usability of the images and the accuracy of analysis. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the purpose of the present invention is to provide a super-resolution method and system for road surface images captured by drones, aiming to solve the above obvious defects existing in the prior art.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] In the first aspect, the present invention provides a super-resolution processing method for road surface images captured by drones, and the method includes the following:

[0009] The road surface target is photographed by drone aerial photography to obtain a road surface image; at the same time, a high-pixel camera is used to photograph the same target as a reference image, and the reference images of all targets form a reference image set. The shooting angles of the reference images and the road surface images can be different;

[0010] Construct an efficient KAN reference-image super-resolution ESR-KAN network model:

[0011] The ESR-KAN network model includes a KAN fusion module, a reconstruction module, and M serial KAN-channel attentions. The KAN fusion module includes a first layer composed of two parallel KANs and a second layer composed of one KAN. The structures of the three KANs are the same. The two KANs in the first layer are respectively used to obtain the feature map information of the input image X and the reference image X R After the processing results of the two KANs in the first layer are concatenated, they enter the KAN in the second layer to obtain the output feature map O 0 ;

[0012] Each KAN-channel attention includes a serial structure composed of a global average pooling GAP and two consecutive KANs, and a multiplication operation and a residual connection that combine the input; in the KAN-channel attention, for the input feature map O iApply global average pooling (GAP) to compress channel features, use two consecutive KANs to learn the importance weights of each channel, generate a new vector from the importance weights of each channel, and use the new vector to process the input feature map O i to weight the channels, obtaining a channel weight vector; multiply the channel weight vector by the input feature map O i , and then add it to the input feature map O i to get the output O i+1 of this KAN-channel attention, and use it as the input for the next KAN-channel attention, and so on to get the output O M of the last KAN-channel attention, where i = 0, 1, …, M - 1;

[0013] The reconstruction module is used to restore the spatial shape of the image. The result after being processed by the reconstruction module is subjected to residual connection with the result of the reference image after upsampling to obtain the final high-resolution image Y.

[0014] Furthermore, the reconstruction module includes a 3×3 convolutional layer, a ReLU activation function, and a 3×3 convolutional layer connected in sequence.

[0015] Furthermore, M = 4 - 6.

[0016] Furthermore, the objective loss function l total during the training of the ESR-KAN network model is: l total = l pred + l sparse

[0017] l pred = |Y - X| 1

[0018]

[0019] where l pred is the prediction loss function; l sparse is the KAN spline sparsity loss function; λ is a parameter that controls the overall regularization amplitude, and μ 1 and μ 2 are relative calculation parameters, both set to 1; φ l represents the activation function of the l-th layer in the last KAN of the last KAN-channel attention, and |φ l | 1 represents the L 1 norm of the activation function, S(φ l ) is the entropy regularization term, L is the number of layers of KAN; X represents the original input image, and Y represents the high-resolution image predicted by the network structure.

[0020] Second aspect, the present invention provides a super-resolution processing system for an unmanned aerial vehicle (UAV) to capture road surface images. The system includes:

[0021] An unmanned aerial vehicle, configured to obtain road surface images of road targets through aerial photography;

[0022] A high-pixel camera, configured to obtain a high-resolution reference image from a certain perspective of the target;

[0023] An image preprocessing module, configured to perform data enhancement processing on the road surface images and the reference images;

[0024] A database, configured to store images and parameters;

[0025] The ESR-KAN network model, configured to perform super-resolution reconstruction on the road surface images.

[0026] Furthermore, the system further includes a network interface, a display device, and an image processing unit. The display device is configured to view the processing results and the image quality evaluation results in real time.

[0027] Furthermore, the system further includes a road defect semantic segmentation model, and the super-resolution image reconstructed by the ESR-KAN network model is used for road defect recognition.

[0028] Furthermore, the super-resolution image is used for monitoring forest cover changes, urban expansion, and agricultural land changes, road surface maintenance, and road surface detection.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] The efficient KAN with reference image super-resolution processing technology ESR-KAN (Efficient Super-Resolution with Reference Image via Kolmogorov-Arnold) mentioned in the present invention shows significant advantages in multiple aspects:

[0031] The three core components of the efficient KAN with reference image super-resolution ESR-KAN network model of the present invention, namely, the KAN fusion module, the KAN-channel attention, and the reconstruction module, work together, enabling the model to significantly improve the detail clarity and overall quality of the low-resolution images of aerial road targets.

[0032] In the present invention, the KAN fusion module serves as the primary link of the model. The KAN fusion module ingeniously utilizes multiple KAN structures distributed in a quasi-triangular shape and adopts a parallel processing mechanism to simultaneously perform in-depth feature extraction on the input low-resolution road surface image and its corresponding high-resolution reference image. This process not only enhances the model's ability to capture image features but also provides richer and more comprehensive information for the subsequent processing flow through the merging of feature maps. In the present invention, the number of reference images is small, and it is only necessary to ensure that each target has a corresponding reference image, with no requirements for shooting angles, etc.

[0033] KAN-channel attention precisely learns the importance weights of each channel through global average pooling and successive KAN layers, achieving refined regulation of the feature map. This not only enhances the model's discriminative ability but also provides a solid foundation for the final image reconstruction. The refined feature regulation ability enables the model to highlight important features and suppress unimportant information, further improving the quality of the reconstructed image.

[0034] The reconstruction module undertakes the important task of converting the optimized feature map into a high-resolution image. By using a combination of 3×3 convolutional layers and ReLU activation functions, it precisely restores the spatial details of the image. At the same time, through the residual connection mechanism, it ensures the consistency between the reconstructed image and the original high-resolution reference image. This process not only improves the accuracy of the reconstructed image but also enhances the clarity and realism of the image, meeting the requirements for high-quality output in the super-resolution task.

[0035] Integrating these three major modules, the ESR-KAN network model demonstrates excellent performance in the field of image super-resolution. This model can not only effectively process low-resolution images but also make full use of the information in high-resolution reference images to achieve high-precision image reconstruction. In addition, considering computational efficiency and parameter optimization, it reduces the scale of computational parameters in the network structure, enabling it to operate efficiently on resource-constrained devices and having broad application prospects. This innovative network structure can accurately capture and reconstruct image details while keeping the model lightweight, providing a new solution for the field of image super-resolution.

[0036] In summary, in the super-resolution processing of road surface images captured by drones, the ESR-KAN network model of the present invention not only improves the image quality, enhances the usability of the image and the accuracy of analysis, providing a new solution for image processing in various application scenarios of drones, but also plays an important role in cost control and improving the accuracy of subsequent computer vision tasks, which will promote the further development and application of drone technology in various fields. The present invention significantly improves the resolution and quality of drone-aerial road surface images, providing more accurate and clear image support for the field of drone vision applications. Brief Description of the Drawings

[0037] Figure 1 This is a schematic diagram of the structure of the ESR-KAN network model in the present invention.

[0038] Figure 2 This is the flowchart of the super-resolution of the actual application road surface of ESR-KAN. Detailed Embodiments

[0039] In order to more clearly describe the technical problems, technical solutions and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scope of the present invention and should not be construed as a limitation of the present invention.

[0040] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0041] The super-resolution processing method for the road surface images captured by the unmanned aerial vehicle of the present invention uses an efficient KAN reference-image super-resolution ESR-KAN (Efficient Super-Resolution with Reference Image via Kolmogorov-Arnold) network model, and includes the following steps:

[0042] Step 1: Obtain the data set

[0043] The road surface target is photographed by the unmanned aerial vehicle to obtain road surface images, and at the same time, the same target is photographed by a high-pixel camera as the reference image. The reference images of all targets form a reference image set, and the obtained road surface image set and reference image set of the unmanned aerial vehicle aerial photography of the road surface are stored in the corresponding files respectively. The reference image can be an image of the same target from a different shooting angle from the road surface image. For multiple road surface images of the same target, one reference image can be corresponding.

[0044] In order to further improve the generalization ability and robustness of the model, data augmentation technology is applied. The road surface images of the road surface targets photographed by the unmanned aerial vehicle in the training set are rotated, scaled, flipped, the brightness and contrast are adjusted, and random noise is added, so as to generate additional variants, expand the diversity of the training data set, enhance the generalization ability of the model, and ensure that the model can better handle various actual situations.

[0045] Step 2: Construct an efficient KAN reference image super-resolution ESR-KAN network model

[0046] The efficient KAN reference image super-resolution ESR-KAN network model has a structural diagram as Figure 1 shown, and is mainly composed of three major modules: a KAN fusion module, a reconstruction module, and M serial KAN-channel attentions.

[0047] The KAN fusion module includes a first layer composed of two parallel KANs and a second layer composed of one KAN. The structures of the three KANs are the same, and the three identical KANs are distributed in a similar triangle shape. The two KANs in the first layer are set in parallel, and are respectively used to obtain the feature map information of the input image X and the reference image X R wherein, X, (H, W represent the height and width of the image, and C represents the number of channels of the image). In this process, the processing processes of the two parallel KANs can be expressed by the formula:

[0048]

[0049] After that, the obtained X 0 and are concatenated, and then input into the second-layer KAN for processing. Using O 0 to represent the output result of the KAN fusion module, it can be expressed by the following formula:

[0050]

[0051] wherein, D is a hyperparameter representation, referring to the channel dimension that the feature map O 0 should maintain. In this embodiment, D is set to 256.

[0052] The KAN-channel attention includes a serial structure composed of a global average pooling GAP and two consecutive KANs, as well as a multiplication operation and a residual connection that combine the input. In the KAN-channel attention, the global average pooling GAP is applied to the input feature map O i for channel feature compression, and two consecutive KANs are used to learn the importance weights of each channel, generating a new vector for the importance weights of each channel. Using the new vector to weight the channels of the input feature map O i to obtain a channel weight vector; multiplying the channel weight vector by the input feature map O i and then adding it to the input feature map O i to obtain the output O i+1 of this KAN-channel attention., and used as the input of the next KAN-channel attention, and so on to obtain the output O of the last KAN-channel attention M , i = 0, 1, …, M - 1.

[0053] The structure of the reconstruction module can be described as: a 3×3 convolutional layer, a ReLU activation function, and a 3×3 convolutional layer at the end. These three components are serially combined to form the reconstruction module. The reconstruction module is used to restore the spatial shape of the image. The result processed by the reconstruction module and the upsampled result of the reference image are added residually to obtain the final output high-resolution image Y. The whole process can be expressed by the following formula:

[0054] Y = Conv(ReLU(Conv(O M )))+UP(X R )

[0055] The overall process of the ESR-KAN network model can be summarized as: the input is the training set image X and the reference image X of the same type R . First, the input image X is used to obtain the feature map information through a KAN in the first layer of the KAN fusion module; at the same time, the reference image X R is upsampled once to get UP(X R ), and then UP(X R ) is input into a KAN in the first layer of the KAN fusion module to obtain the feature map information. After the operation processes of the two KANs in the first layer are completed, the two obtained feature map information are concatenated, and then the output feature map O 0 is processed by the KAN in the second layer of the KAN fusion module.

[0056] Next, M identical serial processing processes of KAN-channel attention are carried out. The feature map O 0 is used as the input of the first KAN-channel attention. First, the feature map O 0 uses global average pooling GAP (Global Average Pooling) to compress the feature of each channel of the input feature map O 0 , then two consecutive KANs are used to learn the importance weight of each channel, and then a weight vector is generated by the learned weights. The channels of the input feature map are weighted by the weight vector to obtain the channel weight vector.

[0057] Finally, the channel weight vector and the input feature map O 0 are multiplied channel by channel, and the result is added to the input feature map O 0 to obtain the input O of the next KAN-channel attention 1Through the operations of the same process, the final output feature map O is obtained after M consecutive serial KAN-channel attentions. M 。

[0058] Finally, using the feature map O M as the input, image reconstruction is performed through a reconstruction module composed of a 3×3 convolutional layer, an activation function ReLU, and a 3×3 convolutional layer to obtain the final high-resolution image Y.

[0059] KAN (Kolmogorov-Arnold Networks) in the present invention is a new type of deep learning model, whose design inspiration comes from the Kolmogorov-Arnold representation theorem. The core feature is that it places the activation functions at the edges (weights) of the network, and these activation functions are learnable, usually parameterized using B-splines. KAN includes the SiLU activation function, linear layers, and a residual structure based on B-spline functions.

[0060] The processing process of KAN can be defined as:

[0061]

[0062] Among them, ω 1 and ω 2 represent weights; c i represents the basis function coefficient; G represents the grid parameter, with a default value of 5; k is a set specified parameter, with a default value of 3; specifically, G is the number of intervals of the original grid, and k is the order of the B-spline, indicating how many node intervals each B-spline function is non-zero on.

[0063] The B-spline function is composed of multiple piecewise polynomials spliced together, and each piecewise polynomial is defined by control points (grid points). Regarding the characteristics of the B-spline function, given a specified domain [t 0 , t G , when using a k-order B-spline function to approximate a one-dimensional function, a coarse grid with G intervals is extended to a fine grid with G + k intervals. A recursive definition is given for the piecewise polynomial during the process of extending to G + k: For the case of k > 0:

[0064]

[0065] For k > 0:

[0066]

[0067] The target loss function of the network model of the present invention combines the prediction loss function l pred and the KAN spline sparsity loss function l sparse, which provides the model with more refined adjustment capabilities for the target loss function l total is defined as follows:

[0068] l pred = |Y - X| 1

[0069] l total = l pred + l sparse

[0070] where X represents the original input image and Y represents the predicted high-resolution image;

[0071] In the present invention, the KAN spline sparse loss function refers to the spline sparse loss of the last KAN of the last KAN-channel attention. The KAN spline sparse loss function l sparse is defined as:

[0072]

[0073] where λ is a parameter controlling the overall regularization amplitude, and μ 1 and μ 2 are relative calculation parameters, both set to 1. φ l represents the activation function of the l-th layer in the last KAN, L is the number of layers of the last KAN, and |φ l | 1 represents the L 1 norm of the activation function, and S(φ l ) is the entropy regularization term.

[0074] The present invention can reduce the number of parameters while maintaining high accuracy, improve the interpretability and sparsity of the network model, and exhibit higher model performance.

[0075] Step 3: Training the efficient KAN reference-image super-resolution ESR-KAN network model

[0076] The training process of the efficient KAN reference-image super-resolution ESR-KAN network model is as follows:

[0077] At the beginning of training, first complete the parameter setting for network initialization: the epoch of the training network is set to 500, the Adam optimizer is used, and the initial learning rate of Adam is set to 5e-5 and 1e-6.

[0078] The training set is sourced from road surface images obtained by drone aerial photography and reference images of the same target taken with a high-pixel camera. These images are processed using data augmentation techniques, including rotation, scaling, flipping, brightness and contrast adjustment, and adding random noise, to enhance the generalization ability and robustness of the model.

[0079] During the training process, the input images of the training set and the reference images are input into the ESR-KAN network model, and these images are read according to the storage path of the training set.

[0080] When the overall network objective loss function l total no longer shows a significant decrease (error ±1e-5), the network training is considered to tend to be stable, and the network training process is completed.

[0081] The ESR-KAN network model has achieved significant cost reduction and efficiency improvement in the field of UAV aerial photography road surface image processing. By improving the processing efficiency and reducing the processing cost, it provides a more economical solution for scientific research and industrial applications, greatly promoting the popularization and application of UAV aerial photography road surface image processing technology.

[0082] Embodiment 1

[0083] The super-resolution processing method for UAV-captured road surface images of the present invention uses an efficient KAN super-resolution with reference image ESR-KAN (Efficient Super-Resolution with Reference Image via Kolmogorov-Arnold) network model to achieve super-resolution processing of UAV aerial photography road surface images through the following steps:

[0084] 1. Data collection and preparation stage

[0085] 1.1 Data collection

[0086] Equipment: Use a UAV to collect target aerial photography road surface images. A high-pixel camera device is used to take high-resolution (4000*3000 pixels) reference images.

[0087] Operation: Take a set of aerial photography road surface images for each sample, and a high-pixel camera takes reference images to ensure that different perspectives and multiple groups of target images are covered.

[0088] For example, in this embodiment, the acquisition target is the area where the disease is located on the road surface. Lock the target area, conduct aerial photography of the target area from different angles, and at the same time use a high-pixel camera to take reference images at any angle. This target area is a macroscopically visible disease on the road surface, and the UAV aerial photography obtains its wide-angle image.

[0089] During the training process of the network model of the present invention, the total number of road surface images of UAV aerial photography road surface targets obtained is 4800. According to the ratio of 7:3, these road surface images are divided into a training set and a test set; the number of reference image sets is 200, and reference images of the same type of target are used for training during training.

[0090] 1.2 Data Processing

[0091] Dataset Acquisition: The collected image data is preliminarily processed, including removing obvious noise and outliers in the image, as well as cropping and resizing the image to ensure the consistency and standardization of the image dataset. The collected images are divided into a training set and a test set in a 7:3 ratio.

[0092] Data Augmentation: Data augmentation operations such as rotation, scaling, flipping, adjusting brightness and contrast are performed on the training set to improve the robustness and generalization ability of the model.

[0093] 2. Network Model Training Phase

[0094] 2.1 Model Construction

[0095] Network Structure: The ESR-KAN network model is constructed, including a KAN fusion module, 5 serial KAN-channel attentions, and a reconstruction module. M = 5.

[0096] The KAN fusion module consists of a first layer composed of two parallel KANs and a second layer composed of one KAN. The structures of the three KANs are the same. The two KANs in the first layer are respectively used to obtain the feature map information of the input image X and the reference image X R ; After the processing results of the two KANs in the first layer are concatenated, they enter the KAN in the second layer to obtain the output feature map O 0 ;

[0097] The processing process of each KAN-channel attention can be mainly divided into two stages. In the first stage, global average pooling GAP is applied to the input feature map O i (i = 0, 1,..., M - 1; O i represents the input of the i-th KAN-channel attention) for channel feature compression, which means that for each channel in the feature map, a value will be obtained to represent its global importance. In the second stage, two consecutive KANs are used to learn the importance weights of each channel. Then, these weights will be used to generate a new vector, which weights the channels of the input feature map O i to obtain a channel weight vector. Finally, the channel weight vector is multiplied by the input feature map O i to accurately adjust the global synthesis ratio of local pixels, and then added to the input feature map O i to obtain the output O i+1 of this KAN-channel attention, and used as the input of the next KAN-channel attention, and so on to obtain the output O M .

[0098] The process of obtaining the final high-resolution image Y after being processed by the reconstruction module is expressed as:

[0099] Y = Conv(ReLU(Conv(O 5 )))+UP(X R )

[0100] Among them, Conv represents a 3×3 convolutional layer; UP represents upsampling; ReLU represents the ReLU activation function.

[0101] The network first extracts the feature map information through the KAN fusion module, and then undergoes serial processing of KAN-channel attention, strengthening the information of the feature map. Each KAN-channel attention enhances the importance of the features through global average pooling and weight learning. Finally, the reconstruction module restores the spatial details of the image through a 3×3 convolutional layer and the ReLU activation function, generating the desired high-resolution output image Y.

[0102] 2.3 Training of the ESR-KAN Network Model

[0103] Training settings: The network model is trained using a deep learning framework. During the training process, the Adam optimizer is adopted, with the initial learning rate set to 5e-5 and 1e-6, and the learning rate and other hyperparameters are adjusted according to the needs of model training.

[0104] Loss function: Combining the prediction loss function l pred and the KAN spline sparse loss function l sparse , calculate the target loss function l total , and optimize the model parameters to improve the image quality.

[0105] When the overall network target loss function l total no longer shows an obvious decrease (error ±1e-5), it is regarded that the network training tends to be stable and the network training process is completed. When conducting network testing, input the prepared test set images, and no additional reference images are required. Import the network weights after training on the training set. For the road surface images taken by drones, when the peak signal-to-noise ratio PSNR and the structural similarity SSIM of the road surface images taken by drones reach 44.89 (error ±0.03) and 0.991 (error ±0.02) respectively, the entire network training ends.

[0106] 3. Super-Resolution Processing and Analysis Stage

[0107] 3.1 Model Deployment

[0108] Deployment: Deploy the trained ESR-KAN network model to the road surface image processing system for real-time processing of the road surface images taken by drones.

[0109] Processing: For the input image, use the trained ESR-KAN network model to generate the corresponding high-resolution image for detail restoration and enhancement.

[0110] 3.2 Result Analysis

[0111] Quality Assessment: Evaluate the quality of the super-resolution image by calculating the PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) of the image.

[0112] 4. Application and Optimization Phase

[0113] 4.1 Image Application

[0114] Scientific Research: The processed high-resolution image is used for scientific research to improve the accuracy of target observation.

[0115] Practical Application: Improve the clarity and details of the road surface images taken by drones to support more accurate target observation and analysis.

[0116] 4.2 System Optimization

[0117] Performance Improvement: According to the feedback in practical applications, optimize the network model and processing algorithm to improve the processing efficiency and image quality of the system.

[0118] User Feedback: Collect user feedback on the image processing effect and further adjust and improve the system settings.

[0119] Hardware Configuration:

[0120] Processor: Responsible for model training and image processing, and requires high-performance computing capabilities suitable for deep learning tasks.

[0121] Memory: Used to store the training dataset, model weights, and processed images, and requires high speed and high reliability.

[0122] Graphics Processing Unit (GPU): Accelerate the training and inference processes of deep learning models to improve computing efficiency.

[0123] Network Interface: Used for data transmission and model update to ensure high speed and stability.

[0124] Display Device: Used to view the processing results and image quality assessment results in real time.

[0125] On the UAV aerial road surface image dataset with the same target, the ESR-KAN network model of the present invention was compared with existing methods. The comparison results show that: compared with existing excellent super-resolution methods such as MASA-SR (Masa-sr: Matching acceleration and spatial adaptation for reference-based image super-resolution), RRSGAN (RRSGAN: Reference-based super-resolution for remote sensing image), and DATSR (Reference-based image super-resolution with deformable attention transformer), the super-resolution method of the present invention has lower computational complexity and better performance. The comparison results obtained by different network models after training for UAV aerial road surface image super-resolution processing are shown in Table 1.

[0126] Table 1 Comparison of indicators of different network models

[0127]

[0128] The results shown in Table 1 indicate that in the super-resolution processing of UAV aerial road surface images, the peak signal-to-noise ratio and structural similarity of the ESR-KAN network model are higher than those of the existing excellent MASA-SR, RRSGAN, and DATSR. Compared with the above-mentioned similar super-resolution methods, the difference is significant, which proves that the reference image super-resolution processing technology of the present invention has superior application performance.

[0129] In order to more objectively demonstrate the superior performance of the application of the model of the present invention, the result images after super-resolution processing by the ESR-KAN network model were saved as a sample dataset and applied to the semantic segmentation task of road defects. Using several typical network structures, the high-resolution images obtained by the present invention show more superior performance when used. The results are shown in Table 2, where mAP represents the mean average precision, AP represents the average precision, and mIoU represents the mean intersection over union.

[0130] Table 2 Performance of the results of using ESR-KAN processed images on different models

[0131]

[0132] Example 2

[0133] Another application example of the present invention is to use the efficient KAN reference-image super-resolution ESR-KAN network model to improve the resolution of the road surface images captured by the drone swarm, in order to support more accurate road surface information system analysis and road surface environment monitoring. The work process is as Figure 2 shown. The following are the specific implementation steps of this application example:

[0134] 1. Data collection and preparation stage

[0135] 1.1 Data collection

[0136] Equipment: Obtain low-resolution images of the target road surface area from the drone swarm and obtain reference images from high-resolution aerial photography.

[0137] Operation: Ensure that the captured road surface images cover different states of the target area to obtain comprehensive surface feature information.

[0138] 1.2 Data processing

[0139] Dataset acquisition: Divide the collected images into a training set and a test set in a ratio of 7:3.

[0140] Data preprocessing: Perform cloud removal, fog removal on the images, and correct radiation and geometric distortions to ensure the quality of the dataset.

[0141] Data augmentation: Perform augmentation operations on the training set images, such as adjusting exposure and simulating different lighting conditions, to improve the adaptability of the model.

[0142] 2. Network model training stage

[0143] 2.1 Model construction

[0144] Network structure: Construct the ESR-KAN network model, including the KAN fusion module, KAN-channel attention, and reconstruction module, which are optimized specifically for the characteristics of drone aerial photography road surface images.

[0145] 2.2 Model training

[0146] Training settings: Use a deep learning framework to train the network model. Adopt the Adam optimizer, set the initial learning rate of Adam to 5e-5 and 1e-6, and adjust the learning rate and other hyperparameters as needed.

[0147] Loss function: Combine the prediction loss function l pred and the KAN spline sparse loss function l sparse , calculate the target loss function l total , and optimize the model parameters to improve the image quality.

[0148] 3. Super-resolution processing and analysis stage

[0149] 3.1 Model Deployment

[0150] Deployment: Integrate the trained ESR-KAN network model into the UAV aerial road image processing workflow to process actual UAV swarm aerial road image data.

[0151] Processing: For the input low-resolution UAV aerial road images, the ESR-KAN network model generates corresponding high-resolution images to enhance the recognizability of road surface features.

[0152] 3.2 Result Analysis

[0153] Quality Assessment: Evaluate the accuracy and reliability of the super-resolution images by comparing with actual road surface survey data.

[0154] 4. Application and Optimization Phase

[0155] 4.1 Image Application

[0156] Environmental Monitoring: The processed high-resolution images are used to monitor environmental indicators such as forest cover changes, urban expansion, and agricultural land changes.

[0157] Analysis: Provide to road information system analysts for more precise road maintenance and road detection.

[0158] 4.2 System Optimization

[0159] Performance Improvement: Optimize the network model and processing algorithm according to the feedback from analysts and the road environment monitoring team to improve the processing efficiency and image quality of the system.

[0160] User Feedback: Collect user feedback on the image processing effect and further adjust and improve the system settings.

[0161] Hardware Configuration:

[0162] Processor: It is required to have high-performance computing capabilities to process large-scale UAV aerial road image datasets.

[0163] Memory: Used to store training datasets, model weights, and processed images, requiring large capacity and high speed.

[0164] Graphics Processing Unit (GPU): Accelerate the training and inference processes of deep learning models to improve computing efficiency.

[0165] Network Interface: Used for data transmission and model updates to ensure high speed and stability.

[0166] Display Device: Used to view the processing results and image quality assessment in real time, requiring high resolution and color accuracy.

[0167] Through these steps, the system of the present invention can effectively improve the resolution of the road surface images captured by the UAV, providing more accurate image data support for road surface maintenance and road surface environment monitoring. In addition, the system allows for further optimization and expansion to adapt to the development of future technologies and new application requirements.

[0168] Compared with traditional methods, the ESR-KAN network model of the present invention not only achieves a significant improvement in image quality in the aspect of super-resolution processing of road surface images captured by the UAV, but also demonstrates great potential in cost optimization and the improvement of the accuracy of subsequent computer vision tasks.

[0169] First, in terms of cost optimization, traditional high-resolution cameras are often expensive, and multi-frame image synthesis technology requires a large amount of shooting and post-processing time, which will significantly increase the cost of the project. In contrast, the ESR-KAN network model can reconstruct high-resolution images from low-resolution images without increasing hardware costs through deep learning technology, thereby restoring more detailed features. This method not only reduces the dependence on high-cost imaging equipment but also avoids the cumbersome operations and time consumption of multi-frame synthesis technology, achieving effective cost control. Second, in terms of improving the accuracy of subsequent computer vision tasks, the ESR-KAN network model can significantly improve the detail clarity and overall quality of the images through its innovative network structure. This is of great significance for subsequent computer vision tasks such as road surface defect detection, target tracking, and analysis. High-resolution images can provide more detailed information, significantly improving the accuracy and reliability of these tasks. For example, by using the images processed by the ESR-KAN network model, targets on the road surface can be more accurately identified and tracked, improving the accuracy and robustness of target tracking. In addition, the high efficiency of the ESR-KAN network model means that it can process images in real time during the flight of the UAV, which is particularly important for application scenarios that require quick responses. Real-time high-resolution image analysis can provide immediate and clear image data for decision-makers, enabling them to make decisions quickly and improving the efficiency of emergency response.

[0170] Matters not described in the present invention are applicable to the prior art.

Claims

1. A super-resolution processing method for road images taken by an unmanned aerial vehicle, characterized in that: The method comprises the following: The road targets are photographed by drones to obtain road images; at the same time, the same targets are photographed by high-pixel cameras as reference images, and the reference images of all targets constitute a reference image set; Construct an efficient KAN reference image super-resolution ESR-KAN network model: The ESR-KAN network model includes a KAN fusion module, a reconstruction module, and M serial KAN-channel attention. The KAN fusion module includes a first layer composed of two KANs in parallel and a second layer composed of one KAN. The three KANs have the same structure. The two KANs in the first layer are used to respectively and reference image Acquire feature map information; after the two KAN processing results of the first layer are spliced, they enter the KAN of the second layer to obtain the output feature map ; Each KAN-channel attention consists of a serial structure consisting of a global average pooling GAP and two consecutive KANs, as well as a multiplication operation and residual connection to combine the inputs; in the KAN-channel attention, the input feature map Apply global average pooling GAP to compress channel features, use two consecutive KANs to learn the importance weight of each channel, generate a new vector for the importance weight of each channel, and use the new vector to input feature maps The channels are weighted to obtain the channel weight vector; the channel weight vector is combined with the input feature map Multiply, and then add to the input feature map Perform the addition operation to get the output of the KAN-channel attention , and used as the input of the next KAN-channel attention, and so on to get the output of the last KAN-channel attention ; The reconstruction module is used to restore the spatial shape of the image. The result processed by the reconstruction module is residually connected with the result of the upsampling of the reference image to obtain the final high-resolution image. ; The objective loss function of the ESR-KAN network model during training for: , , , in, is the prediction loss function; is the KAN spline sparse loss function; is a parameter that controls the overall regularization amplitude, and are relative calculation parameters, all set to 1; represents the last KAN in the last KAN-channel attention The activation function of the layer, Represents the activation function norm, is the entropy regularization term, L is the number of KAN layers; X represents the original input image, and Y represents the high-resolution image predicted by the network structure.

2. The method according to claim 1, characterized in that The reconstruction module includes sequentially connected Convolutional layer, ReLU activation function and Convolutional layer.

3. The method according to claim 1, characterized in that M=4-6。 4. The method according to claim 1, characterized in that: The training process of the ESR-KAN network model is: Initialize the network parameter settings: the epoch of the training network is set to 500, the optimizer uses the Adam optimizer, and the initial learning rate of Adam is set to 5e-5~1e-6; The training set comes from road images obtained by drone aerial photography and reference images of the same target taken by a high-pixel camera. The input images of the training set and the reference images are input into the ESR-KAN network model, and these images are read according to the storage path. When the overall objective loss function When there is no more obvious decline, the network model training is considered to be stable and the network training process is completed; The test set images also come from road images obtained through drone aerial photography. When conducting network testing, the test set images are input, and no additional reference images are needed. The network weights trained with the training set are imported. When the peak signal-to-noise ratio PSNR and structural similarity SSIM of the drone aerial road images reach 44.89±0.03 and 0.991±0.02 respectively, the entire network training is completed.

5. A super-resolution processing system for road images taken by drones, characterized in that: The system comprises: UAVs are used to obtain road surface images of road targets through aerial photography; A high-pixel camera is used to obtain a high-resolution reference image of the target at a certain viewing angle; An image preprocessing module is used to perform data enhancement processing on road surface images and reference images; Database, used to store images and parameters; The ESR-KAN network model described in claims 1-4 is used for super-resolution reconstruction of road surface images.

6. The system according to claim 5, characterized in that The system also includes a network interface, a display device, and an image processing unit. The display device is used to view the processing results and image quality evaluation results in real time.

7. The system according to claim 5, characterized in that The system also includes a road defect semantic segmentation model, and the super-resolution image reconstructed by the ESR-KAN network model is used for road defect recognition.

8. The system according to claim 5, characterized in that The super-resolution images are used to monitor forest cover changes, urban expansion and agricultural land changes, road maintenance and road inspection.

Citation Information

Patent Citations

  • KIA Net network model and image thereof, and high-precision wafer defect detection and segmentation method

    CN118762013A

  • Image processing method and apparatus, portrait super-resolution reconstruction method and apparatus, and portrait super-resolution reconstruction model training method and apparatus, electronic device, and storage medium

    WO2022057837A1