An UAV Perspective Building Image Target Segmentation Method Based on the Kolmogorov-Arnold Algorithm
By introducing the DR-KAN module and DeblurGANv2 network based on the Kolmogorov-Arnold algorithm in architectural image processing from the perspective of the drone, the problems of motion blur, target density, view angle changes and local occlusion in architectural image processing from the perspective of the drone are solved, and higher recognition accuracy and model stability are achieved.
Patent Information
- Application Number
- CN202510288691.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing computer vision models face problems such as motion blur, target density, viewing angle changes and local occlusion when processing architectural images from the drone perspective, resulting in reduced recognition accuracy and unstable model performance.
The DR-KAN module based on the Kolmogorov-Arnold algorithm is adopted, and the image defuzzing is combined with the DeblurGANv2 network, and the improved DRConv module performs hollow feature extraction and feature fusion to improve the model's extraction ability and segmentation generalization ability of high-dimensional feature maps.
It improves the accuracy and stability of building image target segmentation at the drone perspective, can accurately identify buildings in complex environments, and enhances the robustness and prediction performance of the model.
Smart Images

Figure CN119810123B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and unmanned aerial vehicle (UAV) technology, and particularly relates to a method for segmenting building image targets from the perspective of a UAV based on the Kolmogorov-Arnold algorithm. Background Art
[0002] The rapid development of UAV technology has made its applications in fields such as remote sensing, surveillance, and environmental monitoring increasingly widespread. The unique perspective of UAVs can provide unprecedented aerial views, enabling more comprehensive and detailed observation of ground conditions. This perspective is particularly important in natural disaster monitoring, urban planning, agricultural management, etc. For example, during natural disasters, UAVs can quickly fly to the disaster area and provide real-time high-definition images to help emergency response teams accurately assess the disaster situation and plan rescue strategies.
[0003] However, with the continuous complication of UAV application scenarios, how to accurately acquire and analyze image data from the perspective of UAVs has become a key issue. Traditional image segmentation methods often face problems of slow processing speed and insufficient accuracy when dealing with UAV images. This is mainly because the images taken by UAVs often have high resolutions and complex scenes, which makes the details in the images more abundant but also increases the processing difficulty. In addition, UAVs are affected by various environmental factors during flight, such as light changes, weather conditions, etc., and these factors further increase the challenges of image analysis.
[0004] In order to improve the effectiveness of image processing in real-time scene detection, more advanced image segmentation technologies need to be developed. These technologies should be able to handle the complex details of high-resolution images and quickly adapt to dynamically changing environmental conditions. Only through accurate image analysis can the potential of UAVs in various application fields be realized, promoting the further development and application of its technology.
[0005] U-Net is a deep learning model for image segmentation tasks, originally designed for medical image segmentation. The core idea of this model is to achieve accurate image segmentation through a symmetrical encoder-decoder structure. The encoder part extracts high-level features of the image through a series of convolutional and pooling layers, while the decoder part gradually restores the spatial resolution of the image through upsampling and convolution operations. The skip connection between the encoder and decoder allows low-level features to be fused with high-level semantic information, thereby preserving detail information and improving segmentation accuracy. U-Net is characterized by its symmetrical network structure and skip connections, which enable the network to effectively combine local and global information in the image restoration stage. The encoder stage gradually reduces the resolution of the image while extracting more and more abstract features, while the decoder stage gradually restores the resolution of the image and uses skip connections to bring back the details of the encoder stage, ensuring that the final segmentation result retains global semantic information and has sufficient details. In this way, U-Net can not only handle complex segmentation tasks, but also show excellent performance in fields such as medical imaging.
[0006] In recent years, with the advancement of deep learning technology, the application of U-Net network in image segmentation tasks has shown strong performance. However, when U-Net is applied to complex environments under the perspective of drones, it still faces several limitations. These problems may affect the effectiveness of the model in practical applications.
[0007] The design concept of the U-Net network is to achieve efficient feature extraction and fusion through simple sampling and skip connections. Under ideal conditions, this design can effectively combine low-level detail information with high-level semantic information. However, in images taken by drones, the processing of such multi-scale features becomes more complicated. Images from the perspective of drones often contain information of multiple scales from close-ups to distant views and at different depths, and the skip connections of U-Net may not be flexible enough in these cases to fully handle these multi-scale changes. This may result in the model being unable to accurately capture and segment the target in some areas with rich details, thus affecting the overall segmentation effect.
[0008] The processing of high-dimensional features is also an urgent problem to be solved. UAV images are usually of high resolution, and each image contains a large number of pixels and feature information. Although U-Net performs well when processing standard resolution images, the computational complexity of the model increases significantly when facing these high-resolution images. This not only places higher demands on computing resources, but may also lead to a significant increase in training and prediction time. In the case of limited resources, this computational overhead may become a bottleneck restricting real-time applications.
[0009] In addition, motion blur is a prevalent problem in drone images. Due to the relatively high flight speed of drones, slight vibrations and rapid movements during shooting can lead to image blurring. This blurring not only increases the ambiguity of object boundaries but also poses greater challenges for segmentation algorithms in identifying and differentiating objects. In contrast, medical images such as CT and MRI scans are usually acquired in a static state, with high image quality, clear boundaries, and the segmentation process can utilize higher contrast and resolution for more accurate identification.
[0010] Some researchers have proposed an improved U-Net by combining the Kolmogorov-Arnold algorithm with the original U-Net network, which enhances the ability to extract high-dimensional feature maps and improves the generalization ability and accuracy of segmentation. However, this network is mainly applied to medical image segmentation, and for building images from a drone perspective, there are problems such as low feature capture ability and weak model robustness.
[0011] In summary, the characteristics of building images from the perspective of drones pose significant challenges to existing computer vision models. First, due to the relatively high flight speed of drones, especially in the case of low-altitude flight or high-speed movement, motion blur is likely to occur in the images. This blur phenomenon affects the details of the images. Particularly when the drone is making rapid turns or changing altitude, the outlines and textures of buildings may become unclear or even completely distorted. Traditional image processing models often struggle to extract effective features from blurred images, resulting in a significant reduction in the recognition accuracy of buildings. In addition, building images captured by drones often exhibit the characteristic of dense targets, especially in city centers or high-density areas. This dense building layout makes it difficult for traditional object detection models to distinguish multiple adjacent or overlapping buildings in the same image. The small distances between buildings and the overlapping and occlusion in the perspective often lead to misidentifications or missed detections by the model. For example, when shooting in urban blocks, some buildings may be partially occluded by other tall buildings or obstacles, resulting in incomplete appearance information of the buildings and affecting the recognition accuracy of the model. With the continuous change of the drone's flight perspective, perspective transformation is also an issue that cannot be ignored. When shooting at low altitude or at an oblique angle, the projection, perspective distortion, and angle changes of buildings will significantly alter the appearance of buildings. For instance, when looking down from a high altitude, the planar outline of a building usually appears as a simple geometric shape, while when shooting at low altitude or an oblique perspective, the three-dimensional sense and perspective distortion of the building may increase, making it difficult for traditional models to recognize the geometric structure of the building. This instability of the perspective causes the model to be unable to maintain consistent performance at different flight heights and angles. Additionally, local occlusion in building images is also a prominent problem. The shooting angle of the drone may cause parts of the building to be occluded by other objects (such as adjacent buildings, trees, wires, etc.). This local occlusion affects the model's capture of the overall view of the building. Especially in complex urban environments, features such as the facades, windows, and door frames of buildings may be obscured, resulting in the model being unable to accurately extract the complete geometric information of the building. In this case, existing object detection algorithms may only be able to identify some parts of the building or produce inaccurate positioning. Summary of the Invention
[0012] The purpose of the present invention is to disclose a method for segmenting building image targets from the perspective of drones based on the Kolmogorov-Arnold algorithm. According to the existing U-KAN model, a new DR-KAN module is proposed, which can improve the extraction ability of high-dimensional feature maps of buildings captured by drones and enhance the generalization ability and accuracy of segmentation.
[0013] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:
[0014] A method for segmenting building image targets from the perspective of an unmanned aerial vehicle based on the Kolmogorov-Arnold algorithm, the method comprising the following steps:
[0015] Step 1: Preprocess the video data captured by the unmanned aerial vehicle;
[0016] Step 2: Introduce a DRconv module into the KAN module to obtain an improved DR-KAN module, replace the feature extraction module in the encoder of the U-Net network with the improved DR-KAN module, and construct a building image target segmentation model; the input data of the building image target segmentation model is the preprocessed video image, and the output result is the building segmentation image;
[0017] The DRconv module adjusts the number of channels of the feature map by performing a convolution operation on the input image, inputs the adjusted feature map into dilated convolutions with different dilation rates to obtain the dilated feature map; for the concatenated feature maps with different dilation rates, perform a feature fusion operation of ordinary convolution and depth convolution, and then perform a concatenation and Shuffle operation to improve the feature fusion effect of the feature map; finally, input the feature map after feature fusion into a convolutional layer to adjust the number of output channels, and perform a residual connection with the input feature map and then output;
[0018] Step 3: Post-process the segmentation result of the building image target segmentation model;
[0019] Step 4: Compare the post-processed segmentation result with the actual unmanned aerial vehicle video data to evaluate and verify the building image target segmentation model; adjust and optimize the parameters of the DR-KAN module according to the evaluation result.
[0020] As a preferred example, in step 2, the building image target segmentation model includes a first input layer, a feature extraction layer, a dilated feature fitting layer, a feature fusion layer, and a first output layer;
[0021] The feature extraction layer includes a first downsampling layer, a second downsampling layer, and a third downsampling layer connected in sequence. The feature extraction layer extracts features from the building image received by the input layer and outputs a high-dimensional feature map of the building image to the dilated feature fitting layer;
[0022] The dilated feature fitting layer includes a first DR-KAN module, a second DR-KAN module, a third DR-KAN module, and a fourth DR-KAN module connected in sequence. Among them, the output end of the first DR-KAN module is connected to the fourth DR-KAN module. In the dilated feature fitting layer, perform dilated feature extraction and KAN model fitting operations on the high-dimensional feature map, and output the fitted feature map to the feature fusion layer;
[0023] The feature fusion layer includes a first upsampling layer, a second upsampling layer, and a third upsampling layer connected in sequence; the first upsampling layer fuses the atrous fourth DR-KAN module and the first downsampling layer, the second upsampling layer fuses the output results of the first upsampling layer and the second downsampling layer, the third upsampling layer fuses the output results of the second upsampling layer and the first downsampling layer, and then the feature map with enhanced front-back relevance is output through the first output layer.
[0024] As a preferred example, step 1 further includes:
[0025] Smoothing the pixel values in the video image using a spatial domain filter, and processing the spectrum in the video image using a frequency domain filtering method to reduce noise;
[0026] Using a calibration board to estimate and compensate the lens distortion parameters, and aligning the video images captured from different perspectives or at different times through image registration technology;
[0027] Using the DeblurGANv2 network to perform deblurring operation on the image.
[0028] As a preferred example, the DR-KAN module includes a second input layer, a data dimension conversion layer, a KAN fitting layer, a DRconv module, a first stacking layer, a normalization layer, and a second output layer connected in sequence;
[0029] The data dimension conversion layer performs a flattening operation on the high-dimensional feature map sent by the second input layer, inputs the flattened high-dimensional feature map into the KAN fitting layer for feature fitting, and the feature map after fitting is used for atrous feature extraction through the DRConv module; the first stacking layer stacks the feature map extracted by the DRConv module with the original high-dimensional feature map, and after the fused feature after stacking is normalized by the normalization layer, it is output by the second output layer; the calculation formula of the DR-KAN module is:
[0030] .
[0031] As a preferred example, the DRconv module includes a third input layer, a first convolutional layer, a first atrous convolutional layer, a second atrous convolutional layer, a third atrous convolutional layer, a first fusion layer, a second convolutional layer, a second fusion layer, a depth convolutional layer, a Shuffle layer, a second stacking layer, and a third output layer;
[0032] The first convolutional layer Perform a convolution operation to adjust the number of channels of the feature map, and input the adjusted feature maps into the first dilated convolutional layer, the second dilated convolutional layer, and the third dilated convolutional layer with different dilation rates respectively; the first fusion layer splices the feature maps output by the first dilated convolutional layer, the second dilated convolutional layer, and the third dilated convolutional layer to obtain the dilated feature map ; the dilated feature map After sequentially adjusting the output channels through the second convolutional layer and the depth convolutional layer, through the second fusion layer and the Shuffle layer, perform splicing and Shuffle operations with the output result of the second convolutional layer to obtain the fused feature map; after the fused feature map and the input feature map are connected by the second stacking layer residual, output through the third output layer; the calculation formula of the DRconv module is as follows:
[0033] ;
[0034] ;
[0035] Among them, is the number of input channels, is the output channel, n is the number of dilated convolutions, and i is the dilated convolution number
[0036] Step 3 further includes that post-processing the segmentation result of the building image target segmentation model means converting the formats and classifying different parameters in the building segmentation image output in tensor form
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] First, the method for segmenting building images from the perspective of drones based on the Kolmogorov-Arnold algorithm of the present invention proposes a new DR-KAN module according to the existing U-KAN model, which can improve the ability to extract high-dimensional feature maps of buildings captured by drones, enhance the generalization ability and accuracy of segmentation, and enable the network to recognize buildings even when only part of the content (such as occlusion or building overlap) is present
[0039] Second, the method for segmenting building images from the perspective of drones based on the Kolmogorov-Arnold algorithm of the present invention uses the DeblurGANv2 network to perform deblurring operations on the images collected by drones, effectively avoiding the recognition difficulties caused by image blurring due to the relatively fast flight speed of drones and slight vibrations and rapid movements during shooting
[0040] Thirdly, the method for segmenting building image targets from the perspective of an unmanned aerial vehicle (UAV) based on the Kolmogorov-Arnold algorithm of the present invention enables the network to better capture complex features in the data during the non-linear mapping process; this optimization not only improves the expression ability of the network but also enhances its recognition ability for complex data patterns, thereby further improving the prediction performance and training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flow chart of the method for segmenting building image targets from the perspective of an unmanned aerial vehicle (UAV) based on the Kolmogorov-Arnold algorithm of the present invention;
[0042] Figure 2 is a structural diagram of the building image target segmentation model;
[0043] Figure 3 is a schematic structural diagram of the DR-KAN module;
[0044] Figure 4 is a schematic structural diagram of the DRConv module;
[0045] Figure 5 is a schematic diagram of the fusion steps of the DRconv module for the feature map. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following further describes the embodiments of the present invention in detail with reference to the accompanying drawings.
[0047] See Figure 1 , the present invention discloses a method for segmenting building image targets from the perspective of an unmanned aerial vehicle (UAV) based on the Kolmogorov-Arnold algorithm, and the method includes the following steps:
[0048] Step 1: Preprocess the video data captured by the unmanned aerial vehicle (UAV);
[0049] Step 2: Introduce the DRconv module into the KAN module to obtain the improved DR-KAN module, and replace the feature extraction module in the encoder of the U-Net network with the improved DR-KAN module to construct a building image target segmentation model; the input data of the building image target segmentation model is the preprocessed video image, and the output result is the building segmentation image;
[0050] The DRconv module adjusts the number of channels of the feature map by performing a convolution operation on the input image, inputs the adjusted feature map into dilated convolutions with different dilation rates to obtain the dilated feature map; then performs ordinary convolution and depthwise convolution on the dilated feature map respectively to adjust the output channels, and then through concatenation and Shuffle operations to enhance the feature fusion effect of the feature map; finally, inputs the feature map after feature fusion into a convolutional layer to adjust the number of output channels, performs a residual connection with the input feature map and then outputs;
[0051] Step 3: Post-process the segmentation result of the building image target segmentation model;
[0052] Step 4: Compare the post-processed segmentation result with the actual drone video data to evaluate and verify the building image target segmentation model; adjust and optimize the parameters of the DR-KAN module according to the evaluation result.
[0053] In Step 1, the process of preprocessing the video data captured by the drone specifically includes the following steps:
[0054] 1-1: Image denoising. The video data captured by the drone may be affected by various noises, such as sensor noise and environmental interference. The goal of denoising is to eliminate these unnecessary noises in order to extract clear visual information from the video. Denoising methods include spatial domain filters, such as mean filters and median filters, which can smooth the pixel values in the image; and frequency domain filtering methods, such as high-pass filters and low-pass filters, which reduce noise by processing the spectrum of the image.
[0055] 1-2: Image correction. The video captured by the drone often has geometric distortions. Especially in the edge area of the lens, the image may have non-linear deformations. The purpose of correction is to correct these geometric distortions so that the objects in the image can accurately reflect their true shape and position on the two-dimensional plane. The correction process usually includes lens distortion correction, image alignment, and geometric transformation, etc. Specific methods include using a calibration board to estimate and compensate the lens distortion parameters, or aligning images taken from different perspectives or times through image registration techniques.
[0056] 1-3: Image de-motion blur. Motion blur in the drone view means that during the shooting process, due to the rapid movement or vibration of the drone, the boundaries of the objects in the image become blurred. This phenomenon not only affects the clarity of the image, but also poses challenges to image processing, especially image segmentation. In the present invention, the DeblurGANv2 network is used to de-blur the images collected by the drone.
[0057] In step 2, the present invention uses the improved encoder network to extract features from the preprocessed video data. The DR-KAN module realizes more refined feature extraction at each encoding stage by using the improved DRConv, and combines multi-scale features to enhance the model's adaptability to complex scenes and diverse objects.
[0058] Figure 2 It is the structural diagram of the building image target segmentation model. The building image target segmentation model includes a first input layer, a feature extraction layer, a DRConv module (hole feature fitting layer), a feature fusion layer, and a first output layer; the input image received by the first input layer extracts a high-dimensional feature map through the feature extraction layer, and then inputs it into the hole feature fitting layer, where the high-dimensional feature map is subjected to hole feature extraction and KAN model fitting operations. The output feature map is fused with the input feature map through the feature fusion layer to increase the front-back correlation of the feature map, and is output through the first output layer.
[0059] Figure 3 It is the structural schematic diagram of the DR-KAN module. After the DR-KAN module of the present invention downsamples the input image, the obtained high-dimensional feature map is flattened into 2D patches, and the size of each patch is , and there are blocks of patches in total; where and are the scale information of the high-dimensional feature map respectively. Then the flattened feature map is input into the KAN layer for deep feature extraction, and the output feature map is subjected to residual connection with the previous input after passing through the improved convolutional layer, and finally the output feature map is subjected to layer normalization operation. The calculation formula of the DR-KAN module is as follows:
[0060] .
[0061] Figure 4 It is the structural schematic diagram of the DRConv module. The improved DRConv module of the present invention adjusts the number of channels of the feature map by performing a convolution operation on the input image, and then inputs the adjusted feature map into dilated convolutions with different dilation rates to obtain the dilated feature map ; then the output channels of the feature map are adjusted by performing ordinary convolution and depth convolution respectively, and after splicing and Shuffle operations, the feature fusion effect of the feature map is improved; finally, the feature map after feature fusion is input into the convolutional layer to adjust the number of output channels, and is output after performing residual connection with the input feature map. The calculation formula of the DRConv module is as follows:
[0062] ;
[0063] 。
[0064] The fusion steps of the DRconv module for the feature map are as Figure 5 shown.
[0065] The improved DRConv module in the present invention has certain advantages in the task of building image recognition from the perspective of an unmanned aerial vehicle. When traditional convolutional neural networks (CNNs) process building images in dynamic scenes, it is often difficult to effectively extract complex spatial features, especially in the case of morphological deformation, partial occlusion, and low-light conditions of buildings. The DRConv module significantly enhances the network's perception ability of details and edges in building images by introducing dilated convolution and combining it with depth convolution operations. Dilated convolution expands the receptive field, effectively avoiding information loss while maintaining computational efficiency, enabling the model to capture more context information from a larger range. Especially when the building is in a complex background or the perspective changes greatly, the introduction of dilated convolution can help the model better identify the structural features of the building and improve the recognition ability of the building.
[0066] In addition, the DRConv module effectively improves the effect of feature fusion by adjusting the number of channels of the feature map and combining concatenation and Shuffle operations. This process enables the network to more flexibly combine and utilize features at different levels. Especially when facing multiple targets that are dense and highly similar in building images, it can more accurately distinguish the subtle differences between different buildings. For building images taken by an unmanned aerial vehicle, there are often problems such as perspective changes, partial occlusion, and blurred details of the building. The DRConv module can maintain a high recognition accuracy in such a complex environment. Through concatenation and Shuffle operations, the network can effectively exchange and fuse information among features at different scales, thereby enhancing the model's recognition ability for different types of buildings.
[0067] The application of depth convolution also further enhances the performance of the model in high-level feature extraction. Building images from the perspective of an unmanned aerial vehicle often involve different distance scales and different geometric shapes. The depth convolution layer can extract richer spatial information by refining the feature representation, thereby helping the model to be more robust when recognizing buildings. In the case of partial occlusion of the building, depth convolution can compensate for some missing information through higher-dimensional feature extraction, thus reducing the possibility of missed detection.
[0068] Finally, the residual connection design in the module further improves the stability and training efficiency of the network. The building images captured by drones are often affected by factors such as flight dynamics and lighting changes, and there may be certain fluctuations in the quality of the input images. The residual connection effectively alleviates the problem of gradient disappearance during network training by directly adding the input feature map and the output feature map, promoting the effective transmission of information. For building images from the drone's perspective, this design can ensure that the model still maintains relatively accurate recognition ability in complex scenarios, avoiding a decline in recognition accuracy caused by unstable training.
[0069] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0070] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions run by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0071] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0072] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to generate a computer implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in a flowchart Figure 1 one or more flowcharts and / or blocks Figure 1 or steps for implementing the functions specified in a block or blocks.
[0073] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0074] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for segmenting building images from a drone perspective based on the Kolmogorov-Arnold algorithm, characterized in that: The method comprises the following steps: Step 1: Preprocess the video data taken by the drone; Step 2: Introduce the DRconv module into the KAN module to obtain an improved DR-KAN module, replace the feature extraction module in the encoder of the U-Net network with the improved DR-KAN module, and construct a building image target segmentation model; the input data of the building image target segmentation model is the preprocessed video image, and the output result is a building segmentation image; The DRconv module adjusts the number of feature map channels by performing a convolution operation on the input image, and inputs the adjusted feature map into a dilated convolution with different dilation rates to obtain a feature map after dilation; for the spliced feature maps with different dilation rates, a feature fusion operation of ordinary convolution and deep convolution is performed, and then the splicing and shuffle operations are performed to improve the feature fusion effect of the feature map; finally, the feature map after feature fusion is input into the convolution layer to adjust the number of output channels, and is output after performing a residual connection with the input feature map; Step 3: Post-process the segmentation results of the building image target segmentation model; Step 4: Compare the post-processed segmentation results with the actual UAV video data to evaluate and verify the building image target segmentation model; adjust and optimize the parameters of the DR-KAN module based on the evaluation results; The DR-KAN module includes a second input layer, a data dimension conversion layer, a KAN fitting layer, a DRconv module, a first stacking layer, a normalization layer, and a second output layer connected in sequence; The data dimension conversion layer converts the high-dimensional feature map X sent by the second input layer L The flattening operation is performed, and the flattened high-dimensional feature map is input into the KAN fitting layer for feature fitting. The fitted feature map is subjected to hole feature extraction by the DRConv module. The first superposition layer superimposes the feature map extracted by the DRConv module with the original high-dimensional feature map. The superimposed fusion feature is normalized by the normalization layer and output by the second output layer. The calculation formula of the DR-KAN module is: X out,K =Norm(Res(X L ,DRConv(KAN(flatten(X L )))))。 2. The method for segmenting building images from the perspective of an unmanned aerial vehicle based on the Kolmogorov-Arnold algorithm according to claim 1, characterized in that: In step 2, the building image target segmentation model includes a first input layer, a feature extraction layer, a hole feature fitting layer, a feature fusion layer and a first output layer; The feature extraction layer comprises a first downsampling layer, a second downsampling layer and a third downsampling layer connected in sequence, the feature extraction layer extracts features from the building image received by the first input layer, and outputs a high-dimensional feature map of the building image to the hole feature fitting layer; The hole feature fitting layer includes a first DR-KAN module, a second DR-KAN module, a third DR-KAN module and a fourth DR-KAN module connected in sequence, wherein the output end of the first DR-KAN module is connected to the fourth DR-KAN module, and hole feature extraction and KAN model fitting operations are performed on the high-dimensional feature map in the hole feature fitting layer, and the fitted feature map is output to the feature fusion layer; The feature fusion layer includes a first upsampling layer, a second upsampling layer and a third upsampling layer connected in sequence; the first upsampling layer fuses the hole fourth DR-KAN module and the first downsampling layer, the second upsampling layer fuses the output results of the first upsampling layer and the second downsampling layer, the third upsampling layer fuses the output results of the second upsampling layer and the first downsampling layer, and then outputs the feature map with increased front-back correlation through the first output layer.
3. The method for segmenting building images from the perspective of an unmanned aerial vehicle based on the Kolmogorov-Arnold algorithm according to claim 1, characterized in that: Step 1 further comprises: The spatial domain filter is used to smooth the pixel values in the video image, and the frequency domain filtering method is used to process the spectrum in the video image to reduce noise; Use calibration plates to estimate and compensate for lens distortion parameters, and use image registration technology to align video images shot at different viewing angles or times; Use the DeblurGANv2 network to deblur the image.
4. The method for segmenting building images from the perspective of an unmanned aerial vehicle based on the Kolmogorov-Arnold algorithm according to claim 3, characterized in that: The DRconv module includes a third input layer, a first convolutional layer, a first hole convolutional layer, a second hole convolutional layer, a third hole convolutional layer, a first fusion layer, a second convolutional layer, a second fusion layer, a depth convolutional layer, a Shuffle layer, a second stacking layer and a third output layer; The first convolutional layer receives the input feature map X from the third input layer. in A convolution operation is performed to adjust the number of feature map channels, and the adjusted feature maps are respectively input into the first dilated convolution layer, the second dilated convolution layer, and the third dilated convolution layer with different dilated rates; the first fusion layer concatenates the feature maps output by the first dilated convolution layer, the second dilated convolution layer, and the third dilated convolution layer to obtain the feature map X after dilation. d ; Feature map X after hole d After adjusting the output channel through the second convolution layer and the deep convolution layer in turn, the output result of the second convolution layer is spliced and shuffled through the second fusion layer and the shuffle layer to obtain a fused feature map; the fused feature map and the input feature map are connected with the second stacking layer residual and output through the third output layer; the calculation formula of the DRconv module is as follows: Among them, c in is the number of input channels, c out is the output channel, n is the number of atrous convolutions, and i is the atrous convolution number.
5. The method for segmenting building images from the perspective of an unmanned aerial vehicle based on the Kolmogorov-Arnold algorithm according to claim 1, characterized in that: In step 3, post-processing the segmentation results of the building image target segmentation model refers to format conversion and classification of different parameters in the building segmentation image output in the form of tensor.
Citation Information
Patent Citations
Multi-level image restoration method based on partial-to-overall attention mechanism
CN111127346A
Remote sensing image ship segmentation method based on U-KAN network model
CN119600604A