A high-efficiency recognition system for UAV images

By introducing high-resolution cameras, deep learning models and adaptive target tracking algorithms into the drone image recognition system, the problem of unstable image recognition and tracking during high-speed flight is solved, and efficient and accurate target recognition and tracking is achieved.

CN119148755BActive Publication Date: 2025-05-27CADDX US (SHENZHEN) LTD CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411025256.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-05-27
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

Existing drone image recognition systems are difficult to capture clear images during high-speed flights, with low recognition accuracy, low image processing efficiency, unable to respond to dynamic changes in real time, and unstable target tracking.

Method used

A high-efficiency recognition system for image recognition of drones is designed, including a high-resolution camera, an image preprocessing unit, an object detection unit and a target tracking unit. The system performs object detection through a deep learning model, combines Kalman filter and optical flow method to perform adaptive target tracking, and adjusts the drone's flight trajectory through the control unit.

Benefits of technology

It realizes efficient and accurate identification and tracking of target objects under high-speed flight conditions, and improves the autonomous flight and mission execution capabilities of the drone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119148755B_ABST
    Figure CN119148755B_ABST
Patent Text Reader

Abstract

The present invention relates to a high-efficiency recognition system for UAV images, which includes a UAV, an image acquisition unit, an image preprocessing unit, a target detection unit, a target tracking unit, and a control unit; the UAV has the ability to fly at high speed and stability, and the image acquisition unit consists of a high-resolution camera, which is used to capture image data in real time during flight; the image preprocessing unit performs denoising and deblurring processing on the captured image data to generate preprocessed image data; the target detection unit uses a pre-trained deep learning model to achieve efficient recognition of target objects in the preprocessed images; the target tracking unit combines the Kalman filter and the optical flow method to lock and track the target object; the control unit adjusts the flight trajectory of the UAV according to the position and motion information of the target object provided by the target tracking unit to ensure continuous tracking of the target object; through the collaborative work of each unit, the system realizes the efficient and accurate recognition and tracking of target objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a high-efficiency recognition system for unmanned aerial vehicle images. Background Art

[0002] Unmanned aerial vehicles are widely used in various fields, including environmental monitoring, agriculture, logistics, disaster relief, and military. To improve the autonomous flight ability of unmanned aerial vehicles in complex environments, image recognition and target tracking technologies have received increasing attention. Traditional unmanned aerial vehicle image recognition systems usually rely on basic camera devices and simple image processing algorithms to capture and analyze image data on the flight path.

[0003] However, these systems in the prior art have many limitations in high-speed flight. First, traditional camera devices are difficult to capture clear images during high-speed movement, resulting in a decrease in recognition accuracy. Second, existing image processing algorithms are less efficient in processing a large amount of high-resolution image data and cannot respond in real time to dynamic changes during flight. In addition, traditional target tracking methods cannot stably lock targets during high-speed flight, affecting the autonomous flight and mission execution capabilities of unmanned aerial vehicles.

[0004] Therefore, it is very necessary to develop a new high-efficiency recognition system for unmanned aerial vehicle images. Summary of the Invention

[0005] This application provides a high-efficiency recognition system for unmanned aerial vehicle images to achieve efficient and accurate recognition and tracking of target objects.

[0006] This application provides a high-efficiency recognition system for unmanned aerial vehicle images, including:

[0007] An unmanned aerial vehicle for flying stably at high speed;

[0008] An image acquisition unit, including a high-resolution camera placed on the unmanned aerial vehicle, for capturing image data in real time during flight;

[0009] An image preprocessing unit for performing preprocessing operations such as denoising and deblurring on the image data captured by the image acquisition unit to obtain preprocessed image data;

[0010] A target detection unit, including a pre-trained deep learning model, for identifying target objects in the preprocessed images provided by the image preprocessing unit, wherein the deep learning model is optimized based on a convolutional neural network structure to improve recognition speed and accuracy;

[0011] A target tracking unit, which is used to execute an adaptive target tracking algorithm to lock and track the target object identified by the target detection unit. The adaptive target tracking algorithm is implemented based on the fusion of a Kalman filter and an optical flow method, and is used to provide stable target tracking in the case of high-speed flight.

[0012] A control unit, according to the position and motion information of the target object provided by the target tracking unit, adjusts the flight trajectory of the drone to maintain continuous tracking of the target object.

[0013] The present application has the following beneficial technical effects:

[0014] (1) Through the combination of a high-resolution camera and an image preprocessing unit, the system can capture and process image data in real time during high-speed flight. The preprocessing unit performs denoising and deblurring operations, improving the clarity and quality of the images, and providing a more reliable data basis for subsequent target recognition and tracking.

[0015] (2) The target detection unit uses a pre-trained deep learning model optimized based on a convolutional neural network structure. This optimization significantly improves the speed and accuracy of target recognition, and can quickly identify and classify target objects in images in complex environments, thereby enhancing the intelligence level of the drone when performing tasks.

[0016] (3) The target tracking unit combines an adaptive target tracking algorithm of a Kalman filter and an optical flow method to ensure stable locking and tracking of the target object even in the case of high-speed flight. This fusion algorithm improves the stability and accuracy of tracking, enabling the drone to continuously track the target and maintain the effectiveness of tracking even in high-speed movement and complex environments.

[0017] (4) The control unit adjusts the flight trajectory of the drone in real time according to the position and motion information of the target object provided by the target tracking unit. This intelligent adjustment mechanism enables the drone to flexibly respond to the dynamic changes of the target object, ensuring continuous and stable tracking, and enhancing the autonomy and reliability of the drone in various tasks. Description of the Drawings

[0018] Figure 1 It is a schematic diagram of a high-efficiency drone image recognition system provided by the first embodiment of the present application. Detailed Embodiments

[0019] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0020] The first embodiment of this application provides a high-efficiency unmanned aerial vehicle (UAV) image recognition system. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of this application. The following will Figure 1 describe in detail the high-efficiency UAV image recognition system provided by the first embodiment of this application.

[0021] The high-efficiency UAV image recognition system includes a UAV 101, an image acquisition unit 102, an image preprocessing unit 103, a target detection unit 104, a target tracking unit 105, and a control unit 106.

[0022] The UAV 101 is used for high-speed and stable flight.

[0023] The UAV 101 plays a crucial role in the present invention. It is responsible for achieving high-speed and stable flight, ensuring that the system can operate efficiently in various tasks. To ensure high-speed and stable flight, the UAV 101 needs to possess a number of key technologies and design features.

[0024] Firstly, the UAV 101 is equipped with a powerful power system, including a high-performance motor and a lightweight but sturdy fuselage structure. The motor provides sufficient thrust to enable the UAV to maintain stability during high-speed flight, while the lightweight fuselage design helps to improve flight efficiency and reduce energy consumption. The fuselage material is usually made of carbon fiber composite material to ensure its durability and stability during high-speed flight.

[0025] Secondly, the UAV 101 is also equipped with an advanced flight control system, which integrates a variety of sensors and control algorithms to achieve precise flight control. The main sensors include a GPS module, an inertial measurement unit (IMU), a barometer, and a laser altimeter. These sensors provide real-time position, altitude, speed, and attitude information, and the flight control system makes dynamic adjustments based on this data to ensure that the UAV remains stable during high-speed flight.

[0026] The flight control system adopts advanced control algorithms such as PID control, LQR control, and adaptive control. These algorithms combine sensor data to be able to adjust the thrust, control surfaces, and motor speed of the UAV in real time, enabling it to maintain the best state under different flight conditions. Especially during high-speed flight, these control algorithms can effectively cope with sudden airflow changes and environmental disturbances, ensuring flight stability and safety.

[0027] To adapt to different mission requirements, the drone 101 also has a certain degree of autonomous flight ability. The flight control system is built-in with path planning and obstacle avoidance algorithms, which can automatically adjust the flight trajectory according to the preset flight route and real-time environmental information, avoid obstacles, and fly flexibly in complex environments. These algorithms usually combine devices such as visual sensors and lidar, and through image recognition and laser scanning, they can perceive the surrounding environment in real time and make quick responses.

[0028] The communication system of the drone 101 is equally crucial. It is equipped with a high-performance wireless communication module, which can achieve real-time data transmission and command reception with the ground control station. The communication system usually adopts multi-band radio technology to ensure stable signal connections in different environments. Through this communication link, the ground control station can monitor the status of the drone in real time and issue control commands when necessary to further ensure the successful completion of the flight mission.

[0029] In short, through the coordinated work of the power system, flight control system, autonomous flight ability, and communication system, the drone 101 ensures its stability and reliability during high-speed flight.

[0030] The image acquisition unit 102 includes a high-resolution camera placed on the drone, which is used to capture image data in real time during flight.

[0031] The image acquisition unit 102 is a key component in the high-efficiency image recognition system of the drone, responsible for capturing high-quality image data in real time during the high-speed flight of the drone. This unit mainly consists of a high-resolution camera and a related image transmission system to ensure the accuracy and timeliness of the image data.

[0032] The high-resolution camera is the core component of the image acquisition unit, and its performance directly affects the quality of the image data and the effect of subsequent processing. The camera usually has a resolution of no less than 4K to ensure that it can capture image data with rich details. A high frame rate is also required, usually reaching 30 frames per second or higher, to ensure that continuous and clear images can still be captured when the drone is moving at high speed. The optical part of the camera uses a wide-angle lens to expand the field of view and ensure that more areas can be covered and more environmental information can be captured.

[0033] In actual operation, the camera is fixedly installed at the bottom or front of the drone and is stably supported by a shock-proof bracket. The shock-proof bracket uses shock-absorbing materials and mechanisms, which can effectively absorb and buffer the vibrations generated during the flight of the drone, ensuring that the camera can remain stable under various flight conditions. This design greatly reduces the risk of image blurring during flight and improves the clarity and stability of the images.

[0034] The image transmission system is an important part of the image acquisition unit and is responsible for transmitting the image data captured by the camera to the image preprocessing unit in real time. To this end, the system usually adopts a high-speed data interface to ensure that a large amount of image data can be transmitted quickly and stably. The wireless transmission module is also a key part. Especially when the drone is flying at a long distance, the image data needs to be transmitted to the ground control station through a wireless channel. The wireless transmission module uses multi-band radio technology to ensure signal strength and transmission speed in different environments, avoiding data loss or delay caused by unstable signals.

[0035] In addition, the image acquisition unit also includes an intelligent control system, which is responsible for managing the working state of the camera and the data transmission process. This system can dynamically adjust the parameters of the camera, such as exposure time, focal length, and frame rate, according to flight conditions and mission requirements, ensuring that the best images can be captured under different lighting and speed conditions. The intelligent control system also has an image caching function. Through the built-in high-speed cache module, it temporarily stores the image data captured by the camera to cope with emergencies during the data transmission process and ensure the integrity and continuity of the image data.

[0036] Through these designs and configurations, the image acquisition unit 102 can stably and continuously capture high-quality image data during the high-speed flight of the drone, providing reliable basic data for subsequent image preprocessing, target detection, and tracking.

[0037] The image preprocessing unit 103 is used to perform preprocessing operations such as denoising and deblurring on the image data captured by the image acquisition unit to obtain preprocessed image data.

[0038] The image preprocessing unit 103 plays a key role in the high-efficiency drone image recognition system. It is responsible for performing preprocessing operations such as denoising and deblurring on the image data captured by the image acquisition unit 102 to generate high-quality preprocessed image data, ensuring the accuracy and stability of subsequent target detection and tracking.

[0039] First of all, the image preprocessing unit 103 starts to perform preliminary denoising processing by receiving the raw image data transmitted by the image acquisition unit. The purpose of denoising processing is to eliminate the noise introduced into the image due to factors such as environmental light changes and sensor noise. Common denoising algorithms include Gaussian filtering, median filtering, and bilateral filtering, etc. These algorithms can effectively reduce noise while retaining image details. Specifically, when implementing, the image preprocessing unit can dynamically select and adjust the filtering parameters to adapt to different image qualities and noise levels, ensuring the optimization of the denoising effect.

[0040] After the denoising process is completed, the image preprocessing unit then performs deblurring. Since the UAV may cause motion blur during high-speed flight, deblurring is essential. Deblurring generally uses inverse filtering, blind deconvolution, or motion compensation algorithms. These algorithms reconstruct a clear image by analyzing the motion trajectory and blur degree in the image. Implementing these algorithms requires high computing power. Therefore, the image preprocessing unit usually integrates a dedicated image processing chip or uses a high-performance processor to ensure the efficiency and effect of real-time processing.

[0041] After completing the denoising and deblurring processes, the image preprocessing unit further performs image enhancement operations. Image enhancement aims to improve the contrast and clarity of the image, making the target object more prominent. Common methods include histogram equalization, adaptive histogram equalization, and Laplacian enhancement, etc. These methods can effectively enhance the detail and edge information in the image, improving the recognition accuracy of the subsequent target detection unit.

[0042] In addition, the image preprocessing unit is also responsible for geometric correction and normalization of the image. Geometric correction is used to correct image distortion caused by factors such as the camera's perspective and the UAV's flight attitude changes, ensuring the geometric consistency of the image. Normalization processing includes adjusting the image to a fixed size and format to match the input standards required by the target detection unit. This step is usually achieved through operations such as image resampling, rotation, cropping, and normalization.

[0043] The design and implementation of the image preprocessing unit 103 ensure an efficient and smooth transition from image acquisition to image processing. The generated high-quality preprocessed image data provides a reliable basis for subsequent target detection and tracking.

[0044] Furthermore, the image preprocessing unit is specifically used for:

[0045] Perform noise suppression processing on the image data, using the adaptive median filtering algorithm to remove random noise in the image;

[0046] Use the motion estimation algorithm to perform deblurring on the image data and restore the motion blur caused by high-speed flight;

[0047] Apply histogram equalization technology to enhance the contrast and details of the image to improve the recognition accuracy of the subsequent target detection unit.

[0048] The image preprocessing unit is specifically used to perform a series of processes on the captured image data to improve the image quality and the accuracy of subsequent target detection. The specific implementation includes the following steps:

[0049] First, the image preprocessing unit performs noise suppression processing on the image data. The noise suppression processing uses an adaptive median filtering algorithm, which is a commonly used denoising method, especially suitable for removing random noise in images, such as salt noise and pepper noise. The adaptive median filtering algorithm processes the noise in the image by adjusting the size of the filter window. Specifically, the algorithm defines a window around each pixel, calculates the median value of the pixels within the window, and replaces the central pixel value with this median value, thereby effectively removing the noise without losing the details of the image.

[0050] Secondly, the image preprocessing unit uses a motion estimation algorithm to deblur the image data. Motion blur is usually caused by the rapid movement of the camera during the high-speed flight of the drone. The motion estimation algorithm first estimates the motion blur parameters present in the image, such as the motion direction and the motion distance. Then, through inverse filtering technology or blind deconvolution technology, the details in the image are restored, thereby reducing or eliminating the motion blur. This step is crucial for ensuring the clarity of the image and the accuracy of subsequent processing.

[0051] Next, the image preprocessing unit applies histogram equalization technology to enhance the contrast and details of the image. Histogram equalization is a commonly used image processing technology. By adjusting the gray-scale distribution of the image, the brightness and contrast in the image are enhanced, thereby highlighting the detailed parts of the image. The specific operation is to calculate the gray-scale histogram of the image and redistribute the gray-scale values through the cumulative distribution function, so as to improve the overall contrast of the image. The enhanced image is visually clearer and has richer details, which helps the subsequent target detection unit improve the recognition accuracy.

[0052] Through the above steps, the image preprocessing unit can effectively improve the image quality and provide higher-quality input data for subsequent target detection and recognition. The adaptive median filtering algorithm can effectively remove the random noise in the image, making the image cleaner; the motion estimation algorithm can restore the motion blur caused by high-speed flight, making the image clearer; the histogram equalization technology can enhance the contrast and details of the image, making the target object more prominent. These preprocessing steps cooperate with each other to ensure that the image data obtained by the drone during high-speed flight can reach the best quality, thereby improving the recognition accuracy and efficiency of the subsequent target detection unit.

[0053] The target detection unit 104 includes a pre-trained deep learning model for identifying target objects in the preprocessed image provided by the image preprocessing unit, wherein the deep learning model is optimized based on a convolutional neural network structure to improve the recognition speed and accuracy.

[0054] The target detection unit 104 plays a crucial role in the high-efficiency recognition system of UAV images. Its main task is to identify the target objects in the preprocessed images provided by the image preprocessing unit 103 through a pre-trained deep learning model. This unit relies on an advanced convolutional neural network (CNN) structure, and the optimized model can significantly improve the speed and accuracy of target recognition.

[0055] First, the target detection unit receives high-quality image data from the image preprocessing unit. To ensure the efficiency and accuracy of the recognition process, this image data has usually undergone denoising, deblurring, and enhancement processing. After receiving the images, the target detection unit inputs them into the deep learning model for processing.

[0056] The core of the deep learning model is the convolutional neural network (CNN), which has excellent performance in processing image data. The CNN consists of multiple convolutional layers, pooling layers, and fully connected layers. Through layer-by-layer processing, it extracts the feature information in the images. Specifically, the convolutional layer uses a set of trainable filters to perform convolution operations on the input image, extracting low-level features such as edges and textures. The pooling layer reduces the amount of data while retaining important features through downsampling operations, enhancing the model's anti-interference ability. The fully connected layer comprehensively processes the extracted features and outputs the class probabilities and positions of each target object in the image.

[0057] In the target detection unit, the convolutional neural network structure has been specifically optimized to meet the real-time processing requirements during the high-speed flight of UAVs. First, the model adopts a lightweight design, reducing the number of parameters and the amount of computation, enabling it to run at high speed with limited computing resources. Second, by introducing skip connections and residual blocks, the model can more effectively transmit feature information, prevent gradient vanishing, and improve the training efficiency and recognition accuracy.

[0058] To improve the robustness of target detection, the deep learning model also incorporates multi-scale feature fusion technology. Multi-scale feature fusion allows the model to extract and fuse feature information at different levels, thereby better identifying target objects of different scales and complexities. For example, the model can simultaneously consider the global features of large objects and the local details of small objects, significantly improving the detection accuracy and adaptability.

[0059] During the training phase, the model undergoes supervised learning using a large amount of labeled data, and continuously adjusts the network parameters using the backpropagation algorithm to minimize the gap between the prediction results and the true labels. The training dataset covers various flight scenarios and target object types, ensuring that the model has wide adaptability and high accuracy in practical applications.

[0060] During actual operation, the target detection unit not only identifies the category of the target object but also precisely locates its position in the image. Each detection result includes the category label of the target object and the position information (such as the coordinates of the bounding box). This information is transmitted to the target tracking unit and the control unit for subsequent target tracking and flight trajectory adjustment.

[0061] Furthermore, the deep learning model includes a feature extraction part, a feature fusion part, a target detection part, and a post-processing part;

[0062] The feature extraction part is used to extract features from the preprocessed image data to obtain feature maps; among them, the feature extraction part is implemented by an improved convolutional neural network, and the improved convolutional neural network includes multiple convolutional layers, pooling layers, and batch normalization layers, and introduces separable convolutional layers;

[0063] The feature fusion part is used to fuse the feature maps provided by the feature extraction part to obtain fused feature maps; among them, the feature fusion part is implemented by a multi-scale feature fusion network, and the multi-scale feature fusion network adopts a feature map fusion mechanism with different resolutions;

[0064] The target detection part is used to process the fused feature maps provided by the feature fusion part to obtain candidate regions and corresponding confidence scores; the target detection part is implemented by a region proposal network, and an adaptive learning rate adjustment mechanism is introduced during the bounding box regression process to improve the accuracy of bounding box localization;

[0065] The post-processing part is used to process the candidate regions and corresponding confidence scores provided by the target detection part to obtain the position and category information of the target object.

[0066] The feature extraction part is the basic part of the deep learning model. Its main task is to extract features from the preprocessed image data to generate feature maps. This part is implemented by an improved convolutional neural network (CNN), including multiple convolutional layers, pooling layers, and batch normalization layers, and separable convolutional layers are introduced to improve the computational efficiency and reduce the number of parameters. Specifically, the convolutional layer is responsible for extracting low-level features in the image, such as edges and textures; the pooling layer reduces the size of the feature map through downsampling operations while retaining important features; the batch normalization layer is used to normalize the output of the convolutional layer to stabilize the training process and accelerate convergence. The separable convolutional layer decomposes the standard convolution into depthwise convolution and pointwise convolution, significantly reducing the computational complexity while maintaining the performance of the model.

[0067] The role of the feature fusion part is to fuse the feature maps generated by the feature extraction part to generate a fused feature map. This part is implemented using a multi-scale feature fusion network (MSFN). By fusing feature maps of different resolutions, the feature representation ability is enhanced. In the specific implementation, the multi-scale feature fusion network adopts mechanisms such as skip connections and a feature pyramid network (FPN), enabling feature maps from different convolutional layers to be fused with each other, thereby capturing multi-scale object information in the image. Skip connections allow low-level features to be directly combined with high-level features, retaining more spatial information and details. The feature pyramid network further processes these fused feature maps to generate multi-scale fused feature maps, improving the model's detection ability for objects of different sizes.

[0068] The object detection part uses a region proposal network (RPN) to process the fused feature map, generating candidate regions and corresponding confidence scores. The region proposal network generates a series of candidate regions on the feature map by means of a sliding window and extracts the features of these candidate regions using a convolutional neural network. Then, the RPN classifies and regresses the bounding boxes for these features to determine whether each candidate region contains an object and gives the corresponding confidence score. To improve the accuracy of bounding box localization, the object detection part introduces an adaptive learning rate adjustment mechanism. During the bounding box regression process, the learning rate is dynamically adjusted according to the regression error, enabling the model to converge faster and more accurately, improving the localization accuracy of the bounding boxes.

[0069] The post-processing part receives the candidate regions and confidence scores from the object detection part and further processes them to finally obtain the position and category information of the target object. Post-processing usually includes non-maximum suppression (NMS) and classification confidence threshold filtering. Non-maximum suppression is used to eliminate highly overlapping candidate regions, retaining the regions most likely to contain the target object; classification confidence threshold filtering is used to remove candidate regions with low confidence. After post-processing, the system can output the final position and category information of the target object, ensuring the accuracy and reliability of the detection results.

[0070] Furthermore, the improved convolutional neural network specifically includes:

[0071] a) An input layer for receiving preprocessed image data, with a size of ;

[0072] b) A first convolutional layer that performs a convolution operation with a kernel size of and a stride of 1, with an output size of ;

[0073] c) A first batch normalization layer for performing batch normalization on the output of the first convolutional layer, maintaining an output size of ;

[0074] d) The first separable convolutional layer, including depthwise convolution and pointwise convolution. The depthwise convolution uses a convolutional kernel of size and a stride of 1, and the pointwise convolution uses convolution, with an output size of ;

[0075] e) The first max pooling layer, using a max pooling operation with a pooling kernel size of and a stride of 2, with an output size of

[0076] f) Repeat the structure from b) to e), gradually increasing the number of channels and reducing the size of the feature maps to form multi-level feature maps. The output size of each layer gradually decreases, and the number of feature maps gradually increases;

[0077] g) The output layer outputs multi-level feature maps, where the feature maps of each layer are different in terms of resolution and number of channels.

[0078] First, the input layer is used to receive the preprocessed image data, with the size of H×W×C, where H represents the height of the image, W represents the width of the image, and C represents the number of channels. The input layer ensures that the image data can smoothly enter the convolutional neural network for subsequent processing.

[0079] Next, the first convolutional layer performs preliminary processing on the input image data by using a 3×3 convolutional kernel and a convolutional operation with a stride of 1. This convolutional operation can effectively extract low-level features in the image such as edges and textures. After being processed by the first convolutional layer, the output feature map has a size of H×W×32, indicating that the original image size remains unchanged, but the number of channels increases to 32 to capture more feature information.

[0080] To improve the stability of the network and accelerate the training process, the output of the first convolutional layer will be batch-normalized through the first batch normalization layer. Batch normalization can standardize the mean and variance of the output feature map, thus avoiding the problem of internal covariate shift. After batch normalization, the output size remains H×W×32.

[0081] Then, the first separable convolutional layer further processes the batch-normalized feature map. The separable convolutional layer consists of depthwise convolution and pointwise convolution. The depthwise convolution uses a 3×3 convolutional kernel and a stride of 1 operation, and the pointwise convolution uses a 1×1 convolutional operation. The depthwise convolution performs convolution operations on each channel separately, retaining more spatial information, while the pointwise convolution increases the richness of feature expression by linearly combining channels. The output size of the first separable convolutional layer is H×W×64, further increasing the number of channels and capturing richer image features.

[0082] After that, the first max - pooling layer performs max - pooling operations on the output of the separable convolutional layer. This pooling layer uses a 2×2 pooling kernel and a stride of 2. Through downsampling, the size of the feature map is reduced by half. Max - pooling can extract the most significant features in the local area, reducing the size of the feature map while retaining important information. After being processed by the first max - pooling layer, the size of the output feature map is H / 2×W / 2×64.

[0083] To form multi - level feature maps, the above - mentioned convolutional, batch normalization, separable convolutional, and max - pooling operations are repeated in the network structure. Each layer gradually increases the number of channels and reduces the size of the feature map. This hierarchical structure enables the network to extract image features from different scales and levels, abstracting higher - level feature information layer by layer.

[0084] Finally, the output layer outputs multi - level feature maps, and the feature maps of each layer are different in resolution and the number of channels. These multi - level feature maps contain image features from low - level to high - level, and can provide rich information for subsequent object detection and recognition.

[0085] Through this structural design, the improved convolutional neural network can efficiently extract important features in the image, providing a solid foundation for object detection and tracking in the high - efficiency UAV image recognition system.

[0086] The following is the code of a reference implementation, which is used to illustrate the convolutional neural network in the feature extraction part of the high - efficiency UAV image recognition system.

[0087] import tensorflow as tf

[0088] from tensorflow.keras.layers import Input, Conv2D, DepthwiseConv2D,Conv2DTranspose, MaxPooling2D, BatchNormalization, Activation

[0089] from tensorflow.keras.models import Model

[0090] def build_feature_extractor(input_shape):

[0091] # Input layer, used to receive pre - processed image data, with size H×W×C

[0092] inputs = Input(shape=input_shape, name='input_layer')

[0093] # First convolutional layer, using a convolutional operation with a kernel size of 3×3 and a stride of 1, and the output size is H×W×32

[0094] x = Conv2D(filters=32, kernel_size=(3, 3), strides=1, padding='same',name='conv1')(inputs)

[0095] # First batch normalization layer, used to perform batch normalization on the output of the first convolutional layer, keeping the output size as H×W×32

[0096] x = BatchNormalization(name='bn1')(x)

[0097] # Activation function layer, using the ReLU activation function

[0098] x = Activation('relu', name='relu1')(x)

[0099] # First separable convolutional layer, including depthwise convolution and pointwise convolution

[0100] # Depthwise convolution uses a kernel size of 3×3 and a stride of 1

[0101] x = DepthwiseConv2D(kernel_size=(3, 3), strides=1, padding='same',name='dw_conv1')(x)

[0102] # Pointwise convolution uses a 1×1 convolution, and the output size is H×W×64

[0103] x = Conv2D(filters=64, kernel_size=(1, 1), strides=1, padding='same',name='pw_conv1')(x)

[0104] # Batch normalization and activation function

[0105] x = BatchNormalization(name='bn2')(x)

[0106] x = Activation('relu', name='relu2')(x)

[0107] # First max pooling layer, performing max pooling operation with a pooling kernel size of 2×2 and a stride of 2, and the output size is H / 2×W / 2×64

[0108] x = MaxPooling2D(pool_size=(2, 2), strides=2, padding='same', name='pool1')(x)

[0109] # Repeat convolution, batch normalization, depthwise convolution, and max pooling operations to form multi-level feature maps

[0110] # Second convolutional layer, with an output size of H / 2×W / 2×128

[0111] x = Conv2D(filters=128, kernel_size=(3, 3), strides=1, padding='same', name='conv2')(x)

[0112] x = BatchNormalization(name='bn3')(x)

[0113] x = Activation('relu', name='relu3')(x)

[0114] x = DepthwiseConv2D(kernel_size=(3, 3), strides=1, padding='same',name='dw_conv2')(x)

[0115] x = Conv2D(filters=128, kernel_size=(1, 1), strides=1, padding='same', name='pw_conv2')(x)

[0116] x = BatchNormalization(name='bn4')(x)

[0117] x = Activation('relu', name='relu4')(x)

[0118] x = MaxPooling2D(pool_size=(2, 2), strides=2, padding='same', name='pool2')(x)

[0119] # Third convolutional layer, output size is H / 4 × W / 4 × 256

[0120] x = Conv2D(filters=256, kernel_size=(3, 3), strides=1, padding='same', name='conv3')(x)

[0121] x = BatchNormalization(name='bn5')(x)

[0122] x = Activation('relu', name='relu5')(x)

[0123] x = DepthwiseConv2D(kernel_size=(3, 3), strides=1, padding='same',name='dw_conv3')(x)

[0124] x = Conv2D(filters=256, kernel_size=(1, 1), strides=1, padding='same', name='pw_conv3')(x)

[0125] x = BatchNormalization(name='bn6')(x)

[0126] x = Activation('relu', name='relu6')(x)

[0127] x = MaxPooling2D(pool_size=(2, 2), strides=2, padding='same', name='pool3')(x)

[0128] # Output layer, output multi-level feature maps

[0129] outputs = x

[0130] # Build the model

[0131] model = Model(inputs, outputs, name='feature_extractor')

[0132] return model

[0133] # Define the input image size H, W, C

[0134] input_shape = (256, 256, 3)

[0135] # Build the feature extraction model

[0136] model = build_feature_extractor(input_shape)

[0137] # Print the model structure

[0138] model.summary()

[0139] Furthermore, the multi-scale feature fusion network is specifically used for:

[0140] Obtain the multi-level feature maps provided by the feature extraction part;

[0141] Introduce skip connections between feature maps of different resolutions;

[0142] Apply an attention mechanism to the feature maps of different resolutions to calculate the importance weights of each feature map;

[0143] Perform weighted fusion on the feature maps of different resolutions and their corresponding weights, and output the fused feature map.

[0144] The multi-scale feature fusion network specifically implements the following steps. First, it obtains multi-level feature maps from the feature extraction part. These feature maps differ in resolution and number of channels, covering various image features from low-level to high-level. The acquisition of multi-level feature maps is based on the design of the convolutional neural network in the feature extraction part, which generates multiple feature maps through multiple layers of convolution, pooling, and batch normalization operations, and each feature map contains different levels of information.

[0145] Next, skip connections are introduced between feature maps of different resolutions. This skip connection is a commonly used method in deep learning networks for directly connecting non-adjacent layers. Specifically, the skip connection can combine high-level features of low resolution with low-level features of high resolution, thereby enhancing the richness and diversity of feature representation. By introducing skip connections, the model can retain more spatial information and details, thus improving the effect of feature fusion and the overall performance of the network.

[0146] To further improve the effect of feature fusion, the multi-scale feature fusion network also applies an attention mechanism to feature maps of different resolutions. The attention mechanism is a method of dynamic weighting used to calculate the importance weights of each feature map. In the specific implementation process, the attention mechanism will perform weighted processing on each feature map, and the size of the weight reflects the importance of the feature map in the current task. In this way, the model can automatically adjust the contributions of different feature maps, so that important features get higher weights during the fusion process, while unimportant features are weakened.

[0147] After calculating the importance weights of each feature map, the multi-scale feature fusion network performs weighted fusion on the feature maps of different resolutions and their corresponding weights. The process of weighted fusion is to perform weighted summation on each feature map according to the calculated weights, thereby generating a fused feature map. This fused feature map synthesizes information of different resolutions and levels, has stronger representation ability and higher distinguishability, and provides rich feature data for subsequent object detection and recognition tasks.

[0148] Through the above steps, the multi-scale feature fusion network realizes the effective fusion of the multi-level feature maps generated by the feature extraction part. The introduction of skip connections and the attention mechanism not only enhances the diversity and richness of feature representation, but also improves the model's ability to capture important features, thereby enhancing the overall performance of the UAV image high-efficiency recognition system. Such a design ensures that the system can make full use of the feature information at all levels when processing complex image tasks, achieving higher recognition accuracy and efficiency.

[0149] The following is an implementation code used to describe the multi-scale feature fusion part in the UAV image high-efficiency recognition system.

[0150] import tensorflow as tf

[0151] from tensorflow.keras.layers import Input, Conv2D, DepthwiseConv2D,MaxPooling2D, BatchNormalization, Activation, Add, GlobalAveragePooling2D,Multiply, Reshape, Lambda

[0152] from tensorflow.keras.models import Model

[0153] def build_multi_scale_feature_fusion(input_shapes):

[0154] # Create input layers to receive feature maps of different resolutions

[0155] inputs = [Input(shape=shape) for shape in input_shapes]

[0156] # Define skip connection

[0157] def skip_connection(input1, input2, filters):

[0158] # Perform 1x1 convolution on the inputs to match the number of channels

[0159] conv1 = Conv2D(filters, kernel_size=(1, 1), strides=1, padding='same')(input1)

[0160] conv2 = Conv2D(filters, kernel_size=(1, 1), strides=1, padding='same')(input2)

[0161] # Add the two inputs together

[0162] return Add()([conv1, conv2])

[0163] # Define attention mechanism

[0164] def attention_mechanism(input_feature):

[0165] # Use global average pooling layer to aggregate global spatial information of the feature map

[0166] gap = GlobalAveragePooling2D()(input_feature)

[0167] # Reshape the pooling result for pointwise convolution

[0168] gap = Reshape((1, 1, gap.shape[-1]))(gap)

[0169] # Generate attention weights using two pointwise convolutional layers

[0170] dense1 = Conv2D(gap.shape[-1] / / 4, kernel_size=(1, 1), strides=1,padding='same', activation='relu')(gap)

[0171] dense2 = Conv2D(gap.shape[-1], kernel_size=(1, 1), strides=1, padding='same', activation='sigmoid')(dense1)

[0172] # Weight the input feature map

[0173] return Multiply()([input_feature, dense2])

[0174] # Apply the attention mechanism to each input feature map

[0175] attention_outputs = [attention_mechanism(input_feature) for input_feature in inputs]

[0176] # Perform skip connection and fusion operations on feature maps of different resolutions

[0177] fused_features = attention_outputs[0]

[0178] for feature in attention_outputs[1:]:

[0179] fused_features = skip_connection(fused_features, feature, filters=fused_features.shape[-1])

[0180] # Output the final fused feature map

[0181] outputs = fused_features

[0182] # Build the model

[0183] model = Model(inputs, outputs, name='multi_scale_feature_fusion')

[0184] return model

[0185] Furthermore, the region proposal network is specifically used for:

[0186] Generating candidate regions on the fused feature map using a sliding window method;

[0187] Performing bounding box regression on each candidate region, adopting an adaptive learning rate adjustment mechanism, dynamically adjusting the learning rate according to the regression error, and improving the bounding box localization accuracy;

[0188] Calculating a confidence score for each candidate region to determine whether it contains the target object.

[0189] The region proposal network (RPN) is specifically designed to generate and evaluate candidate regions in order to determine the location and presence probability of the target object on the fused feature map. The specific implementation includes the following steps:

[0190] First, the region proposal network generates candidate regions on the fused feature map using a sliding window method. This means sliding a fixed-size window pixel by pixel on the feature map to extract a series of overlapping local regions. These windows can have different scales and aspect ratios to better adapt to the diversity of target objects in the image. The sliding window method ensures that every potential target location is considered, thus generating a large number of candidate regions that cover every part of the entire feature map.

[0191] After generating the candidate regions, the region proposal network performs bounding box regression on each candidate region. The purpose of bounding box regression is to accurately locate the position of the target object by adjusting the boundaries of the candidate regions. In this process, the RPN extracts features for each candidate region through a small convolutional neural network and predicts a more accurate bounding box. In particular, an adaptive learning rate adjustment mechanism is adopted to dynamically adjust the learning rate according to the regression error. This mechanism continuously adjusts the learning rate during the training process to converge to the optimal solution faster and improve the accuracy of bounding box localization. Specifically, when the regression error is large, the learning rate will automatically increase to speed up the learning; when the regression error is small, the learning rate will decrease to avoid over-adjustment. Through this adaptive mechanism, the bounding box regression can more accurately locate the target object.

[0192] Next, to determine whether each candidate region contains the target object, the Region Proposal Network (RPN) calculates a confidence score for each candidate region. This process involves a classification task, which classifies the features of each candidate region to determine whether it contains the target object. The confidence score is a value between 0 and 1, representing the probability that the candidate region contains the target object. The network processes the features of each candidate region through a classifier and outputs the confidence score. If the confidence score of a candidate region is higher than a preset threshold, it is considered that the region may contain the target object; otherwise, it is considered not to contain the target object.

[0193] The Region Proposal Network generates candidate regions through a sliding window, performs bounding box regression on these regions to improve the localization accuracy, and calculates the confidence score to determine whether they contain the target object. Through these steps, the RPN provides high-quality candidate regions for subsequent object detection and recognition, enabling the entire UAV image high-efficiency recognition system to identify and track the target object more accurately and efficiently.

[0194] The code does not provide an adaptive learning rate adjustment mechanism. The adaptive learning rate adjustment mechanism is usually implemented during the model training process. Common methods include using adaptive optimization algorithms (such as Adam, RMSprop) or using learning rate schedulers (such as ReduceLROnPlateau, LearningRateScheduler) to dynamically adjust the learning rate.

[0195] The following is the reference implementation code for the Region Proposal Network (RPN) in the object detection part, using the TensorFlow and Keras libraries.

[0196] import tensorflow as tf

[0197] from tensorflow.keras.layers import Conv2D, Input, Reshape

[0198] from tensorflow.keras.models import Model

[0199] from tensorflow.keras.optimizers import Adam

[0200] from tensorflow.keras.callbacks import ReduceLROnPlateau

[0201] def build_rpn_model(base_layers, num_anchors):

[0202] """

[0203] Build the Region Proposal Network (RPN) model

[0204] :param base_layers: Output of the feature extraction network, i.e., the fused feature map

[0205] :param num_anchors: Number of anchors per sliding window (combinations of different scales and aspect ratios)

[0206] :return: RPN model

[0207] """

[0208] # Convolutional layer for generating candidate regions

[0209] rpn_conv = Conv2D(512, (3, 3), padding='same', activation='relu',kernel_initializer='normal', name='rpn_conv')(base_layers)

[0210] # Bounding box regression layer, outputting the bounding box adjustment values (dx, dy, dw, dh) for each candidate region

[0211] rpn_cls_score = Conv2D(num_anchors * 2, (1, 1), activation='linear',kernel_initializer='uniform', name='rpn_cls_score')(rpn_conv)

[0212] # Confidence score layer, outputting the object existence probability for each candidate region

[0213] rpn_bbox_pred = Conv2D(num_anchors * 4, (1, 1), activation='linear',kernel_initializer='zero', name='rpn_bbox_pred')(rpn_conv)

[0214] # Reshape the classification results into the form of (num_anchors, 2)

[0215] rpn_cls_score_reshape = Reshape((-1, 2))(rpn_cls_score)

[0216] # Reshape the bounding box prediction results into the form of (num_anchors, 4)

[0217] rpn_bbox_pred_reshape = Reshape((-1, 4))(rpn_bbox_pred)

[0218] # Softmax layer to normalize the confidence scores

[0219] rpn_cls_prob = tf.keras.layers.Activation('softmax')(rpn_cls_score_reshape)

[0220] # Build the RPN model

[0221] model = Model(base_layers, [rpn_cls_prob, rpn_bbox_pred_reshape])

[0222] return model

[0223] # Define the size of the input feature map (e.g., 32x32x256)

[0224] input_shape = (32, 32, 256)

[0225] num_anchors = 9 # Assume there are 9 anchors per sliding window

[0226] # Build the input layer

[0227] inputs = Input(shape=input_shape)

[0228] # Build the RPN model

[0229] rpn_model = build_rpn_model(inputs, num_anchors)

[0230] # Compile the model, using the Adam optimizer and cross-entropy loss function

[0231] rpn_model.compile(optimizer=Adam(lr=1e-4),

[0232] loss=['categorical_crossentropy','mean_squared_error'])

[0233] # Set the adaptive learning rate adjustment mechanism, using ReduceLROnPlateau

[0234] reduce_lr = ReduceLROnPlateau(monitor='loss', factor=0.1, patience=10, min_lr=1e-6)

[0235] # Assume there are X_train and Y_train as training data

[0236] # X_train, Y_train_cls, Y_train_bbox are the feature map, classification label, and bounding box regression label respectively

[0237] # Here is just an example, no actual data is provided

[0238] X_train =...

[0239] Y_train_cls =...

[0240] Y_train_bbox =...

[0241] # Train the model

[0242] rpn_model.fit(X_train, [Y_train_cls, Y_train_bbox],

[0243] epochs=50,

[0244] batch_size=32,

[0245] callbacks=[reduce_lr])

[0246] Furthermore, the post-processing part is specifically used for:

[0247] Perform non-maximum suppression on the candidate regions, remove the candidate regions with high overlap and low confidence, and adopt a dynamic threshold adjustment mechanism to dynamically adjust the suppression threshold according to the confidence score;

[0248] Based on the result of non-maximum suppression processing, output the final position and class information of the target object.

[0249] The post-processing part is mainly used to perform non-maximum suppression processing on the candidate regions, and output the final position and class information of the target object according to the processing result. The specific implementation includes the following steps.

[0250] First, the post-processing part performs non-maximum suppression processing on the candidate regions. Non-maximum suppression (NMS) is a commonly used technique aimed at removing candidate regions with high overlap and low confidence. The specific operation process is that among all candidate regions, first select the region with the highest confidence as the benchmark, and then calculate the overlap between this region and all other candidate regions. For those regions whose overlap exceeds the preset threshold, if their confidence is low, they will be removed. This process is repeated until all regions with high overlap and low confidence are removed, thus retaining the most representative candidate regions.

[0251] To further optimize the effect of non-maximum suppression, the post-processing part adopts a dynamic threshold adjustment mechanism. This mechanism dynamically adjusts the suppression threshold according to the confidence score of the candidate regions. Specifically, when the confidence of a certain candidate region is high, the system will correspondingly lower the overlap threshold to ensure that this region can be retained; conversely, when the confidence of a certain candidate region is low, the system will increase the overlap threshold to increase the possibility of removal. Through this dynamic adjustment, the system can more effectively retain candidate regions with high confidence, while removing regions with low confidence and high overlap, further improving the accuracy and stability of target detection.

[0252] After completing the non-maximum suppression processing, the post-processing part will output the final position and class information of the target object according to the processing result. This step mainly further analyzes and processes the remaining candidate regions after non-maximum suppression processing to determine the specific position and class of the target object in each region. In the specific operation process, the system will combine the bounding box and confidence information provided by the previous target detection part to accurately locate each candidate region, and determine the class of the target object according to the output of the classifier. Finally, the system will generate a detailed target detection report, including the position information (such as bounding box coordinates) and class information (such as object class labels) of each target object.

[0253] Through these steps, the post-processing part can not only effectively remove redundant candidate regions, but also accurately determine the position and class of the target object, greatly improving the overall performance and practicality of the UAV image high-efficiency recognition system.

[0254] The following is a reference implementation code for the post - processing part in a deep - learning model, using the TensorFlow and Keras libraries.

[0255] import tensorflow as tf

[0256] import numpy as np

[0257] def non_max_suppression(boxes, scores, max_output_size, iou_threshold):

[0258] """

[0259] Non - maximum suppression (NMS) processing

[0260] :param boxes: Bounding box coordinates [num_boxes, 4]

[0261] :param scores: Confidence scores [num_boxes]

[0262] :param max_output_size: Maximum number of candidate regions to keep

[0263] :param iou_threshold: Overlap threshold (IOU threshold)

[0264] :return: Indices after NMS [max_output_size]

[0265] """

[0266] return tf.image.non_max_suppression(boxes, scores, max_output_size,iou_threshold)

[0267] def dynamic_threshold_nms(boxes, scores, base_iou_threshold = 0.5):

[0268] """

[0269] Dynamic threshold non - maximum suppression processing, dynamically adjusts the suppression threshold according to the confidence

[0270] :param boxes: Bounding box coordinates [num_boxes, 4]

[0271] :param scores: Confidence scores [num_boxes]

[0272] :param base_iou_threshold: Base intersection over union threshold (IOU threshold)

[0273] :return: Bounding boxes and confidence scores after NMS

[0274] """

[0275] sorted_indices = np.argsort(-scores)

[0276] boxes = boxes[sorted_indices]

[0277] scores = scores[sorted_indices]

[0278] keep_boxes = [ ]

[0279] keep_scores = [ ]

[0280] while len(scores) > 0:

[0281] current_box = boxes[0]

[0282] current_score = scores[0]

[0283] keep_boxes.append(current_box)

[0284] keep_scores.append(current_score)

[0285] if len(scores) == 1:

[0286] break

[0287] boxes = boxes[1:]

[0288] scores = scores[1:]

[0289] iou_threshold = base_iou_threshold (1 - current_score)

[0290] ious = compute_iou(current_box, boxes)

[0291] valid_indices = np.where(ious <= iou_threshold)[0]

[0292] boxes = boxes[valid_indices]

[0293] scores = scores[valid_indices]

[0294] return np.array(keep_boxes), np.array(keep_scores)

[0295] def compute_iou(box, boxes):

[0296] """

[0297] Compute the Intersection over Union (IOU) between a bounding box and multiple bounding boxes

[0298] :param box: Bounding box coordinates [4]

[0299] :param boxes: Multiple bounding box coordinates [num_boxes, 4]

[0300] :return: Intersection over Union [num_boxes]

[0301] """

[0302] y1 = np.maximum(box[0], boxes[:, 0])

[0303] x1 = np.maximum(box[1], boxes[:, 1])

[0304] y2 = np.minimum(box[2], boxes[:, 2])

[0305] x2 = np.minimum(box[3], boxes[:, 3])

[0306] inter_area = np.maximum(0, y2 - y1) np.maximum(0, x2 - x1)

[0307] box_area = (box[2] - box[0]) (box[3] - box[1])

[0308] boxes_area = (boxes[:, 2] - boxes[:, 0]) (boxes[:, 3] - boxes[:,1])

[0309] iou = inter_area / (box_area + boxes_area - inter_area)

[0310] return iou

[0311] def post_process(rpn_boxes, rpn_scores, class_predictions, max_output_size=100, base_iou_threshold=0.5):

[0312] """

[0313] The post - processing part processes the candidate regions and confidence scores provided by the RPN, and outputs the final target positions and class information

[0314] :param rpn_boxes: The bounding boxes of candidate regions generated by the RPN [num_boxes, 4]

[0315] :param rpn_scores: The confidence scores of candidate regions generated by the RPN [num_boxes]

[0316] :param class_predictions: The class information output by the classifier [num_boxes, num_classes]

[0317] :param max_output_size: The maximum number of candidate regions to retain

[0318] :param base_iou_threshold: The base intersection - over - union threshold (IOU threshold)

[0319] :return: The final target object positions and class information

[0320] """

[0321] # Perform dynamic threshold non - maximum suppression on candidate regions

[0322] nms_boxes, nms_scores = dynamic_threshold_nms(rpn_boxes, rpn_scores, base_iou_threshold)

[0323] # Obtain the class information for each candidate region

[0324] nms_class_predictions = class_predictions[:len(nms_boxes)]

[0325] final_classes = np.argmax(nms_class_predictions, axis=1)

[0326] # Output the final object positions and class information

[0327] final_boxes = nms_boxes

[0328] final_scores = nms_scores

[0329] final_class_labels = final_classes

[0330] return final_boxes, final_scores, final_class_labels

[0331] # Example data (assuming obtained from RPN and classifier)

[0332] rpn_boxes = np.array([[50, 50, 150, 150], [55, 60, 155, 160], [100, 100, 200, 200]])

[0333] rpn_scores = np.array([0.9, 0.75, 0.6])

[0334] class_predictions = np.array([[0.1, 0.9], [0.2, 0.8], [0.8, 0.2]])

[0335] # Post - processing to obtain the final object positions and class information

[0336] final_boxes, final_scores, final_class_labels = post_process(rpn_boxes, rpn_scores, class_predictions)

[0337] print("Final target locations:", final_boxes)

[0338] print("Final confidence scores:", final_scores)

[0339] print("Final class labels:", final_class_labels)

[0340] To train the deep learning model in the high-efficiency drone image recognition system, specific steps need to be followed. This model includes a feature extraction part, a feature fusion part, an object detection part, and a post-processing part. The following are the detailed training steps.

[0341] First, prepare the dataset. The dataset should contain drone images in various scenarios, with the bounding boxes and class information of the target objects annotated in the images. The dataset should be divided into a training set and a validation set to evaluate the performance of the model.

[0342] Next, construct the feature extraction part of the deep learning model. The feature extraction part uses an improved convolutional neural network, including multiple convolutional layers, pooling layers, and batch normalization layers, and introduces separable convolutional layers. Deep learning frameworks such as TensorFlow or PyTorch can be used to implement this part. The purpose of the feature extraction part is to extract features from the preprocessed image data input and generate feature maps.

[0343] Then, construct the feature fusion part. The feature fusion part is implemented using a multi-scale feature fusion network, which fuses feature maps of different resolutions from the feature extraction part. By introducing skip connections and attention mechanisms, feature maps of different resolutions can be better fused to generate a fused feature map containing multi-level information.

[0344] On this basis, construct the object detection part. The object detection part uses a Region Proposal Network (RPN) to process the fused feature map, generating candidate regions and corresponding confidence scores. The RPN generates candidate regions on the feature map through a sliding window method and calculates the bounding box regression and confidence scores for each candidate region. To improve the accuracy of bounding box localization, an adaptive learning rate adjustment mechanism is introduced to dynamically adjust the learning rate according to the regression error.

[0345] Finally, construct the post - processing part. The post - processing part processes the candidate regions and confidence scores provided by the object detection part. First, perform non - maximum suppression to remove candidate regions with high overlap and low confidence, and adopt a dynamic threshold adjustment mechanism. According to the results of non - maximum suppression, output the final position and class information of the target object.

[0346] The training process is as follows:

[0347] 1. Data pre - processing: Standardize the image data, scale the images to a unified size, and perform data augmentation, such as random cropping, rotation, and flipping.

[0348] 2. Initialize model parameters: Use pre - trained models (such as ResNet, VGG, etc.) to initialize the parameters to accelerate the training convergence speed.

[0349] 3. Define loss functions: Include classification loss and regression loss. The classification loss is used to evaluate the confidence scores of the object detection part, and the regression loss is used to evaluate the accuracy of bounding box regression.

[0350] 4. Select an optimizer: Use Adam or SGD optimizer and set the initial learning rate. Introduce an adaptive learning rate adjustment mechanism to dynamically adjust the learning rate according to the loss change during training.

[0351] 5. Iterative training: In each training iteration, perform the following steps:

[0352] Forward propagation: Pass the input image data through the feature extraction part, feature fusion part, and object detection part to generate candidate regions and confidence scores.

[0353] Calculate the loss: Calculate the classification loss and regression loss according to the predicted candidate regions and confidence scores.

[0354] Backward propagation: Update the model parameters based on the loss to optimize the model.

[0355] 6. Validate the model: After each training iteration, use the validation set to evaluate the performance of the model, including the accuracy of object detection and the precision of bounding box localization. According to the validation results, adjust the model parameters and training strategy.

[0356] 7. Save the model: Regularly save the model parameters during training for subsequent use and further optimization.

[0357] The target tracking unit 105 is used to execute an adaptive target tracking algorithm to lock and track the target object identified by the target detection unit. Among them, the adaptive target tracking algorithm is implemented based on the fusion of the Kalman filter and the optical flow method, and is used to provide stable target tracking in the case of high - speed flight.

[0358] The target tracking unit 105 plays a crucial role in the high-efficiency UAV image recognition system. Its main function is to lock and continuously track the target object identified by the target detection unit 104. This unit relies on an adaptive target tracking algorithm that combines the Kalman filter and the optical flow method to ensure stable target tracking under the condition of high-speed UAV flight.

[0359] First, the target tracking unit receives the position and motion information of the target object identified by the target detection unit. As one of the core components, the Kalman filter is mainly used to predict and correct the state of the target object. Through a series of mathematical operations, the Kalman filter uses the state estimate of the previous moment and the current observation data to predict the position and speed of the target object at the next moment. This prediction process includes two steps: state prediction and update. In the state prediction step, the Kalman filter predicts the position and speed of the target object at the next moment according to its motion model. In the update step, the prediction value is corrected using the observation data at the current moment to obtain a more accurate state estimate.

[0360] At the same time, the optical flow method is used to calculate the motion vector of the target object on the image plane. The optical flow method determines the motion direction and speed of the target object by comparing the pixel changes in the image sequence. This method is particularly suitable for processing continuous image data and can provide high-precision motion detection results. In the specific implementation, the optical flow method calculates the displacement vector of the target object by comparing the pixel intensity changes in adjacent frames, thereby providing real-time motion information.

[0361] To enhance the stability and accuracy of tracking, the target tracking unit combines the prediction result of the Kalman filter with the motion vector of the optical flow method. This fusion method uses the state estimate of the Kalman filter to provide a macroscopic prediction, and at the same time uses the fine motion detection of the optical flow method to correct the prediction error, ensuring that the tracking algorithm can maintain high-efficiency and stable tracking performance under various complex flight conditions.

[0362] In addition, the target tracking unit also has an adaptive adjustment function, which can dynamically adjust the weights of the Kalman filter and the optical flow method according to the actual situation. For example, when the lighting condition is good and the image quality is high, the system can increase the dependence on the result of the optical flow method; while when the image noise is large or the lighting condition is poor, the system can rely more on the prediction result of the Kalman filter. In this way, the target tracking unit can flexibly adjust in various environments to ensure the best tracking effect.

[0363] Through the above process, the target tracking unit 105 can continuously and stably track the target object, regardless of whether the UAV is flying at high speed or the environmental conditions are complex. This design not only improves the overall performance of the UAV system but also significantly enhances its reliability and practicality in actual applications.

[0364] Furthermore, the adaptive target tracking algorithm in the target tracking unit includes the following steps:

[0365] Receive the initial position information of the target object from the target detection unit , where and represent the initial position of the target object, and represent the initial width and initial height of the target object;

[0366] In each frame of the image, according to the image data acquired by the image acquisition unit, use the optical flow method to calculate the motion vector of the target object on the image plane , where and represent the displacement information of the target object between the current frame and the previous frame;

[0367] Use the initial position information and motion vector of the target object, and adopt an adaptive Kalman filter to predict and update the position information of the target object. Among them, the state vector and the observation vector are calculated according to the following formulas 1 and 2 respectively:

[0368] ;

[0369] ;

[0370] Among them, represents the state vector at time , including the position, velocity, and acceleration information of the target object; is the state transition matrix, used to describe the motion model of the target object; is the control matrix, used to describe the influence of the external control input on the motion of the target object; is the process noise, following a Gaussian distribution with a mean of zero; is the observation vector, including the displacement information of the target object obtained from the optical flow method; is the observation matrix, used to map the state vector to the observation space; is the observation noise, following a Gaussian distribution with a mean of zero;

[0371] is the adaptive process noise adjustment coefficient, which is calculated by the following formula 3:

[0372] ;

[0373] where, is the predetermined weight coefficient;

[0374] is the adaptive observation noise adjustment coefficient, which is calculated by the following formula 4:

[0375] ;

[0376] where, is the predetermined weight coefficient;

[0377] Through the prediction and correction steps of the adaptive Kalman filter, the updated position information of the target object in the current frame is calculated and output , where and represent the updated position of the target object, and represent the updated width and updated height of the target object;

[0378] According to the state update result of the adaptive Kalman filter, the motion trajectory of the target object is generated, and the target object is locked and continuously tracked in the image.

[0379] First, the target tracking unit receives the initial position information of the target object from the target detection unit , where and represent the initial abscissa and ordinate positions of the target object, and represent the initial width and initial height of the target object. This information provides the initial position and size of the target object in the image for the algorithm.

[0380] In each frame of the image, the target tracking unit calculates the motion vector of the target object on the image plane using the optical flow method according to the image data acquired by the image acquisition unit . Where, and represent the displacement information of the target object between the current frame and the previous frame. The optical flow method is a commonly used computer vision technique that calculates the motion vector of the target object by analyzing the motion of pixels in the image sequence.

[0381] Next, using the initial position information and motion vectors of the target object, an adaptive Kalman filter is employed to predict and update the position information of the target object. The Kalman filter is a recursive algorithm that provides an optimal estimate of the system state by combining predicted and observed data. The adaptive Kalman filter performs prediction and update through the following steps:

[0382] 1. State vector and observation vector are calculated according to equations (1) and (2) respectively:

[0383] ;

[0384] ;

[0385] where represents the state vector at time k, including the position, velocity, and acceleration information of the target object. is the state transition matrix, which is used to describe the motion model of the target object, is the control matrix, which is used to describe the influence of the external control input on the motion of the target object, is the process noise, which follows a Gaussian distribution with a mean of zero.

[0386] is the observation vector, including the displacement information of the target object obtained from the optical flow method. is the observation matrix, which is used to map the state vector to the observation space, is the observation noise, which follows a Gaussian distribution with a mean of zero.

[0387] 2. Adaptive process noise adjustment coefficient is calculated through equation (3):

[0388] ;

[0389] where γ is a predetermined weight coefficient. This equation dynamically adjusts the magnitude of the process noise based on the difference between the state vector at the previous time step and the state vector at the two previous time steps.

[0390] 3. Adaptive observation noise adjustment coefficient is calculated through equation (4):

[0391] ;

[0392] where is a predetermined weight coefficient. This equation is based on the previous time step's observation vector and the previous time step's state vector The difference in mapping is used to dynamically adjust the magnitude of the observation noise.

[0393] Through the prediction and correction steps of the adaptive Kalman filter, the target tracking unit calculates and outputs the updated position information of the target object in the current frame. . Among them and represent the updated position of the target object, and represent the updated width and updated height of the target object.

[0394] Based on the state update result of the adaptive Kalman filter, the target tracking unit generates the motion trajectory of the target object, and locks and continuously tracks the target object in the image. This process ensures the stable tracking of the target object under the condition of high-speed flight of the UAV, and improves the overall performance and reliability of the system.

[0395] The control unit 106 adjusts the flight trajectory of the UAV according to the position and motion information of the target object provided by the target tracking unit to maintain continuous tracking of the target object.

[0396] The control unit 106 plays a crucial role in the high-efficiency recognition system of UAV images. Its main function is to dynamically adjust the flight trajectory of the UAV according to the position and motion information of the target object provided by the target tracking unit 105 to ensure continuous tracking of the target object. This unit combines real-time data processing and intelligent control algorithms to achieve precise trajectory adjustment and target tracking of the UAV under high-speed flight conditions.

[0397] First of all, the control unit receives real-time data from the target tracking unit. This data includes detailed information such as the current position, motion direction, and speed of the target object. To ensure the immediate adjustment of the flight trajectory, the control unit uses a high-speed processor and efficient algorithms to quickly process and analyze the received data. These algorithms include prediction algorithms and optimization algorithms, which can calculate the optimal flight path of the UAV in advance according to the motion trend of the target object.

[0398] The control unit uses prediction algorithms to estimate the future position of the target object. These algorithms are usually based on Kalman filtering or particle filtering techniques, and combine the historical motion data and current motion state of the target object to predict its position in the future for a period of time. This prediction function enables the control unit to make flight trajectory adjustments in advance, improving the response speed and tracking accuracy of the UAV.

[0399] After determining the predicted position of the target object, the control unit applies an optimization algorithm to calculate the optimal flight trajectory of the UAV. The optimization algorithm takes into account the flight performance, speed, acceleration, and environmental factors of the UAV to ensure that the calculated trajectory can not only accurately track the target object but also maximize energy and time savings. Typical optimization algorithms include algorithms, dynamic programming algorithms, genetic algorithms, etc. These algorithms can efficiently find the optimal path in complex environments.

[0400] After calculating the optimal flight trajectory, the control unit adjusts the attitude and motion parameters of the UAV through the flight control system. The flight control system includes multiple actuators, such as motors, servos, and thrusters. These actuators receive instructions from the control unit and adjust the heading, speed, and altitude of the UAV. Specifically, the control unit adjusts the power of the motor and the angle of the servo to make the UAV fly along the calculated optimal trajectory. To ensure the smoothness and accuracy of the adjustment process, the flight control system is also equipped with sensors such as an inertial measurement unit (IMU) and GPS. These sensors real-time monitor the attitude and position of the UAV and feedback to the control unit for closed-loop control.

[0401] In addition, the control unit also has an adaptive adjustment function, which can dynamically adjust the control parameters according to the real-time flight state and environmental changes. For example, when the UAV encounters sudden strong winds or other obstacles, the control unit can immediately recalculate the flight trajectory and adjust the flight attitude of the UAV to avoid collisions and other accidents. This adaptive ability enables the UAV to maintain high-efficiency target tracking capabilities in various complex and dynamic environments.

[0402] Through the above design and implementation, the control unit 106 ensures that the UAV can continuously and stably track the target object during high-speed flight, improving the overall performance and reliability of the UAV system.

[0403] Furthermore, the control unit is specifically used for:

[0404] Receiving the position and motion information of the target object provided by the target tracking unit;

[0405] Based on the received position and motion information of the target object, calculating the optimal flight path of the UAV in real-time, and using algorithms to ensure the optimality and safety of the path;

[0406] Generating corresponding attitude adjustment instructions according to the calculated optimal flight path to control the heading, speed, and altitude of the UAV.

[0407] The control unit is specifically used for the following tasks. First, the control unit receives the position and motion information of the target object provided by the target tracking unit. The target tracking unit continuously monitors the dynamics of the target object and updates its position and motion data in real time. After the control unit obtains this data, it provides the basic information for subsequent path calculation and flight control.

[0408] Next, based on the received position and motion information of the target object, the control unit calculates the optimal flight path of the UAV in real time. To ensure the optimality and safety of the path, the control unit uses an algorithm for path planning. The algorithm is a graph search algorithm, which is widely used in path planning and navigation problems. It finds the path with the minimum cost by moving between nodes and evaluating the cost of each node (including the actual cost from the starting point to the current node and the estimated cost from the current node to the target node). Specifically, the control unit first takes the current position of the UAV and the position of the target object as the starting point and the end point, and establishes a graph containing all possible flight paths. During the search process, the control unit evaluates the cost of each path, selects the path with the minimum cost and avoids obstacles, ensuring that the UAV approaches the target object in an optimal manner.

[0409] After calculating the optimal flight path, the control unit generates corresponding attitude adjustment instructions according to this path. These instructions are used to control the heading, speed and altitude of the UAV. In the specific implementation process, the control unit decomposes the optimal path into a series of refined flight steps and calculates the required attitude adjustment according to each step. For example, to adjust the heading, the control unit generates corresponding turning instructions; to adjust the speed, the control unit generates acceleration or deceleration instructions; to adjust the altitude, the control unit generates ascending or descending instructions. All these instructions are transmitted to the flight control system of the UAV to adjust the flight attitude of the UAV in real time.

[0410] Through these steps, the control unit ensures that the UAV can approach and continuously track the target object along the optimal path. Receiving the position and motion information of the target object, calculating the optimal flight path, and generating attitude adjustment instructions, these steps are closely linked, enabling the UAV to perform tasks stably and efficiently in a complex and dynamic environment.

[0411] Although this application is disclosed above with preferred embodiments, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application should be determined by the scope defined by the claims of this application.

Claims

1. A high-efficiency drone image recognition system, characterized in that: include: UAV, for high-speed and stable flight; An image acquisition unit, including a high-resolution camera placed on the drone to capture image data in real time during flight; An image preprocessing unit, used for performing preprocessing operations including denoising and deblurring on the image data captured by the image acquisition unit to obtain preprocessed image data; A target detection unit, comprising a pre-trained deep learning model, for identifying a target object in the pre-processed image provided by the image pre-processing unit, wherein the deep learning model is optimized based on a convolutional neural network structure to improve recognition speed and accuracy; A target tracking unit, used for executing an adaptive target tracking algorithm to lock and track the target object identified by the target detection unit, wherein the adaptive target tracking algorithm is implemented based on the fusion of Kalman filter and optical flow method, and is used for providing stable target tracking under high-speed flight conditions; A control unit, which adjusts the flight trajectory of the UAV according to the position and motion information of the target object provided by the target tracking unit to maintain continuous tracking of the target object; The adaptive target tracking algorithm in the target tracking unit includes the following steps: Receive the initial position information of the target object from the target detection unit ,in and represents the initial position of the target object, and Indicates the initial width and initial height of the target object; In each frame of the image, the displacement vector of the target object on the image plane is calculated using the optical flow method based on the image data obtained by the image acquisition unit. ,in and Respectively represent the target object in the current frame and the previous frame Axis and Displacement in the axial direction; Using the initial position information and motion vector of the target object, an adaptive Kalman filter is used to predict and update the position information of the target object, wherein the state vector of the adaptive Kalman filter is and the observation vector Calculate according to the following formula 1 and formula 2 respectively: ; ; in, Indicates at time The state vector includes the position, velocity and acceleration information of the target object; is the state transfer matrix, which is used to describe the motion model of the target object; is the control matrix, which is used to describe the external control input Effects on the motion of the target object; is process noise, which follows a Gaussian distribution with a mean of zero; is the observation vector, including the displacement information of the target object obtained from the optical flow method; is the observation matrix, which is used to map the state vector to the observation space; is the observation noise, which follows a Gaussian distribution with a mean of zero; is the adaptive process noise adjustment factor, calculated using the following formula 3: ; in, is the predetermined weight coefficient is the adaptive observation noise adjustment factor, which is calculated by the following formula 4: ; in, is a predetermined weight coefficient; Through the prediction and correction steps of the adaptive Kalman filter, the updated position information of the target object in the current frame is calculated and output ,in and represents the updated position of the target object, and Indicates the updated width and updated height of the target object; According to the state update result of the adaptive Kalman filter, the motion trajectory of the target object is generated, and the target object is locked and continuously tracked in the image.

2. The high-efficiency drone image recognition system according to claim 1 is characterized in that: The deep learning model includes a feature extraction part, a feature fusion part, a target detection part and a post-processing part; The feature extraction part is used to extract features from the preprocessed image data to obtain a feature map; wherein the feature extraction part is implemented by an improved convolutional neural network, the improved convolutional neural network includes multiple convolutional layers, pooling layers and batch normalization layers, and introduces a separable convolutional layer; The feature fusion part is used to fuse the feature map provided by the feature extraction part to obtain a fused feature map; wherein the feature fusion part is implemented by a multi-scale feature fusion network, and the multi-scale feature fusion network adopts a fusion mechanism of feature maps with different resolutions; The object detection part is used to process the fused feature map provided by the feature fusion part to obtain candidate regions and corresponding confidence scores; the object detection part is implemented by a region proposal network, and an adaptive learning rate adjustment mechanism is introduced in the bounding box regression process to improve the accuracy of bounding box positioning; The post-processing part is used to process the candidate areas and corresponding confidence scores provided by the target detection part to obtain the position and category information of the target object.

3. The high-efficiency drone image recognition system according to claim 2 is characterized in that: The improved convolutional neural network specifically includes: a) Input layer, used to receive preprocessed image data, size is ,in, Represents the height of the image, Represents the width of the image, Represents the number of channels; b) The first convolutional layer uses a convolution kernel size of , a convolution operation with a stride of 1, the output size is ; c) The first batch normalization layer is used to batch normalize the output of the first convolutional layer to keep the output size ; d) The first separable convolution layer includes depthwise convolution and pointwise convolution. The depthwise convolution uses a convolution kernel size of , stride is 1, point-by-point convolution uses Convolution, the output size is ; e) The first maximum pooling layer uses a pooling size of , a maximum pooling operation with a stride of 2, and an output size of ; f) Repeat the structures from b) to e), gradually increase the number of channels and reduce the size of feature maps, to form a multi-level feature map, the output size of each layer gradually decreases, and the number of feature maps gradually increases; g) Output layer, which outputs multi-level feature maps, where the feature maps of each layer are different in resolution and number of channels.

4. The high-efficiency drone image recognition system according to claim 2, characterized in that: The multi-scale feature fusion network is specifically used for: Obtaining a multi-level feature map provided by the feature extraction part; Introduce skip connections between feature maps of different resolutions; Apply attention mechanism to feature maps of different resolutions and calculate the importance weight of each feature map; The feature maps of different resolutions and their corresponding weights are weighted fused to output the fused feature map.

5. The high-efficiency drone image recognition system according to claim 2, characterized in that: The region proposal network is specifically used for: Generate candidate regions on the fused feature map using a sliding window method; Perform bounding box regression on each candidate region and use an adaptive learning rate adjustment mechanism to dynamically adjust the learning rate according to the regression error to improve the bounding box positioning accuracy; Calculate the confidence score for each candidate region to determine whether it contains the target object.

6. The high-efficiency drone image recognition system according to claim 2, characterized in that: The post-processing part is specifically used for: Perform non-maximum suppression on candidate regions to remove candidate regions with high overlap and low confidence, and adopt a dynamic threshold adjustment mechanism to dynamically adjust the suppression threshold according to the confidence score; According to the results of non-maximum suppression processing, the final position and category information of the target object are output.

7. The high-efficiency drone image recognition system according to claim 1, characterized in that: The image acquisition unit uses a high-resolution camera to capture continuous image frames in high-speed flight. The resolution of the high-resolution camera is not less than 4K, and the frame rate is not less than 30 frames per second.

8. The high-efficiency drone image recognition system according to claim 1, characterized in that: The image preprocessing unit is specifically used for: Perform noise suppression on image data and use adaptive median filtering algorithm to remove random noise in the image; Using motion estimation algorithms, image data is deblurred to restore motion blur caused by high-speed flight; The histogram equalization technique is applied to enhance the contrast and details of the image to improve the recognition accuracy of the subsequent target detection unit.

9. The high-efficiency drone image recognition system according to claim 1, characterized in that: The control unit is specifically used for: Receiving the target object position and motion information provided by the target tracking unit; Based on the received target object position and motion information, the optimal flight path of the drone is calculated in real time. Algorithms to ensure optimality and safety of paths; Based on the calculated optimal flight path, corresponding attitude adjustment instructions are generated to control the heading, speed and altitude of the UAV.

Citation Information

Patent Citations

  • Tracking algorithm and tracking system for taking-off and landing of aircraft based on tripod head and camera head

    CN102043964A

  • Aircraft-based target recognition and tracking method and equipment, and storage equipment

    CN108614572A