Unmanned aerial vehicle image recognition and analysis system based on deep learning
By using a deep learning-based UAV image recognition and analysis system, and employing techniques such as adaptive image scaling layers, dynamic convolutional layers, batch normalization layers, and attention mechanism layers, the model training process is optimized. This solves the recognition problem of UAV image recognition systems in small targets and complex scenes, achieving higher recognition accuracy and expanding application scenarios.
Patent Information
- Application Number
- CN202511263050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-14
AI Technical Summary
Existing UAV image recognition systems suffer from insufficient accuracy in recognizing small targets and low accuracy in complex scenarios, which seriously affects their practicality.
A deep learning-based UAV image recognition and analysis system is adopted, including a data acquisition and preprocessing module, a deep learning model building module, a model training module, an image recognition and analysis module, and a result output and feedback module. Through adaptive image scaling layer, dynamically adjusted parameter convolutional layer, batch normalization layer, attention mechanism layer, and multi-task output layer, combined with the weighted sum of cross-entropy loss function and mean squared error loss function, stochastic gradient descent algorithm and adaptive learning rate adjustment strategy, the model is optimized to improve recognition accuracy and precision.
It improves the accuracy of small target recognition and recognition rate in complex scenarios, enhances the practicality and flexibility of the system, supports multi-task output, and improves the user experience.
Smart Images

Figure CN120953853A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition and analysis technology, and more specifically, to a deep learning-based image recognition and analysis system for unmanned aerial vehicles (UAVs). Background Technology
[0002] With the rapid development of UAV technology and remote sensing technology, UAV image recognition and analysis are increasingly being used in fields such as agricultural monitoring, environmental exploration, security patrol, and traffic management.
[0003] High-resolution images acquired by drones enable rapid location, classification, and tracking of ground targets, providing crucial data support for decision-making.
[0004] However, UAV images typically cover a wide area and depict complex scenes. Existing technologies suffer from insufficient accuracy in small target recognition and low accuracy in complex scenarios, severely limiting the practicality of UAV image analysis systems. Therefore, we propose a deep learning-based UAV image recognition and analysis system. Summary of the Invention
[0005] The purpose of this invention is to provide a deep learning-based UAV image recognition and analysis system, which aims to solve the problems of insufficient accuracy in small target recognition and low recognition accuracy in complex scenarios in existing technologies.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a deep learning-based UAV image recognition and analysis system, which includes a data acquisition and preprocessing module, a deep learning model construction module, a model training module, an image recognition and analysis module, and a result output and feedback module; The data acquisition and preprocessing module is used to acquire UAV images under different time, weather, and lighting conditions, and sequentially perform bilateral filtering for noise reduction and histogram equalization enhancement on the acquired images. Then, the UAV images are divided into training set, validation set, and test set in an 80:15:5 ratio. The deep learning model building module is used to build a deep learning model, which includes an adaptive image scaling layer, a convolutional layer that can dynamically adjust parameters, a batch normalization layer, a ReLU activation function, an attention mechanism layer, and an output layer that can perform multi-task output. The model training module is used to train a deep learning model using the training set. It uses the weighted sum of the cross-entropy loss function and the mean squared error loss function as the total loss function. It optimizes the deep learning model by combining the stochastic gradient descent algorithm with an adaptive learning rate adjustment strategy, and adjusts the hyperparameters based on the evaluation results of the validation set. The image recognition and analysis module is used to receive images collected in real time by the UAV and input them into a trained deep learning model to obtain target category information in the UAV images. The result output and feedback module is used to visualize the target category information obtained by the image recognition and analysis module from the UAV image, and to establish a feedback mechanism to receive error samples from users and retrain the deep learning model.
[0007] Preferably, the formula for the bilateral filtering algorithm used in the data acquisition and preprocessing module is as follows: ; In the formula, The original image is in Pixel value at that location, These are the pixel values after noise reduction. These are normalized weights, used to balance the contributions of pixels at different locations. It is a spatial domain Gaussian kernel function used to measure the impact of spatial distance between pixels on filtering. It is a range Gaussian kernel function used to measure the impact of pixel value differences on filtering.
[0008] Preferably, the formula for the histogram equalization algorithm used in the data acquisition and preprocessing module is as follows: ; In the formula, It is the transformed grayscale level. It is grayscale. Number of times it appears It is the total number of pixels in the image. Represents grayscale level The probability of occurrence It represents the total number of gray levels.
[0009] Preferably, in the deep learning model construction module, the normalization formula used for the batch normalization layer is: ; In the formula, This is the normalized output. It is the first in the input batch data One data point, It is the mean of this batch of data. It is the variance of this batch of data. It is a small constant to prevent the denominator from being zero. and These are learnable parameters.
[0010] Preferably, in the deep learning model construction module, the formula for calculating the attention weight vector of the attention mechanism layer is:
[0011] In the formula, It is the vector obtained after global average pooling of the input feature map, used to aggregate global information from the feature map. and This is the weight matrix of the fully connected layer, used for dimensionality reduction and dimensionality expansion of the aggregated information. It is the ReLU activation function, which increases the nonlinear transformation capability. It is the Sigmoid activation function, which maps the output to the interval [0, 1]. It is the attention weight vector, used to weight and adjust different channels of the input feature map.
[0012] Preferably, in the deep learning model construction module, the output layer capable of multi-task output uses a Softmax classifier for the target recognition task, and the classification calculation formula is: ; In the formula, It is input Category The probability, It is a model for categories The output score, The formula represents the total number of categories. It converts the scores output by the model into a probability distribution, ensuring that the probability value of each category is in the interval [0, 1] and the sum of the probabilities of all categories is 1. This makes it easier to intuitively determine the probability that the input image belongs to each category.
[0013] Preferably, in the model training module, the formula for the total loss function is:
[0014]
[0015]
[0016] In the formula, It is the total loss function. It is the cross-entropy loss function. It's a real label. The model predicts that the sample belongs to a category. The probability is used to measure the difference between the predicted class and the true class in a classification task. It is the mean squared error loss function. It is the sample size. These are the actual bounding box coordinates. These are the bounding box coordinates predicted by the model. and It is the weighting coefficient.
[0017] Preferably, in the model training module, the formula for combining the stochastic gradient descent algorithm with the adaptive learning rate adjustment strategy is as follows:
[0018] In the formula, These are the updated model parameters. These are the current model parameters. It is the initial learning rate. Is it up to the The sum of squared historical gradients up to the last step is used to reflect the historical status of parameter updates. It is a small constant to prevent the denominator from being zero. It is the gradient of the loss function under the current parameters.
[0019] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention collects UAV images under different time, weather, and lighting conditions through a data acquisition and preprocessing module, ensuring data coverage of diverse and complex scenarios and providing comprehensive training samples for the model. The adaptive image scaling layer in the deep learning model building module reduces information loss caused by size differences, the convolutional layer that can dynamically adjust parameters improves the targeting of feature extraction, the attention mechanism layer highlights key features and suppresses interference information, and the multi-task output layer enhances the ability to recognize targets. At the same time, the model training module uses the weighted sum of the cross-entropy loss function and the mean squared error loss function as the total loss function, combined with an adaptive learning rate adjustment strategy to optimize the model, thereby improving the model's recognition accuracy for small targets and overall recognition accuracy in complex scenarios.
[0020] 2. In the data preprocessing stage, this invention scientifically divides the training set, validation set, and test set into an 80:15:5 ratio, which are used for model training, hyperparameter tuning, and performance evaluation, respectively. The batch normalization layer stabilizes the distribution of input data, accelerates model convergence, and reduces internal covariate bias. The stochastic gradient descent algorithm combined with an adaptive learning rate adjustment strategy dynamically optimizes the model parameters. Hyperparameters are adjusted based on the validation set results to avoid overfitting, enabling the model to have stable recognition performance on unknown data, making the training process more efficient and stable.
[0021] 3. In the result output and feedback module, this invention provides visual output of target category information, making it convenient for users to intuitively obtain results. The established feedback mechanism can receive erroneous samples from users and reuse them for model training, continuously improving the model's recognition accuracy. In addition, the multi-task output layer can simultaneously support multiple tasks such as target recognition, greatly expanding the application scenarios of the system in different fields and improving the system's practicality and flexibility. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the system architecture of the present invention. Detailed Implementation
[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0024] Example 1 like Figure 1 As shown, a deep learning-based UAV image recognition and analysis system includes a data acquisition and preprocessing module, a deep learning model building module, a model training module, an image recognition and analysis module, and a result output and feedback module. The data acquisition and preprocessing module is used to acquire drone images under different time, weather, and lighting conditions. The acquired images are then subjected to bilateral filtering for noise reduction and histogram equalization for enhancement. The drone images are then divided into training, validation, and test sets in an 80:15:5 ratio. This ensures that the data covers diverse scenarios by acquiring drone images under different time, weather, and lighting conditions, providing comprehensive training samples for subsequent models. Bilateral filtering removes noise while preserving edge details and avoiding feature blurring. Histogram equalization enhances the image grayscale range and improves contrast to highlight target features. The training, validation, and test sets are then divided in an 80:15:5 ratio for subsequent model training, hyperparameter tuning, and performance evaluation, respectively, ensuring the scientific nature and generalization ability of the model training. The deep learning model building module is used to construct deep learning models. These models include adaptive image scaling layers, dynamically adjustable convolutional layers, batch normalization layers, ReLU activation functions, attention mechanisms, and multi-task output layers. The adaptive image scaling layer adapts to input images of different sizes, reducing information loss due to size differences. The dynamically adjustable convolutional layers flexibly optimize the kernel size and number based on image features, improving the targeting of feature extraction. The batch normalization layer stabilizes the input data distribution, accelerating model convergence and reducing internal covariate shifts. The ReLU activation function enhances the model's ability to fit complex features. The attention mechanism layer weights and adjusts feature channels to highlight key features and suppress interference. The multi-task output layer can simultaneously support multiple tasks such as target recognition, expanding the model's application scenarios. The model training module is used to train a deep learning model using the training set. It employs a weighted sum of cross-entropy loss and mean squared error loss as the total loss function. The deep learning model is optimized using stochastic gradient descent combined with an adaptive learning rate adjustment strategy. Hyperparameters are adjusted based on validation set evaluation results. This allows the model to learn the features and patterns in the samples, establishing a mapping relationship between input and output. The cross-entropy loss function measures the difference between the predicted and true classes in classification tasks, while the mean squared error loss function measures the bias in bounding box predictions in regression tasks. The weighted sum of these two loss functions achieves synergistic optimization of classification and regression performance. Furthermore, the stochastic gradient descent algorithm combined with an adaptive learning rate adjustment strategy dynamically optimizes the model parameters to minimize the loss function. Simultaneously, hyperparameters are adjusted based on validation set evaluation results to avoid overfitting and ensure model stability on unknown data. The image recognition and analysis module is used to receive images collected in real time by the UAV and input them into a trained deep learning model to obtain target category information in the UAV images, so as to quickly identify the target category in the image; The results output and feedback module is used to visualize the target category information obtained by the image recognition and analysis module from the UAV image. It also establishes a feedback mechanism to receive error samples from users and retrain the deep learning model. This allows users to intuitively obtain the recognition results by visually outputting the target category information. The feedback mechanism receives error samples from users and reuses them for model training to continuously improve the model's recognition accuracy.
[0025] The formula for the bilateral filtering algorithm used in the data acquisition and preprocessing module is as follows: ; In the formula, The original image is in Pixel value at that location, These are the pixel values after noise reduction. These are normalized weights, used to balance the contributions of pixels at different locations. It is a spatial domain Gaussian kernel function used to measure the impact of spatial distance between pixels on filtering. It is a Gaussian kernel function with a range of values, used to measure the impact of pixel value differences on filtering. This formula balances the weights of spatial distance and pixel value differences, preserving image edge features while removing noise and avoiding the loss of key information.
[0026] The formula for the histogram equalization algorithm used in the data acquisition and preprocessing module is as follows: ; In the formula, It is the transformed grayscale level. It is grayscale. Number of times it appears It is the total number of pixels in the image. Represents grayscale level The probability of occurrence It is the total number of gray levels. This formula improves image contrast by redistributing gray values, enhancing the distinction between the target and the background, and facilitating subsequent feature extraction.
[0027] In the deep learning model construction module, the normalization formula used for the batch normalization layer is: ; In the formula, This is the normalized output. It is the first in the input batch data One data point, It is the mean of this batch of data. It is the variance of this batch of data. It is a small constant to prevent the denominator from being zero. and These are learnable parameters. This formula standardizes the input data, stabilizes the input distribution during network training, accelerates model convergence, and improves training stability.
[0028] In the deep learning model building module, the formula for calculating the attention weight vector of the attention mechanism layer is as follows:
[0029] In the formula, It is the vector obtained after global average pooling of the input feature map, used to aggregate global information from the feature map. and This is the weight matrix of the fully connected layer, used for dimensionality reduction and dimensionality expansion of the aggregated information. It is the ReLU activation function, which increases the nonlinear transformation capability. It is the Sigmoid activation function, which maps the output to the interval [0, 1]. It is the attention weight vector, used to weight and adjust different channels of the input feature map. This formula enhances the role of key features and suppresses interference from irrelevant features by learning the importance weights of feature channels, thereby improving the model's ability to identify important targets.
[0030] In the deep learning model building module, the output layer, capable of multi-task output, uses a Softmax classifier for the target recognition task, with the classification calculation formula as follows: ; In the formula, It is input Category The probability, It is a model for categories The output score, The formula converts the model output score into a probability distribution, ensuring that the probability value corresponding to each category is in the interval [0, 1] and the sum of the probabilities of all categories is 1. This makes it easy to intuitively determine the probability that the input image belongs to each category. The formula converts the model output into a probability distribution, intuitively reflecting the probability that the input image belongs to each category and clearly defining the target category.
[0031] In the model training module, the formula for the total loss function is:
[0032]
[0033]
[0034] In the formula, It is the total loss function. It is the cross-entropy loss function. It's a real label. The model predicts that the sample belongs to a category. The probability is used to measure the difference between the predicted class and the true class in a classification task. It is the mean squared error loss function. It is the sample size. These are the actual bounding box coordinates. These are the bounding box coordinates predicted by the model. and These are weighting coefficients. This formula achieves multi-task collaborative optimization through weighted summation, thereby improving the overall performance of the model.
[0035] In the model training module, the formula for combining the stochastic gradient descent algorithm with the adaptive learning rate adjustment strategy is as follows:
[0036] In the formula, These are the updated model parameters. These are the current model parameters. It is the initial learning rate. Is it up to the The sum of squared historical gradients up to the last step is used to reflect the historical status of parameter updates. It is a small constant to prevent the denominator from being zero. It is the gradient of the loss function under the current parameters. This formula dynamically adjusts the learning rate based on the historical gradient of the parameters, making the parameter updates more accurate and accelerating the model to converge to the optimal state.
[0037] The embodiments disclosed in this invention are preferred embodiments, but are not limited thereto. Those skilled in the art can easily understand the spirit of this invention based on the above embodiments and make different extensions and variations, but as long as they do not depart from the spirit of this invention, they are all within the protection scope of this invention.
Claims
1. A deep learning-based UAV image recognition and analysis system, characterized in that, The system includes a data acquisition and preprocessing module, a deep learning model building module, a model training module, an image recognition and analysis module, and a result output and feedback module. The data acquisition and preprocessing module is used to acquire UAV images under different time, weather, and lighting conditions, and sequentially perform bilateral filtering for noise reduction and histogram equalization enhancement on the acquired images. Then, the UAV images are divided into training set, validation set, and test set in an 80:15:5 ratio. The deep learning model building module is used to build a deep learning model, which includes an adaptive image scaling layer, a convolutional layer that can dynamically adjust parameters, a batch normalization layer, a ReLU activation function, an attention mechanism layer, and an output layer that can perform multi-task output. The model training module is used to train a deep learning model using the training set. It uses the weighted sum of the cross-entropy loss function and the mean squared error loss function as the total loss function. It optimizes the deep learning model by combining the stochastic gradient descent algorithm with an adaptive learning rate adjustment strategy, and adjusts the hyperparameters based on the evaluation results of the validation set. The image recognition and analysis module is used to receive images collected in real time by the UAV and input them into a trained deep learning model to obtain target category information in the UAV images. The result output and feedback module is used to visualize the target category information obtained by the image recognition and analysis module from the UAV image, and to establish a feedback mechanism to receive error samples from users and retrain the deep learning model.
2. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, The formula for the bilateral filtering algorithm used in the data acquisition and preprocessing module is as follows: ; In the formula, The original image is in Pixel value at that location, These are the pixel values after noise reduction. These are normalized weights, used to balance the contributions of pixels at different locations. It is a spatial domain Gaussian kernel function used to measure the impact of spatial distance between pixels on filtering. It is a range Gaussian kernel function used to measure the impact of pixel value differences on filtering.
3. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, The formula for the histogram equalization algorithm used in the data acquisition and preprocessing module is as follows: ; In the formula, It is the transformed grayscale level. It is grayscale. Number of times it appears It is the total number of pixels in the image. Represents grayscale level The probability of occurrence It represents the total number of gray levels.
4. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, In the deep learning model construction module, the normalization formula used for the batch normalization layer is: ; In the formula, This is the normalized output. It is the first in the input batch data One data point, It is the mean of this batch of data. It is the variance of this batch of data. It is a small constant to prevent the denominator from being zero. and These are learnable parameters.
5. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, In the deep learning model construction module, the formula for calculating the attention weight vector of the attention mechanism layer is as follows: ; In the formula, It is the vector obtained after global average pooling of the input feature map, used to aggregate global information from the feature map. and This is the weight matrix of the fully connected layer, used for dimensionality reduction and dimensionality expansion of the aggregated information. It is the ReLU activation function, which increases the nonlinear transformation capability. It is the Sigmoid activation function, which maps the output to the interval [0, 1]. It is the attention weight vector, used to weight and adjust different channels of the input feature map.
6. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, In the deep learning model construction module, the output layer capable of multi-task output uses a Softmax classifier for the target recognition task, and the classification calculation formula is as follows: ; In the formula, It is input Category The probability, It is a model for categories The output score, The formula represents the total number of categories. It converts the scores output by the model into a probability distribution, ensuring that the probability value of each category is in the interval [0, 1] and the sum of the probabilities of all categories is 1. This makes it easier to intuitively determine the probability that the input image belongs to each category.
7. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, In the model training module, the formula for the total loss function is: ; ; ; In the formula, It is the total loss function. It is the cross-entropy loss function. It's a real label. The model predicts that the sample belongs to a category. The probability is used to measure the difference between the predicted class and the true class in a classification task. It is the mean squared error loss function. It is the sample size. These are the actual bounding box coordinates. These are the bounding box coordinates predicted by the model. and It is the weighting coefficient.
8. The UAV image recognition and analysis system based on deep learning according to claim 1, characterized in that, In the model training module, the formula for combining the stochastic gradient descent algorithm with the adaptive learning rate adjustment strategy is as follows: ; In the formula, These are the updated model parameters. These are the current model parameters. It is the initial learning rate. Is it up to the The sum of squared historical gradients up to the last step is used to reflect the historical status of parameter updates. It is a small constant to prevent the denominator from being zero. It is the gradient of the loss function under the current parameters.