Unmanned aerial vehicle target detection and identification method and medium
By constructing the rotational invariant convolution kernel and hyperspherical projection mechanism, the parameters and loss functions of the convolutional neural network are optimized, and the problems of target rotation changes, insufficient category distinction and poor processing of high-dimensional feature information in drone target detection are solved, achieving higher detection robustness and recognition accuracy.
Patent Information
- Application Number
- CN202510572575.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
AI Technical Summary
Existing drone target detection technology is difficult to effectively deal with the problems of target rotation changes, insufficient category distinction, insufficient learning and poor processing of high-dimensional feature information.
By constructing a rotary invariant convolution kernel and hyperspherical projection mechanism, combining convolutional neural network and gradient descent algorithm, the parameters and loss functions of the convolutional neural network are optimized to achieve stable detection and recognition of targets in drone images.
It improves the robustness of rotary target detection, enhances the distinction of target categories, improves the identification accuracy in complex contexts, and improves the detection ability of rare targets.
Smart Images

Figure CN120088688A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to a method and medium for drone target detection and recognition. Background Art
[0002] With the wide application of drones in fields such as monitoring, rescue, and urban management, how to quickly and accurately detect and recognize targets in complex scenarios has become an urgent problem to be solved.
[0003] Traditional target detection techniques often rely on image acquisition at a fixed angle and feature extraction methods with fixed convolutional kernels. When facing the challenges brought by the flight angle, light changes, and multi-angle rotation of targets during drone shooting, it is easy to cause a decline in detection accuracy and insufficient robustness. At the same time, due to the fact that the image data collected by drones is often affected by noise, environmental interference, and background complexity, existing methods have obvious deficiencies in dealing with class imbalance, high-dimensional redundant information, and data adaptability in different environments. Summary of the Invention
[0004] This application provides a method and medium for drone target detection and recognition to solve the problems existing in the existing solutions, such as being difficult to effectively cope with target rotation changes, lacking optimization means for class discrimination, insufficient learning of a small number of target classes, and lacking an effective feature screening mechanism for protecting high-dimensional feature information.
[0005] In a first aspect, this application provides a method for drone target detection and recognition. The method includes:
[0006] Obtain the flight information of the drone through a preset acquisition interface, and then control the drone to execute a flight task; wherein, the flight information at least includes: flight path, flight height, shooting angle, and image overlap rate; obtain a first target image through the drone executing the flight task, and then obtain the first target image with key target information marked.
[0007] Perform noise removal, normalization, and geometric correction on the first target image to obtain a processed target image.
[0008] Use the processed target image as the input sample data of the convolutional neural network and input it into the convolutional neural network.
[0009] Initialize the convolutional neural network structure, construct a rotation-invariant convolutional kernel, and construct a rotation-invariant kernel space containing convolutional kernels at several rotation angles.
[0010] Project the output features of the convolutional neural network into the hypersphere space, and calculate the hypersphere projection interval corresponding to the output features.
[0011] During the training process of the convolutional neural network, obtain the cross-entropy loss, regularization term, class balance loss, and local-global constraint loss;
[0012] Use the gradient descent algorithm to update the parameters of the convolutional neural network that incorporates the rotation-invariant kernel space to minimize the classification error; at the same time, calculate the total loss function using the hypersphere projection margin, cross-entropy loss, regularization term, class balance loss, and local-global constraint loss;
[0013] Repeat the convolutional neural network training process until the preset stop iteration condition is satisfied to obtain a trained convolutional neural network;
[0014] Through the unmanned aerial vehicle performing the flight mission, obtain the second target image, input the second target image into the trained convolutional neural network, and obtain the prediction result.
[0015] In an implementation manner of the present application, initialize the convolutional neural network structure and construct a rotation-invariant convolutional kernel, specifically including:
[0016] Through the formula:
[0017] , construct the kernel function ;
[0018] Wherein, represents the input data of the convolutional neural network; represents the rotation angle; represents the rotation-invariant kernel construction function.
[0019] In an implementation manner of the present application, the rotation-invariant kernel construction function is:
[0020] ;
[0021] Wherein, is the rotation-invariant kernel construction function; represents the input data of the convolutional neural network; represents the rotation angle, represents the rotation matrix for rotating the base kernel function to the angle ;
[0022] represents the base kernel function, the initial features extracted from the original input data; represents the rotation angle space.
[0023] In an implementation manner of the present application, construct a rotation-invariant kernel space including convolutional kernels at several rotation angles, specifically including:
[0024] Through the formula:
[0025] , construct a rotation-invariant kernel space ;
[0026] Among them, is composed of convolution kernels at several rotation angles; T represents the rotation angle space, including all rotation angles; represents the kernel function, represents the input data of the convolutional neural network; represents the rotation angle; represents the rotation-invariant kernel construction function.
[0027] In an implementation manner of the present application, project the output features of the convolutional neural network into the hypersphere space, and calculate the hypersphere projection interval corresponding to the output features, specifically including:
[0028] Through the formula:
[0029] , calculate the hypersphere projection interval ;
[0030] Among them, represents the feature representation extracted by the convolutional neural network from the input data; represents the non-linear mapping function, which is used to map the low-dimensional features to a preset high-dimensional space; represents the center vector of the preset target category , and ; represents the center vector of the non-preset target category , and ;
[0031] Among them, is the number of samples of the preset target category , is the number of samples of the non-preset target category ; is the sample of the th preset target category , is the sample of the th non-preset target category .
[0032] In an implementation manner of the present application, the non-linear mapping function is:
[0033] ;
[0034] Among them, represents the projected preset high-dimensional feature; represents the input data of the non-linear mapping function; represents the A support vector coefficient; Denotes the kernel function, which will Be expressed as , Is the input data of the kernel function, Is the Feature representation of the i-th sample; Denotes the number of support vectors.
[0035] In an implementation manner of the present application, the total loss function is calculated by using the hypersphere projection margin, cross-entropy loss, regularization term, class balance loss, and local-global constraint loss, specifically including:
[0036] Through the formula:
[0037] , the total loss function L is calculated;
[0038] Among them, Denotes the cross-entropy loss; Denotes the class probability predicted by the convolutional neural network model; Denotes the true class label of the target in the UAV image; Denotes the hypersphere projection margin optimization weight coefficient; Denotes the hypersphere projection margin; Denotes the regularization weight coefficient; Denotes the regularization term; Denotes all model parameters in the network; Denotes the class balance loss, Denotes the influence factor of the class balance loss; Denotes the local-global constraint loss, Denotes the influence factor of the local-global constraint loss.
[0039] In an implementation manner of the present application, the gradient descent algorithm is used to update the parameters of the convolutional neural network that fuses the rotation-invariant kernel space, specifically including:
[0040] Through the formula:
[0041] ;
[0042] Among them, Denotes the weight parameter of the convolutional neural network at the t-th iteration; Denotes the weight parameter of the convolutional neural network at the t + 1-th iteration; Denotes the learning rate of the convolutional neural network; Denotes the sum of M samples in the current batch; is the total number of samples input to the convolutional neural network for the current batch; represents the weight coefficient of the ith sample; represents the gradient of the total loss function of the convolutional neural network with respect to the weight parameter; ith input sample in the current batch; represents the true label of the
[0043] ith sample. In an implementation manner of the present application, the weight coefficient of the
[0044] ith sample specifically includes:
[0045] wherein, represents the number of samples of the ith preset target category, and i ∈ [1, y], where y represents the number of preset target categories.
[0046] categories.
[0047] In a second aspect, the present application provides a non-volatile computer storage medium, on which computer instructions are stored, and when the computer instructions are executed, a drone target detection and recognition method as described in any one of the above is implemented.
[0048] From the above technical solutions, it can be seen that the present application has the following advantages:
[0049] 1. Aiming at the problem of arbitrary rotation of targets in drone images, a rotation-invariant convolutional kernel based on a rotation matrix and a basis kernel function is adopted, so that local features can be stably extracted at different rotation angles, thereby improving the robustness of rotation target detection.
[0050] 2. Project the output features of the convolutional neural network into the hypersphere space, and by optimizing the interval between category centers, strengthen the discrimination of target categories and improve the recognition accuracy in complex backgrounds.
[0051] 3. Through the joint optimization mechanism of category balance and local-global constraints, adopt the dynamic category balance loss and the local-global constraint loss, and by adjusting the sample weights and jointly optimizing the local feature aggregation and the global category distribution, overcome the problems of sample imbalance and high-dimensional redundant information, and improve the detection ability of rare targets.
[0052] 4. Adopt an adaptive feature selection and weighted gradient descent update strategy, through the adaptive feature selection mechanism, dynamically screen the features most relevant to the task, and use the weighted gradient descent method to update the network parameters, improving the adaptability and convergence speed of the model in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0054] Figure 1 It is a flowchart of a method for unmanned aerial vehicle target detection and recognition provided by an embodiment of the present application.
[0055] Figure 2 It is a comparison chart of detection accuracy between the conventional target detection method and the present application under different noise environments provided by an embodiment of the present application.
[0056] Figure 3 It is a comparison chart of the change in detection accuracy during the training iteration process provided by an embodiment of the present application.
[0057] Figure 4 It is a graph of the influence change of a rotation angle and the value of a base kernel provided by an embodiment of the present application.
[0058] Figure 5 It is a graph of the influence of the joint optimization of local constraint loss and global constraint loss on the overall loss provided by an embodiment of the present application. Specific embodiments
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0060] Those skilled in the art should understand that the embodiments described below are only the preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through the preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure, rather than to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present disclosure.
[0061] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0062] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0063] An embodiment provides a method for detecting and identifying drone targets, as Figure 1 shown. The method provided in the embodiments of the present application mainly includes the following steps:
[0064] Step 110: Obtain the flight information of the drone through a preset acquisition interface, and then control the drone to execute a flight mission; obtain a first target image through the drone executing the flight mission, and then obtain a first target image with key target information marked; perform noise removal, normalization
[0065] and geometric correction on the first target image to obtain a processed target image; use the processed target image as input sample data of a convolutional neural network and input it into the convolutional neural network.
[0066] It should be noted that the flight information at least includes: flight path, flight altitude, shooting angle, and image overlap rate.
[0067] In some embodiments, the flight state, GPS position, timestamp, and sensor parameters are recorded in real time to facilitate subsequent data correction and quality assessment.
[0068] The collected image data is mainly RGB color images, and the data storage format is a multi-dimensional array (such as a tensor format), and its size is defined as , where represents the image height, represents the image width, represents the number of channels.
[0069] In addition to the image content, metadata is also saved, including information such as timestamp, GPS coordinates, flight altitude, sensor status, and ambient light.
[0070] The collected data is labeled, and the key targets (such as vehicles, pedestrians, buildings, or other specific targets) in the drone mission are determined in advance.
[0071] The labeling method can adopt manual labeling, and a unique label is assigned to each image.
[0072] In addition, the image data collected by the drone often has noise due to reasons such as platform vibration and transmission interference. At the same time, there are significant differences in image brightness, contrast, etc. First, perform data preprocessing and image correction on the first target image, remove noise, normalize, and geometrically correct the image data collected by the drone to standardize the image data, thereby improving the stability of subsequent feature extraction and target detection. The normalization method is expressed as:
[0073] ;
[0074] In the formula, represents the normalized image data; represents the image data collected by the original drone; represents the mean value of the image data; represents the standard deviation of the image data.
[0075] Step 120: Initialize the convolutional neural network structure, construct a rotation-invariant convolutional kernel, and construct a rotation-invariant kernel space containing convolutional kernels at several rotation angles; project the output features of the convolutional neural network into the hypersphere space, and calculate the hypersphere projection interval corresponding to the output features.
[0076] It should be noted that in the images collected by the drone, targets (such as vehicles, pedestrians, or
[0077] other targets) usually appear in arbitrary directions. The fixed directionality of traditional convolutional kernels will cause unstable feature extraction when the rotation angle changes, and may not be robust enough to rotating targets. In this application, by initializing the convolutional neural network structure and constructing a rotation-invariant convolutional kernel, preliminary feature extraction for targets in various orientations is achieved, specifically including:
[0078] Through the formula:
[0079] , construct the kernel function ;
[0080] Among them, represents the input data of the convolutional neural network, that is, the image data collected by the drone; represents the rotation angle, which determines the rotation direction of the convolutional kernel; represents the rotation-invariant kernel construction function, which is used to generate a convolutional kernel with a stable response to rotation changes, ensuring that even if the target presents different rotation states in the image, the convolutional neural network can still capture similar local features.
[0081] In the rotation-invariant kernel space, construct the convolutional kernel through the rotation matrix and the base kernel function. The rotation-invariant kernel construction function is:
[0082] ;
[0083] Among them, is a rotation-invariant kernel construction function; represents the input data of the convolutional neural network; represents the rotation angle, represents the rotation matrix, which is used to rotate the base kernel function to the angle ; represents the base kernel function, the initial features extracted from the original input data; represents the rotation angle space.
[0084] Furthermore, the base kernel function is the unrotated initial convolutional kernel, which is used to generate kernels at different angles through rotation and is a learnable parameter matrix. For example, when the base kernel function is initialized, a 3×3 matrix is randomly generated. During the training process of the convolutional neural network, the base kernel weights are optimized through backpropagation so that the rotated kernels can extract rotation-invariant features. For example, if the base kernel is , a new kernel is generated after rotating 30°.
[0085] Furthermore, the rotation matrix rotates the filter by a given angle, and the specific expression can be:
[0086] ;
[0087] In the formula, represents the cosine value of the angle ; represents the sine value of the angle .
[0088] By constructing a kernel space containing convolutional kernels at each rotation angle, the convolutional operation can stably extract target features in different rotation states to adapt to various orientations of the target in UAV images. The rotation angle space covers all possible directions, ensuring that the network can still maintain a stable response in the face of the diversity of target orientations in the actual UAV-acquired images. Constructing a rotation-invariant kernel space containing convolutional kernels at several rotation angles specifically includes:
[0089] Through the formula:
[0090] , construct the rotation-invariant kernel space ;
[0091] Among them, is composed of convolutional kernels at several rotation angles and is composed of convolutional kernels at several rotation angles; represents the rotation angle space, which contains all possible rotation angles; T represents the rotation angle space, which contains all the rotation angles; denotes the kernel function, denotes the input data of the convolutional neural network; denotes the rotation angle; denotes the rotation-invariant kernel constructor.
[0092] When the target is mixed with the surrounding environment, the distinguishability between the target and the background in the UAV image is often low. By projecting the features extracted by the convolutional network into a hypersphere space and optimizing the interval between classes within this space to strengthen the distinction between target classes, this step projects the output features of the convolutional neural network into the hypersphere space, calculates the hypersphere projection interval corresponding to the output features, and optimizes the hypersphere projection interval. By projecting the output features of the convolutional neural network into the hypersphere space and optimizing the interval between class centers, high distinguishability of the target classes in the UAV image is achieved, specifically including:
[0093] Through the formula:
[0094] , the hypersphere projection interval is calculated , measuring the distance ratio between classes;
[0095] Among them, denotes the feature representation extracted from the input data by the convolutional neural network; denotes the non-linear mapping function for mapping low-dimensional features to a preset high-dimensional space; denotes the preset target class 's central vector, and ; denotes the non-preset target class 's central vector,
[0096] and ;
[0097] Among them, is the number of samples of the preset target class , is the number of samples of the non-preset target class ; is the th sample of the preset target class , is the th sample of the non-preset target class .
[0098] Furthermore, the non-linear mapping function realizes the mapping through the kernel function, maps the low-dimensional features to a higher-dimensional space to better distinguish similar classes, and the non-linear mapping function is:
[0099] ;
[0100] Among them, represents the preset high-dimensional feature after mapping; represents the input data of the non-linear mapping function; represents the th support vector coefficient, reflecting its contribution in the mapping; represents the kernel function, which maps is expressed as , is the input data of the kernel function, is the th sample feature representation; represents the number of support vectors.
[0101] Furthermore, the support vector coefficients are obtained by solving the preset dual problem of the support vector machine, reflecting the importance of samples in classification, and are calculated by optimizing the following optimization problem of the support vector machine through an optimization problem:
[0102] , with the constraint
[0103] In the formula, represents the th support vector coefficient; represents the maximum value of the de-support vector coefficients; represents the class probability predicted by the convolutional neural network model for the th sample; represents the class probability predicted by the convolutional neural network model for the th sample.
[0104] Step 130, during the training process of the convolutional neural network, obtain the cross-entropy loss, regularization term, class balance loss, and local-global constraint loss; use the gradient descent algorithm to update the parameters of the convolutional neural network that integrates the rotation-invariant kernel space to minimize the classification error; at the same time, use the hypersphere projection margin, cross-entropy loss, regularization term, class balance loss, and local-global constraint loss to calculate the total loss function; repeat the convolutional neural network training process until the preset stop iteration condition is met to obtain the trained convolutional neural network.
[0105] Among them, using the hypersphere projection margin, cross-entropy loss, regularization term, class balance loss, and local-global constraint loss to calculate the total loss function specifically includes:
[0106] Through the formula:
[0107] , calculate the total loss function L;
[0108] Among them, Denotes the cross - entropy loss; Denotes the class probability predicted by the convolutional neural network model; Denotes the true class label of the target in the UAV image; Denotes the hyper - spherical projection interval optimization weight coefficient; Denotes the hyper - spherical projection interval; Denotes the regularization weight coefficient; Denotes the regularization term; Denotes all model parameters in the network; Denotes the class - balance loss, Denotes the influence factor of the class - balance loss; Denotes the local - global constraint loss, Denotes the influence factor of the local - global constraint loss. Preferably, Is set to 0.2, Is set to 0.4.
[0109] In the step, the gradient - descent algorithm is used to update the parameters of the convolutional neural network integrated with the rotation - invariant kernel space, specifically including:
[0110] Through the formula:
[0111] ;
[0112] Where, Denotes the weight parameter of the convolutional neural network at the -th iteration; Denotes the weight parameter of the convolutional neural network at the -th iteration; Denotes the learning rate of the convolutional neural network; Denotes the sum of M samples in the current batch; Is the total number of samples input to the convolutional neural network in the current batch; Denotes the weight coefficient of the -th sample; Denotes the gradient of the total loss function of the convolutional neural network with respect to the weight parameter; Denotes the -th input sample in the current batch; Denotes the -th true label of the sample.
[0113] Where, the weight coefficient of the -th sample, the specific calculation process can be:
[0114] ;
[0115] Where, Denotes the The number of samples of a preset target category, and \(i\in[1,y]\), where \(y\) represents the number of preset target categories.
[0116] Those skilled in the art can understand that the weight coefficient of the sample solves the problem of class imbalance.
[0117] For each training sample, a weight is assigned according to the number of samples in its category. The category with fewer samples has a higher weight, so that more attention is given during gradient update, making the convolutional neural network model more sensitive to minority target categories, and effectively suppressing common background noise at the same time, thus showing better robustness and accuracy in the UAV target detection task.
[0118] In addition, the process of obtaining the cross-entropy loss can be as follows:
[0119] The calculation method is expressed as:
[0120] ;
[0121] In the formula, represents the cross-entropy loss, accumulating the prediction errors of each category; represents the total number of target categories; represents the category probability predicted by the convolutional neural network model for the \(i\)-th sample; represents the true category label of the target in the \(i\)-th UAV image sample; represents the logarithmic function.
[0122] In addition, the process of obtaining the regularization term can be as follows:
[0123] The calculation method is expressed as:
[0124] ;
[0125] In the formula, represents the regularization term, used to punish the excessive values of the model parameters; represents the regularization coefficient, controlling the regularization strength; represents the weight parameter of the convolutional neural network. Preferably, is set to 0.3.
[0126] In addition, the process of obtaining the class balance loss can be as follows:
[0127] The calculation method is expressed as:
[0128] ;
[0129] ;
[0130] In the formula, Represents the predicted probability of the th class sample; Represents the one-hot encoding of the true label of the th class; Is the number of samples of the th class, used to assign a larger weight to the class with fewer samples; Is the importance factor of the th class, and its specific value is determined by grid search; Is the sensitivity hyperparameter, used to adjust the sensitivity of class weight update. Preferably, Is set to 0.2.
[0131] In addition, the process of obtaining the local-global constraint loss can be:
[0132] The calculation method of the influence factor of the local-global constraint loss is expressed as:
[0133] ;
[0134] ;
[0135] ;
[0136] In the formula, Is the local constraint loss, which forces similar features to gather by calculating the distance between sample features in the neighborhood, thus ensuring local feature consistency; Is the global constraint loss, which makes different classes have obvious distinctions globally by comparing the mean values of class features; Is the weight hyperparameter of the local constraint; Is the weight hyperparameter of the global constraint; Is the coefficient of the regularization term of the local-global constraint loss; Is the learnable weight parameter matrix in the feature selection module, which is updated by the gradient descent method; Is the neighborhood set of the th sample; Represents the feature representation obtained by the th sample through the feature selection mechanism; Represents the feature representation obtained by the th sample through the feature selection mechanism; Is the mean vector of the features of the th class, Is the global mean vector of all sample features. Preferably, Is set to 0.2, Is set to 0.3, Set to 0.5, the feature selection module uses a fully connected neural network, and the neighborhood set of the samples is obtained by using Euclidean distance measurement.
[0137] Those skilled in the art can understand that high-dimensional features often contain a large amount of redundant information, especially in the complex background of UAV images. This application adopts an adaptive feature selection mechanism based on local-global constraints, and through the joint optimization of local and global constraints, dynamically screens the features most relevant to the classification task, alleviates the redundant information problem in high-dimensional data, and can adaptively screen out the most relevant features in a dynamic environment (UAV real-time monitoring), effectively improving the accuracy and robustness of target recognition, and solving the deficiencies of traditional methods in processing high-dimensional redundant information.
[0138] Step 140: Obtain a second target image through a UAV performing a flight mission, and input the second target image into the trained convolutional neural network to obtain a prediction result.
[0139] Specifically, perform the same preprocessing on the newly collected UAV image data as in the training stage, including noise removal, normalization, and geometric correction, so that the input data meets the distribution requirements during model training; further, input the preprocessed image data into the trained convolutional neural network model, and obtain the high-dimensional feature representation of the image through forward propagation. The rotation-invariant kernel and hypersphere projection optimization mechanism built into the model ensure that features can still be stably extracted and classified under the conditions that the target has an arbitrary rotation angle and a complex background; further, based on the feature information output by the network, use the classification output of the fully connected layer to perform target detection and classification on the entire image, and output the category (such as vehicle, pedestrian, building, or other specific targets) to which the target belongs and the corresponding confidence level.
[0140] Output the detection and classification results in the form of structured data for subsequent monitoring and analysis applications.
[0141] Output the detection and classification results in the form of structured data for subsequent monitoring and analysis applications.
[0142] To verify the advantages of the technology of this application, the following experimental data analysis is carried out:
[0143] As Figure 2 shown, to verify the difference in detection accuracy performance between various conventional target detection methods and the UAV target detection and recognition method proposed in this application under different noise environments, in the experiment, the detection effects of the conventional method and the method of this application were compared under low-noise, medium-noise, and high-noise conditions. The results show that in the case of large noise interference, the method of this application can more effectively resist the influence of noise and maintain a high detection accuracy, thus proving that its robustness and stability in dealing with complex noise environments are superior to traditional technologies.
[0144] As Figure 3As shown in the figure, to analyze the accuracy improvement and convergence of the object detection model during the training process, the performance changes of the conventional method and the method of the present application during the training iteration are compared. In the experiment, by analyzing the curve of the object detection accuracy changing with the number of iterations during the training process, it can be intuitively seen that the method of the present application can quickly improve the detection accuracy in the initial stage of training, and maintain a relatively stable improvement trend in the subsequent iterations, and finally achieve a higher detection accuracy. The test results show the obvious advantages of the method of the present application in terms of model training efficiency and stability, demonstrating the superior performance of the technology of the present application in complex scenarios.
[0145] As Figure 4 shown in the figure, to analyze the influence of the rotation-invariant kernel constructor on the output of the convolutional kernel at different rotation angles, by constructing a graph of the influence change of the rotation angle and the numerical value of the base kernel, the dynamic change of the output of the convolutional kernel in various rotation states is shown. The experimental results show that through the adaptive design of the rotation angle, the technology of the present application enables the convolutional kernel to effectively capture the subtle features under various angle changes during the feature extraction process, thereby ensuring the stability and robustness of the feature extraction, and reflecting the advantages of the technology of the present application in processing multi-angle object detection.
[0146] As Figure 5 shown in the figure, by analyzing the influence of the joint optimization of the local constraint loss and the global constraint loss on the overall loss, the experimental results show the important role of local consistency and global discrimination in improving the model performance. When the technology of the present application simultaneously considers local feature aggregation and global class discrimination, it can make the overall loss reach a lower level, thereby promoting the accurate recognition of the model in complex scenarios.
[0147] In addition, the embodiment of the present application also provides a non-volatile computer storage medium, on which executable instructions are stored. When the executable instructions are executed, a method for unmanned aerial vehicle object detection and recognition as described above is implemented.
[0148] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting and identifying unmanned aerial vehicle targets, characterized in that: The method comprises: The flight information of the UAV is obtained through a preset acquisition interface, and then the UAV is controlled to perform a flight mission; wherein the flight information includes at least: a flight path, a flight altitude, a shooting angle, and an image overlap rate; a first target image is obtained through the UAV performing the flight mission, and then the first target image with key target information marked is obtained; Performing noise removal, normalization and geometric correction on the first target image to obtain a processed target image; The processed target image is used as the input sample data of the convolutional neural network and input into the convolutional neural network; Initialize the convolutional neural network structure, construct a rotationally invariant convolution kernel, and construct a rotationally invariant kernel space containing convolution kernels under several rotation angles; Project the output features of the convolutional neural network to the hypersphere space, and calculate the hypersphere projection interval corresponding to the output features; During the training of convolutional neural networks, cross entropy loss, regularization term, class balance loss and local-global constraint loss are obtained; The gradient descent algorithm is used to update the parameters of the convolutional neural network fused with the rotation-invariant kernel space to minimize the classification error. At the same time, the hypersphere projection margin, cross entropy loss, regularization term, category balance loss and local-global constraint loss are used to calculate the total loss function. Repeat the convolutional neural network training process until the preset stop iteration condition is met to obtain a trained convolutional neural network; The second target image is obtained by the UAV performing the flight mission, and the second target image is input into the trained convolutional neural network to obtain the prediction result.
2. The method for detecting and identifying unmanned aerial vehicle targets according to claim 1, characterized in that: Initialize the convolutional neural network structure and construct a rotation-invariant convolution kernel, including: By formula: , construct kernel function ; in, Represents the input data of the convolutional neural network; Indicates the rotation angle; express Rotationally invariant kernel constructor.
3. The method for detecting and identifying unmanned aerial vehicle targets according to claim 2, characterized in that: Rotationally invariant kernel constructor for: ; in, is the rotationally invariant kernel constructor; Represents the input data of the convolutional neural network; represents the rotation angle, Represents the rotation matrix, which is used to rotate the base kernel function to the angle ; represents the base kernel function, the initial features extracted from the original input data; Represents the rotation angle space.
4. The method for detecting and identifying unmanned aerial vehicle targets according to claim 1, characterized in that: Construct a rotationally invariant kernel space containing convolution kernels under several rotation angles, including: By formula: , construct rotationally invariant kernel space ; in, It is composed of convolution kernels under several rotation angles; T represents the rotation angle space, which includes all rotation angles; represents the kernel function, Represents the input data of the convolutional neural network; Indicates the rotation angle; Represents a rotationally invariant kernel constructor.
5. The method for detecting and identifying unmanned aerial vehicle targets according to claim 1, characterized in that: Project the output features of the convolutional neural network to the hypersphere space, and calculate the hypersphere projection interval corresponding to the output features, including: By formula: , calculate the hypersphere projection interval ; in, Represents the feature representation of the input data extracted by the convolutional neural network; represents a nonlinear mapping function; Indicates the preset target category The center vector of ; Indicates non-preset target category The center vector of ; in, Is the preset target category The number of samples, Non-preset target category The number of samples; For the Preset target categories A sample of For the Unpredictable target categories of the sample.
6. The method for detecting and identifying unmanned aerial vehicle targets according to claim 5, characterized in that: Nonlinear mapping function for: ; in, Represents the preset high-dimensional features after mapping; Represents the input data of the nonlinear mapping function; Indicates Support vector coefficients; represents the kernel function, Expressed as , is the input data of the kernel function, For the Sample feature representation; Represents the number of support vectors.
7. The method for detecting and identifying unmanned aerial vehicle targets according to claim 1, characterized in that: The total loss function is calculated using the hypersphere projection margin, cross entropy loss, regularization term, class balance loss, and local-global constraint loss, including: By formula: , calculate the total loss function L; in, represents the cross entropy loss; Represents the category probability predicted by the convolutional neural network model; Indicates the true category labels of the objects in the drone images; represents the hypersphere projection interval optimization weight coefficient; represents the hypersphere projection interval; represents the regularization weight coefficient; represents the regularization term; Represents all model parameters in the network; represents the class balance loss, Represents the impact factor of class balance loss; represents the local-global constraint loss, Represents the impact factor of the local-global constraint loss.
8. The method for detecting and identifying unmanned aerial vehicle targets according to claim 1, characterized in that: The gradient descent algorithm is used to update the parameters of the convolutional neural network integrated with the rotation invariant kernel space, including: By formula: ; in, Indicates The weight parameters of the convolutional neural network for the iteration; Indicates The weight parameters of the convolutional neural network for the iteration; Represents the learning rate of the convolutional neural network; Indicates the sum of M samples in the current batch; The total number of samples input to the convolutional neural network for the current batch; Indicates The weight coefficient of each sample; Represents the gradient of the total loss function of the convolutional neural network with respect to the weight parameters; Indicates the number of input samples; Indicates The true labels of the samples.
9. The method for detecting and identifying unmanned aerial vehicle targets according to claim 8, characterized in that: No. The weight coefficients of samples include: ; in, Indicates The number of samples of preset target categories, and i∈[1,y], y represents the number of preset target categories.
10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, a method for detecting and identifying a drone target as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Spacecraft visible light image classification method
CN106446965A
Injection molding product surface image defect identification method based on transfer learning
CN110111297A
Fine granularity and incremental learning-based spacecraft classification method in visible light image
CN117576442A
Remote sensing image detection processing method and system based on density mask
CN118279746A
SAR (Synthetic Aperture Radar) image recognition system and method based on distance measurement and full convolutional network
CN118608847A