Unmanned aerial vehicle identification device and method based on link signal and optical feature combination

By combining the multi-dimensional feature recognition network of drone link signals and optical features, the problem of low accuracy in drone identification by single-mounted sensors is solved, and high-accuracy recognition of drones is achieved.

CN120744609APending Publication Date: 2025-10-03SOUTHWEST CHINA RES INST OF ELECTRONICS EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510837200.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In existing technologies, it is difficult for a single sensor to accurately identify drones, especially in optical images where drones are small and easily interfered with by other targets. Silent drones and those susceptible to interference cannot be identified when relying solely on communication link signals.

Method used

A joint recognition method based on link signals and optical features is adopted. By constructing a multidimensional feature recognition network, combining optical images and time-frequency graph features, using deep neural networks for information fusion and residual structure optimization, and designing a residual attention module to select useful features, finally recognition is performed through the cross-entropy loss function.

Benefits of technology

It achieves high-accuracy recognition of UAV targets, makes up for the defect of single-installed sensor's single information acquisition, and improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744609A_ABST
    Figure CN120744609A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle recognition device and method based on link signal and optical feature combination, and belongs to the technical field of target recognition, and the method comprises the steps: constructing a multi-dimensional feature recognition network based on heterogeneous data, and enabling the multi-dimensional feature recognition network to be used for recognizing an unmanned aerial vehicle; the heterogeneous data-based multi-dimensional feature recognition network comprises an input module, a full-connection structure module and an output module, the input module corresponds to two information extraction branches and is respectively used for inputting an optical image and a corresponding time-frequency graph; and information fusion is completed through a full-connection structure module, the recognition accuracy is improved in combination with a residual structure, and final target recognition and classification are completed. According to the invention, high-accuracy identification of the unmanned aerial vehicle target is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target recognition technology, and more specifically, to a drone recognition device and method based on the combination of link signals and optical features. Background Art

[0002] Currently, drone identification is often based on information from a single sensor, such as the characteristics of a drone's optical image or its communication link signal. In optical images, drones are small and occupy only a few pixels, making optical image recognition difficult and inaccurate. Furthermore, the presence of numerous other targets, such as vehicles and pedestrians on the ground and birds, balloons, and civil aircraft in the air, further reduces drone identification accuracy. Furthermore, methods that rely solely on the characteristics of drone communication link signals are unable to identify silent drones or distinguish between other non-signaling targets in the air, such as birds and balloons. Furthermore, they are susceptible to interference from other signals in the same frequency band as the communication link, such as Wi-Fi signals, which can affect recognition accuracy. Summary of the Invention

[0003] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a drone identification device and method based on the combination of link signals and optical features, which can achieve high-accuracy recognition of drone targets.

[0004] The object of the present invention is achieved through the following solutions:

[0005] A drone identification device based on a combination of link signals and optical features, comprising:

[0006] Construction module for building a multi-dimensional feature recognition network based on heterogeneous data;

[0007] The multidimensional feature recognition network based on heterogeneous data includes an input module, a fully connected structure module and an output module; the input module corresponds to two information extraction branches, which are used to input the optical image and its corresponding time-frequency map respectively; the fully connected structure module is used to complete information fusion, and the residual structure is combined to improve the recognition accuracy and complete the final target recognition and classification.

[0008] Furthermore, the input of the multidimensional feature recognition network based on heterogeneous data includes a 256x512x3 optical image and its corresponding 256x512x1 time-frequency map.

[0009] Furthermore, the output of the multi-dimensional feature recognition network based on heterogeneous data includes a 1xN vector, where N is the number of classification categories, which is used to describe the probability that the target belongs to each category.

[0010] Furthermore, it also includes an optical deep feature extraction module, which is used to input the image into the pre-trained Vgg16 network to extract optical features, and take the first 13 layers of the original Vgg16 network as the feature extraction network.

[0011] Furthermore, it also includes a deep feature extraction module for time-frequency graphs, which is used to extract features from time-frequency graphs carrying signal features using a four-layer convolutional network. The convolution kernel sizes are 7x7, 5x5, and two 3x3 kernels. Each convolution layer is activated by the relu function and uses a 2x2 maximum pooling size.

[0012] Furthermore, the information fusion is completed through the fully connected structure module, and the recognition accuracy is improved by combining the residual structure, which specifically includes:

[0013] After the deep features of the optical and time-frequency graph branches are extracted through the corresponding feature networks respectively, the features on both sides are combined together through cascading, and a three-layer residual attention module is designed to enable the network to automatically select more useful features. Among them, a single residual attention module contains a convolutional layer, a channel attention layer, and a spatial attention layer. The channel attention layer is used to help the network select optical or RF domain channels, and the spatial attention layer is used to help the network pay more attention to small targets in the image. The residual structure is then used to mitigate gradient explosion, making the overall network layer deeper and obtaining deeper feature information. Finally, the classification and recognition results are obtained after convolution and straightening the full connection.

[0014] Furthermore, a loss function design module is included to select cross entropy to constrain the classification results output by the network. The formula is as follows:

[0015]

[0016] Where y is the 1xN one-hot label of the target true category, d is the 1xN vector output by the network, and N is the number of classification categories.

[0017] A method for identifying drones based on a combination of link signals and optical features, based on the drone identification device based on a combination of link signals and optical features as described in any of the above items, comprises the following steps:

[0018] The heterogeneous data from two different sensors are input, and the features of the two data in the high-dimensional latent space are extracted through a neural network and combined, and finally the recognition is completed based on the combined features.

[0019] Furthermore, the input of heterogeneous data from two different sensors, extracting features of the two data in a high-dimensional latent space through a neural network and combining them, and finally completing recognition based on the combined features, specifically includes the following sub-steps:

[0020] S1, deep features are extracted from the two branches of optical and time-frequency images through corresponding neural networks;

[0021] S2 combines the features from both sides through cascading and designs a three-layer residual attention module, which enables the network to automatically select the more useful features. A single residual attention module consists of a convolutional layer, a channel attention layer, and a spatial attention layer. The channel attention layer helps the network select optical or RF domain channels, and the spatial attention layer helps the network focus on small objects in the image.

[0022] S3 uses the residual structure to alleviate gradient explosion, making the overall network layer deeper and obtaining deeper feature information;

[0023] S4, after convolution and straightening the full connection, the classification and recognition results are obtained.

[0024] Furthermore, the method further comprises the following sub-steps:

[0025] In the loss function design, cross entropy is selected to constrain the classification results of the network output. The formula is as follows:

[0026]

[0027] Where y is the 1xN one-hot label of the target true category, d is the 1xN vector output by the network, and N is the number of classification categories.

[0028] The beneficial effects of the present invention include:

[0029] The present invention proposes a joint drone identification scheme based on link signals and optical features. The two extracted features are combined in a high-dimensional latent space through a deep neural network, which makes up for the defect that single-mounted sensors can only obtain a single piece of information and cannot fully describe the target characteristics, and achieves high-accuracy identification of drone targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0031] Figure 1 This is the overall network structure diagram of the present invention;

[0032] Figure 2 This is a structural diagram of a single residual attention block of the present invention;

[0033] Figure 3 This is the channel attention layer structure diagram of the present invention;

[0034] Figure 4 This is the structure diagram of the spatial attention layer of the present invention;

[0035] Figure 5 is the optical image used in the embodiment of the present invention;

[0036] Figure 6a This is a time-frequency graph of a certain product model in the time-frequency image of the drone communication link used in the embodiment of the present invention;

[0037] Figure 6b This is a typical interference signal time-frequency image (Bluetooth, WiFi) in the UAV communication link time-frequency image used in the embodiment of the present invention;

[0038] Figure 7 This is the training loss decrease curve and recognition accuracy in the embodiment of the present invention. DETAILED DESCRIPTION

[0039] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.

[0040] The specific implementation process of the present invention is as follows:

[0041] Deep neural networks are widely used in the fields of image and speech recognition. Data often enters the network from shallow layers, and features are extracted from the data through each layer of the deep network. As the number of layers increases, the extracted features become more abstract, and ultimately a deep feature is constructed that is most suitable for completing the target task. Specifically, in order to solve the problems existing in the field of target recognition technology in the background, the present invention uses a deep neural network to combine the drone communication link signal with its optical image features to jointly identify the target. This aims to solve the current technical problem of using single-dimensional sensors in the field of drone recognition, resulting in incomplete acquisition of target features and limited recognition efficiency, and to achieve high-accuracy recognition of drones.

[0042] In the specific implementation scheme, the present invention takes into account that the recognition network needs to process the image transmission signal feature image and the optical feature image at the same time. The two types of sensors are heterogeneous, and the amount of information, information depth and high-dimensional information level contained in the two types of features are different. Therefore, the recognition network needs to be redesigned based on the characteristics of heterogeneous data.

[0043] The characteristic image of the image transmission signal is derived from the drone's downlink signal, which is a two-dimensional frequency variation image obtained after time-frequency analysis. Combined with the inherent properties of the drone downlink signal, it has significant characteristics such as time division and frequency hopping, making it suitable for drone target identification.

[0044] As for UAV optical images, they come from the optoelectronic equipment of the integrated detection node and are affected by the observation distance and weather conditions. Although optical images are close to human intuitive cognition, there are problems such as small targets and cloud cover. The pixel ratio of the UAV target in the image is too small, which is not conducive to the extraction of target information. Therefore, compared with the characteristic image of the image transmission signal, the optical image requires the design of a deeper network convolution structure to meet the actual needs of long-distance warning for key protection.

[0045] In combination with the above design principles, the network structure of the improved multi-dimensional feature recognition network based on heterogeneous data is as follows: Figure 1 The network generally adopts the design idea of ​​"divide first, then summarize". The two inputs correspond to two information extraction branches respectively. Then, the fully connected structure is used to complete the information fusion. The residual structure is combined to further improve the recognition accuracy and complete the target classification.

[0046] Specifically, if Figure 1 As shown in the figure, the network input includes an optical image (256x512x3) and its corresponding time-frequency map (256x512x1). The network output is a 1xN vector, where N is the number of classification categories, describing the probability that the target belongs to each category.

[0047] For optical imaging, the present invention feeds images into a pretrained Vgg16 network to extract optical features. The original Vgg16 network has 16 layers, and the present invention uses the first 13 layers as the feature extraction network. The training process fine-tunes the weights of the Vgg16 network. On the other hand, for time-frequency maps that carry signal features, the present invention uses a four-layer convolutional network for feature extraction, with convolution kernel sizes of 7x7, 5x5, and two 3x3 layers. Each convolution layer uses the ReLU function for activation, and employs 2x2 maximum pooling.

[0048] After the optical and time-frequency graph branches are respectively extracted through the network to obtain deep features, the present invention combines the features on both sides through cascading. Since some categories have greater differences in optical domain features from other categories, while some categories have greater differences in RF domain features from other categories, the present invention designs a three-layer ResAttentionBlock module (residual attention module) so that the network can automatically select those more useful features.

[0049] Among them, the specific structure of a single residual attention module (ResAttentionBlock) is as follows Figure 2 As shown, each residual attention module contains a convolution layer, a channel attention layer (the specific structure is as follows Figure 3 As shown) and a spatial attention layer (the specific structure is as Figure 4 As shown), the channel attention layer can help the network select optical or RF domain channels, while the spatial attention layer can help the network pay more attention to small targets in the image, such as drones in optical images or a small signal in time-frequency graphs. Subsequently, the present invention uses the residual structure to alleviate gradient explosion, so that the overall network layer can be deeper and deeper feature information can be obtained. Finally, the classification and recognition results are obtained after convolution and straightening the full connection. In terms of loss function, the present invention selects cross entropy to constrain the classification results output by the network. The formula is as follows:

[0050]

[0051] Where y is the 1xN one-hot label of the target true category, d is the 1xN vector output by the network, and N is the number of classification categories.

[0052] The technical effect of the method of the present invention is further verified as follows:

[0053] First, collect the required optical and drone link signal data, such as Figure 5 、 Figure 6a-6b As shown, the optical image and the time-frequency image are paired and combined, and the specific definitions of each category and the amount of training and testing data are shown in Table 1 below:

[0054] Table 1. Number of training and test images for each category

[0055]

[0056]

[0057] The Adam optimizer is used for training the network with a learning rate of 1e-5. Figure 7 The training loss and the accuracy on the test set of the multi-dimensional feature recognition network are given respectively. The specific accuracy of each category is shown in Table 2 below:

[0058] Table 2 Test accuracy of each category

[0059] Realistic meaning No. 13600 No. 60000 Drone Type 1 75% 88.5% Drone Type 2 80.7% 100% Drone Type 3 100% 100% Drone Type 4 100% 100% Drone Type 5 100% 100% Drone Type 6 48.1% 100% Wifi, Bluetooth interference 100% 100% No goal 35% 100% Silent drone 65% 100% car 90% 85% Colored aerial targets (kites, balloons) 55.6% 77.8% average 77.2% 95.6%

[0060] It can be seen that the network loss continues to decrease during the training process, proving that the network can achieve convergence. As the number of training times increases, the network's recognition accuracy of the test set gradually increases, and eventually a recognition accuracy of more than 95% can be achieved.

[0061] In contrast, the results of target recognition using only optical information are shown in Table 3 below. As can be seen from Table 3, the recognition probability using only optical information is 26.8%. This is because drones account for a very small percentage of pixels in optical images, making it difficult to accurately distinguish between drone types. However, slightly larger objects such as cars and balloons can be distinguished more effectively.

[0062] Table 3 Test accuracy using only optical images for classification

[0063] Realistic meaning Accuracy Drone Type 1 30.7% Drone Type 2 3.9% Drone Type 3 2% Drone Type 4 19.2% Drone Type 5 5.8% Drone Type 6 26.9% Wifi, Bluetooth interference 3.85% No goal 5% Silent drone 25% car 95% Colored aerial targets (kites, balloons) 77% average 26.8%

[0064] Table 4 shows the results of target recognition using only time-frequency information. Table 4 shows that the recognition probability using only the time-frequency characteristics of drone link signals is 66%. Time-frequency information accurately identifies drone types, due to the significant differences in the time-frequency characteristics of signals from different drone types. However, using only time-frequency information cannot distinguish targets that do not radiate electromagnetic waves, and cannot identify silent drones, cars, balloons, and other targets.

[0065] Table 4 Test accuracy of classification using only link signal time-frequency features

[0066]

[0067] Therefore, it can be seen that the method proposed in the present invention can use multiple sensor features for multi-dimensional comprehensive identification, which can effectively make up for the defect that a single sensor obtains single information and cannot fully describe the target characteristics.

[0068] The units involved in the embodiments of the present invention may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not limit the units themselves.

[0069] According to one aspect of an embodiment of the present invention, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0070] As another aspect, embodiments of the present invention further provide a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not incorporated into the electronic device. The computer-readable medium carries one or more programs, and when executed by the electronic device, the electronic device implements the methods described in the above embodiments.

Claims

1. A drone identification device based on the combination of link signals and optical features, characterized in that: include: Constructing a multidimensional feature recognition network based on heterogeneous data for identifying drones; the multidimensional feature recognition network based on heterogeneous data includes an input module, a fully connected structure module, and an output module; The input module corresponds to two information extraction branches, which are used to input the optical image and its corresponding time-frequency map respectively; information fusion is completed through the fully connected structure module, and the residual structure is combined to improve the recognition accuracy and complete the final target recognition and classification.

2. The drone identification device based on the combination of link signal and optical feature according to claim 1 is characterized in that: The input of the multidimensional feature recognition network based on heterogeneous data includes a 256x512x3 optical image and its corresponding 256x512x1 time-frequency map.

3. The drone identification device based on the combination of link signal and optical feature according to claim 1 is characterized in that: The output of the multi-dimensional feature recognition network based on heterogeneous data includes a 1xN vector, where N is the number of classification categories, which is used to describe the probability that the target belongs to each category.

4. The drone identification device based on the combination of link signal and optical feature according to claim 1 is characterized in that: It also includes an optical deep feature extraction module, which is used to input images into a pre-trained Vgg16 network for optical feature extraction, and takes the first 13 layers of the original Vgg16 network as the feature extraction network.

5. The drone identification device based on the combination of link signal and optical feature according to claim 1 is characterized in that: It also includes a deep feature extraction module for time-frequency graphs, which uses a four-layer convolutional network to extract features from time-frequency graphs that carry signal features. The convolution kernel sizes are 7x7, 5x5, and two 3x3 kernels. Each convolution layer is activated by the ReLU function and uses a 2x2 maximum pooling size.

6. The drone identification device based on the combination of link signal and optical feature according to claim 1 is characterized in that: The information fusion is completed through the fully connected structure module, and the recognition accuracy is improved by combining the residual structure, which specifically includes: After the deep features of the optical and time-frequency graph branches are extracted through the corresponding feature networks respectively, the features on both sides are combined together through cascading, and a three-layer residual attention module is designed to enable the network to automatically select more useful features. Among them, a single residual attention module contains a convolutional layer, a channel attention layer, and a spatial attention layer. The channel attention layer is used to help the network select optical or RF domain channels, and the spatial attention layer is used to help the network pay more attention to small targets in the image. The residual structure is then used to mitigate gradient explosion, making the overall network layer deeper and obtaining deeper feature information. Finally, the classification and recognition results are obtained after convolution and straightening the full connection.

7. The drone identification device based on the combination of link signal and optical feature according to claim 6 is characterized in that: It also includes a loss function design module for selecting cross entropy to constrain the classification results of the network output. The formula is as follows: Where y is the 1xN one-hot label of the target true category, d is the 1xN vector output by the network, and N is the number of classification categories.

8. A drone identification method based on the combination of link signals and optical features, characterized in that: The drone identification device based on the combination of link signals and optical features according to any one of claims 1 to 5 comprises the following steps: The heterogeneous data from two different sensors are input, and the features of the two data in the high-dimensional latent space are extracted through a neural network and combined, and finally the recognition is completed based on the combined features.

9. The drone identification method based on the combination of link signal and optical feature according to claim 8 is characterized in that: The method of inputting heterogeneous data from two different sensors, extracting features of the two data in a high-dimensional latent space through a neural network and combining them, and finally completing recognition based on the combined features, specifically includes the following sub-steps: S1, deep features are extracted from the two branches of optical and time-frequency images through corresponding neural networks; S2 combines the features from both sides through cascading and designs a three-layer residual attention module, which enables the network to automatically select the more useful features. A single residual attention module consists of a convolutional layer, a channel attention layer, and a spatial attention layer. The channel attention layer helps the network select optical or RF domain channels, and the spatial attention layer helps the network focus on small objects in the image. S3 uses the residual structure to alleviate gradient explosion, making the overall network layer deeper and obtaining deeper feature information; S4, after convolution and straightening the full connection, the classification and recognition results are obtained.

10. The drone identification method based on the combination of link signal and optical features according to claim 9 is characterized in that: Also includes sub-steps: In the loss function design, cross entropy is selected to constrain the classification results of the network output. The formula is as follows: Where y is the 1xN one-hot label of the target true category, d is the 1xN vector output by the network, and N is the number of classification categories.