Dynamic open space unmanned aerial vehicle identification method and device based on multi-view feature fusion
Through the multi-view feature fusion and deep feature extraction methods, the accuracy problem of drone recognition in dynamic open scenarios is solved, and the accurate identification and discrimination of known and unknown drones are achieved, which improves the recognition accuracy rate.
Patent Information
- Application Number
- CN202510530674.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Existing drone recognition technology is difficult to accurately identify known drones and unknown drones in dynamic open scenarios, and the method based on a single feature view cannot fully extract the details and local features of the drone signal.
A dynamic open space drone recognition method for multi-view feature fusion is proposed. By generating three different view features from a multi-domain perspective, the time domain, frequency domain, and time frequency domain, and designing a three-branch multi-view deep feature extraction and fusion recognition network, combining the twin network framework and joint loss function for optimization training, the comprehensive and deep feature extraction is achieved.
Accurate identification of known drones and accurate identification of unknown drones are achieved, the accuracy of identification of open space drones is improved, and the category identification and classification accuracy of features are ensured.
Smart Images

Figure CN120071202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV identification, and particularly to a method and device for identifying UAVs in a dynamic open space by fusing multi-view features. Background Art
[0002] With the development of UAV technology, its application fields have gradually expanded and are widely used in various fields such as aerial photography, agricultural monitoring, and entertainment activities. While UAVs provide convenience for people, the characteristics of UAVs such as being small-sized, easy to operate, flying at low altitudes, having a slow speed, and good concealment have become tools for illegal intrusion, illegal snooping, and even destruction.
[0003] Currently, there are mainly four ways to detect and identify UAVs, namely detection based on audio sound waves, detection based on computer vision images, detection based on radar, and detection based on radio frequency signals. The detection method based on audio sound waves identifies UAVs according to the unique propeller rotation noise of different UAVs, that is, audio fingerprints. However, when in a complex environment with a high noise level, and small and medium-sized UAVs have a small sound, the detection effect is poor. The detection method based on visual images detects and identifies UAV targets in the image by taking images of the monitored area, including visible light images and infrared images, etc. However, as small targets, UAVs are prone to missed detection and false detection. When the weather environment is poor and the visibility is low at night, such as in foggy weather, high-definition images of UAVs cannot be captured. Moreover, UAVs fly at a low altitude and have a slow speed, and are battery-powered, making their infrared characteristics not obvious. The active detection method based on radar uses the echo signal generated when the emitted electromagnetic wave encounters a target during propagation to detect the presence of the target. However, the radar detection system is expensive and is only effective for large and fast-moving targets. UAVs fly at a low speed and low altitude and have a small radar cross-section, so the radar detection method has a problem of low detection accuracy for UAV targets. The detection method based on radio frequency signals collects the radio frequency signals emitted by UAVs through radio frequency sensors and analyzes the characteristics of the radio frequency signals to achieve UAV identification. It is a passive detection method, which is more concealed, not affected by occlusion such as vegetation and buildings, and has the advantages of high sensitivity, real-time monitoring, low cost, small computation, and no echo interference. It is not affected by the material characteristics of the UAV body and weather illumination, and the detection conditions are relatively good.
[0004] In practical applications, lawbreakers often select UAV devices of the same brand and the same model as legal UAVs for illegal activities. These illegal UAVs disguise their own ID numbers, MAC addresses and other identity information, bypass the authentication methods based on protocols, and because their appearance is exactly the same as that of legal UAVs, methods based on radar, optoelectronics and infrared, and sound waves cannot successfully detect them.
[0005] In this case, the method for identifying drones based on radio frequency signals has more advantages. As a radiation source device, each drone, as a unique radiation source individual, has unique hardware characteristics hidden in the transmitted signal. The inherent radio frequency fingerprint of each radiation source device is unique, stable, and non-forgeable, that is, it is very difficult to be tampered with and replicated (even forging identity information such as MAC and IP addresses). Therefore, hardware fingerprint information can be extracted from the transmitted radio frequency signal to accurately perceive and characterize the identity of drone devices.
[0006] Currently, most of the research on drone identification based on radio frequency signals is based on the closed-set assumption, that is, the categories of drone devices to be identified belong to the categories in the model training set. However, illegal drone devices are unknown to the identification system. In the application of the real electromagnetic environment, it is mostly a dynamic open scenario, and new unknown drone devices may appear at any time. It is almost impossible to cover all drone categories in the training library. Therefore, it is necessary to break the limitation of the closed-set identification assumption that leads to incorrect identification of unknown radiation sources, and ensure the accurate identification of known drone categories while realizing the identification of new unknown drone categories. In addition, most of the existing drone identification methods rely on feature extraction from a single feature map, and the feature extraction is not comprehensive enough and may lose key and detailed information. There is an urgent need to study a method suitable for accurate perception of drone identity and accurate identification of unknown drones in a dynamic open scenario.
[0007] In the application with the patent application number 202411415469.1 and the patent name "Open-set Identification Method and System for Drone Signals Based on Metric Learning", an open-set identification method for drones based on deep neural networks and metric learning is proposed. However, this application only extracts features from a single feature view of the STFT time-frequency diagram, and the time-frequency diagram obtained by STFT does not have good time-domain and frequency-domain resolutions, lacking in extracting details, local features, and feature comprehensiveness. In addition, the improved KNN unknown discrimination method based on distance thresholds is restricted by the feature space distribution. If the extracted features are not accurate enough or the distributions of different categories in the feature space are relatively dispersed and not aggregated, there is uncertainty in unknown category discrimination and the identification performance cannot be guaranteed.
[0008] "BISSIAM: Bispectrum Siamese Network Based Contrastive Learning for UAV Anomaly Detection" proposed in IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING converts UAV signals into bispectrum features as inputs, uses a contrastive learning model based on the Siamese network to learn vector encodings, and proposes a similarity-based fingerprint matching mechanism to detect unknown UAVs. However, it only utilizes a single feature view, only the bispectrum feature map, and cannot refine the extraction of detailed features. Based on the symmetric loss and mutual information loss of image-enhanced bispectrum pairs of the same sample, it only assists the neural network in extracting class-related features and cannot constrain the spatial distribution of features of different class samples. Moreover, this fingerprint matching method is essentially still based on distance and threshold, assuming that the distances between points within a class will not be greater than the maximum distance. If the feature space distribution of known class samples is not sufficiently concentrated or is irregularly distributed, misjudgments are likely to occur based on the distance threshold, and unknown classes may be misjudged as known classes. Summary of the Invention
[0009] To solve the above technical problems, the present invention proposes a multi-view feature fusion dynamic open space UAV recognition method and device. The specific technical solutions are as follows:
[0010] The multi-view feature fusion dynamic open space UAV recognition method includes:
[0011] Step 1: Characterize the same sample from multiple domain perspectives for the obtained known-class UAV remote control signals to generate three different views: time domain, frequency domain, and time-frequency domain;
[0012] Step 2: Design a three-branch multi-view deep feature extraction and fusion recognition network to extract deep feature information of the three different views respectively, and fuse the output features of the three-branch multi-view deep feature extraction and fusion recognition network;
[0013] Step 3: Use the Siamese network framework to optimize and train the three-branch multi-view deep feature extraction and fusion recognition network by combining center loss, contrastive loss, and classification cross-entropy loss;
[0014] Step 4: Based on the trained three-branch multi-view deep feature extraction and fusion recognition network, propose an open-set recognition algorithm based on the boundary model to achieve accurate discrimination of unknown UAVs.
[0015] The multi-view feature fusion dynamic open space UAV recognition device includes:
[0016] The multi-view feature representation generation module characterizes the same sample from multiple domain perspectives for the obtained known-class UAV remote control signals, generating three different views: time domain, frequency domain, and time-frequency domain.
[0017] The multi-branch multi-view deep feature extraction and fusion recognition module; includes a three-branch multi-view deep feature extraction and fusion recognition network, which respectively extracts deep feature information of three different views, and fuses the output features of the three-branch multi-view deep feature extraction and fusion recognition network.
[0018] The multi-branch network optimization training module based on the siamese network and joint loss function; uses the siamese network framework to optimize and train the three-branch multi-view deep feature extraction and fusion recognition network jointly with center loss, contrast loss, and classification cross-entropy loss.
[0019] The open-set recognition module based on the boundary model; based on the trained three-branch multi-view deep feature extraction and fusion recognition network, proposes an open-set recognition algorithm based on the boundary model to achieve accurate discrimination of unknown UAVs.
[0020] An electronic device includes: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned multi-view feature fusion dynamic open-space UAV recognition method.
[0021] A computer-readable storage medium stores executable instructions thereon, and when the instructions are executed by a processor, the processor implements the above-mentioned multi-view feature fusion dynamic open-space UAV recognition method.
[0022] The present invention has the following beneficial effects:
[0023] (1) A multi-view feature extraction method is proposed. The original signal is transformed to extract features from three different perspectives: time-domain IQ diagram, power spectral density, and Hilbert time-frequency spectrogram, representing UAV signal features from multiple domains and multiple angles in the time domain, frequency domain, and time-frequency domain, achieving comprehensive and multi-level feature characterization. Moreover, the Hilbert time-frequency spectrogram has good time-domain resolution and frequency resolution, capable of finely characterizing details and local characteristics.
[0024] (2) For the three different view features, deep networks are respectively designed for deep feature extraction, and a three-branch deep feature extraction and fusion recognition network of "time-domain IQ diagram - 2DCNN", "power spectral density - 1DCNN", and "Hilbert time-frequency spectrogram - improved lightweight ResNet" is proposed. There are associations and differences between different view features of the same sample. The extraction and fusion of the three view features can realize the full utilization of effective information from different views and the complementary advantages of information.
[0025] (3) It is proposed to use the Siamese network structure for network training, and optimize the feature extraction network by combining the center loss, the Siamese network contrast loss, and the classification cross-entropy loss, so that the feature distribution output by the three-branch feature extraction and fusion network meets the requirements of "minimizing the distance between classes and maximizing the distance within classes", making the feature distributions of each UAV category more concentrated around the category center and the boundaries clearer. At the same time, it is ensured that the features are class-discriminable and the classification is accurate.
[0026] (4) Based on the trained three-branch feature extraction and fusion recognition network, use the deep fusion features to fit the Weibull boundary model, and use the OpenMax algorithm to correct the output scores of each class of the network to obtain the output results of each known class and unknown class, realizing the accurate discrimination of unknown UAVs. At the same time, the recognition ability of known UAVs is ensured, and the recognition accuracy of UAVs in open space is effectively improved. Brief Description of the Drawings
[0027] Figure 1 It is the flowchart of the present invention;
[0028] Figure 2 It is the schematic diagram of the three-branch multi-view deep feature extraction and fusion network;
[0029] Figure 3 It is the schematic diagram of the Siamese network and the combined loss;
[0030] Figure 4 It is the diagram of the multi-view feature representation result;
[0031] Figure 5 It is the multi-view fusion feature distribution diagram of 5 known-class UAVs;
[0032] Figure 6 It is the multi-view fusion feature distribution diagram of 5 known-class UAVs and 3 unknown classes;
[0033] Figure 7 It is the result diagram of the confusion matrix of the recognition results of known-class and unknown-class UAVs;
[0034] Figure 8 It is the recognition result diagram of known-class and unknown-class UAVs under different open-set degrees. Detailed Embodiment
[0035] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above object, the present invention adopts the following technical solutions.
[0036] The present invention proposes a dynamic open space UAV recognition method based on multi-view feature fusion. The specific process is as follows: Figure 1 As shown below, it includes:
[0037] Step 1: For the obtained known-class UAV remote control signals, characterize the same sample from multiple domains to generate three different views: time domain, frequency domain, and time-frequency domain. Characterizing the same sample from multiple domains includes: time-domain IQ diagram, power spectral density, and Hilbert time-frequency spectrum. Figure 3 These different perspective features are used to achieve comprehensive and multi-level feature characterization based on multi-view features.
[0038] Time-domain IQ diagram: The time-domain IQ signal is directly displayed on the amplitude-time plane, and the plane image is directly used as a view in the time domain.
[0039] Power spectral density: The time-domain IQ signal is Fourier-transformed to obtain the frequency-domain representation , and the logarithmic power spectral density is calculated through the following formula :
[0040] (1)
[0041] Hilbert time-frequency spectrum diagram: First, use the variational mode decomposition (VMD) algorithm to decompose the time-domain IQ signal into time-modal components. Perform the Hilbert transform (HT) on each decomposed time-modal component to obtain the Hilbert time-frequency spectrum diagram, as shown in the following formula:
[0042] (2)
[0043] Among them, and are respectively the instantaneous amplitude and instantaneous frequency of the th time-modal component , is the phase function, is 's Hilbert transform. According to the instantaneous amplitude and instantaneous frequency of each time-modal component, the Hilbert spectrum can be obtained, where is the Hilbert spectrum of each time-modal. The operation represents taking the real number operation, represents the time-axis sampling points, is the frequency-axis frequency points.
[0044] Based on its marginal spectrum, obtain the energy concentration region, and intercept the region containing useful information in the Hilbert time-frequency spectrogram to obtain a sub-spectrogram. The Hilbert time-frequency spectrogram has good time-domain resolution and frequency resolution, and can finely characterize details and local characteristics.
[0045] Step 2: Design a three-branch multi-view deep feature extraction and fusion recognition network to extract three different view feature information respectively. Propose a 2D convolutional neural network for the time-domain IQ diagram, a 1D convolutional network for the power spectral density, and an improved lightweight ResNet for the Hilbert time-frequency spectrogram. Fusion the output features of the three-branch multi-view deep feature extraction and fusion recognition network. The three-branch multi-view deep feature extraction and fusion recognition network is as Figure 2 shown. Input the time-domain IQ diagram into the first branch network, which successively includes a first-level convolutional layer (convolution kernel size k1×k1, feature channel number 16), a pooling layer, a batch normalization layer, a RELU activation layer, a second-level convolutional layer (convolution kernel size k2×k2, feature channel number 32), a pooling layer, a batch normalization layer, a RELU activation layer, and a fully connected layer 1-1; Input the power spectral density feature into the second branch network, which successively includes a first-level convolutional layer (convolution kernel size 1×k3, feature channel number 16), a pooling layer, a batch normalization layer, a RELU activation layer, a second-level convolutional layer (convolution kernel size 1×k4, feature channel number 32), a pooling layer, a batch normalization layer, a RELU activation layer, and a fully connected layer 2-1; Input the Hilbert time-frequency spectrogram into the third branch network, which successively includes a first-level convolutional layer (convolution kernel size k5×k5, feature channel number 32), a pooling layer, a batch normalization layer, a RELU activation layer, a residual structure, a pooling layer, a RELU activation layer, and a fully connected layer 3-1. The residual structure includes a second-level convolutional layer (convolution kernel size k6×k6, feature channel number 32), a batch normalization layer, a RELU activation layer, and a third-level convolutional layer (convolution kernel size k7×k7, feature channel number 32). The input of the second-level convolutional layer and the third-level convolutional layer are added as the output of the residual structure; Fusion connect the outputs of the three branch networks to form a multi-view fusion feature. Make full use of the effective information of different views, and then perform UAV category recognition through two fully connected classification layers. Among them, the three-branch multi-view deep feature extraction and fusion network structure is denoted as and the fully connected classification layer is denoted as , see Figure 3 for the multi-branch network part of which contains a two-dimensional convolutional network, a one-dimensional convolutional network, and an improved ResNet network. The views Figure 1 , view Figure 2 , view Figure 3 respectively pass through and the outputs of the three networks in , and then through a fully connected classification layer including a fully connected layer 1 and a fully connected layer 2 , and the predicted label of the UAV category is obtained through the SoftMax function.
[0046] Step 3: Use the siamese network framework to optimize and train the above three-branch multi-view deep feature extraction and fusion recognition network by combining the center loss, contrast loss, and classification cross-entropy loss, making full use of the association and difference information between different views of the UAV signal samples, so that the three-branch multi-view deep feature extraction and fusion network can accurately represent the signal features of the UAV device, and make the feature distributions of each UAV category more concentrated around the category center and the feature distribution boundaries clearer. The siamese network framework and the combined loss function are shown in Figure 3 , the siamese network framework includes two identical core networks. The multi-view features of sample 1 and sample 2 of the known class sample pair are respectively input into the two core networks, and the predicted label 1, predicted label 2, and the fusion features of the two networks obtained through the two core networks , calculate the center loss, siamese network contrast loss, and classification cross-entropy loss.
[0047] Take the above three-branch deep feature extraction and fusion recognition network and the fully connected classification layer as the two identical core networks of the siamese network. Randomly select two known class UAV samples as a sample pair (the two samples may belong to the same category, that is, the positive sample pair label is set to 1, or they may not belong to the same category, that is, the negative sample pair label is set to 0), and input the two samples in the sample pair into the two core networks of the siamese network respectively;
[0048] Calculate the center loss using the multi-view fusion features output from the core network. The expression is as follows:
[0049] (3)
[0050] Where is the number of samples in the training data set, is the multi-view fusion feature output from the core network corresponding to a certain sample, represents the category label of the nth sample, then is the feature center of the nth class.
[0051] Calculate the contrast loss using the fusion features output from the sample pair. The expression is as follows:
[0052] (4)
[0053] Where represents the distance between the features of the two samples in the ith sample pair, that is:
[0054] (5)
[0055] is the number of sample pairs in the training dataset, represents the two samples in the i-th sample pair, represents the label indicating whether the two samples belong to the same class. If they belong to the same class, then the value is 1. If they do not belong to the same class, then the value is 0. is the set threshold. If the distance between the features of two samples from different classes exceeds the threshold, then the loss value is very small and can be regarded as 0.
[0056] Use the fully connected classification layer of the core network and the output of the SoftMax function to calculate the cross-entropy loss. Let the true class label of the n-th sample be , and the predicted label of the core network be . The expression of the classification cross-entropy loss function is as follows:
[0057] (6)
[0058] Among them, is the number of known classes, represents the probability that the n-th sample belongs to the c-th class, satisfying and .
[0059] Jointly optimize and train the core feature extraction network in the siamese network using three loss functions. The joint loss function is expressed as shown in Equation (7). By the center loss and the contrastive loss to constrain the distribution of the feature space, optimize the distance between different samples in the feature space, so that the output feature distribution is as concentrated as possible around the respective class feature centers, satisfying "minimizing the inter-class distance and maximizing the intra-class distance". Through the cross-entropy loss , make the features class-discriminable and ensure accurate classification.
[0060] (7)
[0061] Among them, are all weight factors greater than 0 and less than 1, used to control the proportion of each loss function in the joint loss function.
[0062] Step 4. Based on the trained three-branch feature extraction and fusion recognition network, an open-set recognition algorithm based on the boundary model is proposed to accurately discriminate unknown drones while ensuring the recognition ability of known drones, effectively improving the recognition accuracy of drones in open spaces. The open-set recognition algorithm includes:
[0063] Step 4.1. Based on the feature of known-class drone samples, fit the Weibull boundary distribution and establish the Weibull boundary model;
[0064] Step 4.2. Calculate the classification probability of the test sample, correct the probability of the known class based on the Weibull boundary model, and calculate the probability that it belongs to the unknown class;
[0065] Step 4.3. Discriminate unknown drones and recognize known-class drones according to the probability.
[0066] Specifically, Step 4.1 is as follows: Use the multi-view fusion feature output by the core network to fit the Weibull distribution and construct the Weibull boundary model.
[0067] Activation vector and mean activation vector: The output of the second fully connected layer in the fully connected classification layer in the core network is called the activation vector (AV). For each known class, calculate the mean of the activation vectors of all correctly classified samples in the training set to obtain the mean activation vector (MAV) of each class, which represents the center of the sample feature space distribution of this class, and is expressed as follows:
[0068] (8)
[0069] where, is the mean activation vector of class c, is the number of correctly classified samples in class c, is the activation vector of the j-th sample in class c;
[0070] Distance set and Weibull distribution: Calculate the Euclidean distance , that is, the distance between the activation vector of the j-th sample and the mean activation vector MAV of class c, to form the distance set of this class. Use the Weibull distribution in extreme value theory to fit the distance set of each known class and construct the Weibull boundary model. The Weibull distribution is a probability distribution used to describe extreme value events and can well characterize the extreme values in the distance set. Its probability density function and cumulative distribution function (CDF) are expressed as:
[0071] (9)
[0072] (10)
[0073] Among them, is the scale parameter, is the shape parameter, represents a certain input distance value.
[0074] Step 4.2 is specifically as follows: Calculate the classification probability of the test sample based on the OpenMax algorithm, correct the probability of the known class based on the Weibull boundary model, and calculate the probability that it belongs to the unknown class.
[0075] The OpenMax algorithm is specifically as follows: First, through the trained three-branch feature extraction and fusion recognition network, output the test sample corresponding fully connected layer output score vector , calculate to the center of each known class , and obtain . Based on the Weibull distribution of each known class, calculate the corrected score, and the formula is as follows:
[0076] (11)
[0077] Among them, is the distance between the test sample and the mean activation vector MAV of class c, and are respectively the scale parameter and the shape parameter in the Weibull distribution parameters of class c.
[0078] The score that the test sample belongs to the unknown class is the sum of the difference between the original probability score and the corrected score:
[0079] (12)
[0080] Normalize the scores of predicting the test sample as each class and the unknown class, and map them to classification probabilities through SoftMax, that is:
[0081] (13)
[0082] Step 4.3 is specifically as follows: If the maximum value in the above classification probability is greater than the threshold, then the test sample is judged as this class; if the maximum probability value is less than the threshold, then it is judged as an unknown class. Specific Example 1:
[0084] Step 1: Select the remote control (RC) signals of 17 different drones from the public dataset as the recognition objects. Select 5 categories as known categories, with 900 samples in each category. Obtain their multi-view feature representations, namely time-domain IQ diagram, power spectral density, and Hilbert time-frequency spectrum respectively. Figure 3 Schematic diagrams of different perspective features are as Figure 4 , Figure 4 In (a) of Figure 4 is the time-domain IQ diagram, Figure 4 in (b) of
[0085] is the power spectral density, and Figure 2 in (c) of
[0086] is the Hilbert time-frequency spectrum diagram. Figure 3 Step 2: Design a three-branch multi-view deep feature extraction and fusion recognition network to extract three different view feature information respectively. A 2D convolutional neural network is proposed for the time-domain IQ diagram, a 1D convolutional network is proposed for the power spectral density feature, and an improved lightweight ResNet is proposed for the Hilbert time-frequency spectrum diagram. The output features of the three-branch network are fused. The structure of the three-branch multi-view deep feature extraction and fusion network is as Figure 5 shown. Then, the multi-view fusion features are input into two fully connected layers for classification.
[0087] Step 3: Use the Siamese network framework to optimize and train the above three-branch multi-view deep feature extraction and fusion recognition network by combining center loss, contrast loss, and classification cross-entropy loss, as Figure 6 shown. Make full use of the correlation and difference information between different views of the drone signal samples, so that the three-branch multi-view deep feature extraction and fusion network can accurately represent the signal features of the drone equipment, and make the feature distributions of each drone category gather more around the category center, and the boundaries of the feature distributions are clearer. The feature distributions of 5 known drone categories are as Figure 7 shown. It can be seen that the features of different categories are respectively gathered around their own centers. Figure 8
[0087] Step 4: Based on the trained three-branch multi-view deep feature extraction and fusion recognition network, propose an open-set recognition algorithm based on the Weibull boundary model to achieve accurate discrimination of unknown drones, while ensuring the recognition ability of known drones, and effectively improving the recognition accuracy of drones in the open space. Figure 6 shows the t-SNE multi-view fusion feature distribution diagram of 5 known categories and 3 selected categories as unknown drone categories. It can be seen that the feature boundaries between each known category and unknown category are clear. Figure 7 is the confusion matrix of the recognition results. It can be seen that the samples of each known category can be correctly classified, and the samples of the unknown category are also correctly recognized as unknown. According to the ratio of the number of known drones and unknown drones, set different open-set degrees. The recognition results of known categories and unknown categories under different open-set degrees are as Figure 8As shown, it can be seen that the method of the present invention can maintain a high recognition accuracy under different open-set degrees. As the open-set degree increases, that is, the number of unknown classes increases, the recognition rate of the method of the present invention decreases slightly. When the open-set degree is 0, it means that all test samples belong to known classes and do not contain unknown classes. The open-set degree is calculated as follows:
[0088] (14)
[0089] wherein, represents the number of known classes used for training, represents the number of UAV classes (known + unknown) for testing, is the number of UAV classes to be recognized.
[0090] In the present invention, in addition to the time-domain IQ diagram, power spectral density, and Hilbert time-frequency spectrogram in the multi-view, other feature views can also be used, such as differential constellation diagram (DCTF), texture feature vector based on histogram of oriented gradients (HOG), multi-dimensional entropy feature sequence, etc., which all have a positive effect on accurately characterizing UAV features.
[0091] For the deep feature extraction network for different view features, certain adjustments and changes to the specific structure of each branch network still have the ability to represent features.
[0092] The contrast loss of the dual network structure of the Siamese network is proposed to minimize the intra-class distance and maximize the inter-class distance, and the triplet loss obtained by the triple core network structure (the above multi-view multi-branch deep feature extraction and fusion recognition network as three identical core networks of the Siamese network) can also achieve the constraint effect on the feature space distribution.
[0093] In addition to constructing the boundary model with the Weibull distribution, the Gumbel distribution, Fréchet distribution, etc. in extreme value theory can also be used.
[0094] The present invention also provides a multi-view feature fusion dynamic open space UAV recognition device, including:
[0095] A multi-view feature representation generation module, which represents the same sample from multiple domain perspectives for the obtained remote control signals of known-class UAVs, and generates three different views: time domain, frequency domain, and time-frequency domain;
[0096] A multi-branch multi-view deep feature extraction and fusion recognition module; including a three-branch multi-view deep feature extraction and fusion recognition network, which respectively extracts deep feature information of three different views, and fuses the output features of the three-branch multi-view deep feature extraction and fusion recognition network;
[0097] Multi-branch network optimization training module based on Siamese network and joint loss function; using the Siamese network framework, jointly optimizing and training the three-branch multi-view deep feature extraction and fusion recognition network with center loss, contrastive loss, and categorical cross-entropy loss;
[0098] Open-set recognition module based on boundary model; based on the trained three-branch multi-view deep feature extraction and fusion recognition network, proposing an open-set recognition algorithm based on the boundary model to achieve accurate discrimination of unknown drones.
[0099] The present invention also provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a multi-view feature fusion dynamic open space drone recognition method of the present invention.
[0100] The present invention also provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor implements a multi-view feature fusion dynamic open space drone recognition method of the present invention.
Claims
1. A method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features, characterized in that: include: Step 1: For the known drone remote control signals obtained, characterize the same sample from a multi-domain perspective and generate three different views: time domain, frequency domain, and time-frequency domain; Step 2: Design a three-branch multi-view deep feature extraction and fusion recognition network to extract deep feature information of three different views respectively, and fuse the output features of the three-branch multi-view deep feature extraction and fusion recognition network; Step 3: Using the twin network framework, the center loss, contrast loss and classification cross entropy loss are combined to optimize the training of the three-branch multi-view deep feature extraction and fusion recognition network; Step 4: Based on the trained three-branch multi-view deep feature extraction and fusion recognition network, an open set recognition algorithm based on the boundary model is proposed to achieve accurate identification of unknown drones.
2. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 1 is characterized in that: The three different views of time domain, frequency domain, and time-frequency domain include: Time domain IQ diagram: The time domain IQ signal Directly display in the amplitude-time plane, and use the plane image directly as a view in the time domain; Power spectral density: Perform Fourier transform on the time domain IQ signal to obtain the frequency domain representation , the logarithmic power spectral density is calculated by the following formula : (1) Hilbert time-frequency spectrum: First, the variational mode decomposition algorithm is used to transform the time domain IQ signal Decompose into time modal components, for each time modal component after decomposition Perform Hilbert transform to obtain the Hilbert time spectrum, as shown below: (2) in, and Respectively Time modal components The instantaneous amplitude and instantaneous frequency, is the phase function, yes According to the instantaneous amplitude and instantaneous frequency of each time modal component, the Hilbert spectrum can be obtained. ,in is the Hilbert spectrum of each time mode, Operation means real number operation. Indicates the time axis sampling point, is the frequency point on the frequency axis.
3. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 2 is characterized in that: A 2D convolutional neural network is proposed for the time domain IQ graph, a 1D convolutional network is proposed for the power spectral density, and an improved lightweight ResNet is proposed for the Hilbert time-frequency spectrum graph. The three-branch multi-view deep feature extraction and fusion network is specifically as follows: the time domain IQ graph is input into the 2D convolutional neural network, i.e., the first branch network, which includes the first-level convolution layer, the convolution kernel size is k1×k1, the number of feature channels is 16, the pooling layer, the batch normalization layer, the RELU activation layer, the second-level convolution layer, the convolution kernel size is k2×k2, the number of feature channels is 32, the pooling layer, the batch normalization layer, the RELU activation layer, and the fully connected layer; the power spectral density feature is input into the 1D convolutional network, i.e., the second branch network, which includes the first-level convolution layer, the convolution kernel size is 1×k3, the number of feature channels is 16, the pooling layer, the batch normalization layer, the RELU activation layer, the second-level convolution layer , the convolution kernel size is 1×k4, the number of feature channels is 32, the pooling layer, the batch normalization layer, the RELU activation layer, and the fully connected layer; the Hilbert time-frequency spectrum is input into the light ResNet, i.e., the third branch network, which includes the first-level convolution layer, the convolution kernel size is k5×k5, the number of feature channels is 32, the pooling layer, the batch normalization layer, the RELU activation layer, the residual structure, the pooling layer, the RELU activation layer, and the fully connected layer in sequence, wherein the residual structure includes the second-level convolution layer, the batch normalization layer, the RELU activation layer, and the third-level convolution layer, the convolution kernel size of the second-level convolution layer is k6×k6, the number of feature channels is 32, the convolution kernel size of the third-level convolution layer is k7×k7, the number of feature channels is 32, and the outputs of the second-level convolution layer and the third-level convolution layer are added as the output of the residual structure; the outputs of the three branch networks are fused and connected to form multi-view fusion features.
4. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 3 is characterized in that: In step 3, the twin network framework uses the three-branch multi-view deep feature extraction and fusion recognition network as the core network, including the three-branch multi-view deep feature extraction and fusion network structure and the fully connected classification layer. The three-branch multi-view deep feature extraction and fusion network structure is recorded as , the fully connected classification layer is recorded as ,The fully connected classification layer includes two fully connected layers.
5. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 4 is characterized in that: The three-branch multi-view deep feature extraction and fusion recognition network is used as two completely identical core networks of the twin network, two known drone samples are randomly selected as sample pairs, and the two samples in the sample pairs are respectively input into the two core networks of the twin network; The center loss is calculated using the multi-view fusion features output from the core network, and the expression is as follows: (3) in, is the number of samples in the training dataset, is the multi-view fusion feature output in the core network corresponding to a certain sample, represents the category label of the nth sample, The first The characteristic center of the class; The contrast loss is calculated using the fusion features of the sample pair output. The expression is as follows: (4) in, Represents the distance between the two sample features in the i-th sample pair, that is: (5) is the number of sample pairs in the training dataset, represents the two samples in the i-th sample pair, Indicates whether two samples belong to the same category. If belong to the same category, then The value is 1, if If they do not belong to the same category The value is 0, is the set threshold; Using fully connected classification layer Calculate the cross entropy loss with the output of the SoftMax function. The true category label of the samples is , the core network predicts the label , the classification cross entropy loss function expression is as follows: (6) in, is the number of known categories, Represents the probability that the nth sample belongs to the cth class, satisfying and ; Combine the three loss functions to optimize the core network in the twin network. It is expressed as shown in formula (7); (7) in, are all weight factors.
6. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 1, characterized in that: Open set recognition algorithms include: Step 4.1, based on the known characteristics of the drone samples, fit the Weibull boundary distribution and establish the Weibull boundary model; Step 4.2, calculate the classification probability of the test sample, correct the known class probability based on the Weibull boundary model, and calculate the probability of it belonging to the unknown class; Step 4.3: Identify unknown drones and known drones based on probability.
7. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 6, characterized in that: Step 4.1 includes: The fully connected classification layer The output of the second fully connected layer in is called the activation vector AV. For each known category, the mean of the activation vectors of all correctly classified samples in the training set is calculated to obtain the mean activation vector MAV of each category, which represents the center of the feature space distribution of samples in this category, expressed as follows: (8) in, is the mean activation vector for category c, is the number of correctly classified samples in category c, is the activation vector of the jth sample in category c; Calculate the Euclidean distance between the activation vectors of all correctly classified samples in each known category and the mean activation vector MAV of that category , that is, the distance between the jth sample activation vector and the mean activation vector MAV of category c, forming the distance set of this category. The Weibull distribution is used to fit the distance set of each known category to construct the Weibull boundary model. Its probability density function and cumulative distribution function are expressed as: (9) (10) in, is the scale parameter, is the shape parameter, Represents an input distance value.
8. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 7, characterized in that: In step 4.2, the classification probability of the test sample is calculated using the OpenMax algorithm, which is specifically: First, the test samples are output through the trained three-branch multi-view deep feature extraction and fusion recognition network The corresponding fully connected layer outputs the score vector ,calculate To the center of each known category The distance is , based on the Weibull distribution of each known category, the corrected score is calculated as follows: (11) in, is the mean activation vector MAV distance from the test sample to category c, and are the scale parameter and shape parameter of the Weibull distribution parameters of category c respectively; The fraction of the test sample that belongs to the unknown category is the original probability score Sum the differences from the corrected scores: (12) The scores of the test samples predicted as various categories and unknown categories are normalized and mapped to classification probabilities through SoftMax, that is: (13)。 9. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 8, characterized in that: Step 4.3 includes: Classification probability If the maximum value in is greater than the threshold, the test sample is judged as the category c; if the maximum probability value is less than the threshold, it is judged as an unknown class.
10. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 1, characterized in that: In step 1, the view may also include a differential constellation diagram, a texture feature vector based on a gradient direction histogram, and a multi-dimensional entropy feature sequence.
11. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 1, characterized in that: The twin network can also be a triple core network structure.
12. The method for identifying unmanned aerial vehicles in dynamic open spaces by fusion of multi-view features according to claim 6, characterized in that: Weibull distribution can also be Gumbel distribution or Fréchet distribution.
13. A dynamic open space drone recognition device based on multi-view feature fusion, characterized in that: include: The multi-view feature representation generation module characterizes the same sample from multiple domains for the known drone remote control signals obtained, generating three different views: time domain, frequency domain, and time-frequency domain. Multi-branch multi-view deep feature extraction and fusion recognition module; including a three-branch multi-view deep feature extraction and fusion recognition network, which extracts deep feature information of three different views respectively, and fuses the output features of the three-branch multi-view deep feature extraction and fusion recognition network; A multi-branch network optimization training module based on a twin network and a joint loss function; using a twin network framework, the three-branch multi-view deep feature extraction and fusion recognition network is optimized and trained by combining center loss, contrast loss and classification cross entropy loss; Open set recognition module based on boundary model; based on the trained three-branch multi-view deep feature extraction and fusion recognition network, an open set recognition algorithm based on boundary model is proposed to achieve accurate identification of unknown drones.
14. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a dynamic open space drone identification method with multi-view feature fusion as described in any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by the processor, the processor implements the dynamic open space drone identification method with multi-view feature fusion as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Unmanned aerial vehicle signal open set identification method and system based on metric learning
CN119441717A
Open set radiation source individual identification method based on deep learning
CN111914919A
Radar radiation source individual open set identification method and system
CN114970638A
Air low and slow small target identification method based on time-frequency data analysis
CN116206218A
Video face recognition method and system based on multistage perception self-encoding network
CN116486457A
Cited By
Open set radio frequency fingerprint identification method based on feature space optimization
CN120892930A
Communication signal open set identification method based on three-channel time-frequency fusion
CN121125415A