Multi-modal fusion unmanned aerial vehicle identification method and system
By using multimodal fusion technology to extract and fuse features from radio, image, and sound signals, the problem of low accuracy in visual recognition technology under complex environments has been solved, achieving high reliability and high accuracy in the recognition of UAVs.
Patent Information
- Application Number
- CN202511791112.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-01-09
AI Technical Summary
Existing vision-based drone recognition technology has low accuracy in adverse weather or complex backgrounds, making it difficult to adapt to complex and ever-changing real-world application scenarios.
A multimodal fusion method is adopted to acquire radio signals, image signals and sound signals of UAVs, and use blind source separation algorithm, adaptive filtering, image super-resolution reconstruction and dual-stream feature extraction network, spectrum analysis and harmonic feature extraction to form multimodal fusion features for identification.
It significantly improves the accuracy and reliability of drone identification, enabling precise judgment in complex environments and reducing the probability of missed and false detections.
Smart Images

Figure CN121301855A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drone identification, and more specifically, to a multimodal fusion drone identification method and system. Background Technology
[0002] Driven by the wave of intelligentization and informatization, the drone industry has rapidly emerged. With its flexible, efficient, and low-cost deployment characteristics, it has deeply integrated into various fields such as agricultural plant protection, power line inspection, film and television shooting, and emergency rescue, becoming an important force in promoting the digital transformation of various industries. However, the widespread use of drones has also brought significant security risks. Illegal incursions into airport no-fly zones or the use of drones for espionage and other covert activities seriously threaten aviation safety, security, and public interests. Therefore, achieving rapid and accurate identification of drones has become a crucial issue for ensuring public safety.
[0003] Among existing technologies, vision-based drone identification technology is one of the most widely used solutions. This technology acquires images or video signals of drones through cameras, uses image processing algorithms to extract visual features such as the drone's shape, surface texture, and motion trajectory, and then uses a classification model to identify these features, thereby determining whether it is a drone and its specific category.
[0004] The existing technology has obvious defects: its recognition effect is highly dependent on lighting conditions and environmental scenes. When faced with adverse weather conditions such as night, heavy fog, and sandstorms, or when the drone is in a complex background or an obstructed scene, the clarity of the image signal will drop significantly, resulting in inaccurate visual feature extraction. This leads to a sharp decrease in recognition accuracy and an increase in the rate of missed detections and false detections, making it difficult to adapt to complex and ever-changing real-world application scenarios. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a multimodal fusion method and system for drone identification, which can significantly improve the accuracy, reliability and environmental adaptability of drone identification.
[0006] In a first aspect, embodiments of this application provide a multimodal fusion method for UAV identification, the method comprising: Acquire radio signals, image signals, and sound signals from the drone; Radio features are obtained by extracting features from the radio signal based on blind source separation algorithm and adaptive filtering; Image features are obtained by extracting features from the image signal based on image super-resolution reconstruction and a two-stream feature extraction network. Audio features are obtained by extracting features from the sound signal based on spectrum analysis and harmonic feature extraction. fusing the radio features, the image features and the audio features to form a multi-modal fusion feature; performing recognition based on the multi-modal fusion feature to obtain a UAV category recognition result.
[0007] In a second aspect, the embodiments of the present application provide a multi-modal fusion UAV recognition system, which comprises: a signal acquisition module, configured to acquire radio signals, image signals and sound signals of a UAV; a radio feature extraction module, configured to perform feature extraction on the radio signals based on a blind source separation algorithm and adaptive filtering to obtain radio features; an image feature extraction module, configured to perform feature extraction on the image signals based on image super-resolution reconstruction and a dual-stream feature extraction network to obtain image features; an audio feature extraction module, configured to perform feature extraction on the sound signals based on spectrum analysis and harmonic feature extraction to obtain audio features; a feature fusion module, configured to fuse the radio features, the image features and the audio features to form a multi-modal fusion feature; a category recognition module, configured to perform recognition based on the multi-modal fusion feature to obtain a UAV category recognition result.
[0008] In a third aspect, the embodiments of the present application provide a computer device, which comprises a processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the computer device is running, the processor and the memory communicate through the bus, and the machine readable instructions are executed by the processor to perform the steps of the multi-modal fusion UAV recognition method in any of the optional embodiments of the first aspect.
[0009] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the steps of the multi-modal fusion UAV recognition method in any of the optional embodiments of the first aspect.
[0010] The technical solutions provided by the present application include but are not limited to the following beneficial effects: The present application synchronously acquires radio, image and sound signals of a UAV, which can collect multi-source basic information of the UAV from three dimensions of communication, vision and acoustics, provide comprehensive and diverse data support for subsequent feature extraction, avoid the problem of one-sidedness of single modal signal information, and lay a foundation for multi-dimensional feature construction and fusion.
[0011] The radio signal is extracted based on a blind source separation algorithm and adaptive filtering. The blind source separation algorithm can effectively separate the target signal from the environmental electromagnetic interference in the radio signal, and the adaptive filtering can dynamically suppress the residual noise and enhance the target signal. The two work together to obtain pure and effective radio features, improve the reliability and distinguishability of the radio features, and ensure the accurate extraction of communication-related features.
[0012] The image signal is extracted based on image super-resolution reconstruction and a double-flow feature extraction network. The image super-resolution reconstruction can improve the detail clarity of low-definition images, and the double-flow feature extraction network can adapt to the feature extraction needs of different types of images (visible light, infrared). The combination of the two can accurately capture the visual features of the shape, contour, and texture of the unmanned aerial vehicle, and improve the accuracy of image feature extraction in complex scenes.
[0013] The sound signal is extracted based on spectrum analysis and harmonic feature extraction. Spectrum analysis can convert time-domain sound signals into frequency-domain signals, and harmonic feature extraction can capture periodic acoustic characteristics such as the rotation of the unmanned aerial vehicle propeller. The audio features constructed by the two work together fully reflect the acoustic nature of the unmanned aerial vehicle and enhance the uniqueness and distinguishability of the audio features.
[0014] The three types of features are fused to form multi-modal fusion features. By fusing radio, image, and audio features, the advantages of each modality can be integrated, the shortcomings of single modality features can be compensated for, a complete and three-dimensional image of the unmanned aerial vehicle can be formed, and the features can be more comprehensive and redundant, providing strong feature support for accurate identification.
[0015] Based on the multi-modal fusion features, identification is performed. Relying on multi-dimensional and comprehensive fusion features, the effective information of each modality can be fully utilized, the accuracy of judging the unmanned aerial vehicle category can be greatly improved, the probability of missing and false detection can be reduced, and the unmanned aerial vehicle category identification result can be accurately output, meeting the requirements of actual application for identification accuracy.
[0016] The above steps are progressive and work together, from multi-source signal acquisition to targeted feature extraction, to multi-modal feature fusion and final identification, forming a complete technical link. Through the orderly implementation of each step, the advantages of different modalities are effectively integrated, the limitations of single modalities are compensated for, the accuracy, reliability, and environmental adaptability of unmanned aerial vehicle identification are significantly improved, and accurate judgment of the unmanned aerial vehicle category can still be achieved in complex scenes.
[0017] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are used for reference. BRIEF DESCRIPTION OF DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart of a multimodal fusion drone identification method provided in Embodiment 1 of the present invention is shown; Figure 2 A flowchart of a radio signature determination method provided in Embodiment 1 of the present invention is shown; Figure 3 A flowchart of an image feature determination method provided in Embodiment 1 of the present invention is shown; Figure 4 A flowchart of an audio feature determination method provided in Embodiment 1 of the present invention is shown; Figure 5 A flowchart of a fusion feature determination method provided in Embodiment 1 of the present invention is shown; Figure 6 The flowchart of a multimodal fusion technology provided in Embodiment 1 of the present invention is shown; Figure 7 A flowchart of a fusion weight update method provided in Embodiment 1 of the present invention is shown; Figure 8 A flowchart of a method for determining identification results provided in Embodiment 1 of the present invention is shown; Figure 9 A schematic diagram of the hardware connection and algorithm flow of a multimodal fusion drone identification system provided in Embodiment 1 of the present invention is shown; Figure 10 This diagram illustrates the architecture of a multimodal fusion unmanned aerial vehicle (UAV) identification system provided in Embodiment 1 of the present invention. Figure 11 This diagram illustrates the structure of a multimodal fusion unmanned aerial vehicle (UAV) identification system according to Embodiment 2 of the present invention. Figure 12 A schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention is shown. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0021] Example 1 To facilitate understanding of this application, the following is combined with... Figure 1 The flowchart of the multimodal fusion drone identification method provided in Embodiment 1 of the present invention illustrates the content of Embodiment 1 in detail.
[0022] Embodiment 1 of this application provides a flowchart of a multimodal fusion method for UAV identification, see [link to flowchart]. Figure 1 As shown, Figure 1 The flowchart illustrates a multimodal fusion method for UAV identification provided in Embodiment 1 of the present invention, wherein the method includes steps S101 to S106: S101: Acquires radio signals, image signals, and sound signals from the drone.
[0023] Specifically, radio signals are acquired through the USRP B210 software radio device, covering the commonly used frequency bands of drones, 2.4GHz-2.5GHz and 5.75GHz-5.9GHz, with a sampling rate set to 20MHz; image signals are acquired through a 1080P industrial camera module that supports the MIPI-CSI interface to capture visual information such as the shape and attitude of the drone; and sound signals are acquired through a 4-channel MEMS high-sensitivity microphone array, focusing on capturing the sounds generated by the propeller rotation and motor operation.
[0024] S102: Based on the blind source separation algorithm and adaptive filtering, the radio signal is used to extract radio features.
[0025] Specifically, the blind source separation algorithm is based on the Independent Component Analysis (ICA) theory. It first whitens the multi-channel radio signals (by using eigenvalue decomposition to obtain the whitening matrix), and then uses FastICA to iteratively find the separation matrix to separate the UAV target signal from electromagnetic interference such as WiFi. The adaptive filtering uses the Least Mean Square (LMS) algorithm with a step size ranging from 10K to 48K. It dynamically adjusts the filter weights based on the real-time statistical characteristics of the noise to suppress residual interference.
[0026] S103: Image features are obtained by extracting features from the image signal based on image super-resolution reconstruction and a two-stream feature extraction network.
[0027] Specifically, the image super-resolution reconstruction adopts a lightweight model with an embedded adaptive attention mechanism (AAM), which intelligently focuses on key areas such as the edge and rotor of the UAV through a weight formula, and reduces the number of network parameters by combining reparameterization technology; the dual-stream feature extraction network includes a visible light branch (the backbone network of the MobileNetV3 lightweight convolutional neural network, which adopts depthwise separable convolution) and an infrared branch (the network of the EfficientNet high-efficiency convolutional neural network, which adapts to thermal imaging data through a composite scaling strategy), and the two work together to extract features.
[0028] S104: Based on spectrum analysis and harmonic feature extraction, the sound signal is subjected to feature extraction to obtain audio features.
[0029] Specifically, the spectrum analysis uses Fourier transform to convert the time-domain sound signal into a frequency-domain signal; harmonic feature extraction is performed by detecting parameters such as fundamental frequency, harmonic amplitude distribution, harmonic interval and harmonic energy ratio, while integrating time-domain features such as signal envelope and periodicity.
[0030] S105: The radio features, the image features, and the audio features are fused to form a multimodal fusion feature.
[0031] Specifically, the three types of features are first normalized (the normalization method is not limited), then weights are assigned through the channel attention module and the spatial attention module, and finally the 128-dimensional feature vectors of each modality are concatenated into a 384-dimensional long vector, integrating the visual, auditory and communication features to form a three-dimensional feature representation.
[0032] S106: Based on the multimodal fusion features, the drone category identification result is obtained.
[0033] Specifically, the multimodal fusion features are input into the fully connected layer and mapped to the UAV category dimension space. The Softmax function is then used to convert the features into a category probability distribution, and the confidence scores of each category are output. Finally, the UAV model classification results are obtained by filtering based on the confidence scores, thereby reducing the probability of missed detections and false detections.
[0034] In an optional implementation, see Figure 2 As shown, Figure 2 The flowchart of a radio feature determination method provided in Embodiment 1 of the present invention is shown, wherein the step of extracting radio features from the radio signal based on a blind source separation algorithm and adaptive filtering includes steps S201 to S203: S201: The radio signal is blindly separated using an independent component analysis algorithm to obtain the separated target signal.
[0035] Specifically, the USRP B210 device is used to collect multi-source radio signals (including UAV target signals and electromagnetic interference) in a complex electromagnetic environment. During the collection, the center frequency is flexibly configured according to the UAV signal frequency band (2.4GHz / 5.8GHz frequency sweep coverage) to ensure that it matches the frequency range of UAV image transmission and remote control signals. Then, the collected IQ raw data (complex form) is transmitted to the Rockchip RK3588 host through the USRP UHD driver library and organized into a multi-channel signal matrix (dimension: number of channels × number of sampling points).
[0036] Blind source separation based on Independent Component Analysis (ICA) algorithm: The first step is to perform whitening processing on the multi-channel signal matrix to obtain the whitened multi-channel signal matrix. The whitening treatment formula is: ;in, W is the original multi-channel signal matrix, and W is the whitening matrix, which is obtained by eigenvalue decomposition of the signal covariance matrix, so that the covariance matrix of the whitened signal is the identity matrix.
[0037] The second step involves iteratively searching for the separation matrix A using the FastICA algorithm. The termination condition for the iteration is not explicitly defined (it is adjusted based on the actual signal separation effect, such as reaching a preset threshold or meeting the independence requirements of the separated signals). Finally, the separated independent components are output according to the following formula, from which the UAV target signal is obtained. (Based on the signal frequency band and modulation method, interference components such as WiFi are ruled out): ;in, For the separation matrix, This is the whitened multi-channel signal matrix.
[0038] S202: The target signal is enhanced by adaptive minimum mean square filtering to obtain an enhanced signal.
[0039] Specifically, first determine the input signal and reference noise for the LMS filter: The input signal is the UAV target signal separated by S201 (still containing a small amount of residual interference). The reference noise is extracted from the interference component separated by S201 and used to estimate the statistical characteristics of environmental noise in real time. The weight vector of the LMS filter is then initialized. (Usually initialized as a zero vector or a small random vector).
[0040] The weights are updated iteratively in real time using the LMS algorithm, with the following iterative formula: ;in, Let n be the filter weight vector for the nth iteration. Let be the filter weight vector for the (n+1)th iteration. The step size ranges from 10K to 48K and is dynamically adjusted based on the ambient noise intensity. The stronger the noise, the larger the step size should be to accelerate convergence. For the first Error signal of the next iteration calculate, The desired signal is approximated by the component closest to the UAV target separated by S201; Let be the input signal vector for the nth iteration. for conjugate, The weight vector of the nth iteration filter The conjugate transpose of .
[0041] Through multiple iterations, the filter output signal gradually approximates the pure UAV signal, ultimately yielding an enhanced UAV radio signal, providing a pure data foundation for subsequent feature extraction.
[0042] S203: Extract the frequency band and modulation scheme features of the communication signal from the enhanced signal to obtain the radio characteristics.
[0043] Specifically, time-frequency analysis (such as short-time Fourier transform) is performed on the enhanced signal obtained from S202 to extract the frequency band characteristics of the signal, determine whether the signal falls within the commonly used 2.4GHz-2.5GHz or 5.75GHz-5.9GHz frequency bands of UAVs, and record parameters such as the center frequency and bandwidth of the signal.
[0044] Modulation recognition algorithms (such as those based on constellation diagrams and instantaneous amplitude / phase features) are used to analyze the modulation scheme of the enhanced signal, including ASK amplitude shift keying modulation, FSK frequency shift keying modulation, and PSK phase shift keying modulation (the modulation schemes of image transmission and remote control signals differ among different UAV models and can be used as distinguishing features). The extracted frequency band parameters are combined with the modulation scheme features to form a structured radio feature vector (128-dimensional, matching the dimensions of subsequent image and audio feature vectors for easy fusion).
[0045] In an optional implementation, see Figure 3 As shown, Figure 3 The flowchart of an image feature determination method provided in Embodiment 1 of the present invention is shown, wherein the step of extracting image features from the image signal based on image super-resolution reconstruction and a dual-stream feature extraction network includes steps S301 to S303: S301: A lightweight super-resolution network with an embedded adaptive attention mechanism is used to reconstruct the image signal to obtain a reconstructed image.
[0046] Specifically, the raw image signal captured by the camera is first preprocessed (using existing denoising algorithms such as Gaussian filtering and contrast enhancement algorithms such as histogram equalization, without improvement, only to improve the quality of the raw image).
[0047] The preprocessed image is input into a lightweight super-resolution network, which embeds an adaptive attention mechanism (AAM): AAM calculates the attention weight for each pixel (i,j) of the image using the following formula: ; in, The attention weight at coordinate (i,j) in the image (range 0~1, the larger the value, the more important the position is for detection); This is the feature vector (containing basic information such as edges and textures) of the shallow feature map of the super-resolution network at (i,j). , A learnable weight matrix (used to predict importance from features); the Sigmoid function ensures that the weights are normalized to the range of 0 to 1, avoiding weights that are too large or too small.
[0048] Higher weights are assigned to key areas such as the drone's edges, rotors, and support structures, prioritizing the enhancement of details in these areas. During network training, a multi-branch structure (residual branch + ordinary convolutional branch) is employed to improve reconstruction accuracy. During inference, branches are merged using reparameterization techniques, following a formula. Merge convolutional kernels, where, The fused convolutional kernel, The main branch convolution kernel, , For differential branch convolution kernels, This indicates a convolution operation.
[0049] Simultaneously construct the joint loss function ( For reconstruction losses, such as MSE loss; For detecting loss, such as the loss detection of the YOLO series; , The weighting coefficients are set according to the balance requirements between super-resolution quality and detection accuracy, to achieve collaborative optimization of super-resolution reconstruction and target detection, and finally output a reconstructed image with clear details (focusing on preserving key visual features of the UAV).
[0050] S302: Input the reconstructed image into a dual-stream feature extraction network, and extract features through the visible light branch and the infrared branch respectively to obtain the features extracted by the dual-stream network.
[0051] Specifically, the input to the dual-stream feature extraction network is the reconstructed image obtained from S301. First, the image resolution is standardized (the resolution size is set according to the input requirements of MobileNetV3 and EfficientNet, such as 640×640). The visible light branch uses MobileNetV3 as the backbone network, and depthwise separable convolutions are used to reduce computational complexity and achieve intra-channel feature extraction. Depthwise convolution is defined by the following formula: ;in, To output features for depthwise convolution, For input channel index, It uses a single-channel convolution kernel, which performs convolution on only a single channel of the input feature map; For the input feature map, This is the kernel index (summation range).
[0052] Point convolution is defined by the formula: ; in, The output features are generated by point convolution. It is a 1×1 convolution kernel. Input the number of channels. To output the number of channels, feature fusion between channels is achieved.
[0053] The infrared branch employs the EfficientNet network architecture, precisely balancing the number of channels and image resolution through a composite scaling strategy (simultaneously adjusting network depth, width, and input resolution). This adapts to the low-contrast characteristics of thermal imaging data and avoids blurring between drones and the background in infrared images. Channel pruning techniques are used on the output features of both branches (removing redundant channels and retaining key feature channels) to reduce redundant parameters, ultimately yielding visible light and infrared branch features, which together constitute the features extracted by the two-stream network.
[0054] S303: The features extracted by the dual-stream network are fused, and the irregular contour features of the UAV rotor and support are extracted through deformable convolutional layers to obtain the image features.
[0055] Specifically, the visible light branch features and infrared branch features extracted by the two-stream network are first fused using element-wise addition or channel concatenation to obtain preliminary fused features. These preliminary fused features are then input into a deformable convolutional layer, which predicts the offset of each convolutional kernel sampling point through additional convolutional layers. , (Offset range -1 to 1, ensuring sampling points do not deviate from the effective region of the feature map), adjust the convolution kernel sampling position using the following formula: ;in, To output feature pixels, For convolution kernel weights, To input preliminary fusion features, is the sampling point offset of the convolution kernel (-1 to 1), and k is the convolution kernel size.
[0056] Adaptively captures dynamic contours formed by the rotation of the drone rotor and irregular structures of the support frame, effectively suppressing noise interference from complex backgrounds (such as trees and buildings). Global average pooling is applied to the output features of the deformable convolutional layer, compressing the feature map into a 128-dimensional feature vector as the final image feature (with the same dimension as radio and audio feature vectors, facilitating multimodal fusion).
[0057] In an optional implementation, see Figure 4 As shown, Figure 4 The flowchart of an audio feature determination method provided in Embodiment 1 of the present invention is shown, wherein the step of extracting audio features from the sound signal based on spectrum analysis and harmonic feature extraction includes steps S401 to S404: S401: Perform a Fourier transform on the sound signal to convert the time-domain sound signal into a frequency-domain signal.
[0058] Specifically, the sound signal collected by the microphone array is first preprocessed: the sound data is transmitted to the Rockchip RK3588 platform via the I2S bus, and an adaptive noise suppression algorithm is used to filter out human voices, traffic noise, etc. in the environment, and a bandpass filter is used (the cutoff frequency is set according to the frequency range of the drone propeller sound, such as 200Hz~10kHz) to highlight the target sound.
[0059] Short-time Fourier transform is performed on the preprocessed time-domain sound signal. The number of sampling points for the Fourier transform is set to 512. A sliding window function (the window length is set according to the periodicity of the sound signal to ensure that it covers one cycle of propeller rotation) is used to divide the time-domain signal into multiple short-time segments. Fourier transform is performed on each segment, and finally the time-domain sound signal is converted into a frequency-domain signal, which intuitively presents the frequency distribution characteristics of the sound.
[0060] S402: Based on the frequency domain signal, extract the fundamental frequency and the harmonic feature set composed of harmonic amplitude distribution, harmonic interval and harmonic energy ratio parameters.
[0061] Specifically, the fundamental frequency is extracted using the autocorrelation method: the autocorrelation value of the time domain signal corresponding to the frequency domain signal is calculated at different time intervals τ. When τ is equal to the rotation period of the UAV propeller, the autocorrelation value R(τ) reaches its maximum value, and the corresponding frequency f0=1 / τ is the fundamental frequency.
[0062] The quantitative relationship between the fundamental frequency and the propeller speed was determined by manual feature lookup, with a lookup range of 10K-48K (calibrated based on experimental data of propeller speed and fundamental frequency for different models, without a fixed formula); harmonic features were extracted based on frequency domain signals.
[0063] Based on the fundamental frequency f0, the positions of each harmonic, such as 2f0, 3f0, etc., are identified, and the amplitude values of each harmonic are recorded (forming the harmonic amplitude distribution), the frequency difference between adjacent harmonics (i.e., the harmonic interval, which is theoretically equal to the fundamental frequency f0), and the ratio of each harmonic amplitude to the fundamental frequency amplitude (i.e., the harmonic energy ratio, such as the 2nd harmonic energy ratio = 2f0 amplitude / f0 amplitude). The parameters such as fundamental frequency, harmonic amplitude distribution, harmonic interval, and harmonic energy ratio are combined to form a harmonic feature set.
[0064] S403: Extract the temporal envelope features of the sound signal.
[0065] Specifically, envelope extraction is performed on the preprocessed time-domain audio signal: Hilbert transform or moving average method is used to smooth the amplitude of the time-domain signal to obtain the amplitude envelope curve (reflecting the change of sound signal intensity over time); feature parameters are extracted from the envelope curve, such as the peak value, valley value, mean, and variance of the envelope, as well as the rise time and fall time of the envelope (reflecting the rate of change of sound intensity caused by changes in propeller speed); at the same time, the periodic variation law of the time-domain signal is analyzed, and the periodic stability of the signal (such as the periodic standard deviation) is calculated as a time-domain periodic feature; the amplitude envelope parameters are combined with the periodic features to form the time-domain envelope feature.
[0066] S404: Combine the harmonic feature set with the time-domain envelope feature to construct the audio feature.
[0067] Specifically, the harmonic feature set and the temporal envelope feature are first standardized (eliminating dimensional differences between different features, such as mapping feature values to the [0,1] interval); then the standardized two types of features are concatenated to form a high-dimensional feature vector; the high-dimensional feature vector is reduced in dimensionality using principal component analysis (PCA) or linear discriminant analysis (LDA) to retain features that contribute significantly to UAV recognition, ultimately resulting in a 128-dimensional audio feature vector; this audio feature combines the frequency domain harmonic characteristics of the UAV sound (reflecting differences in the propeller's physical structure) and the temporal variation characteristics (reflecting the UAV's motion state), while also relating the UAV's physical characteristics (propeller speed) and statistical characteristics (probability distribution of harmonic features), possessing high discriminative power and suitable for subsequent multimodal fusion.
[0068] In an optional implementation, see Figure 5 As shown, Figure 5 The flowchart of a fusion feature determination method provided in Embodiment 1 of the present invention is shown, wherein fusing the radio features, the image features, and the audio features to form a multimodal fusion feature includes steps S501 to S504: S501: Normalize the radio features, the image features, and the audio features respectively.
[0069] Specifically, the normalization process is performed on the Rockchip RK3588 platform, based on the Linux system and a deep learning framework (such as TensorFlow). L2 normalization or BatchNorm normalization is used for radio feature vectors (128-dimensional), image feature vectors (128-dimensional), and audio feature vectors (128-dimensional) respectively (the method is not limited and can be selected according to the characteristics of the feature distribution).
[0070] The purpose of normalization is to eliminate the dimensional differences between different modal features. For example, the frequency band parameters (unit: GHz) in radio features are different from the pixel weights (unit: none) in image features and the amplitude values (unit: dB) in audio features. Normalization can ensure that the three types of features have an equal weight contribution basis when fusion, and avoid a certain modal feature from dominating the fusion result due to its excessive magnitude.
[0071] S502: Input the normalized modal feature vectors into the channel attention module and the spatial attention module respectively for processing to obtain the channel weights and spatial weights of each modal feature vector.
[0072] Specifically, the channel attention module and the spatial attention module are deployed in the deep learning inference framework of the RK3588 platform: For image features (including spatial dimensions), the spatial attention module generates a spatial weight map through convolutional layers and the Sigmoid activation function to highlight the spatial region where the drone is located in the image.
[0073] The channel attention module processes features of all three modalities. Taking radio features as an example, it converts the 128-dimensional feature vector into a 1-dimensional vector through global average pooling, and then generates channel weights through a fully connected layer and a sigmoid function to distinguish the importance of each feature channel to recognition. Spatial weights are generated only for image features, while radio and audio features (which have no spatial dimension) only generate channel weights. Finally, the channel weights and spatial weights of image features, the channel weights of radio features, and the channel weights of audio features are obtained.
[0074] S503: Based on the channel weights and spatial weights, the feature vectors of each modality are weighted and fused to obtain weighted radio features, weighted image features, and weighted audio features.
[0075] Specifically, the weighted fusion logic is dynamically adjusted based on environmental conditions and is triggered by the environmental sensing module (deployed on the RK3588 platform, which collects parameters such as light intensity, noise intensity, and electromagnetic interference intensity in real time): For radio characteristics, according to the formula The weighted radio characteristics are obtained by weighting. .in, To normalize radio characteristics, For radio channel weights, (For element-wise multiplication), if the detected electromagnetic interference intensity is higher than a preset threshold, the overall magnitude of the radio channel weight is reduced.
[0076] For audio features, according to the formula The weighted audio features are obtained by weighting. .in, To normalize audio features, The audio channel weights are set such that if the ambient noise intensity is detected to be higher than a preset threshold, the overall magnitude of the audio channel weights is reduced.
[0077] For image features, first apply the formula Channel-weighted image features are obtained by performing channel weighting. Then follow the formula Spatial weighting is performed to obtain weighted image features. .in, To normalize image features, Spatial weights, Image channel weights are used. If the light intensity is detected to be higher than a preset threshold (e.g., 500 lux), the magnitude of the image channel weights and spatial weights is increased; if the light intensity is lower than a preset threshold (e.g., 100 lux), the magnitude of the image weights is decreased. The weighted three types of features still maintain 128 dimensions, and the feature distribution is more focused on information that is effective for identification in the current environment.
[0078] S504: The weighted radio features, weighted image features, and weighted audio features are concatenated to obtain the multimodal fusion features.
[0079] Specifically, the stitching operation is completed in the memory of the RK3588 platform, utilizing the PCIe 4.0 interface to achieve high-speed transmission and caching of feature data; the weighted radio features (128-dimensional), weighted image features (128-dimensional), and weighted audio features (128-dimensional) are sequentially stitched together using the following formula: ); Concat is the concatenation operation. The spliced multimodal fusion features, The weighted radio characteristics, The weighted audio features The weighted image features are stitched together to obtain a fusion feature vector with dimensions of 128+128+128=384. This vector integrates the UAV's communication features (radio), visual features (image), and acoustic features (audio), constructing a complete "feature profile" of the UAV from three dimensions.
[0080] Meanwhile, during the stitching process, the time alignment of the three types of features is ensured by FPGA and Time Sensitive Network (TSN), avoiding fusion deviation caused by the time difference of feature acquisition, and ensuring that the fused features can accurately reflect the state of the UAV at the same moment.
[0081] See Figure 6 As shown, Figure 6 The flowchart of a multimodal fusion technology provided in Embodiment 1 of the present invention is shown. The core logic involves integrating perceptual information from three dimensions—visual, auditory, and communication signals—to extract feature information of the drone from different levels. This information is then processed through multimodal fusion and classification to ultimately achieve accurate determination of whether the drone is a drone and its specific model. The specific modules are detailed below: Input Layer: Multimodal Signal Acquisition and Timing Conversion. The input layer is the foundation of the system's recognition. It requires the acquisition and preprocessing of three types of core signals. Furthermore, to adapt to subsequent timing-dependent capture (such as LSTM networks), all signals must be converted into timing signals beforehand. Image signals: Acquired by an industrial camera module, focusing on the visual feature signals of the drone, including intuitive information such as the drone's shape outline, surface texture, flight attitude and motion trajectory, and converted into a time sequence in the form of "frame time-pixel feature matrix" (each frame image corresponds to a unique timestamp to ensure temporal continuity). Sound signals: Acoustic characteristic signals of the UAV are captured by a high-sensitivity microphone array. The focus is on extracting the periodic sounds generated by the propeller rotation and motor operation. After filtering out environmental background noise, the signals are converted into a time sequence in the form of sampling time-sound pressure amplitude. Radio signals: Acquired through software-defined radio equipment (such as USRP B210), the communication characteristic signals of the UAV are obtained, covering information such as the frequency band and modulation method of UAV remote control and image transmission, and organized into a time sequence in the form of "sampling time - IQ data (amplitude / phase)".
[0082] To achieve efficient and standardized processing of signals of all modalities, all signal branches (image, sound, radio) follow the core process of "feature extraction - temporal capture - weight optimization - dimensionality compression," with adjustments made only in the feature emphasis direction based on modal characteristics. The following details the process using the image branch as an example: CONV (Convolutional Layer): Local Key Feature Extraction. The convolutional layer is the starting point for feature extraction. By sliding the convolutional kernel, it performs weighted summation on local regions of the input temporal image signal to accurately capture the core visual features of the drone (such as edge contours, rotor structure, and surface texture).
[0083] The specific formula is as follows: ; in, It is a local block of the input image signal (corresponding to a local pixel region of a certain frame in the time sequence). ... The local features output by the convolutional layer (reflecting whether the region contains visual features of the target).
[0084] Max Pool: Feature Dimensionality Reduction and Enhancement. To reduce subsequent computation and avoid overfitting, while enhancing key features, the output of the convolutional layer needs to be processed by a max pooling layer: adopting a "sliding window and taking the maximum value" strategy, while reducing the size of the feature map (achieving dimensionality reduction), it retains the most representative features within the window (such as the strong response value of the rotor edge), effectively suppressing noise interference in non-critical areas.
[0085] LSTM (Long Short-Term Memory) network: Time-dependent capture. Since drones are in a dynamic flight state (such as attitude changes and position movements), it is necessary to use an LSTM network to process time-series image sequences and capture the changing patterns of features over time (such as the periodicity of rotor rotation and the continuity of drone flight trajectory).
[0086] The core calculation formula for LSTM is as follows: Input Gate: Used to control the features at the current time. (Max Pool output features) The proportion of cells entering the cell state. For input gate output, Activate the Sigmoid gate (output range 0~1, the larger the value, the more important the feature). This is the hidden state from the previous moment. Here is the learnable weight matrix of the input gate. The learnable bias term for the input gate.
[0087] Cell state: Used for long-term storage of time-series characteristic information. Output for the forget gate (controlling the cell state at the previous moment) (The retention ratio), ⊙ represents element-wise multiplication, and tanh is the activation function that maps feature values to the range of -1 to 1. This represents the cell state at the previous moment. This represents the current state of the candidate cells. Hidden state: This is used to output the temporal characteristics at the current moment. Output the output gate (controlling the proportion of cell state contribution to hidden state). This represents the current hidden state.
[0088] Attention channel (channel attention module): Feature weight optimization. Different feature channels contribute differently to drone recognition (e.g., the rotor contour channel is more important than the background texture channel). The channel attention module needs to dynamically assign weights to each feature channel: give higher weights to channels that are key to drone recognition (e.g., edge and rotor channels), and give lower weights to redundant or interfering channels (e.g., background noise channels), thereby further improving the discriminative power of the features.
[0089] Avg Pool (Average Pooling Layer): Unifies feature dimensions. To meet the dimensionality consistency requirements of subsequent multimodal fusion, the feature maps output by the channel attention module need to be globally compressed using an average pooling layer: a pooling window with the same size as the feature map is used, the average value of the features within the window is calculated, and finally a low-dimensional feature vector is obtained (the output vector of each modality in this system is 128-dimensional), laying the foundation for multimodal fusion.
[0090] It is worth noting that the sound branch focuses on extracting acoustic features such as fundamental frequency and harmonic energy ratio, while the radio branch focuses on extracting communication features such as frequency band and modulation method. The processing flow of both is completely consistent with that of the image branch, with only minor adjustments made to details such as convolution kernel parameters and LSTM time step size based on modal characteristics.
[0091] After extracting features from each modality, the three types of features are integrated through a multimodal fusion and classification module to achieve the final identification of the drone. The specific process is as follows: Feature fusion: Multi-dimensional feature integration. The 128-dimensional feature vectors output from the image, sound, and radio branches are concatenated by channel to form a multi-modal fusion feature vector with dimensions of 128+128+128=384. This vector can simultaneously cover the three core types of information of the UAV: vision, acoustics, and communication, constructing a complete feature profile of the UAV and breaking through the limitations of single-modal information.
[0092] Fully connected layer: Feature mapping to category space. The fully connected layer uses non-linear transformation to map the 384-dimensional multimodal fusion features to a dimensional space corresponding to the number of drone categories preset by the system (e.g., if the system needs to identify 8 drone models, it will be mapped to an 8-dimensional space), realizing the transformation from feature dimension to category dimension, and providing a foundation for subsequent classification.
[0093] Softmax: Category probability distribution output. The Softmax function normalizes the output of the fully connected layer, converting the feature scores of each category into a probability distribution ranging from 0 to 1. The probability value of a certain category is the confidence level of "the target to be identified belongs to this type of drone". The sum of the confidence levels of all categories is 1, which makes it easy to intuitively judge the reliability of the recognition results.
[0094] The core advantage of this system lies in its multimodal complementarity: when a certain modal signal fails or its performance degrades due to environmental interference, other modal signals can promptly supplement the information, ensuring the stable operation of the recognition system. For example, in hazy weather, image signals are blurred (visual modality performance degrades), while sound signals (capturing propeller acoustic features) and radio signals (acquiring communication frequency band information) can continue to provide effective features; in noisy urban environments, when the sound modality is interfered with, image and radio modes can support recognition. This complementarity significantly improves the system's robustness in complex scenarios such as low altitude, multiple electromagnetic interferences, and severe weather, meeting the high requirements of accuracy and stability in practical applications.
[0095] In an optional implementation, see Figure 7 As shown, Figure 7 The flowchart of a fusion weight update method provided in Embodiment 1 of the present invention is shown, wherein the method further includes steps S701 to S705: S701: During the feature fusion stage, environmental condition parameters are obtained through the environmental perception module.
[0096] Specifically, the environmental perception module is integrated into the Rockchip RK3588 platform and works in conjunction with the multi-source sensor module: The system analyzes light intensity from image signals captured by a camera (calculated using the average grayscale value; a grayscale value >150 indicates strong light, <50 indicates weak light). It also calculates noise intensity from ambient noise signals captured by a microphone array (using sound pressure level (SPL) in dB; an SPL >60dB indicates a noisy environment). Furthermore, it analyzes electromagnetic interference (EMI) from electromagnetic signals captured by the USRP B210 device (using signal-to-noise ratio (SNR); an SNR <10dB indicates strong interference). Simultaneously, it acquires ambient temperature using a temperature sensor (integrated into the heat dissipation module) to determine whether the heat dissipation module needs to be activated, indirectly ensuring the stability of feature fusion. The system integrates parameters such as light intensity, noise intensity, EMI intensity, and ambient temperature into an environmental condition parameter set, which is transmitted in real-time to the weight adjustment module.
[0097] S702: Based on the environmental condition parameters, dynamically adjust the weight calculation process in the channel attention module and the spatial attention module.
[0098] Specifically, the weight adjustment module modifies the calculation parameters of the channel attention and spatial attention modules based on the correlation between environmental condition parameters and the recognition performance of each modality. If the electromagnetic interference intensity is high (e.g., SNR < 10dB), the regularization coefficient of the fully connected layer is increased in the radio channel weight calculation (e.g., the L2 regularization coefficient is adjusted from 0.001 to 0.01), reducing the weight of the interference feature channel. If the environmental noise intensity is high (e.g., SPL > 60dB), the window size of the global average pooling is increased in the audio channel weight calculation (e.g., adjusted from 5×5 to 7×7) to smooth the impact of noise on the weight calculation. If the illumination intensity is low (e.g., grayscale mean < 50), the convolution kernel size is increased in the image spatial weight calculation (e.g., adjusted from 3×3 to 5×5) to expand the receptive field of spatial attention and capture drone features over a wider area. The adjustment process is performed in real time, updating the environmental condition parameters every 50ms and simultaneously updating the weight calculation parameters to ensure the module adapts to dynamic environmental changes.
[0099] S703: When the light intensity is higher than a preset threshold, control the attention module to increase the fusion weight of image features.
[0100] Specifically, the preset threshold for light intensity is determined experimentally and is typically set to 500 lux (which can be adjusted according to the camera's sensitivity). When the environment perception module detects that the light intensity is >500 lux, the weight adjustment module sends a control signal to the image channel attention module to increase the overall scaling factor of the image channel weights (e.g., from 1.0 to 1.5), thereby increasing the overall weight value of each channel of the image features by 50%.
[0101] At the same time, control signals are sent to the image spatial attention module to lower the threshold of the Sigmoid activation function (e.g., from 0.5 to 0.3), so that more spatial regions are identified as "important regions" and the coverage of spatial weights is increased.
[0102] Through dual adjustments, the weight of image features in the fusion process is increased from the initial 1 / 3 to about 1 / 2, making full use of the high resolution advantage of image features under sufficient lighting (such as clear details of drone shape and model) to improve recognition accuracy.
[0103] S704: When the light intensity is lower than a preset threshold, control the attention module to increase the fusion weight of sound features.
[0104] Specifically, when the environmental perception module detects that the light intensity is <100 lux (a preset low light threshold that can be calibrated), the weight adjustment module sends a control signal to the audio channel attention module to increase the overall scaling factor of the audio channel weight (e.g., from 1.0 to 1.8), thereby increasing the overall weight value of each audio feature channel by 80%; at the same time, the scaling factor of the radio channel weight is reduced (e.g., from 1.0 to 0.8) (low light environments are often accompanied by fog and haze, which may indirectly enhance electromagnetic interference).
[0105] After adjustment, the weight of audio features in the fusion process increased from the initial 1 / 3 to more than 1 / 2. This leverages the unique characteristics of drone propeller sound (unaffected by lighting conditions) to compensate for the blurriness of image features, ensuring stable recognition in low-light environments. For example, in foggy weather, the image may not clearly show the drone, but the audio features can still distinguish the drone model through harmonic energy ratios.
[0106] S705: When any modal signal is detected to be missing or the signal-to-noise ratio is lower than a set threshold, the attention module is controlled to reallocate the fusion weights of the remaining modal features.
[0107] Specifically, the determination of missing modal signals is based on the sensor data transmission status (e.g., if the camera is disconnected, the image signal is determined to be missing; if the USRP has no data output, the radio signal is determined to be missing). The signal-to-noise ratio (SNR) threshold is set separately for different modalities: the SNR threshold for radio signals is set to 10dB, the SNR threshold for image signals is set to 15dB, and the SNR threshold for audio signals is set to 8dB.
[0108] When a mode failure is detected (such as missing image signal), the weight adjustment module immediately triggers the weight redistribution logic: the weight ratio (1 / 3) of the failed mode is evenly distributed to the remaining two valid modes. For example, when the image is missing, the weight ratio of radio features is increased from 1 / 3 to 1 / 2, and the weight ratio of audio features is increased from 1 / 3 to 1 / 2.
[0109] If two modal failures are detected (such as radio and image missing), the weight percentage (2 / 3) of the failed modality is fully allocated to the remaining valid modality, and the weight scaling factor is adjusted to 3 times. This redistribution logic utilizes the redundancy of multi-source information to reflect the fault tolerance of the system and ensures that recognition can still be completed based on the remaining modality when one or two modal failures occur.
[0110] In an optional implementation, see Figure 8 As shown, Figure 8 The flowchart of a method for determining an identification result provided in Embodiment 1 of the present invention is shown, wherein the step of identifying the UAV category based on the multimodal fusion features includes steps S801 to S803: S801: Input the multi-modal fusion features into a fully connected layer for non-linear transformation to map the fusion features into a dimensional space corresponding to the number of UAV categories.
[0111] Specifically, the fully connected layer is deployed in the deep learning inference framework of the Rockchip RK3588 platform and consists of two hidden layers: the number of neurons in the first hidden layer is set to 256, and the activation function is ReLU; the number of neurons in the second hidden layer is set to 128, and the activation function is also ReLU; the number of neurons in the output layer is the same as the number of UAV categories (for example, if there are 8 categories such as consumer-grade quadcopters, industrial-grade fixed-wing aircraft, multi-rotor plant protection UAVs, etc., then the number of neurons is 8), and there is no activation function; through the non-linear transformation of the fully connected layer, the 384-dimensional fusion features are mapped into an 8-dimensional category space, and each dimension corresponds to the score of a category of UAV.
[0112] S802: Process the output of the fully connected layer through the Softmax function to convert the scores of each category into a probability distribution, and obtain the confidence of each UAV category.
[0113] Specifically, the Softmax function normalizes the score vector z output by the fully connected layer. The normalization result is the confidence that the input fusion feature F_fusion belongs to the k-th category of UAV, with a range of 0-1, and the sum of the confidences of all categories is 1). The confidence value reflects the matching degree between the fusion feature and the feature template of this category of UAV, and the higher the value, the higher the matching degree.
[0114] S803: Select the category with the highest confidence as the final UAV recognition result to complete the classification and recognition of UAV models.
[0115] Specifically, first select the maximum value P from the confidence distribution output by Softmax max and the corresponding category k max ; set a confidence reliability threshold T (calibrated based on experimental data, usually set to 0.3): if P max ≥T, then output "The recognition result is the k-th max category of UAV" (such as "consumer-grade quadcopter UAV"); if P max <T, then output "No valid UAV recognized" or "疑似第k max category of UAV" (to avoid misjudgment due to too low confidence).
[0116] To enhance generalization ability, the training dataset includes multimodal data from different scenarios (strong light, heavy fog, noisy cities, etc.) and different drone models, enabling the model to learn general multimodal feature patterns. After deployment, the weight parameters of the fully connected layer will be updated offline every preset period based on feedback from actual scenarios (such as the deviation in the recognition results of new drone models), ensuring the ability to recognize new drone models and unfamiliar complex environments, and ultimately achieving accurate classification and recognition of drone models in complex low-altitude environments.
[0117] The multimodal fusion UAV recognition method provided in this application has a corresponding hardware architecture and algorithm process carrier, which fully supports the implementation of the method from signal acquisition and feature extraction to cross-modal fusion.
[0118] See Figure 9 As shown, Figure 9 This diagram illustrates the hardware connection and algorithm flow of a multimodal fusion drone identification system provided in Embodiment 1 of the present invention. At the hardware level, a 2.4 / 5.8 GHz dual-band antenna is connected to an Ettus USRP B210, an industrial camera is connected via a MIPI serial communication interface, and a MEMS digital microphone is connected via a USB 3.0 interface. All three are connected to the RK3588 platform, forming the hardware core for signal acquisition and computation. At the algorithm level, the RK3588 platform runs the core algorithms of each module, extracting features from radio, image, and sound signals and forming time series data. After cross-modal fusion, the system outputs a category, completing drone identification.
[0119] See Figure 10 As shown, Figure 10 The diagram shows the architecture of a multimodal fusion drone identification system provided in Embodiment 1 of the present invention. The central processing system (host platform) communicates with the USRP B210 radio receiving device, the microphone array digital microphone, and the 1080P camera module through the communication system. The power supply system provides power support to the central processing system. The modules work together to complete the multimodal identification function of the drone.
[0120] The two diagrams, from hardware architecture to algorithm flow, completely connect the logic of the multimodal fusion drone recognition system: Hardware modules: Figure 10 The "Central Processing System (Host Platform)" is... Figure 9 The "RK3588 platform"; Figure 10 The "Radio Receiver USRP B210" corresponds to Figure 9 The "2.4 / 5.8 dual-band antenna - USRP B210"; Figure 10 The "microphone array digital microphone" corresponds to Figure 9"Digital microphone - USB connection"; Figure 10 The "1080P camera module" corresponds to Figure 9 "Industrial Camera - MIPI Connection".
[0121] Functional Link: Figure 10 The "communication system" is Figure 9 The connection channel between various hardware components and the RK3588 platform (USB, MIPI, etc.) is responsible for data transmission; Figure 10 The "Central Processing System" is running Figure 9 The "core algorithms of each module" achieve a closed loop from hardware acquisition to algorithm recognition by extracting radio, image, and sound features, processing time series data, and fusing cross-modal data to finally output the category.
[0122] Power supply logic: Figure 10 Although the "power supply system" is not in Figure 9 Presented directly in the middle, but for Figure 9 The system provides power to all hardware modules (USRP B210, RK3588 platform, camera, microphone, etc.), which is the foundation for system operation.
[0123] Example 2 See Figure 11 As shown, Figure 11 A schematic diagram of a multimodal fusion unmanned aerial vehicle (UAV) identification system according to Embodiment 2 of the present invention is shown, wherein the system includes: The signal acquisition module 1101 is used to acquire the radio signals, image signals and sound signals of the UAV; The radio feature extraction module 1102 is used to extract radio features from the radio signal based on a blind source separation algorithm and adaptive filtering. Image feature extraction module 1103 is used to extract image features from the image signal based on image super-resolution reconstruction and a two-stream feature extraction network. The audio feature extraction module 1104 is used to extract audio features from the sound signal based on spectrum analysis and harmonic feature extraction. The feature fusion module 1105 is used to fuse the radio features, the image features, and the audio features to form a multimodal fusion feature; The category recognition module 1106 is used to perform recognition based on the multimodal fusion features to obtain the UAV category recognition result.
[0124] In an optional implementation, the step of extracting radio features from the radio signal based on the blind source separation algorithm and adaptive filtering includes: The radio signal is blind source separation is performed using an independent component analysis algorithm to obtain the separated target signal; The target signal is enhanced by applying adaptive least mean square filtering to obtain the enhanced signal; The frequency band and modulation scheme features of the communication signal are extracted from the enhanced signal to obtain the radio characteristics.
[0125] In an optional implementation, the image feature extraction process based on the image super-resolution reconstruction and dual-stream feature extraction network to obtain image features includes: A lightweight super-resolution network with an embedded adaptive attention mechanism is used to reconstruct the image signal to obtain a reconstructed image; The reconstructed image is input into a dual-stream feature extraction network, and feature extraction is performed through the visible light branch and the infrared branch respectively to obtain the dual-stream network extraction. The features extracted by the dual-stream network are fused, and the irregular contour features of the UAV rotor and support are extracted through deformable convolutional layers to obtain the image features.
[0126] In an optional implementation, the step of extracting audio features from the sound signal based on spectrum analysis and harmonic feature extraction includes: Perform a Fourier transform on the sound signal to convert the time-domain sound signal into a frequency-domain signal; Based on the frequency domain signal, the fundamental frequency and the harmonic feature set composed of harmonic amplitude distribution, harmonic spacing and harmonic energy ratio parameters are extracted; Extract the temporal envelope features of the sound signal; The harmonic feature set is combined with the temporal envelope feature to construct the audio feature.
[0127] In an optional implementation, fusing the radio features, the image features, and the audio features to form a multimodal fused feature includes: The radio features, the image features, and the audio features are each normalized. The normalized modal feature vectors are input into the channel attention module and the spatial attention module respectively for processing to obtain the channel weights and spatial weights of each modal feature vector; Based on the channel weights and spatial weights, the feature vectors of each modality are weighted and fused to obtain weighted radio features, weighted image features, and weighted audio features; The weighted radio features, weighted image features, and weighted audio features are concatenated to obtain the multimodal fusion features.
[0128] In an optional implementation, the system further includes a fusion weight update module for: During the feature fusion stage, environmental condition parameters are obtained through the environmental perception module; Based on the environmental condition parameters, the weight calculation process in the channel attention module and the spatial attention module is dynamically adjusted. When the light intensity is higher than a preset threshold, the attention module is controlled to increase the fusion weight of image features; When the light intensity is below a preset threshold, the attention module is controlled to increase the fusion weight of sound features; When any modal signal is detected to be missing or the signal-to-noise ratio is lower than a set threshold, the attention module is controlled to reallocate the fusion weights of the remaining modal features.
[0129] In an optional implementation, the identification based on the multimodal fusion features to obtain the UAV category identification result includes: The multimodal fusion features are input into a fully connected layer for nonlinear transformation to map the fusion features to a dimensional space corresponding to the number of UAV categories. The output of the fully connected layer is processed by the Softmax function to convert the scores of each category into a probability distribution, thereby obtaining the confidence level of each drone category; The category with the highest confidence level is selected as the final drone identification result to complete the classification and identification of drone models.
[0130] Example 3 Based on the same application concept, see [link / reference] Figure 12 As shown, Figure 12 A schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention is shown, wherein, as Figure 12 As shown, the computer device 1200 provided in Embodiment 3 of this application includes: The system includes a processor 1201, a memory 1202, and a bus 1203. The memory 1202 stores machine-readable instructions executable by the processor 1201. When the computer device 1200 is running, the processor 1201 and the memory 1202 communicate via the bus 1203. The machine-readable instructions are executed by the processor 1201 to perform the steps of the multimodal fusion UAV identification method shown in Embodiment 1 above.
[0131] Example 4 Based on the same concept, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the multimodal fusion drone identification method described in any of the above embodiments.
[0132] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0133] The computer program product for multimodal fusion drone identification provided in this embodiment of the invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0134] The multimodal fusion UAV identification system provided in this invention embodiment can be specific hardware on a device or software or firmware installed on the device. The system provided in this invention embodiment has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment section can be referred to the corresponding content in the aforementioned method embodiment. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.
[0135] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some communication interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0137] In addition, the functional units in the embodiments provided by the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0138] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0139] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0140] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A multimodal fusion method for UAV identification, characterized in that, The method includes: Acquire radio signals, image signals, and sound signals from the drone; Radio features are obtained by extracting features from the radio signal based on blind source separation algorithm and adaptive filtering; Image features are obtained by extracting features from the image signal based on image super-resolution reconstruction and a two-stream feature extraction network. Audio features are obtained by extracting features from the sound signal based on spectrum analysis and harmonic feature extraction. The radio features, image features, and audio features are fused to form a multimodal fusion feature; Based on the multimodal fusion features, the drone category identification result is obtained.
2. The method according to claim 1, characterized in that, The process of extracting radio features from the radio signal based on blind source separation algorithm and adaptive filtering includes: The radio signal is blind source separation is performed using an independent component analysis algorithm to obtain the separated target signal; The target signal is enhanced by applying adaptive least mean square filtering to obtain the enhanced signal; The frequency band and modulation scheme features of the communication signal are extracted from the enhanced signal to obtain the radio characteristics.
3. The method according to claim 1, characterized in that, The image features are obtained by extracting features from the image signal using an image super-resolution reconstruction and dual-stream feature extraction network, including: A lightweight super-resolution network with an embedded adaptive attention mechanism is used to reconstruct the image signal to obtain a reconstructed image; The reconstructed image is input into a dual-stream feature extraction network, and feature extraction is performed through the visible light branch and the infrared branch respectively to obtain the dual-stream network extraction. The features extracted by the dual-stream network are fused, and the irregular contour features of the UAV rotor and support are extracted through deformable convolutional layers to obtain the image features.
4. The method according to claim 1, characterized in that, The process of extracting audio features from the sound signal based on spectrum analysis and harmonic feature extraction includes: Perform a Fourier transform on the sound signal to convert the time-domain sound signal into a frequency-domain signal; Based on the frequency domain signal, the fundamental frequency and the harmonic feature set composed of harmonic amplitude distribution, harmonic spacing and harmonic energy ratio parameters are extracted; Extract the temporal envelope features of the sound signal; The harmonic feature set is combined with the temporal envelope feature to construct the audio feature.
5. The method according to claim 1, characterized in that, The process of fusing the radio features, the image features, and the audio features to form a multimodal fusion feature includes: The radio features, the image features, and the audio features are each normalized. The normalized modal feature vectors are input into the channel attention module and the spatial attention module respectively for processing to obtain the channel weights and spatial weights of each modal feature vector; Based on the channel weights and spatial weights, the feature vectors of each modality are weighted and fused to obtain weighted radio features, weighted image features, and weighted audio features; The weighted radio features, weighted image features, and weighted audio features are concatenated to obtain the multimodal fusion features.
6. The method according to claim 5, characterized in that, The method further includes: During the feature fusion stage, environmental condition parameters are obtained through the environmental perception module; Based on the environmental condition parameters, the weight calculation process in the channel attention module and the spatial attention module is dynamically adjusted. When the light intensity is higher than a preset threshold, the attention module is controlled to increase the fusion weight of image features; When the light intensity is below a preset threshold, the attention module is controlled to increase the fusion weight of sound features; When any modal signal is detected to be missing or the signal-to-noise ratio is lower than a set threshold, the attention module is controlled to reallocate the fusion weights of the remaining modal features.
7. The method according to claim 1, characterized in that, The identification based on the multimodal fusion features to obtain the drone category identification result includes: The multimodal fusion features are input into a fully connected layer for nonlinear transformation to map the fusion features to a dimensional space corresponding to the number of UAV categories. The output of the fully connected layer is processed by the Softmax function to convert the scores of each category into a probability distribution, thereby obtaining the confidence level of each drone category; The category with the highest confidence level is selected as the final drone identification result to complete the classification and identification of drone models.
8. A multimodal fusion unmanned aerial vehicle (UAV) identification system, characterized in that, The system includes: The signal acquisition module is used to acquire the radio signals, image signals, and sound signals of the UAV; A radio feature extraction module is used to extract radio features from the radio signal based on a blind source separation algorithm and adaptive filtering. The image feature extraction module is used to extract image features from the image signal based on image super-resolution reconstruction and a two-stream feature extraction network. The audio feature extraction module is used to extract audio features from the sound signal based on spectrum analysis and harmonic feature extraction. The feature fusion module is used to fuse the radio features, the image features, and the audio features to form a multimodal fusion feature; The category recognition module is used to perform recognition based on the multimodal fusion features to obtain the UAV category recognition result.
9. A computer device, characterized in that, include: The system includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the multimodal fusion drone identification method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the multimodal fusion drone identification method as described in any one of claims 1-7.
Citation Information
Cited By
Battery internal temperature monitoring method and system based on multi-feature decoupling and deep learning
CN121541072A