An Unmanned Aerial Vehicle Detection System and Method Based on Multimodal Fusion

By combining radio frequency and acoustic detection methods in the UAV detection system and using deep learning technology to perform multimodal fusion, the problem of difficulty in achieving accurate, effective and reliable UAV detection is solved, and the accurate positioning and wide application of UAVs is achieved.

CN114417908BActive Publication Date: 2025-06-10SHAANXI ZHONGWEI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111446081.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-06-10
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing drone detection technology is difficult to achieve accurate, effective and reliable drone detection in a single mode, especially in complex urban electromagnetic environments.

Method used

The UAV detection system based on multimodal fusion is adopted, combined with RF and acoustic detection methods, and the RF fingerprint and voiceprint extraction, classification recognition and positioning are used to realize modal classification and recognition using deep learning convolutional neural networks, and the fractional layer decision-making method is used to fusion to ultimately achieve accurate positioning of the UAV.

Benefits of technology

It realizes accurate, effective and reliable detection of drones, overcomes the limitations of single-modal detection, and is suitable for a variety of application scenarios, such as airports, prisons and military management areas, and has broad market prospects and actual economic value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114417908B_ABST
    Figure CN114417908B_ABST
Patent Text Reader

Abstract

The present invention discloses a drone detection system based on multimodal fusion, which includes a system host and multiple node modules arranged in the detection area. The node module includes a radio frequency antenna array, a microphone sensor array, a radio frequency signal processing module, an audio signal processing module, a fusion module and a communication module. Among them, the radio frequency antenna array is used to obtain the radio frequency signal of the drone target; the microphone sensor array is used to obtain the audio signal of the drone target; the radio frequency signal processing module is used to obtain the classification recognition and positioning results of the radio frequency signal; the audio signal processing module is used to obtain the classification recognition and positioning results of the audio signal; the fusion module is used to perform multimodal fusion on the classification recognition and positioning results of the radio frequency signal and the classification recognition and positioning results of the audio signal. The present invention solves the limitation problem of single-modal drone detection and the implementation cost problem of multimodal detection, and realizes accurate, effective and reliable detection of drones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) detection, and particularly relates to a UAV detection system and method based on multimodal fusion. Background Art

[0002] In order to meet the control requirements of UAVs in complex urban electromagnetic environments, the detection and early warning technical means of UAVs can be divided into technical categories such as vision, radar, radio frequency (RF), and acoustics. Visual detection generally performs 3D reconstruction of the UAV position through the use of a camera array and multi-camera views to obtain the motion trajectory of the UAV. The effective detection distance of the UAV is within 1 km; radar detection includes active radar detection and passive radar detection. Among them, active radar detection measures the azimuth of the UAV by transmitting electromagnetic waves and detecting the echo reflected from the UAV, while passive radar detection measures the UAV azimuth by directly detecting the environmental electromagnetic waves reflected by the UAV. The radar has a long operating range, but is costly, and civilian small UAVs are small in size and have a small reflection surface. Under the policy requirements of complex urban electromagnetic environments and radiation control, it is very difficult to use radar to detect UAVs; RF detection generally uses an antenna array to receive the communication signals between the UAV and the ground station or remote controller, and realizes the classification and identification of UAVs through RF fingerprint detection methods. The effective operating range can reach more than 3 km, meeting the requirements of most application scenarios. However, after the UAV sets a preset flight plan and is in a radio silent state, the RF detection method will fail; acoustic detection receives the sound emitted by the UAV rotor through a microphone array, and can detect and identify UAVs near the ground or near buildings. Its implementation cost is relatively low, but due to the large attenuation of sound during transmission, the effective detection distance through acoustic detection is generally small.

[0003] In summary, each UAV detection technology has its limitations, and it is difficult to accurately, effectively, and reliably detect UAVs using a single detection modality. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a UAV detection system and method based on multimodal fusion. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] One aspect of the present invention provides a UAV detection system based on multimodal fusion, including a system host and a plurality of node modules arranged in the detection area. The node module includes an RF antenna array, a microphone sensor array, an RF signal processing module, an audio signal processing module, a fusion module, and a communication module. Among them,

[0006] The RF antenna array is used to obtain the RF signals of UAV targets in the detection area;

[0007] The microphone sensor array is used to obtain the audio signal of the UAV target in the detection area;

[0008] The RF signal processing module is used to extract RF fingerprints, classify and identify, and locate the RF signal, and obtain the classification and identification and location results of the RF signal;

[0009] The audio signal processing module is used to extract voiceprints, classify and identify, and locate the audio signal, and obtain the classification and identification and location results of the audio signal;

[0010] The fusion module is used to perform multi-modal fusion on the classification and identification and location results of the RF signal and the classification and identification and location results of the audio signal, and obtain a fusion result including UAV type and location information;

[0011] The communication module is used to send the fusion result to the system host.

[0012] In an embodiment of the present invention, both the RF antenna array and the microphone sensor array are arranged in a planar five-element cross array and are located on the same horizontal plane.

[0013] In an embodiment of the present invention, the RF signal processing module includes an RF signal conversion unit, an RF fingerprint extraction unit, an RF classification and identification unit, and an RF positioning unit, wherein,

[0014] The RF signal conversion unit is used to convert the analog RF signal received by the RF antenna array into a digital RF signal;

[0015] The RF fingerprint extraction unit is used to obtain UAV RF fingerprint information according to the digital RF signal;

[0016] The RF classification and identification unit is used to classify and identify the UAV RF fingerprint information based on a multi-layer convolutional neural network model, and obtain the category information of the UAV target;

[0017] The RF positioning unit is used to obtain the location information of the UAV target according to the digital RF signal.

[0018] In an embodiment of the present invention, the RF fingerprint extraction unit is specifically used for:

[0019] Plot the digital RF signal in a two-dimensional vector diagram to obtain a constellation trajectory diagram; perform differential operation on the digital RF signal at a predetermined interval to obtain a differential constellation trajectory diagram; extract a Haar-like feature vector diagram from the differential constellation trajectory diagram as the UAV RF fingerprint information.

[0020] In one embodiment of the present invention, the audio signal processing module includes an audio signal conversion unit, a voiceprint extraction unit, an audio classification and recognition unit, and an audio positioning unit, wherein,

[0021] The audio signal conversion unit is used to convert the analog audio signal received by the microphone into a digital audio signal;

[0022] The voiceprint extraction unit is used to obtain the drone voiceprint information according to the digital audio signal;

[0023] The audio classification and recognition unit is used to classify and recognize the drone voiceprint information based on a CNN and GRU composite neural network model to obtain the type information of the drone target;

[0024] The audio positioning unit is used to obtain the position information of the drone target according to the digital audio signal.

[0025] In one embodiment of the present invention, the fusion module specifically includes:

[0026] A calculation unit, which is used to calculate the similarity HD between the radio frequency fingerprint feature recognition result and the feature template respectively using the Hamming distance a and the similarity HD between the voiceprint feature recognition result and the feature template b ;

[0027] A weighted fusion unit, which is used to fuse radio frequency data and audio data based on the weighted addition principle at the score layer:

[0028] FS = w a ×(1 - HD a ) + w b ×(1 - HD b )

[0029] wherein, FS represents the similarity with the feature template after fusion, w a , w b respectively represent the weight coefficients of the two modalities of radio frequency and audio, and w a + w b = 1.

[0030] Another aspect of the present invention provides a drone detection method based on multi-modal fusion, including:

[0031] S1: Obtain the radio frequency signal of the drone target in the detection area using a radio frequency antenna array;

[0032] S2: Obtain the audio signal of the drone target in the detection area using a microphone sensor array, wherein both the radio frequency antenna array and the microphone sensor array are arranged in a planar five-element cross array and are located on the same horizontal plane;

[0033] S3: Extract the RF fingerprint, classify and identify, and locate the RF signal to obtain the classification, identification, and location results of the RF signal;

[0034] S4: Extract the voiceprint, classify and identify, and locate the audio signal to obtain the classification, identification, and location results of the audio signal;

[0035] S5: Perform multi-modal fusion on the classification, identification, and location results of the RF signal and the classification, identification, and location results of the audio signal to obtain the fusion results including the UAV type and location information.

[0036] In an embodiment of the present invention, the S3 specifically includes:

[0037] S31: Convert the analog RF signal received by the RF antenna array into a digital RF signal;

[0038] S32: Obtain the UAV RF fingerprint information according to the digital RF signal;

[0039] S33: Classify and identify the UAV RF fingerprint information based on a multi-layer convolutional neural network model to obtain the category information of the UAV target;

[0040] S34: Obtain the location information of the UAV target according to the digital RF signal.

[0041] In an embodiment of the present invention, the S4 includes:

[0042] S41: Convert the analog audio signal received by the microphone into a digital audio signal;

[0043] S42: Obtain the UAV voiceprint information according to the digital audio signal;

[0044] S43: Classify and identify the UAV voiceprint information based on a CNN and GRU composite neural network model to obtain the type information of the UAV target;

[0045] S44: Obtain the location information of the UAV target according to the digital audio signal.

[0046] In an embodiment of the present invention, the S5 includes:

[0047] S51: Use the Hamming distance to calculate the similarity HD between the RF fingerprint feature recognition result and the feature template a and the similarity HD between the voiceprint feature recognition result and the feature template b ;

[0048] S52: Fuse the radio frequency data and the audio data based on the weighted addition principle at the fractional layer:

[0049] FS = w a ×(1 - HD a ) + w b ×(1 - HD b )

[0050] Wherein, FS represents the similarity with the feature template after fusion, w a , w b respectively represent the weight coefficients of the two modalities of radio frequency and audio, and w a + w b = 1.

[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0052] 1. The UAV detection system based on multi-modal fusion of the present invention solves the limitations of single-modal UAV detection and the implementation cost of multi-modal detection, realizes accurate, effective and reliable detection of UAVs, and can be widely promoted and applied in the market.

[0053] 2. The UAV detection method based on multi-modal fusion of the present invention is based on radio frequency and acoustic detection methods. By establishing a UAV radio frequency and audio fingerprint library, it uses a deep learning convolutional neural network to realize the classification and recognition of the two modalities of UAV radio frequency detection and acoustic detection, and uses a fractional layer decision method to realize the fusion of the two modalities; it uses the time difference of arrival (TDOA) positioning method to realize the accurate positioning of UAVs; the research results can be widely applied to application scenarios that require UAV control, such as airports, prisons, military management areas, important event venues, and nuclear industrial facilities, and have broad market prospects and actual economic value.

[0054] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Brief Description of the Drawings

[0055] Figure 1 is a module diagram of a UAV detection system based on multi-modal fusion provided by an embodiment of the present invention;

[0056] Figure 2 is a module schematic diagram of a node module provided by an embodiment of the present invention;

[0057] Figure 3 is an installation structure schematic diagram of a radio frequency antenna array and a microphone array provided by an embodiment of the present invention;

[0058] Figure 4 is a specific structure schematic diagram of a node module provided by an embodiment of the present invention;

[0059] Figure 5 It is a schematic diagram of four basic feature templates of the Haar-like feature algorithm;

[0060] Figure 6 It is a schematic structural diagram of a multi-layer convolutional neural network model provided by an embodiment of the present invention;

[0061] Figure 7 It is a schematic layout diagram of a radio frequency antenna array provided by an embodiment of the present invention;

[0062] Figure 8 It is a schematic structural diagram of a CNN and GRU composite network model provided by an embodiment of the present invention;

[0063] Figure 9 It is a flowchart of a drone detection system based on multi-modal fusion provided by an embodiment of the present invention. Detailed implementation manners

[0064] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following provides a detailed description of a drone detection system and method based on multi-modal fusion proposed according to the present invention in combination with the accompanying drawings and specific implementation manners.

[0065] The foregoing and other technical contents, features, and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific implementation manners, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are only for reference and illustration, and are not used to limit the technical solutions of the present invention.

[0066] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element.

[0067] Embodiment 1

[0068] Please refer to Figure 1 and Figure 2 , Figure 1 It is a module diagram of a drone detection system based on multi-modal fusion provided by an embodiment of the present invention, Figure 2It is a schematic diagram of a node module provided by an embodiment of the present invention. The UAV detection system includes a system host and a plurality of node modules arranged in a detection area. The node module includes a radio frequency antenna array 1, a microphone sensor array 2, a radio frequency signal processing module 3, an audio signal processing module 4, a fusion module 5, and a communication module 6. Among them, the radio frequency antenna array 1 is used to obtain radio frequency signals of UAV targets in the detection area; the microphone sensor array 2 is used to obtain audio signals of UAV targets in the detection area; the radio frequency signal processing module 3 is used to perform radio frequency fingerprint extraction, classification recognition, and positioning on the radio frequency signals to obtain the classification recognition and positioning results of the radio frequency signals; the audio signal processing module 4 is used to perform voiceprint extraction, classification recognition, and positioning on the audio signals to obtain the classification recognition and positioning results of the audio signals; the fusion module 5 is used to perform multi-modal fusion on the classification recognition and positioning results of the radio frequency signals and the classification recognition and positioning results of the audio signals to obtain a fusion result including UAV type and position information; the communication module 6 is used to send the fusion result of the UAV type and position information to the system host. The system host is used to display the dynamic position of the UAV in real time.

[0069] Please refer to Figure 3 , Figure 3 It is a schematic diagram of the installation structure of a radio frequency antenna array and a microphone array provided by an embodiment of the present invention. In this embodiment, both the radio frequency antenna array 1 and the microphone sensor array 2 are arranged in a planar five-element cross array and are located on the same horizontal plane. In this embodiment, for the convenience of installation and maintenance, cost reduction, and volume reduction, the radio frequency antennas and microphone sensors are all installed on the same bracket.

[0070] Further, please refer to Figure 4 , Figure 4 It is a schematic diagram of the specific structure of a node module provided by an embodiment of the present invention. The radio frequency signal processing module 3 includes a radio frequency signal conversion unit 31, a radio frequency fingerprint extraction unit 32, a radio frequency classification recognition unit 33, and a radio frequency positioning unit 34. Among them, the radio frequency signal conversion unit 31 is used to convert the analog radio frequency signal received by the radio frequency antenna array into a digital radio frequency signal; the radio frequency fingerprint extraction unit 32 is used to obtain UAV radio frequency fingerprint information according to the digital radio frequency signal; the radio frequency classification recognition unit 33 is used to perform classification recognition on the UAV radio frequency fingerprint information based on a multi-layer convolutional neural network model to obtain the category information of the UAV target; the radio frequency positioning unit 34 is used to obtain the position information of the UAV target according to the digital radio frequency signal.

[0071] In this embodiment, the radio frequency fingerprint extraction unit 32 is specifically configured to draw the digital radio frequency signal in a two-dimensional vector diagram to obtain a constellation trace figure (CTF); select a certain interval to perform differential operation on the digital radio frequency signal to obtain a differential constellation trace figure (DCTF), and use the DCTF figure to characterize the UAV radio frequency fingerprint.

[0072] Specifically, the digital radio frequency signal is normalized; the normalized signal is subjected to IQ offset to form a discrete signal; the discrete signal is differentially processed to obtain a finite number of points on a two-dimensional plane; grid division is performed on the entire two-dimensional plane, and the grid sizes are uniformly equal, and each grid represents a pixel point; the points in each grid are counted, and the number of points is the gray value of this pixel, and then the gray value is converted into a pseudo-color, so as to obtain a color DCTF figure.

[0073] It should be noted that using the original image pixel values of the DCTF figure to describe the image feature recognition results is easily affected by background noise, while using Haar-like features to describe the image can basically solve the problem of the influence of background noise on the recognition results. Therefore, in this embodiment, Haar-like features are used to implement radio frequency fingerprint classification recognition based on DCTF, that is, a Haar-like feature vector diagram is extracted from the DCTF figure as the UAV radio frequency fingerprint information. The Haar-like feature algorithm defines four basic feature templates, as Figure 5 shown. In this embodiment, extracting a Haar-like feature vector diagram from the differential constellation trace figure specifically includes:

[0074] (a) Select an unused template from the four basic feature templates;

[0075] (b) Calculate the difference between the sum of pixel values in the white-covered area and the sum of pixel values in the black-covered area;

[0076] (c) Traverse the entire image with the current feature template in a one-pixel step. When one traversal ends, the window will be enlarged proportionally in width or length, and then repeat the previous traversal steps until the last ratio is enlarged and then end;

[0077] (d) Determine whether feature operations have been performed on all four templates. If not, repeat steps (a), (b), and (c).

[0078] Next, the radio frequency classification and recognition unit 33 classifies and recognizes the UAV radio frequency fingerprint information based on a multi-layer convolutional neural network model to obtain the category information of the UAV target. Please refer to Figure 6 ,Figure 6 It is a schematic structural diagram of a multi-layer convolutional neural network model provided by an embodiment of the present invention. The multi-layer convolutional neural network model includes 3 convolutional layers and 2 fully connected layers. Import the pre-collected and labeled picture samples of radio frequency fingerprints into the multi-layer convolutional neural network for training, adjust the parameters and train repeatedly to continuously improve the recognition rate. When the recognition rate meets the requirements, a trained neural network model is obtained. Input the Haar-like feature vector diagram extracted from the DCTF into the trained neural network model, and the category information of the UAV target can be obtained.

[0079] Further, the radio frequency positioning unit 34 is used to obtain the position information of the UAV target according to the digital radio frequency signal by using the Time Difference of Arrival (TDOA) positioning method to achieve accurate positioning of the UAV target. Specifically, please refer to Figure 7 , Figure 7 It is a schematic layout diagram of a radio frequency antenna array provided by an embodiment of the present invention. Assume that P1 is the target signal, and an O-XYZ coordinate system is established with the reference point O as the coordinate origin. P1 is the projection of the target P0 in the XOZ plane. S0 is the antenna deployed at the coordinate origin, S1 and S2 are the antennas deployed on the X-axis, S3 and S4 are the antennas deployed on the Y-axis. The distances between the antennas S1, S2, S3, S4 and the coordinate origin are all D. The five antennas S0, S1, S2, S3, S4 in the XOY plane form a five-element cross sensor array, and the coordinates are S0(0,0,0), S1(D,0,0), S2(-D,0,0), S3(0,D,0) and S4(0,-D,0) respectively. Let the rectangular coordinates of the target signal be P 0 (x 0 ,y 0 ,z 0 ), and the spherical coordinates are The propagation time of the target signal reaching the reference antenna S0 is t 0 , and the time delay values to the antennas S1, S2, S3, S4 relative to reaching the reference antenna S0 are τ 10 , τ 20 , τ 30 and τ 40 . The signal generated by the target P0 propagates in a spherical wave. From the geometric relationship and the distance formula between two points, we can get:

[0080]

[0081] The target signal P 0 The conversion relationship between the rectangular coordinates P 0 (x 0 ,y 0 ,z 0 ) and the spherical coordinates is:

[0082]

[0083] Substituting equation (2) into equation (1) gives:

[0084]

[0085] The target signal can be derived from equation (3) The coordinate calculation formula is:

[0086]

[0087] Thus, the position of the signal source point can be obtained according to the above calculation formula, that is, the position coordinates of the UAV can be obtained.

[0088] Continue to refer to Figure 4 , the audio signal processing module 4 of this embodiment includes an audio signal conversion unit 41, a voiceprint extraction unit 42, an audio classification and recognition unit 43, and an audio positioning unit 44. Among them, the audio signal conversion unit 41 is used to convert the analog audio signal received by the microphone into a digital audio signal; the voiceprint extraction unit 42 is used to obtain the UAV voiceprint information according to the digital audio signal; the audio classification and recognition unit 43 is used to classify and recognize the UAV voiceprint information based on a composite neural network model of CNN (Convolutional Neural Network) and GRU (Gated Recurrent Unit) to obtain the type information of the UAV target; the audio positioning unit 44 is used to obtain the position information of the UAV target according to the digital radio frequency signal.

[0089] Specifically, the voiceprint extraction unit 42 of this embodiment extracts voiceprints based on Mel Frequency Cepstral Coefficients (MFCC) feature parameters. MFCC is one of the most commonly used methods for speech feature extraction. The auditory sensitivity of the human ear changes with the change of the sound wave frequency. When the external sound is a low-frequency sound wave, its propagation radius is relatively large, while when the external sound is a high-frequency sound wave, its propagation radius is relatively small. It can be concluded from this that low-frequency sounds are easily masked by high-frequency sounds, and the critical bandwidth of masking at low frequencies is much smaller than that at high frequencies. Mel Frequency Cepstral Coefficients designs a set of band-pass filters with different densities based on the above principle analysis. The density is designed according to the size of the critical bandwidth. After the signal passes through the filter and then undergoes certain processing, the characteristics of the sound signal are formed. The obtained characteristics have no special requirements for the input signal. The voiceprint extraction algorithm based on MFCC feature parameters consists of five main steps:

[0090] (1) Preprocessing of the sound signal: mainly including pre-emphasis, framing and windowing, aiming to improve the quality of the sound signal;

[0091] (2) Fast Fourier Transform: The characteristics of the signal can be analyzed by two methods in the time domain and the frequency domain. However, the signal characteristics are not easily observable in the time domain, while the spectrum obtained through the Fast Fourier Transform can better represent the signal characteristics;

[0092] (3) Obtaining the energy spectrum: Perform a modulus square operation on the obtained spectrum to get the power spectrum;

[0093] (4) Triangular filtering: Take the energy spectrum obtained in the previous step as the input signal. Through a set of Mel-scale triangular filter banks, it can not only eliminate the harmonics, but also make the original voice formants more obvious;

[0094] (5) Calculate the MFCC coefficients, calculate the logarithmic energy output by each filter bank, and obtain the MFCC coefficients through the discrete cosine transform (DCT);

[0095] (6) Select MFCC coefficients of appropriate dimensions to construct a voiceprint feature vector expressing the audio characteristics of the drone.

[0096] Next, the audio classification and recognition unit 43 of this embodiment classifies and recognizes the drone voiceprint information based on the CNN and GRU composite neural network model to obtain the type information of the drone target. Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a CNN and GRU composite network model provided by an embodiment of the present invention. Specifically, the voiceprint vectors extracted from the audio signal are respectively input into the CNN network and the GRU network, and the one-hot codes regarding the drone type are respectively obtained. Perform Concat feature fusion on the outputs of CNN and GRU, and then output through the output layer to obtain the type information of the drone target.

[0097] Furthermore, the audio positioning unit 44 is used to obtain the position information of the drone target according to the digital audio signal. Since both the radio frequency antenna array 1 and the microphone sensor array 2 are arranged in a planar five-element cross array, therefore, the specific process of obtaining the position information of the drone target according to the digital audio signal is the same as the process of obtaining the position information of the drone target according to the digital radio frequency signal described above, and will not be elaborated here.

[0098] The radio frequency detection and acoustic detection modes are relatively independent of each other. The fusion module 5 in this embodiment uses a fractional layer fusion method to achieve the fusion of the two modes, which reduces the computational resource requirements for multi-modal fusion while improving the environmental adaptability and recognition accuracy of the UAV identification and positioning system. The fusion module 5 specifically includes a calculation unit and a weighted fusion unit. The calculation unit is used to calculate the similarity HD between the radio frequency fingerprint feature recognition result and the feature template respectively using the Hamming distance a and the similarity HD between the voiceprint feature recognition result and the feature template b . The feature template mentioned here refers to the existing UAV type template. The similarity between the radio frequency fingerprint feature recognition result and the voiceprint feature recognition result and each UAV type feature template can be calculated through the Hamming distance.

[0099] The weighted fusion unit is used to fuse the radio frequency data and audio data based on the weighted addition principle at the fractional layer

[0100] FS = w a ×(1 - HD a ) + w b ×(1 - HD b )

[0101] where FS represents the similarity with the feature template after fusion, w a , w b represent the weight coefficients of the radio frequency and audio modes respectively, and w a + w b = 1. w a , w b can be determined according to the actual situation or experience. According to the similarity with each feature template after fusion, the type of the UAV target can be finally determined.

[0102] It should be noted that the fusion module 5 in this embodiment can not only fuse the classification results obtained through radio frequency and audio to obtain the fused UAV category information, but also fuse the UAV position information obtained through radio frequency and audio to obtain the fused UAV position information. In this embodiment, by taking the average value of the UAV position coordinates obtained through radio frequency and audio, the position coordinates of the fused UAV can be obtained.

[0103] The UAV detection system based on multi-modal fusion in this embodiment solves the limitations of single-modal UAV detection and the implementation cost problem of multi-modal detection, realizes accurate, effective and reliable detection of UAVs, and can be widely promoted and applied in the market.

[0104] Embodiment 2

[0105] Based on the above embodiments, this embodiment provides a method for detecting unmanned aerial vehicles (UAVs) based on multimodal fusion. Please refer to Figure 9 , Figure 9 which is a flowchart of a UAV detection system provided by an embodiment of the present invention. The UAV detection method includes:

[0106] S1: Obtain the radio frequency (RF) signals of UAV targets in the detection area by using an RF antenna array.

[0107] S2: Obtain the audio signals of UAV targets in the detection area by using a microphone sensor array. Among them, both the RF antenna array and the microphone sensor array are arranged in a planar five-element cross array and are located on the same horizontal plane.

[0108] In this embodiment, for the convenience of installation and maintenance, cost reduction, and volume reduction, both the RF antenna and the microphone sensor are installed on the same bracket.

[0109] S3: Extract the RF fingerprint, classify and identify, and locate the RF signal to obtain the classification, identification, and location results of the RF signal.

[0110] The specific steps of S3 include:

[0111] S31: Convert the analog RF signal received by the RF antenna array into a digital RF signal.

[0112] S32: Obtain the UAV RF fingerprint information according to the digital RF signal.

[0113] Specifically, draw the digital RF signal in a two-dimensional vector diagram to obtain a CTF diagram; perform differential operations on the digital RF signal at a certain interval to obtain a DCTF diagram; extract the Haar-like feature vector diagram from the DCTF diagram as the UAV RF fingerprint information.

[0114] S33: Classify and identify the UAV RF fingerprint information based on a multi-layer convolutional neural network model to obtain the category information of the UAV target.

[0115] The multi-layer convolutional neural network model in this embodiment includes 3 convolutional layers and 2 fully connected layers. Import the pre-collected and labeled picture samples of RF fingerprints into this multi-layer convolutional neural network for training, adjust the parameters and train repeatedly to continuously improve the recognition rate. When the recognition rate meets the requirements, obtain the trained neural network model. Input the Haar-like feature vector diagram extracted from the DCTF into this trained neural network model to obtain the category information of the UAV target.

[0116] S34: Obtain the position information of the drone target based on the digital radio frequency signal. For the specific calculation process, please refer to Embodiment 1 and will not be elaborated here.

[0117] S4: Perform voiceprint extraction, classification recognition, and positioning on the audio signal to obtain the classification recognition and positioning results of the audio signal.

[0118] Further, the S4 includes:

[0119] S41: Convert the analog audio signal received by the microphone into a digital audio signal.

[0120] S42: Obtain the drone voiceprint information based on the digital audio signal.

[0121] In this embodiment, voiceprint extraction is performed based on the Mel-frequency cepstral coefficient feature parameters.

[0122] S43: Classify and recognize the drone voiceprint information based on the CNN and GRU composite neural network model to obtain the type information of the drone target.

[0123] Specifically, input the voiceprint vectors extracted from the audio signal into the CNN network and the GRU network respectively. What are obtained respectively are one-hot codes regarding the drone type. Perform Concat feature fusion on the outputs of CNN and GRU, and then output through the output layer to obtain the type information of the drone target.

[0124] S44: Obtain the position information of the drone target based on the digital radio frequency signal.

[0125] Since both the radio frequency antenna array and the microphone sensor array adopt a planar five-element cross array setting, therefore, the specific process of obtaining the position information of the drone target based on the digital audio signal is the same as the process of obtaining the position information of the drone target based on the digital radio frequency signal above and will not be elaborated here.

[0126] S5: Perform multi-modal fusion on the classification recognition and positioning results of the radio frequency signal and the classification recognition and positioning results of the audio signal to obtain a fusion result including the drone type and position information.

[0127] Specifically, step S5 includes:

[0128] S51: Use the Hamming distance to calculate the similarity HD between the radio frequency fingerprint feature recognition result and the feature template a and the similarity HD between the voiceprint feature recognition result and the feature template b ;

[0129] S52: Fuse the RF data and audio data based on the weighted addition principle at the score level:

[0130] FS = w a ×(1 - HD a ) + w b ×(1 - HD b )

[0131] where FS represents the similarity with the feature template after fusion, w a , w b represent the weight coefficients of the two modalities of RF and audio respectively, and w a + w b = 1.

[0132] It should be noted that this can not only fuse the classification results obtained through RF and audio to obtain the fused UAV category information, but also fuse the UAV position information obtained through RF and audio to obtain the fused UAV position information. In this embodiment, by taking the average value of the UAV position coordinates obtained through RF and audio, the position coordinates of the fused UAV can be obtained.

[0133] The UAV detection method based on multi-modal fusion in this embodiment is based on RF and acoustic detection methods. By establishing a UAV RF and audio fingerprint database, it realizes the classification and recognition of the two modalities of UAV RF detection and acoustic detection using a deep learning convolutional neural network, and realizes the fusion of the two modalities using a score-level decision method; it uses the Time Difference of Arrival (TDOA) positioning method to achieve accurate positioning of the UAV; the research results can be widely applied to application scenarios that require UAV control, such as airports, prisons, military management areas, important event venues, and nuclear industry facilities, and have broad market prospects and actual economic value.

[0134] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. A drone detection system based on multi-modal fusion, characterized in that, it includes a system host and multiple node modules arranged in the detection area. The node module includes a radio frequency antenna array (1), a microphone sensor array (2), a radio frequency signal processing module (3), an audio signal processing module (4), a fusion module (5) and a communication module (6). Among them, the radio frequency antenna array (1) is used to obtain the radio frequency signal of the drone target in the detection area; the microphone sensor array (2) is used to obtain the audio signal of the drone target in the detection area; the radio frequency signal processing module (3) is used to perform radio frequency fingerprint extraction, classification recognition and positioning on the radio frequency signal, and obtain the classification recognition and positioning results of the radio frequency signal; the audio signal processing module (4) is used to perform voiceprint extraction, classification recognition and positioning on the audio signal, and obtain the classification recognition and positioning results of the audio signal; the fusion module (5) is used to perform multi-modal fusion on the classification recognition and positioning results of the radio frequency signal and the classification recognition and positioning results of the audio signal, and obtain a fusion result including the drone type and location information; the communication module (6) is used to send the fusion result to the system host.

2. The drone detection system based on multi-modal fusion according to claim 1, characterized in that, both the radio frequency antenna array (1) and the microphone sensor array (2) are arranged in a planar five-element cross array and are located on the same horizontal plane.

3. The drone detection system based on multi-modal fusion according to claim 1, characterized in that, the radio frequency signal processing module (3) includes a radio frequency signal conversion unit (31), a radio frequency fingerprint extraction unit (32), a radio frequency classification recognition unit (33) and a radio frequency positioning unit (34). Among them, the radio frequency signal conversion unit (31) is used to convert the analog radio frequency signal received by the radio frequency antenna array into a digital radio frequency signal; the radio frequency fingerprint extraction unit (32) is used to obtain the drone radio frequency fingerprint information according to the digital radio frequency signal; the radio frequency classification recognition unit (33) is used to perform classification recognition on the drone radio frequency fingerprint information based on a multi-layer convolutional neural network model, and obtain the category information of the drone target; the radio frequency positioning unit (34) is used to obtain the location information of the drone target according to the digital radio frequency signal.

4. The drone detection system based on multi-modal fusion according to claim 3, characterized in that, the radio frequency fingerprint extraction unit (32) is specifically used for: drawing the digital radio frequency signal in a two-dimensional vector diagram to obtain a constellation trajectory diagram; performing differential operation on the digital radio frequency signal at a predetermined interval to obtain a differential constellation trajectory diagram; extracting a Haar-like feature vector diagram from the differential constellation trajectory diagram as the drone radio frequency fingerprint information.

5. The drone detection system based on multi-modal fusion according to claim 1, characterized in that, The audio signal processing module (4) includes an audio signal conversion unit (41), a voiceprint extraction unit (42), an audio classification and recognition unit (43), and an audio positioning unit (44), where the audio signal conversion unit (41) is configured to convert the analog audio signal received by the microphone into a digital audio signal; the voiceprint extraction unit (42) is configured to obtain the drone voiceprint information according to the digital audio signal; the audio classification and recognition unit (43) is configured to classify and recognize the drone voiceprint information based on a CNN and GRU composite neural network model to obtain the type information of the drone target; the audio positioning unit (44) is configured to obtain the position information of the drone target according to the digital audio signal.

6. The multi-modal fusion-based drone detection system according to claim 1, characterized in that the fusion module (5) specifically includes: A calculation unit for calculating the similarity HD between the radio frequency fingerprint feature recognition result and the feature template respectively using the Hamming distance a and the similarity HD between the voiceprint feature recognition result and the feature template b ; a weighted fusion unit, configured to fuse the radio frequency data and the audio data based on the weighted addition principle at the score layer: FS = w a ×(1 - HD a ) + w b ×(1 - HD b ) Among them, FS represents the similarity with the feature template after fusion, w a , w b respectively represent the weight coefficients of the radio frequency and audio modalities, and w a + w b = 1.

7. A multi-modal fusion-based drone detection method, characterized in that it includes: S1: Obtain the radio frequency signal of the drone target in the detection area by using a radio frequency antenna array; S2: Obtain the audio signal of the drone target in the detection area by using a microphone sensor array, where both the radio frequency antenna array and the microphone sensor array are arranged in a planar five-element cross array and are located on the same horizontal plane; S3: Perform radio frequency fingerprint extraction, classification and recognition, and positioning on the radio frequency signal to obtain the classification and recognition and positioning results of the radio frequency signal; S4: Perform voiceprint extraction, classification and recognition, and positioning on the audio signal to obtain the classification and recognition and positioning results of the audio signal; S5: Perform multi-modal fusion on the classification and recognition and positioning results of the radio frequency signal and the classification and recognition and positioning results of the audio signal to obtain a fusion result including the drone type and position information.

8. The multi-modal fusion-based drone detection method according to claim 7, characterized in that the S3 specifically includes: S31: Convert the analog radio frequency signal received by the radio frequency antenna array into a digital radio frequency signal; S32: Obtain the drone radio frequency fingerprint information according to the digital radio frequency signal; S33: Classify and recognize the drone radio frequency fingerprint information based on a multi-layer convolutional neural network model to obtain the category information of the drone target; S34: Obtain the position information of the drone target according to the digital radio frequency signal.

9. The multi-modal fusion-based drone detection method according to claim 7, characterized in that the S4 includes: S41: Convert the analog audio signal received by the microphone into a digital audio signal; S42: Obtain the drone voiceprint information according to the digital audio signal; S43: Classify and recognize the drone voiceprint information based on a CNN and GRU composite neural network model to obtain the type information of the drone target; S44: Obtain the position information of the drone target according to the digital audio signal.

10. The multi-modal fusion-based drone detection method according to claim 7, characterized in that, the S5 includes: S51: Calculate the similarity HD between the radio frequency fingerprint feature recognition result and the feature template respectively using the Hamming distance a and the similarity HD between the voiceprint feature recognition result and the feature template b ; S52: fusing radio frequency data and audio data based on the weighted addition principle at the fractional layer: FS = w a ×(1 - HD a ) + w b ×(1 - HD b ) Among them, FS represents the similarity with the feature template after fusion, w a , w b respectively represent the weight coefficients of the radio frequency and audio modalities, and w a + w b = 1.