Unmanned aerial vehicle type and flight state recognition method and device based on bimodal fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 36TH RES INST OF CETC
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-10
Smart Images

Figure CN122365339A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) identification technology, and in particular to a method and apparatus for identifying UAV model and flight status based on dual-modal fusion. Background Technology
[0002] Drones are widely used in video shooting, advertising production, urban surveying, and other fields. According to relevant low-altitude airspace control regulations in my country, it is necessary to monitor and manage drone flights, identify the type and flight status of drones appearing in the airspace in real time, and prevent "unauthorized flights" and potential drone attacks.
[0003] Currently, mainstream drone identification methods include: radio frequency (RF) signal-based identification, optical image-based identification, and radar signal-based identification. RF signal-based identification, by receiving and analyzing the RF signals emitted by the drone to uncover hidden model-related features, can achieve beyond-line-of-sight identification and is suitable for effectively identifying drones in environments with numerous obstacles such as tall buildings. However, this method requires collecting massive amounts of RF signals and relies on manual design for signal feature extraction. Optical image-based identification is significantly affected by lighting conditions and shooting distance, and it is difficult to operate stably in heavily obstructed urban environments. Radar signal-based identification is susceptible to multipath interference, has high equipment costs, and is not conducive to large-scale deployment. Furthermore, existing methods have limitations, such as only being able to identify the drone model and not its flight status. Summary of the Invention
[0004] Based on the above analysis, the embodiments of the present invention aim to provide a method and device for identifying UAV model and flight status based on dual-modal fusion, in order to solve the problems of existing technologies such as reliance on manual labor, limited identification capabilities, poor environmental adaptability, and high collection and deployment costs.
[0005] On one hand, embodiments of the present invention provide a method for identifying UAV model and flight status based on dual-modal fusion, including: Collect radio frequency signal data between the drone and the ground controller, and record the corresponding drone model and flight status; The corresponding time-frequency image is obtained by converting the radio frequency signal data; a drone training sample database is constructed based on the radio frequency signal data, the time-frequency image and its corresponding drone model and flight status; Based on the aforementioned UAV training sample database, a dual-modal fusion neural network model is constructed and trained; The system acquires real-time radio frequency signal data between the target UAV and the ground controller, and converts it into a corresponding real-time time-frequency image. Based on the trained dual-modal fusion neural network model and the real-time time-frequency image, the system identifies the model and flight status of the target UAV.
[0006] Furthermore, the radio frequency signal data is a complex baseband signal that simultaneously carries remote control signals and image transmission signals, obtained after quadrature demodulation. Its expression is: ; in, and (t) represents in-phase data and quadrature-phase data, respectively, where N is the number of sampling points, j is the imaginary unit, and t is the time. The signal samples collected from each drone are represented as follows: ,in, This represents a signal sample segment spanning N consecutive time points; These respectively represent the actual drone model label and the flight status label; The training samples in the UAV training sample database are represented in the following form: ; in, This represents the radio frequency signal data of drone i; Indicates according to The converted time-frequency image; This indicates the actual model label of the drone. The label represents the actual flight state of drone i; K represents the total number of training samples.
[0007] Furthermore, the dual-modal fusion neural network model includes at least: multiple layers of dual-modal processing and fusion units; the first... The layer dual-mode processing and fusion unit includes radio frequency signal mode paths. and image visual modal path The dual-path structure, and the radio frequency signal mode path and image visual modal path A cross-modal transformation module is constructed between them; Among them, the radio frequency signal mode path It consists of a single Transformer coding layer, which is used to receive and process data based on radio frequency signals. Obtained radio frequency signal mode data ; The image visual modal path It consists of a residual convolutional block, which is used to receive and process time-frequency based images. Obtained time-frequency image modal data ; The cross-modal transformation module includes at least: a fusion transformation unit for transforming radio frequency signal modal features into image visual modal features. And a fusion transformation unit for transforming image visual modal features into radio frequency signal modal features. .
[0008] Furthermore, the radio frequency signal mode path The specific processing steps include: The current radio frequency signal mode data Input to the radio frequency signal mode path The encoding process is performed to obtain updated radio frequency signal mode data. ; For the image visual modal path Output time-frequency image modal data Perform average pooling to obtain the pooled image feature vector; input the pooled image feature vector into the fusion transformation unit. After nonlinear transformation, the updated radio frequency signal mode data is generated. Image feature data after transformation with the same dimensions ; The transformed image feature data With the updated radio frequency signal mode data The signals are added together to obtain the fused radio frequency signal mode data. ; The fused radio frequency signal mode data The radio frequency signal mode path input to the next layer dual-mode processing and fusion unit RF signal mode data as input to the next layer .
[0009] Furthermore, the image visual modal path The specific processing steps include: Based on the fusion transformation unit For the current radio frequency signal mode data A nonlinear transformation is performed, and based on the radio frequency feature vector after the nonlinear transformation, the current time-frequency image modal data is generated. Radio frequency feature data after transformation with the same dimensions ; The transformed radio frequency feature data With current time-frequency image modal data The data are added together to obtain the fused time-frequency image modal data. ; The fused time-frequency image modal data Input to the image visual modality path Feature extraction is performed to obtain updated time-frequency image modal data. ; The updated time-frequency image modal data Image visual modality path input to the next layer dual-modal processing fusion unit Time-frequency image modal data as input to the next layer .
[0010] Furthermore, the dual-modal fusion neural network model further includes: a prediction output layer for performing UAV type prediction and flight status prediction; the prediction output layer includes at least: a UAV type prediction module. and UAV flight status prediction module .
[0011] Furthermore, the processing procedure of the prediction output layer specifically includes: The radio frequency signal modal data and time-frequency image modal data output from the last dual-modal processing fusion unit are fused to obtain the feature representation vector of UAV i. ; The feature representation vector The data is input into the UAV model prediction module respectively. and the UAV flight status prediction module The drone model was predicted. and flight status .
[0012] Furthermore, the method also includes: based on the actual model label of drone i Real flight status label The UAV model predicted by the dual-modal fusion neural network model. Flight status Calculate the loss function value; Calculate the gradient of the loss function value with respect to the parameters of the bimodal fusion neural network model, and iteratively update the model parameters based on the gradient until the preset training conditions are met, thus obtaining the trained bimodal fusion neural network model.
[0013] On the other hand, embodiments of the present invention provide a device for identifying UAV model and flight status based on dual-modal fusion, comprising: The acquisition module is used to acquire radio frequency signal data between the UAV and the ground controller, and record the corresponding UAV model and flight status; The database construction module is used to convert the radio frequency signal data to obtain the corresponding time-frequency image; and to construct a drone training sample database based on the radio frequency signal data, the time-frequency image and its corresponding drone model and flight status. The model building module is used to build and train a dual-modal fusion neural network model based on the UAV training sample database. The identification module is used to acquire real-time radio frequency signal data between the target UAV and the ground controller, and convert it into a corresponding real-time time-frequency image; based on the trained dual-modal fusion neural network model and the real-time time-frequency image, it identifies the model and flight status of the target UAV.
[0014] Furthermore, the dual-modal fusion neural network model includes at least: multiple layers of dual-modal processing and fusion units; the first... The layer dual-mode processing and fusion unit includes radio frequency signal mode paths. and image visual modal path The dual-path structure, and the radio frequency signal mode path and image visual modal path A cross-modal transformation module is constructed between them; Among them, the radio frequency signal mode path It consists of a single Transformer coding layer, which is used to receive and process data based on radio frequency signals. Obtained radio frequency signal mode data ; The image visual modal path It consists of a residual convolutional block, which is used to receive and process time-frequency based images. Obtained time-frequency image modal data ; The cross-modal transformation module includes at least: a fusion transformation unit for transforming radio frequency signal modal features into image visual modal features. And a fusion transformation unit for transforming image visual modal features into radio frequency signal modal features. .
[0015] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: First, unlike the limitations of recognition capabilities in related technologies, this invention adopts a dual-modal fusion method of radio frequency signals and radio frequency signal time-frequency images. Combining the data processing advantages of short-time Fourier transform and the trained dual-modal fusion neural network model, it can simultaneously achieve high-precision recognition of UAV model and flight status without manual intervention, effectively making up for the limitations of single-modal recognition and significantly improving recognition accuracy.
[0016] Secondly, unlike related technologies which suffer from poor environmental adaptability and high acquisition and deployment costs, this invention is based on the collaborative processing of radio frequency signal modal circuits and image vision modal circuits. This enables the model to capture the temporal characteristics of radio frequency signals and image vision features more comprehensively, and it can still work stably in complex environments, greatly enhancing its adaptability to diverse environments and unknown signals. At the same time, it does not require expensive dedicated equipment, effectively reducing acquisition and deployment costs and possessing greater practicality. It has broad market demand and application prospects in fields such as UAV monitoring.
[0017] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0018] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a flowchart of a method for identifying UAV model and flight status based on dual-modal fusion according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the implementation environment of an embodiment of the present invention; Figure 3 This is a schematic diagram of the radio frequency signal of a drone according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the time-frequency images corresponding to radio frequency signals of various models under different flight states in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the dual-modal fusion neural network model according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the main modules of the UAV model and flight status identification device based on dual-modal fusion according to an embodiment of the present invention. Detailed Implementation
[0019] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0020] A specific embodiment of the present invention discloses a method for identifying the type and flight status of a UAV based on dual-modal fusion, such as... Figure 1 As shown, the steps S1 to S4 are as follows: Step S1: Collect radio frequency signal data between the UAV and the ground controller, and record the corresponding UAV model and flight status.
[0021] Specifically, the embodiments of the present invention can be applied to various scenarios, and the implementation environment involved can include an input / output scenario of a single GPU server, or an interaction scenario between a portable computer and a GPU server. When the implementation environment is an input / output scenario of a single GPU server, the GPU server is the main body for processing and storing UAV radio frequency signal data, the main body for training neural networks, and the main body for inference neural networks; when the implementation environment is an interaction scenario between a portable computer and a GPU server, the schematic diagram of the implementation environment involved in this embodiment can be as follows. Figure 2 As shown. In Figure 2 The schematic diagram of the implementation environment shown includes a radio frequency signal receiver 101, a radio frequency signal receiving antenna 102, a portable computer 103, and a GPU server 104.
[0022] The radio frequency (RF) signal receiver 101 is a signal receiving device with an instantaneous bandwidth of at least 20 MHz. It receives RF signal data radiated by the UAV through the RF signal receiving antenna 102, including control command signals (remote control signals) between the UAV and the ground controller and image transmission signals emitted by the UAV. The RF signal receiver 101 and the RF signal receiving antenna 102 are connected by a fiber optic cable.
[0023] The radio frequency (RF) signal receiver 101 and the portable computer 103 are connected via a wired network. The RF signal receiver 101 transmits the received drone RF signals to the portable computer 103. The portable computer 103 performs a short-time Fourier transform on the received RF signals, converting them into a time-frequency image of the RF signals.
[0024] GPU server 104 is used to store and analyze the time-frequency images of radio frequency signals sent by portable computer 103, or GPU server 104 is used to send data to portable computer 103. Specifically, GPU server 104 can analyze and process the data sent by portable computer 103, train a convolutional neural network, and send the trained convolutional neural network model to portable computer 103.
[0025] In some embodiments, the radio frequency signal data in this invention includes remote control signals and image transmission signals; that is, the radio frequency signal data is a complex baseband signal that simultaneously carries remote control signals and image transmission signals, obtained after quadrature demodulation. Preferably, the radio frequency signal receiver 101 samples the signal at a center frequency Fc = 2.4 GHz and a sampling rate Fs = 50 MHz, and the collected radio frequency signal data can be expressed as: (1) in, and (t) represents in-phase data and quadrature-phase data, respectively, where N is the number of sampling points, j is the imaginary unit, and t is the time. The signal samples collected from each drone are represented as follows: ,in, This represents a signal sample segment spanning N consecutive time points; These respectively represent the actual drone model label and the flight status label; Figure 3 This is a schematic diagram of radio frequency signal data actually collected from a DJI Phantom 3 drone in a preferred embodiment.
[0026] Step S2: Obtain the corresponding time-frequency image based on the radio frequency signal data; construct a drone training sample database based on the radio frequency signal data, the time-frequency image and its corresponding drone model and flight status.
[0027] Specifically, since the UAV radio frequency signal is a burst signal, it is necessary to extract the effective signal from the received signal and perform a short-time Fourier transform to convert the radio frequency signal data into a time-frequency image. The specific steps include the following: (1) Calculate the amplitude of each sampling point in the signal. If the sampling point with the largest amplitude in the signal is lower than the threshold, If the signal contains no valid data, it is considered to be discarded. (2) Divide the effective radio frequency signal into non-overlapping uniform blocks according to the block size M to obtain N / M signal blocks; (3) Perform short-time Fourier transform on each signal block and stack them in the order of sampling time to obtain the radio frequency signal time-frequency image.
[0028] Preferably, the amplitude v(t) of the signal sampling point can be calculated using the following formula: (2) Each sample in the UAV training sample database contains a time-frequency image generated from radio frequency signal data, as well as a corresponding UAV model label and flight status label; the training samples in the UAV training sample database are represented in the following form: (3) in, This represents the radio frequency signal data of drone i. R represents the set of real numbers; Indicates according to The converted time-frequency image, ; This indicates the actual model label of the drone. The label represents the actual flight state of drone i; K represents the total number of training samples.
[0029] It is understood that the drone model in this embodiment refers to the specific product category or model identifier of various drones covered in the training sample database, such as DJI Phantom 3, DJI Phantom 4 Pro, DJI IMARTICE 30T, etc.; the drone flight state includes flight state, photography state, hovering state, etc.; the radio frequency signal time-frequency images radiated by different drone models in different flight states show different patterns, which can be referred to in detail. Figure 4 As shown.
[0030] Step S3: Based on the UAV training sample database, construct and train a dual-modal fusion neural network model.
[0031] refer to Figure 5 As shown, the dual-modal fusion neural network model includes at least: multiple layers of dual-modal processing fusion units, such as 4 layers of dual-modal processing fusion units; the third... l The layer dual-modal processing fusion unit is represented as , Represents a positive integer; the first The layer dual-mode processing and fusion unit includes radio frequency signal mode paths. and image visual modal path The dual-path structure, and the radio frequency signal mode path and image visual modal path A cross-modal transformation module is constructed between them; that is, each layer of dual-modal processing fusion unit includes a dual-path structure of radio frequency signal modal path and image vision modal path, as well as a corresponding cross-modal transformation module. Among them, the radio frequency signal mode path It consists of a single Transformer coding layer, which is used to receive and process data based on radio frequency signals. Obtained radio frequency signal mode data , ; The image visual modal path It consists of a residual convolutional block, which is used to receive and process time-frequency based images. Obtained time-frequency image modal data , The residual convolutional block is composed of a 3×3 convolutional layer, a batch normalization layer, a ReLU activation layer, a 3×3 convolutional layer, and a batch normalization layer stacked in sequence.
[0032] The cross-modal transformation module includes at least: a fusion transformation unit for transforming radio frequency signal modal features into image visual modal features. And a fusion transformation unit for transforming image visual modal features into radio frequency signal modal features. .
[0033] Preferably, the radio frequency signal mode path The specific processing steps include: First, the current radio frequency signal mode data Input to the radio frequency signal mode path The encoding process is performed to obtain updated radio frequency signal mode data. , The formula includes: = ( (4) Secondly, for the image visual modal path Output time-frequency image modal data Averaging pooling is performed to obtain the pooled image feature vector, for example, compressed into a 4096-dimensional vector; the pooled image feature vector is then input into the fusion transformation unit. After nonlinear transformation, the updated radio frequency signal mode data is generated. Image feature data after transformation with the same dimensions , The formula includes: (5) Here, Avgpool represents the average pooling operation.
[0034] Then, the transformed image feature data With the updated radio frequency signal mode data The signals are added together to obtain the fused radio frequency signal mode data. , The formula includes: + (6) Finally, the fused radio frequency signal mode data The radio frequency signal mode path input to the next layer dual-mode processing and fusion unit RF signal mode data as input to the next layer ;Right now .
[0035] Preferably, the image visual modal path The specific processing steps include: First, based on the fusion transformation unit For the current radio frequency signal mode data A nonlinear transformation is performed, and based on the nonlinearly transformed radio frequency feature vector, for example, a 50176-dimensional vector, this vector is then transformed into a 3×224×224 matrix through recombination and copying operations to generate the current time-frequency image modal data. Radio frequency feature data after transformation with the same dimensions , The formula includes: (7) in, Repeat is a recombination operation, while Repeat is a copy operation.
[0036] Secondly, the transformed radio frequency feature data With current time-frequency image modal data The data are added together to obtain the fused time-frequency image modal data. , The formula includes: + (8) Then, the fused time-frequency image modal data Input to the image visual modality path Feature extraction is performed to obtain updated time-frequency image modal data. , The formula includes: (9) Finally, the updated time-frequency image modal data Image visual modality path input to the next layer dual-modal processing fusion unit Time-frequency image modal data as input to the next layer ,Right now .
[0037] It is understandable that the input to the first-layer dual-modal processing fusion unit is... and radio frequency signal time-frequency image ,Right now , Radio frequency signals For example, when l =1 That is, the original radio frequency signal data in the training samples. ,when l >1 hour This refers to the modal characteristics of the radio frequency signal output from the previous layer. The modal characteristics of the time-frequency image are similar and will not be elaborated here.
[0038] Preferably, the dual-modal fusion neural network model further includes: a prediction output layer for performing UAV type prediction and flight status prediction; the prediction output layer includes at least: a UAV type prediction module. and UAV flight status prediction module Both prediction modules consist of a fully connected FC layer, a ReLU activation layer, a fully connected FC layer, and a Softmax layer.
[0039] Preferably, the processing procedure of the prediction output layer specifically includes: The radio frequency signal modal data and time-frequency image modal data output from the last dual-modal processing fusion unit are fused to obtain the feature representation vector of UAV i. ; The feature representation vector The data is input into the UAV model prediction module respectively. and the UAV flight status prediction module The drone model was predicted. and flight status .
[0040] For example, firstly, for a dual-modal fusion neural network model with four layers of dual-modal processing and fusion units, the radio frequency signal mode data output by the fourth layer of dual-modal processing and fusion units is... (The input for layer 4 is) l =4, output is l +1) and time-frequency image modal data By fusing the data, we obtain the feature representation vector of drone i. The formula includes: + (10) MLP stands for two fully connected layers.
[0041] After that, The data is input into the UAV model prediction module respectively. and UAV flight status prediction module Predicting drone models and flight status The formula includes: = ( ) (11) = ( ) (12) In some implementations, the method further includes: based on the actual model label of the drone i Real flight status label The UAV model predicted by the dual-modal fusion neural network model. Flight status Calculate the loss function value; Calculate the gradient of the loss function value with respect to the parameters of the bimodal fusion neural network model, and iteratively update the model parameters based on the gradient until the preset training conditions are met, thus obtaining the trained bimodal fusion neural network model.
[0042] In a preferred embodiment, the training method for the dual-modal fusion neural network model specifically includes the following steps: (1) Divide the UAV training sample database into a training set and a test set in a 7:3 ratio; (2) Optimizer parameter settings, including the initial learning rate Learning rate decay rate Learning rate decay period U, batch size B, maximum number of iterations E, etc.; (3) Randomly initialize the model parameters of the dual-modal fusion neural network model ; (4) Randomly select B signal samples of UAV radio frequency signals + time-frequency image pairs from the training set, and convert the UAV radio frequency signals into signals. and time-frequency images The input is fed into a dual-modal fusion neural network model, and through four layers of dual-modal fusion processing units, the UAV feature representation vector is obtained. Through the drone model prediction module and flight status prediction module and based on Predict the drone model separately and drone flight status ; (5) Based on the time-frequency image of the UAV radio frequency signal Corresponding tags Model predicted by the dual-modal fusion neural network model and drone flight status The loss function value L is calculated using the following formula: (13) (6) Calculate the gradient of the loss function value with respect to the parameters of the dual-modal fusion neural network model. To iteratively update model parameters The specific formulas include: (14) (7) Test the trained neural network model on the test set until the test error is less than the expected value, then stop training and obtain the trained bimodal fusion neural network model.
[0043] Therefore, the embodiments of the present invention adopt a dual-modal fusion method of radio frequency signal and radio frequency signal time-frequency image, and combine the data processing advantages of short time Fourier transform and the trained dual-modal fusion neural network model to effectively make up for the limitations of single-modal recognition and significantly improve the recognition accuracy of target UAV model and flight status.
[0044] Step S4: Acquire real-time radio frequency signal data between the target UAV and the ground controller, and convert it into the corresponding real-time time-frequency image; based on the trained dual-modal fusion neural network model and the real-time time-frequency image, identify the model and flight status of the target UAV.
[0045] Specifically, firstly, the acquired real-time radio frequency signal data is subjected to a short-time Fourier transform to convert it into a corresponding real-time time spectrum. Secondly, the real-time time spectrum is used as input data, and the real-time time spectrum is calculated based on the trained dual-modal fusion neural network model. The data is then processed through the collaborative processing of the radio frequency signal modal path and the image vision modal path of the multi-layer dual-modal processing fusion unit, as well as cross-modal information fusion. Finally, the prediction output layer generates the recognition result, which outputs the specific model of the target UAV and its current flight status.
[0046] Table 1 is a confusion matrix of model and flight status identification results for different drone models shown in a preferred embodiment of the present invention. DJIPhantom3, DJIPhantom4Pro, and DJIMATRICE30T represent three drone models, as detailed in the table below: Table 1: Confusion Matrix for Identifying UAV Models and Flight Status
[0047] As shown in Table 1, taking "DJI Phantom 3 in flight mode" as an example, among all sample data that are actually "DJI Phantom 3 in flight mode", 97.5% of the samples were correctly predicted by the model as "DJI Phantom 3 in flight mode", corresponding to the value of 0.975 in the first row and first column of the matrix; 1% of the samples were misjudged as "DJI Phantom 3 in photography mode", corresponding to the value of 0.01 in the second row and first column; another 1% of the samples were misjudged as "DJI Phantom 3 in hovering mode", corresponding to the value of 0.01 in the third row and first column, etc. It can be seen that the present invention does not show obvious biased misjudgment, can accurately identify the UAV model and its flight state, significantly improve the accuracy and reliability of UAV model and flight state identification, adapt to actual monitoring and control scenarios, and has strong practicality.
[0048] It is understood that the above embodiments are only for ease of understanding and simplification of description, and should not be construed as limiting the present invention. The present invention does not specifically limit the construction, training and application methods of the dual-modal fusion neural network model.
[0049] Therefore, it can be seen that the embodiments of the present invention can achieve at least one of the following beneficial effects: First, this invention adopts a dual-modal fusion method of radio frequency signal and radio frequency signal time-frequency image, combining the data processing advantages of short-time Fourier transform and the trained dual-modal fusion neural network model, which can simultaneously achieve high-precision identification of UAV model and flight status without manual intervention, effectively making up for the limitations of single-modal identification and significantly improving identification accuracy.
[0050] Secondly, this invention is based on the collaborative processing of radio frequency signal mode circuit and image vision mode circuit, which enables the model to capture the temporal characteristics of radio frequency signals and image vision features more comprehensively. It can still work stably in complex environments, greatly enhancing its adaptability to diverse environments and unknown signals. At the same time, it does not require expensive special equipment, effectively reducing the acquisition and deployment costs, and has stronger practicality. It has broad market demand and application prospects in fields such as UAV monitoring.
[0051] In another embodiment of the present invention, a device for identifying UAV model and flight status based on dual-modal fusion is proposed, such as... Figure 6 As shown, it specifically includes the following modules: The acquisition module is used to acquire radio frequency signal data between the UAV and the ground controller, and record the corresponding UAV model and flight status; The database construction module is used to convert the radio frequency signal data to obtain the corresponding time-frequency image; and to construct a drone training sample database based on the radio frequency signal data, the time-frequency image and its corresponding drone model and flight status. The model building module is used to build and train a dual-modal fusion neural network model based on the UAV training sample database. The identification module is used to acquire real-time radio frequency signal data between the target UAV and the ground controller, and convert it into a corresponding real-time time-frequency image; based on the trained dual-modal fusion neural network model and the real-time time-frequency image, it identifies the model and flight status of the target UAV.
[0052] Furthermore, the dual-modal fusion neural network model includes at least: multiple layers of dual-modal processing and fusion units; the first... The layer dual-mode processing and fusion unit includes radio frequency signal mode paths. and image visual modal path The dual-path structure, and the radio frequency signal mode path and image visual modal path A cross-modal transformation module is constructed between them; Among them, the radio frequency signal mode path It consists of a single Transformer coding layer, which is used to receive and process data based on radio frequency signals. Obtained radio frequency signal mode data ; The image visual modal path It consists of a residual convolutional block, which is used to receive and process time-frequency based images. Obtained time-frequency image modal data ; The cross-modal transformation module includes at least: a fusion transformation unit for transforming radio frequency signal modal features into image visual modal features. And a fusion transformation unit for transforming image visual modal features into radio frequency signal modal features. .
[0053] The above-described method and apparatus embodiments are based on the same principle, and their related aspects can be referenced from each other to achieve the same technical effect. For specific implementation processes, please refer to the foregoing embodiments, which will not be repeated here.
[0054] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0055] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying UAV model and flight status based on dual-modal fusion, characterized in that, include: Collect radio frequency signal data between the drone and the ground controller, and record the corresponding drone model and flight status; The corresponding time-frequency image is obtained by converting the radio frequency signal data; a drone training sample database is constructed based on the radio frequency signal data, the time-frequency image and its corresponding drone model and flight status; Based on the aforementioned UAV training sample database, a dual-modal fusion neural network model is constructed and trained; The system acquires real-time radio frequency signal data between the target UAV and the ground controller, and converts it into a corresponding real-time time-frequency image. Based on the trained dual-modal fusion neural network model and the real-time time-frequency image, the system identifies the model and flight status of the target UAV.
2. The method according to claim 1, characterized in that, The radio frequency signal data is a complex baseband signal that simultaneously carries remote control and image transmission signals, obtained after quadrature demodulation. Its expression is: ; in, and (t) represents in-phase data and quadrature-phase data, respectively, where N is the number of sampling points, j is the imaginary unit, and t is the time. The signal samples collected from each drone are represented as follows: ,in, This represents a signal sample segment spanning N consecutive time points; These respectively represent the actual drone model label and the flight status label; The training samples in the UAV training sample database are represented in the following form: ; in, This represents the radio frequency signal data of drone i; Indicates according to The converted time-frequency image; This indicates the actual model label of the drone. The label represents the actual flight state of drone i; K represents the total number of training samples.
3. The method according to claim 2, characterized in that, The dual-modal fusion neural network model includes at least: multiple layers of dual-modal processing and fusion units; the first The layer dual-mode processing and fusion unit includes radio frequency signal mode paths. and image visual modal path The dual-path structure, and the radio frequency signal mode path and image visual modal path A cross-modal transformation module is constructed between them; Among them, the radio frequency signal mode path It consists of a single Transformer coding layer, which is used to receive and process data based on radio frequency signals. Obtained radio frequency signal mode data ; The image visual modal path It consists of a residual convolutional block, which is used to receive and process time-frequency based images. Obtained time-frequency image modal data ; The cross-modal transformation module includes at least: a fusion transformation unit for transforming radio frequency signal modal features into image visual modal features. And a fusion transformation unit for transforming image visual modal features into radio frequency signal modal features. .
4. The method according to claim 3, characterized in that, The radio frequency signal mode path The specific processing steps include: The current radio frequency signal mode data Input to the radio frequency signal mode path The encoding process is performed to obtain updated radio frequency signal mode data. ; For the image visual modal path Output time-frequency image modal data Perform average pooling to obtain the pooled image feature vector; input the pooled image feature vector into the fusion transformation unit. After nonlinear transformation, the updated radio frequency signal mode data is generated. Image feature data after transformation with the same dimensions ; The transformed image feature data With the updated radio frequency signal mode data The signals are added together to obtain the fused radio frequency signal mode data. ; The fused radio frequency signal mode data The radio frequency signal mode path input to the next layer dual-mode processing and fusion unit RF signal mode data as input to the next layer .
5. The method according to claim 3, characterized in that, The image visual modal path The specific processing steps include: Based on the fusion transformation unit For the current radio frequency signal mode data A nonlinear transformation is performed, and based on the radio frequency feature vector after the nonlinear transformation, the current time-frequency image modal data is generated. Radio frequency feature data after transformation with the same dimensions ; The transformed radio frequency feature data With the current time-frequency image modal data The data are added together to obtain the fused time-frequency image modal data. ; The fused time-frequency image modal data Input to the image visual modality path Feature extraction is performed to obtain updated time-frequency image modal data. ; The updated time-frequency image modal data Image visual modality path input to the next layer dual-modal processing fusion unit Time-frequency image modal data as input to the next layer .
6. The method according to claim 3, characterized in that, The dual-modal fusion neural network model further includes a prediction output layer for performing UAV type prediction and flight status prediction; the prediction output layer includes at least a UAV type prediction module. and UAV flight status prediction module .
7. The method according to claim 6, characterized in that, The processing of the prediction output layer specifically includes: The radio frequency signal modal data and time-frequency image modal data output from the last dual-modal processing fusion unit are fused to obtain the feature representation vector of UAV i. ; The feature representation vector The data is input into the UAV model prediction module respectively. and the UAV flight status prediction module The drone model was predicted. and flight status .
8. The method according to claim 7, characterized in that, The method further includes: based on the actual model label of the drone i Real flight status label The UAV model predicted by the dual-modal fusion neural network model. Flight status Calculate the loss function value; Calculate the gradient of the loss function value with respect to the parameters of the bimodal fusion neural network model, and iteratively update the model parameters based on the gradient until the preset training conditions are met, thus obtaining the trained bimodal fusion neural network model.
9. A device for identifying UAV model and flight status based on dual-modal fusion, characterized in that, include: The acquisition module is used to acquire radio frequency signal data between the UAV and the ground controller, and record the corresponding UAV model and flight status; The database construction module is used to convert the radio frequency signal data to obtain the corresponding time-frequency image; and to construct a drone training sample database based on the radio frequency signal data, the time-frequency image and its corresponding drone model and flight status. The model building module is used to build and train a dual-modal fusion neural network model based on the UAV training sample database. The identification module is used to acquire real-time radio frequency signal data between the target UAV and the ground controller, and convert it into a corresponding real-time time-frequency image; based on the trained dual-modal fusion neural network model and the real-time time-frequency image, it identifies the model and flight status of the target UAV.
10. The apparatus according to claim 9, characterized in that, The dual-modal fusion neural network model includes at least: multiple layers of dual-modal processing and fusion units; the first The layer dual-mode processing and fusion unit includes radio frequency signal mode paths. and image visual modal path The dual-path structure, and the radio frequency signal mode path and image visual modal path A cross-modal transformation module is constructed between them; Among them, the radio frequency signal mode path It consists of a single Transformer coding layer, which is used to receive and process data based on radio frequency signals. Obtained radio frequency signal mode data ; The image visual modal path It consists of a residual convolutional block, which is used to receive and process time-frequency based images. Obtained time-frequency image modal data ; The cross-modal transformation module includes at least: a fusion transformation unit for transforming radio frequency signal modal features into image visual modal features. And a fusion transformation unit for transforming image visual modal features into radio frequency signal modal features. .