An intelligent UAV detection and remote monitoring system

A distributed RF and PTZ camera system synchronizes airwave and visual data for robust drone identification, addressing limitations in single-modal systems by fusing RF and visual features for improved accuracy and reliability.

CN120088467BActive Publication Date: 2025-07-15陕西秦泰保安服务有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510567498.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

It is difficult for existing drone detection technology to achieve high-precision and robust target recognition in complex scenarios, especially in the case of environmental interference and occlusion, and it is difficult for a single modal method to effectively explore the deep-level features of drones.

Method used

The distributed RF sensor and PTZ camera are combined, and space-time synchronous correlation analysis is carried out through drone signal feature matching and object detection technology, and deep learning algorithms are introduced for cross-modal interactive inference fusion to construct a joint description of multimodal information.

Benefits of technology

It improves the accuracy and robustness of drone detection and identification, and can effectively identify drone targets in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088467B_ABST
    Figure CN120088467B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of unmanned aerial vehicle (UAV) detection technology. Specifically, it discloses an intelligent UAV detection remote monitoring system, which uses distributed RF sensors and PTZ cameras to collect airborne radio signals and regional monitoring images. It screens suspicious signal sources from the radio frequency signals through a UAV signal feature matching algorithm, and at the same time uses object detection technology to extract the ROI regions of moving objects from the monitoring images. Through spatio-temporal synchronous correlation analysis of the two, a target correspondence relationship is constructed. Then, a deep learning algorithm is further introduced. By performing deep feature extraction and cross-modal interactive reasoning fusion on the associated target signal sources and the ROI images of moving objects, the deep-level association information between the two is mined, and a multi-modal information joint description of the target moving object is constructed, so as to realize the classification and recognition of UAV targets on this basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) detection, and more specifically, to an intelligent UAV detection remote monitoring system. Background Art

[0002] In recent years, with the rapid popularization of UAV technology, its applications in fields such as security monitoring, logistics transportation, and emergency rescue have become increasingly widespread. However, the unregulated flight of UAVs has also brought a series of safety hazards, and its abuse problem poses a serious threat to public safety, privacy protection, and critical infrastructure. Especially in sensitive scenarios such as airports, restricted areas, and large-scale events, the real-time detection and identification of "black flying" UAVs are in urgent need.

[0003] Currently, the UAV detection technology based on radio frequency (RF) signals realizes target perception by capturing the radio communication characteristics between the UAV and the remote controller, which has the advantages of strong penetration and all-weather operation, but has limitations such as being unable to directly associate physical entities and being vulnerable to electromagnetic interference; the vision monitoring technology based on cameras realizes target positioning through image recognition, which has the characteristics of intuitiveness and high resolution, but is significantly affected by environmental factors such as light and weather, and the detection accuracy for low, slow, and small targets needs to be improved.

[0004] In the prior art, due to the limited information dimension of single-modal methods, it is difficult to meet the robustness requirements of complex scenarios. Although some studies have attempted to comprehensively utilize multi-sensor data, due to the significant heterogeneity between RF signals and visual data, the existing methods generally adopt a serial processing mode of RF signal triggering and visual verification. Due to the lack of effective spatio-temporal association and complementary association fusion strategies, the ability to characterize the dynamic behavior of UAVs and their signal-image coupling relationship is insufficient, and it is difficult to mine highly discriminative features, often resulting in inaccurate target matching, especially when the target is occluded or the signal is interfered, the performance drops sharply.

[0005] Therefore, an optimized intelligent UAV detection remote monitoring system is expected. Summary of the Invention

[0006] In order to solve the above technical problems, the present invention is proposed. An embodiment of the present invention provides an intelligent drone detection remote monitoring system, which uses distributed RF sensors and PTZ cameras to collect aerial radio signals and regional monitoring images, and uses drone signal feature matching algorithms to screen suspicious signal sources from radio frequency signals. At the same time, target detection technology is used to extract the mobile object ROI area from the monitoring image, and the target correspondence relationship is constructed by performing spatiotemporal synchronous correlation analysis on the two. Then, a deep learning algorithm is further introduced to perform deep feature extraction and cross-modal interactive reasoning fusion on the successfully associated target signal source and the mobile object ROI image to mine the deep-level correlation information between the two, and construct a multi-modal information joint description of the target mobile object, thereby realizing the classification and recognition of drone targets on this basis. This method can effectively overcome the limitations of single-modality data and improve the accuracy and robustness of drone detection and identification.

[0007] Accordingly, according to one aspect of the present invention, there is provided an intelligent drone detection remote monitoring system, comprising:

[0008] A remote monitoring module for collecting airborne radio signals using distributed RF sensors and collecting regional monitoring images using PTZ cameras;

[0009] A suspicious signal source identification module, used to identify suspicious signal sources that meet the signal characteristics of drones from the aerial radio signals to obtain preliminary RF targets;

[0010] A moving object recognition module, used to recognize moving objects from the regional monitoring image to obtain preliminary photoelectric targets;

[0011] A cross-modal data association fusion module, used for performing cross-modal data interactive reasoning fusion based on spatiotemporal association analysis on the preliminary RF target and the preliminary optoelectronic target to obtain a cross-modal fusion target joint description;

[0012] The drone identification module is used to determine whether the target moving object is a drone based on the cross-modal fusion target joint description.

[0013] Compared with the prior art, the intelligent drone detection remote monitoring system provided by the present invention uses distributed RF sensors and PTZ cameras to collect air radio signals and regional monitoring images, screens suspicious signal sources from radio frequency signals through drone signal feature matching algorithms, and uses target detection technology to extract mobile object ROI areas from monitoring images, and constructs target correspondence relationships by performing spatiotemporal synchronous correlation analysis on the two. Then, a deep learning algorithm is further introduced to perform deep feature extraction and cross-modal interactive reasoning fusion on the successfully associated target signal source and mobile object ROI images to mine the deep-level correlation information between the two, and construct a multi-modal information joint description of the target mobile object, thereby realizing the classification and recognition of drone targets on this basis. This method can effectively overcome the limitations of single-modality data and improve the accuracy and robustness of drone detection and recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0015] Figure 1 4 is a block diagram of an intelligent UAV detection and remote monitoring system according to an embodiment of the present invention.

[0016] Figure 2 Schematic diagram of data flow of a smart drone detection and remote monitoring system according to an embodiment of the present invention.

[0017] Figure 3 4 is a block diagram of a cross-modal data association fusion module in an intelligent UAV detection remote monitoring system according to an embodiment of the present invention.

[0018] Figure 4 The present invention is a block diagram of a correlation data interaction fusion unit in an intelligent UAV detection remote monitoring system according to an embodiment of the present invention.

[0019] Figure 5 The block diagram is a block diagram of an interactive reasoning fusion subunit in an intelligent UAV detection remote monitoring system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] Below, the embodiments of the present invention will be described in more detail in conjunction with the accompanying drawings, and the above and other purposes, features and advantages of the present invention will become more apparent. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein.

[0021] As the technical problem described in the above background technology, the present invention proposes an intelligent drone detection remote monitoring system, which uses distributed RF sensors and PTZ cameras to collect aerial radio signals and regional monitoring images, and uses drone signal feature matching algorithms to screen suspicious signal sources from radio frequency signals. At the same time, target detection technology is used to extract the mobile object ROI area from the monitoring image, and the target correspondence relationship is constructed by performing spatiotemporal synchronous correlation analysis on the two. Then, a deep learning algorithm is further introduced to perform deep feature extraction and cross-modal interactive reasoning fusion on the successfully associated target signal source and mobile object ROI image to mine the deep-level correlation information between the two, and construct a multi-modal information joint description of the target mobile object, so as to realize the classification and recognition of drone targets on this basis. This method can effectively overcome the limitations of single-modality data and improve the accuracy and robustness of drone detection and recognition.

[0022] Figure 1 4 is a block diagram of an intelligent UAV detection and remote monitoring system according to an embodiment of the present invention. Figure 2 This is a data flow diagram of the intelligent drone detection remote monitoring system according to an embodiment of the present invention. Figure 1 and Figure 2 As shown, the intelligent UAV detection remote monitoring system 100 includes: a remote monitoring module 110, which is used to collect airborne radio signals using distributed RF sensors and collect regional monitoring images using PTZ cameras; a suspicious signal source identification module 120, which is used to identify suspicious signal sources that meet the characteristics of UAV signals from the airborne radio signals to obtain preliminary RF targets; a mobile object identification module 130, which is used to identify mobile objects from the regional monitoring images to obtain preliminary optoelectronic targets; a cross-modal data association and fusion module 140, which is used to perform cross-modal data interactive reasoning and fusion based on spatiotemporal association analysis on the preliminary RF target and the preliminary optoelectronic target to obtain a cross-modal fusion target joint description; a UAV identification module 150, which is used to determine whether the target moving object is a UAV based on the cross-modal fusion target joint description.

[0023] In the above intelligent UAV detection and remote monitoring system 100, the remote monitoring module 110 is used to collect aerial radio signals by using distributed RF sensors and collect area monitoring images by using PTZ cameras. It should be understood that the physical characteristics of a single sensor determine that it cannot fully cover the UAV monitoring requirements - although the radio frequency (RF) sensor can penetrate clouds and fog to sense radio signals, it cannot establish an intuitive mapping between the signal and the entity (such as it cannot distinguish whether the signal comes from a low-altitude UAV or a high-altitude relay device); although the PTZ camera (pan-tilt-zoom camera) can capture the visual characteristics of the target, limited by the optical imaging principle, the detection accuracy drops sharply in rainy and foggy weather (visibility < 1 km) or at night (without infrared fill light). Therefore, the present invention constructs a multi-modal perception network, which uses distributed RF sensors and PTZ cameras to synchronously collect aerial radio signals and area monitoring images, so as to form a two-way perception of "electromagnetic space - physical space" through the three-dimensional coverage in space and the complementary advantages of multi-modal data, and achieve all-round and high-precision detection of UAV targets.

[0024] In the specific implementation process, first, it is necessary to design the layout scheme of distributed RF sensors and PTZ cameras according to the characteristics of the monitoring area. Taking the airport clearance area as an example, it is necessary to cover an airspace with a radius of 10 km. Therefore, an RF sensor node is deployed every 2 - 3 km on the perimeter, and the nodes are distributed in a fan shape to eliminate detection blind spots. Each node selects a USRP N321 software-defined radio device, equipped with an omnidirectional dipole antenna array, which can cover the common frequency bands of 2.4 GHz and 5.8 GHz for UAVs, and realizes nanosecond-level clock synchronization through the GPSDO module to ensure the phase consistency of multi-node signals. At the same time, intelligent PTZ cameras are installed at the high points of the airport (such as the control tower and the roof of the terminal building), for example, the Dahua DH-SD-8A8640X-HNI model, which has 23x optical zoom and starlight-level night vision capabilities, and supports linkage with RF sensors through the RS485 interface. The hardware selection should take into account environmental adaptability. For example, the sensor nodes need to be equipped with a lightning protection and grounding system (grounding resistance < 4Ω), and the camera bracket needs to pass the wind and earthquake resistance test (able to withstand a 12-level typhoon) to ensure stable operation under complex climate conditions.

[0025] When deploying RF sensor nodes, first use an electromagnetic environment tester to scan candidate points, avoiding strong interference sources such as Wi-Fi hotspots and high-voltage lines to ensure that the background noise within the UAV frequency band is below -100 dBm. Use RTK-GPS to accurately measure the coordinates of each node (with centimeter-level accuracy), and fix the antenna array with a tripod to ensure that the antenna spacing is 0.5λ (for example, 6.25 cm for the 2.4 GHz frequency band) to support the phase interferometer direction-finding algorithm. When installing the PTZ camera, the optical center coordinates need to be calibrated, the lens distortion is calibrated through a checkerboard calibration board (such as the radial distortion coefficient k1 = -0.012), and the angle between the optical axis direction and the normal of the RF sensor is adjusted to ≤30° to reduce the perspective deviation during spatio-temporal synchronization.

[0026] Next, spatio-temporal synchronization is a key step. The central node obtains the UTC time through GPS and uses the PTP v2 protocol to send synchronization commands to all devices to ensure that the timestamp error between the RF sensor and the camera is < 500 ns. The spatial coordinates uniformly use the ENU coordinate system (with the central node as the origin). The RF sensor calculates the target three-dimensional coordinates through the TDOA and DOA algorithms, and the camera converts the image pixel coordinates into global coordinates through the pinhole imaging model.

[0027] When the RF sensor collects signals, the radio frequency signals received by the antenna are processed by a low-noise amplifier (with a gain of 20 dB) and a band-pass filter (stopband attenuation ≥60 dB), and then converted into I / Q baseband signals through quadrature down-conversion and quantized into 12-bit data at a sampling rate of 40 MS / s. The edge computing unit (such as NVIDIA Jetson Xavier NX) extracts 10-dimensional features such as power spectral density and instantaneous frequency in real time, compresses them and transmits them to the central node through the UDP protocol, and the original IQ data is stored in the local SSD (about 180 GB per hour). The PTZ camera outputs multiple bitstreams: the main bitstream is a 1080P@30fps H.265 video (bit rate 4 - 8 Mbps) for storage and analysis; the sub-bitstream is a 480P@15fps H.264 video for real-time preview. When receiving an RF trigger signal, the camera automatically turns to the target direction, demarcates a dynamic ROI area of 512×512 pixels, and improves the encoding quality (QP = 20) of this area to enhance target details.

[0028] In addition, timestamp compensation and spatial error correction need to be performed on the RF signals and surveillance images. The transmission delay of RF signals (about 2.5 ms) and the encoding / decoding delay of video streams (about 13 ms) need to be compensated using historical statistical values to ensure consistent time bases; the RF positioning error corrects the influence of multipath propagation through Taylor series expansion, and the depth estimation error of the camera is optimized by calculating the parallax of adjacent camera positions. The depth error of targets 500 meters away can be reduced from ±10% to ±3%.

[0029] In the conventional surveillance mode, the RF sensor scans at a frequency band switching period of every 10 seconds. When the detected signal energy exceeds the -85 dBm threshold, the feature matching process is started; the PTZ camera performs a 360° intelligent cruise at a period of 30 seconds to perform low-priority tracking of moving objects. When a certain RF node detects a suspected UAV signal (feature matching degree > 0.7), the system enters the trigger response mode: the central node immediately schedules 3 surrounding sensors to synchronously collect signals, improves the accuracy to within 50 meters through joint TDOA+DOA positioning, and sends the target coordinates to the corresponding azimuth camera, triggering it to turn at a speed of 60° / second and zoom in maximally. At this time, the RF sensor sampling rate is increased to 100 MS / s to capture preamble details, the camera enables the EDSR super-resolution algorithm to synthesize high-definition ROI images, and the raw data is transmitted back to the central node through the fiber optic link with low latency (<50 ms).

[0030] Secondly, parameters need to be dynamically adjusted in complex environments. For example, in rainy and foggy weather, the RF sensor enables 100 ms median filtering to suppress signal fluctuations, and the camera turns on the fog penetration mode to enhance contrast; when the nighttime illumination is insufficient, the camera switches to the infrared mode, and the RF detection threshold is increased by 5 dB to adapt to changes in the electromagnetic environment; in a multi-target scenario, the system sorts by signal strength priority, gives priority to processing targets with strong signals at close range, and weak signal targets enter a 2-second polling queue.

[0031] Based on the above methods, a three-dimensional perception system covering the electromagnetic space and the physical space is constructed through the joint application of distributed RF sensors and PTZ cameras.

[0032] In the above intelligent UAV detection and remote monitoring system 100, the suspicious signal source identification module 120 is used to identify a suspicious signal source that conforms to the UAV signal characteristics from the air radio signals to obtain a preliminary RF target. It should be understood that the originally collected RF signals usually contain a large number of non-UAV signals (such as Wi-Fi, Bluetooth devices, etc.). Direct processing will result in waste of computing resources. Considering that the communication between the UAV and the remote controller has specific frequency bands, modulation methods (such as frequency hopping, QPSK) and protocol characteristics (such as DJI OcuSync), therefore, in order to quickly screen out potential UAV signals from a large number of RF signals, the present invention adopts a UAV signal feature matching algorithm. By constructing a priori knowledge bases (such as signal bandwidths, frequency hopping patterns, preamble characteristics of UAV communication signals of manufacturers such as DJI and Parrot), the present invention calculates the similarity of protocol characteristics of the collected air radio signals to screen out suspicious signal sources that conform to the UAV signal characteristics, provides a preliminary RF target for subsequent UAV target identification, thereby narrowing the target range of data processing and improving the processing efficiency.

[0033] In the above intelligent UAV detection and remote monitoring system 100, the moving object identification module 130 is used to identify a moving object from the area monitoring image to obtain a preliminary optoelectronic target. It should be understood that the area monitoring image may contain various dynamic and static objects, including but not limited to clouds, birds, and UAVs. In order to accurately extract potential UAV targets from a complex background, the present invention uses advanced target detection technologies, such as the YOLO (You Only Look Once) algorithm based on deep learning, to perform real-time processing on the area monitoring image to quickly locate and extract the ROI (Region of Interest) area of the moving object in the image as a preliminary optoelectronic target. Specifically, the YOLO network is an end-to-end single-stage target detection model. By quickly analyzing the global features of the area monitoring image, it can simultaneously predict multiple bounding boxes and their corresponding class probabilities in a single forward pass, thereby realizing the efficient detection and identification of moving object targets in the image and providing important visual information references for subsequent UAV target identification.

[0034] In the above intelligent UAV detection and remote monitoring system 100, the cross-modal data association and fusion module 140 is used to perform cross-modal data interaction reasoning and fusion based on spatio-temporal association analysis on the preliminary RF target and the preliminary optoelectronic target to obtain a cross-modal fusion target joint description. Among them, Figure 3 It is a block diagram of the cross-modal data association and fusion module in the intelligent UAV detection and remote monitoring system according to an embodiment of the present invention. As Figure 3As shown in the figure, the cross-modal data association and fusion module 140 includes: a cross-modal data association unit 141, configured to perform association analysis on the preliminary RF target and the preliminary optoelectronic target based on a timestamp and a spatial position to obtain a successfully associated target signal source and a target moving object ROI image; and an associated data interaction and fusion unit 142, configured to perform cross-modal data association and interaction fusion on the successfully associated target signal source and the target moving object ROI image to obtain the cross-modal fusion target joint description.

[0035] Specifically, the cross-modal data association unit 141 is configured to perform association analysis on the preliminary RF target and the preliminary optoelectronic target based on a timestamp and a spatial position to obtain a successfully associated target signal source and a target moving object ROI image. It should be understood that considering the problems of time asynchrony and spatial misalignment in the acquisition process of RF signals and visual data, in order to avoid misassociating cross-modal data in different spatio-temporal scenarios, the present invention proposes an association analysis method based on a timestamp and a spatial position. Through spatio-temporal consistency verification - if a certain RF signal source and a visual target appear in the same spatial region (such as a horizontal distance < 300 meters and a height difference < 50 meters) within the same time window (such as ±2 seconds), then a corresponding relationship between the preliminary RF target and the preliminary optoelectronic target is constructed. Specifically, first, based on a unified timestamp, a time window (such as ±500 ms) is set to match the appearance times of the preliminary RF target and the preliminary optoelectronic target to exclude data pairs with too large time differences. Secondly, in combination with the spatial position information, the three-dimensional positioning result (by angle of arrival AOA or time difference of arrival TDOA) of the preliminary RF target signal source is mapped to the camera image plane, and the Euclidean distance between the projection of the preliminary RF target signal source on the image plane and the center of the ROI region of the preliminary optoelectronic target is calculated to exclude data pairs with spatial misalignment. Only when the timestamp matches and the spatial position is close (such as the Euclidean distance is less than a preset threshold), it is considered that the preliminary RF target and the preliminary optoelectronic target are successfully associated, forming a pair of a successfully associated target signal source and a target moving object ROI image to ensure that both correspond to the same physical entity.

[0036] Specifically, the associated data interaction and fusion unit 142 is configured to perform cross-modal data association and interaction fusion on the successfully associated target signal source and the target moving object ROI image to obtain the cross-modal fusion target joint description. Among them, Figure 4 is a block diagram of the associated data interaction and fusion unit in the intelligent unmanned aerial vehicle detection and remote monitoring system according to an embodiment of the present invention. As Figure 4As shown, the associated data interaction and fusion unit 142 includes: a signal feature extraction subunit 1421, configured to extract signal features from the target signal source to obtain a target signal source RF feature encoding vector; a moving object visual feature extraction subunit 1422, configured to extract visual features from the ROI image of the target moving object to obtain a target moving object visual feature encoding vector; an interaction reasoning and fusion subunit 1423, configured to perform cross-modal data association and fusion based on local interaction reasoning on the target signal source RF feature encoding vector and the target moving object visual feature encoding vector to obtain a cross-modal fusion target joint feature encoding vector as the cross-modal fusion target joint description.

[0037] More specifically, the signal feature extraction subunit 1421 is configured to extract signal features from the target signal source to obtain a target signal source RF feature encoding vector. In a specific example of the present invention, the signal feature extraction subunit 1421 is configured to: perform signal feature extraction on the target signal source based on a 1D CNN model to obtain the target signal source RF feature encoding vector. That is, considering that traditional signal processing methods such as Fourier transform are difficult to capture the deep features of non-stationary signals (such as frequency-hopping drones), for this reason, the present invention introduces a signal feature extraction method based on a one-dimensional convolutional neural network (1D CNN) to utilize its powerful feature learning ability to automatically extract deep features related to the communication behavior of the target drone from the original RF signal as an important input for subsequent drone target recognition. Specifically, first, the target signal source is converted from a radio frequency signal into time-domain waveform data. Then, a one-dimensional convolutional neural network model is constructed, and through the combination of multiple convolutional layers and pooling layers, the input time-domain waveform data is subjected to layer-by-layer feature extraction to automatically learn discriminative patterns in the signal waveform using sliding convolutional operations, capture key information such as signal intensity changes, pulse rising / falling edges, symbol periods, and periodicity of frequency point switching, thereby generating a high-level feature representation and obtaining the RF feature encoding vector of the target signal source, providing strong feature support for subsequent target recognition tasks.

[0038] More specifically, the moving object visual feature extraction subunit 1422 is used to extract visual features from the target moving object ROI image to obtain the target moving object visual feature encoding vector. In a specific example of the present invention, the moving object visual feature extraction subunit 1422 is used to: perform visual feature extraction based on the ViT model on the target moving object ROI image to obtain the target moving object visual feature encoding vector. It should be understood that since drones often appear in irregular shapes and variable sizes in images, and the background is complex and the lighting changes are diverse, traditional target detection methods based on manual features often find it difficult to achieve ideal recognition effects. In response to this problem, the present invention introduces a visual feature extraction method based on a vision transformer (ViT) model, aiming to automatically extract deep visual features related to the target drone from the target moving object ROI image through its powerful global feature capture capability. Specifically, the target moving object ROI image is first segmented into a series of small image blocks, and then each small image block is linearly embedded into a high-dimensional vector as an input sequence of the ViT model. Then, the self-attention mechanism in the Transformer architecture is used to model the relationship between each image block embedding vector and other image block embedding vectors, so as to capture the dependency between long-distance image blocks in the image and mine the contextual association between component structures, such as the structural association between propellers and fuselage, so as to generate a visual feature encoding vector of the target moving object. The visual feature encoding vector of the target moving object not only contains the local detail information of the target moving object, but also incorporates the global context information of the image based on the self-attention mechanism of the ViT model, which helps to improve the discriminability of feature representation and provide strong visual feature support for subsequent drone target recognition tasks.

[0039] More specifically, the interactive inference fusion subunit 1423 is configured to perform cross-modal data association and fusion based on local interactive inference on the target signal source RF feature encoding vector and the target moving object visual feature encoding vector to obtain a cross-modal fusion target joint feature encoding vector as the cross-modal fusion target joint description. It should be understood that since the radio frequency signal and the visual data come from different physical domains respectively, the information contained in the two has natural complementarity. However, it is often difficult to obtain an effective fusion result by directly splicing the features or simply performing weighted summation on the two. Therefore, in order to make full use of the complementarity of the two-modal data, the present invention proposes a cross-modal data association and fusion method based on local interactive inference, which performs feature deconstruction and local corresponding interactive encoding on the target signal source RF feature encoding vector and the target moving object visual feature encoding vector to achieve fine-grained feature interaction and fusion, and establish potential association points between cross-modal features. Then, a chain inference and an attention mechanism are further introduced to gradually derive a more refined feature interaction pattern along each local association point, and dynamically adjust the weights on the feature interaction path to ensure that key information is effectively strengthened while noise or redundant information is effectively suppressed. By iteratively updating the feature representation, the fusion degree between cross-modal features is continuously deepened to obtain a more discriminative cross-modal fusion target joint description.

[0040] Figure 5 It is a block diagram of the interactive inference fusion subunit in the intelligent UAV detection and remote monitoring system according to an embodiment of the present invention. As Figure 5 shown, the interactive inference fusion subunit 1423 includes: a segmented feature interaction secondary subunit 14231, configured to perform local segmented feature interaction encoding on the target signal source RF feature encoding vector and the target moving object visual feature encoding vector to obtain a set of target signal source - moving object local feature interaction encoding vectors; an attention weight calculation secondary subunit 14232, configured to determine the chain inference attention weights of the respective target signal source - moving object local feature interaction encoding vectors based on the feature distribution characteristics of the respective target signal source - moving object local feature interaction encoding vectors in the set of target signal source - moving object local feature interaction encoding vectors to obtain a set of target signal source - moving object local interactive chain inference attention weights; and a feature saliency modulation inference secondary subunit 14233, configured to perform feature modulation inference based on feature distribution saliency on the set of target signal source - moving object local feature interaction encoding vectors based on the set of target signal source - moving object local interactive chain inference attention weights to obtain the cross-modal fusion target joint feature encoding vector.

[0041] In a specific example of the present invention, the segmented feature interaction secondary subunit 14231 is configured to: respectively perform local implicit feature extraction based on one-dimensional convolutional coding on the target signal source RF feature coding vector and the target moving object visual feature coding vector to obtain a set of target signal source RF local implicit feature coding vectors and a set of target moving object visual local implicit feature coding vectors, which is expressed by the formula:

[0042] ;

[0043] ;

[0044] wherein, represents the target signal source RF feature coding vector, represents one-dimensional convolutional coding with a convolutional kernel scale of , , , and respectively represent the first, second, th, and th target signal source RF local implicit feature coding vectors in the set of target signal source RF local implicit feature coding vectors, is the number of the target signal source RF local implicit feature coding vectors, represents the target moving object visual feature coding vector, , , and respectively represent the first, second, th, and th target moving object visual local implicit feature coding vectors in the set of target moving object visual local implicit feature coding vectors.

[0045] That is, due to the significant heterogeneity between the RF signal features and the target moving object visual features in their original forms, direct global interaction is vulnerable to noise interference and difficult to capture fine-grained associations. Therefore, the present invention respectively extracts the local implicit features of both through one-dimensional convolution, aiming to deconstruct the global features into local pattern units with physical or semantic meanings, providing an alignment basis for subsequent in-depth feature correlation interaction.

[0046] In a specific example of the present invention, the segmented feature interaction secondary subunit 14231 is further configured to: perform single-entity feature interaction on each set of corresponding target signal source RF local implicit feature encoding vectors and target moving object visual local implicit feature encoding vectors in the set of target signal source RF local implicit feature encoding vectors and the set of target moving object visual local implicit feature encoding vectors to obtain the set of target signal source - moving object local feature interaction encoding vectors, which is represented by the formula:

[0047] ;

[0048] Among them, represents the concatenation function, represents the dot product operation, represents the point addition operation, represents the point subtraction operation, represents the bias term, represents the weight matrix, represents the th target signal source - moving object local feature interaction encoding vector in the set of target signal source - moving object local feature interaction encoding vectors, that is, and the target signal source - moving object local feature interaction encoding vector between them.

[0049] Here, in order to mine the potential association between RF signal events (such as remote control instructions) and the visual behavior of the target moving object (such as the motion state) through cross-modal interaction, the present invention designs a single-entity feature interaction module, which jointly models each set of corresponding target signal source RF local implicit feature encoding vectors and target moving object visual local implicit feature encoding vectors by using a deep neural network, and realizes the deep interaction of the two in the feature space through the non-linear transformation and feature mapping of the hidden layer, so as to effectively mine the potential feature association between the two and generate the target signal source - moving object local feature interaction encoding vectors.

[0050] In a specific example of the present invention, the attention weight calculation secondary subunit 14232 is represented by the formula:

[0051] ;

[0052] Among them, represents the normalized exponential function, represents the feature scale value of, represents the calculation of the norm, represents the th target signal source - moving object local feature interaction encoding vector The eigenvalue of a position, denote the target signal source - moving object local interaction chain - reasoning attention weight of

[0053] It should be understood that during the cross - modal feature interaction process, there are differences in the amount of information contained in different local interaction features and their importance to the final fusion result. Therefore, in order to accurately evaluate and dynamically adjust the contribution degrees of each target signal source - moving object local feature interaction coding vector in the subsequent chain - reasoning process, the present invention introduces an attention mechanism. By analyzing the characteristic distribution characteristics of each target signal source - moving object local feature interaction coding vector, such as the distribution range and dispersion degree of eigenvalues, etc., to quantify its contribution potential to the final fusion target description. In this way, it can be ensured that during the subsequent iterative process of feature interaction and fusion, key interaction information (such as the strong correlation information between signal intensity changes and the motion state of the target moving object) receives more attention and reinforcement, while relatively minor or redundant information backgrounds (such as the correlation between clouds and signal noise) are moderately weakened, thereby further improving the discriminability and accuracy of the cross - modal fusion target joint description.

[0054] In a preferred example of the present invention, the feature saliency modulation reasoning secondary subunit 14233 is used for: based on the feature interaction decomposition of each target signal source - moving object local feature interaction coding vector in the set of target signal source - moving object local feature interaction coding vectors, compensating and correcting the set of target signal source - moving object local interaction chain - reasoning attention weights to obtain a set of corrected target signal source - moving object local interaction chain - reasoning attention weights, which is expressed by the formula:

[0055] , , ;

[0056] ;

[0057] ;

[0058] ;

[0059] Among them, and respectively denote and the characteristic interaction translation effect vector statistic between denote and the characteristic interaction fluctuation effect vector statistic between denote the natural logarithm with base e, represents the periodic local rule compensation factor, represents the exponential function with the natural constant as the base, represents the periodic local auxiliary phase change factor, and respectively represent different weighting values, represents the corresponding corrected target signal source - moving object local interaction chain - reasoning attention weight.

[0060] In particular, considering that under the influence of electromagnetic interference, target occlusion, or environmental noise, some local interactions may generate misleading responses due to signal distortion or missing image features, and key interaction features may be submerged. To this end, in order to improve the accuracy of attention weight calculation, the present invention further decomposes the interaction between each corresponding target signal source RF local implicit feature encoding vector and the target moving object visual local implicit feature encoding vector, so as to distinguish the contribution differences of the translational effect (used to reflect the steady - state association of cross - modal features) and the fluctuation effect (used to characterize transient interference or periodic patterns) in the interaction specification space to the reasoning process, dynamically compensates and corrects the target signal source - moving object local interaction chain - reasoning attention weight, and enhances the robustness of the model to noise and the focusing ability on key features.

[0061] In a specific example of the present invention, the feature saliency modulation reasoning secondary sub - unit 14233 is further used for: based on the set of the corrected target signal source - moving object local interaction chain - reasoning attention weights, performing weighted modulation on the set of the target signal source - moving object local feature interaction encoding vectors to obtain a set of modulated target signal source - moving object local feature interaction encoding vectors; performing chain - reasoning based on the forward LSTM model on the set of the modulated target signal source - moving object local feature interaction encoding vectors to obtain the cross - modal fusion target joint feature encoding vector, which is expressed by the formula:

[0062] ;

[0063] ;

[0064] ;

[0065] wherein, 、 、 and respectively represent the first, the second, the th, and the th modulated target signal source - moving object local feature interaction encoding vectors in the set of the modulated target signal source - moving object local feature interaction encoding vectors, represents the forward LSTM model, represents the cross-modal fusion target joint feature encoding vector.

[0066] That is, by using the corrected target signal source - moving object local interaction chain inference attention weight set to scale each target signal source - moving object local feature interaction encoding vector, noise (such as the association between visual misdetection caused by cloud reflection and interference from irrelevant signals) is suppressed, the expression intensity of core correlation features is enhanced, so as to achieve adaptive filtering of the feature space and further highlight the interaction patterns between key features. Furthermore, on this basis, a forward LSTM model is used to perform chain inference on the modulated target signal source - moving object local feature interaction encoding vector set. During the chain inference process, the LSTM unit at each time step updates its internal state according to the output of the previous time step and the input of the current time step, thereby gradually accumulating and transmitting the interaction information between cross-modal features. In this way, the LSTM model can effectively integrate the feature interaction information at different time steps, capture and strengthen the long-term dependence relationship between cross-modal features, gradually construct a global joint description, and obtain the cross-modal fusion target joint feature encoding vector.

[0067] In the above intelligent UAV detection and remote monitoring system 100, the UAV recognition module 150 is used to judge whether the target moving object is a UAV based on the cross-modal fusion target joint description. In a specific example of the present invention, the UAV recognition module 150 is used to: input the cross-modal fusion target joint feature encoding vector into an AI classification model to judge whether the target moving object is a UAV. Specifically, the AI classification model has been trained a lot and has learned the comprehensive differences between UAVs and other moving objects in terms of RF signal features and visual features, and established an accurate classification decision boundary. In the actual application process, the AI classification model uses a deep neural network to perform feature analysis on the input cross-modal fusion target joint feature encoding vector, learns the signal-visual comprehensive feature pattern of the target moving object through multi-layer non-linear transformation and feature mapping, and classifies and judges the target moving object based on this. In this way, rapid and accurate recognition of the target moving object can be achieved, providing strong support for subsequent monitoring, tracking or interception actions.

[0068] In summary, the intelligent UAV detection and remote monitoring system according to the embodiments of the present invention is elucidated. It uses distributed RF sensors and PTZ cameras to collect aerial radio signals and regional monitoring images, screens suspicious signal sources from the RF signals through the UAV signal feature matching algorithm, and at the same time uses object detection technology to extract the ROI area of moving objects from the monitoring images. Then, through spatio-temporal synchronous correlation analysis of the two to construct the target correspondence relationship. Next, a deep learning algorithm is further introduced. By performing deep feature extraction and cross-modal interactive reasoning fusion on the associated target signal source and the ROI image of the moving object, the deep-level association information between the two is mined, and a multi-modal information joint description of the target moving object is constructed, so as to realize the classification and recognition of UAV targets on this basis. This method can effectively overcome the limitations of single-modal data and improve the accuracy and robustness of UAV detection and recognition.

[0069] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present invention are only examples and not limitations. It cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and easy understanding, rather than limitations. The above details do not limit the present invention to necessarily adopt the above specific details to implement.

Claims

1. An intelligent unmanned aerial vehicle detection and remote monitoring system, characterized in that, Including: A remote monitoring module, configured to collect aerial radio signals by using distributed RF sensors and collect regional monitoring images by using a PTZ camera; A suspicious signal source identification module, configured to identify a suspicious signal source conforming to the characteristics of a drone signal from the aerial radio signals to obtain a preliminary RF target; A moving object identification module, configured to identify a moving object from the regional monitoring images to obtain a preliminary optoelectronic target; A cross-modal data association and fusion module, configured to perform cross-modal data interaction reasoning and fusion based on spatio-temporal association analysis on the preliminary RF target and the preliminary optoelectronic target to obtain a joint description of the cross-modal fusion target; A drone identification module, configured to determine whether the target moving object is a drone based on the joint description of the cross-modal fusion target; The cross-modal data association and fusion module includes: A cross-modal data association unit, configured to perform association analysis on the preliminary RF target and the preliminary optoelectronic target based on a time stamp and a spatial position to obtain a target signal source with successful association and an ROI image of the target moving object; An associated data interaction and fusion unit, configured to perform cross-modal data association and interaction fusion on the target signal source with successful association and the ROI image of the target moving object to obtain the joint description of the cross-modal fusion target; The associated data interaction and fusion unit includes: A signal feature extraction subunit, configured to extract signal features from the target signal source to obtain a target signal source RF feature coding vector; A moving object visual feature extraction subunit, configured to extract visual features from the ROI image of the target moving object to obtain a target moving object visual feature coding vector; An interaction reasoning and fusion subunit, configured to perform cross-modal data association and fusion based on local interaction reasoning on the target signal source RF feature coding vector and the target moving object visual feature coding vector to obtain a cross-modal fusion target joint feature coding vector as the joint description of the cross-modal fusion target; The interaction reasoning and fusion subunit includes: A segmented feature interaction secondary subunit, configured to perform local segmented feature interaction coding on the target signal source RF feature coding vector and the target moving object visual feature coding vector to obtain a set of target signal source - moving object local feature interaction coding vectors; An attention weight calculation secondary subunit, configured to determine a chained reasoning attention weight of each target signal source - moving object local feature interaction coding vector based on the feature distribution characteristics of each target signal source - moving object local feature interaction coding vector in the set of target signal source - moving object local feature interaction coding vectors to obtain a set of target signal source - moving object local interaction chained reasoning attention weights; A feature significance modulation reasoning secondary subunit, configured to perform feature modulation reasoning based on feature distribution significance on the set of target signal source - moving object local feature interaction coding vectors based on the set of target signal source - moving object local interaction chained reasoning attention weights to obtain the cross-modal fusion target joint feature coding vector.

2. The intelligent UAV detection remote monitoring system according to claim 1, characterized in that, The signal feature extraction subunit is configured to: Extract the signal features of the target signal source based on the 1D CNN model to obtain the RF feature encoding vector of the target signal source.

3. The intelligent UAV detection remote monitoring system according to claim 2, wherein, The moving object visual feature extraction subunit is used for: Extract the visual features of the ROI image of the target moving object based on the ViT model to obtain the visual feature encoding vector of the target moving object.

4. The intelligent UAV detection remote monitoring system according to claim 3, characterized in that, The segmented feature interaction secondary subunit is used for: Respectively perform local implicit feature extraction based on one-dimensional convolutional encoding on the RF feature encoding vector of the target signal source and the visual feature encoding vector of the target moving object to obtain a set of RF local implicit feature encoding vectors of the target signal source and a set of visual local implicit feature encoding vectors of the target moving object; Perform single-body feature interaction on each group of corresponding RF local implicit feature encoding vectors of the target signal source and visual local implicit feature encoding vectors of the target moving object in the set of RF local implicit feature encoding vectors of the target signal source and the set of visual local implicit feature encoding vectors of the target moving object to obtain a set of local feature interaction encoding vectors of the target signal source - moving object.

5. The intelligent UAV detection remote monitoring system according to claim 4, characterized in that, The feature saliency modulation inference secondary subunit is used for: Based on the decomposition of the feature interaction of each local feature interaction encoding vector of the target signal source - moving object in the set of local feature interaction encoding vectors of the target signal source - moving object, compensate and correct the set of local interaction chain inference attention weights of the target signal source - moving object to obtain a corrected set of local interaction chain inference attention weights of the target signal source - moving object; Based on the corrected set of local interaction chain inference attention weights of the target signal source - moving object, perform weighted modulation on the set of local feature interaction encoding vectors of the target signal source - moving object to obtain a modulated set of local feature interaction encoding vectors of the target signal source - moving object; Perform chain inference based on the forward LSTM model on the modulated set of local feature interaction encoding vectors of the target signal source - moving object to obtain the cross-modal fusion target joint feature encoding vector.

6. The intelligent unmanned aerial vehicle detection remote monitoring system according to claim 5, wherein The UAV recognition module is used for: Input the cross-modal fusion target joint feature encoding vector into the AI classification model to determine whether the target moving object is a UAV.

Citation Information

Patent Citations

  • Multi-source fusion unmanned aerial vehicle intelligent detection system

    CN115508821A

  • Unmanned aerial vehicle comprehensive detection method and system based on multivariate intelligence fusion

    CN117329928A