Method and device for detecting small target in air
By collecting and processing multimodal perception data and using hypergraph models and neural network technology to identify and track small aerial targets, the accuracy and real-time problems of small aerial target detection are solved, and efficient aerial monitoring is achieved.
Patent Information
- Application Number
- CN202510471658.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies make it difficult to accurately identify and track small aerial targets, and cannot meet the real-time and high-efficiency requirements of aerial monitoring.
Multimodal perception data (visible light image data, infrared data, and radar data) is collected, pre-processed, and then subjected to feature-level fusion and fusion representation. Dynamic analysis is performed using hypergraph models and hypergraph neural networks. Spatiotemporal clustering algorithms and collaborative attention mechanisms are combined for collaborative identification and tracking to generate a unified representation of aerial targets. This is then verified and optimized through a real-time tracking algorithm.
It improves the ability to identify and locate small aerial targets in complex environments, enhances detection accuracy and real-time response capabilities, and meets the real-time and high-efficiency requirements of aerial monitoring.
Smart Images

Figure CN120597015A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of aerial monitoring technology, and in particular to a method and device for detecting small aerial targets. Background Art
[0002] Small target detection in the air is a key technology in the field of modern aviation monitoring and security protection, and is widely used in multiple scenarios such as border patrol, ocean surveillance, disaster emergency response, and airspace management.
[0003] With the rapid development of small aerial platforms such as drones and micro air vehicles, traditional detection methods face numerous challenges. These challenges include the small size, high speed, variable flight attitudes, and complex aerial environments of targets, making accurate identification and positioning of small targets particularly difficult. Furthermore, existing detection systems typically rely on a single type of sensor, such as optical cameras or radars, which are subject to fluctuations in detection performance due to external factors such as ambient lighting and weather conditions. While multimodal sensor data fusion technology has improved target detectability to a certain extent, significant technical bottlenecks remain in efficiently integrating multi-source heterogeneous data and achieving accurate identification and real-time tracking. Furthermore, while artificial intelligence technologies such as deep learning and graph neural networks have made significant progress in the field of target detection, their practical application in small aerial target detection systems still faces challenges such as high model complexity and high computational resource consumption, making it difficult to meet the requirements of real-time and high efficiency.
[0004] In summary, existing technologies are difficult to accurately identify and track small targets in the air, and cannot meet the real-time and high-efficiency requirements of aerial monitoring, which needs to be urgently addressed. Summary of the Invention
[0005] The present application provides a method and device for detecting small aerial targets to solve the problems that the existing technology is difficult to accurately identify and track small aerial targets and cannot meet the real-time and high-efficiency requirements of aerial monitoring.
[0006] The first aspect of the present application provides a method for detecting small aerial targets, comprising the following steps: collecting multimodal perception data corresponding to at least one aerial target, wherein the multimodal perception data includes visible light image data, infrared data, and radar data; preprocessing the multimodal perception data, and performing feature-level fusion and fusion representation operations on the preprocessed multimodal perception data to generate a unified aerial target representation for each aerial target in the at least one aerial target; and dynamically parsing the multimodal perception data based on the unified aerial target representation, a pre-built hypergraph model, and a hypergraph neural network structure to obtain aerial target detection data corresponding to each aerial target.
[0007] Optionally, in one embodiment of the present application, after obtaining the aerial target detection data corresponding to each aerial target, it also includes: generating multiple key attribute information corresponding to each aerial target through the aerial target detection data, and converting the multiple key attribute information into standard format key attribute information, and based on the preset data compression and incremental update strategy, transmitting the standard format key attribute information to the preset control center; constructing a multi-target collaborative detection model based on the preset spatiotemporal clustering algorithm, collaborative attention mechanism, joint probability data association strategy and sparse representation strategy; extracting the spatiotemporal features of each aerial target, and based on the spatiotemporal features, the aerial target detection data and the multi-target collaborative detection model, The collaborative identification operation is performed on the target to obtain a collaborative identification result; based on a preset real-time tracking algorithm and the collaborative identification result, each aerial target is dynamically tracked to obtain corresponding dynamic tracking information, and real-time feedback information of the target user is obtained through a preset user feedback interface, so as to use the real-time feedback information and the dynamic tracking information to verify and optimize the collaborative identification result to generate corresponding verification optimization information; according to the verification optimization information, it is determined whether there is a newly appeared new aerial target and / or an abnormal aerial target that meets the preset abnormal behavior requirements, wherein, when the new aerial target and / or the abnormal aerial target exist, the control center is controlled to perform automatic early warning operations on the new aerial target and / or the abnormal aerial target.
[0008] Optionally, in one embodiment of the present application, the collecting of multimodal perception data corresponding to at least one aerial target, wherein the multimodal perception data includes visible light image data, infrared data, and radar data, includes: deploying a non-visual sensor and a visual camera that meets preset resolution requirements on the target ground platform to construct a multimodal perception system based on the non-visual sensor and the visual camera; and using the multimodal perception system to synchronously collect the visible light image data, the infrared data, and the radar data corresponding to the at least one aerial target.
[0009] Optionally, in one embodiment of the present application, the multimodal perception data is preprocessed, and the preprocessed multimodal perception data is subjected to feature-level fusion and fusion representation operations to generate a unified aerial target representation of each aerial target in the at least one aerial target, including: storing the multimodal perception data based on a preset distributed storage strategy, and constructing a computing platform through a preset GPU accelerator card and parallel computing architecture; using the computing platform to preprocess the stored multimodal perception data to obtain standard image data; constructing a data fusion model based on a preset multi-layer convolutional neural network and fusion layer structure, attention mechanism and multimodal collaborative learning strategy, so as to perform feature-level fusion on the standard image data through the data fusion algorithm to obtain a corresponding unified feature vector; constructing a fusion representation model according to a preset deep neural network, regularization strategy and temporal information processing module, so as to integrate and optimize the feature representation of the unified feature vector using the fusion representation model to generate a unified aerial target representation.
[0010] Optionally, in one embodiment of the present application, the multimodal perception data is dynamically parsed based on the unified representation of the aerial targets, the pre-built hypergraph model and the hypergraph neural network structure to obtain the aerial target detection data corresponding to each aerial target, including: based on the unified representation of the aerial targets and the hypergraph model, the multimodal perception data is converted into a node feature vector, and a hypergraph corresponding to the node feature vector is established; and feature extraction and pattern recognition operations are performed on the hypergraph according to the hypergraph neural network structure to obtain corresponding aerial target detection data.
[0011] The second aspect of the present application provides an aerial small target detection device, including: an acquisition module for acquiring multimodal perception data corresponding to at least one aerial target, wherein the multimodal perception data includes visible light image data, infrared data and radar data; a fusion module for preprocessing the multimodal perception data and performing feature-level fusion and fusion representation operations on the preprocessed multimodal perception data to generate a unified aerial target representation for each aerial target in the at least one aerial target; a detection module for dynamically parsing the multimodal perception data based on the unified aerial target representation, a pre-built hypergraph model and a hypergraph neural network structure to obtain aerial target detection data corresponding to each aerial target.
[0012] Optionally, in one embodiment of the present application, it also includes: a transmission module for generating multiple key attribute information corresponding to each aerial target through the aerial target detection data after obtaining the aerial target detection data corresponding to each aerial target, and converting the multiple key attribute information into standard format key attribute information, and based on the preset data compression and incremental update strategy, transmitting the standard format key attribute information to the preset control center; a modeling module for constructing a multi-target collaborative detection model based on a preset spatiotemporal clustering algorithm, a collaborative attention mechanism, a joint probability data association strategy and a sparse representation strategy; a collaborative identification module for extracting the spatiotemporal features of each aerial target, and based on the spatiotemporal features, the aerial target detection data and the multi-target collaborative detection model, for at least one A collaborative identification operation is performed on aerial targets to obtain a collaborative identification result; a verification and optimization module is used to dynamically track each aerial target based on a preset real-time tracking algorithm and the collaborative identification result to obtain corresponding dynamic tracking information, and obtain real-time feedback information of the target user through a preset user feedback interface, so as to use the real-time feedback information and the dynamic tracking information to verify and optimize the collaborative identification result to generate corresponding verification optimization information; an early warning module is used to determine whether there are newly emerged aerial targets and / or aerial abnormal targets that meet the preset abnormal behavior requirements based on the verification and optimization information, wherein, when the new aerial targets and / or the aerial abnormal targets that meet the requirements exist, the control center is controlled to perform automatic early warning operations on the new aerial targets and / or the aerial abnormal targets.
[0013] Optionally, in one embodiment of the present application, the acquisition module includes: a deployment unit for deploying non-visual sensors and visual cameras that meet preset resolution requirements on the target ground platform to build a multimodal perception system based on the non-visual sensors and the visual cameras; an acquisition unit for using the multimodal perception system to synchronously collect the visible light image data, the infrared data and the radar data corresponding to the at least one aerial target.
[0014] Optionally, in one embodiment of the present application, the fusion module includes: a storage unit for storing the multimodal perception data based on a preset distributed storage strategy, and constructing a computing platform through a preset GPU accelerator card and parallel computing architecture; a preprocessing unit for using the computing platform to preprocess the stored multimodal perception data to obtain standard image data; a construction unit for constructing a data fusion model based on a preset multi-layer convolutional neural network and fusion layer structure, attention mechanism and multimodal collaborative learning strategy, so as to perform feature-level fusion on the standard image data through the data fusion algorithm to obtain a corresponding unified feature vector; a generation unit for constructing a fusion representation model according to a preset deep neural network, regularization strategy and temporal information processing module, so as to use the fusion representation model to integrate and optimize the feature representation of the unified feature vector to generate a unified representation of the aerial target.
[0015] Optionally, in one embodiment of the present application, the detection module includes: an establishment unit for converting the multimodal perception data into a node feature vector based on the unified representation of the aerial target and the hypergraph model, and establishing a hypergraph corresponding to the node feature vector; a pattern recognition unit for performing feature extraction and pattern recognition operations on the hypergraph according to the hypergraph neural network structure to obtain corresponding aerial target detection data.
[0016] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for detecting small aerial targets as described in the above embodiment.
[0017] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned method for detecting small aerial targets.
[0018] The fifth aspect of the present application provides a computer program product, including a computer program, which is executed to implement the above-mentioned method for detecting small aerial targets.
[0019] Therefore, the embodiments of the present application have the following beneficial effects:
[0020] The embodiments of the present application can collect multimodal perception data corresponding to at least one aerial target, wherein the multimodal perception data includes visible light image data, infrared data, and radar data; pre-process the multimodal perception data, and perform feature-level fusion and fusion representation operations on the pre-processed multimodal perception data to generate a unified aerial target representation for each aerial target in at least one aerial target; based on the unified aerial target representation, a pre-built hypergraph model, and a hypergraph neural network structure, the multimodal perception data is dynamically parsed to obtain aerial target detection data corresponding to each aerial target. The present application improves the system's ability to identify and locate small targets in complex environments, and enhances detection accuracy and real-time response capabilities by constructing a hypergraph model and a multi-target collaborative detection and tracking algorithm. Thus, it solves the problems that the existing technology is difficult to accurately identify and track small aerial targets and cannot meet the real-time and efficiency requirements of aerial monitoring.
[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1 This is a flow chart of a method for detecting small aerial targets according to an embodiment of the present application;
[0024] Figure 2 A schematic diagram of the logical architecture of a method for detecting small aerial targets provided in one embodiment of the present application;
[0025] Figure 3 This is an example diagram of a device for detecting small aerial targets according to an embodiment of the present application;
[0026] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0027] Among them, 10 is a small target detection device in the air; 100 is an acquisition module, 200 is a fusion module, 300 is a detection module; 401 is a memory, 402 is a processor, and 403 is a communication interface. DETAILED DESCRIPTION
[0028] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0029] The following describes a method and device for detecting small aerial targets according to an embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a method for detecting small aerial targets. In this method, multimodal sensing data corresponding to at least one aerial target is collected, wherein the multimodal sensing data includes visible light image data, infrared data, and radar data; the multimodal sensing data is preprocessed, and feature-level fusion and fusion representation operations are performed on the preprocessed multimodal sensing data to generate a unified aerial target representation for each aerial target in at least one aerial target; based on the unified aerial target representation, a pre-built hypergraph model, and a hypergraph neural network structure, the multimodal sensing data is dynamically parsed to obtain aerial target detection data corresponding to each aerial target. By constructing a hypergraph model and a multi-target collaborative detection and tracking algorithm, the present application improves the system's ability to identify and locate small aerial targets in complex environments, and enhances detection accuracy and real-time response capabilities. This solves the problems that existing technologies have difficulty in accurately identifying and tracking small aerial targets and cannot meet the real-time and efficiency requirements of aerial monitoring.
[0030] Specifically, Figure 1 This is a flowchart of a method for detecting small aerial targets provided in an embodiment of the present application.
[0031] like Figure 1 As shown, the method for detecting small aerial targets includes the following steps:
[0032] In step S101, multimodal perception data corresponding to at least one aerial target is collected, wherein the multimodal perception data includes visible light image data, infrared data, and radar data.
[0033] The embodiment of the present application can first be achieved by installing a high-resolution visual camera (i.e., ground sensor A) and non-visual sensors (such as infrared sensors (i.e., ground sensor B) and radar systems (i.e., ground sensor C)) on the ground platform, such as Figure 2 As shown, visual and non-visual sensors are integrated into ground platforms to form a collaborative multi-modal perception system to collect visible light image data, infrared and radar data of aerial targets.
[0034] Optionally, in one embodiment of the present application, multimodal perception data corresponding to at least one aerial target is collected, wherein the multimodal perception data includes visible light image data, infrared data, and radar data, including: deploying non-visual sensors and visual cameras that meet preset resolution requirements on the target ground platform to build a multimodal perception system based on the non-visual sensors and visual cameras; and using the multimodal perception system to synchronously collect visible light image data, infrared data, and radar data corresponding to at least one aerial target.
[0035] Specifically, the embodiments of the present application can first install a high-resolution visual camera on the ground platform to collect visible light image data of aerial targets. The high-resolution visual camera has a resolution of at least 4K and a high frame rate, and can provide clear images under different lighting conditions; the camera is installed at multiple key positions on the ground platform to ensure all-round coverage of aerial targets, and is equipped with a waterproof and dustproof cover to adapt to various severe weather environments; in addition, the camera realizes automatic tracking function through an intelligent pan-tilt head, and can adjust the shooting angle in real time to maintain continuous monitoring of moving targets.
[0036] Secondly, the embodiments of the present application can install non-visual sensors such as infrared sensors and radar systems on the ground platform to collect infrared and radar data of aerial targets; in actual implementation, the infrared sensor can detect the heat signal of the target at night or in low visibility conditions, and is equipped with a high-sensitivity detector to capture tiny temperature changes; the radar system uses phased array technology, has high-precision distance and speed measurement capabilities, and can penetrate obstacles such as haze to provide stable target detection data; in addition, the non-visual sensor is installed through an independent bracket to ensure its optimal working angle, and works synchronously with the visual camera to achieve coordinated collection of multi-source data.
[0037] Afterwards, the embodiment of the present application can integrate visual and non-visual sensors into the ground platform to form a collaborative multimodal perception system. In the specific implementation process, the visual camera, infrared sensor and radar system can be connected to the central control unit of the ground platform through a high-speed data bus to achieve real-time data transmission and synchronous processing; during the integration process, the embodiment of the present application can adopt standardized interfaces and modular design to ensure compatibility and scalability between the sensors; at the same time, the ground platform is equipped with a power management system and a backup power supply to ensure the continuous operation of the sensors during long-term monitoring tasks; in addition, the multimodal perception system also integrates a preliminary data processing module to filter and pre-process the collected raw data, thereby providing high-quality data input for subsequent in-depth analysis and target recognition.
[0038] In step S102, the multimodal perception data is preprocessed, and feature-level fusion and fusion representation operations are performed on the preprocessed multimodal perception data to generate a unified aerial target representation for each aerial target in at least one aerial target.
[0039] Furthermore, the embodiments of the present application can construct a multimodal perception data storage and computing platform to achieve multi-angle real-time monitoring of aerial data, and design multimodal aerial target representation fusion software (i.e., multimodal aerial target detection software), such as Figure 2 As shown, this is to achieve the preprocessing and fusion representation of multimodal perception data.
[0040] Optionally, in one embodiment of the present application, multimodal perception data is preprocessed, and feature-level fusion and fusion representation operations are performed on the preprocessed multimodal perception data to generate a unified aerial target representation for each aerial target in at least one aerial target, including: storing multimodal perception data based on a preset distributed storage strategy, and constructing a computing platform through a preset GPU accelerator card and parallel computing architecture; using the computing platform to preprocess the stored multimodal perception data to obtain standard image data; constructing a data fusion model based on a preset multi-layer convolutional neural network and fusion layer structure, attention mechanism and multimodal collaborative learning strategy, so as to perform feature-level fusion on the standard image data through a data fusion algorithm to obtain a corresponding unified feature vector; constructing a fusion representation model according to a preset deep neural network, regularization strategy and temporal information processing module, so as to integrate and optimize the feature representation of the unified feature vector using the fusion representation model to generate a unified aerial target representation.
[0041] It should be noted that the specific process of constructing a multimodal perception data storage and computing platform in the embodiment of the present application to perform multi-angle real-time monitoring of aerial data is as follows:
[0042] 1. In embodiments of the present application, a data acquisition module may be established to receive sensor data from a multimodal perception ground platform in real time. The data acquisition module may utilize a high-speed data interface (e.g., USB 3.0, Ethernet, or fiber optic connection) to connect to various sensors on the ground platform to ensure stable and low-latency data transmission. In actual implementation, the data acquisition module includes a built-in real-time data processing unit capable of synchronously receiving multi-source data streams from high-resolution visual cameras, infrared sensors, and radar systems.
[0043] It should be noted that in order to process the data formats and transmission rates of different sensors, the data acquisition module is equipped with a multi-channel signal processor and buffer memory to ensure the orderly reception and preliminary processing of data; in addition, the data acquisition module can also support the timestamp synchronization function, which can ensure the time consistency of multimodal perception data through the Network Time Protocol (NTP) or GPS clock signal, thereby providing an accurate time reference for subsequent data fusion and analysis.
[0044] 2. The embodiments of the present application may be configured with a data storage module and adopt distributed storage technology to ensure efficient storage and management of large-scale data. The data storage module may adopt a distributed file system (such as Hadoop Distributed File System HDFS, Ceph or Apache Cassandra) to achieve efficient storage and reliable management of massive data. In the embodiments of the present application, the system is configured with a multi-node storage architecture to provide data redundancy and fault recovery capabilities to ensure high availability and persistence of data.
[0045] To optimize access performance, the data storage module combines data sharding and load balancing technologies to improve read and write speeds and reduce single-point bottlenecks. At the same time, the data storage module integrates data compression and deduplication algorithms to reduce storage space usage. To support fast data retrieval and indexing, the data storage module introduces data metadata management and tag classification mechanisms to facilitate on-demand access and management of different types of sensor data. In addition, the data storage module is also equipped with secure encryption and access control mechanisms to ensure the security and privacy protection of sensitive data.
[0046] 3. Build a high-performance computing platform equipped with GPU accelerator cards and parallel computing architecture to support real-time processing and analysis of multi-angle data; As a feasible way, the high-performance computing platform in the embodiment of the present application can adopt a distributed computing architecture, equipped with multiple server nodes, each node is equipped with multiple GPU accelerator cards (such as NVIDIA Tesla or A100 series) to meet the needs of large-scale parallel computing; The computing platform runs an efficient parallel computing framework (such as CUDA, OpenCL or TensorFlow distributed) to support parallel processing of multimodal perception data and efficient training and reasoning of deep learning models; The internal network of the platform adopts high-speed interconnection technology (such as InfiniBand or high-speed Ethernet) to ensure data transmission bandwidth and low latency between nodes. In order to optimize resource utilization and computing efficiency, the computing platform integrates task scheduling and resource management systems (such as Kubernetes or Slurm) to achieve dynamic load balancing and task allocation.
[0047] Afterwards, the embodiment of the present application can design a multimodal aerial target representation fusion software system to preprocess and fuse the multimodal perception data. The specific steps are as follows:
[0048] Step 1. Develop an image preprocessing module to perform image denoising, enhancement, scale normalization and other processing steps. Specifically, the image preprocessing module can first use advanced denoising algorithms (such as convolutional neural network driven denoising Autoencoder or non-local mean denoising) to effectively reduce the noise interference collected by the sensor and improve the image quality; secondly, the embodiment of the present application can use image enhancement technology (such as adaptive histogram equalization, contrast-limited adaptive histogram equalization) to enhance the contrast and details of the image to ensure that the target features are more obvious. For aerial targets of different resolutions and scales, the image preprocessing module also introduces scale normalization processing, which unifies the image size through multi-scale pyramid or image pyramid methods to eliminate detection errors caused by changes in target distance and size; in addition, the image preprocessing module integrates an automatic parameter adjustment function, which can dynamically optimize processing parameters according to real-time ambient light and weather conditions to adapt to changing monitoring environments.
[0049] Step 2: Design a data fusion algorithm and use deep learning technology to fuse multimodal perception data from different sensors at the feature level. In an embodiment of the present application, the data fusion algorithm can adopt a multi-layer convolutional neural network and a fusion layer structure. First, the data features from the visual camera, infrared sensor, and radar system are extracted separately. By using a deep learning model with shared weights, such as ResNet or EfficientNet, the feature extraction of different modal data is ensured to be consistent and efficient. Secondly, in the feature-level fusion stage, the embodiment of the present application can use an attention mechanism to perform weighted integration of the features extracted by each sensor, highlighting key features and suppressing redundant information. In addition, the data fusion algorithm introduces a multimodal collaborative learning strategy to optimize the complementarity of different modal features through joint training, thereby improving the expressive power and robustness of the fusion representation.
[0050] Step 3, build a fusion representation model to generate a unified aerial target representation for subsequent recognition and tracking analysis. Specifically, the fusion representation model is based on the unified feature vector after the fusion of multimodal features, and uses a deep neural network (such as a fully connected layer or a graph neural network) to further integrate and optimize the feature representation to generate a high-dimensional, low-redundancy unified representation; in the embodiment of the present application, regularization technology is introduced into the design of the fusion representation model to prevent overfitting and ensure the generalization ability of the representation. In order to support multi-target recognition and tracking, the fusion representation model also integrates a temporal information processing module (such as a long short-term memory network LSTM or a temporal convolutional network) to capture the dynamic change characteristics of the target in continuous frames; in addition, the unified representation output by the model undergoes dimensionality reduction processing to further compress the data dimension and improve the computational efficiency of subsequent recognition and tracking algorithms; ultimately, the generated unified aerial target representation not only has rich feature information, but also has high efficiency and scalability, and can seamlessly connect to subsequent target recognition, classification and trajectory prediction modules, significantly improving the detection and analysis performance of the overall system.
[0051] Therefore, the embodiments of the present application integrate multiple visual and non-visual sensors and adopt multimodal perception data fusion technology based on hypergraphs to comprehensively capture the multi-dimensional feature information of small aerial targets, thereby effectively improving the accuracy of target recognition and classification; in addition, the embodiments of the present application configure high-performance computing platforms and optimized algorithms to achieve real-time collection, processing and analysis of massive multimodal perception data, ensuring rapid response in complex dynamic environments and realizing timely target detection and early warning.
[0052] In step S103, based on the unified representation of aerial targets, the pre-built hypergraph model and the hypergraph neural network structure, the multimodal perception data is dynamically parsed to obtain the aerial target detection data corresponding to each aerial target.
[0053] Furthermore, the embodiments of the present application can also deploy a small target recognition algorithm based on a hypergraph (i.e., a hypergraph target recognition model) to achieve dynamic analysis of aerial multimodal perception data and automatically extract aerial targets (i.e., aerial target detection data) and input them into the control center. Specifically, the embodiments of the present application can first construct a hypergraph model to represent multimodal perception data and its associated relationships as a hypergraph structure; secondly, by designing a deep learning algorithm based on a hypergraph, a hypergraph neural network can be used to perform feature extraction and pattern recognition on aerial multimodal perception data; thereafter, the embodiments of the present application can construct a target extraction module to automatically input the identified aerial small target information into the control center for subsequent processing and decision-making.
[0054] Optionally, in one embodiment of the present application, multimodal perception data is dynamically parsed based on a unified representation of aerial targets, a pre-built hypergraph model, and a hypergraph neural network structure to obtain aerial target detection data corresponding to each aerial target, including: based on a unified representation of aerial targets and a hypergraph model, the multimodal perception data is converted into a node feature vector, and a hypergraph corresponding to the node feature vector is established; feature extraction and pattern recognition operations are performed on the hypergraph according to the hypergraph neural network structure to obtain corresponding aerial target detection data.
[0055] It should be noted that the specific process of dynamic analysis of aerial multimodal perception data in the embodiment of the present application is as follows:
[0056] 1. Build a hypergraph model to represent multimodal perception data and its relationships as a hypergraph structure:
[0057] When constructing a hypergraph model, the embodiment of the present application can first convert multimodal perception data (such as visual images, infrared signals and radar echoes) from different sensors into node feature vectors. Each sensor type corresponds to a specific node category to reflect its unique characteristic attributes; each hyperedge in the hypergraph not only connects multiple nodes, but also can represent the high-order association relationship between these nodes; for example, the visual node, infrared node and radar node are connected at the same time to form a hyperedge to represent the multimodal information of the same target detected at the same time; the hypergraph By vertex set and the hyperedge set ε, where Contains the nodes corresponding to all sensor data, and ε represents the complex associations between these nodes.
[0058] 2. Design a hypergraph-based deep learning algorithm and use the hypergraph neural network to perform feature extraction and pattern recognition on aerial multimodal perception data:
[0059] In the specific implementation process, the hypergraph-based deep learning algorithm (i.e., hypergraph target recognition model) designed in the embodiment of the present application adopts the hypergraph neural network (HGNN) structure to perform in-depth feature extraction and pattern recognition on the constructed hypergraph; HGNN uses the hypergraph convolution operation to convert the node features h of each layer into (k) Updated to:
[0060]
[0061] Where H is the association matrix between nodes and hyperedges; D v and D e Denote the degree matrices of nodes and hyperedges respectively; W denotes the hyperedge weight matrix; Θ (k) is the trainable weight matrix of the kth layer; σ is the activation function.
[0062] The hypergraph target recognition model can learn high-order correlation features at different levels through multi-layer superposition. By introducing the attention mechanism, it further enhances the ability to focus on key features and improves the accuracy of pattern recognition. In addition, the hypergraph target recognition model of the embodiment of the present application improves the training stability and deep expansion capability of the model by adopting batch normalization and residual connection technology. Finally, the node features processed by the multi-layer HGNN are used in the classifier to achieve accurate recognition of small targets in the air.
[0063] Optionally, in one embodiment of the present application, after obtaining the aerial target detection data corresponding to each aerial target, it also includes: generating multiple key attribute information corresponding to each aerial target through the aerial target detection data, and converting the multiple key attribute information into standard format key attribute information, and based on the preset data compression and incremental update strategy, transmitting the standard format key attribute information to the preset control center; constructing a multi-target collaborative detection model based on the preset spatiotemporal clustering algorithm, collaborative attention mechanism, joint probability data association strategy and sparse representation strategy; extracting the spatiotemporal features of each aerial target, and based on the spatiotemporal features, the aerial target detection data and the multi-target collaborative detection model, The collaborative identification operation is performed on the target in the air to obtain a collaborative identification result; based on the preset real-time tracking algorithm and the collaborative identification result, each aerial target is dynamically tracked to obtain corresponding dynamic tracking information, and the real-time feedback information of the target user is obtained through a preset user feedback interface, so as to use the real-time feedback information and dynamic tracking information to verify and optimize the collaborative identification result to generate corresponding verification optimization information; according to the verification optimization information, it is determined whether there is a newly appeared new aerial target and / or an abnormal aerial target that meets the preset abnormal behavior requirements, wherein, when there is a new aerial target and / or an abnormal aerial target that meets the requirements, the control center performs an automatic early warning operation on the new aerial target and / or the abnormal aerial target.
[0064] After obtaining the aerial target detection data, the embodiment of the present application can establish a target extraction module to structure the information of small aerial targets identified by the hypergraph neural network and automatically transmit it to the control center. The specific implementation steps are as follows:
[0065] 1. Generate key attributes such as the target's location, speed, and direction based on the recognition results, and organize this information into a standardized data format (i.e., standard format key attribute information);
[0066] 2. Use message queues to achieve efficient data transmission with the control center to ensure the real-time and reliability of target information; in order to optimize the delay in data transmission, the target extraction module can use data compression and incremental update technology to only transmit the changed data; in addition, the target extraction module integrates data verification and error detection mechanisms to ensure the integrity and accuracy of data during transmission; after the control center receives the target information (i.e., key attribute information in a standard format), it can perform further monitoring, analysis and response through a visual interface or an automated decision-making system; if the data needs to be further processed, the embodiments of the present application can call the back-end database or big data analysis platform to support complex decision-making and emergency response actions.
[0067] After that, the embodiment of the present application also needs to design a multi-target collaborative detection and tracking algorithm (i.e., a multi-target collaborative detection model), and perform real-time verification of the aerial target detection results based on real-time user feedback, and issue early warnings for new targets. The specific process is as follows:
[0068] 1. Develop a multi-target collaborative detection algorithm (i.e., a multi-target collaborative detection model) to collaboratively identify multiple aerial targets by combining spatiotemporal information;
[0069] Specifically, the multi-target collaborative detection algorithm in the embodiment of the present application can first obtain the spatiotemporal characteristics of each target through a multimodal data fusion module. On this basis, the algorithm can use spatiotemporal clustering methods to preliminarily group and identify targets; in order to improve the accuracy of detection and prevent mutual occlusion between targets, the algorithm introduces a collaborative attention mechanism to dynamically adjust the feature weights of each target and enhance the recognition ability of key feature areas; in addition, the algorithm uses a joint probabilistic data association method to solve the multi-target data association problem, and achieves accurate collaborative recognition of multiple aerial targets by maximizing the likelihood estimation of the target trajectory. In order to handle target detection in high-density scenarios, the algorithm also integrates sparse representation technology to reduce computational complexity and improve real-time processing capabilities.
[0070] 2. Design a real-time tracking module and use Kalman filter or particle filter algorithm to dynamically track the detected target:
[0071] As a feasible method, the real-time tracking module constructed by the embodiment of the present application can adopt Kalman filtering or particle filtering algorithm to dynamically track each detected aerial target. Among them, Kalman filtering is applicable to target motion under linear and Gaussian noise assumptions, and can provide optimal recursive estimation; while particle filtering is applicable to nonlinear or non-Gaussian noise conditions, and can handle complex motion models and observation models more flexibly; the real-time tracking module can update the state estimation of the filter according to the detection results, and maintain the consistency and continuity of the target through data association methods. The real-time tracking module also integrates state prediction and uncertainty assessment mechanisms to ensure accurate prediction and robust tracking of target motion; in addition, in order to cope with dynamic environmental changes, the real-time tracking module supports adaptive filter parameter adjustment to improve tracking performance in different scenarios.
[0072] 3. Integrate user feedback interface to verify and optimize test results through real-time user feedback:
[0073] In an embodiment of the present application, an integrated user feedback interface module allows the user to provide verification information of the detection results in real time through a graphical interface or API, including confirming the existence of the target, deleting the falsely detected target, or adding the missed target, etc. The integrated user feedback interface module can adopt incremental learning and online optimization technology to dynamically adjust the parameters and model weights of the detection algorithm according to user feedback. For example, when a user confirms that a certain detection result is a false detection, an embodiment of the present application can record the false detection sample and update the weight of the deep learning model through the back propagation mechanism to reduce the probability of similar false detections. On the contrary, for the real target newly added by the user, the embodiment of the present application will add it to the training set as a positive sample to enhance the model's recognition ability for new targets.
[0074] 4. Implement an early warning mechanism to automatically warn of newly appeared or abnormally behaving aerial targets so as to enable timely response and processing:
[0075] It should be noted that the early warning mechanism of the embodiment of the present application can identify newly appeared or abnormally behaving aerial targets by performing real-time analysis of the detection and tracking results, and automatically trigger an early warning signal; the early warning mechanism is implemented by combining multi-target collaborative detection and tracking algorithms, utilizing real-time data analysis and intelligent decision support systems to ensure that newly appeared or abnormally behaving aerial targets can be warned in a timely and accurate manner, thereby improving the safety and responsiveness of the overall system.
[0076] Therefore, the embodiments of the present application can maintain stable detection performance under various complex environmental conditions (such as low light, bad weather, etc.) through multi-sensor collaboration and data fusion technology, and enhance resistance to target occlusion and interference; in addition, the embodiments of the present application can also handle the detection and tracking tasks of multiple targets at the same time, and perform dynamic verification and optimization based on real-time user feedback, thereby reducing missed detections and false detections and improving overall tracking accuracy.
[0077] According to the method for detecting small aerial targets proposed in the embodiment of the present application, multimodal sensing data corresponding to at least one aerial target is collected, wherein the multimodal sensing data includes visible light image data, infrared data, and radar data; the multimodal sensing data is preprocessed, and feature-level fusion and fusion representation operations are performed on the preprocessed multimodal sensing data to generate a unified aerial target representation for each aerial target in at least one aerial target; based on the unified aerial target representation, a pre-built hypergraph model, and a hypergraph neural network structure, the multimodal sensing data is dynamically parsed to obtain aerial target detection data corresponding to each aerial target. The present application constructs associations between multimodal data and between multiple aerial targets through a hypergraph, and uses an association reasoning enhancement algorithm to identify and track small targets, thereby effectively improving the accuracy, real-time nature, and overall system performance of small aerial target detection, and can be widely used in multiple scenarios such as border security monitoring, marine patrols, disaster emergency response, and airspace management.
[0078] Next, the small aerial target detection device proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0079] Figure 3 Schematic diagram of a small aerial target detection device according to an embodiment of the present application.
[0080] like Figure 3 As shown, the aerial small target detection device 10 includes: an acquisition module 100 , a fusion module 200 and a detection module 300 .
[0081] The acquisition module 100 is used to acquire multimodal perception data corresponding to at least one aerial target, wherein the multimodal perception data includes visible light image data, infrared data and radar data.
[0082] The fusion module 200 is used to preprocess the multimodal perception data and perform feature-level fusion and fusion representation operations on the preprocessed multimodal perception data to generate a unified aerial target representation for each aerial target in at least one aerial target.
[0083] The detection module 300 is used to dynamically analyze the multimodal perception data based on the unified representation of air targets, the pre-built hypergraph model and the hypergraph neural network structure to obtain the air target detection data corresponding to each air target.
[0084] Optionally, in one embodiment of the present application, the aerial small target detection device 10 of the embodiment of the present application further includes: a transmission module, a modeling module, a collaborative recognition module, a verification optimization module and an early warning module.
[0085] Among them, the transmission module is used to generate multiple key attribute information corresponding to each aerial target through the aerial target detection data after obtaining the aerial target detection data corresponding to each aerial target, and convert the multiple key attribute information into standard format key attribute information, and based on the preset data compression and incremental update strategy, transmit the standard format key attribute information to the preset control center.
[0086] The modeling module is used to build a multi-target collaborative detection model based on the preset spatiotemporal clustering algorithm, collaborative attention mechanism, joint probabilistic data association strategy and sparse representation strategy.
[0087] The collaborative recognition module is used to extract the spatiotemporal features of each aerial target and perform a collaborative recognition operation on at least one aerial target based on the spatiotemporal features, aerial target detection data and a multi-target collaborative detection model to obtain a collaborative recognition result.
[0088] The verification and optimization module is used to dynamically track each aerial target based on the preset real-time tracking algorithm and collaborative recognition results to obtain corresponding dynamic tracking information, and obtain real-time feedback information from the target user through a preset user feedback interface, so as to use the real-time feedback information and dynamic tracking information to verify and optimize the collaborative recognition results to generate corresponding verification optimization information.
[0089] The early warning module is used to determine whether there are new aerial targets and / or abnormal aerial targets that meet the preset abnormal behavior requirements based on the verification optimization information. When there are new aerial targets and / or abnormal aerial targets that meet the requirements, the control center performs automatic early warning operations on the new aerial targets and / or abnormal aerial targets.
[0090] Optionally, in one embodiment of the present application, the acquisition module 100 includes: a deployment unit and an acquisition unit.
[0091] Among them, the deployment unit is used to deploy non-visual sensors and visual cameras that meet preset resolution requirements on the target ground platform to build a multimodal perception system based on non-visual sensors and visual cameras.
[0092] The acquisition unit is used to synchronously collect visible light image data, infrared data and radar data corresponding to at least one aerial target using a multimodal perception system.
[0093] Optionally, in one embodiment of the present application, the fusion module 200 includes: a storage unit, a preprocessing unit, a construction unit and a generation unit.
[0094] Among them, the storage unit is used to store multimodal perception data based on a preset distributed storage strategy, and to build a computing platform through a preset GPU accelerator card and parallel computing architecture.
[0095] The preprocessing unit is used to preprocess the stored multimodal perception data using the computing platform to obtain standard image data.
[0096] The construction unit is used to build a data fusion model based on the preset multi-layer convolutional neural network and fusion layer structure, attention mechanism and multimodal collaborative learning strategy, so as to perform feature-level fusion of standard image data through the data fusion algorithm to obtain the corresponding unified feature vector.
[0097] The generation unit is used to build a fusion representation model based on a preset deep neural network, regularization strategy and temporal information processing module, so as to integrate and optimize the feature representation of the unified feature vector using the fusion representation model to generate a unified representation of the aerial target.
[0098] Optionally, in one embodiment of the present application, the detection module 300 includes: an establishment unit and a pattern recognition unit.
[0099] Among them, a unit is established to convert multimodal perception data into node feature vectors based on the unified representation of aerial targets and the hypergraph model, and to establish a hypergraph corresponding to the node feature vectors.
[0100] The pattern recognition unit is used to perform feature extraction and pattern recognition operations on the hypergraph according to the hypergraph neural network structure to obtain corresponding aerial target detection data.
[0101] It should be noted that the above explanation of the embodiment of the method for detecting small aerial targets is also applicable to the device for detecting small aerial targets in this embodiment, and will not be repeated here.
[0102] According to the embodiment of the present application, the aerial small target detection device includes an acquisition module 100 for acquiring multimodal sensing data corresponding to at least one aerial target, wherein the multimodal sensing data includes visible light image data, infrared data, and radar data; a fusion module 200 for preprocessing the multimodal sensing data and performing feature-level fusion and fusion representation operations on the preprocessed multimodal sensing data to generate a unified aerial target representation for each aerial target in at least one aerial target; a detection module 300 for dynamically parsing the multimodal sensing data based on the unified aerial target representation, a pre-built hypergraph model, and a hypergraph neural network structure to obtain aerial target detection data corresponding to each aerial target. The present application constructs associations between multimodal data and between multiple aerial targets through a hypergraph, and uses an association reasoning enhancement algorithm to identify and track small targets, thereby effectively improving the accuracy, real-time performance, and overall system performance of aerial small target detection, and can be widely used in multiple scenarios such as border security monitoring, marine patrols, disaster emergency response, and airspace management.
[0103] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0104] Memory 401 , processor 402 , and computer programs stored in the memory 401 and executable on the processor 402 .
[0105] When the processor 402 executes the program, the method for detecting small aerial targets provided in the above embodiment is implemented.
[0106] Furthermore, the electronic device further includes:
[0107] The communication interface 403 is used for communication between the memory 401 and the processor 402 .
[0108] The memory 401 is used to store computer programs that can be run on the processor 402 .
[0109] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0110] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0111] Optionally, in a specific implementation, if the memory 401 , the processor 402 and the communication interface 403 are integrated on a chip, the memory 401 , the processor 402 and the communication interface 403 can communicate with each other through an internal interface.
[0112] The processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0113] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for detecting small aerial targets.
[0114] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned method for detecting small aerial targets.
[0115] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0116] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0117] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0118] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0119] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0120] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0121] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0122] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for detecting small aerial targets, characterized in that: The following steps are involved: Collecting multimodal sensing data corresponding to at least one aerial target, wherein the multimodal sensing data includes visible light image data, infrared data, and radar data; Preprocessing the multimodal sensing data, and performing feature-level fusion and fusion representation operations on the preprocessed multimodal sensing data to generate a unified aerial target representation for each of the at least one aerial target; Based on the unified representation of the aerial targets, the pre-built hypergraph model and the hypergraph neural network structure, the multimodal perception data is dynamically analyzed to obtain the aerial target detection data corresponding to each aerial target.
2. The method according to claim 1, characterized in that After obtaining the aerial target detection data corresponding to each aerial target, the method further includes: Generate multiple key attribute information corresponding to each aerial target through the aerial target detection data, convert the multiple key attribute information into standard format key attribute information, and transmit the standard format key attribute information to a preset control center based on a preset data compression and incremental update strategy; Based on the preset spatiotemporal clustering algorithm, collaborative attention mechanism, joint probabilistic data association strategy and sparse representation strategy, a multi-target collaborative detection model is constructed; Extracting spatiotemporal features of each aerial target, and performing a collaborative recognition operation on the at least one aerial target based on the spatiotemporal features, the aerial target detection data, and the multi-target collaborative detection model to obtain a collaborative recognition result; Based on a preset real-time tracking algorithm and the collaborative recognition result, each aerial target is dynamically tracked to obtain corresponding dynamic tracking information, and real-time feedback information of the target user is obtained through a preset user feedback interface, and the collaborative recognition result is verified and optimized using the real-time feedback information and the dynamic tracking information to generate corresponding verification optimization information; Based on the verification optimization information, it is determined whether there is a newly emerged new aerial target and / or an abnormal aerial target that meets the preset abnormal behavior requirements. When the new aerial target and / or the abnormal aerial target exist, the control center is controlled to perform automatic early warning operations on the new aerial target and / or the abnormal aerial target.
3. The method according to claim 1, characterized in that The collecting of multimodal perception data corresponding to at least one aerial target, wherein the multimodal perception data includes visible light image data, infrared data, and radar data, includes: Deploying a non-visual sensor and a visual camera that meets preset resolution requirements on a target ground platform to construct a multimodal perception system based on the non-visual sensor and the visual camera; The multimodal perception system is used to synchronously collect the visible light image data, the infrared data, and the radar data corresponding to the at least one aerial target.
4. The method according to claim 3, characterized in that The preprocessing of the multimodal sensing data and performing feature-level fusion and fusion representation operations on the preprocessed multimodal sensing data to generate a unified aerial target representation for each of the at least one aerial target includes: Based on a preset distributed storage strategy, the multimodal perception data is stored, and a computing platform is constructed through a preset GPU accelerator card and parallel computing architecture; Preprocessing the stored multimodal perception data using the computing platform to obtain standard image data; Based on a preset multi-layer convolutional neural network and fusion layer structure, an attention mechanism, and a multimodal collaborative learning strategy, a data fusion model is constructed to perform feature-level fusion on the standard image data through the data fusion algorithm to obtain a corresponding unified feature vector; A fusion representation model is constructed according to a preset deep neural network, regularization strategy and temporal information processing module, so as to utilize the fusion representation model to integrate and optimize the feature representation of the unified feature vector to generate a unified representation of the aerial target.
5. The method according to claim 4, characterized in that The method of dynamically parsing the multimodal perception data based on the unified representation of the aerial targets, the pre-built hypergraph model, and the hypergraph neural network structure to obtain aerial target detection data corresponding to each aerial target includes: Based on the unified representation of aerial targets and the hypergraph model, the multimodal perception data is converted into node feature vectors, and a hypergraph corresponding to the node feature vectors is established; Feature extraction and pattern recognition operations are performed on the hypergraph according to the hypergraph neural network structure to obtain corresponding aerial target detection data.
6. A small target detection device in the air, characterized in that: include: an acquisition module, configured to acquire multimodal sensing data corresponding to at least one aerial target, wherein the multimodal sensing data includes visible light image data, infrared data, and radar data; a fusion module, configured to preprocess the multimodal sensing data and perform feature-level fusion and fusion representation operations on the preprocessed multimodal sensing data to generate a unified aerial target representation for each of the at least one aerial target; A detection module is used to dynamically analyze the multimodal perception data based on the unified representation of the aerial targets, a pre-built hypergraph model and a hypergraph neural network structure to obtain aerial target detection data corresponding to each aerial target.
7. The device according to claim 6, characterized in that Also includes: a transmission module for generating, after obtaining the aerial target detection data corresponding to each aerial target, a plurality of key attribute information corresponding to each aerial target through the aerial target detection data, converting the plurality of key attribute information into standard format key attribute information, and transmitting the standard format key attribute information to a preset control center based on a preset data compression and incremental update strategy; A modeling module is used to build a multi-target collaborative detection model based on a preset spatiotemporal clustering algorithm, collaborative attention mechanism, joint probabilistic data association strategy, and sparse representation strategy; a collaborative recognition module, configured to extract spatiotemporal features of each aerial target, and perform a collaborative recognition operation on the at least one aerial target based on the spatiotemporal features, the aerial target detection data, and the multi-target collaborative detection model to obtain a collaborative recognition result; a verification and optimization module, configured to dynamically track each aerial target based on a preset real-time tracking algorithm and the collaborative recognition result to obtain corresponding dynamic tracking information, and obtain real-time feedback information from the target user through a preset user feedback interface, and to perform verification and optimization on the collaborative recognition result using the real-time feedback information and the dynamic tracking information to generate corresponding verification and optimization information; An early warning module is used to determine whether there are newly emerged new aerial targets and / or abnormal aerial targets that meet preset abnormal behavior requirements based on the verification optimization information, wherein, when the new aerial targets and / or abnormal aerial targets that meet the requirements exist, the control center is controlled to perform automatic early warning operations on the new aerial targets and / or the abnormal aerial targets.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for detecting small aerial targets according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for detecting small aerial targets according to any one of claims 1 to 5.
10. A computer program product comprising a computer program, characterized in that The computer program is executed to implement the method for detecting small aerial targets according to any one of claims 1 to 5.