Low-altitude target recognition-oriented cross-modal solving method and system based on hybrid expert network

By fusing visual, audio, and radar features in a 3D voxel mesh using a hybrid expert network, the problem of spatial geometric feature loss caused by forced 2D dimensionality reduction alignment across modalities is solved, thus improving the accuracy and stability of low-altitude target identification.

CN122432835APending Publication Date: 2026-07-21UESTC (SHENZHEN) ADVANCED RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UESTC (SHENZHEN) ADVANCED RES INST
Filing Date
2026-06-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing low-altitude target perception systems, cross-modal forced 2D dimensionality reduction and alignment leads to the loss of spatial geometric features, reducing the accuracy of target recognition results.

Method used

A hybrid expert network is used for cross-modal computation. Feature fusion is performed on a 3D voxel grid using visual, audio and radar data. Radar-anchored voxels are used as spatial calibration benchmarks, and geometric consistency calibration is performed by combining visual and audio probabilistic voxels to generate a 3D fused voxel grid for input to the Transformer spatiotemporal correlation decoder.

Benefits of technology

It retains the target's altitude, azimuth, and distance information, improving the accuracy and stability of three-dimensional positioning, category identification, and cross-frame trajectory calculation for low-altitude targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432835A_ABST
    Figure CN122432835A_ABST
Patent Text Reader

Abstract

The application discloses a cross-modal solving method and system for low-altitude target identification based on a hybrid expert network, relates to the technical field of low-altitude target identification, and solves the technical problem that the existing technology causes loss of spatial geometric features and reduces the accuracy of target identification results due to forced 2D dimension reduction alignment in cross-modal solving. The method inputs image data, audio data and radar detection data into a solving model to obtain a solving result. The solving model comprises a visual expert network, an audio expert network, a radar expert network, a cross-modal adaptive calibration module, a Transformer space-time correlation decoder and a multi-task solving head. The cross-modal adaptive calibration module performs depth back-projection of visual features to a three-dimensional voxel grid to obtain visual probability voxels, projects audio latent feature probabilities to the three-dimensional voxel grid to obtain audio probability voxels, and fuses visual probability voxels and audio probability voxels that meet geometric consistency conditions in the neighborhood of each radar anchor voxel with the radar anchor voxel to generate a three-dimensional fusion voxel grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-altitude target recognition technology, and in particular to a cross-modal solution method and system for low-altitude target recognition based on a hybrid expert network. Background Technology

[0002] Existing low-altitude target perception systems mostly employ multimodal sensor fusion technology, including visible light and radar, to address highly concealed and maneuverable low-altitude, slow-moving, and small targets (such as micro-drones and model aircraft), thereby improving perception robustness. Specifically, multimodal features, such as visual, radar, and audio features, are extracted separately. Since different modal features not only differ in perception quality but also in spatial representation dimensionality and physical observation attributes, cross-modal spatial alignment and fusion of multimodal features is necessary for subsequent target recognition. However, in cross-modal spatial alignment and fusion, the original 3D radar point cloud is typically dimensionality-reduced and projected onto a 2D image or bird's-eye view plane and then stitched, weighted, or attention-based fused with other modal features. This process suffers from dimensionality collapse, which not only causes irreversible perspective distortion but also leads to the loss of the target's height geometry and 3D point cloud physical topology. When dealing with complex low-altitude environments with tall buildings and undulating terrain, this can easily result in cross-modal spatial misalignment, leading to significant 3D spatial bias in the final target recognition output.

[0003] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art: In existing low-altitude target perception technologies, cross-modal forced 2D dimensionality reduction alignment leads to the loss of spatial geometric features, which reduces the accuracy of target recognition results. Summary of the Invention

[0004] The purpose of this invention is to provide a cross-modal solution method and system based on a hybrid expert network for low-altitude target recognition, thereby solving the technical problem in existing technologies where forced 2D dimensionality reduction and alignment across modalities leads to the loss of spatial geometric features, which reduces the accuracy of target recognition results. The various technical effects of the preferred solutions among the many technical solutions provided by this invention are detailed below.

[0005] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a cross-modal solution method for low-altitude target recognition based on a hybrid expert network, comprising: acquiring and inputting synchronously sampled image data, audio data from a multi-channel microphone array, and radar detection data within the low-altitude defense zone at the current moment into a solution model to obtain the solution result at the current moment; the radar detection data includes radar target detection results; the solution model includes: a visual expert network for extracting visual features from the image data; an audio expert network for extracting latent audio features based on the audio data; and a radar expert network including a fast hierarchical module, which obtains fast hierarchical features of one or more radar targets based on the radar target detection results, and generates each radar target pair in a three-dimensional voxel grid. The system comprises: a radar anchor voxel; a cross-modal adaptive calibration module that back-projects the visual feature depth onto the 3D voxel grid to obtain visual probability voxels, projects the audio latent feature probabilities onto the 3D voxel grid to obtain audio probability voxels, and fuses the visual probability voxels and audio probability voxels that satisfy the geometric consistency condition in the neighborhood of each radar anchor voxel with the radar anchor voxel to generate the 3D fused voxel grid at the current moment; a Transformer spatiotemporal correlation decoder that processes the 3D fused voxel grid at the current moment and the trajectory query set at the previous moment through an attention mechanism to obtain the trajectory query set at the current moment; and a multi-task solution head that obtains the solution result at the current moment based on the trajectory query set at the current moment.

[0006] Preferably, the audio expert network further extracts a sound source direction vector based on the audio data and uses it as the sound source direction vector of the audio probability voxel; the geometric consistency condition is: and ;in, ; Indicates the first The center of the visual probability voxel satisfying the visual candidate condition within the neighborhood of each radar anchor voxel. With the The center of each radar anchor voxel Spatial residuals; Represents the spatial residual threshold; where, ; Indicates the first The sound source direction vector of the audio probability voxel satisfying the audio candidate condition within the neighborhood of each radar anchor voxel. With the Angular residuals between radar target directions Represents the direction vector of the sound source with vector The included angle; This indicates the position of the multi-channel microphone array in the global coordinate system; Represents the cosine function; Indicates the angular residual threshold; This indicates the calculation of the L2 norm.

[0007] Preferably, the visual expert network obtains visual confidence based on the visual features. The audio expert network also obtains the confidence level of the sound source direction estimation based on the sound source direction vector. And obtain rotor harmonic confidence based on the audio data. The radar target detection results include the radar detection confidence score for each radar target; in the cross-modal adaptive calibration module, fusing the visual probability voxels and audio probability voxels that satisfy the geometric consistency condition in the neighborhood of each radar anchor voxel with the radar anchor voxel to generate the current three-dimensional fused voxel mesh includes: generating the first... Local voxel calibration mask corresponding to the neighborhood of each radar anchor voxel : ; will the first The visual probability voxels and audio probability voxels that satisfy the geometric consistency condition within the neighborhood of the first radar anchoring voxel are related to the first... The radar anchoring voxel fusion was performed to obtain the first... One or more 3D fusion voxels corresponding to the neighborhood of each radar anchoring voxel; obtain the nth voxel according to the following formula. Fusion voxel features of three-dimensional fused voxels corresponding to the neighborhood of each radar anchoring voxel. : ;in, This represents the threshold determination function; Indicates the first Radar detection confidence level of a radar target. Indicates will Mapped to multimodal confidence; This represents the attention mechanism. For query, As key, Value; Indicates the first Features of a radar anchoring voxel Indicates the first Features of visual probability voxels that satisfy geometric consistency conditions within the neighborhood of a radar anchoring voxel. Indicates the first Features of audio probability voxels that satisfy geometric consistency conditions within the neighborhood of a radar anchoring voxel. This indicates splicing along the channel dimension. and The obtained splicing features are set to zero for the features of voxels in the three-dimensional voxel mesh that do not belong to the three-dimensional fused voxel mesh, thus obtaining the three-dimensional fused voxel mesh at the current moment.

[0008] Preferably, the radar detection data further includes radar point cloud data and echo data; the radar expert network further includes a full-scale hierarchical module, which includes: a point cloud topology branch, including a preprocessing unit and a point cloud feature extraction network, wherein the preprocessing unit performs ground point filtering, outlier removal, and voxelization processing on the radar point cloud data, and the point cloud feature extraction network is used to extract point cloud topology features from the radar point cloud data processed by the preprocessing unit; a micro-Doppler branch, including a short-time Fourier transform unit and a time-frequency feature extraction network, wherein the short-time Fourier transform unit performs short-time Fourier transform on the echo data to obtain a micro-Doppler time-frequency map, and the time-frequency feature extraction network is used to extract micro-Doppler features from the micro-Doppler time-frequency map; a full-scale fusion unit, which fuses the point cloud topology features and the micro-Doppler features to obtain radar fine discrimination features; the features of each radar anchor voxel include the fast hierarchical features and radar fine discrimination features of the radar target corresponding to the radar anchor voxel.

[0009] Preferably, the solution model further includes a computing power scheduling decision-maker, which determines the control action for the next moment based on the solution result at the current moment, the load state of the method execution entity, and the cross-modal inconsistency. : ;in, This indicates the start / stop flag of the fast-level module. This indicates the start / stop flag for the point cloud topology branch. This indicates the start / stop flag of the micro-Doppler branch. This indicates the voxel grid resolution of the three-dimensional voxel grid. This indicates the radar sampling frequency.

[0010] Preferably, the cross-modal inconsistency is determined jointly by the mean of the spatial residuals, the mean of the angular residuals, and the mean of the confidence residuals of all radar anchoring voxels; wherein, the mean of the confidence residuals is: Audio confidence ; Indicates the number of radar anchoring voxels, and is a positive integer; This indicates taking the absolute value.

[0011] Preferably, the image data includes visible light images and infrared thermal imaging images. The visual expert network includes: a visible light branch, comprising a visible light backbone network and a visible light pyramid network connected in sequence, used to extract multi-scale first enhanced image features from the visible light image; an infrared branch, comprising an infrared backbone network and an infrared pyramid network connected in sequence, used to extract multi-scale second enhanced image features from the infrared thermal imaging image; and a cross-modal feature bidirectional interaction network, including: a first convolution, which performs convolution processing on the second enhanced image features at a preset key scale; a first activation unit, which maps the features output by the first convolution to a thermal target spatial mask; a first multiplication unit, which performs element-wise multiplication of the first enhanced image features at the preset key scale with the thermal target spatial mask to obtain visible light target region features at the preset key scale; a first CSP fusion module, which fuses the visible light target region features at the preset key scale and all first enhanced image features at non-preset key scales to obtain visible light fused features; and a second convolution unit, which performs... The system comprises: a first enhanced image feature at a preset key scale undergoing convolution processing; a second activation unit mapping the features output by the second convolution to a texture contour mask; a second multiplication unit performing element-wise multiplication of the second enhanced image feature at the preset key scale with the texture contour mask to obtain infrared contour features at the preset key scale; a second CSP fusion module fusing the infrared contour features at the preset key scale and all second enhanced image features not at the preset key scale to obtain infrared fused features; a visible light weight acquisition unit comprising a first global average pooling layer and a first fully connected layer connected in sequence, used to map the visible light fused features to visible light weights; an infrared weight acquisition unit comprising a second global average pooling layer and a second fully connected layer connected in sequence, used to map the infrared fused features to infrared weights; a weighted fusion unit fusing the visible light fused features and the infrared fused features based on the visible light weights and the infrared weights to obtain the visual features; and a visual confidence prediction head acquiring visual confidence based on the visual features.

[0012] Preferably, the audio data includes Log-Mel spectral features, inter-channel phase difference, and time-of-arrival difference; the audio expert network includes: a source direction determination unit, which constructs a source direction vector based on the inter-channel phase difference and the time-of-arrival difference; a first audio confidence prediction head, which obtains the source direction estimation confidence based on the source direction vector; a second audio confidence prediction head, which obtains the rotor harmonic confidence based on the Log-Mel spectral features; a spatial branch, which processes the Log-Mel spectral features using a first asymmetric convolution expanding along the time axis to obtain spatial correlation features; a spectral branch, which processes the Log-Mel spectral features using a second asymmetric convolution expanding along the frequency axis to obtain spectral correlation features; and a temporal modeling unit, which obtains audio latent features based on the spatial correlation features and the spectral correlation features.

[0013] Preferably, the training process of the solution model includes: Initial training phase: An initial training model is constructed using the fast hierarchical modules in the visual expert network, audio expert network, and radar expert network within the solution model. This initial training model is trained in batches using the training set, and the network parameters of the initial training model are updated based on the geometric alignment loss until the initial stopping condition is met, thus entering the intermediate training phase; the geometric alignment loss is: ,in, Indicates visual weight, Indicates audio weights, Represents multimodal weights. Indicates the first Mean of spatial residuals of all radar anchor voxels in a sample Indicates the first Mean angle residuals of all radar anchor voxels in a sample Indicates the first The mean of the confidence residuals for each sample. The number of samples in each batch, The sample index is defined for each batch. During the mid-training phase: the solution model is trained in batches using the training set, with only the fast hierarchical module enabled in the radar expert network. The network parameters of the solution model are updated based on the second-stage loss until the mid-training stopping condition is met, entering the late-training phase. The second-stage loss includes geometric alignment loss, classification loss, and 3D localization loss. During the late-training phase: the solution model is trained in batches using the training set, with both the fast hierarchical module and the full hierarchical module enabled simultaneously in the radar expert network. The network parameters of the solution model are updated based on the third-stage loss until the late-training stopping condition is met. The third-stage loss includes geometric alignment loss, classification loss, 3D localization loss, and fine classification loss. The fine classification loss is calculated based on the fine classification result and the true target category in the sample's label. The fine classification result is obtained by processing the radar's fine discriminative features using an auxiliary fine classification head.

[0014] The present invention also provides a low-altitude target perception system, the system comprising: a camera group, a multi-channel microphone array, and a radar device located in the low-altitude defense zone; a data acquisition interface for synchronously sampling image data output by the camera group, audio data output by the multi-channel microphone array, and radar detection data output by the radar device; and a processor connected to the data acquisition interface and executing the cross-modal solution method for low-altitude target recognition based on a hybrid expert network provided by the present invention.

[0015] Implementing one of the above-described technical solutions of this invention has the following advantages or beneficial effects: Simultaneous acquisition of image data, audio data, and radar detection data within the low-altitude defense zone ensures that all three have the same time reference; a cross-modal adaptive calibration module is used for multimodal feature alignment and fusion; the radar expert network constructs radar anchoring voxels on a three-dimensional voxel grid based on radar target detection results; the radar anchoring voxels serve as the physical reference for three-dimensional spatial calibration; visual feature depth is back-projected onto the three-dimensional voxel grid to obtain visual probability voxels; and audio latent feature probabilities are projected onto the three-dimensional voxel grid to obtain audio probability voxels. By fusing visual and audio probability voxels that meet geometric consistency conditions within the neighborhood of the radar anchor voxel, a three-modal geometrically consistent 3D fused voxel is obtained, forming a 3D fused voxel mesh. This fusion method differs from 2D feature stitching or BEV (bird's-eye view) feature stitching. The 3D fused voxel mesh can retain the target's height, azimuth, range, and local 3D neighborhood relationships without losing spatial geometric features. It can be directly used as input to the Transformer spatiotemporal correlation decoder, thereby improving the accuracy and stability of low-altitude target 3D localization, category recognition, and cross-frame trajectory calculation. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a schematic diagram of the internal structure of the solution model in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the audio expert network structure in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the structure of the visual expert network in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the radar expert network structure in Embodiment 3 of the present invention; Figure 5 This is a system block diagram of the low-altitude target sensing system in Embodiment 5 of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments, illustrating various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of the present invention disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of the present invention.

[0018] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0019] To illustrate the technical solution described in this invention, specific embodiments are described below, showing only the parts related to the embodiments of this invention.

[0020] Example 1: This embodiment provides a cross-modal solution method for low-altitude target identification based on a hybrid expert network. The execution entity of this method is not limited to edge computing devices deployed in the low-altitude defense zone, a central server (output data from camera groups, multi-channel microphone arrays, and radar equipment is aggregated to the central server via fiber optic or 5G networks), or a collaborative computing system consisting of edge computing nodes and the central server. The edge computing device can be a processor, directly connected to the camera group, multi-channel microphone array, and radar equipment via Ethernet, USB (Universal Serial Bus), or GMSL (Gigabit Multimedia Serial Link) interfaces. The low-altitude defense zone refers to an airspace area centered on important protected targets (such as nuclear power plants, large event venues, airport airspace, critical infrastructure, etc.), typically ranging in altitude from 0 to 1000 meters, with a horizontal radius set according to actual security needs. Within this airspace, continuous monitoring, identification, and control of various low-altitude aircraft (especially drones, birds, model aircraft, and other low-speed, small targets) are required.

[0021] The cross-modal solution method for low-altitude target recognition based on a hybrid expert network provided in this embodiment is as follows: acquire and input the image data, audio data of the multi-channel microphone array, and radar detection data synchronously sampled at the current moment in the low-altitude defense zone into the solution model to obtain the solution result at the current moment.

[0022] Specifically, the execution process of the above method includes: Step S1: Acquire image data, audio data from a multi-channel microphone array, and radar detection data (referred to as multimodal sensor data) that are synchronously sampled within the low-altitude defense zone.

[0023] In this embodiment, the data acquisition interface continuously and synchronously samples the output data of camera groups, multi-channel microphone arrays, and radar equipment within the low-altitude defense zone according to a preset sampling frequency, and stores the collected multimodal sensor data in a memory or database. The subject of this invention can directly obtain multimodal sensor data from the data acquisition interface, or read multimodal sensor data from a memory or database.

[0024] In this embodiment, preferably, the camera group includes a visible light camera and an infrared camera. The purpose of deploying the camera group, multi-channel microphone array, and radar equipment is to capture the shape texture, thermal radiation, motor harmonics, spatial orientation, and three-dimensional geometric signals of low-altitude targets in different physical dimensions, providing the system with cross-modal physical observation information. To provide a temporal and spatial consistency basis for the subsequent geometric consistency calibration of radar anchoring voxels, visual probability voxels, and audio probability voxels, it is further preferred that the sensor data from each modality be unified to the same time reference and three-dimensional physical coordinate reference, specifically including: Data Acquisition and Time Synchronization: The execution entity employs precise time synchronization technology based on a unified clock source to align the timing of camera arrays, multi-channel microphone arrays, and radar equipment (i.e., multimodal heterogeneous sensors), such as using Precise Time Protocol (PTP) or hardware pulse protocol. To address the issue of inconsistent sampling frequencies among multimodal heterogeneous sensors, a timestamp-based linear interpolation and nearest-neighbor matching method is used to resample visible light images, infrared thermal imaging images, audio data, and radar detection data to a unified time reference point. This ensures timestamp consistency in the multimodal sensor data, reducing time asynchrony errors during cross-modal fusion from the ground up. Specifically, using the radar scan frame or the system master clock as the reference time axis, visible light frames, infrared frames, and audio clips within adjacent time windows are matched to the same target time, forming multimodal sensor data at the same moment.

[0025] Unified basic coordinates and primary geometric alignment: Before system deployment, the visible light camera, infrared camera, multi-channel microphone array, and radar equipment are jointly calibrated offline to obtain the extrinsic parameter matrix of each sensor relative to a unified global coordinate system or radar coordinate system, and to obtain the camera intrinsic parameter matrix and distortion coefficients, thus unifying the multimodal sensor data into the same three-dimensional coordinate frame. During operation, the system uses the above calibration parameters to complete three types of basic mappings: First, the range, azimuth, elevation, and Doppler detection results from the radar target detection results output by the radar front end are converted into three-dimensional candidate positions in a unified coordinate system (not limited to the global coordinate system or radar coordinate system); Second, the pixels in the visible light image and infrared thermal imaging image are converted into camera imaging rays, providing a directional basis for subsequent visual probabilistic voxel depth back-projection; Third, the sound source arrival angle estimated by the multi-channel microphone array is converted into a spatial direction cone in a unified coordinate system, providing a directional basis for subsequent audio probabilistic voxel projection.

[0026] Specifically, for visible light cameras and infrared cameras, the pixel positions can be mapped to imaging rays in the global coordinate system by using their respective camera intrinsic and extrinsic parameters and based on existing depth back-projection algorithms.

[0027] Specifically, for multi-channel microphone arrays, the direction of arrival of the sound source can be converted into a sound source direction vector in the global coordinate system by using array mounting extrinsic parameters.

[0028] In this embodiment, in order to improve the processing effect of the subsequent solution model, it is preferable to further preprocess the acquired multimodal sensor data and input the preprocessed multimodal sensor data into the solution model.

[0029] (1) Image data preprocessing of camera group To address optical distortion and abrupt changes in illumination in complex low-altitude environments, automatic white balance, brightness normalization, or contrast-limited adaptive histogram equalization methods are applied to visible light images and infrared thermal images respectively to enhance local target edges in low-light, high-light, or low-contrast scenes. Finally, the processed visible light and infrared thermal images are uniformly cropped or interpolated and scaled to a fixed resolution, and channel normalization is performed to construct a standard visual tensor flow. Simultaneously, the timestamp, camera intrinsic parameters, and camera pose information corresponding to each frame of the visible light and infrared thermal images are retained, enabling the visual features output by the subsequent visual expert network to be back-projected along the imaging ray depth in the global coordinate system to a 3D voxel mesh, forming visual probabilistic voxels.

[0030] (2) Audio data preprocessing of multi-channel microphone array The raw audio data of the multi-channel microphone array includes the audio signal of each microphone channel, the inter-channel phase difference, and the time difference of arrival. A short-time Fourier transform is performed on the audio signal of each microphone channel to obtain the audio time-frequency graph. Then, a Mel filter bank is used to filter the audio time-frequency graph to obtain the Log-Mel (Logarithmic Mel-spectrogram) spectral characteristics of each microphone channel. This characteristic is used to characterize the UAV rotor harmonics, motor comb spectrum, and their high-frequency energy distribution. The preprocessed audio data includes the Log-Mel spectral characteristics, inter-channel phase difference, and inter-channel time difference of arrival.

[0031] Step S2: Input the current image data, audio data from the multi-channel microphone array, and radar detection data into the solution model to obtain the solution result for the current moment. The solution model is pre-trained.

[0032] It should be noted that the multimodal sensor data from each sampling moment in the video stream output by the camera group in the low-altitude defense zone, the audio data stream output by the multi-channel microphone array, and the radar detection data stream output by the radar equipment can be input into the calculation model to obtain the calculation results at each sampling moment, thereby achieving continuous defense and executing target tracking tasks. The calculation results include target category, target 3D position or target 3D bounding box, target velocity, target trajectory ID, target trajectory confidence, and optional target threat level.

[0033] like Figure 1 As shown, the solution model includes: (1) Visual expert network, extracting visual features from image data.

[0034] Specifically, visual expert networks are not limited to using existing dual-Swin Transformer encoders, such as SalienTR (Saliency Transformer), or existing Transformer-based dual-stream encoders, such as CIGNet (Cross-modal Interaction Guidance Network), as the backbone network. They acquire multi-level features from visible light images and infrared thermal imaging images separately, and then use existing fusion modules such as ComFormer (Complementary Modality Fusion Transformer) and MICM (Modal Information Complementary Module) to fuse the multi-level features of the visible light images and infrared thermal imaging images, thereby obtaining visual features. Transformer represents a transformer.

[0035] (2) Audio expert network, which extracts audio latent features based on audio data. Preferably, the audio expert network also extracts the sound source direction vector based on the audio data and uses it as the sound source direction vector of the audio probability voxel, which facilitates the subsequent projection of the audio latent feature probability onto the three-dimensional voxel grid to obtain the audio probability voxel.

[0036] Existing low-altitude target perception systems fail to effectively decouple the spatial angle of arrival (calculated from the time difference of arrival of the audio channels) from high-frequency harmonics of the rotor in audio feature extraction. This leads to false alarms and misjudgments when facing hovering rotorcraft drones and birds circling overhead, hindering robust classification with high confidence. Furthermore, the audio signal is severely affected by environmental wind noise, wide-angle noise, and sidelobe interference, resulting in a high false alarm rate. Therefore, a preferred approach is to use an audio expert network... Figure 2 The structure shown includes audio data such as Log-Mel spectral characteristics, inter-channel phase differences, and time-of-arrival differences. The audio expert network includes: The sound source direction determination unit constructs a sound source direction vector based on the inter-channel phase difference and arrival time difference. The sound source direction vector is used for the spatial projection of subsequent audio probability voxels.

[0037] Specifically, using existing sound source localization algorithms, such as the time difference of arrival estimation method based on generalized cross-correlation-phase transform, or other well-known sound source localization algorithms in this field, the azimuth angle of the sound source is estimated based on the inter-channel phase difference and time difference of arrival of multiple microphone channels. and pitch angle And construct the source direction unit vector in the global coordinate system: , Represents the cosine function, sin This represents the sine function.

[0038] The first audio confidence prediction head obtains the source direction estimation confidence based on the source direction vector, which is used to constrain the weights of the audio probability voxels.

[0039] Specifically, the first audio confidence prediction head is learnable and is not limited to using the existing multilayer perceptron (MLP). Its output layer is a sigmoid activation function. The multilayer perceptron (MLP) is used to map the sound source direction vector to the sound source direction estimation confidence with values ​​between 0 and 1.

[0040] The second audio confidence prediction head obtains rotor harmonic confidence based on Log-Mel spectral features, which is used to help determine whether the sound source has the acoustic characteristics of a UAV rotor.

[0041] Specifically, the second audio confidence prediction head comprises a Lightweight Convolutional Neural Network (Lightweight CNN), a global pooling layer, a third fully connected layer, and a sigmoid activation function unit (not shown) connected in sequence. The Lightweight CNN extracts features such as the rotor fundamental frequency, higher harmonics, periodic comb spectrum, and frequency band energy distribution from the Log-Mel spectrum. These extracted features are then processed sequentially through the global pooling layer, the third fully connected layer, and the sigmoid activation function unit to obtain rotor harmonic confidence values ​​ranging from 0 to 1.

[0042] The spatial branch uses a first asymmetric convolution dilated along the time axis to process the Log-Mel spectral features to obtain spatial correlation features, represented as: ,in, Represents the Log-Mel spectral characteristics. This represents the first asymmetric convolution dilated along the time axis, with spatially correlated features. It reflects the phase drift and short-time arrival difference characteristics in the audio data.

[0043] The spectral branch uses a second asymmetric convolution dilated along the frequency axis to process the Log-Mel spectral features to obtain spectral correlation features, represented as: ,in, This represents an asymmetric convolution that expands along the frequency axis, with spectral correlation features. It reflects the high-frequency harmonics and frequency band energy distribution in the audio data.

[0044] Temporal modeling unit, based on spatial correlation features Spectral characteristics Obtaining latent audio features To achieve spatially relevant features Spectral characteristics Temporal aggregation and regularization processing. Specifically, the temporal modeling unit is not limited to the existing gated recurrent unit (GRU). The specific processing procedure is as follows: , Indicates spatially related features Spectral characteristics The splicing features are obtained by splicing along the channel dimension. Of course, the temporal modeling unit can also use existing bidirectional gated recurrent units (BiGRU), temporal attention networks (such as Transformer encoders), or temporal pooling layers (such as global average pooling layers).

[0045] The audio expert output definition set is as follows: .

[0046] The aforementioned audio expert network constructs two processing paths: a spatial branch and a spectral branch. The spatial branch focuses on capturing changes in phase difference or time of arrival between microphone channels to estimate the angle of arrival of the sound source; the spectral branch focuses on capturing the periodic comb-like spectral structure formed in the frequency domain by the fundamental frequency and higher harmonics of the UAV rotor and motor to distinguish between rotorcraft UAVs, birds, and environmental noise. This achieves decoupled modeling of audio spatial cues and frequency harmonic cues, avoiding limiting the audio expert to a simple audio classifier. Through a temporal modeling unit, the decoupled spatial and spectral correlation features are aggregated and regularized temporally. Utilizing the continuity of low-altitude target acoustic signals, the network can suppress the influence of environmental wind noise, urban background noise, and transient interference on audio features, obtaining stable latent audio features and reducing the false alarm rate introduced by the audio branch.

[0047] (3) Radar expert network, including a fast hierarchical module, which obtains fast hierarchical features of one or more radar targets based on radar target detection results, and generates radar anchoring voxels corresponding to each radar target in a three-dimensional voxel grid.

[0048] In this embodiment, the radar target detection results include the target range, azimuth, elevation, Doppler velocity, echo intensity, and radar detection confidence level for each of the above and below radar targets. A radar target is a target detected by the radar equipment.

[0049] Specifically, the fast hierarchical module performs the following steps: Step A1, based on the first Target range of radar target Azimuth Pitch angle Calculate the three-dimensional candidate point position of the radar target in the radar coordinate system. : .

[0050] Step A2: Perform coordinate transformation using the extrinsic parameter matrix from the radar coordinate system to the global coordinate system to obtain the first... Global three-dimensional candidate locations of radar targets : ;in, and Let represent the rotation matrix and translation vector from the radar coordinate system to the global coordinate system, respectively. Represents the cosine function. Represents the sine function; This is the radar target index, a positive integer.

[0051] Step A3, splicing the first Global three-dimensional candidate locations of radar targets Doppler velocity echo intensity and radar detection confidence The obtained splicing features The first is obtained through linear mapping or lightweight embedding layer processing. Rapid hierarchical characteristics of individual radar targets ; in, ; Indicates a splicing operation; This represents linear mapping or lightweight embedding layer processing, which can be implemented through a single-layer fully connected network without activation functions, or through a multilayer perceptron (MLP) containing a hidden layer. Its function is to organize radar target detection results into a unified feature dimension, rather than relearning or replacing the target detection function of the radar front end. Therefore, the fast hierarchical module has the characteristics of low computational load, fast response speed, and suitability for continuous cruise and initial spatial anchoring.

[0052] Step A4: Generate the radar anchoring voxel for each radar target in a 3D voxel mesh. This is based on a global coordinate system and a preset voxel mesh resolution. Generate 3D voxel mesh According to the Global three-dimensional candidate locations of radar targets Generate its radar anchoring voxel : ; Indicates the voxelization operation; Indicates the first Radar anchoring voxel corresponding to each radar target Additional radar detection attributes, including , , And the radar target number. It is clear that the rapid hierarchical module does not need to relearn the target location.

[0053] (4) Cross-modal adaptive calibration module: visual feature depth back-projection to three-dimensional voxel grid to obtain visual probability voxels, audio latent feature probability projection to the three-dimensional voxel grid to obtain audio probability voxels, and fusion of visual probability voxels and audio probability voxels that meet geometric consistency conditions in the neighborhood of each radar anchor voxel with the radar anchor voxel to generate the three-dimensional fused voxel grid at the current moment.

[0054] Specifically, this includes, but is not limited to, using visual LSS depth backprojection (Lift-Splat-Shoot, or LSS for short) to backproject visual features depth onto a 3D voxel mesh to obtain visual probability voxels, including: Step B1, for visual features Spatial position point Its corresponding eigenvalues ​​can be expressed as This location point can be determined based on the feature map step size (visual features). The number of pixels in the visible light image and infrared thermal image corresponding to a position point is mapped back to the original image coordinates, and the corresponding camera imaging ray is determined by combining the camera intrinsic and extrinsic parameters.

[0055] Step B2, at location point On the corresponding camera imaging ray, further predict or associate the location point. Depth probability across multiple discrete depth intervals, such as location points In the Discrete depth range The depth probability on is expressed as , For discrete depth range indexes It is learnable during model training and the visual feature values ​​are assigned according to this depth probability. Projected onto a 3D voxel mesh From the voxels corresponding to the spatial locations, the visual probability voxels are obtained. For three-dimensional voxel meshes any voxel in Its visual voxel features It can be obtained by averaging the visual feature values ​​falling within this voxel according to depth probability: ; in, Indicates visual features Upper position point According to discrete depth range The position projected back into three-dimensional space To prevent the use of tiny constants with a denominator of zero. This is necessary if there is no camera imaging ray or no depth sampling point falls into a voxel. If the visual voxel feature (i.e., the feature of the visual probability voxel) is zero, then the voxel can be marked as an empty voxel. Visual confidence. It can be used as an additional visual attribute for all visual probability voxels.

[0056] Visual probabilistic voxels retain two-dimensional visual semantic information and depth uncertainty, thus enabling them to express the possible three-dimensional locations of visually suspected targets. However, their depth estimates may drift, requiring subsequent geometric calibration by radar-anchored voxels.

[0057] The audio probability voxels are obtained by projecting the audio latent feature probability onto a three-dimensional voxel grid, including: since multi-channel microphone arrays typically only provide the direction of the sound source and cannot provide precise distance alone, this embodiment uses the sound source direction vector... An audio probability direction cone is constructed in a three-dimensional voxel space, and audio probability voxels are generated by combining the distance or neighborhood constraints with each radar anchoring voxel. : ; in, This indicates an operation of probabilistic projection along an audio probability direction cone. Unlike other methods that only use audio as a category feature in the stitching process, this invention transforms latent audio features into spatially directional probabilistic voxels, enabling them to participate in subsequent 3D geometric consistency calibration. (Radar anchoring voxel) This is used here to constrain the effective distance range and local spatial neighborhood of the audio direction cone, reducing spatial divergence caused by wind noise, multiple sound sources, and wide-angle positioning errors. This represents the set of radar anchor voxels corresponding to all radar targets. All are audio probability voxels Additional audio attributes. For potential audio features, The direction vector of the sound source. Estimate the confidence level for the direction of the sound source. Confidence level for rotor harmonics.

[0058] For three-dimensional voxel mesh any voxel Let its center position be The position of the multi-channel microphone array in the global coordinate system (its center point position) is: Then it can be based on arrive The voxel direction vector (represented as) ) and the direction vector of the sound source Consistency of angles, and voxels central position Global three-dimensional candidate positions of the i-th radar target Calculate the spatial proximity of the voxel. The audio projection weight for the i-th radar target ; in, This represents the consistency weight of the sound source direction, specifically the sound source direction vector. and arrive voxel direction vector Cosine similarity; This represents the radar anchoring neighborhood constraint weights, specifically, if voxels... center Global three-dimensional candidate position of the i-th radar target Within its neighborhood, then ,like Not located Within its neighborhood, then Alternatively, existing Gaussian radial decay methods can be used to obtain... Continuous values. Voxels. The audio probability voxel feature of the i-th radar target It can be represented as: If the body element If it is not within the neighborhood of the sound source direction cone or the radar anchoring voxel, then the corresponding weight It can be set to zero, that is .

[0059] In this embodiment, the first The neighborhood of a radar anchor voxel is defined by the center of that radar anchor voxel. Centered on the sphere, with a preset spatial residual threshold A spherical spatial region with a radius of 1; or a three-dimensional neighborhood formed by extending outward from the voxel grid cell where the radar anchor voxel is located by a preset number of grid layers (such as 1 or 2 layers).

[0060] In this embodiment, the sound source direction vector extracted by the audio expert network based on the audio data is used as the sound source direction vector of the audio probability voxel, which is an additional audio attribute of the audio probability voxel; preferably, the geometric consistency condition is: and ; in, ; Indicates the first The center of the visual probability voxel satisfying the visual candidate condition within the neighborhood of each radar anchor voxel. With the The center of each radar anchor voxel Spatial residuals; Represents the spatial residual threshold, for example, , This represents the voxel grid resolution of the 3D voxel grid. It should be noted that the visual candidate condition can be that the visual probability voxel feature (also called feature) of the visual probability voxel is not zero, or that the depth probability of the visual probability voxel is greater than a preset depth probability threshold. in, ; Indicates the first The sound source direction vector of the audio probability voxel satisfying the audio candidate condition within the neighborhood of each radar anchor voxel. With the Radar target direction (point) Time The direction is represented as Angular residuals between ) Represents the direction vector of the sound source with vector The included angle, Represents the cosine function; This indicates the position of the multi-channel microphone array in the global coordinate system; Indicates the angular residual threshold; This indicates the calculation of the L2 norm. It should be noted that the audio candidate condition can be that the audio probability voxel feature (also called feature) of the audio probability voxel is not 0, or that the sum of the audio projection weights of the audio probability voxels for all radar targets is greater than the preset audio projection weight threshold.

[0061] It is evident that for visual and audio probability voxels to satisfy the geometric consistency condition, they must first satisfy the visual candidate condition and the audio candidate condition, respectively, and also be located within the neighborhood of the radar anchor voxel. This geometric consistency condition, through dual threshold constraints of spatial and angular residuals, allows only visual and audio probability voxels that maintain a high degree of consistency with the radar anchor voxel in spatial location and sound source direction to participate in cross-modal fusion, thus suppressing visual background misprojection and wide-angle audio noise.

[0062] In a further preferred embodiment, the cross-modal adaptive calibration module fuses the visual probability voxels and audio probability voxels that satisfy the geometric consistency condition within the neighborhood of each radar anchor voxel with the radar anchor voxel to generate a three-dimensional fused voxel mesh for the current moment, including: Step C1, generate the first Local voxel calibration mask corresponding to the neighborhood of each radar anchor voxel : ; Represents the threshold determination function; when the following conditions are met... hour, Equals 1, when not satisfied hour, Equals 0; when satisfied hour, Equals 1, when not satisfied hour, Equals 0; Step C2, will the first The visual probability voxels and audio probability voxels that satisfy the geometric consistency condition within the neighborhood of the first radar anchoring voxel are related to the first... The radar anchoring voxel fusion was performed to obtain the first... One or more three-dimensional fusion voxels corresponding to the neighborhood of each radar anchoring voxel; Step C3, obtain the first according to the following formula Fusion voxel features of three-dimensional fused voxels corresponding to the neighborhood of each radar anchoring voxel. : ; in, Indicates the first Radar detection confidence level of a radar target. Indicates will Mapped to multimodal confidence; This represents the attention mechanism. For query, As key, Value; Indicates the first Features of a radar anchoring voxel Indicates the first Features of visual probability voxels that satisfy geometric consistency conditions within the neighborhood of a radar anchoring voxel. Indicates the first The characteristics of audio probability voxels that satisfy the geometric consistency condition within the neighborhood of a radar anchor voxel (also called audio probability voxel features). This indicates splicing along the channel dimension. and The obtained splicing features; Step C4: Set the features of voxels in the 3D voxel mesh that do not belong to the 3D fused voxel to zero to obtain the 3D fused voxel mesh at the current moment.

[0063] In this embodiment, a local voxel calibration mask is used to allow only visual and audio voxels located in the neighborhood of the radar anchor point and meeting the geometric consistency condition to participate in subsequent cross-modal calibration. The local voxel calibration mask is the key difference between the cross-modal adaptive calibration module and ordinary global cross-attention fusion. Ordinary attention usually calculates feature correlation over a large range, which can easily introduce visual background misprojection or wide-angle audio noise into the fusion result. In contrast, this invention first uses the geometric consistency condition to limit the range of voxels that can participate in fusion, and then performs feature calibration within this local neighborhood.

[0064] In step C3 of this embodiment, under the constraint of local voxel calibration mask, the system uses radar anchor voxel features as query terms and visual and audio probability voxels that satisfy geometric consistency conditions as keys and values ​​to perform local cross-modal attention calibration. During this process, radar anchor voxels provide a defined three-dimensional position reference, visual probability voxels provide texture, thermal radiation, and semantic information, and audio probability voxels provide evidence of sound source direction and rotor harmonics. This allows for the absorption of complementary visual and audio information while preserving the three-dimensional geometric determinism of the radar.

[0065] Fusion voxel features obtained in step C3 It features trimodal geometric consistency, unlike two-dimensional stitching features or BEV (Bird's Eye View) features. The three-dimensional fused voxel mesh preserves the target's height, orientation, distance, and local three-dimensional neighborhood relationships. Visual probability voxels and audio probability voxels achieve cross-modal calibration under the geometric constraints of radar anchored voxels, eliminating visual depth drift, spatial divergence of audio probability cones, and unconstrained global fusion problems of ordinary cross-attention, thereby improving the stability of low-altitude target three-dimensional localization, category recognition, and cross-frame trajectory calculation.

[0066] (5) The Transformer spatiotemporal correlation decoder processes the 3D fused voxel mesh and the trajectory query set from the previous time step through an attention mechanism to obtain the trajectory query set for the current time step. Specifically, the Transformer spatiotemporal correlation decoder can be implemented using an end-to-end Transformer decoder network.

[0067] Most existing target tracking technologies employ a separate paradigm of detection followed by tracking, which is sensitive to physical occlusion and suffers from trajectory fragmentation and frequent jumps. While some schemes based on Transformer decoders (such as end-to-end Transformer decoders) or spatiotemporal query vectors enhance cross-frame correlation capabilities, their query states are still often based on two-dimensional detection boxes, visual appearance, or short-term motion predictions, lacking a three-dimensional identity preservation mechanism supported by radar detection confidence, target velocity, visual information, and audio information. Therefore, it is still difficult to output continuous, stable, and complete 3D flight trajectories in complex occlusion scenarios. Therefore, this invention preferably provides a deep improvement to the MOTR-like (Multiple-Object Tracking with Transformer) architecture, aiming to solve the problems of the traditional detection-follow-correlation paradigm from the underlying logic.

[0068] Therefore, in the preferred embodiment of this example, a trajectory query vector is constructed for each target, and the trajectory query vectors of all targets form a trajectory query set. This represents the target index, which is a positive integer. The current time is the [current time unit]. Trajectory query vector of each target for: ; in, This represents a trajectory query vector embedding mapping. Indicates the previous moment. Trajectory query vector for each target Indicates the current time. Global three-dimensional position features of a target Indicates the current time. The velocity characteristics of the target Indicates visual confidence level. This indicates the confidence level for the sound source direction estimation. This indicates the confidence level of rotor harmonics. , It was initialized. Indicates the first The radar detection confidence of the nth target, if the nth target is... If the global three-dimensional position of a target is located within the neighborhood of the i-th radar anchor voxel, then For the first The radar detection confidence of the nth radar target, if the nth radar target is a radar target with a radar detection confidence of a radar target with If the global three-dimensional position of a target is not within the neighborhood of any radar anchor voxel, then =0, The specific value is obtained after decoding by the Transformer spatiotemporal correlation decoder. By introducing a trajectory query vector, if a target is briefly occluded in the current frame (current moment), the trajectory query vector can still maintain its state based on historical 3D motion status and weak radar / audio observation clues, instead of immediately deleting the trajectory.

[0069] In this preferred embodiment, the trajectory query vectors of all targets at the previous time step constitute the trajectory query set at the previous time step, denoted as: Current moment The 3D fused voxel mesh is represented as The trajectory query vectors of all targets at the current moment constitute the trajectory query set at the current moment, denoted as: The Transformer spatiotemporal correlation decoder processes the trajectory query set from the previous time step through a self-attention mechanism. It is used to model the mutual exclusion, occlusion competition, and neighboring target relationships among different trajectory query vectors. The Transformer spatiotemporal correlation decoder processes the trajectory query set from the previous time step using a self-attention mechanism. For querying, at the current time The 3D fused voxel mesh is represented as Using the key and value, perform cross-attention processing to obtain the trajectory query set at the current time. Cross-attention is performed in a 3D fused voxel grid. The query prioritizes 3D fused voxel regions that are consistent with its historical 3D position, velocity prediction and radar anchoring voxel neighborhood, thereby reducing the interference of background occlusion or visual false detection on trajectory updates.

[0070] In this preferred embodiment, 3D identity preservation under occlusion conditions can be achieved: when a low-altitude target moves against a background of buildings, tree canopies, bridges, or cables, visual confidence may drop suddenly for a short period, but radar position, velocity, or audio direction cues may still provide weak observations. This embodiment provides a trajectory survival state and occlusion memory mechanism: when visual confidence decreases but the radar anchor voxel still exists, or the audio direction is consistent with the historical trajectory direction, the system marks the trajectory as occlusion-preserving and makes a short-term prediction based on historical velocity and global 3D position, instead of immediately generating a new trajectory ID.

[0071] If visual observations are missing at the current moment (current frame), or the visual confidence level is lower than a preset threshold, but the following conditions are met: or ; Then keep the original trajectory query vector The identity ID is used to reduce the confidence level of the trajectory instead of deleting it directly. Indicates by the first The current position is obtained by predicting the historical trajectory of each target. Indicates the first Each radar target corresponds to a radar candidate location. This represents the direction vector of the sound source. express and Spatial distance, Representing vectors and The angle between ray vectors. This mechanism enables the system to reassociate the same target after a short period of occlusion, reducing ID jumps and trajectory fragmentation.

[0072] In this preferred embodiment, for a re-emerging target, not only are spatial distance and motion continuity compared, but also micro-Doppler features, rotor harmonic confidence levels, and visual features are combined to determine identity consistency. For example, for small rotorcraft drones, micro-Doppler sidebands and rotor harmonics have a certain short-term stability, and micro-Doppler features and rotor harmonic confidence levels can be used as auxiliary identity fingerprints in addition to visual features; for birds or clutter targets, the above physical fingerprints usually differ from those of drones, thus reducing the probability of false association.

[0073] This preferred embodiment combines 3D fusion voxels, trajectory query vectors, and cross-modal physical identity fingerprints to achieve a transition from single-frame detection to continuous 3D trajectory calculation. Compared to tracking methods that rely solely on 2D detection boxes, IoU (Intersection over Union) matching, or visual appearance similarity, it can utilize radar 3D position, target velocity, micro-Doppler features, and audio harmonic cues for occlusion preservation and re-association. This is beneficial for outputting continuous, stable, and uniquely identified 3D flight trajectories in complex low-altitude environments, such as those with building occlusion, tree canopy crossings, and other challenging conditions.

[0074] (6) Multi-task solution head, which obtains the solution result at the current time based on the trajectory query set at the current time.

[0075] In this embodiment, the solution results include solution information for each target. This solution information includes target category, target 3D position or target 3D bounding box, target velocity, target trajectory ID, target trajectory confidence level, and an optional target threat level. The target at the current moment The solution information is : .

[0076] in, Indicates the target category. Indicates a 3D bounding box or 3D position. Indicates the target speed. Indicates cross-frame trajectory identifier, Indicates the confidence level of the detection or trajectory. This indicates the threat level. The target category is not limited to drones, birds, clutter, or other low-altitude target types; the target's 3D position or 3D bounding box and target velocity are used to form a continuous trajectory; the target trajectory ID is used to maintain identity consistency across frames; the threat level can be calculated by combining the target's distance from the control center, the target's velocity, the target's approach direction, and the target category confidence level.

[0077] In this embodiment, the multi-task calculation head includes parallel classification, regression, confidence, and threat assessment heads, all of which employ a multilayer perceptron (MLP) structure, differing only in their output layers. The classification head outputs the target category; its output layer can be a softmax activation function to obtain the category confidence for each target type among drones, birds, clutter, or other low-altitude targets. The regression head's output layer is a linear layer, directly outputting the target's 3D position or 3D bounding box and velocity. The confidence head's output layer is also a linear layer, directly outputting the target trajectory confidence. The threat assessment head's output layer is also a linear layer; its inputs are the distance between the target and the control center, the target velocity, the target's approach direction, and the target category confidence, with the output being the target threat level.

[0078] Example 2: This embodiment discloses a cross-modal solution method for low-altitude target recognition based on a hybrid expert network. The difference between this method and Embodiment 1 lies in the different visual expert network structure. To address issues such as visible light failure at night, insufficient infrared detail texture during the day, and interference from camouflage backgrounds, such as... Figure 3 As shown, the visual expert network provided in this embodiment includes: (1) Visible light branch, including a visible light backbone network and a visible light pyramid network connected in sequence, used to extract multi-scale first enhanced image features from visible light images; Specifically, the visible light backbone network adopts the existing CSPDarknet network (Cross-Stage Local Darknet), which is the backbone network of the existing YOLO (You Only Look Once) series. The visible light pyramid network is not limited to using FPN (Feature Pyramid Network) or PAN-FPN (Path Aggregation Feature Pyramid Network). The visible light backbone network extracts multi-scale first image features from the visible light image, and the visible light pyramid network fuses and processes the multi-scale first image features to obtain multi-scale first enhanced image features. (Arbitrary scale) The first enhanced image feature is represented as Key Scale The first enhanced image feature is represented as .

[0079] (2) Infrared branch, including an infrared backbone network and an infrared pyramid network connected in sequence, is used to extract multi-scale second enhanced image features from infrared thermal imaging images; Specifically, the infrared branch can adopt the network structure of the visible light branch, but the network parameters are not shared. The infrared backbone network uses the existing CSPDarknet network; the infrared pyramid network is not limited to using FPN (Feature Pyramid Network) or PAN-FPN (Path Aggregation Feature Pyramid Network). The infrared backbone network extracts multi-scale second image features from infrared thermal imaging images, and the infrared pyramid network fuses and processes these multi-scale second image features to obtain multi-scale enhanced image features. Arbitrary scale. The second enhanced image feature is represented as Key Scale The second enhanced image feature is represented as The key scale is preferably a medium-to-high resolution scale with high spatial resolution and still possessing a certain semantic expressive power, in order to preserve the local texture, edge, and hotspot information of low-altitude small targets.

[0080] (3) Cross-modal feature bidirectional interaction network, including: The first convolution enhances the second image feature at a preset key scale. Perform convolution processing; The first activation unit maps the features output by the first convolution to a hot target space mask. : ; Represents the first convolutional mapping; This represents the first activation unit, preferably the Sigmoid activation function (S-type function). The first multiplication unit will use the first enhanced image features at a preset key scale. With thermal target space mask Element-wise multiplication is performed to obtain the visible light target region features at a preset key scale. : ; This indicates element-wise multiplication; The first CSP fusion module fuses visible light target region features at preset key scales. and all non-preset key scale first enhanced image features Obtain visible light fusion features CSP stands for Cross-Stage Partial; preferably, the first enhanced image features that are not at a preset key scale are first... pass The number of channels is unified by convolution, and then unified to the same spatial resolution by upsampling or downsampling, and then input into the first CSP fusion module respectively; The second convolutional unit enhances the first image features at a preset key scale. Perform convolution processing; The second activation unit maps the features output by the second convolution to a texture contour mask. : ; Indicates the second convolutional mapping; This represents the second activation unit, preferably the Sigmoid activation function; The second multiplication unit will use the second enhanced image features at a preset key scale. With texture contour mask Element-wise multiplication is performed to obtain infrared contour features at preset key scales. : ; The second CSP fusion module fuses infrared contour features at preset key scales. and all second-enhanced image features not based on preset key scales Obtain infrared fusion features Preferably, the second enhanced image features that are not at a preset key scale are first... pass The number of channels is unified by convolution, and then unified to the same spatial resolution by upsampling or downsampling, and then input into the second CSP fusion module respectively; The visible light weight acquisition unit includes a first global average pooling layer and a first fully connected layer connected in sequence, used to map visible light fusion features to visible light weights. ; The infrared weight acquisition unit includes a second global average pooling layer and a second fully connected layer connected in sequence, used to map infrared fusion features into infrared weights. ; The weighted fusion unit fuses the visible light fusion features and infrared fusion features based on visible light weights and infrared weights to obtain visual features. : ; Indicates the convolution operation; Indicates a splicing operation; The reliability of the visible light and infrared modes in the current frame is adaptively estimated by the visible light weight acquisition unit and the infrared weight acquisition unit, and the weighted fusion unit is used to obtain the weighted fused visual features. This enables visible light texture features to supplement the infrared contour in daytime or fully textured scenes; and in nighttime, low-light, or backlight scenes, the gating unit increases the contribution of infrared branches and thermal target guidance features, thereby enhancing all-weather visual perception capabilities.

[0081] (4) Visual confidence prediction head, which obtains visual confidence based on visual features.

[0082] In this embodiment, the visual confidence prediction head is not limited to a multilayer perceptron (MLP) that maps visual features to visual confidence in the range of 0 to 1.

[0083] Example 3: This embodiment discloses a cross-modal solution method for low-altitude target identification based on a hybrid expert network, which improves the radar expert network based on Embodiment 1 or Embodiment 2.

[0084] In this embodiment, the radar detection data includes not only radar target detection results, but also radar point cloud data and echo data; such as Figure 4 As shown, the radar expert network includes not only the fast hierarchical module in Embodiments 1 and 2, but also a full hierarchical module. The full hierarchical module includes: (1) Point cloud topology branch, including a preprocessing unit and a point cloud feature extraction network. The preprocessing unit performs ground point filtering, outlier removal and voxelization on the radar point cloud data. The point cloud feature extraction network is used to extract point cloud topology features from the radar point cloud data processed by the preprocessing unit. Specifically, the point cloud feature extraction network is not limited to sparse 3D convolution, PointNet++ encoder, or lightweight point cloud Transformer. (2) Micro-Doppler branch, including short-time Fourier transform unit and time-frequency feature extraction network. The short-time Fourier transform unit performs short-time Fourier transform on the echo data to obtain micro-Doppler time-frequency map. A time-frequency feature extraction network is used to extract features from micro-Doppler time-frequency maps. Extracting MicroDoppler Features Specifically, the time-frequency feature extraction network is not limited to two-dimensional convolution or time-frequency attention encoder; Represents echo data; This indicates the short-time Fourier transform operation; This represents the processing procedure of the time-frequency feature extraction network. (3) Full fusion unit, fusing point cloud topological features and micro Doppler features Obtaining fine-grained radar discrimination features : ; in, , These represent the first weight and the second weight, respectively, which are learnable parameters; This is not limited to concatenating and mapping two features, attention aggregation, or gated weighting. When the point cloud topology branch or micro-Doppler branch is not activated (i.e., not enabled), the corresponding weights... , It can be set to zero or kept at a low value.

[0085] In this embodiment, the features of each radar anchor voxel include the rapid hierarchical features and the radar fine-grained discrimination features of the radar target corresponding to that radar anchor voxel. When the full-level module is enabled, the first... Features of a radar anchor voxel for: , Indicates the first Each radar anchor voxel corresponds to a rapid hierarchical feature of the radar target; when only the point cloud topology branch is enabled in the full hierarchical module, the first... Features of a radar anchor voxel for: When only the micro-Doppler branch is enabled in the full-level module, the first... Features of a radar anchor voxel for: .

[0086] While the fast hierarchical module in radar expert networks can generate radar anchoring voxels at low cost, it cannot solve the problem of decreased stability in category judgment and trajectory recognition under high-risk scenarios such as increased target threat, insufficient target classification confidence, occlusion recovery, or increased cross-modal geometric inconsistency. Therefore, this embodiment sets up a full-level hierarchical module in the radar expert network. Specifically, this is achieved through point cloud topology features... Describing the spatial distribution, local density, and three-dimensional structural morphology of target point clouds using micro-Doppler features By capturing the periodic micro-Doppler modulation caused by the UAV rotor and distinguishing it from bird flapping, ground clutter, or other non-UAV targets, a fine-grained radar discrimination feature can be obtained. This allows the feature to be used as a category refinement and structural verification feature in subsequent calculations under high-threat, low-confidence, or complex scenarios, improving the stability of category judgment and trajectory recognition. Through the above design, the fast hierarchical module in the radar expert network is responsible for providing 3D position priors and radar anchoring voxels at low cost, while the full-level module supplements point cloud topology features and micro-Doppler features when precise qualitative analysis or complex scene recognition is required. This avoids the waste of computing power caused by high-cost point cloud / echo depth processing for every frame, while retaining the ability to perform fine recognition in high-risk scenarios.

[0087] With the simultaneous introduction of branches such as vision, infrared, radar point cloud, micro-Doppler, audio array, and spatiotemporal decoder, the computational load of the solution model increases significantly. Existing solutions mostly employ fixed computation graphs, keep all experts constantly active, or switch between lightweight and heavy branches only during the training phase. They lack a closed-loop scheduling mechanism driven by the current solution results (target threat level, classification confidence), the load status of the execution subject of the cross-modal solution method based on hybrid expert networks for low-altitude target recognition, and cross-modal inconsistency. This leads to wasted computing power in low-threat scenarios, while in high-risk scenarios such as rapid target approach, uncertain category, or conflicting modal results, the point cloud topology branch and micro-Doppler branch may not be activated in time, making it difficult to balance real-time performance and fine recognition capabilities at the edge.

[0088] Therefore, in this preferred embodiment, the solution model further includes a computing power scheduling decision-maker. The computing power scheduling decision-maker determines the control action for the next moment based on the solution result at the current moment, the load state of the execution entity of the cross-modal solution method based on the hybrid expert network for low-altitude target identification, and the cross-modal inconsistency. : ; in, Indicates the start / stop flag for the fast-level module. Indicates the start and stop flags for point cloud topology branches. Indicates the start and stop indicators of the microDoppler branch. This represents the voxel grid resolution of a 3D voxel grid. This indicates the radar sampling frequency. For example, , , 1 indicates startup, and 0 indicates shutdown.

[0089] In this preferred embodiment, all or part of the parameters can be extracted from the current solution result and input into the computing power scheduling decision unit, such as the target speed, target trajectory confidence, and target threat level in the solution result. The selected target speed, target trajectory confidence, and target threat level are then mapped into a threat score using a linear layer. Threat scores can also be mapped to threat scores through pre-executed threat scoring rules. It is used to determine whether the target needs to be finely identified.

[0090] In this preferred embodiment, the load status of the execution entity is not limited to being determined by factors such as single-frame inference latency, real-time frame rate, GPU (Graphics Processing Unit) / CPU (Central Processing Unit) / NPU (Neural Processing Unit) utilization, video memory utilization, execution entity device temperature, and task queue length. Specifically, the load status can be mapped to a device load score using a linear layer or a predefined load scoring rule. It is used to determine whether the edge end (execution subject) is allowed to open a high-cost branch.

[0091] In this preferred embodiment, the cross-modal inconsistency degree is expressed as: The mean of the spatial residuals, the mean of the angular residuals, and the mean of the confidence residuals of all radar anchor voxels. This is jointly determined to assess whether there are conflicts or drifts in the current cross-modal results.

[0092] Specifically, the mean spatial residual of all radar anchor voxels for: ; Indicates the first Among the visual probability voxels that satisfy the geometric consistency condition within the neighborhood of a radar anchoring voxel, the center of the visual probability voxel is parallel to the first... Radar anchoring voxel center The maximum spatial residual (maximum spatial distance). This represents the number of radar anchor voxels and is a positive integer.

[0093] Specifically, the mean angular residual of all radar anchor voxels for: ; Indicates the first Among the audio probability voxels satisfying the geometric consistency condition within the neighborhood of the first radar anchoring voxel, the sound source direction vector of the audio probability voxel is related to the first... The maximum angular residual between the directions of each radar target.

[0094] Specifically, the mean of the confidence level residuals Audio confidence ; Indicates the number of radar anchoring voxels, and is a positive integer; This indicates taking the absolute value.

[0095] In obtaining the mean of spatial residuals Mean of angular residuals and confidence level residual mean Then, the mean of the spatial residuals is calculated using a linear layer or a pre-defined inconsistency mapping rule. Mean of angular residuals and confidence level residual mean Mapped to cross-modal inconsistency The mean of the spatial residuals can also be used. Mean of angular residuals and confidence level residual mean Multiplication yields cross-modal inconsistency. .

[0096] In one example, the computing power scheduling decision arbiter uses preset scheduling rules based on threat scores. Equipment load rating Cross-modal inconsistency Obtain control action in the next moment The scheduling rules are not limited to: Four scheduling states are set: patrol state, alert state, precision state, and recovery state. Each scheduling state corresponds to a decision condition and a control action. (Threat score...) Equipment load rating Cross-modal inconsistency When the criteria for a certain type of scheduling state are met, the corresponding control action is executed. Cruise state corresponds to low threat and stable confidence (i.e., cross-modal inconsistency). In smaller scenarios, only the radar's rapid hierarchical operation is maintained. , , Alert status corresponds to threat score General modal result conflict (cross-modal inconsistency) In larger scenarios, point cloud topology branches can be enabled locally. , , Or increase the sampling frequency; accurately identify the target's status corresponding to entering the core defense zone, recovery from obstruction, uncertain category, or high threat level (threat score). Larger cross-modal inconsistency In scenarios with larger scale, both point cloud topology branches and micro-Doppler branches can be enabled simultaneously. , , The recovery state corresponds to the transition phase after the target threat has decreased or the confidence level has stabilized. After the cooldown period is met, all levels are gradually turned off. , , ).

[0097] In another example, the computing power scheduling decision maker can use a lightweight multilayer perceptron network to obtain control actions.

[0098] Specifically, threat scoring Equipment load rating Cross-modal inconsistency The state vector is composed of a computational power scheduling decision-maker, which includes one or more fully connected layers and an output layer. The output layer includes components for outputting... , , The softmax activation function for the state, and the output. A linear regression layer is used. The current state vector is input into the computing power scheduling decision-maker to obtain the control action for the next time step.

[0099] This embodiment can reduce average inference latency and power consumption by using only low-cost, fast hierarchical modules in low-threat, long-distance, or stable confidence scenarios. In high-threat, strong occlusion, category uncertainty, or modal conflict scenarios, the system can promptly activate point cloud topology and micro-Doppler full branches to enhance fine recognition capabilities. When the device load is high, the system can maintain real-time performance through branch selection, resolution adjustment, and call frequency control, thereby balancing robust recognition, continuous 3D trajectory calculation, and real-time deployment efficiency at the edge.

[0100] Example 4: This embodiment provides a cross-modal solution method based on a hybrid expert network for low-altitude target recognition. Building upon Embodiment 3, this embodiment provides a training method for the solution model. Specifically, it includes: Step D1 involves constructing a sample set by acquiring multimodal sensor data from multiple consecutive historical moments simultaneously sampled within the low-altitude defense zone. This multimodal sensor data includes image data output from camera arrays, audio data from multi-channel microphone arrays, and radar detection data from radar equipment. Each historical moment's multimodal sensor data is treated as a sample, and labels are manually assigned to each sample. These labels include the target category, target 3D location or target 3D bounding box, target velocity, target trajectory ID, target trajectory confidence level, and target threat level.

[0101] Step D2: Divide the sample set into training set, test set and validation set according to a preset ratio (e.g., 8:1:1). It should be noted that the sampling time of the samples in the training set, test set and validation set is continuous.

[0102] Step D3: Construct the network structure of the solution model according to Example 3, and initialize the network parameters.

[0103] Step D4 involves training the solution model using the training set, specifically including: Initial Training: The initial training model is constructed using fast hierarchical modules from the visual expert network, audio expert network, and radar expert network in the solution model. This initial model is trained in batches using the training set. After each batch is completed, the network parameters of the initial training model are updated based on the geometric alignment loss until the initial stopping condition is met, thus entering the intermediate training phase. The geometric alignment loss is: ,in, Indicates visual weight, Indicates audio weights, Represents multimodal weights. Indicates the first Mean of spatial residuals of all radar anchor voxels in a sample Indicates the first Mean angle residuals of all radar anchor voxels in a sample Indicates the first Mean confidence residuals of each sample. The number of samples in each batch, Index the samples in each batch. , , The calculation method is the same as in Example 3. , , The calculation method is the same and will not be repeated here. The initial stopping condition is not limited to the geometric alignment loss being less than the preset geometric alignment loss threshold.

[0104] Mid-training phase: The solution model is trained in batches using the training set, with only the fast hierarchical module enabled in the radar expert network. The network parameters of the solution model are updated based on the second-stage loss until the mid-training stopping condition is met, entering the late-training phase. The mid-training stopping condition is not limited to the second-stage loss being less than a preset second-stage loss threshold. Second-stage loss Including geometric alignment loss Classification loss and 3D positioning loss ; . and This is used to constrain the target category and the target's 3D location output. Specifically, using the true target category in the sample's label and the target category in the sample's solution result, the classification loss is calculated based on the existing cross-entropy loss function. Using the true 3D location of the target in the sample labels and the 3D location of the target in the sample solution results, the 3D localization loss is obtained based on the existing smoothed L1 loss. .

[0105] In the later stages of training: the solution model is trained in batches using the training set, and both the fast hierarchical module and the full hierarchical module are enabled simultaneously in the radar expert network. The network parameters of the solution model are updated based on the third-stage loss until the later-stage stopping condition is met. The later-stage stopping condition is not limited to the third-stage loss being less than a preset third-stage loss threshold. Third-stage loss Including geometric alignment loss Classification loss 3D positioning loss and fine classification loss , Fine-grained classification loss The radar's fine-classification features are calculated based on the fine-classification results and the true target categories in the sample labels. Specifically, the fine-classification head is an auxiliary classifier whose input is the fine-classification features. The output is a fine-grained classification result (specifically, the probability distribution for each target category). The auxiliary fine-grained classification head is not limited to a multilayer perceptron (MLP); it can be trained uniformly with the solution model or pre-trained. Fine-grained classification loss. The loss function is calculated based on the fine classification results of the samples and the true target category in the sample labels.

[0106] Step D5: Test and validate the solution model using the test set and validation set respectively. If the test and validation pass, end the training. If the test and / or validation fail, adjust the training parameters and return to step D4 until the preset maximum number of training rounds is reached.

[0107] In the early stages of training, the visual expert network and the audio expert network primarily learn how to map their own outputs to a 3D spatial neighborhood near the radar anchor voxel, thereby establishing a basic cross-modal spatial correspondence and utilizing geometric alignment loss. Constrain the consistency between visual and audio probability voxels and radar anchoring voxels. Mean of confidence residuals. Used to measure the consistency of confidence levels among radar, vision, and audio for the same candidate voxel, when When the value is small, it indicates that the multimodal confidence levels for the candidate target are relatively consistent; when... A larger value indicates the presence of modal conflict or low-confidence modal interference.

[0108] Once the visual and audio probabilistic voxels have formed relatively stable spatial responses within the neighborhood of the radar anchor point, the system enters the cross-modal geometric consistency enhancement stage, i.e., the mid-training phase. At this point, the generation process of the local voxel calibration mask continues to be trained using the radar anchor voxel as the physical reference, and visual, audio, and radar feature calibrations are performed within the local spatial neighborhood defined by the local voxel calibration mask. The goal of this stage is to enable the network to no longer rely on global unconstrained attention for fusion, but instead learn to prioritize visual and audio voxels that maintain consistency with the radar anchor point in terms of 3D position, sound source direction, and confidence.

[0109] In the later stages of training, the full-level module of the radar expert network is enabled. The full-level module obtains fine-grained radar discrimination features through point cloud topology branches and micro-Doppler branches. This information is then used as category refinement and structural verification data in subsequent calculations. This stage does not alter the spatial anchoring function of the rapid hierarchical model; rather, it leverages the stable radar anchoring voxels already provided by the rapid hierarchical module. Further distinguish drones from birds, clutter, or other low-altitude targets.

[0110] This embodiment implements progressive cross-modal guided fusion and flexible training. Through this flexible training method, the computational model can adapt in advance to the control of the start and stop of each branch of the radar expert network by the subsequent adaptive computing power scheduling decision-maker, avoiding output instability caused by changes in branch state during the inference stage.

[0111] Example 5: This embodiment provides a low-altitude target sensing system, such as Figure 5 As shown, the system includes: a camera group, a multi-channel microphone array, and radar equipment located in the low-altitude defense zone; a data acquisition interface for synchronously sampling image data output by the camera group, audio data output by the multi-channel microphone array, and radar detection data output by the radar equipment; and a processor connected to the data acquisition interface, which executes the steps of the cross-modal solution method for low-altitude target recognition based on a hybrid expert network provided in Embodiments 1, 2, 3, and 4 of this invention. In this embodiment, the radar equipment is not limited to a phased array radar. The camera group includes visible light cameras and infrared cameras.

[0112] The embodiment is merely a specific example and does not indicate that this is the only way to implement the present invention.

[0113] The above description is merely a preferred embodiment of the present invention. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A cross-modal solution method for low-altitude target recognition based on a hybrid expert network, characterized in that, The image data, audio data from the multi-channel microphone array, and radar detection data that are synchronously sampled in the low-altitude defense zone at the current moment are acquired and input into the solution model to obtain the solution result at the current moment. The radar detection data includes radar target detection results; The solution model includes: A visual expert network extracts visual features from the image data; An audio expert network extracts latent audio features based on the audio data; The radar expert network includes a fast hierarchical module, which obtains fast hierarchical features of one or more radar targets based on the radar target detection results, and generates radar anchoring voxels corresponding to each radar target in a three-dimensional voxel grid. The cross-modal adaptive calibration module back-projects the visual feature depth onto the three-dimensional voxel grid to obtain visual probability voxels, projects the audio latent feature probability onto the three-dimensional voxel grid to obtain audio probability voxels, and fuses the visual probability voxels and audio probability voxels that satisfy the geometric consistency condition in the neighborhood of each radar anchor voxel with the radar anchor voxel to generate the three-dimensional fused voxel grid at the current moment. The Transformer spatiotemporal correlation decoder processes the current 3D fused voxel mesh and the trajectory query set from the previous time step through an attention mechanism to obtain the trajectory query set at the current time step. The multi-task solution head obtains the solution result for the current moment based on the trajectory query set at the current moment.

2. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 1, characterized in that, The audio expert network also extracts the sound source direction vector based on the audio data and uses it as the sound source direction vector of the audio probability voxel; The geometric consistency condition is: and ; in, ; Indicates the first The center of the visual probability voxel satisfying the visual candidate condition within the neighborhood of each radar anchor voxel. With the The center of each radar anchor voxel Spatial residual; Indicates the spatial residual threshold; in, ; Indicates the first The sound source direction vector of the audio probability voxel satisfying the audio candidate condition within the neighborhood of each radar anchor voxel. With the Angular residuals between radar target directions Represents the direction vector of the sound source with vector The included angle; This indicates the position of the multi-channel microphone array in the global coordinate system; Represents the cosine function; Indicates the angular residual threshold; This indicates the calculation of the L2 norm.

3. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 2, characterized in that, The visual expert network obtains visual confidence based on the visual features. The audio expert network also obtains the confidence level of the sound source direction estimation based on the sound source direction vector. And obtain rotor harmonic confidence based on the audio data. The radar target detection results include the radar detection confidence level for each radar target. In the cross-modal adaptive calibration module, the step of fusing the visual probability voxels and audio probability voxels that satisfy the geometric consistency condition in the neighborhood of each radar anchor voxel with the radar anchor voxel to generate a three-dimensional fused voxel mesh at the current moment includes: Generate the first Local voxel calibration mask corresponding to the neighborhood of each radar anchor voxel : ; The first The visual probability voxels and audio probability voxels that satisfy the geometric consistency condition within the neighborhood of the first radar anchoring voxel are related to the first... The radar anchoring voxel fusion was performed to obtain the first... One or more three-dimensional fusion voxels corresponding to the neighborhood of each radar anchoring voxel; The number is obtained according to the following formula. Fusion voxel features of three-dimensional fused voxels corresponding to the neighborhood of each radar anchoring voxel. : ; in, This represents the threshold determination function; Indicates the first Radar detection confidence level of a radar target. Indicates will Mapped to multimodal confidence; This represents the attention mechanism. For query, As key, Value; Indicates the first Features of a radar anchoring voxel Indicates the first Features of visual probability voxels that satisfy geometric consistency conditions within the neighborhood of a radar anchoring voxel. Indicates the first Features of audio probability voxels that satisfy geometric consistency conditions within the neighborhood of a radar anchoring voxel. This indicates splicing along the channel dimension. and The obtained splicing features; Set the features of voxels in the three-dimensional voxel mesh that do not belong to the three-dimensional fused voxel to zero to obtain the three-dimensional fused voxel mesh at the current moment.

4. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 3, characterized in that, The radar detection data also includes radar point cloud data and echo data; The radar expert network also includes a full-level module, which includes: The point cloud topology branch includes a preprocessing unit and a point cloud feature extraction network. The preprocessing unit performs ground point filtering, outlier removal and voxelization on the radar point cloud data. The point cloud feature extraction network is used to extract point cloud topology features from the radar point cloud data processed by the preprocessing unit. The micro-Doppler branch includes a short-time Fourier transform unit and a time-frequency feature extraction network. The short-time Fourier transform unit performs a short-time Fourier transform on the echo data to obtain a micro-Doppler time-frequency map, and the time-frequency feature extraction network is used to extract micro-Doppler features from the micro-Doppler time-frequency map. The full fusion unit fuses the point cloud topological features and the micro-Doppler features to obtain radar fine discrimination features; The features of each radar anchor voxel include the rapid hierarchical features of the radar target corresponding to that radar anchor voxel and the radar fine discrimination features.

5. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 4, characterized in that, The solution model also includes a computing power scheduling decision-maker, which determines the control action for the next moment based on the solution result at the current moment, the load state of the method execution entity, and the cross-modal inconsistency. : ; in, This indicates the start / stop flag of the fast-level module. This indicates the start / stop flag for the point cloud topology branch. This indicates the start / stop flag of the micro-Doppler branch. This indicates the voxel grid resolution of the three-dimensional voxel grid. This indicates the radar sampling frequency.

6. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 5, characterized in that, The cross-modal inconsistency is determined by the mean of the spatial residuals, the mean of the angular residuals, and the mean of the confidence residuals of all radar anchor voxels. The mean of the confidence level residuals is: Audio confidence ; Indicates the number of radar anchoring voxels, and is a positive integer; This indicates taking the absolute value.

7. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 1, characterized in that, The image data includes visible light images and infrared thermal imaging images, and the visual expert network includes: The visible light branch, comprising a visible light backbone network and a visible light pyramid network connected in sequence, is used to extract multi-scale first enhanced image features from the visible light image; The infrared branch, comprising an infrared backbone network and an infrared pyramid network connected in sequence, is used to extract multi-scale second enhanced image features from the infrared thermal imaging image; Cross-modal feature bidirectional interaction network, including: The first convolution process performs convolution processing on the second enhanced image features at a preset key scale; The first activation unit maps the features output by the first convolution to a hot target space mask. The first multiplication unit multiplies the first enhanced image feature at a preset key scale with the thermal target spatial mask element by element to obtain the visible light target region feature at the preset key scale. The first CSP fusion module fuses the visible light target region features at a preset key scale and all first enhanced image features at non-preset key scales to obtain visible light fusion features; The second convolutional unit performs convolution processing on the first enhanced image features at a preset key scale; The second activation unit maps the features output by the second convolution to a texture contour mask; The second multiplication unit performs element-wise multiplication of the second enhanced image feature at a preset key scale with the texture contour mask to obtain the infrared contour feature at the preset key scale. The second CSP fusion module fuses infrared contour features at a preset key scale and all second-enhanced image features at non-preset key scales to obtain infrared fusion features. The visible light weight acquisition unit includes a first global average pooling layer and a first fully connected layer connected in sequence, used to map the visible light fusion features into visible light weights; The infrared weight acquisition unit includes a second global average pooling layer and a second fully connected layer connected in sequence, used to map the infrared fusion features into infrared weights; The weighted fusion unit fuses the visible light fusion feature and the infrared fusion feature based on the visible light weight and the infrared weight to obtain the visual feature; In addition, a visual confidence prediction head is used to obtain visual confidence based on the visual features.

8. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 1, characterized in that, The audio data includes Log-Mel spectral characteristics, inter-channel phase difference, and time of arrival difference; The audio expert network includes: The sound source direction determination unit constructs a sound source direction vector based on the inter-channel phase difference and the arrival time difference; The first audio confidence prediction head obtains the source direction estimation confidence based on the source direction vector; The second audio confidence prediction head obtains rotor harmonic confidence based on the Log-Mel spectral features; Spatial branching: The Log-Mel spectral features are processed by a first asymmetric convolution that expands along the time axis to obtain spatial correlation features; The spectral branch uses a second asymmetric convolution that expands along the frequency axis to process the Log-Mel spectral features and obtain spectral correlation features. The temporal modeling unit obtains potential audio features based on the spatial correlation features and the spectral correlation features.

9. The cross-modal solution method for low-altitude target recognition based on a hybrid expert network according to claim 4, characterized in that, The training process of the solution model includes: Initial Training Phase: An initial training model is constructed using the fast hierarchical modules from the visual expert network, audio expert network, and radar expert network within the solution model. This initial training model is trained in batches using the training set, and its network parameters are updated based on geometric alignment loss until the initial stopping condition is met, thus entering the intermediate training phase. The geometric alignment loss is: ,in, Indicates visual weight, Indicates audio weights, Represents multimodal weights. Indicates the first Mean of spatial residuals of all radar anchor voxels in a sample Indicates the first Mean angle residuals of all radar anchor voxels in a sample Indicates the first The mean of the confidence residuals for each sample. The number of samples in each batch, Index of samples in each batch; Mid-training phase: The solution model is trained in batches using the training set, and the radar expert network only enables the fast hierarchical module. The network parameters of the solution model are updated based on the second-stage loss until the mid-training stopping condition is met, and then the training phase begins. The second-stage loss includes geometric alignment loss, classification loss, and 3D localization loss. In the later stage of training: the solution model is trained in batches using the training set, and the fast hierarchical module and the full hierarchical module are enabled simultaneously in the radar expert network. The network parameters of the solution model are updated based on the third-stage loss until the late-stage stopping condition is met. The third-stage loss includes geometric alignment loss, classification loss, 3D localization loss and fine classification loss. The fine classification loss is calculated based on the fine classification result and the true target category in the label of the sample. The fine classification result is obtained by processing the radar fine discrimination features through the auxiliary fine classification head.

10. A low-altitude target sensing system, characterized in that, The system includes: Camera units, multi-channel microphone arrays, and radar equipment located in the low-altitude defense zone; A data acquisition interface is used to synchronously sample image data output by the camera group, audio data output by the multi-channel microphone array, and radar detection data output by the radar device; The processor is connected to the data acquisition interface and executes the cross-modal solution method for low-altitude target recognition based on a hybrid expert network as described in any one of claims 1-9.