Guidance interception system and method of mixed acoustic visual fusion and neural network
By using a hybrid acoustic-visual fusion and neural network-based guidance and interception system, the problems of target loss and self-noise in anti-drone systems when facing small, high-speed fixed-wing drones have been solved. This system enables continuous tracking and efficient interception of dynamic targets, improving the interception success rate and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 四川腾盾科技有限公司
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing anti-drone systems are prone to target loss and tracking interruption when facing small, high-speed, and highly maneuverable fixed-wing drones. Traditional guidance laws fail to fully utilize the nonlinear flight dynamics of rotary-wing drones, acoustic sensors are subject to self-noise interference, and there is a lack of effective target re-acquisition mechanisms, resulting in a low interception success rate.
The guidance and interception system employs a hybrid acoustic-visual fusion and neural network. It utilizes a dual-position acoustic sensor array and a visible light camera, and achieves acoustic re-acquisition after visual loss through self-noise suppression and deep learning algorithms. Combined with extended Kalman filtering and deep reinforcement learning, it generates flight control commands to improve tracking stability and interception success rate.
It enables uninterrupted tracking and interception of dynamic fixed-wing targets in complex environments, improves the interception success rate and robustness, overcomes the self-noise interference of rotary-wing UAVs, and utilizes asymmetry to build tactical advantages.
Smart Images

Figure CN122015580A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of technology, and in particular to a guidance and interception system and method that combines acoustic and visual fusion with neural networks. Background Technology
[0002] Existing counter-drone technologies, especially in scenarios involving small, high-speed, and highly maneuverable fixed-wing drones, have the following limitations: 1. Traditional interception systems primarily rely on visible light or radar sensors. When a target (such as a high-speed fixed-wing UAV) briefly leaves the field of view (FOV) of the main sensor (such as a camera) due to rapid maneuvering or environmental obstruction, the system is prone to losing the target, leading to tracking interruption and interception mission failure. Existing systems lack an effective mechanism for rapid target re-acquisition using auxiliary information after the main sensor fails, resulting in insufficient tracking stability.
[0003] 2. Traditional guidance laws, such as proportional navigation (PN) and its variants, while effective in many scenarios, are typically based on simplified kinematic models and fail to fully utilize the unique nonlinear flight dynamics of interception platforms (especially rotary-wing UAVs). During the terminal interception phase, rotary-wing UAVs possess high maneuverability, including hovering and lateral translation, which traditional guidance laws cannot fully realize, leading to a decrease in interception success rate when facing highly maneuverable targets.
[0004] 3. Existing technologies deploy acoustic sensors on UAV platforms for passive detection, but face a severe "ego-noise" problem. The interceptor's own rotor generates strong, non-stationary broadband noise, which can easily drown out weak acoustic signals from the target, making it difficult for airborne acoustic sensors to work effectively.
[0005] 4. Rotary-wing UAVs and fixed-wing UAVs exhibit significant asymmetry in flight mechanics (speed, maneuverability, endurance) and acoustic characteristics. Most existing anti-UAV systems are platform-agnostic and do not specifically target the scenario of "rotary-wing intercepting fixed-wing" to leverage these asymmetries to build tactical advantages. Summary of the Invention
[0006] To address the aforementioned issues, this invention provides a guidance and interception system and method that integrates hybrid acoustic-visual fusion with neural networks. When the visual sensor loses control of the target due to its high-speed maneuvering, the system can utilize passive acoustic localization of the target fixed-wing UAV noise to actively guide the camera gimbal for rapid redirection and target reacquisition, thereby achieving uninterrupted closed-loop tracking and interception. This improves the interception success rate and tracking robustness of dynamic fixed-wing targets in asymmetric combat scenarios.
[0007] This invention provides a guidance and interception system that combines acoustic and visual fusion with neural networks for rotary-wing UAVs to intercept fixed-wing UAVs. The specific technical solution is as follows: The system includes a hybrid sensing unit, a self-noise suppression unit, a target state estimation unit, a gimbal redirection unit, and a guidance unit; The hybrid sensing unit includes at least one camera and at least one dual-position acoustic sensor array. The dual-position acoustic sensor array includes a spatially separated first microphone subarray and a second microphone subarray. The camera and the dual-position acoustic sensor array synchronously acquire the visual image sequence and acoustic signal sequence of the target, respectively. The self-noise suppression unit receives the acoustic signal sequence and generates a time-frequency mask in real time, which is multiplied point by point with the acoustic signal sequence to output a noise-reduced acoustic signal. The target state estimation unit receives a visual image sequence and a denoised acoustic signal for processing, obtains the visual and acoustic azimuth and pitch angles respectively, and fuses them to obtain the target's position and velocity state vectors; the target state estimation unit also includes updating the state vectors according to the acoustic azimuth and pitch angles after visual loss, and synchronously outputting the predicted azimuth and pitch angles. The gimbal redirection unit is communicatively connected to the target state estimation unit. After receiving the visual loss signal, it drives the gimbal of the visible light camera to rotate to the predicted target position according to the predicted azimuth and pitch angles. The guidance unit stores an Actor-Critic network model based on deep reinforcement learning. It takes the state vector, line-of-sight angular rate, the motion state of the rotorcraft UAV, and the tracking mode flag representing the current tracking mode as inputs to the state space and outputs the low-level flight control commands of the rotorcraft UAV.
[0008] By employing a dual-microphone array design, the computational baseline of TDOA is maximized, improving the estimation accuracy of DOA for sound sources. Simultaneously, a self-noise suppression unit accurately identifies and filters out rotor noise from the acoustic signal before it enters the DOA estimation process, effectively extracting the acoustic features of distant fixed-wing targets. In tracking mode, where the target is within the camera's field of view, EKF fuses visual and acoustic data to generate a more robust and accurate target state estimate than that of a single sensor. When the system detects visual target loss, it enters pure acoustic mode, utilizing the target DOA vector continuously provided by the acoustic sensor to generate commands to control the camera gimbal to turn towards the predicted target position. Through an active "redirect-search" process, the camera can quickly relock onto the target, transforming sensor failure from "mission failure" to "degraded operation mode," significantly improving tracking stability.
[0009] Furthermore, the first microphone subarray has two microphone devices with a spacing of the first baseline length, and the second microphone subarray has two microphone devices with a spacing of the second baseline length; The length of the first baseline is greater than the length of the second baseline.
[0010] Furthermore, the first baseline has a length of 300 mm, and the second baseline has a length of 150 mm; The first microphone subarray and the second microphone subarray are located at a distance of more than 400 mm from the rotor plane.
[0011] This invention also provides a guidance and interception method that combines acoustic and visual fusion with neural networks, comprising: S1: Acquire visual image data, search for and lock onto the target drone; S2: After locking onto the target drone, enter the normal tracking mode and execute step S3. At the same time, obtain the visual tracking confidence level in real time and compare it with the set threshold to determine the current target tracking status. If the target vision is not lost, maintain the current normal tracking mode; if the target vision is lost, switch to pure acoustic mode and execute step S4. S3: Simultaneously acquire the target's visual image sequence and acoustic signal sequence, analyze the visual and acoustic azimuth and pitch angles respectively, and fuse them through an extended Kalman filter to obtain the target's position and velocity state vector; S4: Acquire acoustic signal sequence, parse to obtain acoustic azimuth and pitch angles, update target position and velocity state vectors, and execute step S5; simultaneously drive gimbal rotation based on current azimuth and pitch angles, and perform acoustic-driven visual recapture. If recapture is completed, return to step S2 and enter normal tracking mode. If not captured, re-execute step S4 to drive gimbal rotation. S5: Using an actor-critic network based on deep reinforcement learning, the system takes the state vector, line-of-sight angular rate, the rotorcraft's own motion state, and tracking mode flags as inputs and outputs low-level flight control commands.
[0012] Furthermore, in step S3, the target is identified and located in the visual image sequence by the target detection algorithm, and the pixel coordinates of the target are output. After distortion correction and normalization of the pixel coordinates, the coordinates are transformed sequentially from the camera coordinate system, the gimbal coordinate system, the body coordinate system to the navigation coordinate system, and finally the azimuth and pitch angles of the target are calculated.
[0013] Furthermore, in steps S3 and S4, before parsing the acquired acoustic signal sequence, noise reduction processing is performed using a trained deep convolutional autoencoder based on the U-Net architecture.
[0014] The network model is trained offline on self-noise data of the interceptor in various flight attitudes. When running in real time, it can generate a time-frequency mask to accurately identify and filter out its own rotor noise before the signal enters the DOA estimation process, thereby effectively extracting the acoustic features of distant fixed-wing targets.
[0015] Furthermore, the time delay difference between each microphone pair is calculated using the GCC-PHAT algorithm on the noise-suppressed acoustic signal. Based on the time delay difference and the geometric layout of the microphone array, the direction of arrival of the sound source is calculated by geometric analysis or least squares optimization, and the acoustic azimuth and pitch angles are output after numerical robustness processing.
[0016] Furthermore, in step S5, the generation of the flight control commands is as follows: S501: Combines the target's position and velocity state vector, line-of-sight angular rate, the rotorcraft's own motion state, and tracking mode flags into an input state vector, and performs normalization processing. S502: Input the normalized state vector into the pre-trained Actor network and output the original action vector; S503: After performing safety trimming and smoothing filtering on the original motion vectors, they are mapped to underlying flight control commands.
[0017] The beneficial effects of this invention are as follows: This invention utilizes a rotary-wing UAV as an interception platform, integrating a hybrid sensing unit consisting of a visible light camera and a dual-position acoustic sensor array, along with a self-noise suppression unit. Through an advanced deep learning self-noise suppression algorithm, it overcomes the interference of the rotary-wing platform's inherent noise. When the target is not lost, target tracking is performed through sound-visual data fusion. When the target is visually lost, an acoustic-driven visual re-acquisition mechanism is used to re-acquire the target using acoustic signals. This solves the problem of direct tracking and interception failure due to dynamic target loss, ensuring the UAV's continuous tracking capability in complex combat environments. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the system framework structure of the present invention.
[0019] Figure 2 This is a schematic diagram of the microphone array layout structure of the present invention.
[0020] Figure 3 This is a schematic diagram of the guidance and interception method of the present invention.
[0021] Figure 4 This is a schematic diagram of the self-noise suppression process of the present invention.
[0022] Figure 5 This is a schematic diagram of the recapture timing of the present invention. Detailed Implementation
[0023] The technical solutions in the embodiments of the present invention are clearly and completely described in the following description. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0024] In the description of the embodiments of the present invention, it should be noted that the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is conventionally placed during use, or the orientation or positional relationship in which those skilled in the art conventionally understand it during use. This is only for the convenience of describing the present invention and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention. Furthermore, the terms "first" and "second" are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] In the description of the embodiments of the present invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0026] Example 1 Embodiment 1 of the present invention discloses a guidance and interception system based on hybrid acoustic-visual fusion and neural networks, used for rotary-wing UAVs to intercept fixed-wing UAVs, such as... Figure 1 As shown, it includes a hybrid sensing unit, a self-noise suppression unit, a target state estimation unit, a gimbal redirection unit, and a guidance unit.
[0027] The hybrid sensing unit includes at least one high frame rate visible light camera and at least one dual-position acoustic sensor array. The dual-position acoustic sensor array includes a spatially separated first microphone subarray and a second microphone subarray. The camera and the dual-position acoustic sensor array synchronously acquire the visual image sequence and acoustic signal sequence of the target, respectively. As a preferred embodiment, such as Figure 2As shown, the first microphone subarray has two microphone devices with a spacing of the first baseline length, and the second microphone subarray has two microphone devices with a spacing of the second baseline length. The length of the first baseline is greater than the length of the second baseline.
[0028] In a preferred embodiment, the first baseline length is 300 mm and the second baseline length is 150 mm; The first microphone subarray and the second microphone subarray are located at a distance of more than 400 mm from the rotor plane.
[0029] The self-noise suppression unit receives the acoustic signal sequence and generates a time-frequency mask in real time, which is multiplied point by point with the acoustic signal sequence to filter out the self-noise generated by the rotor and motor of the rotary-wing UAV and output the noise-reduced acoustic signal. Specifically, the self-noise suppression unit employs a pre-trained deep convolutional autoencoder based on the U-Net architecture, and the training process is as follows: Collect multi-channel microphone data of the interceptor aircraft at different throttle openings, flight speeds and attitudes in an anechoic chamber or open field to build a comprehensive self-noise database. The learning process will take a noisy time-frequency graph as input and a set binary or ratio mask as the target for supervised learning.
[0030] The target state estimation unit receives a visual image sequence and a denoised acoustic signal for processing, obtains the visual and acoustic azimuth and pitch angles respectively, and fuses them to obtain the target's position and velocity state vectors; the target state estimation unit also includes updating the state vectors according to the acoustic azimuth and pitch angles after visual loss, and synchronously outputting the predicted azimuth and pitch angles. Specifically, the target state estimation unit identifies and locates the target in the visual image sequence using a lightweight target detection algorithm (such as YOLOv3), outputs the pixel coordinates of the target, and calculates the azimuth and pitch angles of the target by combining camera intrinsic parameters and attitude information; the target state estimation unit calculates the time difference of arrival (TDOA) of the denoised acoustic signal sequence using the generalized cross-correlation phase transform algorithm (GCC-PHAT), and calculates the direction of arrival (DOA, the azimuth and pitch angles of the sound source signal relative to the sensor array) based on the TDOA value and the geometric layout of the microphone array.
[0031] The gimbal redirection unit is communicatively connected to the target state estimation unit. After receiving the visual loss signal, it drives the gimbal of the visible light camera to rotate to the predicted target position according to the predicted azimuth and pitch angles.
[0032] The guidance unit stores an Actor-Critic network model based on deep reinforcement learning. It takes the state vector, line-of-sight angular rate, the motion state of the rotorcraft UAV itself, and the tracking mode flag representing the current tracking mode as input to the state space and outputs the low-level flight control commands of the rotorcraft UAV. By incorporating the working state of the hybrid sensing unit into the decision input of the guidance system, deep coupling and collaborative evolution of sensing and control are achieved, thereby improving the intelligence level of the system. In this embodiment, the model also includes a reward function to guide the learning process, including a large positive reward for successful interception, a shaping reward that is inversely proportional to the miss distance and the rate of change in line of sight, and an additional reward for keeping the target within the camera's field of view.
[0033] Example 2 Embodiment 2 of the present invention discloses a guidance and interception method that combines acoustic-visual fusion with neural networks, such as... Figure 3 As shown, the details are as follows: S1: Acquire visual image data, search for and lock onto the target drone; S2: After locking onto the target drone, enter normal tracking mode and execute step S3. Simultaneously, acquire the visual tracking confidence level in real time and compare it with a set threshold to determine the current target tracking status. If the target's vision is not lost, maintain the current normal tracking mode; if the target's vision is lost, switch to pure acoustic mode and execute step S4. Figure 5 As shown; S3: Simultaneously acquire the target's visual image sequence and acoustic signal sequence, analyze the visual and acoustic azimuth and pitch angles respectively, and fuse them through an extended Kalman filter to obtain the target's position and velocity state vector; In a preferred embodiment, a target drone is identified and located in a visual image sequence using a target detection algorithm, and the pixel coordinates (u,v) of the bounding box center are output. After distortion correction and normalization of pixel coordinates, the coordinates are transformed sequentially from camera coordinate system, gimbal coordinate system, body coordinate system to navigation coordinate system, and finally the azimuth and pitch angles of the target are calculated. Specifically, if a wide-angle lens is used, the pixel coordinates are first corrected for distortion:
[0034] in, , The radial distortion coefficient is... , The tangential distortion coefficients are all obtained through camera calibration.
[0035] Normalized image coordinates are calculated as follows:
[0036] in, Let these be the coordinates of the camera's principal point. The focal length is the camera's internal parameter.
[0037] Direction vector in camera coordinate system:
[0038] Transform to gimbal coordinate system:
[0039] in, The fixed-mount rotation matrix of the camera relative to the gimbal (obtained through calibration).
[0040] Transform to body coordinate system
[0041] in, The rotation matrix of the gimbal relative to the machine body is calculated based on the joint angles fed back in real time by the gimbal encoder:
[0042] in, For gimbal yaw angle, The tilt angle of the gimbal.
[0043] Transform to the navigation coordinate system (NED)
[0044] in, The rotation matrix from the airframe to the navigation system, and the airframe attitude angles measured by the flight control IMU. Build:
[0045] The azimuth and elevation angles are calculated as follows: set up:
[0046] but:
[0047]
[0048] in, The azimuth angle is measured visually. The pitch angle is measured visually.
[0049] S4: Acquire acoustic signal sequence, parse to obtain acoustic azimuth and pitch angles, update target position and velocity state vectors, and execute step S5; simultaneously drive gimbal rotation based on current azimuth and pitch angles, and perform acoustic-driven visual recapture. If recapture is completed, return to step S2 and enter normal tracking mode; if not captured, re-execute step S4 to drive gimbal rotation.
[0050] As a preferred embodiment, such as Figure 4 As shown, before parsing the acquired acoustic signal sequence, noise reduction processing is performed using a trained deep convolutional autoencoder based on the U-Net architecture.
[0051] The deep convolutional autoencoder outputs a time-frequency mask which is multiplied by the original input signal to suppress noise.
[0052] Specifically, the deep convolutional autoencoder is trained as follows: by collecting multi-channel microphone data of the interceptor at different throttle openings, flight speeds and attitudes in an anechoic chamber or open field, a comprehensive self-noise database is constructed; supervised learning is performed with the time-frequency map of the noisy signal as input and an ideal binary or ratio mask as the target.
[0053] For the noise-suppressed acoustic signal, the time delay difference (TDOA) between each microphone pair is calculated using the GCC-PHAT algorithm, as follows: Signals collected by microphone pair (i,j) and Calculate the generalized cross-correlation function:
[0054] in, and They are respectively and Fourier transform, Indicates conjugate.
[0055] The time delay difference is estimated as the peak position of the cross-correlation function:
[0056] Based on the TDOA value and the geometric layout of the microphone array, the direction of arrival (DOA) of the sound source is calculated using geometric analytical methods or least squares optimization methods, and the acoustic azimuth and pitch angles are output, as follows: The microphone array is selected according to actual needs, and no specific limitation is made here; In this embodiment, a four-element cross array is used as an example for explanation; Let the position of the i-th microphone in the N-element microphone array be... Taking the first microphone as a reference, the relationship between TDOA and the incident direction is as follows: , Where c is the speed of sound The unit vector in the direction of the sound source:
[0057] Using geometric analysis, the microphone positions for a four-element cross array are:
[0058] The relationship between TDOA and the direction angle is:
[0059] Azimuth calculation:
[0060] Pitch angle calculation:
[0061] Due to the influence of noise, the calculated The value may exceed [ The range of 1,1] causes abnormal calculation of the arccos function, therefore, amplitude limiting is required:
[0062]
[0063]
[0064] S5: Using an actor-critic network based on deep reinforcement learning, the system takes the state vector, line-of-sight angular rate, the rotorcraft's own motion state, and the tracking mode flag as inputs and outputs the underlying flight control commands. The generation of the flight control commands is as follows: S501: Combines the target's position and velocity state vector, line-of-sight angular rate, the rotorcraft's own motion state, and tracking mode flags into an input state vector, and performs normalization processing. The state vector Defined as:
[0065] in, Represents the target's relative position vector The normalization process is: divide by the maximum detection distance Rmax; The target relative velocity vector is normalized by dividing by the maximum relative velocity Vmax. , The line-of-sight angular rate (azimuth, pitch) is normalized by dividing by π. , , The attitude angles (roll, pitch, yaw) of a rotary-wing UAV are represented by the normalization method: divide by π. , , The normalization method for representing the body's angular velocity is: divide by the maximum angular velocity. ; The tracking mode flag is normalized as follows: m∈{0,1}, 0=pure acoustic mode, 1=visual / fusion mode; By introducing a tracking mode flag m, the guidance network can perceive the current sensor operating status, achieving deep coupling between perception and control.
[0066] S502: Input the normalized state vector into the pre-trained Actor network and output the original action vector; Specifically: The Actor network adopts a fully connected neural network structure, and its forward propagation process is as follows:
[0067] Output the original action vector:
[0068] in, , , This is a three-axis acceleration command. This is the yaw rate command.
[0069] The Actor network is trained offline in a simulation environment using either the Deep Deterministic Policy Gradient (DDPG) algorithm or the Soft Actor-Critic (SAC) algorithm.
[0070] In this embodiment, a reward function is also included during model training:
[0071] Among them, the miss penalty drives the drone to approach the target:
[0072] Energy consumption penalty to prevent excessive maneuvering:
[0073] The rewards related to this mode are as follows: When m=1 Positive rewards are given for aggressive maneuvers; when m=0, Violent maneuvers are penalized, while smooth flight is rewarded; Through training with the aforementioned reward function, the network learns to adaptively adjust its maneuvering strategy based on the state of the perception system, achieving deep coupling and co-evolution of perception and control.
[0074] S503: After performing safety trimming and smoothing filtering on the original motion vectors, they are mapped to underlying flight control commands; Specifically, the safety cutting procedure is as follows:
[0075] in, and This refers to the preset action boundaries.
[0076] Mode switching smoothing filter, as follows:
[0077] The filter coefficient α is dynamically adjusted according to the mode: when m = 1, α is larger (faster response), and when m = 0, α is smaller (smoother).
[0078] The smoothed motion vectors are mapped to the low-level control commands of the rotary-wing UAV:
[0079] in, For thrust command, , For the desired attitude angle, The yaw rate command. Let g be the mass of the drone, and g be the acceleration due to gravity.
[0080] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.
Claims
1. A guidance and interception system combining hybrid acoustic-visual fusion and neural networks, characterized in that, It includes a hybrid sensing unit, a self-noise suppression unit, a target state estimation unit, a gimbal redirection unit, and a guidance unit; The hybrid sensing unit includes at least one camera and at least one dual-position acoustic sensor array. The dual-position acoustic sensor array includes a spatially separated first microphone subarray and a second microphone subarray. The camera and the dual-position acoustic sensor array synchronously acquire the visual image sequence and acoustic signal sequence of the target, respectively. The self-noise suppression unit receives the acoustic signal sequence and generates a time-frequency mask in real time, which is multiplied point by point with the acoustic signal sequence to output a noise-reduced acoustic signal. The target state estimation unit receives a visual image sequence and a denoised acoustic signal for processing, obtains the visual and acoustic azimuth and pitch angles respectively, and fuses them to obtain the target's position and velocity state vectors; the target state estimation unit also includes updating the state vectors according to the acoustic azimuth and pitch angles after visual loss, and synchronously outputting the predicted azimuth and pitch angles. The gimbal redirection unit is communicatively connected to the target state estimation unit. After receiving the visual loss signal, it drives the gimbal of the visible light camera to rotate to the predicted target position according to the predicted azimuth and pitch angles. The guidance unit stores an Actor-Critic network model based on deep reinforcement learning. It takes the state vector, line-of-sight angular rate, the motion state of the rotorcraft UAV, and the tracking mode flag representing the current tracking mode as inputs to the state space and outputs the low-level flight control commands of the rotorcraft UAV.
2. The guided interception system according to claim 1, characterized in that, The first microphone subarray has two microphone devices with a spacing of the first baseline length, and the second microphone subarray has two microphone devices with a spacing of the second baseline length. The length of the first baseline is greater than the length of the second baseline.
3. The guided interception system according to claim 2, characterized in that, The first baseline has a length of 300 mm, and the second baseline has a length of 150 mm; The first microphone subarray and the second microphone subarray are located at a distance of more than 400 mm from the rotor plane.
4. A guidance and interception method combining hybrid acoustic-visual fusion and neural networks, characterized in that, Based on the guided interception system according to any one of claims 1-3, the method includes: S1: Acquire visual image data, search for and lock onto the target drone; S2: After locking onto the target drone, enter the normal tracking mode and execute step S3. At the same time, obtain the visual tracking confidence level in real time and compare it with the set threshold to determine the current target tracking status. If the target vision is not lost, maintain the current normal tracking mode; if the target vision is lost, switch to pure acoustic mode and execute step S4. S3: Simultaneously acquire the target's visual image sequence and acoustic signal sequence, analyze the visual and acoustic azimuth and pitch angles respectively, and fuse them through an extended Kalman filter to obtain the target's position and velocity state vector; S4: Acquire acoustic signal sequence, parse to obtain acoustic azimuth and pitch angles, update target position and velocity state vectors, and execute step S5; simultaneously drive gimbal rotation based on current azimuth and pitch angles, and perform acoustic-driven visual recapture. If recapture is completed, return to step S2 and enter normal tracking mode. If not captured, re-execute step S4 to drive gimbal rotation. S5: Using an actor-critic network based on deep reinforcement learning, the system takes the state vector, line-of-sight angular rate, the rotorcraft's own motion state, and tracking mode flags as inputs and outputs low-level flight control commands.
5. The guided interception method according to claim 4, characterized in that, In step S3, the target is identified and located in the visual image sequence by the target detection algorithm, and the pixel coordinates of the target are output. After the pixel coordinates are normalized, they are transformed sequentially from the camera coordinate system, the gimbal coordinate system, the body coordinate system to the navigation coordinate system, and finally the azimuth and pitch angles of the target are calculated.
6. The guided interception method according to claim 5, characterized in that, If a wide-angle lens is used, distortion correction processing is also performed before normalizing the pixel coordinates.
7. The guided interception method according to claim 4, characterized in that, In steps S3 and S4, before parsing the acquired acoustic signal sequence, noise reduction processing is performed using a trained deep convolutional autoencoder based on the U-Net architecture.
8. The guided interception method according to claim 7, characterized in that, For the noise-suppressed acoustic signal, the time delay difference between each microphone pair is calculated using the GCC-PHAT algorithm; Based on the time delay difference and the geometric layout of the microphone array, the direction of arrival of the sound source is calculated, and after numerical robustness processing, the acoustic azimuth and pitch angles are output.
9. The guided interception method according to claim 4, characterized in that, In step S5, the generation of the flight control commands is as follows: S501: Combines the target's position and velocity state vector, line-of-sight angular rate, the rotorcraft's own motion state, and tracking mode flags into an input state vector, and performs normalization processing. S502: Input the normalized state vector into the pre-trained Actor network and output the original action vector; S503: After performing safety trimming and smoothing filtering on the original motion vectors, they are mapped to underlying flight control commands.