A method, device and medium for remote sound monitoring based on a drone

By synchronously collecting and processing data such as vibration signals and meteorological parameters, and dynamically compensating for channel characteristics and microphone positions during UAV flight, high-precision remote sound monitoring was achieved, solving the problems of sound source intensity quantization error and positioning accuracy, and improving the monitoring effect.

CN122384971APending Publication Date: 2026-07-14XIAMEN DAO YITAI ELECTRONIC TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN DAO YITAI ELECTRONIC TECHNOLOGY CO LTD
Filing Date
2026-06-11
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing remote sound monitoring methods suffer from significant errors in sound source intensity quantification due to neglecting real-time weather conditions and aircraft vibrations during drone flight, leading to decreased sound source localization accuracy or even failure.

Method used

By synchronously collecting vibration signals, near-field noise reference signals, real-time meteorological parameters, and motion data, a comprehensive prediction signal of UAV self-noise is synthesized in real time. The propagation channel characteristics and frequency-varying absorption coefficient are dynamically calculated, the microphone position is dynamically estimated, beamforming calculation and spatial spectrum analysis are performed, and remote sound monitoring results are obtained.

Benefits of technology

It improves the intelligibility and fidelity of weak sound signals in complex dynamic environments, realizes high-precision positioning and quantitative inversion of sound sources at long distances, and solves the problems of sound source intensity quantization error and positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122384971A_ABST
    Figure CN122384971A_ABST
Patent Text Reader

Abstract

The application discloses a kind of remote sound monitoring method, equipment and medium based on unmanned aerial vehicle, it is related to sound monitoring technical field, including, based on propagation channel characteristics and frequency variable absorption coefficient, the sound signal of preliminary purification is filtered and handled, obtains the enhanced sound signal after propagation correction, according to motion data, the instantaneous spatial position of each microphone in main acoustic array is dynamically calculated, according to the instantaneous spatial position of each microphone, the enhanced sound signal after propagation correction is beamformed and calculated, obtains spatial spectrum, spectrum peak search and tracking are carried out to spatial spectrum, simultaneously calculate the equivalent continuous A metering sound pressure level of sound source, obtain remote sound monitoring result;The application generates remote sound monitoring result, not only improves the intelligibility and fidelity of weak sound signal in complex dynamic environment, and further realizes the high-precision positioning and quantitative inversion of remote sound source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound monitoring technology, and in particular to a remote sound monitoring method, device and medium based on unmanned aerial vehicles (UAVs). Background Technology

[0002] With increasing environmental awareness, stricter industrial noise control standards, and ever-upgrading security monitoring needs, remote sound source localization and monitoring technology has become a core supporting means in key areas such as environmental acoustic assessment and urban noise control. In recent years, advanced digital signal processing technologies, such as beamforming algorithms, have enabled real-time estimation and tracking of sound source locations. Furthermore, the integration of inertial navigation and global positioning has effectively compensated for the impact of UAV flight attitude changes on array geometry, improving monitoring stability in dynamic environments. It has demonstrated enormous application potential and technological advantages in scenarios such as long-distance detection, complex terrain surveying, and emergency monitoring, providing strong technical support for building a comprehensive and three-dimensional intelligent acoustic monitoring network.

[0003] Nevertheless, existing remote sound monitoring methods still have room for improvement. First, ignoring the dynamic coupling of real-time meteorological conditions and using a fixed theoretical propagation model for inversion leads to huge errors in the quantization of sound source intensity. In addition, the high-frequency vibrations and attitude jitter of UAVs during flight can cause instantaneous elastic deformation of the acoustic array geometry. Previous methods based on the rigid array assumption have resulted in severe sidelobe elevation and main lobe shift due to guide vector mismatch, causing a significant decrease in sound source localization accuracy or even failure. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a remote sound monitoring method based on unmanned aerial vehicles (UAVs) to solve the problems of huge quantization errors in sound source intensity and a significant decrease or even failure in sound source localization accuracy.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a remote sound monitoring method based on an unmanned aerial vehicle (UAV), comprising: During the flight of the UAV, the raw sound signal, vibration signal, near-field noise reference signal, real-time meteorological parameters and motion data of the main acoustic array are collected simultaneously. Based on vibration signals and near-field noise reference signals, a comprehensive prediction signal of UAV self-noise is synthesized in real time. The comprehensive prediction signal is removed from the original sound signal to obtain a preliminarily purified sound signal. Based on the pre-purified acoustic signal and real-time meteorological parameters, the propagation channel characteristics of the sound wave from the potential sound source to the main acoustic array are estimated online, and the frequency-varying absorption coefficient under the current atmospheric conditions is dynamically calculated. Based on the propagation channel characteristics and frequency-varying absorption coefficient, the preliminarily purified acoustic signal is filtered to obtain an enhanced acoustic signal after propagation correction. Based on motion data, the instantaneous spatial position of each microphone in the main acoustic array is dynamically calculated; Based on the instantaneous spatial position of each microphone, beamforming calculations are performed on the propagation-corrected enhanced sound signal to obtain the spatial spectrum; The system performs peak search and tracking on the spatial spectrum, and simultaneously calculates the equivalent continuous A-weighted sound pressure level of the sound source to obtain remote sound monitoring results.

[0007] As a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) described in this invention, the real-time meteorological parameters include temperature, atmospheric pressure, relative humidity, and wind speed. The motion data includes the UAV's three-axis angular velocity, airspeed, propeller speed, and three-axis acceleration data.

[0008] In a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) described in this invention, the step of acquiring the preliminarily purified sound signal specifically includes: Using vibration signals, the conducted noise component of the structure is predicted through a pre-constructed transfer function model; Based on the near-field noise reference signal, the aerodynamic noise component is predicted by combining the aerodynamic noise empirical model. The structural conducted noise component and the aerodynamic noise component are superimposed and fused in the time domain to generate a comprehensive prediction signal; At the front end of the signal processing chain, after performing a synchronization alignment operation on the original sound signal and the comprehensive prediction signal, the prediction noise amplitude at the corresponding time point in the comprehensive prediction signal is removed point by point from the sampling sequence of the original sound signal to obtain the preliminarily purified sound signal.

[0009] As a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) described in this invention, the dynamic calculation of the frequency-varying absorption coefficient under the current atmospheric conditions specifically involves: Based on the pre-purified acoustic signal, the main path delay and attenuation factor of the channel impulse response are solved by the generalized cross-correlation method to obtain the propagation channel characteristics; Based on real-time meteorological parameters such as temperature, relative humidity, and atmospheric pressure, and combined with the preliminarily purified acoustic signal, the frequency-varying absorption coefficient of the sound wave under the current atmospheric conditions is calculated.

[0010] As a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) according to the present invention, the step of acquiring the enhanced sound signal after propagation correction specifically includes: The inverse channel response is calculated based on the main path delay and attenuation factor. The frequency-varying absorption coefficient is converted into a frequency domain amplitude compensation factor according to the distance, and then multiplied with the channel inverse response point by point in the frequency domain to synthesize the complete inverse filter frequency response; After converting the frequency response of the inverse filter into the tap coefficients of the time-domain filter through inverse Fourier transform, a convolution operation is performed on the initially purified acoustic signal to obtain the enhanced acoustic signal after propagation correction.

[0011] As a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) according to the present invention, the step of dynamically calculating the instantaneous spatial position of each microphone in the main acoustic array based on motion data specifically includes: Using the calibration geometric center of the main acoustic array as the reference origin, and combining the triaxial angular velocity and triaxial acceleration data, the attitude is calculated through the direction cosine matrix, and the three-dimensional position and orientation of the reference origin are updated in real time. Calculate the additional displacement vector of each microphone relative to the reference origin; Based on the orientation of the reference origin, the nominal position vector of each microphone relative to the reference origin is rotated to the global coordinate system, and then summed with the three-dimensional position of the reference origin to obtain the rigid body motion position vector of each microphone. The instantaneous spatial position of each microphone in the global coordinate system is obtained by superimposing the rigid body motion position vector of each microphone with the corresponding additional displacement vector.

[0012] As a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) according to the present invention, the acquisition of the spatial spectrum specifically includes: Based on the instantaneous spatial position of all microphones, the relative phase delay of each microphone is calculated, the relative phase delay of each microphone is converted into the corresponding complex exponential phase factor, and arranged in sequence to form the beamformer steering vector.

[0013] Extract multi-channel signal snapshots corresponding to all microphones from the propagation-corrected enhanced sound signal and calculate the signal covariance matrix; Substituting the beamformer steering vector and the signal covariance matrix into the weight calculation formula of the adaptive beamformer, the optimal weight vector is solved under the minimum variance distortionless response criterion. Based on the optimal weight vector and the signal covariance matrix, the output power of the main acoustic array in each scanning direction is calculated to obtain the spatial spectrum.

[0014] As a preferred embodiment of the remote sound monitoring method based on unmanned aerial vehicles (UAVs) according to the present invention, the step of obtaining the remote sound monitoring result specifically includes: Search for local maxima points in the spatial spectrum that exceed a preset threshold, and use the azimuth and elevation angles corresponding to all maxima points as candidate sound source azimuths. Correlate and track the spatial spectrum peaks of the candidate sound source azimuths to eliminate transient interference and obtain stable sound source azimuths. For each stable sound source location, the optimal weight vector corresponding to the stable sound source location is used to perform weighted summation of the multi-channel signal snapshots used to generate the spatial spectrum, and the time-domain output signal is obtained. After performing A-weighted filtering on the time-domain output signal, the root mean square value is calculated, and the sound pressure level conversion formula is applied to the root mean square value to obtain the equivalent continuous A-weighted sound pressure level of the stable sound source location. By combining the azimuth, elevation, and corresponding equivalent continuous A-weighted sound pressure level of each stable sound source, the results of remote sound monitoring are obtained.

[0015] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein the computer program, when executed by the processor, implements any step of the remote sound monitoring method based on a drone as described in the first aspect of the present invention.

[0016] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the remote sound monitoring method based on a drone as described in the first aspect of the present invention.

[0017] The beneficial effects of this invention are as follows: Firstly, by constructing a self-noise prediction model to remove UAV self-noise at the time-domain front end, the problem of excessively low signal-to-noise ratio under strong background interference is solved. Secondly, by combining real-time meteorological parameters and channel estimation technology, signal distortion caused by atmospheric absorption and multipath effects is dynamically compensated, overcoming the quantification error of sound source intensity caused by neglecting environmental coupling in previous methods. Thirdly, by dynamically reconstructing the instantaneous geometric configuration of the microphone array using inertial navigation data, the beamforming guide vector mismatch caused by body vibration and attitude jitter is eliminated. This not only improves the intelligibility and fidelity of weak sound signals in complex dynamic environments, but also achieves high-precision positioning and quantitative inversion of long-distance sound sources. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Fig. 1 This is a flowchart of a remote sound monitoring method based on drones.

[0020] Fig. 2 A flowchart for obtaining a preliminary purified acoustic signal.

[0021] Fig. 3 A flowchart for obtaining remote monitoring results.

[0022] Fig. 4 A flowchart for obtaining enhanced acoustic signals. Detailed Implementation

[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0024] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0025] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0026] Reference Figs. 1-4 This is one embodiment of the present invention, which provides a remote sound monitoring method based on a drone, comprising the following steps: S1. During the flight of the UAV, the original sound signal, vibration signal, near-field noise reference signal, real-time meteorological parameters and motion data of the main acoustic array are collected simultaneously.

[0027] It should be noted that, while controlling the UAV to fly within the monitored airspace, all microphones of the main acoustic array are activated to continuously record raw sound signals using a synchronized sampling clock. Simultaneously, vibration sensors installed on the UAV's motors and key fuselage structures collect vibration signals, and a reference microphone adjacent to the plane of the main acoustic array and propeller collects near-field noise reference signals. An onboard weather station simultaneously collects temperature, atmospheric pressure, relative humidity, and wind speed as real-time meteorological parameters, while an inertial measurement unit simultaneously collects the UAV's three-axis angular velocity, airspeed, propeller speed, and three-axis acceleration data as motion data. The timestamps of all sensors and acquisition channels are strictly synchronized using the same clock source to ensure complete time-domain alignment of the raw sound signals, vibration signals, near-field noise reference signals, real-time meteorological parameters, and motion data.

[0028] S2. Based on vibration signals and near-field noise reference signals, synthesize a comprehensive prediction signal of UAV self-noise in real time, remove the comprehensive prediction signal from the original sound signal, and obtain a preliminarily purified sound signal.

[0029] S2.1 It should be noted that, in the test state where the UAV is in static or controlled flight, a known excitation signal is applied to the fuselage structure, and the output signal of the vibration sensor and the acoustic response signal of each microphone in the main acoustic array are simultaneously acquired. Specifically, the UAV is first controlled to be in a fixed, hovering, or low-speed level flight state to keep the attitude of the UAV and the working state of the propeller stable; pulse excitation, short-time knocking excitation, or short-time frequency sweep excitation are applied to the structural parts of the fuselage near the installation location of the vibration sensor as known excitation signals. After the excitation is applied, the output signal of the vibration sensor and the acoustic response signal of each microphone in the main acoustic array are recorded simultaneously using the same synchronous sampling clock as in step S1, and a continuous sampling data segment corresponding to each excitation is taken as a test record. For each test record, the output signal of the vibration sensor and the acoustic response signal of each microphone are time-synchronized and aligned with the excitation start time as a reference. After time synchronization, a response time period within a preset duration following the excitation start time is extracted from the acoustic response signal of each microphone. This preset duration is set according to the vibration propagation path and decay time of the UAV fuselage structure, for example, 50 milliseconds, ensuring that the extracted response time period covers at least the entire process of structural vibration propagating from the vibration sensor installation location to the corresponding microphone location and gradually decaying. To reduce the impact of random disturbances in a single test, multiple excitations and acquisitions are performed under the same test conditions. For the same microphone, the response time periods extracted from multiple tests (e.g., 3 times) are averaged according to the sampling points to obtain the average response sequence for the microphone. After the above processing is completed for each microphone in the main acoustic array, the time-domain response sequence from the vibration sensor installation location to each microphone is obtained. The time-domain response sequence corresponding to each microphone is stored as a vibration-acoustic transfer function model. The response time period refers to the continuous response signal extracted from the microphone acoustic response signal within a preset duration, starting from the excitation start time.

[0030] The vibration signal acquired in step S1 is used as input and convolved into the vibration-acoustic transfer function model to simulate the physical process of vibration energy being transferred through the UAV fuselage structure to the microphone and radiated as noise. The structurally conducted noise components on each channel of the main acoustic array are obtained, as shown in the formula: ; in, Represents the structural conducted noise component. Indicates the distance from the vibration sensor installation location to the... Time-domain response sequence of the vibration-acoustic transfer function model of a microphone This represents the vibration signal collected in step S1.

[0031] S2.2 It should be noted that in controlled flight testing, the UAV is hovered or leveled at stable flight points. Each state point is characterized by a set of key flight parameters, including at least airspeed and propeller speed. At each state point, the near-field noise reference signal of the reference microphone is synchronously acquired. The near-field noise reference signal is divided into multiple continuous signal frames. A fast Fourier transform is performed on each signal frame to obtain the frequency domain amplitude at each frequency point. After all signal frames have completed the frequency domain transform, the frequency domain amplitudes at the same frequency points are statistically averaged to obtain the average amplitude spectrum under the current stable flight state point. The ratio of the linear sound pressure amplitude corresponding to each frequency point in the average amplitude spectrum to the reference sound pressure is calculated. The result of the ratio calculation is taken as the logarithm to the base 10. The logarithm result is multiplied by the value 20 to obtain the sound pressure level value of the corresponding frequency point. The sound pressure level values ​​corresponding to each frequency point are arranged in frequency order to form the sound pressure level spectrum of the reference microphone under the current stable flight state point. The value 20 represents the scaling factor for logarithmic decibel conversion of the amplitude, used to convert the linear sound pressure amplitude corresponding to each frequency point relative to the reference sound pressure to a sound pressure level value; the reference sound pressure is 20 micropascals, which is a commonly used reference value for calculating sound pressure level in air. Aerodynamic noise response spectra of each microphone channel of the main acoustic array at the same flight state point are collected, and based on the relative installation positions and corresponding spectral differences between the reference microphone and each microphone channel of the main acoustic array, the aerodynamic noise frequency domain transfer relationship from the reference microphone to each microphone channel of the main acoustic array is established.

[0032] A mapping relationship is established between the key flight parameter combination at each flight state point and the sound pressure level spectrum calculated by the reference microphone at the corresponding state. All mapping relationships are stored in a lookup table, which constitutes a simplified empirical model of aerodynamic noise.

[0033] In real-time applications, based on the airspeed and propeller speed of the UAV during flight in the monitored airspace in step S1, the sound pressure level spectrum of the reference microphone under the current flight state is obtained by looking up a table (aerodynamic noise empirical model). When the airspeed and propeller speed do not completely correspond to the flight state points in the lookup table, multiple nearby stable flight state points are selected in the parameter space, and the corresponding sound pressure level spectra are interpolated to obtain the sound pressure level spectrum of the reference microphone under the current flight state.

[0034] For the reference microphone sound pressure level spectrum obtained through lookup or interpolation, the sound pressure level values ​​corresponding to each frequency point in the spectrum are read. These values ​​are then converted into linear sound pressure amplitudes corresponding to each frequency point. These linear sound pressure amplitudes are arranged in frequency order to form a linear amplitude spectrum. A phase spectrum is constructed in the complex domain, extracted from the actual phase spectrum of the near-field noise reference signal acquired in step S1. The linear amplitude spectrum and phase spectrum are combined frequency-by-frequency to form the reference complex spectrum of the reference microphone. Based on the aerodynamic noise frequency domain transfer relationship, a channel-by-channel mapping operation is performed on the complex spectrum of the reference microphone to calculate the aerodynamic noise complex spectrum corresponding to each microphone channel of the main acoustic array. The formula is as follows: ; in, Indicates the main acoustic array The complex spectrum of aerodynamic noise corresponding to each microphone. Indicates from the reference microphone To the main acoustic array aerodynamic noise frequency domain transfer function values ​​of each microphone This represents the complex spectrum of the reference noise of the reference microphone. Represents frequency variables. Indicates the reference microphone.

[0035] Furthermore, using a transfer function estimation method based on cross-spectrum and self-spectrum, the transfer function from the reference microphone to the main acoustic array is obtained. The aerodynamic noise frequency domain transfer function value of each microphone is given by the following formula: ; in, Indicates from the reference microphone To the main acoustic array The aerodynamic noise frequency domain transfer function value of each microphone. Indicates the main acoustic array Each microphone channel signal With reference microphone signal In frequency variable The cross power spectrum at the location, Indicates reference microphone signal The self-power spectrum, Represents a frequency variable.

[0036] The frequency variable is used to unify the reference complex spectrum of the reference microphone, the aerodynamic noise complex spectrum of each microphone in the main acoustic array, the aerodynamic noise frequency domain transfer function, the frequency-varying absorption coefficient, and the frequency domain coordinates of the inverse filter frequency response; the value range of the frequency variable is from 100Hz to 10kHz.

[0037] S2.3 It should be noted that an inverse fast Fourier transform is performed on the complex spectrum of aerodynamic noise corresponding to each microphone channel of the main acoustic array. During the inverse fast Fourier transform, the complex spectrum value corresponding to each frequency point in the complex spectrum of aerodynamic noise of the corresponding channel is read. Each complex spectrum value contains both amplitude and phase information of the corresponding frequency point. According to the inverse transform relationship, complex number operations and accumulation are performed on the complex spectrum value of each frequency point to gradually obtain the sequence value in the time-domain aerodynamic noise sequence of each microphone channel. After all sequence values ​​are determined, the time-domain aerodynamic noise sequence of each microphone channel is obtained. The time-domain aerodynamic noise sequence of each microphone channel is the corresponding aerodynamic noise component.

[0038] S2.4 It should be noted that the structural conducted noise component and the aerodynamic noise component are aligned according to the time axis so that the structural conducted noise component and the aerodynamic noise component correspond one-to-one at the same sampling time. After completing the time alignment, the amplitudes of the structural conducted noise component and the aerodynamic noise component at each sampling time are summed point by point to obtain the superposition result corresponding to each sampling time. The superposition results of all sampling times are arranged in the original time order to form a comprehensive prediction signal covering the current processing period.

[0039] At the front end of the signal processing chain, the original sound signal and the comprehensive prediction signal are first aligned along the time axis to establish a correspondence between them at the same sampling time. After time alignment, the signal amplitude corresponding to each sampling time is read point by point from the sampling sequence of the original sound signal, and the prediction noise amplitude corresponding to the same time point is read from the comprehensive prediction signal. The signal amplitude of the original sound signal and the prediction noise amplitude of the comprehensive prediction signal are subtracted point by point. The subtraction results corresponding to all sampling times are arranged to form a preliminary purified sound signal.

[0040] S3. Based on the pre-purified acoustic signal and real-time meteorological parameters, the propagation channel characteristics of the sound wave from the potential sound source to the main acoustic array are estimated online, and the frequency-varying absorption coefficient under the current atmospheric conditions is dynamically calculated.

[0041] S3.1 It should be noted that the pre-purified acoustic signal is segmented according to a sliding time window; for example, the pre-purified acoustic signal is segmented window by window according to a sliding time window with a length of 10 milliseconds and a step size of 5 milliseconds; for each time window, it is detected whether there are bands in the signal waveform that continuously exceed the preset amplitude threshold, and the duration of the bands is counted; when there are bands that continuously exceed the preset amplitude threshold within a certain time window, and the duration of the bands is less than the preset duration threshold, the acoustic signal corresponding to the time window is determined as a transient acoustic event; the transient acoustic event is used for channel estimation. The process involves: calculating reference acoustic events; extracting received signal segments from each microphone in the main acoustic array corresponding to the time period of the transient acoustic events from the pre-purified acoustic signal; pre-setting a set of candidate time delays for any two microphone received signal segments; calculating the generalized cross-correlation function value corresponding to any two microphones for each candidate time delay using a phase-transform weighted method; arranging each candidate time delay as the abscissa and the corresponding generalized cross-correlation function value as the ordinate after calculating the generalized cross-correlation function value for all candidate time delays to obtain the generalized cross-correlation function curve; determining the peak position of the generalized cross-correlation function curve in the time delay domain and defining the time delay corresponding to the peak position as the relative propagation time delay of the sound wave along the main path between the two microphones; uniformly normalizing the peak amplitude of the generalized cross-correlation function curve (proportional normalization) to obtain the relative attenuation factor of the sound wave propagating along the main path; averaging the relative propagation time delay and relative attenuation factor corresponding to multiple transient acoustic events to obtain the main path time delay and attenuation factor of the channel impulse response.

[0042] Furthermore, the amplitude threshold is set to filter out small-amplitude random noise fluctuations, ensuring that the identified burst sound signals have sufficient intensity; in one embodiment, it can be set to 0.03 Pa, which can effectively filter out low-intensity background noise, random disturbances, and small fluctuations remaining after preliminary purification, reducing false detections.

[0043] The duration threshold is set to filter out continuous sound signals with excessively long durations, ensuring that the identified sound signals belong to short-term, transient sound events. In one embodiment, the duration threshold can be set to 20 milliseconds, which can cover the duration range of most short-term, transient sound events, while also excluding continuous sound signals with longer durations, stable background noise, and long-tailed interference signals, thereby improving the accuracy and stability of transient sound event identification.

[0044] Candidate delay is used to determine the correlation between signal segments received by two microphones under different relative delay conditions within a finite and calculable delay search range, thereby finding the delay position with the highest correlation. In one embodiment, the candidate delay is set from -1ms to +1ms, which can cover the main relative delay range formed by the propagation of sound waves between the microphones of the main acoustic array, and can also narrow the delay search interval, reduce the interference of invalid search points on peak positioning, and improve the accuracy and computational efficiency of relative propagation delay estimation.

[0045] S3.2 It should be noted that, based on the real-time meteorological parameters of temperature, relative humidity, and atmospheric pressure, the frequency-varying absorption coefficient of sound waves under the current atmospheric conditions is calculated using the following formula: ; in, This represents the frequency-varying absorption coefficient of sound waves under current atmospheric conditions. This represents the proportionality coefficient that is weakly correlated with frequency. Indicates atmospheric pressure. Represents frequency variables. Indicates temperature and relative humidity Environmental correction factors working together.

[0046] Furthermore, a proportionality coefficient weakly correlated with frequency is used to perform amplitude calibration on the output of the simplified absorption calculation formula, ensuring that the frequency-varying absorption coefficient remains consistent with actual atmospheric absorption characteristics while preserving the influence trends of frequency, temperature, relative humidity, and atmospheric pressure; an exemplary value is... This allows the frequency-varying absorption coefficient output by the simplified absorption calculation formula to remain on the same order of magnitude as the results of common atmospheric absorption, thus balancing computational simplicity with engineering applicability.

[0047] S4. Based on the propagation channel characteristics and frequency-varying absorption coefficient, the preliminarily purified acoustic signal is filtered to obtain an enhanced acoustic signal after propagation correction.

[0048] S4.1 It should be noted that the channel frequency response is calculated based on the main path delay and attenuation factor, using the following formula: ; in, Indicates the channel frequency response. Indicates the number of primary paths. Represents frequency variables. Indicates the first The attenuation factor of the main path, Indicates the first The delay of the main path, Represents pi (π). Represents the imaginary unit. It represents the base of the natural logarithm.

[0049] Taking the reciprocal of the channel frequency response yields the inverse channel response.

[0050] S4.2 It should be noted that the speed of sound under current atmospheric conditions is calculated based on the temperature from real-time meteorological parameters, using the following formula: ; in, Indicates the speed of sound. The value represents temperature, and 331.4 represents the reference speed of sound in air at 0°C under standard atmospheric pressure, derived from commonly used engineering data on the speed of sound in air. The temperature coefficient representing the change of sound speed with temperature is derived from a commonly used engineering approximation of the dry air sound speed formula after a linear approximation near 0℃.

[0051] The equivalent propagation distance is obtained by multiplying the speed of sound by the main path delay (or the earliest path delay when multiple paths exist); the cumulative absorption attenuation is obtained by multiplying the frequency-varying absorption coefficient by the equivalent propagation distance; the cumulative absorption attenuation is then converted back to the frequency domain amplitude compensation factor using the following formula: ; in, This represents the frequency domain amplitude compensation factor. This indicates the cumulative absorption attenuation. This indicates the base used in logarithmic conversion, specifically base 10 for exponentiation. It is a fixed coefficient corresponding to the conversion of amplitude in decibels.

[0052] The frequency response of the inverse filter is obtained by multiplying the frequency domain amplitude compensation factor and the channel inverse response at each frequency point.

[0053] S4.3 It should be noted that an inverse fast Fourier transform is performed on the frequency response sequence of the inverse filter to calculate the corresponding time-domain sequence. The absolute values ​​of each sequence value in the time-domain sequence are compared to determine the position of the sequence point with the largest absolute value, and the corresponding position is determined as the main energy concentration position. A cyclic shift operation is performed on the time-domain sequence according to the main energy concentration position, so that the main energy concentration position is moved to the beginning of the time-domain sequence. After the cyclic shift is completed, a sequence value of a preset length (such as 512 points) is extracted from the beginning of the shifted time-domain sequence to form the time-domain filter tap coefficient sequence. The preset length is set according to the sampling rate and the expected propagation compensation duration. In one embodiment, when the sampling rate is 48kHz, the preset length can be set to 512 points, which can cover the main propagation correction components and the main energy range of the inverse filter in the UAV remote sound monitoring scenario, thereby preserving the compensation capability for propagation attenuation and time delay distortion more completely.

[0054] The pre-purified acoustic signal is convolved with the tap coefficient sequence of a time-domain filter. In the convolution operation, the time-domain filter tap coefficient sequence is used as a sliding convolution kernel, sliding point-by-point across the sampling sequence of the pre-purified acoustic signal. At each sliding position, the sum of the products of the time-domain filter tap coefficients and the corresponding sampling point of the pre-purified acoustic signal is calculated. This calculation process is performed continuously on the sampling sequence of the pre-purified acoustic signal, generating an output signal sequence; the output signal sequence is the enhanced acoustic signal after propagation correction.

[0055] S5. Based on motion data, dynamically calculate the instantaneous spatial position of each microphone in the main acoustic array.

[0056] S5.1 It should be noted that the calibration geometric center of the main acoustic array is taken as the reference origin, and the three-dimensional position of the reference origin is initialized to (0,0,0). The initial direction cosine matrix of the reference origin is set as the identity matrix. The collected three-axis angular velocity and three-axis acceleration data are read. Based on the three-axis angular velocity, the angular increment vector of the calibration geometric center of the main acoustic array at the current sampling time relative to the previous sampling time is calculated using the following formula: ; in, Represents the angular increment vector. Indicates the angular velocity of the three axes. This indicates the sampling time interval between the current sampling time and the previous sampling time.

[0057] The direction cosine matrix at the reference origin is updated using the angular increment vector, resulting in the updated direction cosine matrix, as shown in the formula: ; in, This represents the direction cosine matrix corresponding to the current sampling time (i.e., the updated direction cosine matrix). This represents the direction cosine matrix at the previous sampling time. Represents the identity matrix. Represents the angular increment vector. The antisymmetric matrix representing the angular increment vector. This indicates the sampling sequence number corresponding to the current sampling time.

[0058] Using the updated direction cosine matrix, the three-axis acceleration data is rotated to the global coordinate system. Then, time integration is performed on the rotated three-axis acceleration data to obtain the velocity change. Time integration of the velocity change yields the three-dimensional displacement of the reference origin in the global coordinate system. The three-dimensional displacement is then summed with the three-dimensional position of the reference origin at the previous sampling time to obtain the updated three-dimensional position of the reference origin. The updated direction cosine matrix represents the orientation of the reference origin.

[0059] S5.2 It should be noted that when the UAV is assembled and in a stationary horizontal state, a body coordinate system is established with the calibration geometric center of the main acoustic array as the origin. Using a three-dimensional coordinate measuring machine or laser tracker, the three-dimensional coordinates of the acoustic center of each microphone in the established body coordinate system are accurately measured. The measured three-dimensional coordinates are recorded as the three-dimensional position vector corresponding to each microphone. The three-dimensional position vector is the nominal installation position vector of each microphone in the body coordinate system. The angular increment vector is cross-multiplied with the nominal installation position vector of each microphone to obtain the additional displacement vector of each microphone relative to the reference origin at the current sampling time.

[0060] S5.3 It should be noted that the direction cosine matrix describing the orientation of the reference origin at the current sampling time is read, and the nominal position vector of each microphone in the body coordinate system is read; for the nominal position vector of each microphone, a matrix multiplication operation with the direction cosine matrix is ​​performed to obtain the rotated position vector of each microphone relative to the reference origin in the global coordinate system; the three-dimensional position vector of the reference origin in the global coordinate system is added to the rotated position vector of each microphone, and the sum is the rigid body motion position vector of each microphone in the global coordinate system.

[0061] The rigid body motion position vector of each microphone in the global coordinate system, along with its corresponding additional displacement vector, is read. In the superposition operation, the three components of the additional displacement vector are algebraically added component-by-component to the corresponding components of the rigid body motion position vector to generate a new three-dimensional position vector. This vector represents the instantaneous spatial position of the microphone in the global coordinate system, taking into account both rigid body motion and the elastic deformation of the fuselage. The superposition operation is performed independently for each microphone in the main acoustic array.

[0062] S6. Based on the instantaneous spatial position of each microphone, perform beamforming calculations on the propagation-corrected enhanced sound signal to obtain the spatial spectrum.

[0063] S6.1 It should be noted that for the propagation-corrected enhanced acoustic signal, a time-domain signal segment of the same length (e.g., 0.1s) is synchronously sampled on each microphone channel at a sampling rate of 48kHz. This time-domain signal segment contains multiple sampling points, forming a multi-channel signal snapshot matrix. Each row of the multi-channel signal snapshot matrix corresponds to one microphone channel, and each column corresponds to one sampling time. For each column vector of the multi-channel signal snapshot matrix, the outer product of the column vector and its conjugate transpose is calculated to obtain an instantaneous covariance matrix. The summation of the instantaneous covariance matrices for all sampling times is then divided by the total number of sampling points in the time-domain signal segment (the product of sampling time and sampling rate) to obtain the average covariance matrix. The average covariance matrix is ​​the signal covariance matrix calculated from the propagation-corrected enhanced acoustic signal.

[0064] Furthermore, the sampling rate is used to determine the discrete sampling density of the propagation-corrected enhanced acoustic signal on the time axis, thereby ensuring that the calculation of the multi-channel signal snapshot matrix and signal covariance matrix has a uniform time resolution. Setting the sampling rate to 48kHz can cover a sufficiently wide effective acoustic bandwidth while taking into account both time resolution and computational complexity, which is beneficial to improving the accuracy of multi-channel acoustic signal analysis and beamforming processing.

[0065] S6.2 It should be noted that the calibration geometric center of the main acoustic array is taken as the reference origin, and the front of the main acoustic array is taken as the zero-degree reference direction; then, based on the expected observation range of the UAV in the monitoring airspace, the starting and ending values ​​of the azimuth angle, as well as the starting and ending values ​​of the pitch angle, are determined; the range of azimuth and pitch angles are discretized according to the preset angle step size (2°) to obtain multiple sequentially arranged scanning azimuths; all scanning azimuths together constitute the scanning azimuth range of the main acoustic array, and the scanning azimuth corresponds to a normalized wavefront direction vector, expressed as: ; in, This represents the wavefront direction vector corresponding to the scanning azimuth. This indicates the azimuth angle corresponding to the scanning direction. This represents the transpose operator. This indicates the elevation angle corresponding to the scanning azimuth.

[0066] The angle step size is set to control the discrete division accuracy of the scanning azimuth. In one embodiment, the angle step size is set to 2, which enables the azimuth and elevation angle ranges to be finely discretely divided, thereby improving the main acoustic array's ability to distinguish different incoming wave directions and reducing the angle error of the sound source azimuth estimation.

[0067] For each microphone's instantaneous spatial position, calculate the dot product with the wavefront direction vector to obtain the projected distance from the wavefront to the microphone; then, quotient the sound velocity with the frequency variable to obtain the sound wavelength. Based on the microphone's projected distance and the sound wavelength, calculate the microphone's absolute phase delay using the following formula: ; in, Indicates absolute phase delay. Indicates the wavelength of the sound wave. Represents pi (π). Indicates the projected distance.

[0068] Randomly select a microphone as the phase reference; remove the absolute phase delay of the reference microphone from the absolute phase delay of each microphone to obtain the relative phase delay of each microphone.

[0069] The relative phase delay of each microphone is converted into a corresponding complex exponential phase factor, as shown in the formula: ; in, This represents the complex exponential phase factor of the microphone. The base of the natural logarithm. Represents the imaginary unit. This indicates the relative phase delay of the microphone.

[0070] The complex exponential phase factors of all microphones are sequentially arranged to form a complex column vector; this column vector is the beamformer steering vector constructed for each scanning azimuth. The beamformer steering vector is calculated sequentially for each scanning azimuth.

[0071] S6.3. It should be noted that the inverse of the signal covariance matrix is ​​calculated. The inverse of the signal covariance matrix is ​​multiplied by the beamformer steering vector for each scanning azimuth to obtain the intermediate vector. The dot product of the conjugate transpose of the beamformer steering vector and the intermediate vector is calculated to obtain the scalar denominator. The quotient of the intermediate vector and the scalar denominator is then performed to obtain the optimal weight vector for each scanning azimuth.

[0072] Read the optimal weight vector and signal covariance matrix corresponding to the scanning azimuth, and calculate the output power of the main acoustic array at each scanning azimuth. The formula is as follows: ; in, This indicates the output power of the main acoustic array in the scanning orientation. This represents the conjugate transpose of the optimal weight vector for the scanning orientation. Represents the signal covariance matrix. This represents the optimal weight vector for the scanning orientation.

[0073] Establish a correlation record between the scanning azimuth and the corresponding main acoustic array output power; repeat the beamformer steering vector construction and output power calculation for the remaining scanning azimuths in the same way to obtain the main acoustic array output power for all scanning azimuths; arrange the main acoustic array output power of all scanning azimuths in sequence according to the order of the scanning azimuths to form a spatial spectrum.

[0074] S7. Perform spectral peak search and tracking on the spatial spectrum, and simultaneously calculate the equivalent continuous A-weighted sound pressure level of the sound source to obtain remote sound monitoring results.

[0075] S7.1 It should be noted that, on the two-dimensional angle grid of the spatial spectrum, for each scanning azimuth, the corresponding spatial spectrum value is compared one by one with the spatial spectrum values ​​of each adjacent scanning azimuth in the neighborhood; when the corresponding spatial spectrum value is greater than the spatial spectrum values ​​of each adjacent scanning azimuth in the neighborhood, the corresponding scanning azimuth is determined to be a local maximum point. The spatial spectrum value of the local maximum point is compared with a preset limit threshold (-10dB), and local maximum points whose spatial spectrum values ​​exceed the limit threshold are retained. For each retained local maximum point, the corresponding azimuth and elevation coordinates are extracted to form candidate sound source azimuths. The azimuth of candidate sound sources is compared moment by moment along multiple consecutive time points. When the azimuth difference between the azimuth of two adjacent time points is less than the preset azimuth threshold (4°) and the pitch difference is less than the preset pitch threshold (5°), the azimuth of the candidate sound sources at two adjacent time points is determined to belong to the same sound source azimuth trajectory. The azimuth of candidate sound sources is repeatedly determined for multiple consecutive time points, and the azimuth trajectory of the sound source that appears continuously for a duration that reaches the preset duration threshold (4 frames) is determined as the stable sound source azimuth.

[0076] Furthermore, the threshold is used to limit whether the spatial spectrum value is significant enough, thereby determining which local maxima can be retained as candidate sound source locations. A threshold of -10dB can filter out weak spurious peaks and background noise peaks in the spatial spectrum, retaining only candidate sound source locations with significant energy, thereby reducing the probability of false detection.

[0077] The azimuth threshold is set to determine whether the difference in azimuth angle between candidate sound source locations at adjacent times is small enough to ensure the continuity of the sound source azimuth trajectory. A azimuth threshold of 4° can preserve the continuous change of the true sound source location while avoiding the misconnection of unrelated peaks with excessively large azimuth angle jumps into the same sound source azimuth trajectory.

[0078] The pitch angle threshold is set to determine whether the difference in pitch angle between candidate sound source locations at adjacent times is small enough, thus determining the stability of the sound source location trajectory in the pitch angle. A pitch angle threshold of 5° can ensure that the change in candidate sound source location in the pitch direction is continuous, thereby improving the reliability of stable sound source location determination.

[0079] The time threshold is used to determine whether the candidate sound source directional trajectory lasts for a sufficient period of time to ensure the stability of the sound source directional and to eliminate transient interference. A time threshold of 4 frames can eliminate transient interference peaks that only appear briefly at individual moments, while retaining the real sound source directional trajectory that exists stably over continuous time.

[0080] S7.2 It should be noted that the optimal weight vector corresponding to the stable sound source location and the multi-channel signal snapshot matrix are read. At each sampling time, the column vector corresponding to that time in the multi-channel signal snapshot matrix is ​​extracted, and the column vector is multiplied by the optimal weight vector to obtain the weighted output scalar for the corresponding sampling time. The dot product operation is repeated for all column vectors of the multi-channel signal snapshot matrix, i.e., all sampling times, to obtain a time-domain scalar sequence. The time-domain scalar sequence is the time-domain output signal for the stable sound source location.

[0081] Based on the start and end times of the sound source trajectory corresponding to the stable sound source location on the time axis, the analysis period of the time-domain output signal corresponding to the stable sound source location is determined. According to the sampling positions corresponding to the start and end times, continuously arranged sampled values ​​within the analysis period corresponding to the stable sound source location are read from the time-domain output signal. The continuously arranged sampled values ​​are truncated and retained according to their original time order to obtain the continuous signal within the analysis period corresponding to the stable sound source location. A-weighted filtering is performed on the continuous signal to ensure that the continuous signal satisfies the A-weighting characteristics in its frequency response, resulting in the A-weighted time-domain output signal. The root mean square value of the A-weighted time-domain output signal is calculated using the following formula: ; in, This represents the root mean square value of the time-domain output signal after A-weighted filtering. This represents the total number of sampling points within the analysis period. This indicates that the time-domain output signal after A-weighting is in the [missing information] stage. The amplitude at each sampling point Substituting the root mean square value into the sound pressure level conversion formula and performing a logarithmic conversion, we obtain the equivalent continuous A-weighted sound pressure level corresponding to the location of a stable sound source, as shown in the formula: ; in, This represents the equivalent continuous A-weighted sound pressure level corresponding to the location of a stable sound source. Represents the reference sound pressure level (2×10). 5 Pa (a uniform, fixed value). This represents the root mean square value of the time-domain output signal after A-weighted filtering.

[0082] For each stable sound source, the corresponding azimuth angle, elevation angle, and equivalent continuous A-weighted sound pressure level are read. The azimuth angle, elevation angle, and equivalent continuous A-weighted sound pressure level are associated according to the correspondence of the same stable sound source and combined into a complete set of sound source monitoring parameters. The combination process is repeated for all stable sound sources to obtain structured result data composed of multiple stable sound source monitoring parameters, i.e., remote sound monitoring results.

[0083] This embodiment also provides a computer device applicable to the remote sound monitoring method based on unmanned aerial vehicles (UAVs), including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the remote sound monitoring method based on UAVs as proposed in the above embodiment.

[0084] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0085] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the remote sound monitoring method based on a drone as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0086] In summary, this invention solves the problem of low signal-to-noise ratio under strong background interference by constructing a self-noise prediction model to remove UAV self-noise at the time-domain front end. Secondly, by combining real-time meteorological parameters and channel estimation technology, it dynamically compensates for signal distortion caused by atmospheric absorption and multipath effects, overcoming the sound source intensity quantization error caused by ignoring environmental coupling in previous methods. Furthermore, by dynamically reconstructing the instantaneous geometric configuration of the microphone array using inertial navigation data, it eliminates beamforming guide vector mismatch caused by body vibration and attitude jitter. This not only improves the intelligibility and fidelity of weak sound signals in complex dynamic environments, but also achieves high-precision positioning and quantitative inversion of long-distance sound sources.

[0087] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A remote sound monitoring method based on unmanned aerial vehicles (UAVs), characterized in that: include, During the flight of the UAV, the raw sound signal, vibration signal, near-field noise reference signal, real-time meteorological parameters and motion data of the main acoustic array are collected simultaneously. Based on vibration signals and near-field noise reference signals, a comprehensive prediction signal of UAV self-noise is synthesized in real time. The comprehensive prediction signal is removed from the original sound signal to obtain a preliminarily purified sound signal. Based on the pre-purified acoustic signal and real-time meteorological parameters, the propagation channel characteristics of the sound wave from the potential sound source to the main acoustic array are estimated online, and the frequency-varying absorption coefficient under the current atmospheric conditions is dynamically calculated. Based on the propagation channel characteristics and frequency-varying absorption coefficient, the preliminarily purified acoustic signal is filtered to obtain an enhanced acoustic signal after propagation correction. Based on motion data, the instantaneous spatial position of each microphone in the main acoustic array is dynamically calculated; Based on the instantaneous spatial position of each microphone, beamforming calculations are performed on the propagation-corrected enhanced sound signal to obtain the spatial spectrum; The system performs peak search and tracking on the spatial spectrum, and simultaneously calculates the equivalent continuous A-weighted sound pressure level of the sound source to obtain remote sound monitoring results.

2. The remote sound monitoring method based on unmanned aerial vehicles as described in claim 1, characterized in that: The real-time meteorological parameters include temperature, atmospheric pressure, relative humidity, and wind speed; The motion data includes the UAV's three-axis angular velocity, airspeed, propeller speed, and three-axis acceleration data.

3. The remote sound monitoring method based on unmanned aerial vehicles as described in claim 2, characterized in that: The acquisition of the preliminarily purified acoustic signal specifically involves: Using vibration signals, the conducted noise component of the structure is predicted through a pre-constructed transfer function model; Based on the near-field noise reference signal, the aerodynamic noise component is predicted by combining the aerodynamic noise empirical model. The structural conducted noise component and the aerodynamic noise component are superimposed and fused in the time domain to generate a comprehensive prediction signal; At the front end of the signal processing chain, after performing a synchronization alignment operation on the original sound signal and the comprehensive prediction signal, the prediction noise amplitude at the corresponding time point in the comprehensive prediction signal is removed point by point from the sampling sequence of the original sound signal to obtain the preliminarily purified sound signal.

4. The remote sound monitoring method based on unmanned aerial vehicles as described in claim 3, characterized in that: The dynamic calculation of the frequency-varying absorption coefficient under current atmospheric conditions is specifically as follows: Based on the pre-purified acoustic signal, the main path delay and attenuation factor of the channel impulse response are solved by the generalized cross-correlation method to obtain the propagation channel characteristics; Based on real-time meteorological parameters such as temperature, relative humidity, and atmospheric pressure, and combined with the preliminarily purified acoustic signal, the frequency-varying absorption coefficient of the sound wave under the current atmospheric conditions is calculated.

5. The remote sound monitoring method based on unmanned aerial vehicles as described in claim 4, characterized in that: The acquisition of the propagation-corrected enhanced acoustic signal specifically involves: The inverse channel response is calculated based on the main path delay and attenuation factor. The frequency-varying absorption coefficient is converted into a frequency domain amplitude compensation factor according to the distance, and then multiplied with the channel inverse response point by point in the frequency domain to synthesize the complete inverse filter frequency response; After converting the frequency response of the inverse filter into the tap coefficients of the time-domain filter through inverse Fourier transform, a convolution operation is performed on the initially purified acoustic signal to obtain the enhanced acoustic signal after propagation correction.

6. The remote sound monitoring method based on unmanned aerial vehicles as described in claim 2, characterized in that: The step of dynamically calculating the instantaneous spatial position of each microphone in the main acoustic array based on motion data is as follows: Using the calibration geometric center of the main acoustic array as the reference origin, and combining the triaxial angular velocity and triaxial acceleration data, the attitude is calculated through the direction cosine matrix, and the three-dimensional position and orientation of the reference origin are updated in real time. Calculate the additional displacement vector of each microphone relative to the reference origin; Based on the orientation of the reference origin, the nominal position vector of each microphone relative to the reference origin is rotated to the global coordinate system, and then summed with the three-dimensional position of the reference origin to obtain the rigid body motion position vector of each microphone. The instantaneous spatial position of each microphone in the global coordinate system is obtained by superimposing the rigid body motion position vector of each microphone with the corresponding additional displacement vector.

7. The remote sound monitoring method based on a drone as described in claim 6, characterized in that: The acquisition of the spatial spectrum specifically involves: Based on the instantaneous spatial position of all microphones, the relative phase delay of each microphone is calculated, the relative phase delay of each microphone is converted into the corresponding complex exponential phase factor, and arranged in sequence to form the beamformer steering vector. Extract multi-channel signal snapshots corresponding to all microphones from the propagation-corrected enhanced sound signal and calculate the signal covariance matrix; Substituting the beamformer steering vector and the signal covariance matrix into the weight calculation formula of the adaptive beamformer, the optimal weight vector is solved under the minimum variance distortionless response criterion. Based on the optimal weight vector and the signal covariance matrix, the output power of the main acoustic array in each scanning direction is calculated to obtain the spatial spectrum.

8. The remote sound monitoring method based on unmanned aerial vehicles as described in claim 7, characterized in that: The specific steps for obtaining the remote sound monitoring results are as follows: Search for local maxima points in the spatial spectrum that exceed a preset threshold, and use the azimuth and elevation angles corresponding to all maxima points as candidate sound source azimuths. Correlate and track the spatial spectrum peaks of the candidate sound source azimuths to eliminate transient interference and obtain stable sound source azimuths. For each stable sound source location, the optimal weight vector corresponding to the stable sound source location is used to perform weighted summation of the multi-channel signal snapshots used to generate the spatial spectrum, and the time-domain output signal is obtained. After performing A-weighted filtering on the time-domain output signal, the root mean square value is calculated, and the sound pressure level conversion formula is applied to the root mean square value to obtain the equivalent continuous A-weighted sound pressure level of the stable sound source location. By combining the azimuth, elevation, and corresponding equivalent continuous A-weighted sound pressure level of each stable sound source, the results of remote sound monitoring are obtained.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the remote sound monitoring method based on a drone as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the remote sound monitoring method based on unmanned aerial vehicles as described in any one of claims 1 to 8.