A method and system for modeling environmental noise of a low altitude aircraft

By transforming sound source directivity data and 3D environmental information in spherical coordinates, and combining multimodal latent feature fusion and diffusion models, the problem of single sound source processing and neglect of 3D influencing factors in existing low-altitude aircraft noise modeling methods is solved, achieving high-precision noise modeling and generalization capabilities.

CN120930268BActive Publication Date: 2025-12-26SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511447589.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-12-26
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing deep learning-based environmental noise modeling methods for low-altitude aircraft are limited in their source processing and ignore three-dimensional influencing factors. Furthermore, U-Net struggles to fully capture the coupling relationship between noise and multiple factors such as terrain and atmosphere, resulting in poor noise modeling accuracy.

Method used

A spherical coordinate system is constructed for the environment under test. The sound source directivity data of the low-altitude aircraft at different positions in the spherical coordinate system are converted into a sound source directivity tensor. Combined with three-dimensional environmental terrain information and building geometric information, a noise intensity map is obtained through multimodal latent feature fusion and diffusion model.

Benefits of technology

It achieves generalization for any complex aircraft, accurately reflects the propagation characteristics of noise in three-dimensional space, significantly improves the accuracy and efficiency of noise modeling, and can better meet the complex needs of urban low-altitude aircraft noise management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930268B_ABST
    Figure CN120930268B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of noise modeling, and particularly relates to a low-altitude aircraft environmental noise modeling method and system. The sound source directivity data of a low-altitude aircraft at different positions in a spherical coordinate system is re-projected to a hemisphere, and then converted into a sound source directivity tensor. The environmental terrain information and building geometry information are converted into a terrain layout tensor. The high-level features of the sound source directivity tensor, the high-level features of the terrain layout tensor, the frequency features of the low-altitude aircraft at the current moment, and the flight height features of the low-altitude aircraft at the current moment are fused to obtain multi-modal latent features at the current moment. The multi-modal latent features at the current moment are input into a diffusion model to output a noise intensity map of the low-altitude aircraft at the current moment. The present application can accurately capture and fuse the complex multi-modal conditions of the low-altitude aircraft noise source characteristics, terrain layout, and flight height, effectively improving the efficiency, accuracy, and generalization ability of noise modeling.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of noise modeling, in particular to a low-altitude aircraft environmental noise modeling method and system. BACKGROUND

[0002] With the rapid development of aviation technology, low-altitude aircraft are increasingly widely used, but the environmental noise generated by low-altitude aircraft not only causes interference to residents' lives and ecological environment, but also may affect the acoustic environment of sensitive areas such as airports around the airport. Therefore, noise problems are increasingly concerned. Low-altitude aircraft environmental noise modeling, as a core technology to address this problem, accurately simulates and predicts the noise intensity map of aircraft noise at different environments and different locations.

[0003] Traditional low-altitude aircraft environmental noise modeling methods mainly rely on array microphones and computer numerical simulation. Array microphone measurement is a traditional noise measurement method, which measures the noise level when the aircraft flies over by arranging microphone arrays at designated locations. This method requires a large number of hardware devices, and each measurement needs to reconfigure the equipment, which makes it costly and less flexible. For example, when measuring noise in urban environments, multiple microphone arrays need to be arranged at different locations to capture the noise generated by the aircraft at different locations. However, this method can only measure at limited locations and cannot fully capture the noise distribution of the aircraft along the entire flight path. Moreover, due to the limitations of geographical conditions and environmental factors, it is difficult to conduct large-scale noise measurement in complex urban environments.

[0004] Computer numerical simulation methods simulate sound wave propagation through computers to calculate the sound pressure level at any observation point. Common numerical simulation methods include wave-type methods and geometric-type methods. Wave-type methods such as finite element method, boundary element method, and finite difference time domain method can accurately simulate the propagation of sound waves, but these methods require solving a large number of equation systems, with extremely high computational cost, especially when simulating large-scale urban environments, which requires a large amount of time and computing resources. Geometric-type methods such as Gaussian beam tracing method calculate noise distribution by simulating the geometric propagation path of sound waves, which is more efficient than wave-type methods, but still has high computational complexity and high computational cost when dealing with complex geometric environments and dynamic scenarios.

[0005] To solve the problems of high cost and poor flexibility of hardware measurement, high complexity and difficulty in large-scale promotion of numerical simulation in traditional methods, the development of deep learning technology brings new opportunities for low-altitude aircraft environmental noise modeling. Some studies have used U-Net architecture, taking fixed sound source model and urban building 2D map as input, and outputting ground noise distribution. In the field of communication propagation, similar ideas also focus on 2D ground signal mapping, and the sound source form is limited to simple models such as monopole and dipole.

[0006] Although deep learning technology brings new opportunities, existing deep learning methods still have limitations in dealing with multi-modal conditions such as different aircraft types, flight altitudes, terrain layouts, and atmospheric conditions, making it difficult to accurately capture complex acoustic characteristics and environmental factor coupling effects. First, existing deep learning methods can only handle preset sound sources and cannot generalize to any complex aircraft. This is because the sound source data used in the training of existing deep learning models is often a specific type of preset sound source with fixed parameters, and its noise characteristics are relatively simple and fixed. However, actual low-altitude aircraft, due to differences in aircraft type (such as helicopters, fixed-wing drones, multi-rotor drones, etc.), engine type, number and speed of propellers, flight attitude, etc., cannot accurately capture the noise characteristics of new or complex aircraft when they are not present in the training data, making it difficult to generalize. Second, existing methods often use urban building 2D maps as input. This two-dimensional information can only reflect the planar position and general outline of the building. However, noise propagates in three dimensions, and each facade of the building will reflect noise. The reflection coefficient and reflection direction of facades with different angles and materials differ, resulting in complex three-dimensional characteristics of reflected noise in space. Since the model does not consider these three-dimensional factors and only calculates based on two-dimensional information, it cannot accurately simulate the strength changes of noise in three-dimensional space due to shielding and reflection, and the generated 2D ground map cannot accurately reflect the actual noise distribution. In addition, the network structure design of U-Net focuses more on spatial mapping learning of input data, and lacks a special mechanism to model and integrate the nonlinear coupling relationship between multiple factors. Therefore, U-Net cannot fully capture the coupling relationship between noise and terrain, atmosphere, and other factors, and cannot fully exploit the complex relationship between terrain elevation, atmospheric changes, and noise propagation characteristics, resulting in poor accuracy of the generated noise distribution map. SUMMARY

[0007] To this end, the technical problem to be solved by the present application is to overcome the defects of the existing low-altitude aircraft environmental noise modeling method based on deep learning, which is single in sound source processing, ignores three-dimensional influencing factors, and U-Net is difficult to fully capture the coupling relationship between noise and terrain, atmosphere and other factors, resulting in poor modeling accuracy of low-altitude aircraft environmental noise.

[0008] To solve the above technical problems, the present application provides a low-altitude aircraft environmental noise modeling method, comprising:

[0009] A spherical coordinate system of the environment to be tested is constructed, and the sound source directivity data of the low-altitude aircraft at different positions in the spherical coordinate system is converted into a sound source directivity tensor;

[0010] The three-dimensional environmental terrain information and three-dimensional building geometry information of the environment to be tested are converted into a terrain layout tensor;

[0011] The high-level features of the sound source directivity tensor and the terrain layout tensor, the frequency feature and the flight height feature of the low-altitude aircraft at the current moment are extracted respectively;

[0012] The high-level features of the sound source directivity tensor and the terrain layout tensor, the frequency feature and the flight height feature of the low-altitude aircraft at the current moment are fused to obtain the multi-modal latent feature at the current moment;

[0013] Based on the multi-modal latent feature at the current moment, the noise intensity map of the low-altitude aircraft at the current moment is obtained.

[0014] Preferably, the process of obtaining the multi-modal latent feature at the current moment comprises:

[0015] After the sound source directivity tensor and the terrain layout tensor are spliced, the output features of each encoder layer of the first U-Net model are obtained through the encoder of the first U-Net model;

[0016] The frequency and flight height of the low-altitude aircraft at the current moment are respectively extracted through the corresponding learnable embedding layer to obtain the frequency feature and flight height feature of the low-altitude aircraft at the current moment;

[0017] After the output features of the last encoder layer of the first U-Net model, the frequency feature and the flight height feature of the low-altitude aircraft at the current moment are fused through element-level addition, the bottleneck layer feature of the first U-Net model is obtained through the bottleneck layer of the first U-Net model;

[0018] The bottleneck layer feature of the first U-Net model and the output features of each encoder layer of the first U-Net model are obtained through the decoder of the first U-Net model to obtain the multi-modal latent feature at the current moment.

[0019] Preferably, the current time multi-modal latent feature is used to obtain the noise intensity map of the low-altitude aircraft at the current time, comprising:

[0020] The current time multi-modal latent feature is used as a condition feature of the diffusion model, and the noise intensity map of the low-altitude aircraft at the current time is obtained through a reverse diffusion process of time steps; wherein, is a preset total number of time steps.

[0021] Preferably, the current time multi-modal latent feature is used as a condition feature of the diffusion model, and the noise intensity map of the low-altitude aircraft at the current time is obtained through a reverse diffusion process of time steps, comprising:

[0022] Generate a time embedding vector for each time step through sinusoidal position encoding;

[0023] The noise intensity map of the first time step, the time embedding vector of the first time step, and the current time multi-modal latent feature are fused through the reverse diffusion process of the first time step to obtain the noise intensity map of the first time step, comprising:

[0024] Generate a time embedding vector for each time step through sinusoidal position encoding;

[0025] The noise intensity map of the first time step, the time embedding vector of the first time step, and the current time multi-modal latent feature are fused through the reverse diffusion process of the first time step to obtain the noise intensity map of the first time step, comprising:

[0026] The noise intensity map of the first time step, the time embedding vector of the first time step, and the current time multi-modal latent feature are fused through element-level addition to obtain the fusion feature map of the first time step; wherein, , is a time step index, and the noise intensity map of the first time step is obtained by sampling a standard normal distribution;

[0027] The fusion feature map of the first time step is input into the encoder of the second U-Net model to obtain the target feature map output by the first encoder layer of the second U-Net model; wherein, ​total number of layers of the encoder of the second U-Net model;

[0028] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step;

[0029] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step;

[0030] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step;

[0031] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step;

[0032] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step;

[0033] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step;

[0034] respectively, the target feature map output by the second U-Net model at the i th time step through the multi-head self-attention block, to obtain the multi-head self-attention feature map corresponding to the j th encoder layer of the second U-Net model at the i th time step; ​​​​​​​​​​The target fusion feature map of the first decoder layer of the second U-Net model at the second time step;

[0035] The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The decoder layer index is denoted as i.

[0036] The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step.

[0037] The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step. The target fusion feature map of the first decoder layer of the second U-Net model at the second time step is obtained by upsampling the target fusion feature map of the first decoder layer of the second U-Net model at the second time step, and then inputting the upsampling result into the first decoder layer of the second U-Net model at the second time step.

[0038] Preferably, the aircraft noise evaluation model is constructed based on a multi-modal conditional encoder and a diffusion model.

[0039] The loss function of the aircraft noise evaluation model includes:

[0040] The diffusion denoising loss between the noise intensity map of the low-altitude aircraft at the current moment and the real noise intensity map.

[0041] The smooth L1 loss between the multi-modal latent feature at the current moment and the real noise intensity map.

[0042] Preferably, the sound source directivity data includes frequency characteristics and intensity characteristics of sound.

[0043] ​​​​​​​​​​Preferably, a spherical coordinate system of the environment to be measured is constructed, and the sound source directivity data of the low-altitude aircraft at different positions in the spherical coordinate system is converted into a sound source directivity tensor after being re-projected onto a hemispherical surface.

[0044] Preferably, the noise intensity map of the low-altitude aircraft at each time is obtained, and for each noise intensity map belonging to the same flight altitude, noise footprint processing is performed to obtain a plurality of 2D images belonging to the same flight altitude.

[0045] The plurality of 2D images belonging to the same flight altitude are generated into a 3D noise distribution visualization result through interpolation.

[0046] The application also provides a low-altitude aircraft environmental noise modeling system, comprising:

[0047] The first quantization module is configured to construct a spherical coordinate system of the environment to be measured, and convert the sound source directivity data of the low-altitude aircraft at different positions in the spherical coordinate system into a sound source directivity tensor.

[0048] The second quantization module is configured to convert the three-dimensional environmental terrain information and the three-dimensional building geometry information of the environment to be measured into a terrain layout tensor.

[0049] The feature extraction module is configured to extract high-level features of the sound source directivity tensor and the terrain layout tensor, frequency features and flight altitude features of the low-altitude aircraft at the current time.

[0050] The fusion module is configured to fuse the high-level features of the sound source directivity tensor and the terrain layout tensor, the frequency features and the flight altitude features of the low-altitude aircraft at the current time to obtain multi-modal latent features at the current time.

[0051] The prediction module is configured to obtain a noise intensity map of the low-altitude aircraft at the current time based on the multi-modal latent features at the current time.

[0052] The above technical solutions of the application have the following beneficial effects compared with the prior art:

[0053] This invention discloses a method and system for modeling environmental noise of low-altitude aircraft. The invention constructs a spherical coordinate system for the environment under test, reprojects the sound source directivity data of the low-altitude aircraft at different locations within this coordinate system onto a hemisphere, and converts it into a two-dimensional tensor. This encompasses the complex noise characteristics arising from different aircraft types and flight attitudes, achieving generalization for any complex aircraft. Simultaneously, environmental terrain information and building geometry information are converted into a terrain layout tensor, fully incorporating information such as terrain undulations and the three-dimensional structure of buildings. This accurately reflects the three-dimensional occlusion and reflection effects of different facades during noise propagation, avoiding biases caused by two-dimensional information input. Through the encoder of the first U-Net model of the aircraft noise assessment model, high-level geometric features are extracted from the concatenation of the sound source directivity tensor and the terrain layout tensor, combined with the current-moment frequency characteristics extracted through a learnable embedding layer. The characteristics of flight altitude, after element-level additive fusion, are passed through a bottleneck layer and then input into the decoder of the first U-Net model along with the output characteristics of the encoder layer to obtain multimodal latent features. Finally, these features are input into the diffusion model to output a noise intensity map. This method can deeply explore the complex relationship between environmental factors such as terrain and atmosphere and noise propagation characteristics, accurately model the nonlinear coupling relationship between multiple factors, and input the multimodal latent features at the current moment into the diffusion model of the aircraft noise assessment model to output the noise intensity map of the low-altitude aircraft at the current moment. The diffusion model conforms to the physical logic of noise spreading outward from the sound source and being gradually modulated by environmental factors. Through the backward denoising process, it gradually learns the mapping from noise to the real distribution, optimizes the calculation process of noise estimation, significantly improves the efficiency of noise modeling, and can effectively improve the prediction accuracy of the noise intensity map of low-altitude aircraft. Attached Figure Description

[0054] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein:

[0055] Figure 1 This is a flowchart illustrating a method for modeling environmental noise of low-altitude aircraft according to the present invention.

[0056] Figure 2 This is a schematic diagram of the modeling and visualization results of drone noise tracking in urban Shanghai. Detailed Implementation

[0057] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0058] Reference Figure 1 As shown in the figure, this embodiment provides a method for modeling environmental noise of low-altitude aircraft, including:

[0059] In view of the deficiencies of the prior art in processing multi-modal conditions, the present application designs a unique multi-modal perception module and multi-modal condition coding mechanism, which can accurately capture and fuse complex multi-modal conditions such as low-altitude aircraft noise source characteristics, terrain layout, and flight height, effectively improving the precision and generalization ability of noise modeling, and better meeting the complex needs of urban low-altitude aircraft noise management. The multi-modal perception module is shown in steps S1-S2.

[0060] Step S1: Construct a spherical coordinate system for the environment to be tested. After projecting the sound source directivity data of the low-altitude aircraft at different positions in the spherical coordinate system to the hemispherical surface, convert it to a sound source directivity tensor.

[0061] After obtaining the sound source directivity data of the low-altitude aircraft at different spatial positions through the spherical coordinate system, use the re-projection technique to map it to the hemispherical surface, achieving dimension reduction processing of the aircraft sound source spatial characteristics. This is aimed at matching the dimensions with the input data of another modality (128*128 specification city model). This method innovatively converts high-dimensional spatial characteristics into a multi-layer image format (128*128*2) and inputs it into a deep learning model for subsequent processing.

[0062] In this embodiment, specifically, the sound source directivity data includes the frequency characteristics and intensity characteristics (such as sound pressure level, sound power) of the sound.

[0063] This embodiment extracts the sound source characteristics of the low-altitude aircraft based on its structural parameters (including rotor, fuselage, and other key components) and operating conditions (such as flight speed, load conditions, etc.), and represents them from two dimensions of sound source directivity and frequency characteristics:

[0064] Based on the original sound source model data in spherical coordinate format, map it to the hemispherical surface through spatial re-projection technology; further use projection algorithm to convert the hemispherical surface data to 128x128 dimension tensor to quantitatively describe the directional distribution of the sound source in space.

[0065] The frequency characteristic representation then directly models the frequency parameters of the sound source in the subsequent scalar form as an independent dimension input, used to reflect the spectral characteristics of the sound source.

[0066] The traditional technology has significant limitations in describing the sound source characteristics of low-altitude aircraft: only the sound source information is abstracted as isolated scalars (such as "5m / s speed in the positive x direction" is represented by only two scalars "5m / s" and "+1"), this simplification completely loses the radiation directionality details of the sound source in three-dimensional space, and cannot reflect the intensity differences of noise in different directions under different aircraft models and flight attitudes (such as elevation angle and side deviation) (for example, the noise radiation intensity of a propeller aircraft to the front may be several times that to the back).

[0067] The present application reprojects the sound source directivity data of low-altitude aircraft at any position to a hemisphere by constructing a spherical coordinate system of the environment to be measured, forming a three-dimensional universal directivity spherical model. The model quantifies the radiation intensity of the sound source in each direction in three-dimensional space through the sound source directivity tensor, without relying on preset parameters of specific aircraft models, and can adapt to any aircraft (fixed-wing, multi-rotor, helicopter, etc.), truly realizing the universal and fine characterization of sound source characteristics.

[0068] The present application selects to input the frequency characteristics in scalar form, mainly based on the dual considerations of the dynamic characteristics of aircraft noise and engineering application practice. Although the frequency characteristics of the aircraft change dynamically during flight, this change can be fully reflected through the discretization processing of the time dimension - when the aircraft moves along the predetermined trajectory, the model makes independent prediction based on the frequency characteristics of the current position at each time, and the input frequency scalar strictly corresponds to the noise characteristics at that time. The frequency parameters of adjacent time do not conflict because they belong to different time nodes. For the broadband characteristics of aircraft noise, the model uses a single input single frequency processing method, which not only conforms to the conventional practice of dividing broadband characteristics into 1-20 discrete frequencies for analysis in industry and academia, but also has good flexibility. If you need to fully evaluate the noise characteristics in the broadband range, you can achieve this by running a corresponding number of model instances simultaneously, each instance focusing on processing a discrete frequency component, and finally integrating the prediction results at each frequency to obtain the full-band analysis conclusion.

[0069] This processing method not only ensures accurate capture of dynamic changes in frequency, but also takes into account the integrity of broadband characteristic analysis, while maintaining consistency with conventional methods in engineering practice, embodying the organic unity of theoretical rigor and application feasibility.

[0070] The sound generated by low-altitude aircraft has three-dimensional characteristics of time, space, and frequency. The sound source directivity can be represented by a directivity sphere to indicate the propagation characteristics of spherical waves in spatial direction. The time characteristics are implied in the aircraft coordinates, which affect the sound propagation distance and radiation range. As the aircraft coordinates move, the surrounding urban environment moves accordingly (the aircraft is always located directly above the city center). The frequency characteristics are taken separately as a dimension because different frequencies of sound have different attenuation factors in the atmosphere and different wall absorption factors (i.e., different reflection coefficients). These multi-modal inputs are closely related to acoustics.

[0071] Step S2: Convert the three-dimensional environmental terrain information and three-dimensional building geometry information of the environment to be tested into a terrain layout tensor;

[0072] In this embodiment, the three-dimensional environmental terrain information is obtained using the digital elevation model of the geographic information system.

[0073] The distribution, shape, and height characteristics of buildings are extracted from public datasets such as OpenStreetMap, Google Maps, and Open3Dhk to obtain three-dimensional building geometry information.

[0074] The above two types of information are combined into a geometric form city model in.stl,.vtk, etc. format using modeling software. The model is then cut into equal size ranges along the horizontal direction to form a standardized model. Finally, it is converted into a 128x128 horizontal 2D terrain layout tensor through quantization operation. ;

[0075] At the same time, the terrain height is projected onto the ground plane to form a structured representation for the aircraft noise evaluation model processing.

[0076] Step S3: Extract the high-level features of the sound source directivity tensor and the terrain layout tensor, as well as the frequency characteristics and flight height characteristics of the low-altitude aircraft at the current time.

[0077] The quantized terrain data is encoded to extract key geometric features and form a compact representation for noise propagation modeling. These features include the distribution, height of buildings, and terrain undulations, which can reflect the influence of terrain on sound wave propagation.

[0078] Step S4: Fuse the high-level features of the sound source directivity tensor and the terrain layout tensor, as well as the frequency characteristics and flight height characteristics of the low-altitude aircraft at the current time, to obtain the multi-modal latent features at the current time.

[0079] Step S5: Based on the multi-modal latent features at the current time, obtain the noise intensity map of the low-altitude aircraft at the current time.

[0080] In this embodiment, preferably, the acquisition process of the multi-modal latent feature at the current time includes:

[0081] As shown in Figure 1 , Figure 1 The multi-modal conditional encoder is the encoder of the first U-Net model, and the multi-modal conditional decoder is the decoder of the first U-Net model.

[0082] The sound source directivity tensor , the terrain layout tensor After splicing, the output features of each encoder layer of the first U-Net model are obtained through the encoder of the first U-Net model.

[0083] The frequency and flight height of the low-altitude flying object at the current time are respectively extracted through the corresponding learnable embedding layer to obtain the frequency feature and flight height feature of the low-altitude flying object at the current time.

[0084] After the output feature of the last encoder layer of the first U-Net model, the frequency feature and flight height feature of the low-altitude flying object at the current time are fused through element-level addition, the bottleneck layer feature of the first U-Net model is obtained through the bottleneck layer of the first U-Net model, and the formula is:

[0085] ,

[0086] Wherein, is the bottleneck layer feature of the first U-Net model, is the encoder of the first U-Net model, is element-level addition, is the learnable embedding layer corresponding to the frequency of the low-altitude flying object at the current time, is the frequency of the low-altitude flying object at the current time, is the learnable embedding layer corresponding to the flight height of the low-altitude flying object at the current time, is the flight height of the low-altitude flying object at the current time, is the splicing operation, represents the sound source directivity tensor, represents the terrain layout tensor.

[0087] The bottleneck layer feature of the first U-Net model and the output features of each encoder layer of the first U-Net model are obtained through the decoder of the first U-Net model to obtain the multi-modal latent feature at the current time.

[0088] In this embodiment, preferably, the acquisition process of the multi-modal latent feature at the current time includes:

[0089] Using the current multimodal latent features as conditional features of the diffusion model, through... The reverse diffusion process at each time step yields the noise intensity map of the low-altitude aircraft at the current moment; where... This represents the preset total number of time steps.

[0090] The diffusion model adopts an encoder-decoder structure based on the U-Net model, which includes an encoder. and decoder with residual connection The diffusion model uses a predefined sequence. The noise intensity map is gradually added to the step size for forward diffusion; during backward diffusion, a neural network is trained to gradually denoise the noise to restore the original noise intensity map.

[0091] In this embodiment, specifically, the step of inputting the multimodal latent features at the current moment into the diffusion model and outputting the noise intensity map of the low-altitude aircraft at the current moment includes:

[0092] Using the current multimodal latent features as conditional features of the diffusion model, through... The reverse diffusion process at each time step yields the noise intensity map of the low-altitude aircraft at the current moment; where... This represents the preset total number of time steps.

[0093] In this embodiment, preferably, the current multimodal latent features are used as conditional features of the diffusion model, through... The reverse diffusion process at each time step yields the noise intensity map of the low-altitude aircraft at the current moment, including:

[0094] like Figure 1 As shown, Figure 1 The noise estimation encoder is the encoder of the second U-Net model, and the noise estimation decoder is the decoder of the second U-Net model with residual connections.

[0095] A temporal embedding vector for each time step is generated using sinusoidal position encoding;

[0096] The first Noise intensity map at each time step , No. The time embedding vector at each time step Current multimodal latent features Through the first The reverse diffusion process at the nth time step yields the nth time step. The noise intensity map at each time step includes:

[0097] The first Noise intensity map at each time step , the time embedding vector of the first time step , the current time multi-modal latent feature , the fusion feature map of the first time step is obtained by element-level addition fusion; wherein, , is a time step index, and the noise intensity map of the first time step is obtained by standard normal distribution sampling;

[0098] The fusion feature map of the first time step is input into the encoder of the second U-Net model to obtain the target feature map output by the first encoder layer of the second U-Net model of the first time step; wherein, is the total number of encoder layers of the second U-Net model, and the formula is:

[0099] The formula is:

[0100] ,

[0101] wherein, is the target feature map output by the first encoder layer of the second U-Net model of the first time step, is a noise estimation encoder, is the noise intensity map of the first time step, is the current time multi-modal latent feature, is the time embedding vector of the first time step.

[0102] The target feature map output by the first encoder layer of the second U-Net model of the first time step is input into a multi-head self-attention block to enhance the modeling of spatial dependency relationship, associate each spatial position with the region related to the condition, and better learn the acoustic interference texture, so as to obtain the multi-head self-attention feature map corresponding to the first encoder layer of the second U-Net model of the first time step, and the formula is:

[0103] ,

[0104] wherein, is a multi-head self-attention block, is the target feature map output by the first encoder layer of the second U-Net model of the first The multi-head self-attention feature map corresponding to each encoder layer. For the query weight matrix, This is the key weight matrix. Value is the weight matrix.

[0105] The first The second U-Net model at the first time step The multi-head self-attention feature map corresponding to the encoder layer is passed through the bottleneck layer of the second U-Net model to obtain the first encoder layer. Bottleneck feature map of the second U-Net model at each time step;

[0106] The first Bottleneck feature map of the second U-Net model at each time step The multi-head self-attention feature map corresponding to the encoder layer is passed through the decoder of the second U-Net model with residual connections to obtain the first encoder layer. The noise intensity map at each time step is given by the following formula:

[0107] ,

[0108] in, For the first Noise intensity map at each time step, The denoising process for the decoder of the second U-Net model with residual connections. For the first The second U-Net model at the first time step Multi-head self-attention feature maps corresponding to each encoder layer.

[0109] The noise intensity map at time step 0 This is a noise intensity diagram of a low-altitude aircraft at the current moment.

[0110] In this embodiment, preferably, the loss function of the aircraft noise assessment model includes:

[0111] The diffusion denoising loss between the noise intensity map of the low-altitude aircraft at the current moment and the actual noise intensity map; the diffusion denoising loss is used to optimize the model's ability to predict the noise components added in each forward diffusion step.

[0112] The smoothed L1 loss between the multimodal latent features and the true noise intensity map at the current time is expressed by the following formula:

[0113] ,

[0114] in, a multi-modal condition alignment loss, i.e., a smooth L1 loss between the multi-modal latent feature at the current time and the real noise intensity map, a height of the real noise intensity map at the current time, a width of the real noise intensity map at the current time, a height index of a pixel point, a width index of a pixel point, a pixel value in the i-th row and the j-th column of the multi-modal latent feature at the current time, a pixel value in the i-th row and the j-th column of the real noise intensity map at the current time, a pixel value in the i-th row and the j-th column of the multi-modal latent feature at the current time, a pixel value in the i-th row and the j-th column of the real noise intensity map at the current time, a pixel value in the i-th row and the j-th column of the multi-modal latent feature at the current time, a pixel value in the i-th row and the j-th column of the real noise intensity map at the current time, a smooth L1 loss function.

[0115] The multi-modal condition alignment loss is used to ensure that the encoded conditions accurately capture the physically meaningful patterns in the input modalities, so as to force structural consistency with .

[0116] In this embodiment, the training process of the aircraft noise evaluation model includes:

[0117] The parameters of the diffusion model are initialized, including the weights of the second U-Net model encoder and the decoder with residual connection, the parameters of the learnable embedding layer, etc.

[0118] In the forward diffusion process, for the input noise intensity map, Gaussian noise is gradually added through a pre-defined sequence step size, to obtain the noise intensity map corresponding to each time step after adding Gaussian noise.

[0119] Based on the sound source directivity tensor, the terrain layout tensor, the frequency of the low-altitude aircraft at the current time, and the flight height output by the multi-modal perception module, the multi-modal latent feature at the current time is obtained.

[0120] Starting from the noise intensity map corresponding to each time step after adding Gaussian noise, through the backward diffusion process of steps, the noise is gradually removed to generate the final noise intensity map .

[0121] The diffusion denoising loss and the multi-modal condition alignment loss are calculated, the diffusion denoising loss is used to measure the difference between the predicted noise component and the real noise component, and the multi-modal condition alignment loss is used to measure the difference between the predicted multi-modal latent feature at the current time and the real noise intensity map ​​structural consistency between the images;

[0122] According to the calculated loss value, the parameters of the diffusion model are updated, including the weights of the U-Net encoder and decoder, the parameters of the learnable embedding layer, etc.

[0123] The steps of forward diffusion, conditional encoding, backward diffusion, loss calculation and parameter updating are repeatedly performed until the aircraft noise evaluation model converges, that is, the loss value reaches a preset threshold or the training reaches a preset number of iterations. The trained aircraft noise evaluation model is used as the target aircraft noise evaluation model.

[0124] In this embodiment, preferably, the noise intensity map of the low-altitude aircraft at each time is obtained, and for each noise intensity map belonging to the same flight altitude, noise footprint processing is performed to obtain multiple 2D images belonging to the same flight altitude.

[0125] Noise footprint refers to the effective impact range of the noise of a low-altitude aircraft flying at a certain altitude and at a certain time on the ground (or observation plane). The core function of noise footprint mapping is to convert the values in the noise intensity map into the noise impact range in space, that is, by setting a threshold, the outline and distribution of the effective noise area are extracted from the original noise intensity map to generate a 2D image focusing only on the noise impact range.

[0126] The multiple 2D images belonging to the same flight altitude are interpolated to generate a 3D noise distribution visualization result.

[0127] In this embodiment, the noise footprint reconstruction module is used to generate detailed 2D and 3D images of noise distribution visualization, providing intuitive decision-making basis for urban planning and noise mitigation strategies. The 2D visualization displays the noise levels at different altitude layers, where H=0 layer represents ground noise, and higher layers display the noise levels on building surfaces and in the air. For noise data at different altitude layers, separate visualization processing is performed to obtain 2D noise distribution images. H=0 layer represents ground noise, and H=50, 100, 150, and 200 layers represent noise levels at different altitudes.

[0128] 3D visualization generates a spatial representation of 3D noise distribution by interpolating layered noise data (2D noise data) onto a 3D environment model. Through 3D visualization, the propagation and distribution of noise in space can be more intuitively displayed, and the visualization result of noise distribution can be analyzed to identify noise hotspots, i.e., areas with high noise levels. These hotspots are crucial for evaluating the effectiveness of noise mitigation strategies and developing targeted noise control measures.

[0129] The present application aims to solve multiple key problems in the field of low-altitude aircraft environmental noise modeling. First, to overcome the shortcomings of traditional hardware measurement methods, such as high cost and poor flexibility, the present application proposes a low-altitude aircraft environmental noise modeling method, aiming to effectively reduce the measurement cost and improve the flexibility of modeling. Second, to address the high computational complexity of numerical simulation methods, the present application innovatively introduces a diffusion model to optimize the calculation process of noise estimation, significantly improving the modeling efficiency and making it more suitable for large-scale urban environmental noise simulation.

[0130] Through the innovative multi-modal perception module and diffusion model, the present application optimizes the calculation process of noise estimation and significantly improves the efficiency of noise modeling. Compared with traditional numerical simulation methods, the present application can complete large-scale urban environmental noise simulation in a shorter time, providing a more efficient solution for urban low-altitude aircraft noise management. The present application effectively reduces the cost of noise modeling, as traditional hardware measurement methods require a large number of microphone arrays and complex device configurations, which are costly. However, the present application utilizes the powerful computing capacity of deep learning models and diffusion models to reduce the dependence on hardware devices and lower the measurement cost. At the same time, this method also improves the flexibility and scalability of the model, making it better adapt to different application scenarios and requirements.

[0131] The multi-modal perception module designed in the present application can accurately capture the characteristics of low-altitude aircraft noise sources, terrain layout, and flight height, as well as other complex multi-modal conditions. By effectively fusing and encoding these conditions, the model can more accurately simulate the noise propagation process, improving the accuracy and generalization ability of noise modeling. This enables the present application to better adapt to the complex and variable requirements of urban low-altitude aircraft noise management, providing more reliable data support for developing effective noise control strategies.

[0132] The noise footprint reconstruction module proposed in the present application can generate detailed 2D and 3D images of noise distribution visualization, providing intuitive decision-making basis for urban planning and noise mitigation strategies. By analyzing the visualization results of noise distribution, noise hotspots can be clearly identified, the effectiveness of noise mitigation strategies can be evaluated, and targeted noise control measures can be developed. This is of great significance to urban managers and policymakers, and helps to achieve fine-grained management of urban low-altitude aircraft noise.

[0133] In this embodiment, the noise hotspot area is set as the area with noise level higher than 70dB, and a low-altitude aircraft environmental noise modeling method is used to model the unmanned aerial vehicle noise tracking in Shanghai urban area. As shown in Figure 2 , it is a schematic diagram of the unmanned aerial vehicle noise tracking modeling and visualization results in Shanghai urban area. Figure 2

[0134] ​The embodiment two of the application provides a low-altitude aircraft environmental noise modeling system, comprising:

[0135] The first quantization module is used for constructing a spherical coordinate system of the environment to be measured, and converting sound source directivity data of the low-altitude aircraft at different positions in the spherical coordinate system into a sound source directivity tensor;

[0136] The second quantization module is used for converting three-dimensional environmental terrain information and three-dimensional building geometry information of the environment to be measured into a terrain layout tensor;

[0137] The feature extraction module is used for extracting high-level features of the sound source directivity tensor and the terrain layout tensor, frequency features and flight height features of the low-altitude aircraft at the current moment.

[0138] The fusion module is used for fusing the high-level features of the sound source directivity tensor and the terrain layout tensor, the frequency features and the flight height features of the low-altitude aircraft at the current moment, to obtain multi-modal latent features at the current moment.

[0139] The prediction module is used for acquiring a noise intensity map of the low-altitude aircraft at the current moment based on the multi-modal latent features at the current moment.

[0140] Those skilled in the art should understand that the embodiments of the application can be provided as a method, a system, or a computer program product. Therefore, the application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the application can adopt a computer program product in the form of one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.

[0141] The application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks

[0142] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0144] Obviously, the above-described embodiments are only examples and are not intended to limit the present application. Based on the above description, one of ordinary skill in the art can make other variations and changes without departing from the present application. It is not necessary to recite all of the embodiments, and obvious changes or variations are not included herein. However, such obvious changes or variations are still within the scope of the present application.

Claims

1. A method of modeling ambient noise for a low altitude aircraft, the method comprising: The method comprises the following steps: A spherical coordinate system of a to-be-tested environment is constructed, and sound source directivity data of a low-altitude flying vehicle at different positions in the spherical coordinate system is converted into a sound source directivity tensor; Three-dimensional environmental terrain information and three-dimensional building geometric information of the to-be-tested environment are converted into a terrain layout tensor; High-level features of the sound source directivity tensor and the terrain layout tensor, frequency features and flight height features of the low-altitude flying vehicle at the current moment are extracted respectively; The high-level features of the sound source directivity tensor and the terrain layout tensor, the frequency features and the flight height features of the low-altitude flying vehicle at the current moment are fused to obtain multi-modal latent features at the current moment, which comprise: The sound source directivity tensor and the terrain layout tensor are spliced, and then the output features of each encoder layer of a first U-Net model are obtained through an encoder of the first U-Net model; The frequency and flight height of the low-altitude flying vehicle at the current moment are respectively input into corresponding learnable embedding layers to extract the frequency features and the flight height features of the low-altitude flying vehicle at the current moment; The output features of the last encoder layer of the first U-Net model, the frequency features and the flight height features of the low-altitude flying vehicle at the current moment are fused through element-level addition, and then the bottleneck layer features of the first U-Net model are obtained through a bottleneck layer of the first U-Net model; The bottleneck layer features of the first U-Net model and the output features of each encoder layer of the first U-Net model are input into a decoder of the first U-Net model to obtain the multi-modal latent features at the current moment; Based on the multi-modal latent features at the current moment, a noise intensity map of the low-altitude flying vehicle at the current moment is obtained.

2. The method of claim 1, wherein, The method comprises the following steps: Using the current multimodal latent features as conditional features of the diffusion model, through... The reverse diffusion process at each time step yields the noise intensity map of the low-altitude aircraft at the current moment; where... This represents the preset total number of time steps.

3. The method of claim 2, wherein, The current time multi-modal latent feature is taken as a conditional feature of the diffusion model, by The reverse diffusion process of the time step is obtained, and the noise intensity map of the low-altitude aircraft at the current time is obtained. A time embedding vector of each time step is generated through sinusoidal position encoding; The method comprises the following steps: The noise intensity map of the first time step, the time embedding vector of the first time step, and the current time multi-modal latent feature are obtained through the reverse diffusion process of the first time step. The noise intensity map of the first time step comprises the following steps: The noise intensity map of the first time step comprises the following steps: The noise intensity map of the first time step comprises the following steps: The first Noise intensity map at time step 1, the first time step 2 The temporal embedding vector at each time step and the multimodal latent features at the current time step are fused together using element-wise addition to obtain the first... The fused feature maps at each time step; among them... , For time step index, the first The noise intensity map at each time step was obtained by sampling from a standard normal distribution; The first The fused feature map at the nth time step is passed through the encoder of the second U-Net model to obtain the nth time step. The second U-Net model at the first time step The target feature map output by each encoder layer; where... This represents the total number of encoder layers in the second U-Net model. The first The second U-Net model at the first time step The target feature map output by the encoder layer is passed through a multi-head self-attention block to obtain the first... The second U-Net model at the first time step Multi-head self-attention feature maps corresponding to each encoder layer; The first U-Net model is used for feature extraction of the first time step, and the second U-Net model is used for feature extraction of the second time step. The multi-head self-attention feature map corresponding to the first encoder layer is input into the first U-Net model, and a first bottleneck feature map of the first U-Net model is obtained. The multi-head self-attention feature map corresponding to the second encoder layer is input into the second U-Net model, and a second bottleneck feature map of the second U-Net model is obtained. The multi-head self-attention feature map corresponding to the second encoder layer is input into the second U The bottleneck feature map of the second U-Net model at the first time step, The multi-head self-attention feature map corresponding to the first encoder layer is input into the decoder of the second U-Net model with a residual connection to obtain the noise intensity map at the first time step. The bottleneck feature map of the second U-Net model at the second time step, The multi-head self-attention feature map corresponding to the second encoder layer is input into the decoder of the second U-Net model with a residual connection to obtain the noise intensity map at the second time step. The noise intensity map of the 0th time step is taken as the noise intensity map of the low-altitude flying vehicle at the current moment.

4. The method of claim 3, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The first The bottleneck feature map of the second U-Net model after upsampling at the first time step and the first time step The second U-Net model at the 1st time step The multi-head self-attention feature maps corresponding to the encoder layers are concatenated along the channel dimension and then passed through the first decoder layer of the second U-Net model to obtain the first... Decoding fusion feature map of the first decoder layer of the second U-Net model at each time step; The decoding fusion feature map of the first decoder layer of the second U-Net model at the first time step, the up-sampled bottleneck feature map of the second U-Net model at the first time step, and the target fusion feature map of the first decoder layer of the second U-Net model at the first time step are obtained through element-level addition. The decoding fusion feature map of the first decoder layer of the second U-Net model at the first time step, the up-sampled bottleneck feature map of the second U-Net model at the first time step, and the target fusion feature map of the first decoder layer of the second U-Net model at the first time step are obtained through element-level addition.​​ The first The second U-Net model at the 1st time step The target fusion feature map upsampled by the decoder layer and the first decoder layer The second U-Net model at the 1st time step The output of the encoder layer is the first... After the target feature maps at each time step are concatenated along the channel dimension, they are processed by the second U-Net model. The decoder layer is obtained to get the first... The second U-Net model at the 1st time step Decoding fusion feature maps of each decoder layer; where... , For decoder layer index; The first The second U-Net model at the 1st time step Decoding fusion feature map of the decoder layer, the first The second U-Net model at the 1st time step The target fusion feature map upsampled from the decoder layer is used to obtain the th element-wise addition. The second U-Net model at the 1st time step Target fusion feature map of each decoder layer; The first The second U-Net model at the 1st time step The target fusion feature map of the decoder layer is used as the first decoder layer. Noise intensity map at each time step.

5. The method of claim 2, wherein, A flying vehicle noise evaluation model is constructed based on a multi-modal conditional encoder and a diffusion model; The loss function of the flying vehicle noise evaluation model comprises: A diffusion denoising loss between the noise intensity map of the low-altitude flying vehicle at the current moment and a real noise intensity map; A smooth L1 loss between the multi-modal latent features at the current moment and the real noise intensity map.

6. The method of claim 1, wherein, The sound source directivity data comprises frequency characteristics and intensity characteristics of sound.

7. The method of claim 1, wherein, A spherical coordinate system of a to-be-tested environment is constructed, and sound source directivity data of a low-altitude flying vehicle at different positions in the spherical coordinate system is converted into a sound source directivity tensor.

8. The method of claim 1, wherein, Noise intensity maps of the low-altitude flying vehicle at different moments are obtained, and for each noise intensity map belonging to the same flight height, noise footprint processing is performed to obtain multiple 2D images belonging to the same flight height; The multiple 2D images belonging to the same flight height are input into an interpolation module to generate a 3D noise distribution visualization result.

9. A low altitude aircraft ambient noise modeling system characterized by, The method comprises the following steps: A first quantization module is configured to construct a spherical coordinate system of a to-be-tested environment, and convert sound source directivity data of a low-altitude flying vehicle at different positions in the spherical coordinate system into a sound source directivity tensor; A second quantization module is configured to convert three-dimensional environmental terrain information and three-dimensional building geometric information of the to-be-tested environment into a terrain layout tensor; The feature extraction module is configured to extract high-level features of a sound source directivity tensor and a terrain layout tensor, frequency features and flight height features of the low-altitude flying object at the current time, respectively. The fusion module is configured to fuse the high-level features of the sound source directivity tensor and the terrain layout tensor, the frequency features and the flight height features of the low-altitude flying object at the current time, to obtain multi-modal latent features at the current time, including: After the sound source directivity tensor and the terrain layout tensor are spliced, the output features of each encoder layer of the first U-Net model are obtained through an encoder of the first U-Net model. The frequency and the flight height of the low-altitude flying object at the current time are respectively input into corresponding learnable embedding layers to extract frequency features and flight height features of the low-altitude flying object at the current time. After the output features of the last encoder layer of the first U-Net model, the frequency features and the flight height features of the low-altitude flying object at the current time are fused through element-level addition, the bottleneck layer features of the first U-Net model are obtained through a bottleneck layer of the first U-Net model. The bottleneck layer features of the first U-Net model and the output features of each encoder layer of the first U-Net model are input into a decoder of the first U-Net model to obtain the multi-modal latent features at the current time. The prediction module is configured to obtain a noise intensity map of the low-altitude flying object at the current time based on the multi-modal latent features at the current time.

Citation Information

Patent Citations

  • Interactive music intelligent generation method and system adaptive to scene space

    CN119558356A

  • Sound source localization method and sound source localization apparatus based coherence-to-diffuseness ratio mask

    US20190228790A1