A spherical harmonic coefficient order raising method and sound field description method based on sound pressure map learning
Through the adversarial generative network learned from sound pressure maps, and using the generator and discriminator of the fully convolutional cascade structure, the mapping of spherical harmonic coefficients from low-order to high-order is realized, which solves the performance bottleneck of increasing the order of spherical harmonic coefficients in the existing technology and improves the frequency range and accuracy of sound field expression.
Patent Information
- Application Number
- CN202210650517.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-06-09
AI Technical Summary
When the order of spherical harmonic coefficients is increased, the performance of the compressed sensing method decreases significantly for high-order problems, while the performance of the fully connected neural network-based method in expressing high-order vectors of the sound field needs to be improved.
A generative adversarial network based on sound pressure map learning is adopted. Through a fully convolutional cascade generator and discriminator, the Wasserstein generative adversarial network (WGAN) is used to realize the mapping of low-order sound pressure maps to high-order sound pressure maps, thereby improving the order of spherical harmonic coefficients.
It effectively improves the expression frequency range and accuracy of spherical harmonic coefficients, and improves the spatial resolution and accuracy of sound field expression.
Smart Images

Figure CN115096432B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sound field analysis based on spherical harmonic function analysis, and specifically relates to a method for increasing the order of spherical harmonic coefficients of an adversarial generative network based on sound pressure map learning. Background Art
[0002] Spherical harmonic coefficient analysis is widely used in array signal processing. Spherical harmonics are a set of basis functions that can represent and manipulate exponential and Legendre functions defined on the unit sphere. Sound field reconstruction is one of the key applications of spherical harmonic coefficient analysis. The goal of sound field reconstruction is to reconstruct the original continuous sound field within a spatial region by sampling the sound field from a finite region. One of the most commonly used methods for acoustic analysis based on spherical harmonic coefficients is high-order spherical harmonic coefficient analysis. In high-order spherical harmonic coefficient analysis, given a cutoff order, the sound pressure at a single point can be expressed as a finite basis sum of spherical harmonic functions corresponding to the cutoff order. The cutoff order limits the reconstruction region and the available frequency bandwidth for sound field analysis. This effective region is often referred to as the optimal description region. For tasks such as sound field analysis, multiple spherical microphone arrays are often required to obtain finite-order spherical harmonic coefficients to accurately represent the sound field in the control region. By obtaining higher-order sound field representations, the number of microphone arrays can be reduced while maintaining the same accuracy in sound field description. Improving the expression of the truncation order of spherical harmonic coefficients helps to accurately describe sound pressure information over a larger range, further improving the accuracy and spatial resolution of sound field expression.
[0003] Currently, there are two main approaches for increasing the order of spherical harmonic coefficients: compressed sensing-based methods and neural network-based methods. Compressed sensing-based methods for increasing the order of spherical harmonic coefficients rely on the sparsity of the sound field. Typically, methods based on matrix optimization and plane wave decomposition are used to achieve sound field order increase within the compressed sensing framework. This sparsity-based approach is effective for low-order increases, but performance degrades significantly for higher-order problems due to the quadratic relationship between the spherical harmonic coefficient representation and the truncation order. The nonlinear mapping capabilities of neural networks play a role in this task. In a data-driven approach, supervised learning-based neural network methods use low-order spherical harmonic coefficients as input and high-order spherical harmonic coefficients as output. Compared to compressed sensing-based methods, methods based on deep neural networks offer further performance improvements. However, current neural network-based order increase efforts primarily utilize fully connected networks. More efficient network structures are needed to further improve the performance of neural networks in high-order vector representations of the sound field. Summary of the Invention
[0004] In view of the shortcomings of the existing methods, the purpose of the present invention is to provide a method for raising the spherical harmonic coefficients of an adversarial generative network based on sound pressure map learning. The present invention uses low-order sound pressure map information as system input, and learns the mapping relationship from low-order sound pressure map to high-order sound pressure map through a generative model, which can effectively improve the performance of spherical harmonic coefficient raising under the plane wave assumption. Using the obtained high-order sound pressure map information, the sound field description corresponding to the raised spherical harmonic coefficient expression can be obtained. The sound pressure map is obtained by the expanded expression of the spherical harmonic coefficients. In the expression of high-order spherical harmonic coefficients, in the area centered on the coordinate origin, within a certain range, the sound pressure value of any point is obtained by multiplying and superimposing the radial function, spherical harmonic function and high-order spherical harmonic coefficients at the location. Under ideal conditions, the radial function selects the spherical Bessel function, and the value of this function is only related to the frequency domain and position. The result of the spherical harmonic function is only related to the pitch angle and horizontal angle of the location, and has nothing to do with the radius.
[0005] The technical solution of the present invention is:
[0006] A spherical harmonic coefficient order raising method based on sound pressure map learning, comprising the following steps:
[0007] 1) Mapping the expansion of the low-order spherical harmonic coefficients of the sound field to be described into a low-order sound pressure map;
[0008] 2) Inputting the low-order sound pressure map into the generative model to obtain the corresponding high-order sound pressure map, i.e., the sound pressure map of the high-order spherical harmonic coefficients,
[0009] Complete the order upgrading of spherical harmonic coefficients.
[0010] Furthermore, the generative model is a generative adversarial network based on a fully convolutional cascade.
[0011] Furthermore, the GAN is a Wasserstein GAN; in the GAN, minimizing the approximate value of the "land moving distance" is used as a metric, and the low-order spherical harmonic coefficients and the high-order spherical harmonic coefficients are two probability distributions involved in minimizing the approximate value of the "land moving distance".
[0012] Furthermore, the adversarial generative network includes a generator and a discriminator; wherein, the generator is used to increase the order of low-order spherical harmonic coefficients to generate high-order sound field pressure maps; and the discriminator is used to determine whether the sound field pressure maps generated by the generator play a real role.
[0013] Furthermore, the discriminator is a classification network with a fully connected structure.
[0014] Furthermore, the generator adopts a deep neural network with a U-shaped structure.
[0015] Furthermore, the U-shaped deep neural network includes an encoding layer and a decoding layer; and a feature splicing operation of the same level is selected in the same level between the encoding layer and the decoding layer.
[0016] A sound field description method, comprising the steps of:
[0017] 1) Mapping the expansion of the low-order spherical harmonic coefficients of the sound field to be described into a low-order sound pressure map;
[0018] 2) Inputting the low-order sound pressure map into the generation model to obtain the corresponding high-order sound pressure map, i.e., the sound pressure map of the high-order spherical harmonic coefficients, thereby completing the order of the spherical harmonic coefficients;
[0019] 3) Describing the sound field to be described using the high-order sound pressure map.
[0020] The present invention proposes a spherical harmonic coefficient enhancement method based on a generative adversarial network learning of sound pressure maps. A generative network based on full convolution cascade is used to generate a mapping scheme from low-order sound pressure maps to high-order sound pressure maps, thereby improving the performance of spherical harmonic coefficient expansion.
[0021] The technical problem to be solved by the present invention is a method for improving the order of spherical harmonic functions. By improving the order of spherical harmonic coefficient expression, the frequency range and effective range of sound field expression can be effectively improved, as well as the accuracy of sound field expression. The technical solution adopted by the present invention is a spherical harmonic coefficient order expansion scheme based on a deep neural network. In the present invention, by mapping the expansion of the spherical harmonic coefficients into a sound pressure map, the mapping from the low-order sound pressure map to the high-order sound pressure map is achieved through the adversarial generative network, the order of the spherical harmonic coefficients is completed, and the sound pressure map expression of the high-order spherical harmonic coefficients is obtained. Using the sound pressure map result, a wider range of sound pressure description can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a traditional spherical harmonic coefficient expansion scheme.
[0023] Figure 2 This is the sound pressure map-based order expansion solution proposed in the present invention.
[0024] Figure 3 It is a sound pressure map mapping scheme based on adversarial generative networks.
[0025] Figure 4 It is a generator model based on a U-shaped fully convolutional network.
[0026] Figure 5 It is a discriminator model based on a fully connected network. DETAILED DESCRIPTION
[0027] The following describes a spherical harmonic coefficient order expansion solution provided by the present invention in conjunction with the accompanying drawings and embodiments:
[0028] 1. Sound field expression based on spherical harmonic decomposition
[0029] In the spatial sound field, in order to simplify the description, the single-frequency sound field condition is used as the expression. Under broadband or multi-frequency conditions, the corresponding single-frequency sound field can be obtained by Fourier transform of the corresponding sound pressure. In free three-dimensional space, the sound pressure solution that satisfies the homogeneous sound field wave equation can be converted into the time-independent Helmholtz equation. In the presence of a directional amplitude a(k,θ k ,φ k ) space, the sound pressure value can be described as
[0030]
[0031] Among them, the wave number It is positively correlated with the frequency, c is the speed of sound, and is a fixed value under normal temperature conditions. r = (r, θ, φ) is a spherical coordinate system, which represents the radius, pitch angle, and horizontal angle respectively. In the spherical coordinate system, the sound pressure value of a far-field plane wave at r is expressed as
[0032]
[0033] Among them, a nm (k) is the direction amplitude a(k,θ k ,φ k ) of the sphere. n (x) is the spherical Bessel function of the first kind, is the spherical coordinate basis for the decomposition of the spherical harmonics, where n is the order of the spherical harmonic coefficients.
[0034] In practical applications, such as acoustic applications based on microphone arrays, due to the distribution of microphones and the orthogonality of the spherical harmonic basis, only spherical harmonic coefficients of a finite number of cutoff orders can be obtained. In practical applications, the theoretical infinite integral under the condition of a cutoff order of N can be expressed as
[0035]
[0036] In order to make the sound field expression based on spherical harmonic function decomposition close to the actual sound field condition, the spherical harmonic coefficient order-raising scheme starts from the currently obtained N-order spherical harmonic coefficients and obtains higher-order spherical harmonic coefficients.
[0037] 2. Spherical Harmonic Coefficient Enhancement Scheme Based on Generative Adversarial Network
[0038] With the improvement of computing performance, the nonlinear mapping and learning capabilities of deep neural networks are reflected in the task of upgrading. In recent years, the spherical harmonic coefficient estimation method based on deep neural networks has attracted attention. This type of method mainly uses a fully connected network to learn the mapping of low-order spherical harmonic coefficients to high-order spherical harmonic coefficients. In the present invention, the purpose of the spherical harmonic coefficients is to reconstruct the sound pressure map. For the task of reconstructing the sound field, the purpose of the final high-order sound field expression is to obtain a continuous sound field. Based on this, we suggest using an image-based method to learn the mapping from low-order sound pressure maps to high-order sound pressure maps. The comparison of the two methods is as follows. Figure 1 、 Figure 2 In the traditional coefficient method, a fully connected network is used to directly obtain high-order coefficients, which are then converted into corresponding sound pressure images. In contrast, in this work, we implement image-based order-up, thereby designing a network structure that matches the desired result.
[0039] During the application phase, the limited-range sound pressure map obtained by expanding the low-order spherical harmonic coefficients is used as input. Through the generative network model of this work, a high-order sound pressure map corresponding to the expansion of the high-order spherical harmonic coefficients is obtained. This high-order sound pressure map can be used to obtain a wider range and higher resolution sound pressure representation, realizing the sound field representation of the area expanded by the spherical harmonic coefficients.
[0040] Theoretically, the perturbation of the elastic medium in the local area of the sound field will drive the vibration of the nearby area at the equilibrium position. The change law of the sound pressure in the local area is consistent with the property that the convolutional network can learn the characteristics of the local area. Based on this idea, this work introduces a fully convolutional network. The sound field representation of high-order spherical harmonic signals contains low-order information. On the contrary, deriving high-order sound field information from low-order sound field representation is a generative task. The overall framework of this work is based on the following Figure 3 Generative Adversarial Networks for the generative task shown.
[0041] The GAN is a two-part deep neural network consisting of a generator and a discriminator. This is an unsupervised learning task. The generator and discriminator models are trained together in a zero-sum approach until the generator is able to generate reasonable examples. The Wasserstein GAN (GAN) is a GAN that minimizes an approximation of the Earth Mover's Distance (EMD) instead of the JS divergence in the original GAN formula. EMD is a measure of the distance between two probability distributions on a region. Through gradient penalty, the GAN implements the Lipschitz continuity condition, which replaces weight clipping with a constraint on the gradient norm to achieve stable training of the GAN model. In this work, the low-order and high-order datasets can be viewed as two probability distributions, and the GAN model is then used to learn the mapping relationship from low-order sound pressure maps to high-order sound pressure maps.
[0042] The role of the generator is to increase the order of the spherical harmonic coefficients. In tasks based on sound pressure map learning, the goal is to describe the sound field more accurately in a larger area. In order to predict the two-dimensional plane sound pressure, the present invention proposes a deep neural network based on a U-shaped structure. The U-shaped deep neural network was first introduced into the biomedical image segmentation task, and since then, many applications of the U-shaped structure network have achieved significant improvements. Figure 4 The U-shaped network structure shown in the figure can capture sound pressure at different scales. The neural network takes the low-level sound pressure map as input and outputs a high-level sound pressure map corresponding to the low-level sound pressure map. The frequency range is 500Hz to 3000Hz, and the sound pressure map describes a 6m*6m square plane. The coordinate resolution is 0.1m, resulting in a 61*61 two-dimensional plane containing the origin and symmetrically distributed about the xy plane. During the neural network operation, to facilitate efficient network operation and upsampling / downsampling, zero padding is performed on both sides of the two-dimensional plane. The resulting two-dimensional sound pressure map is 64*64, which serves as the input and output feature dimensions of the neural network. The two symmetrical parts represent the encoding and decoding layers of the network, respectively. The encoder halves the feature map with a stride of 2 and doubles the filter kernel size with each downsampling operation. The decoder behaves symmetrically. Feature concatenation operations at the same level between the encoder and decoder are performed. This operation conforms to the physical principle that low-order sound field information is contained in high-order sound fields. All downsampling and upsampling operations only contain convolution operations before the nonlinear layer to simulate acoustic principles.
[0043] The discriminator plays an important role in judging whether the high-level ambient sound pressure is real. It is designed as a classification network. A fully connected network is selected. Figure 5 The fully connected architecture shown here fully considers the high-order sound pressure at every pixel. Furthermore, choosing a fully connected architecture for the discriminator helps stabilize the network without conflicting with the generator. The sound pressure input is first converted into a one-dimensional sequence. Each fully connected layer is followed by a LeakyReLU activation layer. The final layer is a sigmoid layer for binary classification problems.
[0044] 3. Experimental setup for spherical harmonic coefficient raising task based on generative adversarial networks
[0045] For the two-dimensional sound pressure plane, a plane within 1.5m from the coordinate origin is selected as the description range of the sound pressure map, and the resolution of each grid point is 0.05m. Therefore, a 61*61 grid is obtained. In order to enable the network to be trained effectively, zero padding operations are performed on both sides of the grid. Therefore, the input training data dimension of the neural network is 64*64, which is an exponential multiple of 2. It improves the training speed and network convergence performance. In the experiment, the decomposition order of the low-order spherical harmonic coefficients is 4, and the truncation order of the high-order spherical harmonic coefficients is 8, which means that the number of low-order coefficients that need to be generated is 25 and the number of high-order coefficients is 81.
[0046] For the training set, a single-frequency far-field signal with a frequency of 500-3000Hz is used as input, and its spherical harmonic coefficient expression is located at the coordinate origin. The direction information of the sound source is randomly generated, that is, the resolution of the horizontal angle and elevation angle is 1 degree, respectively, so there are 360 and 180 possible sound source directions, respectively. Considering that in actual scenes, due to interference from reflections or noise, there is only one direction without a sound source. Therefore, sound sources from different directions are added, the directions are randomly selected, and the intensity of the sound source is randomly selected according to the above data generation method. The maximum number of sound sources is set to 4. GAN and the baseline method used a total of 60,000 training data, and the training data for the four sound source cases were evenly distributed.
[0047] For the test data, we selected frequency points starting at 500Hz and increasing in 250Hz increments up to 3000Hz. For each frequency and number of sound sources, we had 100 test data points, for a total of 4400 test data points. In addition to the test data mentioned above, we also introduced a room model to test the room's performance.
[0048] During the training phase, all experimental hyperparameters for the GANs used in this paper followed the default values reported in the original paper. The gradient penalty coefficient was set to 10, and the critical number of iterations per generator iteration was 5. The Adam optimizer was used for optimization. The batch size for both the training and validation datasets was 64. All models were trained using two Titan RTX GPUs for distributed training.
[0049] 4. Experimental Evaluation Metrics
[0050] The performance of the algorithm is directly related to the acoustic pressure maps of the high-order spherical harmonics generated by the network. For evaluation, the high-order network output must be compared with the actual high-order results. During this evaluation process, a variety of objective metrics are used to quantitatively assess the quality of the generated high-order acoustic pressure images.
[0051] Mean Square Error (MSE) refers to the expected value of the square of the pixel difference between two images, and is mainly used to evaluate the degree of change in data. The smaller the minimum mean square error value, the more similar the restored image is to the real image. The calculation formula for the minimum mean square error is:
[0052]
[0053] Here, i represents the pixel points of image X and Y.
[0054] Peak Signal to Noise Ratio (PSNR) is a widely used objective image evaluation indicator. The larger the PSNR value, the smaller the distortion between the restored image and the true image, and the better the image quality. The calculation formula of Peak Signal to Noise Ratio is:
[0055]
[0056] Here, MAX represents the maximum value of the image pixel values.
[0057] Structural Similarity Index Measure (SSIM) is an indicator that measures the structural similarity between two images. The larger the value, the smaller the image distortion and the better the image quality. SSIM uses the mean as an estimate of brightness, the standard deviation as an estimate of contrast, and the covariance as a measure of structural similarity. The specific calculation process is as follows:
[0058]
[0059]
[0060]
[0061] SSIM(X,Y)=L(X,Y)*C(X,Y)*S(X,Y) (9)
[0062] μ x and μ y are the standard deviations of the sound pressure graph, σ represents the covariance matrix, and C1, C2, and C3 are constants.
[0063] 5. Experimental Results and Analysis
[0064] The experimental results of different model structures and different training methods are shown in the table. The evaluation results listed in the table are classified by the number of sound sources. In the experiment, different model structures and training methods were used for comparison. In the baseline method, the high-order coefficients output by the fully connected network are connected in series with the low-order coefficients input. The sound field is described by the expansion of the spherical harmonic coefficients to obtain a sound pressure image similar to that based on the U-type cascade method. The supervised method based on the U-shaped structure convolutional neural network is compared with the method proposed in the present invention. The U-shaped convolutional neural network based on the minimum mean square error training method is trained using the same data set. The results show that the U-shaped convolutional neural network method has better performance than the fully connected network method. Both supervised learning and the adversarial generative network based on the U-shaped convolutional neural network have achieved significant results. By further analyzing the two U-shaped convolutional neural network methods, when the sound sources do not exceed two, the performance of the supervised learning method and the adversarial generative network method is consistent. As the number of sound sources increases, the performance of the adversarial generative network model proposed in the present invention is further improved.
[0065] Table 1 Mean values of evaluation indicators under different sound source numbers
[0066]
[0067] The sound pressure image results under the conditions of single sound source, dual sound source and six sound sources are compared. The results are, in order, low-order (4th order) sound pressure image input, high-order (8th order) label, sound pressure image results of fully connected network method, supervision method based on minimum mean square error and the method proposed in this invention. The frequency of single source data is 750Hz. The azimuth angle of the sound source is 100°. The frequency of dual sound sources is 1000Hz. The azimuth angles are 70° and 85° respectively. The frequency of six sound sources is 1500Hz. The azimuth angles are 36°, 60°, 84°, 205°, 216° and 228°. The image-based results are consistent with the evaluation indicators. The fully connected network method can only calculate the results when the number of sound sources is small, and the error of high-order spherical harmonic coefficients has an erroneous effect on the sound pressure image. The method based on U-shaped fully convolutional network can obtain results similar to the ideal high-order sound pressure image. Further comparison found that for scenarios with more sound sources, in the image-based spherical harmonic coefficient expansion work, the supervision method based on minimum mean square error is not as good as the learning method based on the adversarial generative network proposed in this invention.
[0068] Table 2 Mean values of first-order reflection condition evaluation indicators
[0069] Evaluation indicators MSE PSNR SSIM Fully connected network 0.10 15.70 0.49 U-shaped structure minimum mean square error 0.06 16.78 0.57 U-shaped structure adversarial generative network 0.06 16.70 0.64
[0070] Real-world sound sources generate sound fields in their vicinity, and their behavior lends itself to modeling as simple point sources or combinations of these sources. In indoor scenes, in addition to direct sound, there are also early and late reflections. Early reflections are strongly correlated with direct sound. In this experiment, the performance of the model was explored in the presence of direct room sound and first-order reflections generated by a mirror model. Room sizes ranged from 3*3*3 to 8*8*6m. 3 The wall reflection coefficient ranges from 0.7 to 0.9. The results are shown in the table. The results show that the presence of first-order reflection will reduce the performance. However, the experimental results show that the method based on the U-shaped fully convolutional network has better performance than the fully connected network method, and the inventive method based on the adversarial generative network has a performance close to that of supervised learning in terms of MMSE and PSNR. In the calculation of SSIM, this method also maintains a relatively obvious advantage. Under the experimental conditions corresponding to the figure, the room size is 6*6*4m 3 , frequency 1000Hz, point source locations [2.0, 2.0, 1.0], and a wall reflection coefficient of 0.8. In the sound pressure-based results, the U-shaped fully convolutional network-based method remains effective, while the fully connected method's performance is severely affected. Furthermore, experimental results based on the U-shaped fully convolutional network demonstrate that this method outperforms the supervised learning method based on minimum mean square error in detail, which is consistent with the results of multi-source far-field experiments.
[0071] 6. Summary
[0072] In this work, we propose a generative adversarial network (GAN) based on a U-shaped fully convolutional network as a generator for the task of expanding the order of spherical harmonic coefficients. Through theoretical analysis and experimental verification, this approach demonstrates superior performance compared to existing methods using fully connected networks. Results demonstrate that a neural network can learn the relationship between spherical harmonic coefficients and sound pressure using sound pressure images. Results for plane waves and point sound sources demonstrate the advantages of convolutional networks for sound pressure analysis.
[0073] Finally, it should be noted that the purpose of disclosing the embodiments is to facilitate a further understanding of the present invention. However, those skilled in the art will appreciate that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the contents disclosed in the embodiments; the scope of protection claimed by the present invention shall be determined by the scope defined in the claims.
Claims
1. A spherical harmonic coefficient order raising method based on sound pressure map learning, comprising the following steps: 1) Mapping the expansion of the low-order spherical harmonic coefficients of the sound field to be described into a low-order sound pressure map; 2) Inputting the low-order sound pressure map into the generative model to obtain the corresponding high-order sound pressure map, i.e., the sound pressure map of the high-order spherical harmonic coefficients, The order of spherical harmonic coefficients is increased; the generative model is a generative adversarial network based on a fully convolutional cascade, and the generative adversarial network includes a generator and a discriminator; wherein the generator is used to increase the order of low-order spherical harmonic coefficients and generate a high-order sound field pressure map; the discriminator is used to determine whether the sound field pressure map generated by the generator plays a real role; The discriminator is a classification network with a fully connected structure; the generator adopts a U-shaped deep neural network, which includes an encoding layer and a decoding layer, and a feature splicing operation of the same level is selected in the same level between the encoding layer and the decoding layer.
2. The method according to claim 1, characterized in that The adversarial generative network is a Wasserstein adversarial generative network; the adversarial generative network uses minimizing the approximate value of the "land' moving distance" as a metric, and the low-order spherical harmonic coefficients and the high-order spherical harmonic coefficients are two probability distributions involved in minimizing the approximate value of the "land' moving distance".
3. A sound field description method, comprising the steps of: 1) Mapping the expansion of the low-order spherical harmonic coefficients of the sound field to be described into a low-order sound pressure map; 2) Inputting the low-order sound pressure map into a generative model to obtain a corresponding high-order sound pressure map, i.e., a sound pressure map of high-order spherical harmonic coefficients, thereby completing the order increase of the spherical harmonic coefficients; the generative model is a generative adversarial network based on a full convolutional cascade, and the generative adversarial network includes a generator and a discriminator; wherein the generator is used to increase the order of the low-order spherical harmonic coefficients to generate a high-order sound field sound pressure map; the discriminator is used to determine whether the sound field sound pressure map generated by the generator plays a real role; The discriminator is a classification network with a fully connected structure; the generator adopts a U-shaped deep neural network, which includes an encoding layer and a decoding layer, and a feature splicing operation of the same level is selected in the same level between the encoding layer and the decoding layer; 3) Describing the sound field to be described using the high-order sound pressure map.
4. The method according to claim 3, characterized in that The generative model is a Wasserstein generative adversarial network.