Parametric array loudspeaker nonlinear distortion modeling and compensation method based on deep learning
By using deep learning and FIR filter models to model and compensate for nonlinear distortion in parametric array loudspeakers, the problem of severe nonlinear distortion in parametric array loudspeakers is solved, and a significant improvement in sound quality is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2024-11-29
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies are insufficient to effectively reduce the nonlinear distortion of parametric array loudspeakers, which limits their application.
A deep learning-based approach is used to model and compensate for nonlinear distortion in a parametric array loudspeaker. The system is modeled and the inverse filter is trained using a deep neural network F and an FIR filter model Flin. An inverse filter G is then constructed for signal preprocessing to compensate for nonlinear distortion.
It significantly reduces total harmonic distortion and intermodulation distortion of parametric array loudspeakers, outperforming existing Volterra filter methods and improving sound quality.
Smart Images

Figure CN119584016B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of loudspeaker technology and proposes a method for modeling and compensating nonlinear distortion of parametric array loudspeakers based on deep learning. Background Technology
[0002] The working principle of a parametric array loudspeaker is to modulate the audio signal to be played onto an ultrasonic carrier wave. The modulated sound signal is amplified by an amplifier and then emitted by an ultrasonic transducer. The audio signal is demodulated in the air by the nonlinear effect of the air and heard by the human ear. Compared with traditional loudspeakers, parametric array loudspeakers can generate strong directivity at low frequencies, which is beneficial for the transmission of directional sound and improves the confidentiality of speech. However, due to its special sound generation mechanism, parametric array loudspeakers have serious nonlinear distortion, which greatly limits their application. The most advanced distortion suppression method based on Volterra filters (Mu Y, Ji P, Ji W, et al. Modeling and compensation for the distortion of parametric loudspeakers using aone-dimension Volterra filter[J].IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2014, 22(12): 2169-2181.) is still difficult to reduce the nonlinear distortion of parametric array loudspeakers to a low level. Summary of the Invention
[0003] To address the severe nonlinear distortion problem in existing parametric array loudspeakers, this invention proposes a deep learning-based method for modeling and compensating nonlinear distortion in parametric array loudspeakers.
[0004] The technical solution adopted in the method of this invention is as follows:
[0005] A deep learning-based method for modeling and compensating nonlinear distortion in parametric array loudspeakers includes the following steps:
[0006] (1) Record the audio signal output by the parametric array loudspeaker system and the input signal of the system at the listening position where nonlinear distortion needs to be suppressed, and create a dataset for system modeling;
[0007] (2) The parametric array loudspeaker system is modeled as a whole using a deep neural network F. The neural network F is trained using the dataset obtained in step (1). The FIR filter model F is then used. lin Model the linear part of the system;
[0008] (3) Using the trained neural network F and the FIR filter model F lin Train another deep neural network G as the inverse filter of the system;
[0009] (4) The audio signal to be reproduced is preprocessed by a deep neural network G and then input into the parametric array speaker system to achieve compensation for the nonlinearity of the system.
[0010] Furthermore, in step (1), the electrical signal u[n] of the input parametric array loudspeaker system and the acoustic signal y[n] collected at the listening position are recorded simultaneously. The recorded input and output signals are aligned and then segmented into fragments for deep neural network training.
[0011] Furthermore, in step (2), the input of the deep neural network F is the recorded system input signal u[n], and the output is its estimated system output signal. To reduce The difference between y[n] and y[n] is used to construct the loss function for the objective. To train the network.
[0012] Further, in step (1), after training the neural network F, it is calculated whether the error between the total harmonic distortion and intermodulation distortion of the system estimated by the neural network F and its measured value meets the threshold condition. If it does, the modeling of the system is considered successful.
[0013] Furthermore, in step (1), model F lin The input is the recorded system input signal u[n], and the output is its estimated system linear output signal. By reducing The error between y[n] and y[n] is used to update the model F. lin The parameters are adjusted until convergence.
[0014] Further, in step (3), the input of the deep neural network G is the signal q[n] to be reproduced, and the output is the preprocessed signal z[n]. By passing q[n] sequentially through the deep neural networks G and F, the output signal of the parametric array loudspeaker system with the preprocessed signal z[n] as input can be obtained by the network F. Add a certain delay to q[n] and then pass it through model F lin We can obtain model F lin Estimated linear output signal of parametric array loudspeaker system To reduce and The difference between them is used to construct a loss function for the objective. To train network G.
[0015] Furthermore, in step (4), the signal q[n] to be reproduced is preprocessed offline by the trained neural network G and then input into the parametric array speaker system to achieve compensation for the nonlinearity of the system.
[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention can significantly reduce the total harmonic distortion and intermodulation distortion of the audio sound generated by parametric array loudspeakers, and its performance is significantly better than the most advanced methods based on Volterra filters. This invention provides a new method for improving the sound reproduction quality of parametric array loudspeakers. Attached Figure Description
[0017] Figure 1 This is a flowchart of the method of the present invention.
[0018] Figure 2 This is a schematic diagram of the WaveNet neural network used in the embodiments of the present invention.
[0019] Figure 3 These are comparison diagrams of (a) Total harmonic distortion (THD) and the measured values of the system estimated by the WaveNet neural network F obtained from system modeling in this embodiment of the invention, and (b) Interdistortion distortion (IMD) and the measured values.
[0020] Figure 4 The figures show a comparison of (a) linear response sound pressure level (SPL), (b) THD, and (c) IMD of the system before and after compensation using the method of this invention and the most advanced Volterra filter-based method. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings.
[0022] The process of the method of the present invention is as follows: Figure 1 As shown, the process includes steps such as dataset recording, overall system modeling, linear part modeling, inverse filter training, and compensation. The specific implementation method is as follows:
[0023] (1) First, determine the listening position where nonlinear distortion needs to be suppressed, and record the audio input and output signals of the parametric array loudspeaker system at that position. In this embodiment, the signals input to the parametric array loudspeaker system include various types of signals such as speech, music, and ambient sound, with a total duration of more than 2 hours used for overall modeling of the deep learning system (taken from the public dataset NIGENS: Ivo Trowitzsch, Jalil Taghia, Youssef Kashef, and Klaus Obermayer (2019). The NIGENS general sound events database. Technische Berlin, Tech.Rep.arXiv:1902.08314[cs.SD]), with band-limited white noise for modeling the linear part of the system, and swept-frequency signals and single-frequency plus swept-frequency signals for calculating the total harmonic distortion (THD) and interdistortion (IMD) of the system.
[0024] The electrical signal u[n] of the input parametric array loudspeaker system and the acoustic signal y[n] collected by the microphone at the listening position were simultaneously recorded by the PULSE system in an anechoic chamber at a sampling rate of 44100Hz. The time delay between u[n] and y[n] was obtained by calculating their cross-correlation and then aligned. The signal used for overall modeling of the deep learning system was divided into sample segments with a length of 65536 sampling points. 70% of these sample segments were used as the training set, 20% as the validation set, and 10% as the test set. In addition, the input-output pairs of the system input being a swept frequency signal and a single-frequency plus swept frequency signal were also included as part of the test set to test the accuracy of the model's estimation of the system's THD and IMD.
[0025] (2) Next, based on the prepared dataset, a deep learning neural network F is trained to perform an overall modeling of the entire parametric array loudspeaker system (including linear and nonlinear parts), such as... Figure 1 As shown in (a), the input of network F is the recorded system input signal u[n], and the output is its estimated system output signal. The loss function used for training is
[0026] In this embodiment, network F specifically adopts a feedforward variant of the WaveNet neural network, and its structure is as follows: Figure 2As shown, the network's input and output are the system input signal sequence and its estimated system output signal sequence, respectively. The input signal first passes through a pointwise one-dimensional convolutional layer to adjust the number of channels from 1 to C; then the signal sequentially passes through a series of residual modules, each consisting of a dilated one-dimensional convolutional layer, a gated Tanh activation layer, and a pointwise one-dimensional convolutional layer; each residual module has two outputs: one is the output signal v of the pointwise one-dimensional convolutional layer. k [n], this signal is input to the "linear mixing layer"; another output signal u k+1 [n] The input u of this module is connected by residual connections. k [n] is the output v added to the pointwise one-dimensional convolutional layer. k The signal obtained from [n] serves as the input to the next residual module; the "linear mixing layer" is implemented by a pointwise one-dimensional convolutional layer, and the output of the "linear mixing layer" is the output signal of the WaveNet network, which is the input of each residual module. In this embodiment, the WaveNet structure used to implement network F contains 9 residual modules, in which the dilation factors of the dilated convolutional layers are 1, 2, 4, ..., 256, the number of channels C of all convolutional layers is 16, and the kernel size is 16.
[0027] In this embodiment, the loss function used to train network F is:
[0028]
[0029] Where y and These are the measured system output signal and model output signal, respectively, Y and These are the time spectra of the two, respectively.
[0030] If, after training, the errors between the total harmonic distortion (THD) and intermodulation distortion (IMD) estimated by the network F and the measured values are less than the acceptable upper limit (e.g., 2%), then the system modeling is considered successful. In this embodiment, the comparison between the THD and IMD estimated by the trained WaveNet neural network F and the measured values is shown below. Figure 2 As shown, the average errors of the two are 1.08% and 0.34% respectively, indicating that the overall system modeling of this embodiment is successful.
[0031] (3) Next, model the linear part of the system, such as Figure 1 As shown in (b). In this embodiment, based on the recorded system input signal u[n] and output signal y[n] which are band-limited white noise, the least mean square (LMS) algorithm is used to reduce the FIR filter model F. lin Output Error update model F between y[n] linThe parameters are adjusted until convergence. In this embodiment, the length of the FIR (finite impulse response) filter is set to 160.
[0032] (4) Then use the network F obtained from step (2) and the FIR filter F obtained from step (3). lin Training the inverse filter G, such as Figure 1 As shown in (c), the input of network G is the signal q[n] to be reproduced, and the output is the preprocessed signal z[n]. By passing q[n] sequentially through two neural networks, G and F, the output signal of the parametric array loudspeaker system with the preprocessed signal z[n] as input can be obtained by network F. Add a certain time delay to q[n] and then pass it through the FIR filter model F lin We can obtain model F lin Estimated linear output signal of parametric array loudspeaker system To reduce and The difference between them is used to construct a loss function for the objective. To train network G; the loss function used for training is
[0033] In this embodiment, network G is also implemented using a feedforward variant of the WaveNet neural network. Furthermore, the amplitude of signal z[n] needs to be within the dynamic range of the real system's input; therefore, a Tanh activation function is added after its "linear mixing layer" to limit the network's output amplitude. The network contains 24 residual modules, where the dilation factors of the dilated convolutional layers are 1, 2, 4, ..., 2048, 1, 2, 4, ..., 2048, respectively. All convolutional layers have 24 channels and a kernel size of 4.
[0034] In this embodiment, the loss function used to train network G is:
[0035]
[0036] in The linear output signal of the system is the desired signal plus a time delay, passed through the linear model F. lin The result is obtained. In this embodiment, the latency is set to 100 sampling points (approximately 2.3ms). The system output signal estimated by model F is obtained by sequentially passing the desired signal through models G and F. and They are respectively and The time spectrum.
[0037] (5) Finally, the inverse filter G obtained from the above training is used to perform offline preprocessing on the audio signal to be reproduced, and then input into the parametric array speaker system to achieve compensation for the nonlinearity of the system, thereby reducing the nonlinear distortion of the system.
[0038] In this embodiment, the method of the present invention and the most advanced Volterra filter-based method (Mu Y, Ji P, Ji W, et al. Modeling and compensation for the distortion of parametric loudspeakers using a one-dimension Volterra filter[J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2014, 22(12): 2169-2181.) are used to calculate the system linear response sound pressure level (SPL), THD, and IMD before and after compensation. Figure 4 As shown, the method of this invention can reduce the average THD and average IMD of a parametric array loudspeaker to below 5% and 3%, respectively, and the THD and IMD compensated by the method of this invention are superior to those of the Volterra filter-based method at all frequencies. It is evident that the method of this invention can significantly reduce the total harmonic distortion and intermodulation distortion of a parametric array loudspeaker system without affecting the linear response, and its performance is significantly better than the state-of-the-art Volterra filter-based methods currently available.
Claims
1. A method for modeling and compensating nonlinear distortion of parametric array loudspeakers based on deep learning, characterized in that, The method includes the following steps: (1) Record the audio signal output by the parametric array loudspeaker system and the input signal of the system at the listening position where nonlinear distortion needs to be suppressed, and create a dataset for system modeling; (2) Using deep neural networks The parametric array loudspeaker system is modeled as a whole, and the neural network is trained using the dataset obtained in step (1). Neural Networks A feedforward variant of the WaveNet neural network was used; and an FIR filter model was employed. Model the linear part of the system; (3) Utilize the trained neural network and FIR filter model Training another deep neural network As the inverse filter of the system; deep neural network The input is the signal to be reproduced. The output is the preprocessed signal. ;Will sequentially through deep neural networks and Available from the network Estimated preprocessed signal The output signal of the input parametric array loudspeaker system ;Give Add a certain delay and then pass through the model It can be obtained from the model Estimated linear output signal of parametric array loudspeaker system To reduce and The difference between them is used to construct a loss function for the objective. To train the network ; (4) The audio signal to be reproduced is processed by a deep neural network. After preprocessing, the signal is input into the parametric array loudspeaker system to compensate for the system's nonlinearity.
2. The method for modeling and compensating nonlinear distortion of parametric array loudspeakers based on deep learning as described in claim 1, characterized in that, In step (1), the electrical signal of the input parametric array loudspeaker system is... Acoustic signals collected at the listening position The input and output signals are recorded simultaneously, and after being aligned, they are segmented into fragments for deep neural network training.
3. The method for modeling and compensating nonlinear distortion of parametric array loudspeakers based on deep learning as described in claim 2, characterized in that, In step (2), the deep neural network The input is the recorded system input signal. The output is its estimated system output signal. ; To reduce and The difference between them is used to construct a loss function for the objective. To train the network.
4. The method for modeling and compensating nonlinear distortion of parametric array loudspeakers based on deep learning as described in claim 2, characterized in that, In step (2), the neural network is trained. Then, compute the neural network. If the estimated total harmonic distortion and intermodulation distortion of the system are compared with the measured values, the error meets the threshold condition. If it does, the system modeling is considered successful.
5. The method for modeling and compensating nonlinear distortion of parametric array loudspeakers based on deep learning as described in claim 2, characterized in that, In step (2), the model The input is the recorded system input signal. The output is its estimated system linear output signal. By reducing and The error between them is used to update the model. The parameters are adjusted until convergence.
Citation Information
Patent Citations
Adaptive correction of loudspeaker using recurrent neural network
CN108024179A