Loudspeaker system and compensation method thereof
By using recurrent neural networks for signal compensation in speaker systems, the problem of complexity and high cost of nonlinear compensation in existing technologies is solved, enabling simplified and efficient speaker system design and improving playback quality.
Patent Information
- Application Number
- CN202111293120.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-29
- Filing Date
- 2021-11-03
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2041-11-03
AI Technical Summary
Existing loudspeaker systems are complex and expensive in terms of nonlinear compensation, making it difficult to effectively reduce distortion.
A recurrent neural network (RNN) is used to compensate based on the source signal and the sensed signal. The desired playback effect is generated through frequency domain transformation and feature extraction. The signal is processed and adjusted using a processor, amplifier and sensing circuit.
It achieves effective nonlinear compensation for the speaker system, simplifies system design, reduces costs, and improves playback quality.
Smart Images

Figure CN114697813B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a loudspeaker device, system and method thereof. In particular, embodiments of the present disclosure employ a neural network to compensate for distortions caused by a playback path of a loudspeaker system. BACKGROUND
[0002] Loudspeakers are often subject to linear or nonlinear distortions that result in incorrect playback. Most products currently provide linear compensation, such as filtering, equalization, and / or automatic gain control. Only a few products provide effective nonlinear compensation. However, nonlinear compensation requires in-depth knowledge of the physical characteristics of each component of the loudspeaker system. Therefore, existing nonlinear compensation systems are complex and expensive.
[0003] Therefore, there is a need for improved methods and systems that address the above issues. SUMMARY
[0004] In some embodiments of the present disclosure, a recurrent neural network employed in a loudspeaker system compensates for distortions of the loudspeaker system based on a source signal (content) and a sensed signal (signal context) of a sensing circuit. A frequency domain transform is selected to provide a mapping between the source signal and the recorded signal. And an ideal playback effect is reconstructed. Various sensing-related features and source-signal-related features are derived to be used as side information. Therefore, a desired content is generated based on the original content and the signal context.
[0005] Embodiments of the present disclosure provide a loudspeaker system for playing a sound signal. The loudspeaker system includes a processor, an amplifier, and a loudspeaker. The processor is configured to receive a source signal and generate a processed signal; the amplifier is configured to amplify the processed signal to provide an amplified signal. The loudspeaker is configured to receive the amplified signal and generate an output signal. In a deployment phase, the processor is configured to compensate the source signal using a recurrent neural network (RNN) and trained parameters to generate the processed signal. The RNN is trained according to the source signal and the output signal to generate the trained parameters.
[0006] According to some embodiments of the present disclosure, a speaker system is provided. The speaker system includes a speaker, an amplifier, a sensing circuit, and a processor. The speaker is configured to play a sound signal according to an amplified signal. The amplifier is connected to the speaker and configured to receive a justified source signal, generate the amplified signal according to the justified source signal, and transmit the amplified signal to the speaker. The sensing circuit is connected to the amplified signal and configured to measure a voltage and a current of the amplified signal, generate a sensing signal including the measured voltage and the measured current. The processor is configured to receive a source signal and the sensing signal, derive a sensing-related feature from the sensing signal, convert the source signal to a reconstructable frequency domain representation, derive a source signal-related feature from the source signal, deploy a trained recurrent neural network (RNN) to convert the reconstructable frequency domain representation to a justified frequency domain representation according to the sensing-related feature and the source signal-related feature, inverse convert the justified frequency domain representation to the justified source signal, and transmit the justified source signal to the amplifier.
[0007] According to some embodiments of the present disclosure, in the speaker system of the present disclosure, the sensing-related feature includes impedance, conductance, differential impedance, differential conductance, instantaneous power, and root mean square power.
[0008] According to some embodiments of the present disclosure, in the speaker system of the present disclosure, the reconstructable frequency domain representation is selected from a fast Fourier transform (FFT), a discrete Fourier transform (DFT), a modified discrete cosine transform (MDCT), a modified discrete sine transform (MDST), a constant Q transform (CQT), and a variable Q transform (VQT) using a filterbank distribution according to equivalent rectangular bandwidth (ERB) or Bark scale.
[0009] According to some embodiments of the present disclosure, in the speaker system of the present disclosure, the source signal-related feature includes at least one of a mel-frequency cepstral coefficient (MFCC), a perceptual linear prediction (PLP), a spectral centroid, a spectral flux, a spectral roll-off, a zero-crossing rate, a peak frequency, a kurtosis, an energy entropy, a mean amplitude, a root mean square value, a skewness, a kurtosis, and a maximum amplitude.
[0010] According to some embodiments of the application, in the loudspeaker system of the application, the recurrent neural network is a Gated Recurrent Unit (GRU).
[0011] According to some embodiments of the application, in the loudspeaker system of the application, the recurrent neural network is a Long Short-Term Memory (LSTM).
[0012] According to some embodiments of the application, in the loudspeaker system of the application, the recurrent neural network comprises memory elements storing parameters of the recurrent neural network.
[0013] According to some embodiments of the application, in the loudspeaker system of the application, the recurrent neural network is trained on a device comprising a microphone, a first delay device, a second delay device, and a neural network training device. The microphone is configured to convert the sound signal played by the loudspeaker to a recorded signal. The first delay device is configured to synchronize the source signal with the recorded signal. The second delay device is configured to synchronize the sensed signal with the recorded signal. The neural network training device is configured to receive the source signal and the sensed signal, derive the sensed-related features from the sensed signal, convert the source signal to a first frequency-domain representation, derive the source- related features from the source signal, convert the recorded signal to a second frequency- domain representation, and train the parameters of the recurrent neural network based on the first frequency-domain representation, the second frequency-domain representation, the source-related features, and the sensed-related features. During the training phase, the trained recurrent neural network is bypassed and the adjusted source signal is the source signal.
[0014] According to some embodiments of the application, in the loudspeaker system of the application, the recurrent neural network is trained by a forward training mechanism in which the first frequency-domain representation is designated as input and the second frequency-domain representation is designated as the desired output.
[0015] According to some embodiments of the application, in the loudspeaker system of the application, the recurrent neural network is trained by a backward training mechanism in which the second frequency-domain representation is designated as input and the first frequency-domain representation is designated as the desired output.
[0016] According to some embodiments of the application, there is provided a method for playing a sound signal in a speaker system, the speaker system comprising a processor configured to receive a source signal and generate a processed signal, an amplifier configured to amplify the processed signal to provide an amplified signal, and a speaker configured to receive the amplified signal and generate an output signal, the method comprising, in a training phase, training a recurrent neural network (RNN) based on the source signal and the output signal to generate trained parameters; and in a deployment phase, compensating the source signal using the RNN and the trained parameters to generate the processed signal.
[0017] According to some embodiments of the application, the method of the application further comprises, in the training phase, sensing the amplified signal to generate a sensed signal, deriving sensed-related features based on the sensed signal; converting the output signal played by the speaker into a recorded signal using a microphone; converting the source signal into a first frequency-domain representation; deriving source signal-related features based on the source signal; converting the recorded signal of the output signal into a second frequency-domain representation; training the RNN based on the first frequency-domain representation, the second frequency-domain representation, the source signal-related features, and the sensed-related features to generate the trained parameters.
[0018] According to some embodiments of the application, the method of the application further comprises, in the deployment phase, receiving the source signal and the sensed signal; deriving sensed-related features based on the sensed signal; converting the source signal into a reconstructable frequency-domain representation; deploying the trained RNN and the trained parameters to convert the reconstructable frequency-domain representation into a compensated frequency-domain representation based on the features derived from the source signal and the sensed signal; inversely converting the compensated frequency-domain representation into a compensated source signal; and transmitting the compensated source signal to the amplifier.
[0019] According to some embodiments of the application, in the method of the application, the recurrent neural network is trained by a forward training mechanism in which the first frequency-domain representation is designated as input and the second frequency-domain representation is designated as desired output.
[0020] According to some embodiments of the application, in the method of the application, the recurrent neural network is trained by a backward training mechanism in which the second frequency-domain representation is designated as input and the first frequency-domain representation is designated as desired output. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 A schematic diagram of an audio system of the application.
[0022] Figure 2 Spectrogram of recorded sweeping tone signal and sensed current voltage signal for some embodiments of the present invention.
[0023] Figure 3 Structural schematic diagram of a loudspeaker system embodiment of the present invention.
[0024] Figure 4 Dynamic waveform diagram of recorded sweeping tone, corresponding IV detection signal and derived features for an embodiment of the present invention.
[0025] Figure 5 Modified discrete cosine transform (MDCT) diagram of source signal and recorded signal for an embodiment of the present invention.
[0026] Figure 6 Constant Q transform (CQT) diagram of source signal and recorded signal for an embodiment of the present invention.
[0027] Figure 7 Structural schematic diagram of a two-layer feedforward neural network that can be used in an area function based detection module for an embodiment of the present invention.
[0028] Figure 8 Computational unit example for a gated recurrent unit (GRU).
[0029] Figure 9 Computational unit example for a long short-term memory neural network operation layer (LSTM) neural network.
[0030] Figure 10 Forward training mechanism for an embodiment of the present invention.
[0031] Figure 11 Reverse training mechanism for an embodiment of the present invention.
[0032] Figure 12 Simplified flowchart of a method of playing a sound signal by a loudspeaker system for an embodiment of the present invention.
[0033] Figure 13 Simplified structural schematic diagram of an apparatus for implementing an embodiment of the present invention.
[0034] The above figures are schematic and not drawn to scale. Relative dimensions and proportions of parts on the drawings have been shown exaggerated or reduced, for purposes of simplicity and clarity, and are not necessarily drawn to scale with respect to each other. The dimensions and the proportions of the parts on the drawings are arbitrary and not limiting.
[0035] Symbol Explanation:
[0036] 101: audio input signal
[0037] 103: analog-to-digital converter
[0038] 104: digital signal processing unit
[0039] 105: digital-to-analog converter
[0040] 106: audio amplifier unit
[0041] 107: analog signal
[0042] 109: audio output signal
[0043] 110: loudspeaker
[0044] 300: loudspeaker system
[0045] 301: processor
[0046] 302: recurrent neural network
[0047] 304: amplifier
[0048] 306: sensing circuit
[0049] 308: loudspeaker
[0050] 312: sensing signal
[0051] 313: source signal
[0052] 315: adjusted source signal
[0053] 317: amplified signal
[0054] 321: sound signal
[0055] 331: microphone
[0056] 332: recorded signal
[0057] 333: neural network training unit
[0058] 335, 337: delay device
[0059] 700: feedforward neural network
[0060] 710: input port
[0061] 720: hidden layer
[0062] 730: output layer
[0063] 740: output port
[0064] 800: unit of GRU neural network
[0065] 900: Units of an LSTM neural network
[0066] 1000: Forward Training Mechanism
[0067] 1010: Training Phase
[0068] 1020: Deployment Phase
[0069] 1100: Reverse Training Mechanism
[0070] 1110: Training Phase
[0071] 1120: Deployment Phase
[0072] 1200~1220: Steps and Procedures
[0073] 1300: Computer System
[0074] 1310: Monitor
[0075] 1320: Computer
[0076] 1330: User output device
[0077] 1340: User input device
[0078] 1350: Communication Interface
[0079] 1360: Processor
[0080] 1370: RAM
[0081] 1380: Disk drive
[0082] 1390: Bus Subsystem Detailed Implementation
[0083] The following will describe in detail the implementation of the present invention with reference to the accompanying drawings and embodiments, so that the process of how the present invention uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.
[0084] When used here, unless otherwise expressly stated in the text, the singular forms such as “a,” “the,” and “this” also include the plural forms.
[0085] Figure 1 This is a schematic diagram of the audio system of the present invention. Figure 1As shown, the audio system 100 is configured to receive an audio input signal Vin(101) and provide an audio output signal Vout(109) to the loudspeaker 110. The audio system 100 includes an analog-to-digital converter (ADC) 103, a digital signal processing unit 104, a digital-to-analog converter (DAC) 105, and an audio amplifier unit 106. The output signal from the DAC is an analog signal Va(107) that is fed into the loudspeaker 110. The detailed functions of these elements are well known to those skilled in the art and will not be described in detail here.
[0086] Figure 2 A recording of a sweeping tone and the sensed current and voltage signals according to some embodiments of the present application. In Figure 2 In (a) of the figure, the spectrum of the recorded signal of a sweeping tone play is shown. The horizontal axis represents time, from zero to three seconds; the vertical axis represents the frequency of the detected signal. The dark main line 201, highlighted by the dashed line (from 40 Hz slightly later than 0 seconds to 20000 Hz slightly later than 3 seconds), is the expected response of the sweeping tone. The secondary lines 202 above the main line represent the induced harmonics. The dark section at the bottom 203 shows the low frequency noise. The two smaller horizontal lines at about 7000 Hz (line 204) and 13000 Hz (line 205) indicate resonances due to the system size.
[0087] In Figure 2 In (a) of the figure, various distortions in an audio system are illustrated. In some embodiments of the present application, a recurrent neural network is employed to compensate for the distortions of the system based on the source signal (content), the sensed or recorded output signal (signal context) with the sensing circuit. In one example, a frequency domain transform is chosen to provide a mapping between the source signal and the recorded signal and to produce a reconstruction of the desired playback. Various sensing related features and source signal related features are derived to be used as side information. Thus, the required content is generated from the original content and the signal context using a recurrent neural network. In some embodiments, machine learning is employed to determine the mapping between the source signal and the distorted playback described above, so that the source signal can be adjusted or compensated to produce the desired playback result.
[0088] Embodiments of the present invention provide a loudspeaker system for playing sound signals. The loudspeaker system includes a processor, an amplifier, and a loudspeaker. The processor receives a source signal and generates a processed signal, the amplifier amplifies the processed signal to provide an amplified signal, and the loudspeaker receives the amplified signal and generates an output signal. During the training phase, the loudspeaker system trains a recurrent neural network (RNN) to generate trained parameters based on the source signal and the output signal, and during the deployment phase, the RNN and the trained parameters are used to compensate for the source signal to operate the loudspeaker system.
[0089] Figure 3 This is a schematic diagram of a loudspeaker system according to various embodiments of the present invention. Figure 3 As shown, the loudspeaker system 300 includes a processor 301, an amplifier 304, a sensing circuit 306, and a loudspeaker 308. The processor 301 deploys a neural network (NN) with trained parameters, such as a recurrent neural network (RNN) 302, to convert a source signal (v) 313 into a justified source signal u (315). The justified source signal u (315) is also referred to as a compensation signal or a preprocessed signal. The amplifier 304 amplifies the justified source signal u (315) to generate an amplified signal p (317), which is fed to the loudspeaker 308 for playback. The loudspeaker 308 generates an output sound signal q (321). In some embodiments, the RNN described above may include a memory element storing multiple parameters of the RNN.
[0090] The sensing circuit 306 measures the voltage and current of the amplified signal p (317) and sends the sensing signal s (312) (which includes the aforementioned measured signal) to the processor 301. Figure 2 These are examples of the sensed current and voltage signals described above. Please refer to... Figure 2 Part (a) shows the spectrum of the recorded signal during frequency sweep playback. The horizontal axis represents time, from zero to three seconds; the vertical axis represents the frequency of the detected signal. Figure 2 In the diagram, section (b) shows the spectrum of the sensed current signal I-sense, and section (c) shows the spectrum of the sensed voltage signal V-sense. The spectra of the current and voltage signals (I-sense and V-sense) are similar to those of the original sweep frequency 201 and exhibit similar distortion characteristics. The inventors of this invention have identified useful values from the measured current and voltage signals and their spectra, which can be used to enable a neural network to learn the mapping between the source signal v 313 and the adjusted source signal u (315).
[0091] Figure 4Time waveforms of recorded sweeps, corresponding IV-sense signals, and derived features according to some embodiments of the invention. In Figure 4 the horizontal axis shows time, from zero to three seconds. The (a) part shows the amplitude of the sweep signal with a logarithmic scale, increasing from 40 Hz to 20 kHz in three seconds. The solid line represents the source signal, and the dashed line represents the recorded signal. In the time domain, Figure 4 the varying amplitude recorded in the (a) part of Fig. 1 shows that the frequency response of the system needs to be equalized. In Figure 4 the (b) part shows the IV-sense signals, where the solid line shows the voltage sense signal (V-sense), and the dashed line shows the current sense signal (I-sense). Furthermore, the (c) part shows the derived resistance from the sense outputs (V / I), where the dashed line represents the instantaneous resistance, and the thick dashed line represents the frame root mean square (RMS) resistance. Other parameters can also be used in the neural network. These parameters can include conductance (I / V), differential resistance (dV / dI), or differential conductance (dl / dV). The (d) part shows the power (IV) versus time plot, where the dashed line represents the instantaneous power, and the thick dashed line represents the frame root mean square (RMS) power. These derived features (as a non-linear signal context of physical meaning) can help the neural network learn.
[0092] In some cases, it is easier to observe the distortion between the source signal and the recorded signal in the frequency domain. In these cases, it can be advantageous to convert the time domain waveforms into a frequency domain representation so that the neural network can learn more meaningfully. Many transforms can be applied in various audio applications. In some embodiments of the invention, reconstructable transforms can be used. For example, a fast Fourier transform (FFT) can be employed to achieve reconstruction. Figure 2 The example shown in Fig. 2 is derived using an FFT. For the FFT, if a window of 1024 samples is used, then the number of frequency bins is 512, each represented by a complex number, i.e. 1024 real numbers to learn. In some embodiments, a discrete Fourier transform (DFT) can be used.
[0093] Other reconstructable transforms can also be used in embodiments of the invention. Figure 5 Modified discrete cosine transform (MDCT) of the sweep source signal and recorded signal. In the (a) part and the (b) part, the horizontal axis shows time in seconds (s), and the vertical axis shows frequency in Hz. In Figure 5In the diagram, (a) shows the MDCT transformation of the swept source signal, while (b) shows the MDCT transformation of the recorded signal. Given the same settings, each bin is represented by a single real number, meaning only 512 real numbers need to be learned. Similar to MDCT, the Modified Discrete Sine Transform (MDST) can also be applied.
[0094] Figure 6 This is a constant Q-transform (CQT) for the sweep source signal and the recorded signal. Figure 6 Part (a) shows the CQT transformation of the swept source signal, while part (b) shows the CQT transformation of the recorded signal. Constant Q-transformation (CQT) is another transformation suitable for perfect reconstruction, but its frequency points are logarithmically distributed along the frequency axis. Given a frequency range of 40 Hz to 20 kHz, which is approximately nine octaves, with each octave having a resolution of 12 divisions, each division represented by a complex number, only 9 x 12 x 2 = 216 real numbers need to be learned. For near-perfect reconstruction, variable Q-transformation (VQT) can be applied, where the frequency distribution can correspond to an equivalent rectangular bandwidth (ERB) or a Bark scale.
[0095] Some non-reconstructable frequency domain representations, such as Mel-frequency cepstral coefficients (MFCC) or perceptual linear prediction (PLP), can provide auditory-relative cues suitable for source signal-related features to enhance learning. Other suitable frequency-based source signal-related features may include spectral centroid, spectral flux, spectral roll-off, spectral variability, spectral entropy, zero-crossing rate, and / or peak frequency. In the time domain waveform, useful features include mean amplitude, root mean square value, skewness, kurtosis, maximum amplitude, crest factor, and / or energy entropy. These source signal-related features provide a variety of audio properties as signal context, allowing neural networks to allocate more resources to learning additional mapping rules between them.
[0096] Please refer to Figure 3According to various embodiments of the application, the loudspeaker system 300 can comprise a loudspeaker 308 and an amplifier 304. The loudspeaker 308 plays a sound signal v(313) based on an amplified signal p(317). The amplifier 304 is connected to the loudspeaker 308, configured to receive a modified source signal u(315) and generate an amplified signal p(317) from the adjusted source signal u(315) and send the amplified signal p(317) to the loudspeaker 308. The adjusted source signal u(315) is also referred to as a compensation signal or a pre-processed signal. The loudspeaker system 300 further comprises a sensing circuit 306 connected to the amplified output signal p(317). The sensing circuit 306 is configured to measure the voltage and current of the amplified signal p(317) and generate a sensing signal s(312) comprising the measured voltage and current.
[0097] The loudspeaker system 300 further comprises a processor 301 configured to receive the source signal v(313) and the sensing signal s(312). The processor 301 is further configured to derive sensing-related features based on the sensing signal s(312) and to convert the sensing signal s(312) into a reconfigurable frequency domain representation. The processor 301 is further configured to derive source signal-related features. The processor 301 further deploys a trained recurrent neural network (RNN) 302 according to the plurality of features derived from the source signal and the sensing signal to convert the frequency domain representation into a justified frequency domain representation. The processor 301 further inverse converts the justified frequency domain representation into the justified source signal u(315) and sends the justified source signal u(315) to the amplifier.
[0098] Figure 3 A training device is also shown (as indicated by the dashed part), comprising a microphone 331, a neural network training apparatus 333, and two delay apparatuses 335 and 337. The microphone converts a sound signal q(321) played by the loudspeaker into a recorded signal r(332). The delay apparatuses 335 and 337 synchronize the source signal v(313) and the sensing signal s(312) with the recorded signal r(332). According to the source signal v(313), the sensing signal s(312), the recorded signal r(332), and the features derived from the source signal v(313) and the sensing signal s(312), the neural network training apparatus (e.g. a computer) trains the parameters W(311) of the recurrent neural network 302.
[0099] As mentioned above, a neural network can be used to compensate an input source signal to reduce output distortion. In some embodiments, the neural network can be applied to perform offline machine learning. An example of a neural network is described as follows. Please refer to Figure 7 Examples of the described general neural network, and Figure 8 and Figure 9Two examples of recurrent neural networks are described.
[0100] Figure 7 A schematic diagram of an exemplary two-layer feedforward neural network according to an embodiment of the present application is shown in FIG. 7. The exemplary two-layer feedforward neural network can also be used to construct an area-function-based detection module. In the example shown, the feedforward neural network 700 includes an input port 710, a hidden layer 720, an output layer 730, and an output port 740. In the network, information moves in only one direction, from input nodes forward through hidden nodes and output nodes. In the example shown, the input port 710 is connected to the hidden layer 720, the hidden layer 720 is connected to the output layer 730, and the output layer 730 is connected to the output port 740. Figure 7 In the example shown, the feedforward neural network 700 includes an input port 710, a hidden layer 720, an output layer 730, and an output port 740. In the network, information moves in only one direction, from input nodes forward through hidden nodes and output nodes. In the example shown, the input port 710 is connected to the hidden layer 720, the hidden layer 720 is connected to the output layer 730, and the output layer 730 is connected to the output port 740. Figure 7 In the example shown, W represents a weight vector and b represents a bias parameter.
[0101] In some embodiments, the hidden layer 720 can have Sigmoid neurons, and the output layer 730 can have Softmax neurons. A Sigmoid neuron has an output relationship defined by a Sigmoid function, which is a mathematical function that has a characteristic S-shaped curve or Sigmoid curve. Depending on the application, the Sigmoid function has a domain of all real numbers, returning values that most often increase unidirectionally from 0 to 1, or from -1 to 1. A wide variety of Sigmoid functions can be used as activation functions for artificial neurons, including logistic and hyperbolic tangent functions.
[0102] In the output layer 730, a Softmax neuron has an output relationship defined by a Softmax function. The Softmax function, or normalized exponential function, is a generalization of the logistic function that compresses a K-dimensional vector of arbitrary real values z into a K-dimensional vector of real values σ(z), where each entry lies in the range (0, 1) and the sum of all entries is 1. The output of the Softmax function can be used to represent a categorical distribution, i.e., a probability distribution over K different possible outcomes. The Softmax function is commonly used in the last layer of a neural network-based classifier. In the example shown, the input port 710 is connected to the hidden layer 720, the hidden layer 720 is connected to the output layer 730, and the output layer 730 is connected to the output port 740. Figure 7 In the example shown, W represents a weight vector and b represents a bias parameter.
[0103] To achieve reasonable classification, at least 10 neurons should be allocated in the first hidden layer. If more hidden layers are used, any number of neurons can be used in the additional hidden layers. When given more computational resources, more neurons or layers can be allocated. Providing enough neurons in its hidden layers can improve performance. More complex networks (e.g., convolutional neural networks or recurrent neural networks) can also be applied to achieve better performance. As long as there are enough neurons in its hidden layers, the vector can be well classified.
[0104] In embodiments of the application, recurrent neural networks (RNNs) process sequential data for prediction. Suitable RNNs include simple recurrent neural networks (RNNs), gated recurrent units (GRUs), as shown in Figure 8 Figure 9 The GRU uses fewer tensor operations than the LSTM. Thus, the GRU trains somewhat faster than the LSTM. On the other hand, the LSTM can provide the most controllability and thus can provide better results, but at the same time comes with more complexity and operational cost.
[0105] Figure 8 The unit 800 for a GRU neural network, where x t is the input, h t is the output, h t-1 is the previous output, and the hyperbolic tanh function is used as the activation function to help regulate the values passing through the network The GRU unit has a reset gate to decide how much past information to forget (r t ), and an update gate to decide what information to discard (1-z t ) and what new information to add (z t ), where the reset coefficient (r t ) and the update coefficient (z t ) are determined by the Sigmoid activation (σ).
[0106] Figure 9 The unit 900 for an LSTM neural network, where x t is the input, h t is the output, h t-1 is the previous output, C t is the cell state, C t-1 is the state of the previous cell, This is the state of the regulating unit activated by the first tanh function. The LSTM unit has three different gates to regulate the information flow: a forget gate to determine which information should be discarded or retained (f...). t An input gate is used to determine which information is relevant to the output from the first tanh. Maintain (i) t It is important that an output gate is used to determine the hidden state, which should be output from the second tanh output (tanh C). t What information does it carry? t The above factors are determined by the Sigmoid activation (σ).
[0107] Figure 10 This is a forward training mechanism 1000 according to various embodiments of the present invention. In the training phase 1010, the training device designates the source signal as the original content, designates the sensed output (i.e., the aforementioned sensed signal) together with the features derived from the source signal and the sensed signal as a signal context, and designates the recorded signal as the desired output, thereby training parameters (W). In the deployment phase 1020, the trained neural network with the trained parameters predicts an inferred signal (s) based on the original content and the signal context. Distortion (d) can be obtained by subtracting the source signal from the inferred signal. Finally, the adjusted signal u (also called the compensation signal) can be obtained by adding the source signal and the inverted distortion. Figure 10 In the diagram, the inverse is represented by the "-" operator.
[0108] Figure 11 This is a reverse training mechanism 1100 according to various embodiments of the present invention. In the training phase 1110, the training device receives a recorded signal (as content), a sensed output (i.e., the aforementioned sensed signal) and its derived features (as signal context), and a source signal (as the desired output) as training parameters (W). In the deployment phase 1120, the trained neural network directly predicts the adjusted signal u based on the content and signal context. The mechanism configures the neural network to infer the source signal based on the recorded signal during the training phase. Since optimal playback is the source signal, the trained neural network infers the adjusted signal that will produce the desired playback.
[0109] Figure 12 This is a simplified flowchart of a method for playing sound signals in a loudspeaker system according to various embodiments of the present invention. Figure 3 This illustrates an example of a loudspeaker system. For example... Figure 3The speaker system 300 includes a processor 301 for receiving the source signal v (313) and generating a processed signal u (315), an amplifier 305 for amplifying the processed signal u (315) to provide an amplified signal p (317), and a speaker 308 for receiving the amplified signal p (317) and generating an output signal q (321). Reference is made to Figure 3 The method 1200 includes the following steps. In step 1210, in a training phase, a recurrent neural network (RNN) is trained to generate trained parameters from a source signal and an output signal. The method 1200 also includes step 1220, in a deployment phase, using the RNN with the trained parameters to compensate a source signal to operate a speaker system.
[0110] In the training phase, in step 1210, the method includes deriving a sensing-related feature from a sensing signal, using a microphone to convert a sound signal played by the speaker into a recorded signal, converting the source signal into a first frequency-domain representation, converting the recorded signal into a second frequency-domain representation, and training the RNN to generate trained parameters from the first frequency-domain representation, the second frequency-domain representation, and the feature derived from the source signal and the sensing signal. The description of the training phase is made with reference to Figure 3 . The description of the neural network is made with reference to Figure 7 . The description of the example of the RNN is made with reference to Figure 8 and Figure 9 .
[0111] In the deployment phase, in step 1220, the method includes receiving the source signal and sensing the amplified signal, deriving a sensing-related feature from the sensing signal, converting the source signal into a reconfigurable frequency-domain representation, deploying the trained RNN using the trained parameters, thereby converting the reconfigurable frequency-domain representation into a compensated frequency-domain representation from the feature derived from the source signal and the sensing signal, converting the compensated frequency-domain representation into a compensated source signal, and sending the compensated source signal to the amplifier. The description of the processing in the deployment phase is made with reference to Figure 3 .
[0112] In some embodiments, the recurrent neural network is trained by a forward training scheme, in which the first frequency-domain representation is designated as an input and the second frequency-domain representation is designated as a desired output. The description of the example of the training process is made with reference to Figure 10 .
[0113] In some embodiments, the recurrent neural network is trained by a backward training scheme, in which the second frequency-domain representation is designated as an input and the first frequency-domain representation is designated as a desired output. The description of the example of the training process is made with reference to Figure 11 .
[0114] Figure 13 A simplified block diagram of an apparatus according to the present application that can be used to implement various embodiments. Figure 13 This description of embodiments is not intended to be exhaustive or to limit the scope of the application, as claimed in the application, to any one embodiment. Many modifications, variations, and alternatives to one or more embodiments of the present application are possible. In one embodiment, computer system 1300 typically includes a monitor 1310, a computer 1320, a user output device 1330, a user input device 1340, and a communication interface 1350, among other things.
[0115] Figure 13 A computer system that can embody the present application. For example, speaker system 300 can be implemented using a system similar to system 1300. The functions of processor 301 and neural network training unit 333 can be performed by one or more processors as shown. Speaker 308, microphone 331, and sensing circuit 306 can be peripheral devices in a system similar to system 1300 as shown. Moreover, offline training of a machine learning system can be performed in a system similar to system 1300 as shown. Figure 13 Figure 13 Figure 13
[0116] As shown, computer 1320 can include a processor 1360 in communication with a number of peripheral devices via a bus subsystem 1390. These peripheral devices can include a user output device 1330, a user input device 1340, a communication interface 1350, and storage subsystems such as a random access memory (RAM) 1370 and a disk drive 1380. Figure 13
[0117] User input device 1340 can include all possible types of devices and mechanisms used to input information to computer system 1320. These can include a keyboard, a keypad, a touch screen incorporated into the display, audio input devices such as voice recognition systems, a microphone, and other types of input devices. In various embodiments, user input device 1340 is typically implemented as a computer mouse, a trackball, a trackpad, a joystick, a wireless remote, a drawing tablet, a voice command system, an eye tracking system, and the like. User input device 1340 typically allows a user to select objects, icons, or text on the display of monitor 1310 by means of a pointing action.
[0118] User output device 1330 includes all possible types of devices and mechanisms used to output information from computer 1320. These can include a display (e.g., monitor 1310), non-visual displays such as audio output devices, etc.
[0119] The communication interface 1350 provides an interface to other communication networks and devices. The communication interface 1350 can be used to receive data from and transmit data to other systems. Implementation of the communication interface 1350 typically includes an Ethernet card, a modem (telephone, satellite, cable or ISDN), a (analog) telephone line, a Fire Wire interface, a USB interface, etc. For example, the communication interface 1350 can be coupled to a computer network, Fire Wire, etc. In other embodiments, the communication interface 1350 can be physically integrated in the motherboard of the computer 1320 and can be a software program, such as a software DSL, etc.
[0120] In various embodiments, the computer system 1300 can also include software that enables communications over the network, such as the hypertext transfer protocol (HTTP), the transmission control protocol and the network protocol (TCP / IP), the real-time streaming protocol and the real-time transport protocol (RTSP / RTP) protocols, etc. In other embodiments of the application, other communication software and transfer protocols can be used, such as the Internet packet exchange (IPX), or the user datagram protocol (UDP), etc. In some embodiments, the computer 1320 includes one or more Intel Xeon microprocessors as the processor 1360. Additionally, in one embodiment, the computer 1320 includes a UNIX-based operating system. The processor 1360 can also include special purpose processors such as digital signal processors (DSPs) and / or a reduced instruction set computer (RISC).
[0121] The RAM 1370 and the disk drive 1380 are examples of tangible storage media that store data, such as data that represents instructions of embodiments of the application, which can include executable computer code, human readable code, etc. Other types of tangible storage media that can be used in the absence of computer storage media include floppy disks, removable hard disks, optical storage media (e.g., CD-ROMs, DVDs, and bar codes), semiconductor memory (e.g., flash memory), read-only memory (ROMS), battery backed-up random access memory (RAM), networked storage devices, etc. The RAM 1370 and the disk drive 1380 can be configured to store the basic programming and data constructs that provide the functionality of the present application.
[0122] Software modules and instructions that provide the functionality of the present application can be stored in the RAM 1370 and the disk drive 1380. These software modules and instructions can be executed by the processor 1360. The RAM 1370 and the disk drive 1380 can also provide a repository for storing data used in accordance with the present application.
[0123] RAM 1370 and disk drive 1380 can include a number of memories including a main random access memory (RAM) for storage of instructions and data during program execution and a fixed or non-removable read only memory (ROM) where a basic input / output system (BIOS) is stored. RAM 1370 and disk drive 1380 can include a file storage subsystem that provides persistent (non-volatile) storage for program and data files. RAM 1370 and disk drive 1380 can further include removable storage systems such as a floppy disk drive for floppy disk storage, or a CD ROM drive for CD ROM or DVD
[0124] Bus subsystem 1390 provides a mechanism for letting the various components and subsystems of computer 1320 communicate with each other as intending. Although bus subsystem 1390 is shown as a single bus, alternative embodiments of the bus subsystem can utilize multiple buses.
[0125] Figure 13 An example of a computer system that can embody the present application is depicted. It will be apparent to those of ordinary skill in the art that many other hardware and software configurations can be substituted for the present application. For example, the computer can be a desktop computer, notebook computer, computer kiosk, or tablet computer. Additionally, the computer can be a plurality of networked computers. Furthermore, other microprocessors can be used, such as AMD's Pentium TM or Itanium TM microprocessors, Opteron TM or Athlon XP TM microprocessors, etc. Moreover, other types of operating systems can be used, such as Microsoft's Windows etc., Sun Microsystems' Solaris, LINUX, UNIX, etc. In other embodiments, the above technologies can be implemented on a chip or auxiliary processing board.
[0126] Various embodiments of the present application can be implemented in software, hardware, or a combination of both. The aforementioned logic can be stored in non-transitory computer-readable or machine-readable storage medium as an array of commands that instruct a processor of a computer system to perform a set of steps to implement the techniques described in the embodiments of the present application. The aforementioned logic can form a portion of a computer program product that can be executed by a processor of a computer system to implement the techniques described in the embodiments of the present application. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will appreciate other ways and / or methods to implement the present application.
[0127] The data structures and code described herein can be partially or entirely stored on a computer-readable storage medium and / or hardware modules and / or hardware apparatus. Computer-readable storage media include, but are not limited to, random access memory, non-volatile memory, magnetic and optical storage devices such as disk drives, magnetic tape, CDs (compact discs), DVDs (digital versatile discs) or other storage devices now known or later developed that are capable of storing code and / or data. Hardware modules or apparatus described herein include, but are not limited to, special-purpose integrated circuits (ASICs), field-programmable gate arrays (FPGAs), special-purpose or shared processors, and / or other hardware modules or apparatus now known or later developed.
[0128] The methods and processes described herein can be partially or entirely embodied in code and / or data stored in a computer-readable storage medium or device, such that when a computer system reads and executes the code and / or data, the computer system performs the associated methods and processes. The methods and processes can also be partially or entirely embodied in hardware modules or apparatus, such that when the hardware modules or apparatus are activated, they perform the associated methods and processes. Combinations of code, data, and hardware modules or apparatus can be used to embody the methods and processes disclosed herein.
[0129] While the application has been disclosed in connection with the foregoing embodiment, this application is not intended to be limited to the embodiments set forth herein, but is intended to encompass any and all variations which fall within the spirit and scope of the application hereinbefore set forth and claimed. Accordingly, the patentable scope of the application is defined only by the claims that follow.
Claims
1. A loudspeaker system, characterized in that, Include: A loudspeaker is used to play sound signals based on an amplified signal to generate an output signal; An amplifier, connected to the speaker, is used to: Receive the adjusted source signal; The amplified signal is generated based on the adjusted source signal; And transmit the amplified signal to the speaker; A sensing circuit, connected to the amplified signal, is used to: Measure the voltage and current of the amplified signal; And generate a sensing signal, the sensing signal including the measured voltage and the measured current; as well as The processor is used to receive source signals and generate processing signals, and during the deployment phase: Receive the source signal and the sensing signal, wherein the sensing signal is obtained by sensing the amplified signal; Based on the sensing signals, sensing-related features are derived; Convert the source signal to a reconstructable frequency domain representation; Based on the source signal, relevant features of the source signal are derived; Deploy a trained recurrent neural network and trained parameters to convert the reconstructable frequency domain representation into an adjusted frequency domain representation based on the sensed correlation features and the source signal correlation features; The adjusted frequency domain representation is inversely converted to the adjusted source signal; And transmit the adjusted source signal to the amplifier, wherein the adjusted source signal is the processing signal; The recurrent neural network is trained based on the source signal and the output signal to generate the trained parameters; The recurrent neural network is trained using a device, the device comprising: A microphone for converting the sound signal played by the speaker into a recording signal; A first delay device is used to synchronize the source signal and the recording signal; A second delay device is used to synchronize the sensing signal and the recording signal; as well as Neural network training device, used for: Receive the source signal and the sensing signal; The sensing-related features are derived from the sensing signals; The source signal is converted into a first frequency domain representation; Derive the relevant features of the source signal based on the source signal; Convert the recorded signal into a second frequency domain representation; and Based on the first frequency domain representation, the second frequency domain representation, the source signal correlation features, and the sensing correlation features, multiple parameters of the recurrent neural network are trained. The trained recurrent neural network described during the training phase is bypassed.
2. The loudspeaker system as claimed in claim 1, characterized in that, The sensing-related features include impedance, conductance, instantaneous power, and root mean square power.
3. The loudspeaker system as claimed in claim 1, characterized in that, The reconstructable frequency domain representation is selected from Fast Fourier Transform, Discrete Fourier Transform, Modified Discrete Cosine Transform, Modified Discrete Sine Transform, Constant Q Transform, and Variable Q Transform, wherein the Variable Q Transform uses a filtered channel distribution based on an equivalent rectangular bandwidth or a Bark scale.
4. The loudspeaker system as claimed in claim 1, characterized in that, The source signal related features include at least one of the following: Mel frequency cepstral coefficients, spectral centroid, spectral flux, zero-crossing rate, peak frequency, crest factor, energy entropy, average amplitude, root mean square value, skewness, kurtosis, and maximum amplitude.
5. The loudspeaker system as claimed in claim 1, characterized in that, The recurrent neural network is a gated recurrent unit.
6. The loudspeaker system as claimed in claim 1, characterized in that, The recurrent neural network is a long short-term memory network.
7. The loudspeaker system as claimed in claim 1, characterized in that, The recurrent neural network includes a memory element that stores multiple parameters of the recurrent neural network.
8. The loudspeaker system as claimed in claim 1, characterized in that, The recurrent neural network is trained using a forward training mechanism, in which the first frequency domain representation is designated as the input and the second frequency domain representation is designated as the desired output.
9. The loudspeaker system as claimed in claim 1, characterized in that, The recurrent neural network is trained using a reverse training mechanism, in which the second frequency domain representation is designated as the input, and the first frequency domain representation is designated as the desired output.
10. A loudspeaker system, characterized in that, Include: A processor is used to receive source signals and generate processing signals; An amplifier is used to amplify the processed signal to provide an amplified signal; And a speaker, for receiving the amplified signal and generating an output signal; During the deployment phase, the processor uses a trained recurrent neural network with trained parameters to compensate for the source signal to generate the processed signal. The recurrent neural network is trained based on the source signal and the output signal to generate the trained parameters. During the deployment phase, Receive the source signal and the sensing signal obtained by sensing the amplified signal; Based on the sensing signals, sensing-related features are derived; Convert the source signal into a reconstructable frequency domain representation; Based on the source signal, relevant features of the source signal are derived; The trained recurrent neural network and the trained parameters are deployed to convert the reconstructable frequency domain representation into an adjusted frequency domain representation based on the sensed correlation features and the source signal correlation features. The adjusted frequency domain representation is inversely converted into the adjusted source signal; And transmit the adjusted source signal to the amplifier, wherein the adjusted source signal is the processing signal; The loudspeaker system further includes: During the training phase, training is performed using a device, which includes: A microphone is used to convert the sound signal played by the speaker into a recording signal; A first delay device is used to synchronize the source signal and the recording signal; A second delay device is used to synchronize the sensing signal and the recording signal; as well as Neural network training device, used for: Receive the source signal and the sensing signal; The sensing-related features are derived from the sensing signals; The source signal is converted into a first frequency domain representation; Derive the relevant features of the source signal based on the source signal; Convert the recorded signal into a second frequency domain representation; and Based on the first frequency domain representation, the second frequency domain representation, the source signal correlation features, and the sensing correlation features, multiple parameters of the recurrent neural network are trained. The trained recurrent neural network described during the training phase is bypassed.
11. The loudspeaker system as claimed in claim 10, characterized in that, The recurrent neural network is trained using a forward training mechanism, in which the first frequency domain representation is designated as the input and the second frequency domain representation is designated as the desired output.
12. The loudspeaker system as claimed in claim 10, characterized in that, The recurrent neural network is trained using a reverse training mechanism, in which the second frequency domain representation is designated as the input and the first frequency domain representation is designated as the desired output.
13. A compensation method, characterized in that, Applicable to a loudspeaker system, the loudspeaker system comprising a processor, an amplifier, and a loudspeaker, the processor being configured to receive a source signal and generate a processed signal, the amplifier being configured to amplify the processed signal to provide an amplified signal, and the loudspeaker being configured to receive the amplified signal and generate an output signal, the method comprising: During the training phase, a recurrent neural network is trained based on the source signal and the output signal to generate a trained recurrent neural network and trained parameters. And during the deployment phase, the source signal is compensated using the trained recurrent neural network and the trained parameters to generate the processed signal; During the deployment phase, Receive the source signal and the sensing signal obtained by sensing the amplified signal; Based on the sensing signals, sensing-related features are derived; Convert the source signal into a reconstructable frequency domain representation; Based on the source signal, relevant features of the source signal are derived; The trained recurrent neural network and the trained parameters are deployed to convert the reconstructable frequency domain representation into an adjusted frequency domain representation based on the sensed correlation features and the source signal correlation features. The adjusted frequency domain representation is inversely converted into the adjusted source signal; And transmit the adjusted source signal to the amplifier, wherein the adjusted source signal is the processing signal; The compensation method further includes: During the training phase, training is performed using a device, the device comprising: A microphone is used to convert the sound signal played by the speaker into a recording signal; A first delay device is used to synchronize the source signal and the recording signal; A second delay device is used to synchronize the sensing signal and the recording signal; as well as Neural network training device, used for: Receive the source signal and the sensing signal; The sensing-related features are derived from the sensing signals; The source signal is converted into a first frequency domain representation; Derive the relevant features of the source signal based on the source signal; Convert the recorded signal into a second frequency domain representation; and Based on the first frequency domain representation, the second frequency domain representation, the source signal correlation features, and the sensing correlation features, multiple parameters of the recurrent neural network are trained. The trained recurrent neural network described during the training phase is bypassed.
14. The compensation method as described in claim 13, characterized in that, The recurrent neural network is trained using a forward training mechanism, in which the first frequency domain representation is designated as the input and the second frequency domain representation is designated as the desired output.
15. The compensation method as described in claim 13, characterized in that, The recurrent neural network is trained using a reverse training mechanism, in which the second frequency domain representation is designated as the input and the first frequency domain representation is designated as the desired output.
Citation Information
Patent Citations
Non-linear control of loudspeakers
CN106664481A