Signal transformation based on unique key-value-based network guidance and regulation
By configuring key-value space and signal transformation mapping based on key-value space, and using machine learning neural networks to conduct dynamic key-value guidance signal transformation, the problem of inefficiency of static machine learning models when dealing with multiple signal transformations is solved, and efficient modeling and transformation of time-varying signal transformation is realized.
Patent Information
- Application Number
- CN202080105682.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-31
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-07-31
Smart Images

Figure CN116324802B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to performing key-value guided signal transformation. Background Art
[0002] Static machine learning (ML) networks can model and learn fixed signal transformation functions. When there are multiple different signal transformations or in the case of continuous time-varying transformations, static ML models tend to learn, for example, suboptimal random average transformations. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 is a high-level block diagram of an example system configured with a trained neural network model to perform dynamic key-value guided signal transformation.
[0004] Figure 2 It is used for training Figure 1 Flowchart of a first example training process of a neural network machine learning (ML) model of a system to perform signal transformation.
[0005] Figure 3 is a flow chart of a second example training process for training an ML model to perform signal transformation.
[0006] Figure 4 is a block diagram of an example high-level communication system in which a neural network, once trained, can be deployed to perform inference-stage key-value guided signal transformation.
[0007] Figure 5 is a flow chart of a first example transmitter process performed in a transmitter of a communication system to produce a bitstream compatible with an ML model when the ML model is trained with a non-encoded input signal.
[0008] Figure 6 is a flow chart of a second example transmitter process performed in a transmitter of a communication system to produce a bitstream compatible with an ML model when the ML model is trained with an encoded input signal.
[0009] Figure 7 is a flow chart of an example inference phase receiver process performed in a receiver of a communication system.
[0010] Figure 8 is a flow chart of an example method for performing key-value guided signal transformation using a neural network previously trained to be configured by key-value parameters to perform signal transformation.
[0011] Fig. 9 is a block diagram of a computer device configured to implement the embodiments presented herein. DETAILED DESCRIPTION
[0012] Exemplary Embodiments
[0013] Embodiments presented herein provide key-based machine learning (ML) neural network conditioning to model time-varying signal transformations. Embodiments relate to configuring a "key-value space" and mapping signal transformations for different applications based on the key-value space. Applications range from audio signal synthesis and speech quality improvement to encryption and authentication.
[0014] These embodiments achieve at least the following high-level features:
[0015] a. Identifying a suitable key-value space associated with a signal transformation of an input signal, generating key-value parameters that uniquely represent or characterize the signal transformation and are fixed over a period of time (e.g., a frame of the input signal), and configuring a machine learning neural network using the key-value parameters corresponding to the frame of the input signal to synthesize an output signal of the transformed input signal. The key-value space associated with the signal transformation defines or contains a finite number of key-value parameters and a range of values for the key-value parameters suitable for configuring the neural network to perform the associated signal transformation.
[0016] b. During the training process of the neural network, a cost minimization criterion is adjusted or selected based on at least the characteristics of the input signal frame, the training frame, and the unique key value corresponding to the frame, so that the neural network learns to be configured by the unique key value to achieve signal transformation.
[0017] refer to Figure 1 , a high-level block diagram of an example system 100 is presented, which is configured with a trained neural network model to perform dynamic key-guided / key-based signal transformation. The system 100 is presented as a construction for describing the concepts employed in the different embodiments presented below. Therefore, not all components and signals presented in the system 100 apply to all different embodiments, as will be apparent from the subsequent description.
[0018] The system 100 includes a key value generator or estimator 102, and a key value guidance signal transformer 104, which can be deployed in a transmitter (TX) / receiver (RX) (TX / RX) system. In an example, the key value estimator 102 receives key value generation data, which may include at least an input signal, a target or expected signal, a transformation index or a signal transformation map. Based on the key value generation data, the key value estimator 102 generates or estimates a set of transformation parameters KP, also referred to as "key value parameters" KP. The key value estimator 102 can estimate the key value parameters KP frame by frame or over a group of frames, as described below. The key value parameters KP parameterize or represent the expected / target signal characteristics of the target signal, such as the spectral / frequency-based characteristics or time / time base characteristics of the target signal. In a TX / RX system, the key value parameters KP are estimated at the transmitter TX and then transmitted to the receiver RX together with the input signal.
[0019] At the receiver RX, the signal transformer 104 receives the input signal and the key parameters KP sent by the transmitter TX. The signal transformer 104 performs a desired signal transformation of the input signal based on the key parameters KP to generate an output signal having output signal characteristics similar to or matching the desired / target signal characteristics of the target signal.
[0020] The signal transformer 104 includes a previously trained neural network model configured to perform the desired KP driven signal transformation. The neural network (NN) may be a convolutional neural network (CNN) comprising a series of neural network layers having convolution filters having weights or coefficients configured based on a conventional stochastic gradient based optimization algorithm. In another example, the neural network may be based on a recursive neural network (RNN) model. In one embodiment, the neural network includes a machine learning (ML) model that is trained to be uniquely configured by key-value parameters KP to perform dynamic key-value guided signal transformation of an input signal to produce an output signal such that one or more output signal characteristics match or follow one or more desired / target signal characteristics. For example, the key-value parameters KP configure the ML model of the neural network to perform a signal transformation such that the spectral or temporal characteristics of the output signal match the corresponding desired / target spectral or temporal characteristics of the target signal. The above-described processing performed by the signal transformer 104 of the system 100 is referred to as an "inference phase" processing because the processing is performed by the neural network after the neural network of the signal transformer has been trained.
[0021] In an example where the input signal and the target signal include respective sequences of signal frames, such as respective sequences of audio frames, the key value estimator 102 estimates the key value parameters KP on a frame-by-frame basis to generate a frame-by-frame sequence of key value parameters, and the ML model of the neural network of the signal transformer 104 is configured by the key value parameters to perform a signal transformation of the input signal to the output signal frame by frame. That is, due to the frame-specific key value parameters used to guide the transformation of a given input frame, the neural network generates a uniquely transformed output frame for each given input frame / corresponding to each given input frame. Therefore, since the expected / target signal characteristics vary dynamically from frame to frame, and the estimated key value parameters representing the expected / target signal characteristics vary correspondingly from frame to frame, the key value-guided signal transformation will vary correspondingly from frame to frame, so that the output frame has a signal characteristic that tracks the signal characteristic of the target frame. In this way, the neural network of the signal transformer 104 performs a dynamic, key value-guided signal transformation on the input signal to generate an output signal that matches the target signal characteristics over time. In the subsequent description, the signal transformer 104 is also referred to as a "neural network" 104.
[0022] In various embodiments, the input signal may represent a preprocessed input signal representing the input signal, and the target signal may represent a preprocessed target signal representing the target signal, such that the key value estimator 102 estimates the key value parameter KP based on the preprocessed input signal and the preprocessed target signal, and the neural network 104 performs a signal transformation on the preprocessed input signal. In another embodiment, the key value parameter KP may represent an encoded key value parameter, such that the encoded key value parameter configures the neural network 104 to perform a signal transformation on the input signal or the preprocessed input signal. In addition, the input signal may represent an encoded input signal or an encoded preprocessed input signal, such that the key value estimator 102 and the neural network 104 each operate on the encoded input signal or the encoded preprocessed input signal. All of these and further variations are possible in various embodiments, some of which will be described below.
[0023] By way of example, various aspects of the system 100 are described in the context of an input signal and a target signal, respectively, being an audio signal, i.e., "input audio" and "target audio". It should be understood that the embodiments presented herein are equally applicable to other contexts, such as contexts in which the input signal and the target signal include respective radio frequency (RF) signals, images, videos, etc. In the audio context, the target signal may be a speech or audio signal, for example, sampled at 32 kHz and buffered as frames of 32 ms corresponding to 1024 samples per frame. Similarly, the input signal may be a speech or audio signal, for example:
[0024] a. Sample at the same sampling rate as the target signal (e.g. 32kHz), or at a different sampling rate (e.g. 16kHz, 44.1kHz, or 48kHz).
[0025] b. Buffering with the same frame duration as the target signal (e.g. 32 ms) or a different duration (e.g. 16 ms, 20 ms or 40 ms).
[0026] c. A bandwidth-limited version of the target signal. For example, the target signal is a full-band audio signal, including frequency content up to the Nyquist frequency of, for example, 16kHz, while the input signal is bandwidth-limited, having audio frequency content less than the target signal, for example, up to 4kHz, 8kHz, or 12kHz. This "bandwidth-limited" scenario is referred to as "Example A".
[0027] d. A distorted version of the target signal. For example, the input signal contains unwanted noise or a temporal / spectral distortion of the target signal. This “distorted” scenario is referred to as “Example B”.
[0028] e. Not perceptibly or intelligibly related to the target signal. For example, the input signal includes speech / conversation, while the target signal includes music; or the input signal includes music from instrument-1, while the target signal includes music from another instrument, etc. This "perceptual" scenario is called "Example C".
[0029] In one embodiment, the input signal and the target signal may each be preprocessed to produce a preprocessed input signal and a preprocessed target signal, on which the key value estimator 102 and the neural network 104 operate. Example preprocessing operations that may be performed on the input signal and the target signal include one or more of the following: resampling (e.g., downsampling or upsampling); direct current (DC) filtering to remove low frequencies, such as below 50 Hz; pre-emphasis filtering to compensate for spectral tilt in the input signal; and / or adjusting gain so that the input signal is normalized before its subsequent signal transformation.
[0030] As described above, the key value estimator 102 estimates key value parameters KP, which are used to guide / configure the neural network 104 to perform signal transformation on the input signal. In order to estimate the key value parameters KP, the key value estimator 102 can perform various different analysis operations on the input signal and the target signal to generate corresponding different sets of key value parameters KP. In one example, the key value estimator 102 performs linear prediction (LP) analysis on at least one of the target signal, the input signal, or an intermediate signal generated based on the target signal and the input signal. The LP analysis produces LP coefficients (LPC) and line spectrum frequencies (LSF), which generally compactly represent the wider spectrum envelope of the base signal (i.e., the target signal, the input signal, or the intermediate signal). The LSF compactly represents the LPC, which exhibits good quantization and inter-frame interpolation properties. In both Example A and Example B, the LSF of the target signal (i.e., which represents a reference or true value) serves as a good representation for the neural network 104 to learn or simulate the spectral envelope of the target signal (i.e., the target spectral envelope) and apply a spectral transform to the spectral envelope of the input signal (i.e., the input spectral envelope) to produce a transformed signal (i.e., the output signal) having the target spectral envelope. Therefore, in this case, the key-value parameters KP represent or form the basis of a "spectral envelope key" including spectral envelope key-value parameters. The spectral envelope key configures the neural network 104 to transform the input signal into an output signal so that the spectral envelope of the output signal (i.e., the output spectral envelope) matches or follows the target spectral envelope. In a specific non-limiting example of generating the key-value parameters, the input signal is transformed according to a whitening filter (e.g., a 2-pole filter) represented by a linear prediction polynomial of LPC order L=2 to produce an output signal. The LPC of the linear prediction polynomial is estimated during training to achieve an estimated LPC that drives the output signal to match the target signal (e.g., based on any of a variety of error / loss functions associated with the whitening desire of the input signal). The estimated LPC is then converted to LSFs (ranging from 0 to pi) and quantized using a 6-bit scalar quantizer per LSF to generate key parameters. The 6-bit scalar quantizer produces a total of 12 bits or 4096 possible combinations of unique key values; however, in this example, there are 2 key values corresponding to 2 pole locations.
[0031] In another example, the key value estimator 102 performs frequency harmonic analysis on at least one of the target signal, the input signal, or an intermediate signal generated based on the target signal and the input signal. The harmonic analysis generates a representation of a subset of dominant tone harmonics that are present in the target signal and present / missing in the input signal as a key value parameter KP. The key value estimator 102 estimates the dominant tone harmonics using, for example, a search for spectrum peaks or a sinusoidal analysis / synthesis algorithm. In this case, the key value parameter KP represents or forms the basis of a "harmonic key value" including a harmonic key value parameter. The harmonic key value configures the neural network 104 to transform the input signal into an output signal so that the output signal includes spectral features that are present in the target signal but not in the input signal. In this case, the signal transformation can represent a signal enhancement of the input signal to produce an output signal with a perceptually improved signal quality, which can, for example, include frequency bandwidth expansion (BWE). The above-mentioned LP analysis and harmonic analysis that generate LSF are both examples of spectral analysis.
[0032] In yet another example, the key value estimator 102 performs a temporal analysis (i.e., a time domain analysis) on at least one of the target signal or the intermediate signal generated based on the target signal and the input signal. The temporal analysis produces a key value parameter KP as a parameter that, for example, compactly represents the temporal evolution (e.g., gain variation) in a given frame, or a wide temporal envelope (commonly referred to as a "temporal amplitude" characteristic) of the target signal or the intermediate signal. In both the bandwidth limited example-A and the distorted example-B, the temporal characteristics of the target signal (i.e., a reference value or true value) are used as a good prototype for the neural network 104 to learn or simulate the temporal fine structure of the target signal (i.e., the desired temporal fine structure) and to apply this temporal feature conversion to the input signal. In this case, the key value parameter KP represents or forms the basis of a "temporal key" that includes the temporal key value parameter. The temporal key configures the neural network 104 to transform the input signal into an output signal so that the output signal has a desired temporal envelope.
[0033] The above-mentioned key value estimation / generation and inference phase processing relies on the trained ML model of the neural network 104. The various processes for training the ML model of the neural network 104 to perform dynamic key value guidance signal transformation will be described below in conjunction with Figure 2 and 3 Describe. Figure 2 , which is a flowchart of a first example training process 200 for training an ML model using various training signals. For example, the training signal includes a training input signal (e.g., training input audio), a training target signal (e.g., training target audio), and training key-value parameters having signal characteristics / attributes that are generally similar to the input signal and the target signal, and key-value parameters KP for inference phase processing in the system 100; however, the training signal and the inference phase signal are not the same signal. Figure 2 In the example of , the training process 200 uses an unencoded version of the input signal to train the ML model. In addition, the training process 200 operates on a frame-by-frame basis, that is, the training process operates on each frame of the input signal and the corresponding concurrent frame of the target signal.
[0034] At 202, a training process preprocesses an input signal frame to produce a preprocessed input signal frame. Example input signal preprocessing operations include resampling; DC filtering to remove low frequencies, such as below 50 Hz; pre-emphasis filtering to compensate for spectral tilt in the input signal; and / or adjusting gain so that the input signal is normalized before subsequent signal transformation. Similarly, at 204, a training process preprocesses a corresponding target signal frame to produce a preprocessed target signal frame. Target signal preprocessing may perform all or a subset of the operations performed by input signal preprocessing.
[0035] At 206, the training process estimates a corresponding set of key parameters for the input signal frame, which will guide subsequent signal transformations of the (preprocessed) input signal frame. In order to estimate the key parameters, the training system can perform various different analysis operations on the input signal frame and the corresponding target signal frame to generate corresponding different sets of key parameters in the manner described above in conjunction with the key estimation / generation and inference stage processing. For example, the training system can perform the above-mentioned LP analysis, frequency harmonic analysis, and / or time analysis on at least one of the input signal frame, the corresponding target signal frame, and the intermediate signal frame based on the input signal and the corresponding target signal frame to generate spectrum envelope keys, harmonic keys, and / or time keys for the input signal frame, respectively.
[0036] At 208, the training system encodes the key parameters to generate encoded key parameters KPT, i.e., an encoded version of the key parameters for the input signal frame. The encoding of the key parameters may include, but is not limited to, quantizing at least one or a subset of the key parameters, and encoding the key parameters using a scalar or vector quantizer codebook.
[0037] At 210, the ML model of the neural network 104 receives the preprocessed input signal frame and the encoded key parameter KPT for the input signal frame. In addition, the preprocessed target signal frame is provided to the cost minimizer CM for training. The encoded key parameter KPT configures the ML model to perform a signal transformation on the preprocessed input signal frame to produce an output signal frame. The cost minimizer CM implements a loss function to generate a current cost / error based on the difference or similarity between the output signal frame and the target signal frame. The error can represent the deviation of the desired signal characteristic of the target signal frame from the corresponding signal characteristic of the input signal frame. By updating the weights of the neural network to minimize the loss function using, for example, any known or later developed back propagation technique, the weights of the ML model are updated / trained based on the error to reduce the deviation. The loss function can be implemented using any known or later developed technology for implementing a loss function to be used for training the ML model. For example, the implementation of the loss function may include estimating the mean square error (MSE) or absolute error between the target signal and the model output signal generated by the signal transformer (model). The target signal and the model output signal can be in the time domain, the spectrum domain, or the key parameter domain. The domains here correspond to representations of the target and model output signals, where the spectral domain corresponds to a frequency domain (e.g., discrete Fourier transform (DFT)) representation of the signal, and the key-value parameter domain corresponds to a parameter representation of the signal (e.g., LPC, pitch, spectral tilt factor, and / or prediction gain) known to those skilled in the art (e.g., linear prediction coefficients, pitch, spectral tilt factor, prediction gain). In another example embodiment, the loss function can be implemented as a weighted combination of multiple errors estimated in the time domain, spectral domain, and / or key-value parameter domain.
[0038] Operations 202-210 are repeated for consecutive input and corresponding target signal frames so that the key-value parameters configure the ML model over time to perform a signal transformation on the input signal such that the output signal characteristics of the output signal match the target signal characteristics that are the target of the signal transformation. Once the ML model is trained on many frames of the input signal, the trained ML model (i.e., the trained ML model of the neural network 104) can be deployed to perform (inference phase) inference stage processing of the input signal based on the (inference phase) key-value parameters.
[0039] refer to Figure 3, which is a flow chart of a second example training process 300 for training an ML model. Training process 300 is similar to training process 200, except that training process 300 uses an encoded version of the input signal to train the ML model. The description of the above operations 202-208 is generally common to training processes 200 and 300, and should be sufficient to describe their respective functions in training process 300, and therefore will not be repeated. However, training process 300 includes an additional encoding operation 302. Encoding operation 302 encodes the input signal to produce an encoded input signal. Encoding operation 302 can encode the input signal using any known or later developed waveform-preserving audio compression technology. Signal preprocessing operation 202 then performs its preprocessing on the encoded input signal to produce an encoded, preprocessed input signal. Signal preprocessing operation 202 provides the encoded, preprocessed input signal to the ML model for training operation 310, which is performed in a manner similar to operation 210.
[0040] refer to Figure 4 , which is a block diagram of an example high-level communication system 400 in which a trained neural network 104 may be deployed to perform an inference stage key-value guided signal transformation. The communication system 400 includes a transmitter (TX) 402 in which a key-value estimator 102 may be deployed and a receiver (RX) 404 in which the trained neural network 104 is deployed. At a high level, the transmitter 402 generates a bitstream including an input signal and key-value parameters (e.g., key-value parameters KP) used to guide the transformation of the input signal, and transmits the bitstream via a communication channel. The receiver 404 receives the bitstream from the communication channel and recovers the input signal and the key-value parameters from the bitstream. The trained neural network 104 of the receiver 404 performs its inference processing and transforms the input signal recovered from the bitstream based on the key-value parameters recovered from the bitstream to produce an output signal. The following is combined with Figure 5-7 The key value estimation / generation and inference phase processing performed in the transmitter 402 and the receiver 404 is described.
[0041] refer to Figure 5 , which is a flow chart of a first example transmitter process 500 performed by the transmitter 402 to produce a bitstream compatible with an ML model of a neural network 104 previously trained with an uncoded input signal (e.g., trained according to the training process 200). The transmitter process 500 operates on a corpus of signals, such as input signals, target signals, and key-value parameters KP, that have similar statistical properties to the corresponding training signals of the training process 200. In addition, the transmitter process 500 employs many of the operations employed by the training process 200. The above description of operations 202-208 that are substantially common to both the training process 200 and the transmitter process 500 should be sufficient for the transmitter process and will not be repeated in detail.
[0042] The transmitter process 500 includes operations 202 and 204 to provide the pre-processed input signal and the pre-processed target signal to the key value estimation operation 206, respectively. Next, the key value estimation operation 206 and the key value encoding operation 208 jointly generate the encoded key value parameters KP from the pre-processed input signal and the target signal. Next, the encoding operation 502 encodes the input signal to produce the encoded / compressed input signal. Finally, the bitstream multiplexing operation 504 multiplexes the encoded input signal and the encoded key value parameters into a bitstream (i.e., a multiplexed signal) so that it is transmitted by the transmitter 402 through a communication channel.
[0043] refer to Figure 6 , which is a flow chart of a second example transmitter process 600 performed by the transmitter 402 to produce a bitstream that is compatible with an ML model of a neural network 104 that was previously trained with an encoded input signal (e.g., trained according to the training process 300). The transmitter process 600 operates on a corpus of signals having statistical properties similar to the training signals of the training process 300. In addition, the transmitter process 600 employs many of the operations employed by the training process 300. The above description of operations 202-208 and 302 that are substantially common to both the training process 300 and the transmitter process 600 should be sufficient for the transmitter process and will not be repeated in detail.
[0044] The transmitter process 600 includes operations 302 and 202, which together provide the encoded pre-processed input signal to the key value estimation operation 206 and the bit stream multiplexing operation 504. In addition, the operation 204 provides the pre-processed target signal to the key value estimation operation 206. Next, the key value generation operations 206 and 208 jointly generate the encoded key value parameters KP based on the encoded pre-processed input signal and the pre-processed target signal. Finally, the bit stream multiplexing operation 504 multiplexes the encoded input signal and the encoded key value parameters into a bit stream for the transmitter 402 to transmit through a communication channel.
[0045] See also Figure 7 , which is a flow chart of an exemplary inference phase receiver process 700 performed by the receiver 404. The receiver process 700 receives a bitstream transmitted by the transmitter 402. The receiver process 700 includes a demultiplexer-decoder operation 702 (also simply referred to as a "decoder" operation) to demultiplex and decode the encoded input signal and the encoded key-value parameters from the bitstream to recover a local copy / version of the input signal and the key-value parameters (in Figure 7 are marked as "decoded input signal" and "decoded key-value parameter" respectively).
[0046] Next, the optional input signal preprocessing operation 704 preprocesses the input signal from the bitstream demultiplexer-decoder operation 702 to produce a preprocessed version of the input signal, which represents the input signal. Based on the key-value parameters, the ML model of the neural network 104 performs the desired signal transformation on the preprocessed version of the input signal to produce an output signal (in Figure 7 In the embodiment where the preprocessing operation 704 is omitted, the ML model of the neural network 104 directly performs the desired signal transformation on the input signal. Both the processed version of the input signal and the input signal may be more generally referred to as a "signal representing the input signal".
[0047] The receiver process 700 may also include an input-output mixing operation 710 to mix the pre-processed input signal with the output signal. The input-output mixing operation 710 may include one or more of the following operations performed on a frame-by-frame basis:
[0048] a. Constant overlap-add (COLA) windowing, e.g., 50% skip and overlap-add of two consecutive windowed frames.
[0049] b. Mixing the preprocessed input signal and the windowed / filtered version of the output signal to generate the desired signal, the purpose of the mixing is to control the characteristics of the desired signal in the spectral overlap region between the output signal and the preprocessed signal. Mixing can also include post-processing the output signal based on key-value parameters to control the overall tone and noise in the output signal.
[0050] In summary, process 700 includes (i) receiving input audio and key-value parameters representing target audio characteristics, and (ii) configuring a neural network 104 previously trained to be configured by the key-value parameters using the key-value parameters to cause the neural network to perform a signal transformation of audio representing the input audio (e.g., the input audio or a preprocessed version of the input audio) to produce output audio having output audio characteristics that match the target audio characteristics. The key-value parameters may represent target spectral characteristics as the target audio characteristics, and the configuring includes configuring the neural network 104 using the key-value parameters to cause the neural network to perform a signal transformation of the input spectral characteristics of the input audio to the output spectral characteristics of the output audio that match the target spectral characteristics.
[0051] refer to Figure 8 , which is a flow chart of an example method 800 for performing key-value guided signal transformation using a neural network (e.g., neural network 104) that has been previously trained to be configured using key-value parameters to perform signal transformation, i.e., perform signal conversion based on key-value parameters.
[0052] At 802, a key value estimator receives input audio and target audio having target audio characteristics. Both the input audio and the target audio may include a sequence of audio frames. The key value estimator estimates key value parameters representing the target audio characteristics based on one or more of the target audio and the input audio. The key value estimator may perform spectral and / or temporal analysis of the input and target audio to generate key value parameters, as described above. The key value estimator may estimate the key value parameters frame by frame to generate a sequence of frame by frame key value parameters. The key value estimator provides the key value parameters to a first input of a trained neural network.
[0053] At 804, the trained neural network also receives input audio at a second input of the neural network. The key-value parameters configure the trained neural network to perform a desired signal transformation. In response to the key-value parameters, the trained neural network performs a desired signal transformation on the input audio (i.e., input audio characteristics of the input audio) to produce output audio having output audio characteristics that match the target audio characteristics. That is, the signal transformation transforms the input audio characteristics into output audio characteristics that match or are similar to the target audio characteristics. The trained neural network can be configured on a frame-by-frame basis by the sequence of frame-by-frame key-value parameters to convert each input audio frame into a corresponding output audio frame to produce the output audio as a sequence of output audio frames (one output audio frame for each input audio frame and each set of frame-by-frame key-value parameters).
[0054] In the prior training stage, the neural network is trained to perform signal transformation so as to minimize the error between the output audio and the target audio. For example, the neural network is trained by training the weights of the neural network so that the neural network performs signal transformation on the training input audio in response to the training key-value parameters to generate the training output audio so as to minimize the error.
[0055] Reference Fig. 9 , which is a block diagram of a computer device 900 configured to implement the embodiments presented herein. There are many possible configurations of the computer device 900, and Fig. 9Just one example. Examples of computer device 900 include tablet computers, personal computers, laptop computers, mobile phones such as smart phones, etc. Computer device 900 includes one or more network interface units (NIU) 908, and memory 914, each coupled to a processor 916. One or more NIUs 908 may include wired and / or wireless connection capabilities that allow processor 916 to communicate over a communication network. For example, NIU 908 may include an Ethernet card that communicates over an Ethernet connection, a wireless RF transceiver that wirelessly communicates with a cellular network in a communication network, an optical transceiver, etc., as understood by those of ordinary skill in the art. Processor 916 receives sampled or digitized audio and provides the digitized audio to one or more audio devices 918, as is known. Audio device 918 may include a microphone, a speaker, an analog-to-digital converter (ADC), and a digital-to-analog converter (DAC).
[0056] The processor 916 may include a collection of microcontrollers and / or microprocessors, for example, each configured to execute corresponding software instructions stored in the memory 914. The processor 916 may implement an ML model of a neural network. The processor 916 may be implemented in one or more programmable application specific integrated circuits (ASICs), firmware, or a combination thereof. Portions of the memory 914 (and the instructions therein) may be integrated with the processor 916. As used herein, the terms "acoustics," "audio," and "sound" are synonymous and interchangeable.
[0057] The memory 914 may include a read-only memory (ROM), a random access memory (RAM), a disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible (e.g., non-transitory) memory storage device. Thus, typically, the memory 914 may include one or more computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (by the processor 916), it is operable to perform the operations described herein. For example, the memory 914 stores or encodes instructions for the control logic 920 to implement modules that are configured to perform the operations described herein related to ML models of neural networks, key-value estimators, input / target signal preprocessing, input signal and key-value encoding and decoding, cost minimization, bitstream multiplexing and demultiplexing, input-output mixing (post-processing), etc., and the above-mentioned methods.
[0058] Additionally, the memory 914 stores data / information 922 used and generated by the processor 916 , including key-value parameters, input audio, target audio, and output audio, as well as coefficients and weights employed by the ML model of the neural network.
[0059] In summary, in one embodiment, a method is provided, including: receiving input audio and target audio having target audio characteristics; estimating key-value parameters representing the target audio characteristics based on one or more of the input audio and the target audio; and configuring a neural network, which is trained to be configured by the key-value parameters, and the key-value parameters cause the neural network to perform signal transformation of the input audio to generate output audio having output audio characteristics that correspond to and match the target audio characteristics.
[0060] In another embodiment, a device is provided, comprising: a key-value estimator for receiving input audio and target audio having target audio characteristics; and estimating key-value parameters representing the target audio characteristics based on one or more of the input audio and the target audio; and a neural network trained to be configured by the key-value parameters to perform signal transformation of the input audio to produce output audio having output audio characteristics that correspond to and match the target audio characteristics.
[0061] In yet another embodiment, a non-transitory computer-readable medium is provided. The medium is encoded with instructions that, when executed by a processor, cause the processor to perform: receiving input audio and target audio having target audio characteristics; estimating key-value parameters representing the target audio characteristics based on one or more of the input audio and the target audio; and configuring a neural network (implemented by the instructions) that is trained to be configured by the key-value parameters, which cause the neural network to perform a signal transformation of the input audio to generate an output audio having an output audio characteristic that corresponds to and matches the target audio characteristic.
[0062] In another embodiment, an apparatus is provided, comprising: a decoder for decoding encoded input audio and encoded key-value parameters to produce input audio and key-value parameters, respectively; and a neural network trained to be configured by the key-value parameters to perform a signal transformation of audio representing the input audio (e.g., the input audio itself or a preprocessed version of the input audio) to produce output audio. The key-value parameters represent target audio characteristics, and the neural network is trained to be configured by the key-value parameters to perform a signal transformation of input audio characteristics of the input audio to output audio characteristics of the output audio matching the target audio characteristics.
[0063] In another embodiment, a method is provided, comprising: receiving input audio and key-value parameters representing target audio characteristics; and configuring a neural network previously trained to be configured by the key-value parameters using the key-value parameters so that the neural network performs a signal transformation of audio representing the input audio to produce output audio having output audio characteristics that match the target audio characteristics.
[0064] In another embodiment, a non-transitory computer-readable medium is provided. The medium is encoded with instructions that, when executed by a processor, cause the processor to perform: receiving input audio and key-value parameters representing target audio characteristics; and configuring a neural network previously trained to be configured by the key-value parameters using the key-value parameters so that the neural network performs a signal transformation of audio representing the input audio to generate output audio having output audio characteristics that match the target audio characteristics.
[0065] While the technology herein has been illustrated and described as embodied in one or more specific examples, it is not intended to be limited to the details shown, since various modifications and structural changes may be made within the scope and equivalents of the claims.
[0066] Each claim presented below represents a separate embodiment, and embodiments that combine different claims and / or different embodiments are within the scope of the disclosure and will be apparent to those of ordinary skill in the art upon reading this disclosure.
Claims
1. A method comprising: receiving input audio and target audio having target audio characteristics, wherein the input audio and the target audio are received as separate signals; estimating key-value parameters representing characteristics of the target audio based on one or more of the input audio and the target audio; and configuring a neural network that is trained to be configured by key-value parameters that cause the neural network to perform a signal transformation of the input audio to produce output audio having output audio characteristics that correspond to and match the target audio characteristics, wherein The estimating includes performing a temporal analysis to generate temporal key-value parameters representing target temporal characteristics of the target audio as the key-value parameters; The configuration includes configuring a neural network with a time key parameter so that the neural network performs a signal transformation as a transformation of a time characteristic of an input audio to a time characteristic of an output audio that matches a target time characteristic, wherein The input audio and the target audio include respective audio frame sequences; Estimating the key-value parameters includes estimating the key-value parameters frame by frame; and Configuring the neural network includes configuring the neural network with the key-value parameters estimated frame by frame, so that the neural network performs signal transformation frame by frame to generate output audio as a sequence of audio frames.
2. The method of claim 1, wherein: The target time characteristic and the time characteristic of the output audio are both their respective time amplitude characteristics.
3. The method of claim 1, wherein: The estimating key-value parameter includes estimating a time key-value parameter representing a time amplitude characteristic of the target audio as the time key-value parameter; and the estimating further includes at least one of the following: Spectral envelope key parameters, including line spectral frequencies (LSF) or LP coefficients (LPC) representing a target spectral envelope of the target audio; and Harmonic key-value parameter, indicating the harmonics present in the target audio.
4. The method of claim 1, wherein: The input audio includes encoded input audio.
5. The method of claim 1, wherein: Key-value parameters include encoded key-value parameters.
6. A device comprising: A decoder for decoding the encoded input audio and the encoded key-value parameters in the bit stream from the transmission channel to generate the input audio and the key-value parameters respectively; A neural network trained to be configured by key-value parameters generated by the decoder to perform a signal transformation of audio representing input audio to produce output audio; wherein: The key-value parameter represents a target temporal audio characteristic as a target audio characteristic; and The neural network is trained to be configured with key-value parameters to perform a signal transformation of input temporal audio characteristics of an input audio to output temporal audio characteristics of an output audio matching target temporal audio characteristics, wherein The audio representing the input audio includes a sequence of audio frames; The key-value parameters include a frame-by-frame key-value parameter sequence that represents the target audio characteristics frame by frame; and The neural network is configured by the frame-by-frame sequence of key-value parameters to perform a signal transformation of audio representing the input audio on a frame-by-frame basis to produce output audio as a sequence of output audio frames. 7 . The apparatus of claim 6 , further comprising a preprocessor for preprocessing the input audio to generate preprocessed input audio as the audio representing the input audio.
8. The device of claim 6, wherein: The audio representing the input audio includes the input audio.
9. The device according to claim 6, wherein: The decoder is further configured to demultiplex the encoded input audio and the encoded key-value parameters from the multiplexed signal and then perform decoding of the encoded input audio and the encoded key-value parameters.
10. The apparatus of claim 6, further comprising: A mixing unit that provides mixing operations to mix the decoded input signal with the output audio generated by the neural network.
11. A method comprising: receiving input audio and key-value parameters representing target audio characteristics in a multiplexed and encoded bitstream in which both the input audio and the key-value parameters are encoded; Demultiplexing and decoding the encoded input audio and the encoded key-value parameters to restore the input audio and the key-value parameters; as well as configuring a neural network previously trained to be configured by the key-value parameters using the decoded key-value parameters so that the neural network performs a signal transformation of the audio representing the input audio to produce output audio having output audio characteristics matching the target audio characteristics, wherein: The key-value parameter represents a target temporal audio characteristic as a target audio characteristic; and The neural network is trained to be configured with key-value parameters to perform a signal transformation of input temporal audio characteristics of an input audio to output temporal audio characteristics of an output audio matching target temporal audio characteristics, wherein The input audio and the audio include respective sequences of audio frames; The key-value parameters represent target audio characteristics on a frame-by-frame basis; and The neural network is configured by key-value parameters to perform signal transformation frame by frame to produce output audio as a sequence of output audio frames. 12 . The method of claim 11 , further comprising preprocessing input audio to generate preprocessed input audio as the audio.
Citation Information
Patent Citations
Split-domain speech signal enhancement
US20190341067A1