Identifying salient features of generative networks
By training an encoder that extracts significant features, the problem of difficulty in reproducing inputs in a noisy environment is solved, and the robustness of the generation network in a noisy environment is realized and the naturalness of the output of the generation network is realized.
Patent Information
- Application Number
- CN201980055674.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-08-02
AI Technical Summary
There are difficulties in generating a network to reproduce input outside the manifold, especially when the noise level rises, which can easily lead to phoneme errors.
By training an encoder that extracts significant features from the input, significant features are more robust and less redundant, which can serve as an improvement to the generation network for noise robustness, resulting in fewer artifacts and more realistic outputs.
The generated network using significant feature conditioning is more robust in noisy environments, and the errors are generated more natural, and it can effectively process the distorted input signal without errors.
Smart Images

Figure CN112639832B_ABST
Abstract
Description
Background Art
[0001] Generative networks such as WaveNet, WaveRNN, and Generative Adversarial Networks have produced very good results in audio / visual synthesis such as speech synthesis and image generation. Such models have the property of being confined to a manifold or topological space. Thus, for example, WaveNet is confined to producing natural speech, i.e., confined to the speech manifold. However, such systems have difficulty reproducing input outside the manifold. For example, non-speech sounds tend to cause phoneme errors as the noise level increases. Summary of the invention
[0002] The embodiment provides an encoder for extracting significant features from input. The significant features are more robust and less redundant, and can act as a boost that makes the generating network robust to noise, which produces fewer artifacts and more realistic outputs. The significant features can also act as an effective compression technique. The embodiment trains the clone encoder to identify the significant features of different but equivalent inputs. The embodiment can also use significant features to condition the generating network. The embodiment does not attempt to approximate a clean signal, so that the generation properties of the network are not limited during conditioning. In order to train the encoder, the embodiment can generate several equivalent signals for a clean input signal. The equivalent signal shares significant features with the clean input signal, but is modified (for example, in some ways) according to the clean signal. As used herein, a modified signal refers to any information added to a clean signal that the designer considers acceptable or any change to a clean signal. Therefore, a modified signal refers to noise, distortion (for example, phase shift, delay of the modified signal), artifacts, information, etc. of a modified clean signal outside the target manifold of the clean signal. In other words, as used herein, a modified signal can be any change to a clean signal that is not considered significant for a human designer. The embodiment trains the encoder to filter out (e.g., ignore) information that is not considered significant from a set of equivalent inputs. The clone encoders sharing weights take different signals as input from a set of equivalent signals. The clone encoders all use a global loss function, which promotes equivalence in the significant features extracted by each clone encoder and independence within the significant features extracted by the individual encoders. In some embodiments, the global loss function may also promote sparsity in the extracted significant features, and / or may promote mapping of the extracted features to a shared target signal. The embodiment may use a set of clone decoders of the mirror encoder to reconstruct the target signal from the extracted significant features. In some embodiments, the system may extract 12 significant signals per input. When used in inference mode, the trained encoder may extract significant features for the input sequence and use the significant features to condition the generation network.
[0003] According to one aspect, a method for identifying features of a generative network includes obtaining a set of inputs for each clean input in an input batch, the set of inputs including at least one modified input, each modified input being a different modified version of the clean input. The method also includes training an encoder with weights to provide features of the inputs by performing the following operations for each set of inputs in the input batch: providing the set of inputs to one or more clone encoders, each clone encoder sharing weights, and each of the one or more clone encoders receiving a different corresponding input in the set of inputs, and modifying the weights to minimize a global loss function. The global loss function has a first term and a second term, the first term maximizing the similarity between features of the set of inputs, the second term maximizing independence and unit variance within features generated by the encoder, the encoder being one of the one or more encoders. The method may include extracting features of a new input using the encoder and providing the extracted features to the generative network. The method may include compressing the features of the new input and storing the features.
[0004] According to one aspect, a method includes receiving an input signal and extracting salient features of the input signal by providing the input signal to an encoder trained to extract salient features. The salient features may be independent and have a sparse distribution. The encoder may be configured to generate nearly identical features from two input signals that a system designer considers equivalent. The method may include conditioning a generative network using the salient features. The method may include compressing the salient features.
[0005] In one general aspect, a computer program product embodied on a computer readable storage device includes instructions that, when executed by at least one processor formed in a substrate, cause the computing device to perform any disclosed method, operation, or process. Another general aspect includes a system and / or method for learning how to identify independent salient features from an input that can be used to condition a generative neural network or for compression, generally as shown in and / or described in conjunction with at least one of the accompanying drawings, and as more fully set forth in the claims.
[0006] One or more embodiments of the subject matter described herein can be implemented to achieve one or more of the following advantages. As an example, an embodiment provides a new type of improvement in which a conditioned generative network is focused on a specific feature in any input - a significant feature. Such conditioning makes the generative network so conditioned robust to noise. As a result, the generative network conditioned using significant features produces more natural errors compared to the errors caused by a generative network that has not been similarly conditioned. In addition, compared to other networks that have not been similarly conditioned, the generative network conditioned using significant features handles more distorted input signals without generating errors. For example, as the noise level in the input signal increases, the generative speech network conditioned using the significant features disclosed herein generates fewer errors, and if any errors are generated, they are also natural sounding errors that are in the speech manifold. In contrast, the artifacts generated by other generative networks deviate from the speech manifold and the sound is more unnatural, and therefore are more easily noticed by the listener as the noise level in the input increases.
[0007] As another example, salient features are compact. For example, an embodiment may extract a small number (e.g., 10 or less, 12 or less, 15 or less, 20 or less) of features from an input, but because the features represent salient and therefore perceptually important information and are independent of each other, the decoder is able to use the salient features to produce realistic output based on the features. In other words, salient features ignore features that are not perceptually relevant but use significant memory allocation. This is in contrast to most conventional encoders such as variational autoencoders or VAEs that do not consider perceptual importance. Salient features are features that can be used for storage or transmission of signals, or for manipulation of signal properties. For example, salient features can be used for robust coding of speech, encoding specific types of images (e.g., faces, handwriting, etc.), changing the identity of the speaker, resynthesizing speech signals without noise, and so on. Although some previous methods have attempted to identify salient features, such previous methods either require explicit knowledge of discarding redundant variables or require equivalent signal pairs for training; such methods are not as scalable as the disclosed embodiments.
[0008] As another example, the disclosed embodiments do not interfere with the generative nature of the generative network, because the embodiments do not attempt to reconstruct the ground truth, i.e., clean speech. Traditional lifting measures used for evaluation use benchmarks without considering other solutions, which are effective for networks like feedforward networks and recurrent neural networks, which try to find a good approximation of the clean speech waveform based on noisy observations and available prior knowledge. In contrast, generative networks (e.g., generative convolutional neural networks, generative deep neural networks, adversarial networks, etc.) are limited to manifold or topological spaces. For example, generative speech networks are limited to producing natural speech, and generative image networks are limited to producing images. Generative networks can use random processes to generate perceptually irrelevant complex details instead of trying to accurately reproduce the input signal. For example, when reproducing an image of leaves, it should look correct (e.g., color, shape), but the leaves will change in number or position compared to the original state. In this example, the number and position are perceptually irrelevant. In other words, the generative network can provide a solution that is quite different from the ground truth but equivalent to it. The current system for lifting the generative network is at least partially optimized to reconstruct the ground truth input, which limits the generation aspect of the generative network. In other words, traditional lift metrics do not take these other solutions into account.Since the disclosed embodiments do not rely on reconstruction of ground truth, the disclosed embodiments do not tend to limit the generative aspects and thus tend to produce more natural or realistic output.
[0009] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 An example system for training a salient feature encoder in accordance with the disclosed subject matter is illustrated.
[0011] Figure 2 An example system for inference in accordance with the disclosed subject matter is illustrated.
[0012] Figure 3 is a flow chart of an example process for identifying and using salient features in accordance with the disclosed subject matter.
[0013] Figure 4 A flow chart of an example process for training an encoder to recognize salient features in accordance with the disclosed subject matter.
[0014] Figures 5A to 5C The benefits provided by the disclosed embodiments are demonstrated.
[0015] Figure 6 An example of a computer device that can be used to implement the described techniques is shown.
[0016] Figure 7 An example of a distributed computer device that can be used to implement the described techniques is shown.
[0017] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION
[0018] Embodiments provide enhancements to generative networks by learning to extract salient features from input signals. Salient features are features shared by signals defined by the system designer as equivalent. The system designer provides qualitative knowledge for feature extraction. This qualitative knowledge enables the encoder to ignore perceptually irrelevant features (e.g., noise, pauses, distortion, etc. that do not affect meaning or content) and only extract those features that affect the content or meaning of the input. In other words, perceptually relevant features are features and other features that affect the ability of humans to grasp the essence of the input. Features that can be perceived but do not affect the essence are perceptually irrelevant. Therefore, salient features are described as being perceptually important to humans. Such features may be small in quantity for each input, but produce better and more realistic reconstructions as perceived by the user.
[0019] Figure 1 is a block diagram of a salient feature extraction system according to an example embodiment. System 100 can be used to train an encoder to extract salient features from an input. Salient features capture perceptually relevant features from an input signal in a scalable manner. Salient features can be used to store or transmit an encoded signal. Salient features can be used to condition a generative network. Salient feature extraction system 100 jointly trains a set of clone encoders 115. Each encoder, e.g., 115(1), 115(2), ..., 115(N), receives as input a different input signal from a set of equivalent signals 110, e.g., 110(1), 110(2), ..., 110(N). The objective function used by the clone encoders 115 encourages the encoders (115(1) to 115(N)) to map their respective inputs to a set of unit variance features that are the same across the clone encoders 115. Training can be supervised or unsupervised. Conventionally, supervised training uses labeled inputs during training. As used herein, supervised training does not refer to such conventional techniques. In contrast, as used herein, supervised training refers to using the reconstructed target term as an additional optimization term during training. Thus, in supervised training, the system 100 includes a clone decoder 125 that maps salient features to a shared target signal. For ease of description, Figure 1 The description of system 100 in FIG. 1 is sometimes described as processing speech input (eg, mel spectrum), but the embodiments are not limited thereto. For example, Figure 1The system 100 can process image input, video input, music input, and the like.
[0020] The salient feature extraction system 100 may be one or more computing devices in the form of a variety of different devices, such as a standard server, a group of such servers, or a rack-mounted server system, etc. In addition, the system 100 may be implemented in a personal computer, such as a laptop computer. The system 100 may be, for example, Figure 6 The computer device 600 depicted in Figure 7 An example of a computer device 700 is depicted in .
[0021] Although Figure 1 , but the system 100 may include one or more processors formed in a substrate configured to execute one or more machine executable instructions or software, firmware, or a combination thereof. The processor may be semiconductor-based—that is, the processor may include semiconductor materials capable of executing digital logic. The processor may be a dedicated processor, such as a graphics processing unit (GPU). The system 100 may also include an operating system and one or more computer memories, such as main memory, configured to store one or more data temporarily, permanently, semi-permanently, or in a combination of the above. The memory may include any type of storage device that stores information in a format that can be read and / or executed by one or more processors. The memory may include volatile memory, non-volatile memory, or a combination thereof, and a storage module that performs certain operations when executed by one or more processors. In some embodiments, the module may be stored in an external storage device and loaded into the memory of the system 100.
[0022] The clone encoder 115 represents a computational model or encoder for multiple machine learning. In machine learning, the computational model is organized as connected nodes, and the nodes are organized into layers. The node performs a mapping function on the input provided to produce a certain output. The first node layer obtains the input provided to the model, that is, the input from an external source. The output of the first node layer is provided as input to the second node layer. The nodes in the second layer provide input to the subsequent layers, and so on, until the final layer is reached. The final or output node layer provides the output of the model. In the case of system 100, the output of the encoder is a feature vector. A vector is usually an array of numbers, where each position in the array represents a different property. The number of array positions is called the dimension of the vector. The value in each array position can be an integer or a decimal. In some embodiments, the value can represent the percentage, probability, possibility, etc. of the existence of the property. In some embodiments, the value can represent the actual value of the property. The layer can be fully connected or partially connected. In a fully connected model, each node in the layer sends its output to each node in the next layer. In a partially connected network, each node in the layer sends its output to some nodes in the next layer.
[0023] The function performed by the node on the input value maps the input to the output. The function uses parameters to perform the mapping. The mapping can be a surjective mapping. The model needs to be trained to determine the parameters, which can start with random values. The parameter is also called weight. For the purpose of this application, the weight can be expressed as ψ. The training process uses an objective function to determine the optimal parameters. The objective function identifies the purpose of the mapping and helps the model modify the parameters through iterative training rounds until the optimal set of parameters is reached. Once the optimal parameters are identified, the model is considered to be trained and can be used in inference mode. In inference mode, the model uses parameters to provide or predict output based on given input. Each machine learning model is trained for a specific task, such as prediction, classification, encoding, etc. The task performed by the computational model is determined by the provided input, mapping function, and desired output.
[0024] exist Figure 1 In the example of , clone encoder 115 includes multiple encoders. Each encoder has its own layer and receives a separate input, but each encoder shares the same set of weights ψ with other encoders. Therefore, during training, weights ψ are adjusted identically for all encoders in clone encoder 115. Since the encoders share weights, they can be called clones. In other words, each encoder in the clone encoder effectively represents the same encoder and will produce the same features when given the same input. However, different inputs are given during training each encoder in clone encoder 115, but each input is considered by the system designer to be substantially equivalent. The encoders of clone encoder 115 use the same objective function.
[0025] The input provided to the clone encoder 115 represents an equivalent signal. In some embodiments, an equivalent signal is a signal considered equivalent by a system designer. Therefore, a system designer can select the type of modification performed on a clean input. In this regard, a system designer can supervise the generation of a set of equivalent inputs for processing. A clean input is any input that is modified to generate a set of equivalent signals. Generally speaking, a clean input represents an original file, an expected output, and the like. The set of equivalent signals 110 represents different modifications performed for a clean signal. As used herein, a modification may include any modification to a clean signal that a designer considers acceptable or information added to a clean signal. The modified signal may include noise or artifacts added to a clean signal. The modified signal may include clean signal distortion, such as all-pass filtering (i.e., relative delays, phase shifts, etc. of different frequency bands). The modified signal may include information outside the target manifold added to a clean signal. In other words, a modification may be any modification performed for an input that is not considered significant by a human designer. For example, if the clean input 110(1) is a speech sample, the modified input 110(2) may be the same speech sample with traffic noise added, the modified input 110(3) may be the relative delay of the frequency band of the speech signal, the modified input 110(4) may be the same speech sample with restaurant noise added, the modified input 110(5) may be the same speech sample with reverberation added, the input 110(6) may be the same speech sample with microphone distortion added, and so on.
[0026] In some embodiments, the system designer decides what kind of modification to make and then uses the modification engine 105 to apply the modification to the clean signal and generate a set of equivalent signals 110. The modification engine 105 can use the training data 103 as a source of clean data. The modification engine 105 can use the training data 103 as a source of one or more modifications, such as to generate one or more of the modified inputs. For example, the training data can be a data set of clean signals, such as a data set of images, a data set of voice data spoken by a local speaker, a professionally recorded musical work, and the like. In some embodiments, the modification engine 105 can be configured to provide input within a signal-to-noise ratio (SNR) range. For example, if the signal noise provided to the clone encoder 115 is too much / modified too much, the objective function of minimizing the reconstruction error may promote the removal of significant properties (perceived by humans and related to understanding and / or quality). On the other hand, if the noise is insufficient, the significant features may be more sensitive to noise. The number N of signals in the set of equivalent signals 110 depends on the implementation. Generally, this number depends on the type of modification that can be applied to the input and the processing power and training time of the hardware of the system 100. In some implementations, the choice of N is a trade-off between performance and quality. In some implementations, the set of equivalent signals 110 may include 32 inputs (eg, N=32).
[0027] Each encoder 115(1)-115(N) in clone encoder 115 receives a different equivalent signal from the set of equivalent signals 110. Clone encoder 115 includes an encoder for each different equivalent signal. For example, encoder 115(1) may receive clean input 110(1), while encoder 115(2) may receive first modified input 110(2), and so on.
[0028] Each of the encoders in the clone encoder 115 uses an objective function to learn parameters (weights) that enable the encoder to extract a feature vector that is similar across all encoders despite different inputs and includes the information needed to reconstruct a representation of the clean input. The resulting extracted feature vectors - salient features - generally lack information related to modifications and are therefore robust to modifications. In an embodiment of unsupervised learning, the objective function includes two terms. The first term promotes similarity across salient features, and the second term promotes independence and unit variance. The second term can also promote sparsity. In some embodiments, the objective function can include a third term that encourages the clone encoder 115 to find salient features that are mapped to a shared target signal via a decoder, such as minimizing decoder loss. In some embodiments, the decoder can be one of a set of clone decoders. In some embodiments, the second and optional third terms can be weighted. In some embodiments, the objective function can be expressed as Dglobal =D E +λ MMD D MMD +λ D D D , where D global is the global loss that training tries to minimize, D E is the first term that maximizes the similarity between the salient features output by the clone encoder, D MMD is the second term that maximizes independence, unit variance, and optionally sparsity, λ MMD is the weighting factor applied to the second term, D D is the third term that facilitates the reconstruction of the target signal, and λ D is the weighting factor applied to the third term. The goal of training is to determine the equation that makes D global Weights as close to zero as possible. In some embodiments, the system 100 defines the global loss as the expectation on the data distribution. In some embodiments, the system 100 defines the global loss as the average value over the observed batches of m data. In other words, training occurs on batches of m clean inputs and their corresponding sets of equivalent signals 110. The system 100 can use stochastic gradient descent on batches of m data points to optimize on the empirical distribution associated with the training data 103. The training data 103 can be any data that matches the input to be used in the inference mode. In the field of speech, the input data can be a block or frame of speech. For example, the training data can be 40ms of speech. The training data 103 can be a block of speech (e.g., 20ms, 40ms) converted to a spectral representation on the mel frequency scale. In the field of images, the input data can be a block of pixel data. For example, the training data can be a 32×32 pixel block. The training data 103 can be 40ms of music, and so on.
[0029] The first term of the objective function maximizes the similarity among the salient features generated by the encoders in the clone encoders 115. The similarity between the salient features extracted by different encoders in the clone encoders 115 can be maximized by minimizing the L2 norm (or square root of the square) between the salient features generated by the first clone and the salient features generated by the remaining clones (to reduce the computational workload) or between the salient features generated by each clone and the remaining clones. For example, in some embodiments, the first term can be expressed as where m is the number of inputs (sets of equivalent signals 110), N is the number of encoders in clone encoder 115, ‖·‖ is the L2 norm, and z represents the salient features extracted by the encoder, while i marks the specific features. In other words, as an example, for one of the m sets of equivalent signals 110, encoder 115(1) may extract feature z(1) , encoder 115(2) can extract feature z (2) , etc. The salient feature z from the above expression is (1) is called the datum feature, and z is extracted from its corresponding input (1) The encoder of clone encoder 115 is the reference encoder. It is to be understood that any encoder in clone encoder 115 can be selected as the reference encoder. Therefore, embodiments are not limited to comparing a salient feature from a clean signal (i.e., salient feature 150(1)) with other salient features (i.e., salient features 150(2) to 150(N)). Rather, any one of salient features 150(1) to 150(N) can be selected as the reference feature. Some embodiments compare only the reference feature with the remaining salient features to reduce computational workload.
[0030] In some embodiments, the system 100 may compare all significant features of the set of equivalent signals 110 with all other significant features of the set of equivalent signals 110. In such embodiments, each encoder becomes a reference encoder. In one embodiment, a subset, such as two, three, or five of the encoders, may be selected as reference encoders. In some embodiments, the system minimizes the 1-norm instead of the L2-norm. The L2-norm reduces larger differences, which is advantageous when determining significant features, but embodiments may use the 1-norm. The first term of the objective function may be calculated by the equivalence optimizer 130. The equivalence optimizer 130 may be configured to select features that describe the information components shared between the signals in the set of equivalent signals (e.g., inputs 110(1) to 110(N)) with relatively high fidelity. In some embodiments, the equivalence optimizer 130 may calculate the 1-norm. The equivalence optimizer 130 may determine which weight Ψ to adjust to minimize the difference in significant features generated for the set of equivalent signals.
[0031] The second term of the objective function that maximizes the independence and variance of the significant features can be determined using any method that promotes / forces the significant features to have a specified distribution. In other words, the system can force a specified distribution on the features. The specified distribution promotes independence and a given variance. Examples of such methods include chi-square test, earth-moving distance re-expressed via Kantorovich-Rubinstein duality, and maximum mean deviation (MMD). MMD measures the distance between two distributions and requires a desired (specified) distribution for comparison. Example distributions include Gaussian, uniform, normal, Laplace, etc. The selection of the distribution determines sparsity. Gaussian and normal distributions do not promote sparsity. Laplace distribution promotes sparsity. In a sparse vector, most dimensions have low values (e.g., values close to zero) and only some have large values. In some embodiments, the system 100 can promote that the significant features are sparse, such as having large values of 10, 12, 15, etc. in the dimensions. The dimension of a vector is the number of different properties it represents.
[0032] When using the MMD metric, the second term of the objective function can be expressed as Where m is the number of batches (eg, the number of sets of different equivalent signals 110), y i From the chosen distribution, z i is the set of salient features generated by the reference encoder, and k(·,·) is a kernel of desired size and shape. In some embodiments, the kernel is a multivariate quadratic kernel. In some embodiments, when the MMD metric is used, the second term can be expressed as In order for MMD to perform correctly, m must be large enough and depend on the desired accuracy. For example, m can be in thousands. The benchmark encoder of the second item of the objective function can be the encoder identical with the benchmark encoder for the first item of the objective function. The benchmark encoder of the second item can be different from the benchmark encoder for the first item. In some embodiments, system 100 can use more than one benchmark encoder. For example, the system can compare the significant features extracted by two, three, five or other different encoders with the distribution. The second item of the objective function can be calculated by independence optimizer 140, which is configured to measure independence and variance summarized above, and determines which weight Ψ to adjust to, for example, minimize the difference between distribution and the selected distribution. In some embodiments, the second item can be weighted.
[0033] The third term of the objective function is an optional item that promotes the mapping of different significant features to a shared target signal. In other words, the third term involves the reconstruction of the target signal. The target signal can be derived from a clean signal (e.g., clean input 110 (1)). The target signal may not attempt to directly approximate the clean signal. In some embodiments, the target signal can characterize the short-term spectral characteristics of the clean signal. Such an embodiment reduces the use of computing resources. In some embodiments, the target signal can be one of the equivalent modified inputs. In some embodiments, the target signal can be a representation of a clean signal with an appropriate standard. For example, in the field of speech, a mel spectrum representation of a clean speech signal with an L2 norm (i.e., squared error) can be used as the target signal. As another example, in the field of music, a mel spectrum representation of a clean music signal with an L2 norm can be used as the target signal. As another example, in the field of images, a wavelet representation of a selected resolution can be used as the target signal. In some embodiments, the target signal can be a clean signal.
[0034] System 100 can use clone decoder 125 to reproduce the target signal based on the salient features. In some embodiments, clone decoder 125 can have the same number of decoders as encoders in clone encoder 115. In such embodiments, each decoder in clone decoder 125 receives a different one of the salient feature vectors, e.g., decoder 125(1) receives salient features 120(1) as input, decoder 125(2) receives salient features 120(2) as input, and so on. Figure 1 In some embodiments (not shown in the figure), the clone decoder 125 may include a single decoder. In some embodiments, the clone decoder 125 may have fewer clone decoders than the clone encoder. The clone decoder takes the salient features from the encoder as input and maps the salient features to the output. The output represents a reconstruction of the input provided to the encoder. In some embodiments, the configuration of the decoder may mirror the configuration of the encoder. For example, the encoder may each use two long short-term memory (LSTM) node layers and a fully connected node layer, each layer having 800 nodes, and each clone encoder may use a fully connected layer followed by two LSTM layers, each layer having 800 nodes. As another example, the encoder may use four LSTM layers and two fully connected node layers, and the decoder mirrors the configuration. In some embodiments, the decoder may not mirror the configuration of the encoder. The encoder may not be limited to these exact configurations and may include feed-forward layers, convolutional neural networks, ReLU excitations, and the like.
[0035] The clone decoder 125 may include any suitable decoder. The clone decoder 125 may be optimized by using a loss function with an L2 norm. In some embodiments, the third term may be expressed as where m is the number of sets of equivalent signals, N is the number of cloned encoders, is the salient feature vector generated by one of the cloned encoders, is a suitable signal representation of the clean input (i.e., the target signal) with dimension P, f φ : is a cloned decoder network 125 with learned parameters φ, which will have dimension Q (e.g., ) to a vector of dimension P (e.g., v i ), and where the summation is over a set of m equivalent inputs. In some embodiments, the number of clone encoders (N) used to calculate the third term can be less than the number of clone encoders 115. In some embodiments, the third term can be weighted. In some embodiments, the weight of the third term can be higher than the weight of the second term. For example, when the weight of the second term is 1.0, the weight of the third term can be 18.0. In some embodiments, the weight of the third term can be zero. In such an embodiment, the training is unsupervised. The third term of the objective function can be calculated by a decoder loss optimizer 150, which is configured to measure the similarity of the decoded output to the target signal as outlined above and determine which weights Ψ to adjust to, for example, minimize the difference between the reconstructed signal and the target signal. The parameter φ can also be learned during training.
[0036] The system 100 may include or communicate with other computing devices (not shown). For example, other devices may provide a collection of training data 103, modification engine 105, and / or equivalent signals 110. In addition, the system 100 may be implemented in multiple computing devices that communicate with each other. Therefore, the salient feature extraction system 100 represents an example configuration, and other configurations are possible. In addition, the components of the system 100 may be combined or distributed in a manner different from that shown.
[0037] Figure 2 An example system 200 for inference in accordance with the disclosed subject matter is illustrated. System 200 is an example of how salient features may be used. In the example of system 200, salient features are used to conditionally generate network 225. System 200 is a computing device or multiple devices in the form of multiple different devices, such as a standard server, a group of such servers, a rack-mounted server system, two computers communicating with each other, and so on. In addition, system 200 may be implemented in a personal computer, such as a desktop or laptop computer. System 200 may be such as Figure 6 The computer device 600 shown or Figure 7 An example of a computer device 700 is shown.
[0038] Although Figure 2 , but the system 200 may include one or more processors formed in a substrate, which are configured to execute one or more machine executable instructions or multiple software, firmware, or a combination thereof. The processor may be semiconductor-based—that is, the processor may include semiconductor materials capable of executing digital logic. The processor may be a dedicated processor, such as a graphics processing unit (GPU). The system 100 may also include an operating system and one or more computer memories, such as main memory, configured to store one or more data temporarily, permanently, semi-permanently, or in a combination of the above. The memory may include any type of storage device that stores information in a format that can be read and / or executed by one or more processors. The memory may include volatile memory, non-volatile memory, or a combination thereof, and stores modules that perform certain operations when executed by one or more processors. In some embodiments, the module may be stored in an external storage device and loaded into the memory of the system 200.
[0039] System 200 includes encoder 215. Encoder 215 represents Figure 1 One of the encoders in the clone encoder 115 of . In other words, encoder 215 is a trained encoder having weights Ψ that have been optimized to produce salient features 220 based on a given input 210. Since the encoders in the clone encoder 115 share the same weights Ψ, they are the same encoder and any encoder (115(1) to 115(N)) can be used as encoder 215 in inference mode. In other words, since the encoder uses weights to map inputs to outputs, only one set of weights is used in the clone encoder 115, with the weights Ψ representing the encoder. Therefore, the weights Ψ determined by system 100 enable encoder 215 to map input 210 to salient features 220. Input 210 is the same as that used to train Figure 1The clone encoder 115 generates a signal of the same format as the signal of the clone encoder 115. For example, the input 210 can be a mel frequency block of 40ms speech. The encoder 215 maps the input 210 to salient features 220. The system 200 can provide the salient features 220 to the generation network 225. In some embodiments, the system 200 can compress and / or store the salient features 220 and transmit the features 220 to the generation network 225. The generation network 225 then uses the salient features 220 and the input 210 for conditioning. Conditioning is a method of providing features to the network 225 to produce specific characteristics. The salient features 220 provide better features for conditioning. In the example of the system 200, the generation network 225 is conditioned to focus on the salient features of the input 210, which makes the network 225 robust to modifications such as noise and distortion. Although Figure 2 2, but the system 200 in inference mode processes many different inputs as inputs 210. For example, the system 200 may split a long audio recording into frames, such as 20ms or 40ms, and process each frame in the recording as a separate input 210, providing corresponding salient features 220 for each input 210 for conditioning.
[0040] System 200 is one example use of salient features 220, but salient features may be used in other ways. For example, salient features 220 may be used to store or transmit data in a compressed format. When salient features 220 are sparse, they may be compressed to a significantly smaller size and transmitted using a smaller bandwidth than the original signal. A decoder such as that used in training an encoder may regenerate the compressed salient features. In addition, although the invention has been generally described with respect to speech pair Figure 1 and Figure 2 Discussion is made, but the embodiments are not limited thereto. For example, the embodiments can be adapted for use with input from files of images, music, videos, etc.
[0041] Figure 3 is a flow chart of an example process for identifying and using salient features in accordance with the disclosed subject matter. Process 300 may be performed by, for example, Figure 1 System 100 and Figure 2The process 300 can be performed by a significant feature system such as the system 200 of . The process 300 can be started by obtaining a set of inputs for a batch of clean inputs (305). The clean inputs can come from a database, such as a database of hundreds of hours of speech, a collection of image libraries, etc. The clean inputs do not need to be clean in the conventional sense, only signals for which equivalent signals are generated. In this sense, the clean signal represents a data point in a batch of data points. For each clean input, the system also obtains multiple equivalent inputs. The system designer can choose the type of modification made to the clean input to obtain the equivalent input. The system thus obtains a set of equivalent inputs. In some embodiments, the set may include a clean input and multiple equivalent inputs. In some embodiments, the set may not include a clean input, but is still referred to as being associated with a clean input. The number of equivalent inputs depends on the implementation and is a trade-off between training time, computing resources, and accuracy. In general, less than 100 equivalent inputs are used. In some embodiments, less than 50 equivalent inputs can be used. In some embodiments, less than 20 equivalent inputs can be used. Each input in the set of equivalent inputs is based on a clean input. For example, different artifacts can be added to the clean input. Different distortions can be made to the clean input. Different noises can be added to the clean input. In general, the modified input is any modification made to the clean input, but is still considered equivalent to the clean input, for example, in terms of content and understanding. The clone encoder learns to ignore this extra information. The system can train a collection of clone encoders, i.e., multiple encoders that share weights, to extract salient features from a collection of equivalent inputs (310). This process is about Figure 4 Described in more detail. Once the clone encoder is trained, i.e., the set of optimized weights is determined, the weights represent the trained encoder, which is also referred to as a salient feature encoder. The system uses the salient feature encoder (i.e., using the optimized weights) to extract salient features for the input (315). The input is of a type similar to the training input used in step 305. The input is also referred to as a conditioned input signal. In some embodiments, the input signal can be parsed into several inputs, such as multiple time series of an audio file, multiple pixel blocks of a video file, and so on. Each parsed component, such as each time series, can be used as a separate input by the system. The system can use the salient features and the conditioned inputs to condition the generation network (320) or perform compression (325). In some embodiments, the system can compress and store the salient features (325) before transmitting the features to the generation network for conditioning (320). Although in Figure 3315 for one input, but the system can repeat step 315 and any of steps 320 or 325 any number of times. For example, the system can repeatedly perform step 315 for frames in a video or audio file or for blocks of pixels in an image file. Therefore, it is to be understood that process 300 includes repeating step 315 and any of steps 320 or 325 with different inputs as needed. Therefore, this portion of process 300 can be used on time series, large images, etc.
[0042] Figure 4 is a flow chart of an example process 400 for training an encoder to recognize salient features in accordance with the disclosed subject matter. Process 400 can be performed by, for example, Figure 1 The process 400 may be performed by the system 100. The process 400 may start with one of the clean inputs (405). The clean input represents a training data point. The training data may be a frame of an audio or video file, an audio or video file of a specified time (e.g., 20ms, 30ms, 40ms), a pixel block of a specified size from an image, etc. The clean input is associated with a set of equivalent inputs. As described with respect to Figure 3As discussed in step 305 of , the set of equivalent inputs includes clean inputs and one or more modified inputs. The system provides each encoder in the set of clone encoders (multiple clone encoders) with a corresponding input from the set of equivalent inputs (410). Therefore, each encoder in the set receives different inputs from the set of equivalent inputs. The clone encoders share weights. Each encoder provides an output based on the shared weights (415). The output of the encoder represents the salient features of the corresponding input. The system can repeat the process of generating salient features for different clean inputs (420, yes) until the clean input batch has been processed (420, no). The batch can have a size (e.g., m) sufficient to enable the distribution measurement to be performed correctly. The system can then adjust the shared weights to minimize the global loss function, which maximizes equivalence, independence, variance, and optionally sparsity and / or signal reconstruction (425). The global loss function has a first term that maximizes the similarity of the salient features extracted by each encoder for the set of equivalent inputs. Therefore, each encoder is facilitated to extract the same features as other encoders for a given set of equivalent inputs. The global loss function has a second term that maximizes independence and variance. The second term forces the salient features to have a specific distribution. The distribution can be a sparse distribution, for example, promoting sparsity of the salient features. The second term favors disentangled features. Some embodiments may include a third term in the global loss function that ensures that the salient features can be mapped to a target input. The target input can be derived from a clean input for a set of equivalent inputs. The target input can be a clean input. The target input can be any input in the set of equivalent inputs. The objective function is about Figure 1 The system repeats process 400 with the newly adjusted weights until convergence, e.g., the weights result in a mapping that minimizes the objective function to an acceptable degree, or until a predetermined number of training iterations have been performed. When process 400 is complete, the optimal weights for the encoder have been determined and can be used in inference mode.
[0043] Figures 5A to 5C The benefits provided by the disclosed embodiments are demonstrated. Figure 5A is a graph illustrating listening test results for various implementations compared to conventional systems. Figure 5AIn the example of , the training database includes 100 hours of speech and 200 speakers and a mixed corpus of noise. The mixed corpus of noise includes static and non-static noise from about 1000 recordings captured in various environments including busy streets, coffee shops and swimming pools. The input to the encoder in the cloning encoder is a set of equivalent inputs including 32 different versions of the signal containing speech. The set of equivalent inputs includes clean speech and versions with noise added from 0 to 10dB signal-to-noise ratio (SNR).
[0044] exist Figure 5A In the example of , the signal is preprocessed into an oversampled logarithmic mel spectrum representation. In some embodiments, a single window (sw) method is used. The single window (sw) method uses 40ms, which has a time offset of 20ms and a resolution of 80 coefficients for each time offset. In some embodiments, a double window (dw) method is used. In the double window (dw) method, each 20ms offset is associated with a 40ms window and two 20ms windows, and the two 20ms windows are located at 5-25 and 15-45ms of the 40ms window. The 20ms window is described using 80 logarithmic mel spectrum coefficients, for a total of 240 coefficients (dimensions) for each 20ms offset. The implementation marked as SalientS uses a decoder (supervised) during training. The implementation marked as SalientU uses unsupervised training. When using a decoder (supervised training), the mel spectrum of the clean signal is used as the target signal. Each implementation (e.g., supervised / unsupervised, sw / dw) is used to condition different WaveNets. The conditioned WaveNet was provided with clean (-clean) and noisy (-noisy) inputs, and the output was evaluated using a MUSHRA-like listening test.
[0045] Figure 5A The conventional system illustrated as a baseline in uses a feature set based on principal component analysis (PCA) that extracts 12 features from a 240-dimensional vector of dual window (dw) data. The PCA is calculated for the signal used as input to the clone encoder during training. The PCA that extracts four features is also illustrated.
[0046] Figure 5AThe diagram shows that WaveNet conditioned using the disclosed embodiments is more robust to noise, significantly outperforming the baseline system. More specifically, unsupervised learning (SalientU-sw) using a single window provides good speaker identification for natural speech quality, but errors for short duration phonemes are quite frequent. The number of errors for clean (SalientU-sw-clean) input signals is lower than for noisy (SalientU-sw-noisy) input signals. Supervised learning reduces the error of noisy input for clean input. For noisy input, the error is further reduced by using a double window, almost reaching the quality obtained with clean input.
[0047] Figure 5B and 5C is a graph illustrating a comparison of the disclosed embodiments with other benchmark systems. Figure 5B Comparisons between various embodiments and SEAGAN are illustrated. Figure 5C Figure 2 shows the comparison between various implementations and the denoising WaveNet. Figure 5B and 5C In the example, two sizes of models are used. The first implementation (SalientS and SalientP) has two layers of LSTM units and one fully connected layer, each with 800 nodes. The second implementation (SalientL) has 4 LSTM layers and 2 layers of fully connected nodes, also with 800 nodes per layer. Figure 5B and 5C All implementations illustrated in use supervised training, where the decoder mirrors the encoder configuration. Figure 5B and 5C In the example implementation, the training process is a sequence of 16kHz input from six 40ms double-window mel frequency frames overlapped by 50%. For a total of 240 mel frequency bins per frame, each double-window frame includes 80 mel frequency bins from a 40ms window and two 20ms windows (located at 5-25 and 15-45ms of the 40ms window). The clonal encoder outputs 12 values per frame as significant features (e.g., 12 significant features per frame) using linear excitation. Clean speech is used for teacher forcing and negative log-likelihood loss, and significant features are inferred on the complete speech from the training set to create conditioned training data for WaveNet. SalientL and SalientS examples are trained on the VoiceBank-DEMAND speech enhancement dataset (provided by Valetini et al.). In addition, the SalientP example is pre-trained on the WSJ0 dataset and switched to VoiceBank-DEMAND in the middle of training.
[0048] Figure 5B It is shown that in listening tests (MUSHRA-like), embodiments match or exceed the performance of SEGAN. Figure 5C The illustrated embodiment outperforms the denoised WaveNet at all SNR ranges. SEGAN and denoised WaveNet are examples of conventional systems that develop generative networks but are at least partially optimized to reconstruct ground truth waveforms. Figure 5B and 5C We demonstrate that implementations that do not attempt to reconstruct the ground truth outperform such approaches.
[0049] Figure 6 An example of a general computer device 600 that can be used with the techniques described herein is shown. Figure 1 System 100 or Figure 2 The system 200 of the present invention is shown in Figure 20. The computing device 600 is intended to represent various example forms of computing devices, such as laptops, desktops, workstations, personal digital assistants, cellular phones, smart phones, tablet computers, servers, and other computing devices, including wearable devices. The components shown herein, their connections and relationships, and their functions are intended only as examples and are not intended to limit the implementation of the inventions described and / or claimed herein.
[0050] The computing device 600 includes a processor 602, a memory 604, a storage device 606, and an expansion port 610 connected via an interface 608. In some embodiments, the computing device 600 may include, among other components, a transceiver 646, a communication interface 644, and a GPS (global positioning system) receiver module 648 connected via the interface 608. The device 600 may communicate wirelessly, if necessary, through the communication interface 644, which may include digital signal processing circuitry. Each of the components 602, 604, 606, 608, 610, 640, 644, 646, and 648 may be mounted on a common motherboard, or in other suitable ways.
[0051] The processor 602 can process instructions for execution within the computing device 600 to display graphical information for a GUI on an external input / output device such as a display 616, including instructions stored in the memory 604 or on the storage device 606. The display 616 can be a monitor or a flat touch screen display. In some embodiments, multiple processors and / or multiple buses, as well as multiple memories and multiple types of memories, can be used if appropriate. Moreover, multiple computing devices 600 can be connected with each device providing a portion of the necessary operations (e.g., as a server group, a blade server group, or a multi-processor system).
[0052] The memory 604 stores information within the computing device 600. In one embodiment, the memory 604 is one or more volatile memory units. In another embodiment, the memory 604 is one or more non-volatile memory units. The memory 604 may also be other forms of computer-readable media, such as a disk or an optical disk. In some embodiments, the memory 604 may include an extended memory provided by an expansion interface.
[0053] Storage device 606 can provide mass storage for computing device 600. In one embodiment, storage device 606 can be or can include computer readable media, such as floppy disk devices, hard disk devices, optical disk devices or tape devices, flash memory or other similar solid-state memory devices, or device arrays, including devices in storage area networks or other configurations. A computer program product can be presented in such a computer readable medium in a tangible manner. The computer program product can also include instructions that, when executed, implement one or more methods such as those described above. The computer or machine readable medium is a storage device such as memory 604, storage device 606, or a memory on processor 602.
[0054] The interface 608 can be a high-speed controller that manages bandwidth-intensive operations of the computing device 600 or a low-speed controller that manages less bandwidth-intensive operations, or a combination of such controllers. An external interface 640 can be provided to enable the device 600 to communicate with other devices in a near-area manner. In some embodiments, the controller 608 is coupled to the storage device 606 and the expansion port 614. The expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or a router, for example, through a network adapter.
[0055] As shown, computing device 600 can be implemented in various forms. For example, it can be implemented as a standard server 630, or multiple implementations in a group of such servers. It can also be implemented as a part of a rack server system. In addition, it can be implemented in a personal computer such as a laptop 622 or a smart phone 636. The entire system can include multiple computing devices 600 that communicate with each other. Other configurations are possible.
[0056] Figure 7 An example of a general computer device 700 that can be used with the techniques described herein is shown. Figure 1 System 100 or Figure 2The computing device 700 is intended to represent various example forms of large-scale data processing devices, such as servers, blade servers, data centers, mainframes, and other large-scale computing devices. The computing device 700 may be a distributed system with multiple processors, which may include network-attached storage nodes interconnected by one or more communication networks. The components shown herein, their connections and relationships, and their functions are intended only as examples and are not intended to limit the embodiments of the inventions described and / or claimed herein.
[0057] Distributed computing system 700 may include any number of computing devices 780. Computing devices 780 may include servers or rack-mounted servers, mainframes, etc. communicating via a local or wide area network, dedicated optical links, modems, bridges, routers, switches, wired or wireless networks, etc.
[0058] In some embodiments, each computing device may include multiple racks. For example, computing device 780a includes multiple racks 758a-758n. Each rack may include one or more processors, such as processors 752a-752n and 762a-762n. The processor may include a data processor, a network-attached storage device, and other computer-controlled devices. In some embodiments, a processor may operate as a master processor and control scheduling and data distribution tasks. The processors may be interconnected through one or more rack switches 758, and the racks may be connected through switches 778. Switches 778 may handle communications between multiple connected computing devices 700.
[0059] Each rack may include memory such as memory 754 and memory 764, and storage such as 756 and 766. Storage 756 and 766 may provide mass storage and may include volatile or non-volatile storage, such as an attached network disk, a floppy disk, a hard disk, an optical disk, a tape, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. Storage 756 or 766 may be shared between multiple processors, multiple racks, or multiple computing devices, and may include a computer-readable medium storing instructions that can be executed by one or more processors. Memories 754 and 764 may include, for example, one or more volatile memory units, one or more non-volatile memory units, and / or other forms of computer-readable media such as magnetic or optical disks, flash memory, cache, random access memory (RAM), read-only memory (ROM), and combinations thereof. Memories such as memory 754 may also be shared between processors 752a-752n. Data structures such as indexes may be stored across storage 756 and memory 754, for example. The computing device 700 may include other components not shown, such as a controller, a bus, input / output devices, a communication module, and the like.
[0060] An overall system such as system 100 may include multiple computing devices 700 in communication with each other. For example, device 780a may communicate with devices 780b, 780c, and 780d, and these may be collectively referred to as system 100. As another example, Figure 1 The system 100 may include one or more computing devices 700. Some computing devices may be located close to each other geographically, while others may be geographically far away. The layout of the system 700 is only an example and the system may adopt other layouts or configurations.
[0061] According to one aspect, a method for identifying features of a generative network includes obtaining a set of inputs for each clean input in an input batch, the set of inputs including at least one modified input, each modified input being a different modified version of the clean input. The method also includes training an encoder with weights to provide features of the inputs by performing the following operations for each set of inputs in the input batch: providing the set of inputs to one or more clone encoders, each clone encoder sharing weights, and each of the one or more clone encoders receiving a different corresponding input in the set of inputs, and modifying the weights to minimize a global loss function. The global loss function has a first term and a second term, the first term maximizing the similarity between features of the set of inputs, the second term maximizing independence and unit variance within features generated by the encoder, the encoder being one of the one or more encoders. The method may include extracting features of a new input using the encoder and providing the extracted features to the generative network. The method may include compressing the features of the new input and storing the features.
[0062] These and other aspects may include, alone or in combination, one or more of the following. For example, providing the extracted features to the generation network may include decoding the compressed features and / or transmitting the features. As another example, the first term may measure the difference between a feature extracted from a first input of a set of inputs and a feature extracted from each remaining input of the set of inputs. As another example, the first term may measure the difference between a feature extracted from each input of a set of inputs and a feature extracted from each remaining input of the set of inputs. As another example, the second term may minimize the maximum mean deviation between a specified distribution and a distribution of the first feature extracted by the encoder over a batch of inputs. In some embodiments, the specified distribution is a Laplace distribution. In some embodiments, the specified distribution is a Gaussian distribution. In some embodiments, the second term may be expressed as where M is the batch size, k(·,·) is the kernel, z represents the salient features of the input in the input set, and y is obtained according to the specified distribution.
[0063] As another example, the second term can further maximize the sparsity within the extracted features. As another example, the global loss function has a first term, a second term, and a third term that maximizes the similarity between the decoded input and the target input, which is a decoded version of the extracted features. In some embodiments, maximizing the similarity between the decoded input and the target input includes providing features generated by a decoder in one or more clone decoders to the decoder; and adjusting the weights of the decoder to match the target input. In some embodiments, the target input can characterize short-term spectral characteristics of a clean input associated with a set of inputs. In some embodiments, the decoder is one of one or more clone decoders, and the clone decoder has a one-to-one relationship with the clone encoder.
[0064] According to one aspect, a method includes receiving an input signal and parsing the input signal into a plurality of time series. The method also includes extracting features of the time series by providing the time series to an encoder that is one of one or more clone encoders and is trained to extract features for each of the multiple time series. The clone encoders share weights and minimize a global loss function during training. The global loss function maximizes the similarity between the features output by each of the clone encoders and maximizes independence, unit variance, and sparsity within the features generated by the encoder. The method may include compressing the extracted features. The method may include transmitting and / or storing the extracted features. The storage and / or transmission may be of compressed features. The method may include using feature conditioning to generate a network. The method may include decompressing the features and using feature conditioning to generate a network.
[0065] These and other aspects may include one or more of the following, either individually or in combination. For example, during training, each encoder in the clone encoder may receive a corresponding input signal from a set of signals, each input signal from the set of signals representing a clean input signal or a different modification of a clean input signal. As another example, independence, unit variance, and sparsity may be maximized using a forced Laplace distribution. As another example, a conditioned generative network may produce a speech waveform and a time series may include a mel frequency slot. In some such embodiments, the conditioning results in a generative network that produces artifacts in a speech manifold. As another example, a conditioned generative network may produce an image.
[0066] According to one aspect, a method includes obtaining a set of inputs for each clean input in an input batch, the set of inputs for the clean inputs including at least one modified input, each modified input being a different modified version of the clean input. The method also includes training an encoder with weights to generate features for each set of inputs in the input batch. The training includes, for each set of inputs, providing corresponding inputs from the set of inputs to one or more clone encoders, the clone encoders sharing weights, wherein each of the one or more clone encoders receives a different corresponding input in the set of inputs, and modifying the shared weights to minimize a global loss function. The global loss function has a first term and a second term, the first term maximizing the similarity between features generated by the clone encoders for the set of inputs, the second term maximizing independence and unit variance within the features generated by the encoder, the encoder being one of the one or more clone encoders. The method may also include extracting salient features of a new input using the encoder.
[0067] These and other aspects may include one or more of the following, alone or in combination. For example, the method may include compressing the salient features extracted for the new input, and storing the compressed salient features extracted for the new input as compressed input. In some embodiments, storing the features may include transmitting the compressed features to a remote computing device that stores the compressed features. In some embodiments, the remote computing device decompresses the features and uses a decoder to generate a reconstructed input based on the decompressed features. As another example, the global loss function may include a third term that minimizes the error in the reconstructed target input. In such an embodiment, the method may further include providing the features generated by the cloned encoder to the decoder and adjusting the shared weights of the encoder using the target input to minimize the reconstruction loss.
[0068] According to one aspect, a method includes receiving an input signal and extracting salient features of the input signal by providing the input signal to an encoder trained to extract salient features. The salient features are independent and have a sparse distribution. The encoder can be configured to generate almost identical features based on two input signals that the system designer considers equivalent. The method also includes conditioning the generative network using the salient features. In some embodiments, the method also includes extracting multiple time series from the input signal and extracting salient features for each time series.
[0069] According to one aspect, a non-transitory computer-readable storage medium stores weights of a salient feature encoder, which has been determined by a training process including an operation, the operation including minimizing a global loss function having a first term and a second term for an input batch using a clone encoder with shared weights, the input batch being a set of multiple equivalent inputs, each clone encoder receiving a different one of the set of equivalent inputs. The first term of the global loss function maximizes the similarity between the outputs of the clone encoders. The second term of the global loss function maximizes the independence and variance between the features in the output of at least one clone encoder. The set of equivalent inputs may include clean inputs and different modified versions of the clean inputs. These and other aspects may include one or more of the following, alone or in combination. For example, the output of each clone encoder may be a sparse feature vector. In such an embodiment, the second term of the global loss function may maximize sparsity, for example, by applying (forcing) a Laplace distribution. As another example, the global loss function may include a third term related to the reconstruction of the target input. In such an embodiment, the third term may minimize the difference between the output of a decoder and a target signal, the decoder receiving the output of the clone encoder as input. The decoder may mirror the encoder.
[0070] According to one aspect, a system includes at least one processor, a memory storing a plurality of clone encoders, the encoders in the plurality of clone encoders sharing weights, and means for training the clone encoders to generate sparse feature vectors based on inputs. In some embodiments, the system may include means for using the encoders with shared weights to generate sparse feature vectors based on new inputs and using the sparse feature vectors to condition a generative network. In some embodiments, the system may include means for using the encoders with shared weights to generate salient features for inputs, means for compressing the generated salient features, and means for storing the salient features.
[0071] According to one aspect, a system includes at least one processor, means for obtaining a conditioned input signal, means for extracting sequences from the conditioned input signal, and means for extracting salient features from each sequence. Some embodiments may include means for conditioning a generative network using the salient features. Some embodiments may include means for compressing the salient features and storing the salient features.
[0072] According to one aspect, a system includes at least one processor and a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods disclosed herein.
[0073] Various embodiments may include embodiments of one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be special purpose or general purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0074] These computer programs (also referred to as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium", "computer-readable medium" refer to any non-transitory computer program product, apparatus and / or device (e.g., magnetic disk, optical disk, memory (including read-access memory), programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor.
[0075] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0076] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0077] A variety of embodiments have been described. However, various modifications may be made without departing from the spirit and scope of the invention. In addition, the logic flow depicted in the figures does not require the particular order or sequential order shown to achieve the desired results. In addition, other steps may be provided, or steps may be eliminated from the described flow, and other components may be added to the described system or components may be removed therefrom. Therefore, other embodiments are within the scope of the following claims.
Claims
1. A method for identifying salient features of a generative network, the method comprising: obtaining a set of inputs for each original input in the input batch, the set of inputs comprising at least one modified input, each modified input being a different modified version of the original input; An encoder with weights is trained to provide features of the inputs by performing the following operations for each set of inputs in the input batch: providing the set of inputs to two or more clone encoders, each of the two or more clone encoders sharing weights and receiving a different respective input from the set of inputs, the encoder being one of the two or more clone encoders, and modifying the weights to minimize a global loss function having a first term, a second term, and a third term, wherein the first term maximizes similarity between features of the set of inputs, the second term maximizes independence and unit variance within the features generated by the encoder, and the third term minimizes a reconstruction loss by mapping the features generated by the encoder to an output representing a reconstruction of the input via a decoder to utilize a target input derived from the original input; extracting salient features of a new input using the encoder as extracted salient features; as well as The extracted salient features are provided to the generation network. 2 . The method of claim 1 , wherein the first term measures the difference between a feature extracted from a first input in the set of inputs and a feature extracted from each remaining input from the set of inputs.
3. The method of claim 1, wherein the first term measures the difference between features extracted from each input in the set of inputs and features extracted from each remaining input in the set of inputs.
4. The method of claim 1, wherein the second term minimizes a maximum mean deviation between a specified distribution and a distribution of first features extracted by the encoder over the input batch. The method of claim 4 , wherein the specified distribution is a Laplace distribution. The method of claim 4 , wherein the specified distribution is a Gaussian distribution.
7. The method according to claim 4, wherein the second term is expressed as: in: M is the size of the input batch, k(·,·) is the kernel, z represents the salient features of the input in the set of inputs, and y is obtained according to the specified distribution. The method of claim 1 , wherein the second term further maximizes sparsity within the extracted features.
9. The method of claim 1 , wherein minimizing the reconstruction loss using the target input comprises: providing the features generated by the encoders of the two or more clone encoders to the decoder; as well as The weights are adjusted to match the target input.
10. The method of claim 1, wherein the target input characterizes short-term spectral characteristics of original inputs associated with the set of inputs.
11. The method of claim 9, wherein the decoder is one of two or more clone decoders, the two or more clone decoders having a one-to-one relationship with the two or more clone encoders.
12. A method for identifying salient features of a generative network, comprising: receiving an input signal; For each time series of a plurality of time series extracted from the input signal: Extracting salient features of the time series as extracted salient features by providing the time series to an encoder that is one of two or more clone encoders trained to extract salient features, wherein the two or more clone encoders share weights and minimize a global loss function during training, the global loss function: maximizing the similarity between the salient features extracted by each of the two or more clone encoders; Maximizing independence, unit variance, and sparsity within the salient features; as well as minimizing a reconstruction loss by mapping the salient features generated by the encoder to an output representing a reconstruction of the time series to utilize a target input derived from the input signal; and The generative network is conditioned using the salient features.
13. The method of claim 12, wherein during training, each of the two or more clone encoders receives a respective input signal from a set of signals, each input signal from the set of signals representing an original input signal or a different modification of the original input signal.
14. The method of claim 12, wherein independence, unit variance, and sparsity are maximized using an enforced Laplace distribution.
15. The method of claim 12, wherein the input signal is a speech input and the conditioned generative network produces a speech waveform and the time series comprises mel frequency bins.
16. The method of claim 15, wherein conditioning the generative network causes the generative network to produce artifacts within a speech manifold.
17. A method for identifying salient features of a generative network, comprising: obtaining a set of inputs for each original input in the batch of inputs, the set of inputs for the original inputs comprising at least one modified input, each modified input being a different modified version of the original input; An encoder with weights is trained to generate features for each set of inputs in the input batch by performing the following operations for each set of inputs: providing respective inputs from the set of inputs to two or more clone encoders, the two or more clone encoders sharing the weights, wherein each of the two or more clone encoders receives a different respective input from the set of inputs, the encoder being one of the two or more clone encoders, and modifying the weights to minimize a global loss function having a first term, a second term, and a third term, the first term maximizing similarity between features generated by the two or more clone encoders for the set of inputs, the second term maximizing independence and unit variance within the features generated by the encoders, and the third term minimizing a reconstruction loss by mapping the features generated by the encoders to outputs representing reconstructions of the corresponding inputs via a decoder to utilize a target input derived from the original input; extracting salient features of a new input using the encoder; compressing the salient features extracted for the new input as compressed salient features; storing the compressed salient features extracted for the new input as compressed input; as well as The generative network is conditioned using the salient features.
18. The method of claim 17, wherein storing the compressed salient features comprises: The compressed salient features are transmitted to a remote computing device, which stores the compressed salient features.
19. The method of claim 18, wherein the remote computing device is configured to: The salient features are decompressed as decompressed salient features and a reconstruction input is generated according to the decompressed salient features.
20. The method according to any one of claims 17 to 19, further comprising: providing the features generated by the two or more cloned encoders to the decoder; as well as The weights are adjusted using the target input to minimize the reconstruction loss.
21. A system for identifying salient features of a generative network, comprising: at least one processor; and A memory storing instructions which, when executed by the at least one processor, cause the system to perform a method according to any one of claims 1 to 20.