Structural Retention Attention Mechanism in Sequence-to-Sequence Neural Models
By generating secondary attention vectors in the attention decoder of seq2seq artificial neural network, the problem of modifying the alignment matrix at inference time is solved, and the robust control of output generation and the retention of the alignment matrix structure is achieved.
Patent Information
- Application Number
- CN202080065832.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-19
- Filing Date
- 2020-09-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-09-18
AI Technical Summary
In some seq2seq generation applications, it is difficult to modify the alignment matrix at inferred time to control output generation, and the prior art lacks a robust alignment control mechanism.
By generating a sequence of primary concern vectors in the focus decoder of the trained seq2seq artificial neural network, a focus vector candidate is generated for each primary concern vector, the similarity between the candidate and the required focus vector structure is evaluated, and the secondary focus vector is generated using a soft selection ANN, and the output sequence is finally generated based on the secondary concern vector.
More robust control of output generation is realized, the structural retention ability of the alignment matrix is improved, and the alignment convergence performance of the model at inferred time is improved.
Smart Images

Figure CN114424209B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sequence-to-sequence (seq2seq) artificial neural networks (ANNs). Background Art
[0002] The use of neural models (i.e., ANNs) for seq2seq learning and inference was first introduced by I. Sutskever, O. Vinyals, and Q. V. Le in 2014 in "Sequence to Sequence learning with neural networks" in Advances in Neural Information Processing Systems 27 (NIPS 2014). The seq2seq neural model is capable of mapping an input sequence to an output sequence without prior knowledge of the input sequence length. Nowadays, seq2seq neural models are used for tasks such as machine translation, speech recognition, text-to-speech (TTS), video captioning, text summarization, text entailment, question answering, chatbots, etc.
[0003] Seq2seq neural models typically use an encoder-decoder architecture. Generally, the encoder and the decoder each include a recurrent neural network (RNN), such as a long short-term memory (LSTM) or a gated recurrent unit (GRU) network. In the encoder, the input sequence is encoded into a compact representation commonly referred to as a "state vector" or "context vector". These are used as inputs to the decoder, which generates a suitable output sequence. The decoder operates in separate iterations ("time steps"), and in one such time step, outputs each symbol of the output sequence.
[0004] The attention mechanism plays an important role in seq2Seq neural models. In many tasks, in order to generate the correct output sequence, not all symbols in the input sequence should be treated equally. For example, in machine translation, words that appear in the input sequence can have several different meanings; in order to translate them into the correct words in the second language, it is necessary to infer their correct meanings from other words in the input sequence based on the context. The attention mechanism can guide the seq2seq neural model to rely on the correct words in the input sequence to understand the context of the problematic words. This is typically performed by biasing the context vector before it is processed by the decoder. An attention weight vector (also referred to as an "alignment vector") is provided, and each attention weight vector determines the relative attention (also referred to as "alignment") of an output symbol of the decoder to the entire input sequence of the decoder. The context vector is represented as a linear combination of the encoded input sequence vector and its corresponding weights taken from the alignment vector, and the context vector is then processed by the decoder.
[0005] The above examples of the prior art and the limitations associated therewith are intended to be illustrative rather than exclusive. Other limitations in the relevant art will become apparent to those skilled in the art upon reading the specification and studying the drawings. In some seq2seq generation applications, it is highly desirable to modify the alignment matrix at inference time to control output generation. To this end, a mechanism for obtaining robust alignment control (i.e., preserving the structure of the alignment matrix) is needed.
[0006] Therefore, there is a need in the art to solve the above problems. Summary of the Invention
[0007] From a first aspect, the present invention provides a method, the method comprising: in a trained attention decoder of a trained sequence-to-sequence (seq2seq) artificial neural network (ANN), the method comprising: obtaining an encoded input vector sequence; generating a primary attention vector sequence using a trained primary attention mechanism of the trained attention decoder; for each primary attention vector in the primary attention vector sequence: generating a set of attention vector candidates corresponding to the respective primary attention vector, for each attention vector candidate in the set of attention vector candidates, evaluating a structural fitness metric that quantifies the similarity of the respective attention vector candidate to a desired attention vector structure, using a trained soft selection ANN to generate a secondary attention vector based on the evaluation and based on state variables of the trained attention decoder; and using the trained attention decoder to generate an output sequence based on the encoded input vector sequence and the secondary attention vector.
[0008] From another aspect, the present invention provides a system, comprising: (i) at least one hardware processor; and (ii) a non-transitory computer-readable storage medium having program code embodied therein, the program code being executable by the at least one hardware processor to perform the following instructions in a trained attention decoder of a trained sequence-to-sequence (seq2seq) artificial neural network (ANN): obtaining a sequence of encoded input vectors, generating a sequence of primary attention vectors using a trained primary attention mechanism of the trained attention decoder, for each primary attention vector in the sequence of primary attention vectors: generating a set of candidate attention vectors corresponding to the respective primary attention vector, for each candidate attention vector in the set of candidate attention vectors, evaluating a structural fitness metric that quantifies the similarity of the respective candidate attention vector to a desired attention vector structure, generating a secondary attention vector using a trained soft-selection ANN based on the evaluation and based on state variables of the trained attention decoder; and generating an output sequence using the trained attention decoder based on the sequence of encoded input vectors and the secondary attention vectors.
[0009] From another aspect, the present invention provides a computer program product for a sequence-to-sequence artificial neural network, the computer program product comprising a computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method for performing the steps of the present invention.
[0010] From another aspect, the present invention provides a computer program stored on a computer-readable medium and loadable into an internal memory of a digital computer, the computer program comprising software code portions for performing the steps of the present invention when the program is run on the computer.
[0011] The following embodiments and aspects thereof are described and illustrated in connection with systems, tools, and methods, which are intended to be exemplary and illustrative and not limiting in scope.
[0012] An embodiment relates to a method that includes: in a trained attention decoder of a trained sequence-to-sequence (seq2seq) artificial neural network (ANN): obtaining a sequence of encoded input vectors; generating a sequence of primary attention vectors using a trained primary attention mechanism of the trained attention decoder; for each primary attention vector in the sequence of primary attention vectors: (a) generating a set of attention vector candidates corresponding to the respective primary attention vector, (b) for each attention vector candidate in the set of attention vector candidates, evaluating a structural fitness metric that quantifies the similarity of the respective attention vector candidate to a desired attention vector structure, (c) using a trained soft selection ANN to generate secondary attention vectors based on the evaluation and based on state variables of the trained attention decoder; and generating an output sequence using the trained attention decoder based on the sequence of encoded input vectors and the secondary attention vectors.
[0013] Another embodiment relates to a system that includes: (i) at least one hardware processor; and (ii) a non-transitory computer-readable storage medium having program code embodied therein that is executable by the at least one hardware processor to perform the following instructions in a trained attention decoder of a trained sequence-to-sequence (seq2seq) artificial neural network (ANN): obtaining a sequence of encoded input vectors; generating a sequence of primary attention vectors using a trained primary attention mechanism of the trained attention decoder; for each primary attention vector in the sequence of primary attention vectors: (a) generating a set of attention vector candidates corresponding to the respective primary attention vector, (b) for each attention vector candidate in the set of attention vector candidates, evaluating a structural fitness metric that quantifies the similarity of the respective attention vector candidate to a desired attention vector structure, (c) using a trained soft selection ANN to generate secondary attention vectors based on the evaluation and based on state variables of the trained attention decoder; and generating an output sequence using the trained attention decoder based on the sequence of encoded input vectors and the secondary attention vectors.
[0014] A further embodiment relates to a computer program product that includes a non-transitory computer-readable storage medium having program code embodied therewith, the program code being executable by at least one hardware processor to perform the following instructions in a trained attention decoder of a sequence-to-sequence (seq2seq) artificial neural network (ANN) during training: obtain a sequence of encoded input vectors; generate a sequence of primary attention vectors using a trained primary attention mechanism of the trained attention decoder; for each primary attention vector in the sequence of primary attention vectors: (a) generate a set of candidate attention vectors corresponding to the respective primary attention vector, (b) for each candidate attention vector in the set of candidate attention vectors, evaluate a structural fitness metric that quantifies the similarity of the respective candidate attention vector to a desired attention vector structure, (c) generate a secondary attention vector using a trained soft selection ANN based on the evaluation and based on state variables of the trained attention decoder; and generate an output sequence using the trained attention decoder based on the sequence of encoded input vectors and the secondary attention vectors.
[0015] In some embodiments, generating the output sequence includes: generating an input context vector based on the sequence of encoded input vectors and based on the secondary attention vectors; and generating the output sequence using the trained attention decoder based on the input context vector.
[0016] In some embodiments, generating a set of candidate attention vectors includes: obtaining at least one of: a current primary attention vector, a set of previous primary attention vectors, and a set of previous secondary attention vectors; and augmenting at least one of the obtained vectors with additional attention vectors by shuffling and shifting at least one of the contents of the obtained vectors.
[0017] In some embodiments, generating a set of candidate attention vectors includes: obtaining at least one of: a current primary attention vector, a set of previous primary attention vectors, and a set of previous secondary attention vectors; and augmenting at least one of the obtained vectors with additional attention vectors by calculating additional attention vectors to conform to a desired attention vector structure.
[0018] In some embodiments, the structural fitness metric is based on at least one of: smooth maximum, kurtosis, skewness, entropy, ratio between L2 norm and L1 norm.
[0019] In some embodiments, generating a secondary attention vector includes: applying a scalar mapping to an evaluated structural fitness metric to produce a mapped structural fitness metric vector; providing a trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer; applying the trained sequential ANN to a state variable of a trained attention decoder, and adding the output vector of the application to the mapped structural fitness metric vector to produce an intermediate vector; providing the intermediate vector to a softmax layer to produce weights for a set of attention vector candidates; and forming a secondary attention vector by combining a set of attention vector candidates according to the weights of the set of attention vector candidates.
[0020] In some embodiments, generating a secondary attention vector includes: applying a scalar mapping to an evaluated structural fitness metric to produce a mapped structural fitness metric; defining a plurality of subsets of attention vector candidates and their corresponding mapped structural fitness metrics; for each subset of the plurality of subsets: (a) providing a trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer, (b) applying the trained sequential ANN to a state variable of a trained attention decoder, and adding the output vector of the application to the mapped structural fitness metric of the corresponding subset to produce an intermediate vector, (c) providing the intermediate vector to a softmax layer to produce weights for the subset of attention vector candidates, (d) forming a subset attention vector candidate by combining the attention vector candidates of the corresponding subset according to the weights of the attention vector candidates of the corresponding subset, (e) for the subset attention vector candidate, evaluating a subset structural fitness metric that quantifies the similarity of the subset attention vector candidate to a desired attention vector structure, and (f) applying a scalar mapping to the evaluated subset structural fitness metric to produce a mapped subset structural fitness metric; providing an additional trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer; applying the additional trained sequential ANN to a state variable of a trained attention decoder, and adding the output vector of the application of the additional trained sequential ANN to the vector of the mapped subset structural fitness metrics to produce an intermediate vector; providing the intermediate vector to a softmax layer to produce weights for the subset attention vector candidates; and forming a secondary attention vector by combining the subset attention vector candidates according to the weights of the subset attention vector candidates.
[0021] In some embodiments, the trained primary attention mechanism is an additive attention mechanism.
[0022] In some embodiments, the seq2seq ANN is configured for a text-to-speech task, and the method further includes, or the instructions further include: operating a vocoder to synthesize speech from the output sequence; and modifying the secondary attention vector before or during generating the output sequence to affect at least one prosodic parameter of the synthesized speech.
[0023] In some embodiments, the at least one prosodic parameter is selected from the set consisting of: intonation, stress, speed, rhythm, pause, and chunking.
[0024] In some embodiments, the method further comprises, or the instructions further comprise: receiving from the user a definition of a desired attention vector structure.
[0025] In addition to the above-exemplified aspects and embodiments, further aspects and embodiments will become apparent by reference to the drawings and by study of the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Exemplary embodiments are shown in the accompanying drawings. The dimensions of the components and features shown in the figures are generally chosen for convenience and clarity of presentation and are not necessarily shown to scale. The drawings are listed below.
[0027] Figure 1 is a block diagram of an exemplary system for operating a seq2seq ANN according to an embodiment.
[0028] Figure 2 is a flowchart of an exemplary method for operating a seq2seq ANN according to an embodiment.
[0029] Figure 3 is a diagram of an exemplary encoder-decoder configuration of a seq2seq ANN model according to an embodiment.
[0030] Figure 4 is a graph comparing the average alignment vector entropy of two attention mechanisms according to experimental results. DETAILED DESCRIPTION
[0031] Disclosed herein is a structure-preserving secondary attention mechanism for a seq2seq ANN model (hereinafter referred to as "the model"). The secondary attention mechanism (which can be used to replace a pre-existing ("primary") attention mechanism of the model) can improve the alignment stability of the model and is particularly beneficial in controlling various parameters of the output sequence of the model during inference. It is also beneficial for improving alignment convergence during the learning (also referred to as "training") of the model.
[0032] Advantageously, the structure-preserving secondary attention mechanism is capable of providing a secondary attention vector that, on the one hand, biases the context vector of the decoder differently from the primary attention mechanism of the model and, on the other hand, preserves a certain desired structure. The preservation of the desired structure is important for ensuring accurate prediction of the output sequence by the model. A secondary attention mechanism generated without attachment to an appropriate structure is unlikely to improve the primary attention mechanism.
[0033] The desired qualitative structure of the attention matrix can be defined by a user (e.g., a developer of a model who understands the characteristics of the attention mechanism suitable for the current situation) or hard-coded. As an example, for generating a high-quality output sequence, a text-to-speech model may require a sparse and monotonic structure for its attention matrix. This structure requires sparse and unimodal matrix rows (i.e., vectors), where each row has its peak position (e.g., argmax) that is not lower than the peak position of the previous row.
[0034] Considering the desired structure of the attention matrix, and thus for its vectors, a corresponding structure fitness metric can be evaluated for a set of generated attention vector candidates. The structure fitness metric quantifies the fitness of each candidate to the desired qualitative attention vector structure. The structure fitness metric should be differentiable so as to be integrated into the main model. For example, the output of a soft-maximum operator (e.g., LogSumExp) can be used as a rough structure fitness metric for unimodal sparse attention vectors.
[0035] To name just a few examples, the candidates in the set can be obtained, for example, from the current main attention vector, a set including one or more previous main attention vectors, and / or a set including one or more previous secondary attention vectors (i.e., vectors selected as secondary attention vectors in one or more previous iterations of the decoder).
[0036] Optionally, one or more additional attention vectors are employed to augment these obtained candidates to increase the number of candidates available for later evaluation. The additional attention vectors can be generated, for example, by shuffling and / or shifting the content of one or more of the obtained candidates. Another option is to compute one or more additional attention vectors from scratch based on the desired structure.
[0037] Next, the secondary attention vector can be obtained through soft selection, i.e., generated as a convex linear combination of the obtained candidate vectors and optionally also the augmented candidates. The weights of the convex linear combination can be generated by a trained sequential ANN that is fed at least with the structure fitness metric of the evaluated candidates. This ANN can be jointly trained with the main seq2seq network, thus preserving the original training loss of the main network.
[0038] The secondary attention vector generated according to this technique can then be used by the model for learning and / or inference, substituting the main attention vector into the calculation of the input context vector, which is fed to the rest of the decoder and used to generate the output sequence.
[0039] Now refer to Figure 1, which shows a block diagram of an exemplary system 100 for operating a seq2seq ANN according to an embodiment. The system 100 may include one or more hardware processors 102, random access memory (RAM) 104, and one or more non-transitory computer-readable storage devices 106.
[0040] The storage device 106 may have program instructions and / or components stored thereon that are configured to operate the hardware processor 102. The program instructions may include one or more software modules, such as the seq2seq ANN module 108. An operating system with various software components and / or drivers is also included for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), facilitating communication between different hardware and software components, and running the seq2seq ANN module 108.
[0041] The system 100 may be operated by loading the instructions of the seq2seq ANN module 108 into the RAM 104 for execution by the processor 102. The instructions of the seq2Seq ANN module 108 may cause the system 100 to receive an input sequence 110, process the input sequence 110, and generate an output sequence 112.
[0042] The system 100 described herein is only an exemplary embodiment of the present invention and may actually be implemented only in hardware, only in software, or in a combination of both hardware and software. The system 100 may have more or fewer components and modules than shown, may combine two or more of the components, or may have a different configuration or arrangement of the components. The system 100 may include any additional components that enable it to function as an operable computer system, such as a motherboard, data bus, power supply, network interface card, etc. (not shown). The components of the system 100 may be co-located or distributed (e.g., in a distributed computing architecture).
[0043] Now refer to Figure 2 the flowchart of Figure 2 to discuss the instructions of the seq2seq ANN module 108.
[0044] The steps of method 200 may be performed in the order in which they are presented or in a different order (or even in parallel), provided that the order allows the necessary input for a step to be obtained from the output of an earlier step. Additionally, unless otherwise specifically stated, the steps of method 200 (e.g., by Figure 1 system 100) are performed automatically.
[0045] In step 202, a sequence of encoded input vectors may be obtained from the encoder of the seq2seq ANN, as is known in the art.
[0046] In step 204, a sequence of primary attention vectors may be generated using the trained primary attention mechanism (optionally of the additive type) of the seq2seq ANN, as is known in the art.
[0047] The following steps, numbered 206, 208, and 210, may be repeated for each primary attention vector in the sequence of primary attention vectors generated in step 204. Steps 206, 208, and 210 may together constitute the secondary attention mechanism of the present embodiment for generating the output sequence, and this secondary attention mechanism is employed in step 212.
[0048] In step 206, a set of attention vector candidates (hereinafter referred to as "candidates") may be generated for each primary attention vector of the sequence of primary attention vectors using the trained primary attention mechanism. This is done so that a candidate that is most suitable for the desired attention vector structure can be selected later.
[0049] Generating the set of candidates may include obtaining one or more of the following vectors to be used as members of the set: the current primary attention vector (i.e., the corresponding primary attention vector of the current repetition, which is provided by the primary attention mechanism at the current time step of the operation of the decoder); a set of one or more previous primary attention vectors (i.e., one or more previously repeated primary attention vectors, which are provided by the primary attention mechanism at one or more previous time steps of the operation of the decoder); and a set of one or more previous secondary attention vectors (i.e., secondary attention vectors that are selected as secondary attention vectors in one or more previous repetitions (see step 210 below), which are provided by the secondary attention mechanism at previous time steps of the operation of the decoder).
[0050] In sub-step 206a, the obtained candidates are optionally enhanced with one or more additional attention vectors to increase the number of candidates available for evaluation in the next step of the method. One option for generating the additional attention vectors is to shuffle and / or shift (using circular shifting or zero-padding) the content of one or more of the obtained candidates. As a simple example, the content of the vector (9, 12, 23, 45) can be randomly shuffled to (23, 9, 12, 45), or linearly shifted (by zero-padding) by one index position to (0, 9, 12, 23). Another option is to generate the additional attention vectors by adding random noise to the content of one or more of the obtained candidates. Another option for generating the additional attention vectors is to recompute the additional attention vectors such that the additional attention vectors conform to the required attention matrix structure. For example, if the required structure is sparse and monotonic, the computed additional attention vectors can be sparse and unimodal, where the peak position (arg-max) of each vector is not lower than the peak position of the previous vector of that vector in the attention matrix.
[0051] In step 208, a structural fitness metric can be evaluated for each of the set of candidates obtained and optionally enhanced in steps 206 and 206a (and as previously mentioned, this is done separately for each set of candidates generated for each primary attention vector of the primary attention vector sequence). The structural fitness metric can be a mathematical formula that quantifies the similarity of the corresponding candidate to the required attention matrix structure. For example, the structural fitness metric can indicate how closely each candidate in the candidates conforms to the required structure, e.g., sparse and monotonic.
[0052] The evaluated structural fitness metric can be given on any numerical scale, such as, by way of example only, on a scale of [0, 1] (from completely dissimilar to completely identical).
[0053] In step 210, the secondary attention vectors can be generated based on the various candidates, the evaluation results of their structural fitness metrics, and one or more state variables of the decoder. Since steps 206, 208, and 210 are repeated for each primary attention vector in the primary attention vector sequence, the overall execution of step 210 results in multiple secondary attention vectors.
[0054] One way to generate a secondary attention vector is through soft selection. The following example section describes two variants of the soft selection module, namely a single-stage selection module and a hierarchical selection module, each of which is an embodiment of the present invention. Generally, the two variants can use one or a series of trained sequential ANNs that ultimately perform a convex linear combination of the obtained candidate vectors (and optionally also enhanced candidates), where the weights of the convex linear combination are generated by feeding the structural fitness metric of the evaluated candidates and the decoder state variables (e.g., the previous input context vector, the hidden state vector of the decoder, etc.) to the trained sequential ANN. The sequential ANN can be trained jointly with the main seq2seq network so as to preserve the original training loss of the main network.
[0055] As an alternative to the hierarchical selection module, the secondary attention vector can be generated by hierarchically applying a binary gating mechanism to pairs of candidates based on their respective structural fitness metrics.
[0056] Alternatively, the secondary attention vector can be selected or generated according to any criteria provided by the user of method 200 that take into account the evaluated structural fitness metrics.
[0057] Finally, in step 212, the decoder can generate an output sequence based on the secondary attention vector generated in step 210 and the sequence of encoded input vectors obtained in step 202. Depending on the current task (the tasks listed in the background section and other tasks), the output sequence can include any type of digital output, such as text, synthetic speech, media (images, videos, audio, music), etc. In some types of tasks, the output sequence requires another processing step to turn it into something meaningful to the user. For example, in the TTS task, the output sequence is typically a sequence of spectral audio features (represented by computer code) that requires vocoder processing, as is known in the art, in order to produce an audible waveform. To produce other types of media (such as images, videos, audio, music, etc.), other types of encoders can be used to process the output sequence into the desired type of media.
[0058] Optionally, step 212 also includes a sub-step 212a of controlling one or more characteristics of the decoder's output sequence during inference. That is, before or during the generation of the output sequence, the user can change one or more parameters to cause a corresponding modification of the secondary attention vector and, subsequently, a corresponding modification of the output sequence. This control can be achieved by implementing a sub-mechanism in the secondary attention mechanism that can receive parameters from a source external to the decoder and modify the secondary attention vector accordingly.
[0059] For example, in a seq2seq neural TTS task, control over the output sequence characteristics of the decoder can be beneficial, where a user may attempt to modify prosodic parameters of the speech synthesized from the output sequence. Prosody can reflect different characteristics of the speaker or utterance: the emotional state of the speaker; the form of the utterance (statement, question, or command); the presence of irony or sarcasm; emphasis, contrast, and focus. It can additionally reflect other elements of the language that may not be encoded by grammar or by the choice of vocabulary. Exemplary prosodic parameters include intonation (pitch, tenseness, accent, pitch range, tone), stress (pitch salience, length, loudness, timbre), speed, rhythm, pauses, and chunking. As an addition to or alternative to prosody, other types of audio characteristics of the synthesized speech can be controlled.
[0060] Example
[0061] The following provides an exemplary algorithm for implementing Figure 2 method 200. It is not intended to limit method 200 to any details, but rather provides such details as additional embodiments of the various steps of the method.
[0062] Reference Figure 3 describes an exemplary algorithm, Figure 3 is a block diagram showing a seq2seq ANN 300 having an encoder 302 and a decoder 304.
[0063] Let t be the current time step of the operation of the decoder. The exemplary algorithm employs a secondary attention mechanism 304b in place of the primary attention mechanism 304a of the decoder. The secondary attention mechanism 304b derives the t-th alignment vector from W previously obtained alignment vector candidates (e.g., a consecutive set ), depending on the decoder state variable 304c and the encoded input sequence.
[0064] Optionally, there is an additional "backward" candidate, which is equal to the initial t-th alignment vector candidate c 0 = ai n i t,t .
[0065] From each alignment vector candidate , a set of enhanced alignment vector candidates is generated, for example, by shuffling or shifting its components. For example, the enhancement by linear shifting can be: where n is the input sequence index, and the boundary conditions for the shift are set appropriately (e.g., by zero-padding). The enhancement can be random or determined based on existing knowledge about the desired attention weight structure.
[0066] Then, for each alignment vector candidate (including a "backward" candidate, the augmented set of which is simple, i.e., including only the original candidate), evaluate the differentiable structure fitting metric s j,k = f(c j,k ). In a variant of the exemplary algorithm, the structure fitting metric is determined only by the original alignment vector candidates (before augmentation), i.e., s j,k = f(c j ).
[0067] In another variant, the structure fitting metric includes the LogSumExp smoothed maximum operator combined with a "peak" criterion evaluated by the ratio of the L2 norm to the L1 norm (which is equal to the L2 norm since for alignment candidates, the L1 norm is always equal to unity). This also ensures that the combined criterion of the structure fitting metric is within the range [0, 1], where 1 represents a perfect fit and 0 represents the worst fit. The exemplary formula proposed for f(c) is given by:
[0068] f(c) = Thresh(f 1 (c) f 2 (c)),
[0069] where
[0070]
[0071]
[0072] and
[0073]
[0074] This criterion favors the maximum sparse unimodal probability distribution (i.e., the delta function). Another known alternative to the "peak" criterion is kurtosis, which can be used as an alternative.
[0075] Feed the fully augmented set of alignment vector candidates into a trainable and differentiable candidate selection module that outputs the final alignment vector a t . The candidate selection module is conditioned on the decoder state variables. It also utilizes the evaluated structure fitting metric to favor appropriate structured candidates. Variants of the exemplary algorithm include a single-stage selection module or a hierarchical selection module, both of which deploy the following alignment vector structure fitting adjustment:
[0076] Let be the restricted log(x), e.g., Then, for a set of candidate structure fitting metrics s j,k = f(c j,k ), a set of candidate structure fitting adjustment components is defined:
[0077]
[0078] The evaluated structural fitness metric is mapped from its original [0, 1] range to a wider range [-100, 0] in a predefined manner. Of course, other wider ranges can also be used.
[0079] As an alternative to this predefined mapping, the mapping can be performed by feeding the structural fitness metric into a trainable scalar mapping (implemented by an ANN), which is trained jointly with other parts of the decoder. Then, this trained scalar mapping ANN generates the structural fitness adjustment component vector S j,k 。
[0080] As a further alternative, the structural fitness metric can itself be formed such that it provides an evaluation result in a range wider than [0, 1], so that no additional mapping is required.
[0081] The variant with a single-level selection module can operate as follows:
[0082] Let K be the quantity of all alignment vector candidates c j,k That is And S be the vector of the corresponding candidate structural fitness adjustment components. Then, there are K candidate selection weights {α j,k}, and a trained multi-layer sequential ANN (with alternating linear and non-linear layers and a terminating linear layer) fed by the decoder state variables is used to evaluate these K candidate selection weights. The evaluated K-dimensional vector output (specifically, the output of the terminating linear layer) is added to the adjustment vector S, and the resulting intermediate vector is fed into a softmax layer with K weight outputs {α j,k}. Then, the secondary alignment vector is formed through a soft selection operation:
[0083] a t =∑ j,k α j,k c j,k
[0084] Such that the set of attention vector candidates c j,k is combined according to their weights.
[0085] The variant with a hierarchical selection module can operate as follows:
[0086] Define W separate subsets of the alignment vector candidates Which are from the corresponding enhanced alignment vector sets Select from among them. For each of the W subsets, a process similar to that of the single-level selection module is performed; however, instead of ending with a secondary attention vector, a single attention vector candidate (referred to as the "subset" attention vector candidate) representing the optimal structural fit of the subset is finally formed for each subset. This also involves evaluating the structural fit metric of the subset attention vector candidate. Then, an additional trained sequential ANN is used to process the structural fit metrics of all subsets, and the output of this ANN is fed to a softmax layer to determine the weights of the subset attention vector candidates. Finally, the secondary attention vector is formed by combining the intermediate attention vector candidates according to their weights.
[0087] More specifically, the j-th soft selection module (among the W such modules) employs a multi-layer sequential ANN (with alternating linear and non-linear layers and a terminating linear layer) to predict the K j selection weights {β k}, which is fed by the decoder state variables, and its output is further added to the corresponding subset of the structural fit adjustment S and passed through a softmax layer to obtain the soft selection weight β k for the j-th subset. Further, the j-th soft selection of the intermediate vector candidates is performed as follows:
[0088] d j = ∑ k β k c j,k
[0089] In addition, d 0 = c 0 .
[0090] After all W soft selection modules are completed, a single attention vector candidate is selected from the W + 1 intermediate attention vector candidates. Let (W + 1) be the quantity of the intermediate attention vector candidate d j , and S be the vector of the corresponding candidate structural fit adjustment components {S j}:
[0091]
[0092] Then, there are W + 1 final candidate selection weights {γ j}, and a multi-layer sequential ANN with alternating linear and non-linear layers and a terminating linear layer fed by the decoder state variables is used to evaluate these W + 1 final candidate selection weights. The (W + 1)-dimensional output of the terminating linear layer is added to the corresponding adjustment S, and the resulting output vector is fed to a softmax layer with (W + 1) outputs {γ j}. Finally, the secondary alignment vector is formed as follows:
[0093] a t = ∑ j γd j 。
[0094] In the simplified usage of the hierarchical selection module, where W = 1 and K 1 = 2, sigmoid can be used instead of softmax as follows:
[0095] Here, f(c) is the structural fitness metric, and β 1 and γ 1 are scalar conversion probabilities, which are predicted by a separate multi-layer sequential ANN, fed by the decoder state variables and terminated with a sigmoid layer.
[0096] Then,
[0097] d 1 = β 1 c 1,1 +(1 - β 1 )c 1,0
[0098] a t = (1 - f(c 0 )(1 - f(c 1 )))γ 1 d 1 + f(c 0 )(1 - f(c 1 ))(1 - γ 1 )c 0
[0099] Experimental results
[0100] The disclosed structure-preserving secondary attention mechanism was successfully tested in the seq2seq neural TTS task and showed good alignment convergence during training and high MOS scores during user control of two TTS prosody parameters (speech rate, speech pitch) at inference time.
[0101] The experimental task follows the "Tacotron2" architecture ("Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions" proposed by Shen, Jonathan, et al. in the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) in 2018), including a sequence-to-sequence network with recurrent attention for spectral feature prediction, which is similar to Wavenet (Van Den Oord, cascaded with the neural vocoder "WaveNet: A generative model for raw audio" proposed by Tamamori, Akira et al. at SSW125 in 2016 and "Speaker-Dependent WaveNet Vocoder" proposed by Tamamori, Akira et al. at INTERSPEECH in 2017, all having various advantageous modifications aimed at improving the quality of synthetic speech, training convergence, and sensitivity to prosody control mechanisms.
[0102] A corpus of male and female utterances with a sampling rate of 22050 Hz is used. The male dataset contains 13 hours of speech and the female dataset contains 22 hours of speech. Both are generated and recorded by native speakers of American English in a professional recording studio. Audio is recorded utterance by utterance, with most utterances containing a single sentence.
[0103] To facilitate control of prosody parameters, appropriate training is performed based on prosody observations extracted from the recordings, and a mechanism for controlling these parameters using per-component offsets in the range [-1, 1] is embedded. At inference time, the prosody parameters are predicted from the encoded output sequence, and the user can deliberately offset the prosody parameter for generating the output sequence and the subsequent output waveform.
[0104] Knowing the required alignment matrix structure for this specific TTS task (monotonic alignment evolution), a set of alignment vector candidates is derived from the previous alignment vector in addition to the main current alignment vector. Then, soft selection is applied to obtain the secondary alignment vector in a way that preserves its desired required structure (i.e., a unimodal shape with a spike).
[0105] Let b t be the initial alignment vector as evaluated by the initial attention module, and a t [n] be the secondary alignment vector at the output time step t. Assuming the monotonic attention of "Online and linear-time attention by enforcing monotonic alignments" proposed by Raffel, Colin et al. in the proceedings of the 34th International Conference on Machine Learning, volume 70, organized by the Journal of Machine Learning Research in 2017, without skipping input symbols, by adding the previous alignment vector a t-1 [n] and its shifted version a t-1 [n - 1] to the current initial alignment b t at the current time step t to create a candidate set
[0106]
[0107] This enhancement assumes that at the current time, the output remains aligned with the previous input symbol or moves to the next one.
[0108] With this set of candidates, only the soft selection can be trained to propose secondary alignment vectors. However, as a precautionary measure, the experiment aims to ensure that the soft selector prefers appropriately structured candidates to eliminate temporarily occurring attention disruptions. For this purpose, a scalar structure metric is used as a structure fitting metric, which evaluates the unimodality and peak sharpness of the alignment vector candidates. This metric combines the LogSumExp softmax evaluation with an additional peak sharpness metric derived from the common "peak" metric (McCree AV, Barnwell TP, "A mixed excitation LPC vocoder model for low bitrate speech coding", IEEE Transactions on Speech and Audio Processing, Vol. 3, No. 4, July 1995, pp. 242 - 50), i.e., the L2 norm divided by the L1 norm, noting that the L1 norm is always equal to unity for the alignment vector and the squared L2 norm of the worst - case flat alignment vector is equal to 1 / N. The boost constant is experimentally set to 1.67 to reduce the sensitivity of this metric.
[0109] The combined structure metric used in the experiment is given by:
[0110]
[0111] where
[0112]
[0113]
[0114] and, the threshold operator |x| α is defined as:
[0115]
[0116] The added thresholding operation (with an experimentally set near - zero threshold of 0.12) ensures that poor alignment vector candidates are not suitable for soft selection.
[0117] The structure - preserving soft selection of the alignment vector is carried out in two stages. The first stage is given by:
[0118] d = αa t-1 [n - 1]+(1 - α)a t-1 [n] (5)
[0119] where α is a scalar initial stage selection weight, which is generated by a single fully-connected layer, fed with the concatenated decoder state variables (x c , h c ) and terminated with a sigmoid layer. Looking at the first stage selection (5), it can be noted that it provides explicit phoneme conversion control through the embedded prosodic parameters, which are part of the input context vector.
[0120] The final stage of the selection process utilizes the structural metric f(c):
[0121] a t = (1 - γ)βd + γ(1 - β)b t (6)
[0122] where β is a scalar final stage selection weight, which is generated by a single fully-connected layer, fed with the input context vector x c and terminated with a sigmoid layer, and γ = f(b t )(1 - f(d)) is the structural preference score. This multiplicative structural preference score ensures that the initial attention vector is only considered if its structure is better than other candidates.
[0123] In the experiment, at inference time, a Wavenet-type vocoder (Van Den Oord et al. and Tamamori et al., ibid.) was used to generate the output waveform from the spectral features predicted by the model.
[0124] This experiment shows an improvement in the alignment convergence during training, as visible in Figure 4 . The figure shows the average alignment vector entropy of a 100-sentence validation set during training on a data corpus of 13,000 sentences. The mini-batch size is 48. The average alignment vector entropy with the structure-preserving attention mechanism of the present invention is lower than that of only the conventional attention mechanism (the previously existing attention mechanism of the model).
[0125] To evaluate the quality and expressiveness of the output waveforms created with the structure-preserving attention mechanism, two formal MOS listening tests were performed on 40 synthetic sentences (one for each speech corpus, male and female). Each test rated four systems: a first system using this structure-preserving attention mechanism (denoted as AugAttn in the following tables), and three baseline systems: the original, unmodified speech recording (denoted as PCM); the output waveforms of the same model but with only the primary attention mechanism (denoted as RegAttn); and the output waveforms of the "WORLD" system ("WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications" by Morise Masanori, Fumiya Yokomori, and Kenji Ozawa in IEICE Transactions on Information and Systems 99.7 (2016): 1877-1884) (denoted as WORLD). The AugAttn system was rated three times: once without any prosody control (speech rate and pitch 0,0), and twice with different speech rate and pitch controls. Each of the synthetic sentences was rated by 25 different subjects.
[0126] Tables 1 and 2 list the MOS evaluation results for naturalness and expressiveness of female and male utterances, respectively. A significance analysis of the results in Table 1 revealed that most of the cross-system expressiveness differences were statistically significant, except for the differences between RegAttn and AugAttn (0,0), and between RegAttn (-0.1,0.5) and AugAttn (0.15,0.6). In terms of naturalness, all enhanced attention systems performed like RegAttn (no significant difference), except that RegAttn (-0.1,0.5) performed slightly better (p = 0.046). Thus, for female utterances, prosody control can significantly improve the perceived expressiveness while preserving the original quality and naturalness.
[0127] Similarly, a significance analysis of the male utterances (Table 2) revealed that only the pairs of RegAttn and AugAttn (0,0), and RegAttn (0.2,0.8) and AugAttn (0.5,1.5) were equivalent in terms of perceived expressiveness. In terms of naturalness, both RegAttn (0.2,0.8) and AugAttn (0.5,1.5) provided significant improvements compared to RegAttn (0,0) and RefAttn. That is, for male utterances, prosody control can fully and significantly improve expressiveness, quality, and naturalness.
[0128] Table 1. MOS Evaluation of Naturalness and Expressiveness of American English Female Speech (μ±95%)
[0129]
[0130] Table 2. MOS Evaluation of Naturalness and Expressiveness of American English Male Speech (μ±95%)
[0131]
[0132]
[0133] In summary, the experiments have revealed that the structure-preserving attention mechanism of the present invention applied in a seq2seq neural TTS system maintains high quality and naturalness with and without prosody control during inference. Those skilled in the art will recognize that similar results are likely to be achievable in other types of seq2seq neural tasks such as machine translation, speech recognition, video captioning, text summarization, text entailment, question answering, chatbots, etc.
[0134] The above techniques used in the experiments are considered embodiments of the present invention.
[0135] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0136] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable) or an electrical signal transmitted through a wire. Instead, a computer-readable storage medium is a non-transitory (i.e., non-volatile) medium.
[0137] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or an external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0138] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by using the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform aspects of the present invention.
[0139] The present invention will be described below with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0140] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that directs a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium which stores the instructions comprises an article of manufacture including instructions which implement aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0141] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices so that a series of operational steps are performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0142] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. It should also be noted that each box in the block diagrams and / or flowchart, and combinations of boxes in the block diagrams and / or flowchart, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts or combinations of special-purpose hardware and computer instructions.
[0143] The description of a numerical range should be considered to have specifically disclosed all possible sub-ranges as well as the individual numerical values within that range. For example, the description of a range from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as the individual numbers within that range, for example 1, 2, 3, 4, 5, and 6. This applies regardless of the width of the range.
[0144] The description of various embodiments of the invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or technical improvement over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method in a trained attention decoder of a trained sequence-to-sequence seq2seq artificial neural network ANN, the method comprises: obtaining an encoded sequence of input vectors; generating a sequence of primary attention vectors using a trained primary attention mechanism of the trained attention decoder; for each primary attention vector in the sequence of primary attention vectors: (a) generating a set of candidate attention vectors corresponding to the respective primary attention vector, (b) for each candidate attention vector in the set of candidate attention vectors, evaluating a structural fitness metric that quantifies the similarity of the respective candidate attention vector to a desired attention vector structure, (c) using a trained soft selection ANN to generate a secondary attention vector based on the evaluation and based on state variables of the trained attention decoder; and using the trained attention decoder to generate an output sequence based on the encoded sequence of input vectors and the secondary attention vector.
2. The method according to claim 1, wherein, generating the output sequence comprises: generating an input context vector based on the encoded sequence of input vectors and based on the secondary attention vector; and using the trained attention decoder to generate the output sequence based on the input context vector.
3. The method according to claim 1 or 2, wherein, generating the set of candidate attention vectors comprises: obtaining at least one of: a current primary attention vector, a set of previous primary attention vectors, and a set of previous secondary attention vectors; and augmenting the at least one obtained vector with additional attention vectors by at least one of content shuffling and shifting of the at least one obtained vector.
4. The method according to claim 1 or 2, wherein, generating the set of candidate attention vectors comprises: obtaining at least one of: a current primary attention vector, a set of previous primary attention vectors, and a set of previous secondary attention vectors; and augmenting the at least one obtained vector with additional attention vectors by calculating additional attention vectors that conform to the desired attention vector structure.
5. The method according to claim 1 or 2, wherein, the structural fitness metric is based on at least one of: smooth maximum, kurtosis, skewness, entropy, ratio between L2 norm and L1 norm.
6. The method according to claim 1 or 2, wherein, generating the secondary attention vector comprises: applying a scalar mapping to the evaluated structural fitness metric to produce a mapped structural fitness metric vector; providing a trained sequential ANN having: alternating linear layers and non-linear layers, and a terminating linear layer; applying the trained sequential ANN to the state variables of the trained attention decoder and adding the output vector of the application to the mapped structural fitness metric vector to produce an intermediate vector; providing the intermediate vector to a softmax layer to produce weights for the set of candidate attention vectors; and forming the secondary attention vector by combining the set of candidate attention vectors according to the weights of the set of candidate attention vectors.
7. The method according to claim 1 or 2, wherein, generating the secondary attention vector includes: applying a scalar mapping to the evaluated structural fit metric to produce a mapped structural fit metric; defining multiple subsets of attention vector candidates and their corresponding mapped structural fit metrics; for each of the multiple subsets: providing a trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer, applying the trained sequential ANN to the state variables of the trained attention decoder, and adding the output vector of the application to the mapped structural fit metric of the corresponding subset to produce an intermediate vector; providing the intermediate vector to a softmax layer to produce the weights of the subset of the attention vector candidates; forming a subset attention vector candidate by combining the attention vector candidates of the corresponding subset according to the weights of the attention vector candidates of the corresponding subset; evaluating a subset structural fit metric that quantifies the similarity between the subset attention vector candidate and the desired attention vector structure for the subset attention vector candidate; and applying a scalar mapping to the evaluated subset structural fit metric to produce a mapped subset structural fit metric; providing an additional trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer; applying the additional trained sequential ANN to the state variables of the trained attention decoder, and adding the output vector of applying the additional trained sequential ANN to the vector of the mapped subset structural fit metric to produce an intermediate vector; providing the intermediate vector to a softmax layer to produce the weights of the subset attention vector candidates; and forming the secondary attention vector by combining the subset attention vector candidates according to the weights of the subset attention vector candidates.
8. The method according to claim 1 or 2, wherein, the trained primary attention mechanism is an additive attention mechanism.
9. The method according to claim 1 or 2, wherein, the seq2Seq ANN is configured for a text-to-speech task, and the method further includes: operating a vocoder to synthesize speech from the output sequence; and modifying the secondary attention vector before or during generating the output sequence to affect at least one prosodic parameter of the synthesized speech.
10. The method according to claim 9, wherein, the at least one prosodic parameter is selected from the group consisting of intonation, stress, speed, rhythm, pause, and chunking.
11. The method according to claim 1 or 2, further comprising receiving a definition of the desired attention vector structure from a user.
12. A computer system, comprising: (i) at least one hardware processor; and (ii) A non-transitory computer-readable storage medium having program code embodied therein, the program code being executable by the at least one hardware processor to perform the following instructions in a trained attention decoder of a trained sequence-to-sequence seq2seq artificial neural network ANN: Obtain an encoded sequence of input vectors, Generate a sequence of primary attention vectors using a trained primary attention mechanism of the trained attention decoder, For each primary attention vector in the sequence of primary attention vectors: (a) Generate a set of candidate attention vectors corresponding to the respective primary attention vector, (b) For each candidate attention vector in the set of candidate attention vectors, evaluate a structural fitness metric that quantifies the similarity of the respective candidate attention vector to a desired attention vector structure, (c) Generate a secondary attention vector using a trained soft selection ANN based on the evaluation and based on a state variable of the trained attention decoder; And Generate an output sequence using the trained attention decoder based on the encoded sequence of input vectors and the secondary attention vector.
13. The system according to claim 12, Wherein, Generating the output sequence includes: Generating an input context vector based on the encoded sequence of input vectors and based on the secondary attention vector; and Generating the output sequence using the trained attention decoder based on the input context vector.
14. The system according to claim 12 or 13, Wherein, Generating the set of candidate attention vectors includes: Obtaining at least one of: a current primary attention vector, a set of previous primary attention vectors, and a set of previous secondary attention vectors; and Enhancing the at least one obtained vector with additional attention vectors by at least one of shuffling and shifting the content of the at least one obtained vector.
15. The system according to claim 12 or 13, Wherein, Generating the set of candidate attention vectors includes: Obtaining at least one of: a current primary attention vector, a set of previous primary attention vectors, and a set of previous secondary attention vectors; and Enhancing the at least one obtained vector with the additional attention vectors by calculating additional attention vectors that conform to the desired attention vector structure.
16. The system according to claim 12 or 13, Wherein, The structural fitness metric is based on at least one of: smooth maximum, kurtosis, skewness, entropy, ratio between L2 norm and L1 norm.
17. The system according to claim 12 or 13, Wherein, Generating the secondary attention vector includes: Applying a scalar mapping to the evaluated structural fitness metric to produce a mapped structural fitness metric vector; Providing a trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer; Applying the trained sequential ANN to the state variable of the trained attention decoder and adding the output vector of the application to the mapped structural fitness metric vector to produce an intermediate vector; Providing the intermediate vector to a softmax layer to produce weights for the set of candidate attention vectors; and Forming the secondary attention vector by combining the set of candidate attention vectors according to the weights of the set of candidate attention vectors.
18. The system according to claim 12 or 13, wherein, generating the secondary attention vector includes: Applying a scalar mapping to the evaluated structural fitness metric to produce a mapped structural fitness metric; Defining a plurality of subsets of candidate attention vectors and their corresponding mapped structural fitness metrics; For each of the plurality of subsets: Providing a trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer, Applying the trained sequential ANN to the state variables of the trained attention decoder and adding the output vector of the application to the mapped structural fitness metric of the corresponding subset to produce an intermediate vector; Providing the intermediate vector to a softmax layer to produce weights for the subset of the candidate attention vectors; Forming a subset candidate attention vector by combining the candidate attention vectors of the corresponding subset according to the weights of the candidate attention vectors of the corresponding subset, Evaluating, for the subset candidate attention vector, a subset structural fitness metric that quantifies the similarity of the subset candidate attention vector to a desired attention vector structure, and Applying a scalar mapping to the evaluated subset structural fitness metric to produce a mapped subset structural fitness metric; Providing an additional trained sequential ANN having: alternating linear and non-linear layers, and a terminating linear layer; Applying the additional trained sequential ANN to the state variables of the trained attention decoder and adding the output vector of applying the additional trained sequential ANN to the vector of the mapped subset structural fitness metric to produce an intermediate vector; Providing the intermediate vector to a softmax layer to produce weights for the subset candidate attention vectors; and Forming the secondary attention vector by combining the subset candidate attention vectors according to the weights of the subset candidate attention vectors.
19. The system according to claim 12 or 13, wherein, The trained primary attention mechanism is an additive attention mechanism.
20. The system according to claim 12, wherein, The program code is further executable by the at least one hardware processor to execute the following instructions: Receiving from the user a definition of the desired attention vector structure.
21. A computer program product for a sequence-to-sequence artificial neural network, the computer program product comprising: A computer-readable storage medium that can be read by a processing circuit and stores instructions for execution by the processing circuit to perform the method according to any one of claims 1 to 11.
22. A computer-readable storage medium stores computer instructions that can be loaded into the internal memory of a digital computer. The computer instructions include a software code portion that, when the computer instructions are run on a computer, is configured to perform the method of any one of claims 1 to 11.
Citation Information
Patent Citations
Systems and methods for neural text-to-speech using convolutional sequence learning
US20190122651A1
End-to-end text-to-speech conversion
WO2018183650A2