Speech synthesis learning device, speech synthesis device, speech synthesis leaning method, speech synthesis method and program

By introducing symbolic representation and discrete expression techniques into the speech synthesis learning device, the problem of difficult to interpret and express the speaker's personality and emotional information in speech synthesis in the prior art is solved, and user-interpretable speech expression and more natural speech synthesis effects are achieved.

JP2025073546APending Publication Date: 2025-05-13NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023184449
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing pronunciation techniques are difficult to effectively express and interpret the personality and emotional information of the speakers estimated during pronunciation, making it difficult for users to understand and interpret this information.

Method used

By introducing symbolic representations in the speech synthesis learning device, the speech feature vector of the continuous value is converted into discrete symbolic expressions and a discrete speech expression vector is generated, which facilitates user understanding and control.

Benefits of technology

The conversion of the speaker's personality and emotional information estimated during the speech synthesis process into user-interpretable symbolic representations is realized, which reduces the user's error correction burden and improves the naturalness and expressiveness of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073546000001_ABST
    Figure 2025073546000001_ABST
Patent Text Reader

Abstract

To provide a speech synthesis learning device and method, a speech synthesis device and method, and a program for expressing an expression of speech estimated during speech synthesis in a form interpretable by a user.SOLUTION: A speech synthesis device 10 includes: a speech expression estimation unit that outputs any of a plurality of symbols corresponding to an expression of speech for a second text including a first text indicating contents of speech; a discrete extraction unit that generates a discretized speech expression vector, which is a vector corresponding to the expression of speech by converting a vector corresponding to the expression of speech based on speech features related to the speech into any of the symbols; and a learning unit that updates parameters of the discrete expression extraction unit, a series conversion unit, and the speech expression estimation unit so that the speech features generated by the series conversion unit are closer to those of the teacher speech features, by using learned data that is a pair of the teacher speech features and the speech information and the series conversion unit that generates the speech features based on a language information vector based on the speech information including the first text and the discretized speech expression vector.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a voice synthesis training device, a voice synthesis device, a voice synthesis training method, a voice synthesis method, and a program. [Background technology]

[0002] In the field of speech synthesis, a speech synthesis technique based on Deep Neural Networks (DNNs) has been proposed. This technique is known to be capable of generating synthetic speech of higher quality than conventional methods.

[0003] On the other hand, when comparing the synthetic voice by the above technology with the voice of a narrator reading a picture book or the like, there is a large difference in the naturalness of intonation, etc. One of the reasons for this is that the voice synthesis technology generates synthetic voice only from linguistic information such as reading and accent obtained from the text of the picture book, etc. In contrast, when a narrator reads a picture book, etc., the voice is produced by utilizing not only the reading and accent obtained from the text, but also information such as the speaker's personality and emotion of the character inferred from the text and its surrounding long-term context, etc. For this purpose, Non-Patent Document 1 models the speaker's personality and emotion contained in the voice of the training data (voice data of a voice actor reading a picture book, etc. and the text), and improves the expressiveness of the synthetic voice by inferring it from text information using a large-scale language model such as BERT. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] LEI, Shun, et al., "Towards Expressive Speaking Style Modeling with Hierarchical Context Information for Mandarin Speech Synthesis", arXiv preprint arXiv:2203.12201, 2022 Summary of the Invention [Problem to be solved by the invention]

[0005] However, since the modeled speaker characteristics and emotion information are expressed as vectors of continuous values ​​of tens to hundreds of dimensions, it is difficult for users to interpret the speaker characteristics and emotions estimated during speech synthesis.

[0006] The present invention has been made in consideration of the above points, and has an object to make it possible to express an utterance expression estimated during speech synthesis in a format that can be interpreted by the user. [Means for solving the problem]

[0007] In order to solve the above problem, the speech synthesis training device includes an utterance expression guessing unit configured to output one of a plurality of symbols corresponding to an utterance expression for a second text including a first text indicating the content of an utterance; a discrete expression extraction unit configured to convert a vector corresponding to an utterance expression based on speech features related to the utterance into one of the symbols and generate a discretized utterance expression vector, which is a vector corresponding to the utterance expression, from the symbol; a sequence conversion unit configured to generate speech features based on a language information vector based on utterance information including the first text and the discretized utterance expression vector; and a set of teacher speech features and utterance information, and a learning unit configured to update parameters of the discrete expression extraction unit and the sequence conversion unit using training data obtained by training the speech information so that the discretized utterance expression vector generated by the discrete expression extraction unit for the teacher speech feature and the speech feature generated by the sequence conversion unit for the utterance information approach the teacher speech feature, and to update parameters of the utterance expression prediction unit so that the discretized utterance expression vector generated by the discrete expression extraction unit from the symbols output by the utterance expression prediction unit that has been trained for a text that partly includes a text related to the utterance information and the speech feature generated by the sequence conversion unit that has been trained for the utterance information approach the teacher speech feature. Effect of the Invention

[0008] It is possible to make it possible to present the speech expression estimated during speech synthesis in a form that can be interpreted by the user. [Brief description of the drawings]

[0009] [Figure 1] 1 is a diagram illustrating an example of a hardware configuration of a voice synthesizer 10 according to an embodiment of the present invention. [Diagram 2] 1 is a diagram illustrating an example of a functional configuration of a speech synthesis device 10 in a learning phase according to an embodiment of the present invention. [Diagram 3] 10 is a flowchart illustrating an example of a first processing procedure executed by the speech synthesizer 10 in a learning phase. [Figure 4] FIG. 11 is a diagram for explaining the discrete expression extraction unit 133 in the first processing procedure of the learning phase. [Diagram 5] 10 is a flowchart illustrating an example of a second processing procedure executed by the speech synthesizer 10 in the learning phase. [Figure 6] FIG. 11 is a diagram for explaining the discrete expression extraction unit 133 in a second processing procedure of the learning phase. [Figure 7] 2 is a diagram illustrating an example of a functional configuration of the speech synthesis device 10 in the estimation phase according to the embodiment of the present invention. FIG. [Figure 8] 10 is a flowchart illustrating an example of a processing procedure executed by the speech synthesizer 10 in the estimation phase. [Figure 9] FIG. 13 is a diagram showing an example of a screen for receiving a selection of an utterance expression symbol from a user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 is a diagram showing an example of a hardware configuration of a voice synthesizer 10 in an embodiment of the present invention. The voice synthesizer 10 in Fig. 1 has a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, an interface device 105, etc., which are mutually connected by a bus B.

[0011] A program for implementing processing in the speech synthesizer 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the program does not necessarily have to be installed from the recording medium 101, but may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program as well as necessary files, data, etc.

[0012] When an instruction to start a program is received, the memory device 103 reads out the program from the auxiliary storage device 102 and stores it. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the speech synthesizer 10 in accordance with the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.

[0013] In this embodiment, the process executed by the speech synthesizer 10 is roughly divided into two phases: a learning phase and an estimation phase.

[0014] In the learning phase, the speech synthesis device 10 models speech expressions such as speaker characteristics (e.g., characteristics of a particular individual (speaker)'s voice and speaking style (habits)) and emotions contained in the speech of the learning data (audio data of a voice actor or the like reading a picture book or other book and its text) as discrete information (speech expression symbols described below), and infers the speaker characteristics and emotions from the text information of the learning data.

[0015] In the estimation phase, the speech synthesizer 10 estimates the speaker characteristics and emotions to be expressed in the synthetic speech as discrete information from the text to be read aloud (the subject of speech synthesis) based on a model that has been trained in advance. The speech synthesizer 10 also presents a plurality of discrete pieces of information to the user during speech synthesis, thereby reducing the burden on the user during speech synthesis (the burden on the user to correct errors in the speaker characteristics and emotions estimated during actual use of speech synthesis).

[0016] Each phase will be explained in detail below.

[0017] [Learning Phase] FIG. 2 is a diagram showing an example of a functional configuration of a speech synthesizer 10 in a learning phase according to an embodiment of the present invention. In FIG. 2, the speech synthesizer 10 includes a vector representation acquisition unit 11, a long-term context information extraction unit 12, a speech feature generation unit 13, and a learning unit 14. These units are realized by a process in which one or more programs installed in the speech synthesizer 10 are executed by a processor 104. The speech feature generation unit 13 is one or more machine learning models (e.g., a neural network (DNN (Deep Neural Network))) including a sequence conversion unit 131, an utterance expression extraction unit 132, a discrete expression extraction unit 133, and an utterance expression prediction unit 134. The sequence conversion unit 131 includes an encoder layer 1311 and a decoder layer 1312.

[0018] In the learning phase, the speech feature generating unit 13 is trained (calculation of trainable parameters that can be used by the speech feature generating unit 13). A plurality of training data consisting of pairs of speech features and speech information is used for the training. The speakers and speech information related to the speech features of each training data may be different. One training data corresponds to, for example, a text in a book divided into predetermined processing units. In other words, in this embodiment, reading aloud of a text corresponding to a predetermined processing unit is called a speech. The division into the predetermined processing units may be performed by a machine or may be performed manually. The length of the predetermined processing unit may be measured by the number of sentences or may be in paragraph units. In addition, the predetermined processing unit may be the same or different between the utterance sentence and the narration part as long as it complies with a predetermined standard. It is desirable that the predetermined processing unit is the same in the learning phase and the estimation phase.

[0019] The voice features are, for example, parts corresponding to one utterance among voice parameters (pitch parameters (fundamental frequency, etc.), spectral parameters (mel spectrogram, cepstrum, mel cepstrum, etc.)) obtained by performing voice processing on the voice signal of voice data recorded in advance for learning. The voice data may be recorded by a narrator or the like speaking the main text related to book information. The voice features constituting the learning data are particularly referred to as "teacher voice features."

[0020] Book information refers to information on books such as picture books and picture story shows that are the source of speech required to create voice data (such as stage directions written in picture story shows), and includes text information that is the main text.

[0021] The speech information is information about pronunciation and the like that is assigned to the speech content indicated by the audio data. One piece of speech information is generated for one utterance. The speech information includes at least text information corresponding to the utterance that is included in the book information (i.e., text information indicating the speech content in the audio data). The speech information may also include accent information (accent type, accent phrase length), part of speech information, and information on the start time and end time of each phoneme (phoneme segmentation information). The start time and end time of each phoneme are the elapsed time when the start point (start time point) of each utterance is set to 0 [seconds]. Note that a sound that is composed of one or more accent phrases and has a pause is called a pause phrase, and a sound composed of one or more pause phrases becomes one utterance. Therefore, one piece of speech information may include multiple pieces of accent information.

[0022] Note that information other than the text information added to the speech information (accent information, part of speech information, and information on the start and end times of each phoneme) may be automatically extracted from the voice data.

[0023] FIG. 3 is a flowchart illustrating an example of a first processing procedure executed by the speech synthesizer 10 in the learning phase.

[0024] In step S101, the vector representation acquisition unit 11 converts the speech information constituting the learning data into a language information vector that is an expression (numerical expression) that can be used by the speech feature generation unit 13, and outputs the language information vector. The language information vector is a vector obtained by quantifying characters, morphemes, readings, accents, etc. A known conversion method can be used to convert the speech information into a language information vector. For example, when text information (characters) is used as the speech information, one-hot representation may be used to convert into a language information vector. A one-hot representation vector is a vector whose number of dimensions is the number of characters N included in the speech information, and whose dimension corresponding to the characters constituting the speech information is 1 and whose other dimensions are 0. When phonemes or accents are used as the speech information, the vector representation acquisition unit 11 may convert phonemes, accents, etc. into numerical vectors in the same manner as in Reference 1 (reference information for each reference will be described later). In addition, when characters are used as speech information, the vector expression acquisition unit 11 can also convert phonemes, accents, etc. into numerical vectors in a manner similar to that of Reference 1 by using phoneme and accent information output from the speech information using text analysis.

[0025] Note that information other than the text information added to the speech information (accent information, part of speech information, and information on the start and end times of each phoneme) may be extracted by the learning unit 14 from teacher voice features, for example.

[0026] Next, the speech expression extraction unit 132 converts the teacher speech features into speech expression vectors using trainable parameters (S102). The speech expression vectors are vectors of continuous values ​​with tens to hundreds of dimensions, and correspond to speech expressions such as speaker characteristics and emotions in speech (Reference 1).

[0027] Next, the discrete expression extraction unit 133 uses the learnable parameters to discretize and decode the utterance expression vector output from the utterance expression extraction unit 132 by quantization, thereby generating a discretized utterance expression vector that is a vector corresponding to the utterance expression (S103). For example, a VQ-VAE (vector quantized variational autoencoder) (Reference 2) can be used for this process.

[0028] 4 is a diagram for explaining the discrete representation extraction unit 133 in the first processing procedure of the learning phase. During learning, the discrete representation extraction unit 133 as a VQ-VAE includes an encoder network 1331, a quantizer 1332, and a decoder network 1333.

[0029] In VQ-VAE, the encoder network 1331 acquires (calculates) a latent variable for an input vector or the like. The quantizer 1332 discretizes the latent variable into K symbols (codes) by vector quantization (acquires a codebook consisting of K appropriate symbols for vector quantization). The codebook is a correspondence table between the K symbols (codes) and representative points of the vector as a latent variable. The quantizer 1332 converts the latent variable into a symbol corresponding to a representative point that is closest to the latent variable output from the encoder network 1331 using the codebook. In this embodiment, the symbol is an utterance expression symbol. The decoder network 1333 restores the input vector from the utterance expression symbol. In this embodiment, the restoration result of the input vector by the decoder network 1333 (output from the decoder network 1333) is a discretized utterance expression vector. Therefore, the discretized utterance expression vector is a vector of continuous values, like the utterance expression vector that is the input vector.

[0030] This enables quantization with reduced information loss of the input vector. In this embodiment, the output of the utterance expression extraction unit 132 (a utterance expression vector of continuous values ​​of several tens to several hundreds of dimensions) can be handled as K symbols, so that the user can intuitively control the utterance expression.

[0031] Next, the sequence conversion unit 131 generates (calculates) speech features using trainable parameters based on the linguistic information vector generated in step S101 and the discretized utterance expression vector output from the discrete expression extraction unit 133 (S104). More specifically, the encoder layer 1311 encodes the linguistic information vector. The decoder layer 1312 generates (calculates) speech features based on the output from the encoder layer 1311 and the discretized utterance expression vector.

[0032] Next, the learning unit 14 learns the speech feature generation unit 13 so that the speech feature output from the decoder layer 1312 approaches the teacher speech feature (so that loss based on the speech feature and the teacher speech feature is reduced) (S105). Learning refers to calculating and updating parameters that can be used by the speech feature extraction unit. However, in step S105, parameters that can be used by the sequence conversion unit 131, the utterance expression extraction unit 132, and the discrete expression extraction unit 133 are subject to update, and parameters that can be used by the utterance expression prediction unit 134 are not updated. Steps S101 to S105 are executed for each training data.

[0033] When the learning by the first processing procedure is completed (when the loss based on the speech features output from the decoder layer 1312 and the teacher speech features becomes sufficiently small), the second processing procedure is executed.

[0034] FIG. 5 is a flowchart illustrating an example of a second processing procedure executed by the speech synthesizer 10 in the learning phase.

[0035] Step S201 is the same as step S101 in FIG.

[0036] Next, the long-term context information extraction unit 12 converts text information including the text information of the utterance information of the learning data and text information surrounding the utterance information in the book information (for example, text information of a certain length before and after the utterance information) into a long-term context information vector, and outputs the long-term context information vector (S202). That is, the long-term context information extraction unit 12 converts text information including part of the text information of the utterance information of the learning data into a long-term context information vector indicating the context of the text information. The certain length before and after is, for example, 10 sentences before and after. The conversion from the text information to the long-term context information vector may be performed by using a large-scale language model such as BERT that has been trained in advance from a large amount of text (for example, Reference 3). In this case, the long-term context information extraction unit 12 performs a forward propagation process from the input text information in the same manner as Reference 3 when converting into a vector. The long-term context information extraction unit 12 outputs the information of the output layer finally obtained as a long-term context information vector. Note that the text information surrounding the utterance information is not essential, but is added for the purpose of predicting a speech expression that is more suitable for the utterance information by looking at the long-term context.

[0037] Next, the utterance expression prediction unit 134 uses the learnable parameters to predict (calculate) the likelihood or the posterior probability after softmax (hereinafter referred to as "score") of each of the K utterance expression symbols, which are intermediate information of the discrete expression extraction unit 133, based on the long-term context information vector (S203). The network structure, etc. of the utterance expression prediction unit 134 can be the same as that of Reference 1.

[0038] Next, the discrete expression extraction unit 133 converts the utterance expression symbol with the highest score into a discretized utterance expression vector using the learned parameters (S204). At this time, as shown in Fig. 6, the decoder network 1333 of the discrete expression extraction unit 133 converts the utterance expression symbol into a discretized utterance expression vector.

[0039] Next, the sequence conversion unit 131 uses the learned parameters to generate (calculate) speech features based on the linguistic information vector generated in step S201 and the discretized utterance expression vector output from the discrete expression extraction unit 133 (S205). More specifically, the encoder layer 1311 encodes the linguistic information vector. The decoder layer 1312 generates (calculates) speech features based on the output from the encoder layer 1311 and the discretized utterance expression vector.

[0040] Next, the learning unit 14 updates the parameters used by the utterance expression prediction unit 134 so that the speech features output from the decoder layer 1312 approach the teacher speech features (so that the loss based on the speech features and the teacher speech features becomes small) (S206). Here, the parameters of each unit of the speech feature generation unit 13 other than the utterance expression prediction unit 134 are fixed.

[0041] Steps S201 to S206 are repeated for each piece of learning data.

[0042] In Reference 1, the neural network is trained so that the utterance expression prediction unit 134 can predict a vector of continuous values ​​of tens to hundreds of dimensions (utterance expression vector) output from the utterance expression extraction unit 132. In contrast, in the present embodiment, the discrete expression extraction unit 133 discretizes the vector of continuous values ​​of tens to hundreds of dimensions (utterance expression vector) output from the utterance expression extraction unit 132 to extract a discretized utterance expression vector, and the utterance expression prediction unit 134 is trained to predict an utterance expression symbol, which is different from Reference 1.

[0043] [Estimation Phase] 7 is a diagram showing an example of a functional configuration of the speech synthesizer 10 in the estimation phase according to the embodiment of the present invention. In Fig. 7, the same parts as those in Fig. 2 are given the same reference numerals, and their description will be omitted.

[0044] 7, the speech synthesizer 10 further includes a preprocessing unit 15, a speech output unit 16, a query unit 17, and a reception unit 18. Each of these units is realized by a process executed by the processor 104 of one or more programs installed in the speech synthesizer 10.

[0045] On the other hand, the speech synthesizer 10 does not include the speech expression extraction unit 132 and the learning unit 14.

[0046] 7 is a model that has been trained in the training phase (the second processing procedure has been executed after the first processing procedure). Different computers may be used in the training phase and the estimation phase.

[0047] 8 is a flowchart for explaining an example of a processing procedure executed by the speech synthesizer 10 in the estimation phase. In the estimation phase, as input information for synthesizing speech, a part of text information (hereinafter referred to as "input text") contained in a book such as a picture book or a picture story show to be subjected to speech synthesis (hereinafter referred to as "target book") and text information surrounding the input text in the target book (hereinafter referred to as "input surrounding information") are input. In addition, the book information of the target book is stored in, for example, the auxiliary storage device 102.

[0048] In step S301, the preprocessing unit 15 performs text analysis on the input text, and acquires information equivalent to the utterance information (hereinafter referred to as "target utterance information") by referring to the book information of the target book.

[0049] Next, the vector expression acquisition unit 11 converts the target utterance information into a linguistic information vector that is an expression (numerical expression) that can be used by the speech feature generation unit 13, and outputs the linguistic information vector (S302).

[0050] Next, the long-term context information extraction unit 12 converts the text information including the input text and the input peripheral information into a long-term context information vector, and outputs the long-term context information vector (S303).

[0051] Next, the utterance expression prediction unit 134 predicts (calculates) the score of each of the K utterance expression symbols based on the long-term context information vector using the learned parameters (S304).

[0052] Next, the discrete expression extraction unit 133 converts the utterance expression symbol with the highest score into a discretized utterance expression vector using the learned parameters (S305). The configuration of the discrete expression extraction unit 133 at this time is as shown in FIG.

[0053] Next, the sequence conversion unit 131 uses the learned parameters to generate (calculate) speech features based on the linguistic information vector generated in step S302 and the discretized utterance expression vector output from the discrete expression extraction unit 133 (S306). More specifically, the encoder layer 1311 encodes the linguistic information vector. The decoder layer 1312 generates (calculates) speech features based on the output from the encoder layer 1311 and the discretized utterance expression vector.

[0054] Next, the voice output unit 16 obtains synthetic voice by generating a voice waveform based on the voice feature quantity, and outputs the synthetic voice (S307). For generating the voice waveform, for example, a voice waveform generating method using signal processing as in Reference 4 or a voice waveform generating method using a neural network as in Reference 5 may be used.

[0055] Next, the inquiry unit 17 inquires of the user whether the output voice is appropriate (S308). The appropriateness of the voice refers to whether the voice is as expected by the user. The inquiry to the user may be made using, for example, a GUI (Graphical User Interface).

[0056] When the query unit 17 receives input from the user indicating that the voice is inappropriate (different from expected) (No in S309), the receiving unit 18 receives input of any of K speech expression symbols (a speech expression symbol that is considered to be close to the speech expression expected by the user) from the user (S310).

[0057] In addition, regarding what kind of speech expression each of the K speech expression symbols corresponds to, for example, it is sufficient to assign meaning (labeling) in advance by experiments or the like, and associate the symbols with labels that the user can understand (for example, "joy", "anger", etc.). The user may specify the speech expression symbol by inputting a label. Alternatively, the accepting unit 18 may present the user with a screen 510 including a list of K labels as shown in FIG. 9, and accept the selection of any label from the list. FIG. 9 shows an example of a screen in which the labels, which are information corresponding to each speech expression symbol, are options of radio buttons. At this time, the labels corresponding to the speech expression used in the voice synthesis (i.e., the labels that were not appropriate) may be presented so that the user can understand them. The list may be sorted based on descending order of the score. Alternatively, the list may not include the labels corresponding to the speech expression used in the voice synthesis. The selection of the label is accepted when the user presses the OK button after selecting a radio button. However, the selection of the label may be accepted when the user selects a radio button. When the user directly operates the voice synthesizer 10, the reception unit 18 may display the screen 510 on a display device of the voice synthesizer 10. When the user operates a terminal connected to the voice synthesizer 10 via a network, the reception unit 18 may transmit display data of the screen 510 (e.g., a Web page, etc.) to the terminal. In this case, the reception unit 18 receives the result of the selection made by the user from the terminal.

[0058] In addition to the speech information, the screen 510 may display related information used in the process, such as book information, so that the user can check it together. In this case, the text of the speech information may be modified on the spot by the user, and the synthetic speech may be generated again from the rewritten speech information. In addition, in FIG. 9, the screen 510 includes a "play speech synthetic speech" button. When the "play speech synthetic speech" button is pressed, the voice output unit 16 outputs the synthetic speech at a predetermined time point. For example, the voice output unit 16 may re-output the synthetic speech output in the last step S307 (i.e., the latest synthetic speech at the current stage), or may temporarily store synthetic speeches output in the past and re-output one of the synthetic speeches according to a user's instruction. In this way, the user can re-check the past synthetic speech.

[0059] Next, steps S305 and after are repeated. In step S305, which is executed following step S310, the discrete expression extraction unit 133 converts the utterance expression symbol corresponding to the user's input into a discretized utterance expression vector. Therefore, in this case, it is highly likely that a voice corresponding to the utterance expression expected by the user will be synthesized. In other words, the user can correct the utterance expression and generate a desired synthetic voice while checking the synthetic voice generated as a result.

[0060] As described above, according to this embodiment, information indicating an utterance expression is represented in a form discretized into utterance expression symbols, rather than as a continuous-valued utterance expression vector of tens to hundreds of dimensions. Therefore, it is possible to express the utterance expression estimated during speech synthesis in a form that can be interpreted by the user. As a result, even if an error occurs in the utterance expression prediction unit 134 during speech synthesis and a synthetic speech having an appropriate expression cannot be obtained, the user can make corrections so as to obtain a synthetic speech having an intuitively appropriate expression. In addition, it is possible to generate a synthetic speech that reflects the speaker's characteristics, emotions, etc. of the character, taking into account long-term context, etc. Although the above describes an example of voice synthesis being performed on books such as picture books and paper theaters, this embodiment can also be applied to text information other than books as long as the text information requires reading that expresses the speaker's characteristics and emotions based on the context.

[0061] In the present embodiment, speech synthesis device 10 in the learning phase is an example of a speech synthesis training device.

[0062] [Reference information] [Reference 1] Zen, Heiga, Andrew Senior, and Mike Schuster, "Statistical parametric speech synthesis using deep neural networks", Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013. [Reference 2] VAN DEN OORD, Aaron, et al., "Neural discrete representation learning. Advances in neural information processing systems", 2017, 30. [Reference 3] DEVLIN, Jacob, et al., "Bert: Pre-training of deep bidirectional transformers for language understanding", arXiv preprint arXiv:1810.04805, 2018. [Reference 4] Imai et al., "Mel-Log Spectral Approximation (MLSA) Filter for Speech Synthesis," IEICE Transactions on Information and Communication Engineers, Vol. J66-A, No. 2, pp. 122-129, Feb. 1983. [References 5] Oord, Aaron van den, et al., "Wavenet: A generative model for raw audio", arXiv preprint arXiv:1609.03499 (2016). The following supplementary notes are further disclosed regarding the above embodiment.

[0063] (Additional note 1) Memory, at least one processor coupled to the memory; Including, The processor, outputting one of a plurality of symbols corresponding to the speech expression for a second text including a first text indicating the content of the speech; converting a vector corresponding to an utterance expression based on the speech feature quantity related to the utterance into any one of the symbols, and generating a discretized utterance expression vector, which is a vector corresponding to the utterance expression, from the symbol; generating speech features based on a linguistic information vector based on speech information including the first text and the discretized speech expression vector; using learning data which is a pair of teacher speech features and speech information, updating parameters for generating the discretized utterance expression vector and parameters for generating the speech features so that the discretized utterance expression vector generated for the teacher speech feature and the speech feature generated for the utterance information approach the teacher speech feature; and updating parameters for outputting any one of a plurality of symbols according to the speech expression so that the discretized utterance expression vector generated from the symbol output based on learned parameters for a text partly including a text related to the utterance information and the speech feature generated based on learned parameters for the utterance information approach the teacher speech feature; A speech synthesis training device comprising:

[0064] (Additional note 2) Memory, at least one processor coupled to the memory; Including, The processor, outputting one of a plurality of symbols corresponding to an expression of the utterance from a second text including a first text indicating the content of the utterance; generating a discretized utterance expression vector, which is a vector according to the utterance expression, from the output symbols; generating speech features based on a linguistic information vector based on speech information including the first text and the discretized speech expression vector; During learning, a vector corresponding to an utterance expression based on speech features related to an utterance is converted into any one of the symbols, and the discretized utterance expression vector is generated from the symbol; the parameters for generating the discretized utterance expression vector and the parameters for generating the speech feature are updated by using learning data which is a pair of a teacher speech feature and utterance information, so that the discretized utterance expression vector generated for the teacher speech feature and the speech feature generated for the utterance information approach the teacher speech feature; and the parameters for outputting one of a plurality of symbols according to the utterance expression are updated by using learning data which is a pair of a teacher speech feature and utterance information, so that the discretized utterance expression vector generated from the symbol output based on parameters already learned for a text which partly includes a text related to the utterance information and the speech feature generated based on parameters already learned for the utterance information approach the teacher speech feature. A speech synthesis device comprising:

[0065] (Additional note 3) an utterance expression guessing step of outputting one of a plurality of symbols corresponding to the utterance expression for a second text including a first text indicating the content of the utterance; a discrete expression extraction step of converting a vector corresponding to an utterance expression based on the speech feature quantity related to the utterance into any one of the symbols, and generating a discretized utterance expression vector, which is a vector corresponding to the utterance expression, from the symbol; a sequence transformation step of generating speech features based on a linguistic information vector based on utterance information including the first text and the discretized utterance expression vector; a learning procedure of updating parameters of the discrete expression extraction procedure and the sequence conversion procedure using training data which is a pair of teacher speech features and utterance information so that the discretized utterance expression vector generated by the discrete expression extraction procedure for the teacher speech features and the speech feature generated by the sequence conversion procedure for the utterance information approach the teacher speech features, and updating parameters of the utterance expression prediction procedure so that the discretized utterance expression vector generated by the discrete expression extraction procedure from the symbol output by the utterance expression prediction procedure which has been trained for a text which partly includes a text related to the utterance information and the speech feature generated by the sequence conversion procedure which has been trained for the utterance information approach the teacher speech features; A computer-readable recording medium having a program recorded thereon for causing a computer to execute the above.

[0066] (Additional note 4) an utterance expression guessing step of outputting one of a plurality of symbols corresponding to the utterance expression for a second text including a first text indicating the content of the utterance; a discrete expression extraction step of generating a discretized utterance expression vector, which is a vector corresponding to an utterance expression, from the symbols output by the utterance expression guessing step; a sequence transformation step of generating speech features based on a linguistic information vector based on utterance information including the first text and the discretized utterance expression vector; Run the following on your computer: The discrete expression extraction step includes, during learning, converting a vector corresponding to an utterance expression based on a speech feature amount related to an utterance into any one of the symbols, and generating the discretized utterance expression vector from the symbol; parameters of the discrete expression extraction procedure and the sequence conversion procedure are updated using training data which is a pair of teacher speech features and speech information so that the discretized utterance expression vector generated by the discrete expression extraction procedure for the teacher speech features and the speech feature generated by the sequence conversion procedure for the utterance information approach the teacher speech features, and parameters of the utterance expression estimation procedure are updated so that the discretized utterance expression vector generated by the discrete expression extraction procedure from the symbol output by the utterance expression estimation procedure which has been trained for a text which partly includes a text related to the utterance information and the speech feature generated by the sequence conversion procedure which has been trained for the utterance information approach the teacher speech features. A computer-readable recording medium having a program recorded thereon, Although the embodiment of the present invention has been described in detail above, the present invention is not limited to such specific embodiment, and various modifications and variations are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]

[0067] 10. Voice synthesizer 11 Vector Representation Acquisition Unit 12 Long-term context information extraction part 13 Speech feature generation unit 14 Learning Department 15 Pretreatment section 16 Audio output section 17 Enquiry Section 18 Reception Department 100 Drive device 101 Recording media 102 Auxiliary storage 103 Memory device 104 processors 105 Interface device 131 Series Conversion Unit 132 Speech expression extraction unit 133 Discrete expression extraction part 134 Speech Expression Prediction Unit 1311 Encoder Layer 1312 Decoder Layer 1331 Encoder Network 1332 Quantization section 1333 Decoder Network B Bus

Claims

1. an utterance expression guessing unit configured to output one of a plurality of symbols corresponding to an utterance expression for a second text including a first text indicating the content of an utterance; a discrete expression extraction unit configured to convert a vector corresponding to an utterance expression based on a speech feature quantity related to the utterance into any one of the symbols and generate a discretized utterance expression vector, which is a vector corresponding to the utterance expression, from the symbol; a sequence conversion unit configured to generate speech features based on a linguistic information vector based on utterance information including the first text and the discretized utterance expression vector; a learning unit configured to update parameters of the discrete expression extraction unit and the sequence conversion unit using training data that is a pair of teacher speech features and utterance information so that the discretized utterance expression vector generated by the discrete expression extraction unit for the teacher speech features and the speech feature generated by the sequence conversion unit for the utterance information approach the teacher speech features, and to update parameters of the utterance expression prediction unit so that the discretized utterance expression vector generated by the discrete expression extraction unit from the symbol output by the utterance expression prediction unit that has been trained for a text that partially includes a text related to the utterance information and the speech feature generated by the sequence conversion unit that has been trained for the utterance information approach the teacher speech features; A voice synthesis training device comprising:

2. an utterance expression guessing unit configured to output one of a plurality of symbols corresponding to an utterance expression for a second text including a first text indicating the content of an utterance; a discrete expression extraction unit configured to generate a discretized utterance expression vector, which is a vector according to an utterance expression, from the symbol output by the utterance expression prediction unit; a sequence conversion unit configured to generate speech features based on a linguistic information vector based on utterance information including the first text and the discretized utterance expression vector; having the discrete expression extraction unit is configured to convert a vector corresponding to an utterance expression based on a speech feature amount related to an utterance into any one of the symbols during learning, and generate the discretized utterance expression vector from the symbol; parameters of the discrete expression extraction unit and the sequence conversion unit are trained using training data which is a pair of teacher speech features and utterance information so that the discretized utterance expression vector generated by the discrete expression extraction unit for the teacher speech features and the speech feature generated by the sequence conversion unit for the utterance information approach the teacher speech features, and parameters of the utterance expression prediction unit are trained so that the discretized utterance expression vector generated by the discrete expression extraction unit from the symbol output by the utterance expression prediction unit which has been trained for a text which partly includes a text related to the utterance information and the speech feature generated by the sequence conversion unit which has been trained for the utterance information approach the teacher speech features. A speech synthesis device comprising:

3. a reception unit configured to receive an input of information corresponding to any one of the symbols from a user, the discrete expression extraction unit is configured to generate a discretized utterance expression vector according to an utterance expression from the symbol corresponding to the information accepted by the acceptance unit.

3. The speech synthesis device according to claim 2.

4. an output unit configured to output a synthetic speech based on the speech features generated by the sequence conversion unit; having the receiving unit is configured to receive information corresponding to any one of the symbols from a user after the synthetic voice is output.

4. The speech synthesis device according to claim 3.

5. an utterance expression guessing step of outputting one of a plurality of symbols corresponding to the utterance expression for a second text including a first text indicating the content of the utterance; a discrete expression extraction step of converting a vector corresponding to an utterance expression based on the speech feature quantity related to the utterance into any one of the symbols, and generating a discretized utterance expression vector, which is a vector corresponding to the utterance expression, from the symbol; a sequence transformation step of generating speech features based on a linguistic information vector based on utterance information including the first text and the discretized utterance expression vector; a learning procedure of updating parameters of the discrete expression extraction procedure and the sequence conversion procedure using training data which is a pair of teacher speech features and utterance information so that the discretized utterance expression vector generated by the discrete expression extraction procedure for the teacher speech features and the speech feature generated by the sequence conversion procedure for the utterance information approach the teacher speech features, and updating parameters of the utterance expression prediction procedure so that the discretized utterance expression vector generated by the discrete expression extraction procedure from the symbol output by the utterance expression prediction procedure which has been trained for a text which partly includes a text related to the utterance information and the speech feature generated by the sequence conversion procedure which has been trained for the utterance information approach the teacher speech features; A speech synthesis training method comprising:

6. an utterance expression guessing step of outputting one of a plurality of symbols corresponding to the utterance expression for a second text including a first text indicating the content of the utterance; a discrete expression extraction step of generating a discretized utterance expression vector, which is a vector corresponding to an utterance expression, from the symbols output by the utterance expression guessing step; a sequence transformation step of generating speech features based on a linguistic information vector based on utterance information including the first text and the discretized utterance expression vector; The computer executes The discrete expression extraction step includes, during learning, converting a vector corresponding to an utterance expression based on a speech feature amount related to an utterance into any one of the symbols, and generating the discretized utterance expression vector from the symbol; parameters of the discrete expression extraction procedure and the sequence conversion procedure are trained using training data which is a pair of teacher speech features and utterance information so that the discretized utterance expression vector generated by the discrete expression extraction procedure for the teacher speech features and the speech feature generated by the sequence conversion procedure for the utterance information approach the teacher speech features, and parameters of the utterance expression estimation procedure are trained so that the discretized utterance expression vector generated by the discrete expression extraction procedure from the symbol output by the utterance expression estimation procedure which has been trained for a text which partly includes a text related to the utterance information and the speech feature generated by the sequence conversion procedure which has been trained for the utterance information approach the teacher speech features.

13. A speech synthesis method comprising:

7. 6. A program for causing a computer to execute the voice synthesis training method according to claim 5.

8. 7. A program for causing a computer to execute the speech synthesis method according to claim 6.