Editing device, voice synthesis device, and program
The editing device and speech synthesis system address the challenge of accurately estimating accents in Japanese voice synthesis by enabling users to intuitively edit accent information for ruby characters, resulting in improved speech synthesis quality.
Patent Information
- Application Number
- JP2021074758
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-04-27
AI Technical Summary
Existing voice synthesis technologies face challenges in accurately estimating accents for Japanese text data due to the language's complex readings and lack of regularity in accents, leading to errors in speech synthesis.
An editing device and speech synthesis system that allow users to easily edit accent information for ruby characters in text data by displaying characters and prosodic symbols in a manner that indicates accent nuclei, allowing for intuitive correction of accents.
Enables accurate and efficient editing of accent information, improving the quality of speech synthesis by allowing users to intuitively correct accent positions and types, thereby enhancing the overall accuracy and naturalness of synthesized speech.
Smart Images

Figure 0007685869000001 
Figure 0007685869000002 
Figure 0007685869000003
Abstract
Description
Technical Field
[0001] The present invention relates to an editing device, a voice synthesis device, and a program.
Background Art
[0002] There is a technique for performing voice synthesis based on text data of an intermediate language in which reading kana and prosodic symbols are described (see, for example, Patent Document 1 and Non-Patent Document 1). In particular, voice synthesis using deep learning is being used in practice (see, for example, Non-Patent Document 2). For generating the intermediate language used in voice synthesis, for example, the results of language analysis according to the prior art can be used. On the other hand, when creating pronunciation information from text, there is a technique for interactively correcting the pronunciation information (see, for example, Patent Document 2).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] The accents estimated by language analysis contain quite a few errors. This is because Japanese has multiple readings and the accents have no regularity and are difficult to estimate. In order to perform high-quality speech synthesis using the text data of the intermediate language as input, it is necessary to correct the text data so that it represents appropriate accents for the intermediate language obtained by language analysis. The technique of Patent Document 2 provides a user interface for moving the accent nucleus of the text, but the relationship between the high and low accents and the ruby characters is not clear.
[0006] The present invention has been made in consideration of such circumstances, and provides an editing device, a speech synthesis device, and a program that can easily edit accent information for ruby characters described in text data used for speech synthesis processing.
Means for Solving the Problems
[0007] [1] One aspect of the present invention is based on text data in which characters representing the reading of speech content, first prosodic symbols representing accents, and second prosodic symbols representing reading breaks are described, and the characters representing the reading and the prosodic display objects representing the second prosodic symbols are displayed on a display unit in the order of appearance in the text data, and an accent display object that is the accent of the reading represented by the characters and that represents the accent indicated by the first prosodic symbol is displayed so as to overlap or be associated with the characters displayed on the display unit. A display control unit, and when information for selecting any of the characters displayed on the display unit as an accent nucleus is input, among the partial text data composed of the characters and the first prosodic symbols delimited by the second prosodic symbol in the text data, the partial text data including the selected character that is the character selected as the accent nucleus is set as processing target data, and a rewriting unit that rewrites the first prosodic symbol included in the processing target data so as to represent an accent corresponding to the position of the selected character in the processing target data. An editing device characterized by comprising:
[0008] [2] One aspect of the present invention is the above-described editing device, wherein the display control unit displays the characters indicated as accent nuclei based on the first prosodic symbol on the display unit in a manner indicating that they are accent nuclei.
[0009] [3] One aspect of the present invention is the above-described editing device, wherein the accent display object representing a low accent is a line displayed below a predetermined height in the display of the characters, and the accent display object representing a high accent is a line displayed above the predetermined height.
[0010] [4] One aspect of the present invention is the above-described editing device, wherein the display control unit connects and displays the lines that are the accent display objects corresponding to the characters for each partial text data in the order of appearance of the characters.
[0011] [5]One aspect of the present invention is the above-described editing device, wherein the rewriting unit receives an input of information indicating a change in reading or a delimiter of reading, and rewrites the character or the second prosodic symbol included in the text data based on the input information.
[0012] [6]One aspect of the present invention is the above-described editing device, wherein the first prosodic symbol represents an accent rise or an accent fall, and the second prosodic symbol represents an accent delimiter, an end of a sentence, or a pause.
[0013] [7]One aspect of the present invention is a speech synthesis device including: a display control unit that displays, in the order of appearance in text data, characters representing a reading of speech content, first prosodic symbols representing accents, and second prosodic symbols representing reading delimiters, and displays, superimposed on or associated with the characters displayed on the display unit, an accent display object representing the accent of the reading represented by the characters and indicated by the first prosodic symbols; a rewriting unit that, when information for selecting any of the characters displayed on the display unit as an accent nucleus is input, sets, as processing target data, partial text data including the selected character as the accent nucleus and consisting of the characters and the first prosodic symbols delimited by the second prosodic symbols in the text data, and rewrites the first prosodic symbols included in the processing target data so as to represent an accent corresponding to the position of the selected character in the processing target data; an acoustic feature amount estimation unit that estimates an acoustic feature amount based on the text data rewritten by the rewriting unit; and a vocoder unit that estimates a speech waveform using the acoustic feature amount estimated by the acoustic feature amount estimation unit.
[0014] [8]One aspect of the present invention is a program for causing a computer to function as any of the above-described editing devices. [Advantages of the Invention]
[0015] According to the present invention, it is possible to easily edit accent information for the ruby written in the text data used for speech synthesis processing.
Brief Description of Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0017] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0018] FIG. 1 is a diagram showing the configuration of a voice synthesis device 1 according to an embodiment of the present invention. The voice synthesis device 1 is an example of an editing device. The voice synthesis device 1 can be realized by, for example, a server computer, a personal computer, a tablet terminal, a smartphone, smart glasses, a smartwatch, an embedded device, or the like. The voice synthesis device 1 includes a language analysis unit 2, an editing unit 3, a display unit 4, an input unit 5, and a voice synthesis unit 6.
[0019] The language analysis unit 2 converts text data of a Japanese text with a mixture of kana and kanji into an intermediate language using kana and prosodic symbols. This conversion can be performed by existing techniques such as morphological analysis. Hereinafter, the text data of a Japanese text with a mixture of kana and kanji will be referred to as original text data, and the text data of the intermediate language will be referred to as intermediate language data. Kana is an example of a character representing pronunciation and corresponds to a mora. The kana representing pronunciation in the intermediate language data is also referred to as reading kana. In the present embodiment, the case of using katakana as kana is described, but hiragana, alphabet, pronunciation symbols, or symbols representing phonemes may be used instead of kana. The prosodic symbols used in the intermediate language data are characters or symbols representing prosody. Characters different from the characters representing pronunciation are used as the characters representing prosody. The prosodic symbols include a first prosodic symbol and a second prosodic symbol. The first prosodic symbol represents accent. The second prosodic symbol represents, for example, the boundary of an accent phrase, the end of a sentence, a pause, or other boundaries of pronunciation. The text data of the intermediate language consisting of kana and the first prosodic symbol separated by the second prosodic symbol is referred to as accent phrase intermediate language data. The reading kana included in the accent phrase intermediate language data corresponds to an accent phrase.
[0020] The editing unit 3 includes a storage unit 31, a display control unit 32, and a rewriting unit 33. The editing unit 3 may be provided, for example, by a web browser or as a computer application. The storage unit 31 stores intermediate language data. The storage unit 31 may further store source text data. The display control unit 32 displays the ruby text and the rhythm display object representing the second rhythm symbol on the display unit 4 in the order of appearance in the intermediate language data. The rhythm display object may be a character different from the kana of the character representing the reading, a symbol, or a graphic. Further, the display control unit 32 displays an accent display object representing the pitch of the accent represented by the first rhythm symbol superimposed on or associated with the ruby text displayed on the display unit 4. In addition, the display control unit 32 displays the ruby text indicated as the accent nucleus by the first rhythm symbol in a manner indicating that it is the accent nucleus. Specifically, an accent nucleus display object indicating that it is the accent nucleus may be displayed superimposed on or associated with the ruby text having the accent nucleus, or the color, thickness, or background color of the ruby text at the position where the accent nucleus is present may be displayed differently from other ruby texts. Also, the display control unit 32 may display the source text data corresponding to the intermediate language data on the display unit 4.
[0021] When accent nucleus selection information for selecting any of the ruby texts displayed on the display unit 4 as the accent nucleus is input by the input unit 5, the rewriting unit 33 sets the accent phrase intermediate language data including the selected ruby text as the data to be processed. Hereinafter, the ruby text selected as the accent nucleus will be referred to as the selected character. The rewriting unit 33 changes the first rhythm symbol included in the data to be processed so as to represent an accent determined according to the position of the selected character in the data to be processed, and rewrites the intermediate language data. Also, when the rewriting unit 33 receives an input of information indicating a change in the ruby text, it rewrites the ruby text included in the intermediate language data to the input ruby text. Further, when the rewriting unit 33 receives an input of information indicating a change in the reading delimiter, it changes the position and type of the second rhythm symbol included in the intermediate language data, or inserts or deletes the second rhythm symbol according to the input information.
[0022] The display unit 4 displays data. The display unit 4 is, for example, an image display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro Luminescence) display. The display unit 4 may also be a head-mounted display, a retinal projection display, or the like. The display unit 4 may be an interface for connecting an image display device to the voice synthesizer 1. In this case, the display unit 4 generates a video signal for displaying data and outputs the video signal to the image display device connected to itself. Further, the display unit 4 may display data on an information processing device connected to the voice synthesizer 1.
[0023] The input unit 5 inputs a user's instruction. The input unit 5 is configured using existing input devices such as a keyboard, a pointing device (mouse, tablet, etc.), buttons, or a touch panel. The input unit 5 is operated by the user when inputting the user's instruction to the voice synthesizer 1. Further, the input unit 5 may input the user's voice by voice recognition. The input unit 5 may be an interface for connecting an input device to the voice synthesizer 1. In this case, the input unit 5 inputs an input signal generated according to the user's input in the input device to the voice synthesizer 1. Further, the input unit 5 may receive an instruction input by the user from an information processing device connected to the voice synthesizer 1.
[0024] The speech synthesis unit 6 performs speech synthesis using the intermediate language data as input data. For the speech synthesis unit 6, for example, in addition to the techniques described in Patent Document 1, Non-Patent Documents 1 and 2, the technique described in Japanese Unexamined Patent Application Publication No. 2018-146803, and Reference 1 "Yoshi Hashimoto, Shinji Takagi, 'Statistical Speech Synthesis Based on Deep Learning', Journal of the Acoustical Society of Japan, Vol. 73, No. 1, 2017, pp. 55-62", and Reference 2 "Kiyoshi Kurihara, Nobumasa Kiyama, Tadashi Kumano, 'Effectiveness of a Sequence-to-Sequence Acoustic Feature Estimation Method That Does Not Require Labeling Work', The Institute of Electronics, Information and Communication Engineers, Technical Report, Vol. 119, No. 321, SP2019-37, 2019" can be used. The speech synthesis unit 6 includes an acoustic feature amount estimation unit 61 and a vocoder unit 62. The acoustic feature amount estimation unit 61 estimates an acoustic feature amount based on the intermediate language data input from the editing unit 3. The vocoder unit 62 estimates a speech waveform using the acoustic feature amount estimated by the acoustic feature amount estimation unit 61.
[0025] FIG. 2 is a diagram showing prosodic symbols used in the intermediate language data of the present embodiment. The prosodic symbols shown in FIG. 2 are information obtained by modifying the prosodic symbols described in Reference 3, "Acoustic Input / Output Method Standardization Committee, JEITA Standard IT-4006 Symbols for Japanese Text-to-Speech Synthesis, The Institute of Electronics, Information and Communication Engineers, 2010, p. 4-10". The prosodic information includes types such as specification of accent position, specification of clause / phrase delimiter, specification of end-of-sentence intonation, and specification of pause. The prosodic symbols representing the specification of accent position include an accent rise symbol "^" and an accent fall symbol "!". The accent rise symbol "^" indicates that the accent rises on the kana (mora) immediately following the symbol. The accent fall symbol "!" represents that the accent falls on the kana (mora) immediately following the symbol. The accent rise symbol "^" and the accent fall symbol "!" are the first prosodic symbols. For the specification of clause / phrase delimiter, a prosodic symbol "#" representing the delimiter of an accent clause is used. For the specification of end-of-sentence intonation, a prosodic symbol "=" representing a normal end of sentence, a prosodic symbol "(" representing an end of sentence with a noun stop, and a prosodic symbol "?" representing an end of sentence with a question are used. For the specification of pause, a prosodic symbol "," representing a pause is used. The prosodic symbol "#" representing the delimiter of an accent clause, the prosodic symbol "=" representing a normal end of sentence, the prosodic symbol "(" representing an end of sentence with a noun stop, the prosodic symbol "?" representing an end of sentence with a question, and the prosodic symbol "," representing a pause are the second prosodic symbols representing the reading delimiter. Note that these prosodic symbols are examples, and other symbols may be used.
[0026] FIG. 3 is a diagram showing an example of intermediate language data according to the present embodiment. FIG. 3(a) shows the original data of a sentence with a mixture of kana and kanji. Actually, the original data does not include the information on the pitch of the accent, but in FIG. 3(a), the correct pitch of the accent is superimposed and shown by the height of the line. FIG. 3(b) shows the intermediate language data obtained by the language analysis unit 2 of the speech synthesis device 1 performing morphological analysis on the original data shown in FIG. 3(a). For example, the language analysis unit 2 converts the sentence with a mixture of kanji and kana shown in the original data into full context label data by an existing technique. The full context label data includes information on phonemes in the utterance, information on phonemes before and after the phoneme, accent phrase information of the phoneme, and the like. The accent phrase information indicates features related to the accent phrase in which the current phoneme is included in the utterance, and features related to the accent phrase adjacent to the accent phrase. The language analysis unit 2 extracts information on phonemes from the full context label data, and adds a prosody symbol representing the prosody obtained based on the phonemes and accent phrase information shown in the full context label data to a character string consisting of reading kana corresponding to the readings represented by the phonemes in the order of appearance, thereby generating intermediate language data.
[0027] When the accents and the like shown in the intermediate language data generated by the language analysis unit 2 are different, the user corrects the intermediate language data. The editing unit 3 of the speech synthesis device 1 corrects the intermediate language data shown in FIG. 3(b) according to the user input, and generates the intermediate language data shown in FIG. 3(c). The dotted line part indicates that the correction has been made. The speech synthesis device 1 outputs the corrected intermediate language data shown in FIG. 3(c) to the speech synthesis unit 6. The speech synthesis unit 6 performs speech synthesis using the intermediate language data shown in FIG. 3(c) as an input.
[0028] FIG. 4 is a diagram showing an accent display example of an accent correction interface provided by conventional speech synthesis software. As shown in FIGS. 4(a) and 4(b), in the conventional accent correction interface, the reading kana and the accent display object representing the pitch of the accent of the reading kana are displayed in different columns. The accent display object in FIG. 4(a) is a line connecting lines displayed at a high position or a low position corresponding to the pitch of the accent for each accent phrase. The vertical line segment between the high-position line and the low-position line represents that the pitch of the accent changes. The accent display object in FIG. 4(b) is a circle displayed at a high position or a low position corresponding to the pitch of the accent. The line segment between the high-position circle and the low-position circle represents that the pitch of the accent changes. The circles between the accent phrases are displayed with a space in between.
[0029] FIG. 5 is a diagram showing an accent correction operation using a conventional correction interface. In the conventional correction interface, the user corrects the accent by performing an operation of changing the position of the accent display object corresponding to the reading kana of the accent correction target from a high position to a low position or from a low position to a high position for each reading kana. For example, the accent of the accent phrase "arayuru" shown in FIG. 5(a) is changed to the accent shown in FIG. 5(b). In this case, the user uses a mouse or the like to correct the accent display object corresponding to "a" from a high position to a low position as shown by reference sign A1, corrects the accent display object corresponding to "ra" from a low position to a high position as shown by reference sign A2, and further corrects the accent display object corresponding to "yu" from a low position to a high position as shown by reference sign A3.
[0030] FIG. 6 is a diagram showing a display example of the accent correction interface provided by the speech synthesizer 1 of the present embodiment. FIG. 6(a) shows the display before the intermediate language data is corrected, and FIG. 6(b) shows the display after the intermediate language data is corrected. The display control unit 32 of the speech synthesizer 1 of the present embodiment divides the intermediate language data by the second prosodic symbol in order to display the reading kana separately for each accent phrase, and generates accent phrase intermediate language data. The accent phrase intermediate language data includes the reading kana and the first prosodic data. The display control unit 32 extracts the reading kana in the order of appearance from the accent phrase intermediate language data to obtain an accent phrase, and further obtains the second prosodic symbol set immediately after the accent phrase intermediate language data from the intermediate language data. The display control unit 32 displays the accent phrase obtained from the accent phrase intermediate language data and the delimiter object representing the second prosodic symbol set immediately after the accent phrase intermediate language data in the order of appearance in the intermediate language data. In FIGS. 6(a) and (b), a delimiter object "_" corresponding to the prosodic symbol "#" representing the delimiter of the accent phrase is displayed immediately after the accent phrase "arayuru". Also, a delimiter object "," corresponding to the prosodic symbol "," representing a pause is displayed immediately after the accent phrase "genjitsu o".
[0031] Furthermore, the display control unit 32 determines the accent nucleus of each accent phrase. When the accent phrase intermediate language data includes an accent drop symbol "!", the accent nucleus is the reading kana immediately before the accent drop symbol. When the accent phrase intermediate language data does not include an accent drop symbol, the accent nucleus is the last kana of the accent phrase. The display control unit 32 displays an accent nucleus display object B1 indicating the accent nucleus of each accent phrase above the reading kana having the accent nucleus.
[0032] Reference 4, "Nobuaki Mine Matsumatsu, 'OJAD and Voice Guidance Using It', [online], <URL:https: / / www.gavo.t.u-tokyo.ac.jp / ~mine / japanese / acoustics / OJAD_workshop_long.pdf>", and Reference 5, "Hiroya Fujisaki and Keikichi Hirose, 'Analysis of voice fundamental frequency contours for declarative sentences of Japanese', 1984, [online], <URL: https: / / www.jstage.jst.go.jp / article / ast1980 / 5 / 4 / 5_4_233 / _pdf / -char / en>", describe the principle for identifying the pitch accents of the Tokyo dialect (standard language) of Japanese. This principle shows that the accent type, which is the pattern of the pitch of the accents for each mora in the accent phrase, is uniquely determined by which mora in the accent phrase has the accent nucleus. This is based on the following rules: (1) the pitch of the accents of the first mora and the second mora in the accent phrase are different; (2) the mora with the accent nucleus has a high accent, and the mora following the accent nucleus has a low accent; and (3) once the accent becomes low in the accent phrase, the accent does not rise in that accent phrase. That is, there are as many accent types as the number of moras, and the accent type is uniquely determined by the position of the mora with the accent nucleus. Therefore, the display control unit 32 superimposes and displays an accent display object B2 representing the accent of the accent type determined according to the accent nucleus on the display of each accent phrase.
[0033] When the user changes the accent of an accent phrase, the user inputs to the speech synthesizer 1 the position of the mora with the accent nucleus among the furigana in the accent phrase. For example, the accent of the accent phrase "arayuru" shown in Fig. 6(a) is changed to the accent shown in Fig. 6(b). In this case, the user clicks on the display area B3 of "yu" which has the accent nucleus in the accent phrase "arayuru" or the area B4 above the display area B3 with a mouse or the like for selection. It is also possible to select the mora of the accent nucleus with the area B5 below the display area B3. The rewriting unit 33 of the speech synthesizer 1 identifies the accent type in the accent phrase including the selected mora based on the position of the selected mora with the accent nucleus. The rewriting unit 33 rewrites the first prosodic symbol included in the intermediate language data according to the identified accent type. Further, the display control unit 32 changes the display of the accent nucleus display object B1 and the accent display object B2 as shown in Fig. 6(b) with the rewritten intermediate language data.
[0034] In this way, the speech synthesizer 1 realizes the control of the accent described in the intermediate language data used for speech synthesis by the user clicking once on the display of the mora of the accent nucleus or the upper part thereof. Further, since the speech synthesizer 1 displays the furigana and the accent display object, which were conventionally displayed in two lines as shown in Fig. 4, in one line, a lot of information can be displayed on the screen.
[0035] Fig. 7 is a diagram showing the determination of the accent type by the speech synthesizer 1. In Fig. 7, the case where the number of moras of the accent phrase is 5 is shown as an example. Fig. 7(a) shows the accent types of type 1 to type 5 represented by the conventional accent correction interface. Fig. 7(b) shows the accent nuclei corresponding to the accent types of type 1 to type 5. Fig. 7(c) shows an example of the accent display object displayed by the accent correction interface of the speech synthesizer 1 when the accent nucleus is specified as shown in Fig. 7(b).
[0036] As shown in FIG. 7, the pitch of the accent in the accent phrase is determined by the position of the accent nucleus. Therefore, the user inputs, via the input unit 5, designation information for designating the character in the reading in which the accent nucleus exists, with respect to the character string of the accent phrase displayed on the display unit 4. The rewriting unit 33 sets, as target data for processing, the accent phrase intermediate language data including the reading kana indicated by the designation information, and determines an accent type based on the position of the character within the accent phrase designated by the user as the accent nucleus in the target data for processing. The rewriting unit 33 changes the intermediate language data according to the determined accent type. That is, the rewriting unit 33 changes the descriptions of the accent rising symbol and the accent falling symbol included in the target data for processing according to the determined accent type.
[0037] Specifically, the rewriting unit 33 deletes the accent falling symbol and the accent rising symbol from the target data for processing. When the accent nucleus is the first reading kana (type 1), the rewriting unit 33 inserts an accent falling symbol immediately after the first reading kana. When the accent nucleus is neither the first reading kana nor the last reading kana in the target data for processing (types 2 to 4), the rewriting unit 33 inserts an accent rising symbol immediately after the first reading kana of the target data for processing, and further inserts an accent falling symbol immediately after the reading kana in which the accent nucleus is designated. When the accent nucleus is the last reading kana of the target data for processing (type 5), the rewriting unit 33 inserts an accent rising symbol immediately after the first reading kana of the target data for processing, and does not describe an accent falling symbol. The display control unit 32 changes the display of the accent display object based on the target data for processing after rewriting.
[0038] FIG. 8 is a flowchart showing the speech synthesis process of the speech synthesizer 1. The language analysis unit 2 of the speech synthesizer 1 acquires original data of a sentence including kana and kanji representing the utterance content (step S110). The language analysis unit 2 may receive the original data from the outside, may read it from a recording medium, or may acquire the original data input by the user via the input unit 5. The sentence representing the utterance content may be one sentence or a plurality of sentences.
[0039] The language analysis unit 2 performs morphological analysis on the sentence indicated by the acquired original text data, and converts the sentence representing the utterance content into intermediate language data described by a character string using reading kana and prosodic symbols (step S120). The language analysis unit 2 outputs the intermediate language data to the editing unit 3. The language analysis unit 2 may output the original text data and the intermediate language data to the editing unit 3. In this case, the language analysis unit 2 adds information for associating each sentence included in the original text data with the sentence of the intermediate language data corresponding to that sentence to the original text data. The editing unit 3 stores the intermediate language data and the original text data in the storage unit 31.
[0040] The display control unit 32 of the editing unit 3 displays the intermediate language data on the display unit 4. The display control unit 32 may further display the original text data associated with the intermediate language data on the display unit 4. The editing unit 3 modifies the intermediate language data stored in the storage unit 31 according to the instruction input by the user via the input unit 5 (step S130). Details of the processing will be described later with reference to FIG. 9.
[0041] The editing unit 3 determines whether the user has input an instruction for voice synthesis via the input unit 5 (step S140). When the editing unit 3 determines that no instruction for voice synthesis has been input (step S140: NO), it performs the processing of step S160. On the other hand, when the editing unit 3 determines that an instruction for voice synthesis has been input (step S140: YES), it reads out the modified intermediate language data from the storage unit 31 and outputs it to the voice synthesis unit 6. When the user inputs an instruction for voice synthesis specifying a part of the intermediate language data via the input unit 5, the editing unit 3 reads out the specified part of the intermediate language data from the storage unit 31 and outputs it to the voice synthesis unit 6. The voice synthesis unit 6 performs voice synthesis using the intermediate language data input from the editing unit 3 (step S150).
[0042] The voice synthesis device 1 determines whether the user has input an end (step S160). When the voice synthesis device 1 determines that no end has been input (step S160: NO), it repeats the processing from step S130, and when it determines that an end has been input (step S160: YES), it ends the processing of FIG. 8.
[0043] FIG. 9 is a flowchart showing the intermediate language data correction process of the speech synthesis device 1. FIG. 9 shows the detailed process of the speech synthesis device 1 in step S130. The display control unit 32 reads the intermediate language data from the storage unit 31 and displays it on the display unit 4 (step S210). That is, the display control unit 32 divides the intermediate language data into accent phrase intermediate language data by the second prosodic symbol. The display control unit 32 displays the accent phrase of the reading kana included in the accent phrase intermediate language data and the prosodic display object representing the second prosodic symbol on the display unit 4 in the order of appearance in the intermediate language data. Further, the display control unit 32 determines the accent nucleus and the pitch of the accent of each accent phrase based on the first prosodic symbol included in each accent phrase intermediate language data.
[0044] Specifically, when there is an accent drop symbol immediately after the first reading kana of the accent phrase intermediate language data, the display control unit 32 determines that the first reading kana is the accent nucleus, and determines that the first reading kana has a high accent, and the subsequent reading kana from the second reading kana to the last reading kana has a low accent.
[0045] Also, when there is an accent drop symbol immediately after the second and subsequent reading kana of the accent phrase intermediate language data, the display control unit 32 determines that the reading kana immediately before the accent drop symbol is the accent nucleus, and determines that the first reading kana has a low accent, the reading kana from the second reading kana to the reading kana immediately before the accent drop symbol has a high accent, and the reading kana from the next reading kana of the accent drop symbol to the last reading kana has a low accent. When the accent nucleus is in the second and subsequent reading kana from the last to the second of the accent phrase intermediate language data, there is an accent rise symbol immediately before the second reading kana. Therefore, the display control unit 32 may determine that the reading kana before the accent rise symbol has a low accent, and the reading kana from the next reading kana of the accent rise symbol to the reading kana immediately before the accent drop symbol has a high accent.
[0046] In addition, when there is no accent drop symbol in the accent phrase intermediate language data, the display control unit 32 determines that the last reading kana of the accent phrase is the accent nucleus, and determines that the first reading kana has a low accent and the second to the last reading kana have high accents. When the last reading kana is the accent nucleus, although there is no accent drop symbol in the accent phrase intermediate language data as described above, there is an accent rise symbol immediately before the second reading kana. Therefore, the display control unit 32 may determine that the reading kana before the accent rise symbol has a low accent, and the reading kana from the next reading kana of the accent rise symbol to the reading kana immediately before the second prosody symbol has a high accent.
[0047] The display control unit 32 displays an accent display object representing the accent of the reading kana, superimposed on or associated with the reading kana. Further, the display control unit 32 displays the character indicated as being the accent nucleus in a manner representing that it is the accent nucleus. The display control unit 32 may further display the original text data on the display unit 4. Note that when the user inputs information designating a part of the original text data as a display target through the input unit 5, the display control unit 32 may perform the process of step S210 on the intermediate language data corresponding to the part of the original text data that is the display target.
[0048] The rewriting unit 33 determines whether accent nucleus selection information for selecting a reading kana having an accent nucleus is input by the input unit 5 (step S220). When the rewriting unit 33 determines that the accent nucleus selection information has been input (step S220: YES), the accent phrase intermediate language data including the selected character which is the reading kana indicated by the accent nucleus selection information is set as the data to be processed. The rewriting unit 33 determines the accent type based on the position of the selected character in the data to be processed. The rewriting unit 33 rewrites the intermediate language data stored in the storage unit 31 so as to describe the first prosodic symbol in the data to be processed according to the determined accent type (step S230). That is, the rewriting unit 33 deletes the accent descending symbol and the accent ascending symbol from the data to be processed. When the accent nucleus is the first reading kana, the rewriting unit 33 inserts an accent descending symbol immediately after the first reading kana. When the accent nucleus is the last reading kana of the data to be processed, the rewriting unit 33 inserts an accent ascending symbol immediately after the first reading kana of the data to be processed. When the accent nucleus is a reading kana from the second to the second last, the rewriting unit 33 inserts an accent ascending symbol immediately after the first reading kana of the data to be processed, and inserts an accent descending symbol immediately after the reading kana where the accent nucleus is specified. The display control unit 32 displays the intermediate language data in which the rewriting in step S230 has been performed by the same process as in step S210 (step S240).
[0049] The rewriting unit 33 determines whether an instruction to end the correction has been input (step S250). When the display control unit 32 determines that the instruction to end the correction has not been input (step S250: NO), the process from step S220 is repeated.
[0050] When the rewriting unit 33 determines that no accent nucleus selection information has been input (step S220: NO), it determines whether a correction of the reading kana or the second prosodic symbol has been input by the input unit 5 (step S260). When the rewriting unit 33 determines that a correction of the reading kana or the second prosodic symbol has been input (step S260: YES), it corrects the intermediate language data stored in the storage unit 31 based on the input content (step S270). For example, when the rewriting unit 33 is input with the reading kana to be corrected and the corrected reading kana, it corrects the intermediate language data by replacing the reading kana to be corrected with the corrected reading kana. Also, when the rewriting unit 33 is input with the second prosodic symbol to be corrected and the destination position, it corrects the intermediate language data by moving the second prosodic symbol to be corrected to the destination position. When the rewriting unit 33 is input with the second prosodic symbol to be corrected and a deletion instruction, it corrects the intermediate language data by deleting the second prosodic symbol to be corrected. Also, when the rewriting unit 33 is input with the second prosodic symbol to be corrected and the type of the corrected prosodic symbol, it corrects the intermediate language data by replacing the second prosodic symbol to be corrected with the prosodic symbol of the corrected type. When the rewriting unit 33 is input with the second prosodic symbol to be added and the addition position, it corrects the intermediate language data by inserting the second prosodic symbol to be added at the addition position. The editing unit 3 performs the processing from step S240.
[0051] When the rewriting unit 33 determines that no correction of the reading kana or the second prosodic symbol has been input (step S260: NO), the editing unit 3 performs processing according to the content input by the input unit 5 (step S280). For example, when the display control unit 32 is displaying intermediate language data corresponding to a part of the original text data specified by the user and receives a specification of another part of the original text data, it may perform the same processing as in step S210. The editing unit 3 performs the processing from step S250. Then, when the editing unit 3 determines that an end has been input in step S250 (step S250: YES), it ends the processing of FIG. 9. The voice synthesis instruction may be used as an end-of-correction instruction.
[0052] FIG. 10 is a diagram showing an example of a display of an intermediate language data correction screen of the speech synthesis device 1. The display control unit 32 displays, on the intermediate language data correction screen, sentences including a mixture of kana and kanji included in the original text data in different lines one by one. The display control unit 32 displays a detailed display button C1 at the head of each sentence. When the user sets it to the display state by clicking the detailed display button C1 with a mouse or the like, the display control unit 32 reads out the intermediate language data corresponding to the kana-kanji mixed sentence displayed in the same line as the detailed display button C1 and displays it below the kana-kanji mixed sentence. The display control unit 32 displays the reading kana included in the read intermediate language data and the delimiter object representing the second prosodic symbol in the order of appearance in the intermediate language data. Further, the display control unit 32 displays an accent nucleus display object C2 representing the accent nucleus above the reading kana having the accent nucleus, and displays an accent display object C3 representing the accent height on the character string of the accent phrase. A high accent is represented by a line at a position higher than the center height of the display of the character string, a low accent is represented by a line at a position lower than the center of the display of the character string, and the accent display object C3 is a line connecting those lines in units of accent phrases. Further, the display control unit 32 further shows an object C4 indicating the accent nucleus before correction.
[0053] When the user changes the accent of an accent phrase, the user places the cursor on the display area of the reading kana of the correct accent nucleus in the accent phrase or the area above it and left-clicks the mouse. Also, when the user gives other instructions, the user right-clicks the mouse to display a menu. The user selects character correction, pause correction, end-of-sentence correction, etc. from the displayed menu. For example, the delimiter object C5 represents a normal end-of-sentence prosody symbol. The user selects end-of-sentence correction with the mouse and selects the changed end-of-sentence intonation. Thereby, the rewriting unit 33 rewrites the normal end-of-sentence prosody symbol described in the currently displayed intermediate language data with a prosody symbol representing the changed end-of-sentence intonation. The display control unit 32 changes the delimiter object C5 to a delimiter object representing the rewritten prosody symbol. Also, the user can instruct voice synthesis using the corrected intermediate language data by clicking the play button C6 with the mouse.
[0054] When instructing the high and low accents of each accent phrase by left-clicking above the reading kana of the accent nucleus, it can also be used to specify a character string that corrects the left-click operation on the display of the reading kana.
[0055] FIG. 11 is a diagram showing another display example of the intermediate language data correction screen of the speech synthesis device 1. Similar to FIG. 10, the display control unit 32 displays the intermediate language data corresponding to the kana-kanji mixed sentence in the display state by the user using the detailed display button C1 below the kana-kanji mixed sentence. Further, in the intermediate language data correction screen in FIG. 11, the display control unit 32 displays an edit button D1 for instructing character editing in front of the character at the head of each accent phrase. When the user selects the edit button D1 with the input unit 5, the display control unit 32 enables the reading kana of the accent phrase following the edit button D1 to be edited. The rewriting unit 33 replaces the reading kana of the corresponding accent phrase in the intermediate language data with the input character. Further, the display control unit 32 superimposes and displays an accent nucleus display object D2 representing the accent nucleus on the display of the reading kana having an accent nucleus, and displays an accent display object D3 representing the height of the accent of the reading kana for each reading kana. The accent display object D3 with a high accent is a line at a position higher than the center height of the string display, and the accent display object D3 representing a low accent is a line at a position lower than the center of the string display. In FIG. 11, the accent display object D3 is displayed above and below the display of the reading kana.
[0056] Next, an example of the speech synthesis unit 6 will be described. FIG. 12 is a diagram showing an example of a speech synthesis algorithm using the acoustic feature amount generation model 80 and the speech waveform generation model 90. The acoustic feature amount generation model 80 is an example of the acoustic feature amount estimation unit 61. The acoustic feature amount generation model 80 is a DNN to which the technique shown in Reference 6 "Shen et al., [online], February 2018, "Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions", arXiv:1712.05884v2, Internet <URL:https: / / arxiv.org / pdf / 1712.05884.pdf>" is applied. The speech waveform generation model 90 is an example of the vocoder unit 62. The speech waveform generation model 90 is a DNN that inputs data of acoustic feature amounts and outputs speech waveforms.
[0057] The acoustic feature generation model 80 includes an encoder 82 and a decoder 85. The encoder 82 uses a CNN (Convolutional Neural Network) and an RNN (Recurrent Neural Network) to generate feature amounts of character strings for the utterance content in the sentence indicated by the input intermediate language data, taking into account the context before and after the utterance content in the sentence indicated by the intermediate language data. The decoder 85 uses an RNN to generate, one frame at a time, the acoustic feature amounts of the predicted voice corresponding to the utterance content indicated by the input intermediate language data, based on the feature amounts generated by the encoder 82 and the acoustic feature amounts generated in the past.
[0058] The encoder 82 is composed of a character string conversion process 811, a convolutional network 812, and a bidirectional LSTM network 813. In the character string conversion process 811, each of the reading kana and prosody symbols used in the intermediate language data is converted into a numerical value, and the intermediate language is converted into a vector representation. The convolutional network 812 is a neural network in which a plurality of layers (for example, three layers) of convolutional layers are connected. In each convolutional layer, convolutional processing is performed on the vector representation of the intermediate language using a plurality of filters having a size corresponding to a predetermined number of characters, and further, batch normalization and ReLU (Rectified Linear Units) activation are performed. As a result, the context of the utterance content is modeled. For example, the filter size of the three-layer convolutional layer is [5, 0, 0], and the number of filters is 512. In order to generate the feature amounts of the character string to be input to the decoder 85, the output of the convolutional network 812 is input to the bidirectional LSTM network 813. The bidirectional LSTM network 813 is a single bidirectional LSTM of 512 units (256 units in each direction). The bidirectional LSTM network 813 can generate the feature amounts of the character string considering the context before and after in the sentence described in the input text data. LSTM is one type of RNN (Recurrent Neural Network).
[0059] The decoder 85 is an autoregressive RNN. The decoder 85 is composed of an attention network 851, a preprocessing network 852, an LSTM network 853, a first linear transformation process 854, a postprocessing network 855, an addition process 856, and a second linear transformation process 857.
[0060] The attention network 851 is a network that adds an attention function to the autoregressive RNN, and outputs a fixed-length context vector obtained by summarizing the entire output from the encoder 82 for each frame. The attention network 851 inputs the output from the bidirectional LSTM network 813 (encoder output). For each frame, the weights for extracting data from the encoder output to generate a summary vary according to the data position in the encoder output. The attention network 851 uses the data extracted from the encoder output and the data with features added using the context vector generated at the timing of the previous decoding to generate the context vector (attention network output) for the current frame.
[0061] The preprocessing network 852 inputs the data output by the first linear transformation process 854 in the previous time step. The preprocessing network 852 is a neural network including a plurality (e.g., two) of fully connected layers each consisting of 256 hidden ReLU units. The layer consisting of ReLU units outputs zero if the value of each unit is less than zero, and outputs the original value if it is greater than zero. The LSTM network 853 is a neural network in which a plurality (e.g., two layers) of unidirectional LSTMs having 1024 units are combined, and inputs the data obtained by combining the output from the preprocessing network 852 and the output from the attention network 851. Since the acoustic feature amount of a frame is affected by the acoustic feature amount of the previous frame, by combining the output from the preprocessing network 852 with the feature amount of the current frame output from the attention network 851, features based on the acoustic feature amount of the previous frame are added.
[0062] The first linear transformation process 854 linearly transforms the data output from the LSTM network 853 and generates a context vector which is the data of the mel spectrogram for one frame. The first linear transformation process 854 outputs the generated context vector to the preprocessing network 852, the postprocessing network 855, and the addition process 856.
[0063] The postprocessing network 855 is a neural network in which a plurality of layers (e.g., five layers) of convolutional networks are combined. For example, the five-layer convolutional network has a filter size of [5, 0, 0] and the number of filters is 1024. In each convolutional network, convolutional processing, batch normalization, and tanh activation are performed except for the last layer. The output from the postprocessing network 855 is used to improve the overall quality after wavelength conversion. In the addition process 856, the context vector generated by the first linear transformation process 854 and the output from the postprocessing network 855 are added.
[0064] In parallel with the above spectrogram frame prediction, in the second linear transformation process 857, after projecting the concatenation of the output of the LSTM network 853 and the attention context into a scalar, sigmoid activation is performed to output a stop token used to determine whether the output sequence is complete.
[0065] The acoustic feature generation model 80 inputs intermediate language data, generates a mel spectrogram, which is an acoustic feature for each frame, and outputs it to the speech waveform generation model 90. The speech waveform generation model 90 inputs the mel spectrogram for each frame into the speech waveform generation model, inversely transforms it into a time-domain waveform to generate and output speech waveform data. For the speech waveform generation model 90, for example, WaveNet using a multi-layer convolutional network is used, but other vocoders may also be used. The estimated speech waveform is output by the speech data or by a speech output unit (not shown) such as a speaker.
[0066] In addition to Tacotron 2 described in Reference 6, the acoustic feature quantity generation model 80 can use a sequence-to-sequence + attention method such as Deep Voice 3 or Transformer-based TTS. Deep Voice 3 is described, for example, in Reference 7 "Wei Ping et al., [online], February 2018, 'Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning', arXiv:1710.07654v3, Internet <URL:https: / / arxiv.org / pdf / 1710.07654.pdf>". Transformer-based TTS is described, for example, in Reference 8 "Naihan Li et al., [online], January 2019, 'Neural Speech Synthesis with Transformer Network', arXiv:1809.08895v3, Internet <URL:https: / / arxiv.org / pdf / 1809.08895.pdf>".
[0067] Note that the voice synthesis device 1 updates the acoustic feature quantity generation model 80 by learning using the intermediate language data and the correct acoustic feature quantities as learning data. That is, the acoustic feature quantity generation model 80 inputs the intermediate language data of the learning data and estimates the mel spectrogram. The voice synthesis unit 6 updates the acoustic feature quantity generation model 80 so that the difference between the mel spectrogram of the correct acoustic feature quantity and the estimated mel spectrogram becomes small.
[0068] According to the speech synthesis device 1 of the embodiment described above, based on the intermediate language data before being input to the speech synthesis unit 6, an object representing an accent nucleus, a high-low accent, an accent separation position, a phrase separation position, a punctuation mark, end-of-sentence information (usually, noun-ending, interrogative form, etc.) is displayed over the display of the characters representing the reading, or at a peripheral position corresponding to the display position of the characters representing the reading. A user interface that allows these to be modified can be provided. Therefore, the speech synthesis device 1 does not need to display high-low accent information by an operation different from the display of the characters representing the reading. Also, when the user inputs the designation of an accent nucleus, the speech synthesis device 1 changes the description of the intermediate language data so that the high-low accent is determined by the designated accent nucleus. Thus, it is not necessary for the user to specify an accent for each character representing the reading.
[0069] An accent nucleus is a part where the pitch is consciously raised or lowered during pronunciation. Therefore, an accent nucleus is as intuitively explicit as the stress accent in English. By simply clicking on the part that the user perceives as this accent, the speech synthesis device 1 determines the high-low accent of the accent phrase. Thus, the speech synthesis device 1 can assist the user in easily performing accent correction work. In the above embodiment, the speech synthesis device 1 displays the accent display object superimposed on the reading kana, but the accent display object may be displayed on a line different from the reading kana. In this case, the speech synthesis device 1 displays an accent display object representing the high-low accent of the reading in association with the reading represented by the reading kana.
[0070] Incidentally, the above-described speech synthesis device 1 has a computer system inside. And the process of the operation of the speech synthesis device 1 is stored in a computer-readable recording medium in the form of a program, and the above processing is performed by the computer system reading and executing this program. The computer system mentioned here includes a CPU (central processing unit), various memories, an OS (Operation System), and hardware such as peripheral devices. Also, all or part of the functions of the speech synthesis device 1 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0071] Also, the "computer system" shall include a web page providing environment (or display environment) if the WWW system is used. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built into the computer system. Furthermore, the "computer-readable recording medium" also includes something that dynamically holds a program for a short time, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and something that holds a program for a certain time, like a volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the functions described above, and may further be something that can be realized in combination with a program already recorded in the computer system for the functions described above.
[0072] The speech synthesis device 1 can be realized by, for example, one or more computer devices. When the speech synthesis device 1 is realized by a plurality of computer devices, it can be arbitrary which functional unit is realized by which computer device. For example, the editing unit 3 and the speech synthesis unit 6 may be realized by different computer devices. Further, the editing unit 3, or the language analysis unit 2 and the editing unit 3 may be realized by an editing device external to the speech synthesis device 1.
[0073] According to the embodiment described above, the editing device includes a display control unit and a rewriting unit. The display control unit displays, on the display unit in the order of appearance in the text data, the characters representing the reading of the utterance content, the first prosodic symbols representing the accents, and the prosodic display objects representing the second prosodic symbols indicating the breaks in the reading. Further, the display control unit displays, superimposed on or associated with the characters displayed on the display unit, accent display objects representing the accents of the reading represented by the characters. The accents of the reading represented by the characters are indicated by the first prosodic symbols. The display control unit may display, on the display unit, in a manner indicating that it is an accent nucleus, the character indicated as being an accent nucleus based on the first prosodic symbol. When information for selecting any of the characters displayed on the display unit as an accent nucleus is input, the rewriting unit uses, as the processing target data, the partial text data composed of the characters separated by the second prosodic symbols and the first prosodic symbols and including the selected character, which is the character selected as the accent nucleus, and rewrites the first prosodic symbols included in the processing target data so as to represent accents corresponding to the position of the selected character in the processing target data. The partial text data is, for example, the accent phrase intermediate language data of the embodiment. The display control unit updates the display on the display unit based on the text data rewritten by the rewriting unit.
[0074] An accent display object representing a low accent is a line displayed below a predetermined height such as the center height of the character display, and an accent display object representing a high accent is a line displayed above the predetermined height. The display control unit may connect and display, for each partial text data, the lines that are accent display objects corresponding to the characters in the order of appearance of the characters. Further, the rewriting unit may receive an input of information indicating a change in reading or reading separation, and rewrite the characters or second prosodic symbols included in the text data based on the input information.
[0075] The first prosodic symbol represents an accent rise or an accent fall, and the second prosodic symbol represents an accent separation, the end of a sentence, or a pause.
[0076] The speech synthesis device may have the functions of the above-described editing device. The speech synthesis device includes an acoustic feature amount estimation unit that estimates an acoustic feature amount based on the text data rewritten by the rewriting unit, and a vocoder unit that estimates a speech waveform using the acoustic feature amount estimated by the acoustic feature amount estimation unit.
[0077] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.
Explanation of Reference Numerals
[0078] 1... Speech synthesis device 2... Language analysis unit 3... Editing unit 31... Storage unit 32... Display control unit 33... Rewriting unit 4... Display unit 5... Input unit 6... Speech synthesis unit 61... Acoustic feature amount estimation unit 62... Vocoder unit
Claims
1. Based on intermediate language text data in which characters representing the reading of speech content, first prosodic symbols representing accents, and second prosodic symbols representing reading breaks are described, the characters representing the reading and the prosody display object representing the second prosodic symbols are displayed on the display unit in the order of appearance in the intermediate language text data, and an accent display object representing the accent of the reading represented by the characters, which is the accent indicated by the first prosodic symbol, is superimposed on each of the characters displayed on the display unit. A display control unit that performs a display process of displaying an accent nucleus image, which is an image representing that it is an accent nucleus, above each character shown to be an accent nucleus based on the first prosodic symbol among the characters displayed on the display unit; When a click operation for selecting any of the characters displayed on the display unit as an accent nucleus is performed in an area above the display area of the selected character, which is the selected character, among the partial text data consisting of the characters and the first prosodic symbols delimited by the second prosodic symbol in the intermediate language text data, the partial text data including the selected character is set as the processing target data, and a rewriting unit that rewrites the first prosodic symbol included in the processing target data to represent an accent corresponding to the position of the selected character in the processing target data; comprising The display control unit performs the display process based on the rewritten processing target data. The accent display object representing a low accent is a line displayed below a predetermined height in the display of the character, and the accent display object representing a high accent is a line displayed above the predetermined height. An editing apparatus characterized by the above.
2. The display control unit arranges text data in a mixture of kana and kanji representing the speech content including a plurality of sentences in different lines one by one and displays them on the display unit. When a selection of any of the sentences is input, between the display of the original text data of the selected sentence and the display of the original text data of the next sentence of the selected sentence, the result of the display process based on the intermediate language text data corresponding to the selected sentence is displayed. The editing apparatus according to claim 1, characterized by the above.
3. The display control unit connects and displays the lines, which are accent display objects corresponding to the characters, for each of the partial text data in the order of appearance of the characters. The editing device according to claim 1, characterized in that.
4. The rewriting unit receives an input of information indicating a change in reading or a reading delimiter, and rewrites the character or the second prosodic symbol included in the intermediate language text data based on the input information. The editing device according to any one of claims 1 to 3, characterized in that.
5. The first prosodic symbol represents an accent rise or an accent fall. The second prosodic symbol represents an accent delimiter, the end of a sentence, or a pause. The editing device according to any one of claims 1 to 4, characterized in that.
6. Based on intermediate language text data, which is text data in which characters representing the reading of utterance content, first prosodic symbols representing accents, and second prosodic symbols representing reading delimiters are described, the characters representing the reading and the prosody display objects representing the second prosodic symbols are displayed on the display unit in the order of appearance in the intermediate language text data, and an accent display object, which is the accent of the reading represented by the character and is represented by the first prosodic symbol, is displayed superimposed on each of the characters displayed on the display unit. A display control unit that performs a display process of displaying an accent nucleus image, which is an image representing that it is an accent nucleus, above each character shown to be an accent nucleus based on the first prosodic symbol among the characters displayed on the display unit. When a click operation for selecting any one of the characters displayed on the display unit as an accent nucleus is performed in an area above the display area of the selected character, which is the selected character, among the partial text data consisting of the characters and the first prosodic symbols delimited by the second prosodic symbol in the intermediate language text data, the partial text data including the selected character is set as processing target data, and the first prosodic symbol included in the processing target data is rewritten to represent an accent corresponding to the position of the selected character in the processing target data. A rewriting unit. An acoustic feature amount estimation unit that estimates an acoustic feature amount based on the text data rewritten by the rewriting unit. A vocoder unit that estimates an audio waveform using the acoustic feature amount estimated by the acoustic feature amount estimation unit; comprising; The display control unit performs the display process based on the processed data that has been rewritten; The accent display object representing a low accent is a line displayed below a predetermined height in the display of the character, and the accent display object representing a high accent is a line displayed above the predetermined height. A voice synthesis device characterized by this.
7. A program for causing a computer to function as the editing device according to any one of Claims 1 to 5.
Citation Information
Patent Citations
Pronunciation information creating method and device therefor
JP1997171392A
Information processor, its control method and storage medium
JP2001290491A
Device, method, and program for voice synthesis editing
JP2002351486A
Information processing method
JP2004271854A
Accent adjustable speech synthesizer
JP2008268478A