Speech synthesis system, speech synthesis method, and program

The speech synthesis system addresses flexibility issues by enabling users to set and adjust speaker properties and edit parameters, resulting in customizable synthetic speech generation.

JP7729682B2Active Publication Date: 2025-08-26NTT TECHNOCROSS CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021119548
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-20
Publication Date
2025-08-26
Estimated Expiration
2041-07-20

AI Technical Summary

Technical Problem

Conventional speech synthesis technologies lack flexibility in setting multiple speakers for the same line, adjusting speaker properties, and comparing/selecting desired synthetic speech.

Method used

A speech synthesis system that allows users to set and adjust speaker information, properties, and edit synthetic speech parameters through various screens, enabling flexible generation of synthetic speech.

Benefits of technology

Enables flexible creation of synthetic speech by allowing multiple speakers and customizable properties, enhancing user control over the output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007729682000001
    Figure 0007729682000001
  • Figure 0007729682000002
    Figure 0007729682000002
  • Figure 0007729682000003
    Figure 0007729682000003
Patent Text Reader

Abstract

To flexibly create a synthesized speech.SOLUTION: A text-to-speech system in one embodiment includes a display control unit that displays a first screen on which one or more speaker information can be set for each of one or more texts to create a synthesized speech for the text, and a speech synthesis unit that creates the synthesized speech for the text by using one of one or more speaker information set for the text, in response to a user operation.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice synthesis system, a voice synthesis method, and a program. [Background technology]

[0002] A technology called speech synthesis has been known for some time. For example, Non-Patent Document 1 discloses a technology for creating a synthetic speech that reads a text aloud by a speaker set for the text. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] AIVOICE, Internet<URL:https: / / aivoice.jp / > Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional technologies sometimes lack the flexibility to create synthetic speech. For example, it is not possible to set multiple speakers for the same line, adjust the properties of each speaker (e.g., voice quality, speech rate, volume, intonation, pitch, etc.), compare the multiple synthetic speeches, and select the desired synthetic speech from among them.

[0005] An embodiment of the present invention has been made in view of the above points, and aims to flexibly generate synthetic speech. [Means for solving the problem]

[0006] To achieve the above object, a speech synthesis system according to one embodiment includes a display control unit that displays a first screen on which one or more pieces of speaker information can be set for each of one or more texts to create synthetic speech for the text, and a speech synthesis unit that, in response to a user operation, creates synthetic speech for the text using one piece of speaker information from the one or more pieces of speaker information set for the text. [Effects of the Invention]

[0007] It allows for flexible creation of synthesized speech. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of a speech synthesis system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of a transition relationship between screens. [Figure 3] FIG. 10 is a diagram illustrating an example of a main screen. [Figure 4] FIG. 10 is a diagram showing an example of a reading text editing screen. [Figure 5] FIG. 10 is a diagram illustrating an example of editing an accent phrase. [Figure 6] FIG. 10 is a diagram illustrating an example of a combination of accent phrases. [Figure 7] FIG. 10 is a diagram illustrating an example of editing an accent using a text editing field. [Figure 8] FIG. 10 is a diagram illustrating an example of an SSML editing screen. [Figure 9] FIG. 10 is a diagram illustrating an example of SSML editing. [Figure 10] FIG. 10 is a diagram illustrating an example of GUI switching. [Figure 11] FIG. 10 is a diagram showing an example of a prosody curve editing screen. [Figure 12] FIG. 10 is a diagram illustrating an example of a phoneme region. [Figure 13] FIG. 10 is a diagram illustrating an example of changing a pause section. [Figure 14]FIG. 10 is a diagram illustrating an example of changing the length of an accent phrase. [Figure 15] FIG. 10 is a diagram illustrating an example of a change in pitch. [Figure 16] FIG. 10 is a diagram illustrating an example of changing intonation. [Figure 17] FIG. 10 is a diagram illustrating an example of an actor editing screen. DETAILED DESCRIPTION OF THE INVENTION

[0009] An embodiment of the present invention will be described below, which will be described with reference to a voice synthesis system 1 that can flexibly generate synthetic voices.

[0010] <Overall configuration of speech synthesis system 1> The overall configuration of a speech synthesis system 1 according to this embodiment is shown in Fig. 1. As shown in Fig. 1, the speech synthesis system 1 according to this embodiment includes a speech synthesis server 10 and a client terminal 20. The speech synthesis server 10 and the client terminal 20 are connected to each other so as to be able to communicate with each other via a communication network 30 including, for example, the Internet.

[0011] The speech synthesis server 10 is a server that provides the client terminal 20 with various screens for creating synthetic speech from text and synthetic speech created using speech synthesis technology.

[0012] The client terminal 20 is a terminal that displays various screens provided by the speech synthesis server 10 and transmits user operations on these screens to the speech synthesis server 10. The client terminal 20 may be, for example, a PC (personal computer), a smartphone, a tablet terminal, or the like.

[0013] Here, the speech synthesis server 10 has a UI providing unit 101, a setting editing unit 102, and a speech synthesis engine unit 103. Each of these units is realized by, for example, processing executed by one or more programs (e.g., speech synthesis programs) installed in the speech synthesis server 10, which is executed by a computing device such as a CPU (Central Processing Unit).

[0014] The speech synthesis server 10 also includes a storage unit 104. The storage unit 104 is implemented by an auxiliary storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The storage unit 104 may also be implemented by an external storage device (such as a database server) that can communicate with the speech synthesis server 10.

[0015] The UI providing unit 101 provides various screens (such as a main screen, an actor editing screen, a reading text editing screen, an SSML editing screen, and a prosody curve editing screen, which will be described later) to the client terminal 20.

[0016] The setting and editing unit 102 sets and edits various information used to create synthetic speech in response to user operations on various screens.

[0017] The speech synthesis engine unit 103 creates synthetic speech from text in response to user operations on various screens.

[0018] The storage unit 104 stores various types of information used to create synthetic speech. Examples of this information include the following: However, in addition to these, the storage unit 104 also stores various other types of information used to create synthetic speech.

[0019] Speaker model Tone model ·dictionary Actor Information A speaker model is a model that represents a speaker's synthetic speech (including a mixed model that mixes synthetic speech from multiple speakers). A tone model is a model that represents a speaker's tone. A dictionary is information for replacing a specific word, sentence, text, or speech with another word, sentence, text, or speech, respectively. That is, dictionaries include a word dictionary, a sentence dictionary, a text replacement dictionary, a speech replacement dictionary, and a prosody dictionary. A word dictionary is information for replacing a specific word with another word when outputting synthetic speech. Similarly, a sentence dictionary is information for replacing a specific sentence with another sentence when outputting synthetic speech. A text replacement dictionary is information for replacing a specific text with another text when outputting synthetic speech. A speech replacement dictionary is information for replacing a specific speech with another speech when outputting synthetic speech. A prosody dictionary is information for replacing a specific text with a different prosody. In addition to these, a dictionary may also include information indicating the pronunciation of specific words, terms, etc.

[0020] Actor information is information consisting of an actor name, a speaker model, a tone model, speaker properties (e.g., parameters such as voice quality, speech rate, volume, intonation, pitch, etc.), and zero or more dictionaries. One piece of actor information represents one speaker who reads text aloud. Since this speaker represents an abstract actor or voice actor, in this embodiment, they are also referred to as actors. Note that the speaker model, tone model, and dictionary are information that is also used in existing speech synthesis technologies, and are created using known methods.

[0021] The client terminal 20 also includes a UI control unit 201. The UI control unit 201 is realized, for example, by processing that one or more programs (e.g., a web browser, etc.) installed in the client terminal 20 cause a computing device such as a CPU to execute.

[0022] The UI control unit 201 displays various screens provided by the speech synthesis server 10 on a display, and transmits information indicating user operations on these various screens to the speech synthesis server 10.

[0023] 1, the speech synthesis server 10 and the client terminal 20 are connected to each other so as to be able to communicate with each other via a communication network 30. However, the present invention is not limited to this, and the speech synthesis server 10 and the client terminal 20 may be integrated into one unit. That is, the client terminal 20 may include a UI providing unit 101, a setting editing unit 102, a speech synthesis engine unit 103, and a storage unit 104.

[0024] <Transitions between screens> The transition relationship between the screens provided by the UI providing unit 101 of the speech synthesis server 10 (in other words, the screens displayed by the UI control unit 201 of the client terminal 20) is shown in Fig. 2. As shown in Fig. 2, a speech editing area exists within the main screen 1000, and within this speech editing area, one of a reading text editing screen 2000, an SSML editing screen 3000, or a prosodic curve editing screen 4000 is displayed. At this time, the reading text editing screen 2000, the SSML editing screen 3000, and the prosodic curve editing screen 4000 are displayed so that transitions between them are possible.

[0025] The main screen 1000 and the actor edit screen 5000 are displayed so that transition between them is possible.

[0026] Here, the main screen 1000 is the starting screen when synthesizing text to speech, and allows the user to set the text (dialogue) to be output as synthetic speech and the actor who will read the dialogue. The main screen 1000 is, for example, the screen that is first displayed when the client terminal 20 logs in to the speech synthesis server 10. The reading text editing screen 2000, SSML editing screen 3000, and prosody curve editing screen 4000 are all screens for editing information about synthetic speech (for example, accent, pause interval, prosody, etc.). On the other hand, the actor editing screen 5000 is a screen for creating and editing actor information.

[0027] Hereinafter, it is assumed that the above-mentioned various screens are provided to the client terminal 20 by the UI providing unit 101 and displayed on the display of the client terminal 20 by the UI control unit 201. It is also assumed that user operations on the above-mentioned various screens (for example, operations using a keyboard, mouse, etc.) are accepted by the UI control unit 201, and information representing the user operations is transmitted to the speech synthesis server 10.

[0028] In addition, when various information is set or edited in accordance with user operations on the various screens described above, the setting editing unit 102 stores the information in the memory unit 104 or updates the information stored in the memory unit 104.

[0029] <Main screen 1000> An example of the main screen 1000 is shown in Fig. 3. As shown in Fig. 3, the main screen 1000 includes a dialogue management area 1100, a voice editing area 1200, a property editing area 1300, and an operation history area 1400.

[0030] The dialogue management area 1100 is an area for managing the dialogue to be output as synthetic voice and the actors set for those dialogues. The dialogue management area 1100 includes a dialogue number column 1110, a dialogue ID column 1120, a dialogue column 1130, a pattern column 1140, an actor column 1150, a comment column 1160, and a delete icon 1170.

[0031] The line number field 1110 is set with the line number. The line ID field 1120 is set with an ID that identifies the line. The line field 1130 is set with the line (text) that is to be output as synthetic speech. The pattern field 1140 is set with a symbol that represents a pattern. The actor field 1150 is set with the actor name of the actor information selected by the user. The comment field 1160 is set with any comment entered by the user.

[0032] The user sets a desired line in the line column 1130, and then selects the desired actor information from one or more pieces of actor information stored in the storage unit 104. As a result, the actor name of the actor information selected by the user is set in the actor column 1150, and a symbol representing a pattern is set in the pattern column 1140. At this time, the user can set multiple pieces of actor information for one line. Note that in the pattern column 1140, alphabets such as "A," "B," and "C" are set in order from top to bottom, for example.

[0033] 3, for the line with line ID "E001", actor information with the actor name "default" is set as pattern A, and actor information with the actor name "relaxed" is set as pattern B. Note that the user can delete the actor information corresponding to the delete icon 1170 by pressing the delete icon 1170.

[0034] Here, the user may set lines in the line column 1130 by uploading an electronic file containing text, or may set lines by directly entering text into the line column 1130. In addition, the user can press the line operation button 1131 to, for example, add blank lines to the line column 1130 or duplicate text.

[0035] Furthermore, the user can, for example, add a pattern immediately below that does not have actor information set, or reset a pattern, by pressing the pattern operation button 1141. Resetting a pattern means, for example, resetting (returning to the initial value) the edited properties in the property edit area 1300 described below.

[0036] Furthermore, by pressing the play button 1142, the user can play (read aloud) the lines in the synthetic voice of the actor corresponding to this play button 1142. The playback of this synthetic voice is realized by the voice synthesis engine unit 103 (similarly, hereinafter, playback of synthetic voice is assumed to be realized by the voice synthesis engine unit 103). Furthermore, by pressing the download button 1143, the user can download to the client terminal 20 an audio file that plays back the lines in the synthetic voice of the actor corresponding to this download button 1143.

[0037] The speech editing area 1200 is an area in which a reading text editing screen 2000, an SSML editing screen 3000, and a prosodic curve editing screen 4000 are displayed in a manner that allows for mutual transition. The example shown in FIG. 3 shows a case in which the reading text editing screen 2000 is displayed. The speech editing area 1200 includes a play button 1210 and a display menu button 1220. The play button 1210 is a button for playing back the entire dialogue displayed in the speech editing area 1200 by speech synthesis. The display menu button 1220 is a button for displaying a menu containing options for selectively displaying one of the reading text editing screen 2000, the SSML editing screen 3000, and the prosodic curve editing screen 4000. By selecting a desired option from the menu, the user can display the screen corresponding to the selected option (one of the reading text editing screen 2000, the SSML editing screen 3000, and the prosodic curve editing screen 4000) in the speech editing area 1200.

[0038] The property editing area 1300 is an area for editing the properties of the actor information of a pattern selected by the user from one or more patterns set for lines. Note that the properties refer to actor (speaker) parameters, such as voice quality, speaking speed, volume, intonation, and voice pitch. In addition, if the actor's speaker model is a hybrid model (i.e., a model that combines multiple speaker models), the blend ratio may also be included in the properties.

[0039] In the example shown in FIG. 3, the properties of actor information of pattern B with dialogue ID "E001" (actor name "soothing") are displayed. For example, the user can adjust (edit) the voice quality, which is one of the properties of the actor information, by sliding the value of voice quality adjustment bar 1310 left or right. Similarly, the user can adjust (edit) the speech rate, volume, intonation, and voice pitch, which are properties of the actor information, by sliding the values ​​of speech rate adjustment bar 1320, volume adjustment bar 1330, intonation adjustment bar 1340, and pitch adjustment bar 1350 left or right, respectively.

[0040] The operation history area 1400 is an area where the operation history of the user on the main screen 1000 is displayed.

[0041] In this way, the main screen 1000 allows multiple lines to be set, and multiple actors to be set for each line. Furthermore, the properties of each of these multiple actors can be adjusted as appropriate. Therefore, the user can set various actor and property patterns as variations for a single line. This allows the user to create the desired synthetic voice by comparing the synthetic voices resulting from various variations for each line.

[0042] Although not shown in Figure 3, in addition to downloading a synthetic voice audio file for each pattern, it may also be possible to download, for example, a synthetic voice audio file for multiple lines selected by the user and their patterns. In particular, multiple lines and their patterns arranged in chronological order are also called a timeline, which means that a synthetic voice audio file for each timeline may be downloaded.

[0043] <Reading text editing screen 2000> An example of the reading text editing screen 2000 displayed in the speech editing area 1200 of the main screen 1000 is shown in FIG. 4. As shown in FIG. 4, the reading text editing screen 2000 includes accent phrase boxes 2110-2150 that represent accent phrases of the text in the dialogue field 1130 selected by the user, and pause intervals 2210-2250 that represent the intervals (in milliseconds) between accent phrases. An accent phrase is a phrase that contains one accent. Dividing dialogue (text) into one or more accent phrases can be achieved using known natural language processing techniques.

[0044] On the reading text editing screen 2000 shown in FIG. 4, various edits can be made to the accent phrases by performing operations on each accent phrase box.

[0045] As an example, editing an accent phrase represented by the accent phrase box 2110 will be described with reference to Figure 5. The accent phrase box 2110 includes a play button 2111, a split button 2112, an accent change field 2113, a devoicing switch button 2114, an inflection raise switch button 2115, and a text edit field 2116.

[0046] By pressing the play button 2111, the user can have the accent phrase corresponding to this play button 2111 played (read aloud) by synthetic voice.

[0047] The split button 2112 is a button for splitting an accent phrase, and the user can split the accent phrase into smaller accent phrases by pressing the split button 2112. For example, the accent phrase "mukashi" can be split into "mu" and "kashi," or into "muka" and "shi."

[0048] The accent change field 2113 is a field for changing the accent position, and when the user presses the accent change field 2113, an accent position change window 2510 is displayed. This accent position change window 2510 allows the user to select the accent position of an accent phrase. In the example shown in FIG. 5, the user can select a desired accent phrase from three accent phrases 2511 to 2513 with different accent positions. At this time, by pressing the play button 2515, each of the accent phrases 2511 to 2513 with different accent positions can be played back by synthetic speech.

[0049] The devoicing switch button 2114 is a button for devoicing part of an accent phrase (the characters directly below the devoicing switch button 2114). When the user presses the devoicing switch button 2114, a devoicing switch window 2520 is displayed. In this devoicing switch window 2520, the user can select whether or not to devoice the corresponding characters. In the example shown in FIG. 5, the user can select the desired accent phrase from among an accent phrase 2521 that is not to be devoiced and an accent phrase 2522 that is devoiced. At this time, by pressing the play button 2525, the accent phrases 2521 to 2522 can each be played back by synthetic speech.

[0050] The inflection-raising switch button 2115 is a button for switching whether or not to raise the endings of accent phrases. When the user presses the inflection-raising switch button 2115, an inflection-raising switch window 2530 is displayed. In this inflection-raising switch window 2530, the user can select whether or not to raise the endings of accent phrases. In the example shown in FIG. 5, the user can select a desired accent phrase from accent phrase 2531 without inflection and accent phrase 2532 with inflection. At this time, by pressing the play button 2535, accent phrases 2531 to 2532 can each be played back by synthetic speech.

[0051] The text edit field 2116 is a field for editing the accent of the accent phrase by adding or changing accent marks.

[0052] Here, accent phrases can be combined by dragging the mouse on the reading text editing screen 2000. For example, as shown in Figure 6, by dragging accent phrase box 2120 to the left and pressing it against accent phrase box 2110, a new accent phrase box 2160 is created by combining these accent phrase boxes. This combines accent phrase box 2110 representing "mukashi" with accent phrase box 2120 representing "mukashi," creating accent phrase box 2160 representing "mukashimukashi."

[0053] When editing accents in the text editing field, the user adds, changes, or deletes accent marks in the text editing field. For example, as shown in Figure 7, assume that "mukashimukashi [.00]" is displayed in the text editing field 2166 of the accent phrase box 2160. [.00] is the accent mark.

[0054] In this case, for example, if the accent mark is changed from [.00] to [.02], the accent of the accent phrase will change. In the example shown in Figure 7, when the accent mark is [.00], the accent remains flat after the first "ka" is pronounced, but when the accent mark is changed to [.02], only the first "ka" is pronounced high.

[0055] Also, for example, if an accent mark [.00] is added after the first "mukashi" in "mukashimukashi[.00]," the accent phrase is split into two. In the example shown in Figure 7, accent phrase box 2160 is split into accent phrase box 2110 and accent phrase box 2120.

[0056] Furthermore, when the accent symbol [.00] is deleted from the text editing field 2116 of the accent phrase box 2110, it is combined with the next accent phrase box 2120. In the example shown in FIG. 7, the accent phrase box 2110 and the accent phrase box 2120 are combined to create the accent phrase box 2160.

[0057] Note that the accent symbol is expressed in the form of [xnn]. Here, x represents the relationship with the next accent phrase or the pitch rise at the end of a word, and nn represents that the accent phrase is given an accent of type nn. For example, [.00] indicates that the interval from the next accent phrase is 0.5 ms and an accent of type 00 is given. Also, for example, [ / 01] indicates that there is no interval from the next accent phrase and an accent of type 01 is given. When x is "?", it represents a pitch rise at the end of a word.

[0058] In this way, on the reading text editing screen 2000, various edits can be made to the accent phrases that make up the lines selected by the user on the main screen 1000. Moreover, at this time, since the accent phrases can be reproduced by synthesized voice in units of accent phrases, it is possible to compare and consider the voices between the options when setting the accent position, voicelessness, pitch rise at the end of a word, etc.

[0059] Also, when combining adjacent accent phrases, a drag operation can be performed on the accent phrase box, and the combination of accent phrases can be intuitively performed.

[0060] <SSML Editing Screen 3000> An example of the SSML editing screen 3000 displayed in the voice editing area 1200 of the main screen 1000 is shown in FIG. 8. Note that SSML is an abbreviation of Speech Synthesis Markup Language and is also called a voice synthesis markup language. In SSML, by defining SSML tags for the text, it becomes possible to change the pause interval, voice intensity, volume, speed, etc. of the synthesized voice.

[0061] As shown in Figure 8, the SSML editing screen 3000 includes a text display field 3100 in which the text in the dialogue field 1130 selected by the user is displayed, an edit button 3200 for performing various edits on the synthesized speech corresponding to this text, and a GUI switch button 3300 for switching the display format of the text display field 3100.

[0062] When editing the synthesized speech corresponding to the text in text display field 3100, the user selects the desired range of text to edit and then presses the desired edit button from edit buttons 3200. For example, as shown in the upper diagram of FIG. 9, the user selects "Good morning" as range 3110 and presses the "Louder" button from edit buttons 3200. This makes it possible to edit the synthesized speech so that the "Good morning" part is spoken louder.

[0063] At this time, the text of the "Good morning" portion is displayed in a format corresponding to the "Larger" button (i.e., a larger font size), as shown in the SSML editing screen 3000 in the lower diagram of Figure 9. Note that even when another edit button 3200 is pressed, the text within the selected range is displayed in a format corresponding to that edit button 3200. Such a change in the display format can be realized, for example, by HTML / CSS functions. This allows the user to visually confirm how the text will be pronounced based on the display format.

[0064] Furthermore, for example, when the user presses the GUI switching button 3300 on the SSML editing screen 3000 shown in the upper diagram of Fig. 10, SSML tags are displayed in the text display field 3100. SSML tags are displayed in the text display field 3100 of the SSML editing screen 3000 shown in the lower diagram of Fig. 10. This allows the user to check the actual SSML tags. Note that the user may edit or delete SSML tags in the text display field 3100 of the SSML editing screen 3000 shown in the lower diagram of Fig. 10, or may add SSML tags to this text display field 3100.

[0065] In this way, the SSML editing screen 3000 allows users to intuitively add SSML tags to the dialogue text selected by the user on the main screen 1000. Furthermore, users with knowledge of SSML can also directly add, delete, edit, etc. SSML tags. Therefore, users without knowledge of SSML can add SSML tags simply and easily, while users with knowledge of SSML can add SSML tags in detail.

[0066] <Prosody curve editing screen 4000> Fig. 11 shows an example of the prosodic curve editing screen 4000 displayed in the speech editing area 1200 of the main screen 1000. As shown in Fig. 11, the prosodic curve of the text in the dialogue field 1130 selected by the user is displayed in an editable manner on the prosodic curve editing screen 4000. In the example shown in Fig. 11, the prosodic curve of the text "Hello, everyone. I hope you have a good day" is displayed in an editable manner.

[0067] The prosodic curve can be edited in predetermined units (sentence, pause phrase, accent phrase, mora, phoneme). The user can edit the prosodic curve by performing mouse operations such as dragging on the areas representing these units.

[0068] A sentence unit is a unit that represents one sentence, such as "Konnichiwa Kyowayoi Tenki desu ne." A pause phrase unit is a unit that represents a pause phrase separated by pauses, such as "Konnichiwa" and "Kyowayoi Tenki desu ne." An accent phrase unit is a unit that represents an accent phrase, such as "Konnichiwa," "Kyowa," "Yoi," and "Tenki desu ne." A mora unit is a unit that represents the moras "ko," "n," "ni," "chi," "wa," "kyo," "-," "wa," "yo," "i," "te," "n," "ki," "de," "su," and "ne." A phoneme unit is a unit that represents phonemes, such as "k," "o," "n," "n," "i," "c," "h," "i," "w," and "a."

[0069] The prosodic curve editing screen 4000 shown in Fig. 11 includes a sentence area 4100 for editing a prosodic curve on a sentence-by-sentence basis. Furthermore, this sentence area 4100 includes pause phrase areas 4200-4300 for editing a prosodic curve on a pause phrase-by-pause basis. The pause phrase area 4200 includes an accent phrase area 4210 for editing a prosodic curve on an accent phrase basis, and the pause phrase area 4300 includes accent phrase areas 4310-4330. The accent phrase area 4210 includes mora areas 4211-4215 for editing a prosodic curve on a mora-by-mora basis. Similarly, the accent phrase area 4310 includes mora areas 4311-4313, the accent phrase area 4320 includes mora areas 4321-4322, and the accent phrase area 4330 includes mora areas 4331-4336.

[0070] 12, for example, mora area 4311 includes phoneme areas 4511 to 4512 for editing prosodic curves on a phoneme-by-phoneme basis. Similarly, mora area 4313 includes phoneme areas 4521 to 4522. Note that phoneme area 4511 is an area representing the phoneme "k," phoneme area 4512 is an area representing the phoneme "y," and phoneme area 4513 is an area representing the phoneme "o." Similarly, phoneme area 4521 is an area representing the phoneme "w," and phoneme area 4522 is an area representing the phoneme "a."

[0071] The user can perform various edits (such as changing pause intervals, changing the length of accent phrases, changing pitch, changing intonation, etc.) by performing mouse operations (such as dragging) on ​​each of the above areas. An example of these edits is described below.

[0072] Fig. 13 shows an example of changing the pause interval between the pause phrase "Hello" and the pause phrase "You're welcome." As shown in Fig. 13, by dragging the pause phrase area 4300 to the left, the pause interval between "Hello" and "You're welcome." It is possible to lengthen the pause interval by dragging the pause phrase area 4300 to the right.

[0073] Fig. 14 shows an example of extending the accent phrase "KYOWA." As shown in Fig. 14, by expanding the accent phrase area 4310 horizontally with the mouse, the speech period of the accent phrase "KYOWA" can be lengthened. Note that by reducing the accent phrase area 4310 horizontally, the speech period of the accent phrase "KYOWA" can be shortened.

[0074] Figure 15 shows an example of lowering the pitch of each of the accent phrases "kyowa," "yoi," and "tenki desu ne." As shown in Figure 15, by dragging each of the accent phrase areas 4310 to 4330 downward with the mouse, the pitch of each of the accent phrases "kyowa," "yoi," and "tenki desu ne" can be lowered. Note that by dragging any of the accent phrase areas 4310 to 4330 upward with the mouse, the pitch of that accent phrase can be raised.

[0075] Fig. 16 shows an example of changing the intonation of the pause phrases "Konnichiwa" and "Kyouwayoitenki desune". As shown in Fig. 16, by vertically expanding the pause phrase area 4200 with a mouse and vertically shrinking the pause phrase area 4300 with a mouse, it is possible to strengthen the intonation of the pause phrase "Konnichiwa" and weaken (suppress) the intonation of the pause phrase "Kyouwayoitenki desune".

[0076] 14 to 16 are merely examples, and the prosodic curve can be edited (for example, changing the pause interval, the length of the accent phrase, the pitch, or the intonation) by operating the mouse or the like on an area representing a predetermined unit (sentence area, pause phrase area, accent phrase area, mora area, phoneme area). In particular, since the prosodic curve can be edited even on a phoneme-by-phoneme basis, it becomes possible to edit the prosodic curve of the synthesized speech in extremely fine units.

[0077] In this way, the prosodic curve editing screen 4000 allows editing of the prosodic curve when the text of the dialogue selected by the user on the main screen 1000 is converted into synthetic speech. Moreover, editing can be done intuitively by mouse operation in various units such as sentence units or accent phrase units.

[0078] <Actor Editing Screen> 17 shows an example of the actor edit screen 5000, which can be transitioned to and from the main screen 1000. For example, the actor edit screen 5000 is displayed by performing a display operation for the actor edit screen 5000 on the main screen 1000 (for example, by right-clicking to open a menu and then selecting "Edit Actor" from the menu). Similarly, the main screen 1000 is displayed by performing a display operation for the main screen 1000 on the actor edit screen 5000 (for example, by right-clicking to open a menu and then selecting "Manage Dialogue" from the menu).

[0079] As shown in FIG. 17, the actor editing screen 5000 includes an actor name setting field 5100, a speaker model selection field 5200, a tone model selection field 5300, a property setting field 5400, and a dictionary selection field 5500.

[0080] An arbitrary actor name is set in the actor name setting field 5100. A speaker model selection field 5200 displays speaker models that the user can select. A tone model selection field 5300 displays tone models that the user can select. A property setting field 5400 displays properties that the user can set. A dictionary selection field 5500 displays dictionaries that the user can select.

[0081] The user enters an arbitrary actor name in the actor name setting field 5100, selects a desired speaker model from the speaker model selection field 5200, and selects a desired tone model from the tone model selection field 5300. Furthermore, the user sets values ​​for various properties in the property setting field 5400 and selects a desired dictionary from the dictionary selection field 5500. After that, by pressing the save button 5600, actor information consisting of the actor name, the speaker model, the tone model, the properties, and the dictionary is created and saved. The user can select two or more dictionaries from the dictionary selection field 5500, or can select no dictionary at all.

[0082] The user can input any text into the text input field 5700 and then press playback test 5800 to play back the text in the synthesized voice of the actor represented by the actor information. In other words, the user can perform a playback test of the actor information.

[0083] In this way, the actor editing screen 5000 allows the user to create actor information consisting of an actor name, a speaker model, a tone model, properties, and a dictionary. This allows the user to create a desired actor (speaker). Note that while Fig. 17 illustrates the case of creating new actor information, it may also be possible to edit actor information that has already been created (i.e., change any of the actor name, speaker model, tone model, properties, or dictionary).

[0084] In particular, since a dictionary can be set for each actor, it becomes possible to output different synthesized voices for each actor for the same text, for example.

[0085] For example, suppose that actor A's actor information includes Dictionary A, which replaces the string "●branch name●" with "Tokyo branch" and the string "●start time●" with "9:00 AM." Actor B's actor information includes Dictionary B, which replaces the string "●branch name●" with "Italy branch" and the string "●start time●" with "11:00 AM."

[0086] In this case, if the text "The opening time of the ● branch name ● is ● opening time ●" is read aloud by the synthesized voices of Actor A and Actor B, the synthesized voice output from Actor A will say "The opening time of the Tokyo branch is 9:00 AM," and the synthesized voice output from Actor B will say "The opening time of the Italy branch is 11:00 AM."

[0087] In this way, by providing a dictionary for each actor, it becomes possible to output different synthesized speech for each actor even when the same text is used.

[0088] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims. [Explanation of symbols]

[0089] 10 Speech synthesis server 20 Client Terminals 30 Communication Network 101 UI provision department 102 Settings Editorial Department 103 Speech synthesis engine 104 Storage section 201 UI control section

Claims

1. a display control unit that displays, in a manner that allows transitions between, a first screen on which, for each of one or more texts, multiple pieces of speaker information can be set to create synthetic speech for the text and parameters included in each of the multiple pieces of speaker information can be changed; a second screen for editing synthetic speech for each accent phrase included in the text; a third screen for editing synthetic speech by adding, deleting, or changing SSML tags for the text; and a fourth screen for editing synthetic speech by changing a curve representing the prosody of the text; a speech synthesis unit that generates synthetic speech for the text edited on the second screen, the third screen, and the fourth screen, respectively, using two or more pieces of speaker information among the plurality of pieces of speaker information set for the text on the first screen in response to a user operation; and The speech synthesis system, wherein the parameters include at least one of voice quality, speech rate, volume, intonation, and pitch.

2. 2. The speech synthesis system of claim 1, wherein the speaker information includes a speaker model representing a virtual speaker who speaks the text using synthetic speech, a tone model representing the tone of the speaker, the parameters, and a dictionary for replacing specific words, sentences, text, and speech with other words, sentences, text, and speech, respectively.

3. The speech synthesis system of claim 2 , wherein the dictionary includes information indicating how to read specific words or terms.

4. The display control unit The speech synthesis system according to claim 1 , wherein a fifth screen for creating or editing the speaker information is displayed so as to be transitionable between the fifth screen and the first screen.

5. In the second screen, 2. The speech synthesis system according to claim 1, wherein editing of the synthesized speech can include at least changing the accent of each accent phrase and combining adjacent accent phrases.

6. In the third screen, 2. The speech synthesis system according to claim 1, wherein said synthetic speech can be edited by specifying a range of said text and by changing the stress, volume and speed of the synthetic speech of said text within said range.

7. In the fourth screen, 2. The speech synthesis system according to claim 1, wherein, in response to a user operation, at least one of the following can be performed: changing the curve on a sentence-by-sentence basis, changing the curve on a pause phrase-by-pause phrase basis, changing the curve on an accent phrase basis, changing the curve on a mora-by-more basis, or changing the curve on a phoneme-by-phoneme basis.

8. a display control procedure for displaying, in a manner that allows transitions between, a first screen on which, for each of one or more texts, multiple pieces of speaker information for creating synthetic speech for the text can be set and parameters included in each of the multiple pieces of speaker information can be changed; a second screen for editing the synthetic speech for each accent phrase included in the text; a third screen for editing the synthetic speech by adding, deleting, or changing SSML tags for the text; and a fourth screen for editing the synthetic speech by changing a curve representing the prosody of the text; a speech synthesis procedure for creating synthetic speech for the text edited on the second screen, the third screen, and the fourth screen, respectively, using two or more pieces of speaker information among the plurality of pieces of speaker information set for the text on the first screen in response to a user operation; The computer executes The speech synthesis method, wherein the parameters include at least one of voice quality, speech rate, volume, intonation, and pitch.

9. A program that causes a computer to function as the speech synthesis system according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Personal computer, and method for managing voice attribute parameter

    JP1998133852A

  • Information processor, content providing method, and terminal device

    JP2004233709A

  • Device, method, and program for speech editing

    JP2005345699A

  • SSML editor system

    JP2008059586A

  • Speech learning device, speech learning method and program

    JP2014038209A