A device that provides an editor for speech synthesis.

JP7917886B2Active Publication Date: 2026-09-09REMEM CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022108856
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2026-09-09
Estimated Expiration
2042-07-06

AI Technical Summary

Benefits of technology

【0008】 本発明では、読み上げ用テキストと表音文字記号列(両者あわせて「音声合成用テキスト」ということもある)の編集作業中、常に元原稿が表示されている。しかも音声合成は、原稿を小単位に分割した区分ごとに行われるので、原稿の区分ごとの文字表示と音声表示とを連動させることが可能となる。そのため、元になる原稿のどの部分を文字と音声で表示しているかを把握しながら編集作業ができる。 また、2つの編集エリアがあり、それぞれのエリアでの修正内容を特化することで、効率的に編集ができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007917886000001
    Figure 0007917886000001
  • Figure 0007917886000002
    Figure 0007917886000002
  • Figure 0007917886000003
    Figure 0007917886000003
Patent Text Reader

Abstract

To provide a tool which generates voice synthesis data with satisfactory visibility, high work efficiency and excellent operability for generating natural synthesized voice from text data of a target document.SOLUTION: An editor providing device for voice synthesis displays three types of areas of a document display area, a first editing area and a second editing area side by side on a display screen part 16. A document to be voiced is displayed on the document display area and the document is subjected to a voice synthesis process. When outputted voice naturally reads the document, the voice is stored. When there is a problem in the voice, a reading text is edited on the first editing area, and the text is submitted to the voice synthesis process to be listened. When outputted voice is natural voice, the voice is stored. When there is the problem in voice, a phonogram symbol string is edited on the second editing area, and the voice is stored when there is no problem in the synthesized voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus that provides a tool for auditioning and correcting synthesized speech, for generating natural synthesized speech from text data of a target manuscript.

Background Art

[0002] Japanese morphological analysis used in speech synthesis has limited capabilities, and the Japanese used in a manuscript may be grammatically incorrect. For this reason, the speech generated by a speech synthesis process using morphological analysis may not result in the intended natural reading manner. For example, determining whether the characters "今日は" should be read as "kyouha" or "konnichiwa" may require judgments beyond grammar, such as the position of the phrase in the sentence, the flow of the text, and the context before and after the phrase, making the determination difficult. Therefore, in speech synthesis, manual work of correcting the pronunciation of an input sentence is indispensable. Accordingly, in order to improve the quality of synthesized speech, inventions aiming to allow manual editing without completely relying on an automatic speech synthesis process have been proposed in Patent Document 1, Patent Document 2, and the like.

[0003] Patent Document 1 discloses an invention that aims to correct prosody (accent) on a phrase-by-phrase basis. The content disclosed therein is no different from mainstream speech synthesis creation tools even as of the filing date of the present application (2022). Patent Document 2 discloses an invention that allows manual editing, and improves operability in portions of reading symbol strings (corresponding to the "phonogram symbol string" of the present invention), and is particularly specialized for correction / editing of prosody.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

[0005] The invention described in Patent Document 1 does not particularly consider improving visibility or operability for the editor. This can be seen from the fact that the original text is not displayed. Since the original text to be synthesized is not displayed, even if it deviates from the original text as edited, it is not clear which parts have changed. Furthermore, the correction of prosody is also done by dragging with the mouse. This mouse dragging can be an effective correction method if the prosody is intentionally made to be in an unusual form for the purpose of characterizing it. However, if the goal is natural pronunciation that follows the normal rules of Japanese prosody, there is no need to perform complex and subtle adjustments such as adjusting the mouse and sliders. Rather, it is more efficient to make corrections using text editing operations such as adjusting word boundaries and correcting accent marks, as these can be completed using the keyboard. The invention described in Patent Document 2 aims to improve the efficiency of correction work in order to produce the most accurate reading possible in the speech synthesis process, but this work is done entirely using a sequence of reading symbols. Since this correction is done using "reading symbols" which are a mixture of half-width katakana characters and special symbols, it is very difficult to read and the work efficiency is poor. Also, similar to the invention in Patent Document 1, it focuses only on generating speech, and there is no particular data link between the original manuscript text and the generated speech.

[0006] In view of these conventional problems, the present invention aims to provide a tool for generating speech synthesis data that is highly visible, highly efficient, and easy to operate. [Means for solving the problem]

[0007] The speech synthesis editor providing device of the present invention The display screen shows the document display area, the first editing area, and the second editing area arranged side by side. A document input display means for displaying the document to be synthesized in the document display area, A document analysis means for analyzing the document in order to perform speech synthesis and for dividing the document into sections that will be units for speech synthesis processing, A speech synthesis means that performs speech synthesis for each of the aforementioned categories and outputs the speech, A synthesized speech confirmation means inputs the section being processed on the document display area to the speech synthesis means and outputs the synthesized speech; A first editing means that edits text for reading aloud corresponding to the section being processed on the first editing area, inputs the section being edited to the speech synthesis means and outputs the synthesized speech, A second editing means that edits a sequence of phonetic characters corresponding to the section being processed on the second editing area, inputs the section being edited to the speech synthesis means and outputs the synthesized speech, The system is characterized by comprising: speech synthesis data storage means for synthesizing the final output audio for each section output by the speech synthesis means and storing it as speech data of the manuscript. [Effects of the Invention]

[0008] In this invention, the original manuscript is always displayed during the editing of the text for reading aloud and the phonetic character sequence (sometimes referred to collectively as "text for speech synthesis"). Furthermore, since speech synthesis is performed for each small section into which the manuscript has been divided, it is possible to synchronize the text display and audio display for each section of the manuscript. Therefore, editing can be performed while understanding which part of the original manuscript is being displayed in text and audio. Furthermore, there are two editing areas, allowing for efficient editing by specializing in the type of modifications needed in each area. [Brief explanation of the drawing]

[0009] [Figure 1] This diagram shows the three areas displayed on the display screen of the first embodiment and illustrates the general outline of the processes performed by the operator on each area. [Figure 2] This diagram illustrates the types of data displayed in each area of ​​the first embodiment. [Figure 3]FIG. 1 is a hardware configuration diagram of an apparatus according to a first embodiment. [Figure 4] FIG. 1 is a functional block configuration diagram of an apparatus according to a first embodiment. [Figure 5] FIG. 1 is a diagram for explaining a phonetic character symbol string to be edited according to the first embodiment. [Figure 6] FIG. 1 is a diagram for explaining a GUI (Graphical User Interface; a pop-up menu in this embodiment) for each area displayed on a display screen unit of the first embodiment. [Figure 7] FIG. 1 is a processing flow diagram from manuscript input to generation of speech as a final product in the first embodiment. MODE FOR CARRYING OUT THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0011] <<First Embodiment>> A first embodiment of the present invention (hereinafter referred to as "the present embodiment") will be described in the following order. [1. Overview of the present embodiment] [1-1. Three areas displayed on the screen] [1-1a. Manuscript display area] [1-1b. First editing area] [1-1c. Second editing area] [1-2. Text displayed in each area] [2. Configuration of the present embodiment] [2-1. Hardware configuration] [2-2. Functional block configuration] [3. Operation of the present embodiment] [3-1. Description of the GUI used] [3-2. Processing flow]

[0012] [1. Overview of the present embodiment] [1-1. Three areas displayed on the screen] To enable interactive use of the tools implemented in the device of this embodiment (hereinafter referred to as "this device"), three areas are displayed side-by-side on the display screen of this device. These are the document display area AreaA, the first editing area AreaB, and the second editing area AreaC in Figure 1. They will be described in order below.

[0013] [1-1a. Manuscript display area] On the document display area AreaA, the operator displays the target document tx1 (Step S1). The manuscript is analyzed and divided into sections that will be processed (Step S2). A waveform for speech synthesis is generated for each section and output from the speaker (Step S3). From here on, processing is carried out for each section up to Step S9 in Figure 1. If there are no problems with the output audio, this audio will be the final output for the section being processed. Therefore, no further work is required in the first editing area AreaB and the second editing area AreaC for this section.

[0014] [1-1b. Editorial Area 1] On the first editing area, AreaB, the text tx2 for reading aloud, generated in the manuscript display area, AreaA, is displayed within the area (step S4). Editing work is performed (first stage of editing, step S5). Editing in this area includes correcting pronunciation, correcting intonation, and correcting pause times in pronunciation. The text being edited for speech synthesis is converted into an audio waveform and output through the speaker (step S6). In other words, it is listened to in order to determine whether the editing is appropriate. Editing (Step S5) and listening (Step S6) are performed at least once. If the result is as intended, this synthesized audio becomes the final output for the processing section. If there are any issues that cannot be corrected in the first editing area, Area B, they will be addressed in the next editing area, Area C.

[0015] [1-1c. Second Editorial Area] On the second editing area AreaC, the phonetic character symbol sequence tx3 generated by the speech synthesis process in the first editing area AreaB is displayed within the area (step S7). Editing work is performed (second stage of editing, step S8). In this area, any issues that could not be corrected in the first editing area, Area B, are corrected to the phonetic character symbol sequence. The phonetic character sequence being edited is converted into an audio waveform and output through the speaker (step S9). In other words, it is listened to in order to determine whether the editing is appropriate. Editing (step S8) and listening (step S9) are performed at least once. If the result is as intended, this synthesized audio becomes the final output for the section being processed. Once the final output is obtained for all sections, these are combined and saved as audio data for the target document (step S10).

[0016] [1-2. Text data displayed in each area] Figure 2(1) shows an example of the data displayed in each area. As shown in Figure 2(2), the manuscript display area AreaA displays the entire manuscript text tx1 in principle, but also indicates the section currently being processed so that it can be seen at a glance. For example, in Figure 2(2), the section currently being processed in the manuscript display area AreaA, "The descriptions of the exhibited works in this building utilize REMEM technology in addition to conventional panel displays," is underlined. On the other hand, the first editing area AreaB and the second editing area AreaC may display only the parts that are underlined in the document display area AreaA, as shown in Figure 2(2).

[0017] As described above, the editing process involves two stages: first, listening to the speech synthesis result of the target manuscript and, if there are any problems, editing in the first editing area, Area B; and second, listening to the result and, if there are any problems, editing in the second editing area, Area C. Looking at each section individually, in most cases, the functions of the speech synthesis process eliminate the need for correction, and it is not necessary to display the data in both editing areas. Even when editing is necessary, the correction can usually be completed by working only in the first editing area, so the work required in the most complex second editing area is very limited. The above is an overview of [1. This Embodiment]. Next, we will explain [2. The Configuration of This Embodiment].

[0018] [2. Configuration of this embodiment] In this embodiment, the operator uses an information processing device such as a personal computer or tablet terminal to perform a series of processes from inputting the original manuscript text tx1 to generating and saving the final synthesized speech. This document describes the hardware configuration and the functional block configuration of the control unit of this information processing device (hereinafter referred to as "this device").

[0019] [2-1. Hardware Configuration] Since this embodiment is implemented by installing software on a general-purpose personal computer, the device 1 has the hardware configuration shown in Figure 3. This is the same as that of a typical personal computer. In other words, it comprises a control unit 11 such as a CPU that controls the entire device, a storage unit 12 such as a ROM, RAM, HDD, or CD drive that stores various data and programs, an input operation unit 13 such as a keyboard or mouse that accepts input from the operator, a communication interface unit 14 that controls communication with external devices, an audio output unit 15 such as a speaker, a display screen unit 16 that displays characters and GUI components, and a bus 17 that connects these.

[0020] [2-2. Functional Block Configuration] When focusing on their functions, the control unit 11 is composed of the following blocks as shown in Figure 4. Specifically, it comprises a manuscript input and display unit 111, a manuscript analysis unit 112, a speech synthesis unit 113, a synthesized speech confirmation unit 114, a first editing unit 115, a second editing unit 116, and a speech data storage unit 117. The system also includes a memory control unit (not shown) that stores manuscripts, text for reading aloud, phonetic character sequences, and audio waveforms in the memory unit 12, and retrieves them from the memory unit 12. Each function of the control unit 11 is realized by reading various programs stored in the memory unit 12 into memory and having the CPU execute them. However, some of these functions may be executed by hardware. The following describes each block of the control unit 11.

[0021] (Manuscript input and display section) The document input and display unit 111 displays the document tx1 to be synthesized in the document display area AreaA. The operator retrieves the document tx1 specified by the operator via the input operation unit 13 from the storage unit 12 or obtains it from an external database via the communication network and displays it in the document display area AreaA of the display screen unit 16. Furthermore, the text to be processed by the device is not limited to existing manuscripts as described above; the operator may also input text into the manuscript display area (AreaA) via a keyboard or other means. In other words, the manuscript to be processed by this device can be anything from a large volume of data, such as an entire book, to short sentences entered by the operator each time. The basic function of the document display area, AreaA, is to display the target document and the text tx1, such as the strings typed in that area, exactly as they are. However, this does not mean that the document tx1 cannot be edited at all. For example, if "internal hard disk" is misspelled as "internal hard disk," it is natural that this should be corrected in the document display area, AreaA. Thus, in practice, the text to be displayed is checked and final edits are made in the document display area, AreaA. This is the process of preparing the document, which was previously saved separately, into a final form for text display.

[0022] (Manuscript Analysis Department) The manuscript analysis unit 112 analyzes the manuscript tx1 input by the manuscript input and display unit 111. Specifically, it performs morphological analysis on the manuscript by referring to language dictionary data consisting of word pronunciations, parts of speech, etc. Through morphological analysis, it divides the manuscript into morphemes, classifies those morphemes as independent or non-independent words, and determines their parts of speech. The language dictionary consists of a standard dictionary and dictionaries added by the user of this device 1 (for example, registering REMEM as "rimemu"). If there are multiple dictionaries, the standard dictionary is used first, and then, if there are additional dictionaries, the information from those dictionaries is reflected by overwriting them.

[0023] Furthermore, the manuscript analysis unit 112 divides the manuscript into processing units called sections in order to synchronize it with the audio. This is because the manuscript tx1 includes long texts, such as those in a novel, and also long sentences that span multiple lines. Long texts are divided into appropriate paragraphs, and long sentences are divided in the middle of the sentence, and each of these divided units (called "sections") is input into the speech synthesis process. As a result, the unit in which the speech synthesis process is performed becomes relatively short, processing is completed in a short time, and the synchronization between the audio and the displayed text is also improved. This makes it possible to play back the audio in a synchronized manner with the highlighted text (for example, by flashing, changing the font or color, or underlining) when listening to the narration while displaying the manuscript on a device for listening to audiobooks. The process of dividing the document into sections can be done automatically or manually, and should be performed on the document display area, AreaA. The subsequent work will basically be carried out in smaller, divided sections.

[0024] (Speech synthesis unit) The speech synthesis unit 113 converts the input segment into a speech waveform using an existing speech synthesis algorithm and outputs it from the speaker 15.

[0025] (Synthesized voice verification section) The synthesized speech verification unit 114 receives a judgment from the operator when the speech synthesis unit 113 outputs the speech for the currently being verified category from the speaker 15. The output at this stage is the speech before editing by the operator. If the operator who listens to the speech determines that the speech synthesis was done as intended, this speech is saved in the storage unit 12 as the final output for the target category, and the speech synthesis process for the next category is moved on.

[0026] (First Editorial Department) The first editing unit 115 uses the display screen unit 16 and the input operation unit 13 to display the text tx2 for reading aloud of the section being processed in the first editing area AreaB and makes corrections to it in the same way as regular text data. The text tx2 for reading aloud before correction by the first editing unit 115 is a copy of the original text tx1 and is identical to the original text tx1. In other words, the text is played back using speech synthesis on the original display area AreaA, the section to be corrected is displayed as text tx2 in the first editing area AreaB, and the editing work begins. For example, if the script says "the northern direction" and the speech synthesis process reads it as "kita no ho," if you want it to read as "kita no kata," you can do so without changing the script itself by simply changing the corresponding part of the text to be read aloud tx2 to "kita no kata." As this example shows, modifications to the text to be read aloud tx2 are solely for the purpose of speech. This is a significant difference from modifications to the script text tx1 in the script display area AreaA, which are solely for display purposes.

[0027] It might seem that the target for editing in the first editing area, AreaB, should be the phonetic character sequence tx3. However, making corrections to the phonetic character sequence tx3, such as correcting pronunciation or adding pauses, would be burdensome for the worker, so we decided to edit the text for reading aloud, tx2, which is as close as possible to the original text tx1. Only when editing the text for reading aloud tx2 is inappropriate, such as when changing accent placement, should the phonetic character sequence tx3 be modified in the second editing area, AreaC.

[0028] The first editing unit 115 then inputs the section being edited into the speech synthesis unit 113, and once the audio is output from the speaker 15, it accepts the operator's judgment. The audio output here is the audio after corrections have been made to the correct pronunciation and adjustments to the pause time in pronunciation. If the operator who listens to this audio determines that the speech synthesis has been performed as intended, this audio is saved in the storage unit 12 as the final output of the section being edited, and the speech synthesis process for the next section begins. Editing and listening to the synthesized speech output in the first editing section 115 are performed continuously, and the process of editing, listening, editing, listening, etc. is repeated until the synthesized speech as intended is generated. The phonetic character symbol sequence tx3 in the relevant section that could not be corrected by the first editing section 115 becomes the target of editing in the second editing area AreaC.

[0029] (Second Editorial Department) The second editorial department 116 displays the phonetic character sequence tx3, which corresponds to the latest edit by the first editorial department 115, in the second editing area AreaC, and then modifies it in the same way as regular text data. The most common corrections made by the 2nd Editorial Department (116) are for inappropriate meter. Furthermore, they also make intentional modifications to intonation and pitch.

[0030] Here, we will explain the phonetic character sequence tx3 with reference to Figure 5. In general, the JEITA standard TT-6004 phonetic synthesis symbols are often used as phonetic character symbols. In the TT-6004 phonetic synthesis symbols, half-width katakana characters represent pronunciation, "%" represents vowel devoicing, "'" represents the position of the accent nucleus, ":" represents a short pause in a sentence, a space represents an accent break, and "." represents a pause at the end of a sentence. The information displayed in the second editing area, AreaC (phonetic character symbol sequence tx3) is often written in half-width katakana (v1 in Figure 5), as described above, making it difficult for operators to recognize. Therefore, in this embodiment, the phonetic character symbol sequence tx3 is changed to a full-width display (v2 in Figure 5), and further replaced with full-width hiragana (v3 in Figure 5). In addition, the control symbols are also changed to be clearer and easier to understand (v4 in Figure 5) to improve visibility.

[0031] Furthermore, the second editing unit 116 inputs the section being edited into the speech synthesis unit 113, and when the audio is output from the speaker 15, it accepts a judgment from the worker who listens to it. The worker listens to the audio each time they edit the phonetic character symbol sequence tx3, and if they obtain a natural-sounding audio to the extent they intended, they save this audio as the final output of that section into the recording unit 12 and proceed to the speech synthesis processing of the next section.

[0032] (Voice Data Storage Unit) The audio data storage unit 117 stores the final output for each section in the storage unit 12. The final output for all sections is the audio data corresponding to the entire target document. The above describes [2. Configuration of this embodiment]. Next, we will explain [3. Operation of this embodiment].

[0033] [3. Operation of this embodiment] When using this device 1, the worker performs tasks via a GUI provided in each area. Therefore, let's start by explaining the GUI.

[0034] [3-1. Description of the GUI used] The GUI components that serve as the interface between this device 1 and the operator include various types such as buttons and checkboxes, but here we will use the pop-up menu that appears when you right-click with the mouse within an area. Figure 6 shows examples of pop-up menus for each area.

[0035] As shown in Figure 6(1), the pop-up menu menuA provided on the manuscript display area AreaA includes the following items: "Get Manuscript," "Listen," "Save," and "Proceed to First Edit." The pop-up menu menuB, provided on the first editing area AreaB, includes the following items, as shown in Figure 6(2): "Preview," "Save," and "Proceed to Second Edit." The pop-up menu menuC, located on the second editing area AreaC, includes the items "Preview" and "Save," as shown in Figure 6(3). When performing any task in each area, select the appropriate menu item. In addition to the above, menu items such as "Temporarily Save Edits" and "Clear Edits" may also be included. The goal is simply to create an environment where users can work smoothly.

[0036] [3-2. Processing Flow] Figure 7 is a flowchart of the speech synthesis process using this device 1. When the menu item "Get Manuscript" in Figure 6(1) is clicked, the manuscript input display unit 111 displays the manuscript tx1 to be used for speech synthesis in the manuscript display area AreaA (Step F1). Next, the operator checks the document in the document display area AreaA, confirms the sentence breaks to be displayed in sync with the audio, and divides the document into sections (Step F2). From this point onward, the verification and correction work is basically done section by section, and continues until there are no unprocessed sections (Step F3, Yes). Next, the operator selects a category to work on and clicks the menu item "Preview". The speech synthesis unit 113 then synthesizes the text of the specified category into speech and outputs a waveform (step F4). If the previewed speech sounds natural to the operator as intended (Yes in step F5), the operator clicks the menu item "Save" to save the synthesized speech to the memory unit 12 (step F6), and returns to step F3 to proceed to checking the next category. If the operator is not satisfied with this speech (No in step F5), the operator clicks the menu item "Proceed to First Edit" to display the speech synthesis text tx2 in the first editing area AreaB (step F7).

[0037] The user begins modifying the text to be read aloud, tx2, displayed in the first editing area, AreaB (Step F8). When the user clicks the menu item "Preview," the speech synthesis unit 113 outputs the voice generated from the text to be read aloud, tx2, being edited, through the speaker 15 (Step F9). If the previewed voice is as intended by the user (Yes in Step F10), the user clicks the menu item "Save" to save the synthesized voice to the memory unit 12 (Step F11), and returns to Step F3 to proceed to the next section. If the user is dissatisfied with this voice (No in Step F10) and the modifications in the first editing area, AreaB, are insufficient (Yes in Step F12), they return to Step F8. In other words, the quality of the synthesized voice is improved as much as possible by repeating the editing and previewing process. When it is determined that no further editing is possible in the first editing area, AreaB (No in step F12), click the menu item "Proceed to second editing" to display the phonetic character symbol sequence tx3 in the second editing area, AreaC (step F13).

[0038] Modify the phonetic character sequence tx3 displayed in the second editing area, AreaC (step F14). For example, in the manuscript "Shibuya no club", modify the phonetic character sequence "shibuya no kurabu" (with accent on "ku") to "shibuya no kurabu" (with accent on "ra") in step F14. Next, when the operator clicks the menu item "Preview," the speech synthesis unit 113 converts the phonetic character sequence tx3 being edited into a speech waveform and outputs synthesized speech from the converted waveform (step F15). If the previewed speech is as intended by the operator (Yes in step F16), they click the menu item "Save" to save the synthesized speech to the memory unit 12 (step F17), and return to step F3 to proceed to the next section. If they are not satisfied with this speech (No in step F16), they return to step F14 and repeat the correction process. Once processing for all sections of the original text tx1 is complete (in step F3, none), the final output of each section (saved in either step F6, F11, or F17) is combined and saved as audio data for the target document (step F18), and processing for that document is completed. The above describes [3. Operation of this embodiment]. The following summarizes the distinctive advantages, concluding the description of this embodiment.

[0039] The advantages of this embodiment, which provides three areas on the screen, are as follows: By adding a new document display area, which was not present in Patent Documents 1 and 2, the first and second editing areas became dedicated work areas, independent of display. In the first editing area, users can freely edit the text for the sole purpose of correcting the audio, without being constrained by the original document notation, while always referring to the "original text." Only items that cannot be corrected in the first editing area can be further edited in the second editing area. In this way, users can freely perform correction work in two dedicated editing areas, improving work efficiency.

[0040] Many conventional speech synthesis tools focus solely on speech generation, treating the script and the generated speech as independent entities with no particular data linkage. In this embodiment, the script is divided into sections, speech synthesis is performed for each section, and the text and its corresponding speech are saved for each section. This makes it possible to link the display of the script's text with the speech reading. As a result, it is possible to display the text of the reading portion on listening devices such as smartphones, or to highlight the reading portion along with the surrounding text, achieving a "text + speech" expression. Incidentally, conventional methods have involved adding synchronization information to synchronize text and speech, but in this embodiment, such extraneous information is unnecessary. Furthermore, even if character display is linked with audio, the contents do not need to be identical. For example, while displaying a manuscript in old kana orthography as "梅の花にほひ", it is acceptable to read the audio in modern kana usage as "梅の花におい". In addition, it is not limited to kana orthography, and it is also possible to provide an expression method different from conventional ones by intentionally reproducing audio with content different from the original text. For example, if the original manuscript describes "小鬼", this corresponds to the case of pronouncing it as "ゴブリン".

[0041] When the original text is a long passage, if editing is performed while displaying text corresponding to the entire long passage in the first and second editing areas, work efficiency will be affected. Therefore, as shown in FIG. 2(2), it is preferable to display only the sentence or paragraph to be edited in the editing area, while displaying the entire content in the original text display area, and make the currently edited portion identifiable at a glance by means of highlighting or the like. This allows the editor to concentrate on the part being edited, and at the same time, immediately grasp which part of the entire manuscript is being edited. It is preferable to provide a vertical scroll bar so that any area can display a large amount of text, and can accommodate the progression of text.

[0042] <<Second Embodiment>> In the first embodiment, a personal computer at the worker's hand performs a series of processes from manuscript reading to speech synthesis. However, the same processing as in the first embodiment can also be executed by a client computer (hereinafter referred to as "client") at the worker's hand and a server computer (hereinafter referred to as "server") connected to the client via a communication network. Basically, the client only includes a device for displaying and inputting characters in three areas, a user interface such as function buttons and pop-up menus, and a device for reproducing audio, and all functional processes related to manuscript analysis and speech synthesis are provided on the server side.

[0043] Although several embodiments of the present invention have been described, these are merely examples and do not limit the scope of the present invention. The present invention can be implemented in various forms other than the above embodiments without departing from the gist of the present invention described in the claims. [Industrial applicability]

[0044] This invention, which can efficiently generate accurate and natural-sounding speech from diverse manuscripts containing a large amount of text, such as books, has many potential applications. Furthermore, by separating the characters for display and the characters for speech, it opens up new possibilities for expression, such as synchronizing and playing back speech that does not necessarily match the display. [Explanation of Symbols]

[0045] 1: Editor provider 11: Control Unit 111: Manuscript input and display unit, 112: Manuscript analysis unit, 113: Speech synthesis unit, 114: Synthesized Voice Verification Department, 115: First Editorial Department, 116: Second Editorial Department, 117: Audio Data Storage Department 12: Storage part 13: Input operation section 15: Audio output section 16: Display screen section Area A: Manuscript display area, Area B: First editing area, Area C: Second editing area tx1: Manuscript text, tx2: Text for reading aloud, tx3: Phonetic character sequence

Claims

1. The display screen shows the document display area, the first editing area, and the second editing area arranged side by side. A document input display means for displaying the document to be synthesized in the document display area, A document analysis means for analyzing the document in order to perform speech synthesis and for dividing the document into sections that will be units for speech synthesis processing, A speech synthesis means that performs speech synthesis for each of the aforementioned categories and outputs speech, A synthesized speech confirmation means inputs the section being processed on the document display area to the speech synthesis means and outputs the synthesized speech; A first editing means that edits text for reading aloud corresponding to the section being processed on the first editing area, inputs the section being edited to the speech synthesis means and outputs the synthesized speech, A second editing means that edits a sequence of phonetic characters corresponding to the section being processed on the second editing area, inputs the section being edited to the speech synthesis means, and outputs the synthesized speech. A speech synthesis editor providing device, characterized by comprising: speech synthesis data storage means for synthesizing the final output speech for each section output by the speech synthesis means and saving it as speech data of the manuscript.

2. The speech synthesis editor providing device according to claim 1, characterized in that, for each of the aforementioned sections, the characters and audio in the first editing area and / or the second editing area are output in conjunction with the characters highlighted in the document display area.

3. The speech synthesis editor providing device according to claim 1, characterized in that, if the output speech of the original manuscript by the speech synthesis means is not natural speech, the text for reading aloud generated by the manuscript analysis means is corrected by the first editing means.

4. The speech synthesis editor providing device according to claim 1, characterized in that the second editing means corrects matters that the first editing means could not correct, including correcting the prosody of the currently processed section.

5. The speech synthesis editor providing device according to claim 1, wherein the phonetic character sequence consists of characters representing pronunciation and control symbols which are symbolic characters used for controlling the speech synthesis process, the characters representing pronunciation and the control symbols are basically displayed in full-width characters, and the characters representing pronunciation and the control symbols can be modified by deletion and insertion via the input operation unit.

6. The speech synthesis editor providing device according to any one of claims 1 to 5, comprising a server computer that operates as the speech synthesis means and a client computer that outputs characters and sounds for each of the three areas of the display screen and accepts editing via an input operation unit, wherein the server computer and the client computer are connected via a network to provide the functions of a speech synthesis editor.

7. The process involves arranging and displaying the document display area, the first editing area, and the second editing area side by side on the display screen. A step of displaying the document to be synthesized in the document display area, The process involves analyzing the manuscript in order to perform speech synthesis, and dividing the manuscript into sections that will be units for speech synthesis processing. The process involves performing speech synthesis for each of the aforementioned sections and outputting the speech, The process of synthesizing the section being processed on the aforementioned document display area and outputting the audio, The process includes editing the text for reading aloud corresponding to the section being processed in the first editing area, and outputting the speech by synthesizing the section being edited, The process includes editing a sequence of phonetic characters corresponding to the section being processed in the second editing area, and outputting speech by synthesizing the section being edited, A method for providing a speech synthesis editor, characterized by comprising the steps of synthesizing the final output audio for each section output by speech synthesis and saving it as speech data of the document.

8. A speech synthesis editor providing device that displays a document display area, a first editing area, and a second editing area side-by-side on the display screen, A step of displaying the document to be synthesized in the document display area, The process involves analyzing the manuscript in order to perform speech synthesis, and dividing the manuscript into sections that will be units for speech synthesis processing. The process involves performing speech synthesis for each of the aforementioned sections and outputting the speech, The process of synthesizing the section being processed on the aforementioned document display area and outputting the audio, The process includes editing the text for reading aloud corresponding to the section being processed in the first editing area, and outputting the speech by synthesizing the section being edited, The process includes editing a sequence of phonetic characters corresponding to the section being processed in the second editing area, and outputting speech by synthesizing the section being edited, A speech synthesis editor program that enables the process of synthesizing the final output audio for each section produced by speech synthesis and saving it as speech data of the aforementioned manuscript.

Citation Information

Patent Citations

  • Voice synthesizing device and recording medium

    JP2001265374A

  • Voice synthesizer, correction method for phrase units therein, rhythm pattern editing method therein, sound setting method therein, and computer-readable recording medium with voice synthesis program recorded thereon

    JP2002023781A

  • Method and device for synthesizing voices

    JP2006349787A

  • Reading symbol string editing device and reading symbol string editing method

    JP2016105210A

  • Voice synthesis method and program

    JP2019101094A