Voice processing assistance device, voice processing assistance method, and voice processing assistance program

By combining the input receiving and recording units of the speech processing auxiliary device with speech dictionary data, the emotional shift parameters of the speech data are set in detail, which solves the problem of insufficient emotional expression in the existing technology and realizes fine control and emotional expression of speech data on the time axis.

CN120858404APending Publication Date: 2025-10-28KK TOSHIBA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480018372.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-01
Filing Date
2024-08-26
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to set parameters related to the emotional shifts over time, resulting in insufficient refinement of emotional expression in voice data.

Method used

The speech processing aid uses an input receiving unit and a recording unit to receive and record parameters of multiple emotion types and mixing ratios. Combined with speech dictionary data, it generates synthesized speech data, enabling detailed editing and reproduction of the speech data.

Benefits of technology

It enables precise control over the emotional shift of voice data over time, reducing the user's input and editing burden and improving the emotional expression of voice data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120858404A_ABST
    Figure CN120858404A_ABST
Patent Text Reader

Abstract

The invention relates to a voice processing assistance device, a voice processing assistance method, and a voice processing assistance program. A voice processing assistance device (10) is provided with an input reception unit (20B) and a recording unit (20F). The input reception unit (20B) receives an input of a parameter including at least a plurality of different types of feelings and a mixing ratio of the plurality of types of feelings during reproduction of voice data to be edited. The recording unit (20F) associates and records the input-accepted parameter with the reproduction timing at which the input of the parameter is accepted in the voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to a speech processing aid device, a speech processing aid method, and a speech processing aid program. Background Technology

[0002] Techniques are known for synthesizing speech with a different sound quality than existing speech by mixing (deforming) speech data. For example, techniques for generating synthesized speech data by combining multiple speech data according to a pre-specified deformation ratio have been disclosed.

[0003] However, while existing technologies can set parameters such as the overall deformation ratio of voice data, it is difficult to set parameters related to the shift of emotions over time in detail.

[0004] Prior art literature Patent documents Patent Document 1: Japanese Patent Application Publication No. 9-50295 Summary of the Invention

[0005] The problem that the invention aims to solve The problem to be solved by the present invention is to provide a speech processing aid device, a speech processing aid method, and a speech processing aid program that can set in detail parameters related to the progression of emotions that change over time.

[0006] Methods for solving problems The speech processing aid device of this embodiment includes an input receiving unit and a recording unit. The input receiving unit receives parameters that include at least a plurality of different types of emotion and a mixing ratio of said plurality of emotions during the reproduction of speech data of the object being edited. The recording unit records the received parameters in a corresponding manner with the reproduction timing of the speech data in which the input parameters were received. Attached Figure Description

[0007] Figure 1 This is a diagram illustrating an example of a speech processing aid device according to an embodiment.

[0008] Figure 2 This is a schematic diagram of an example of a display screen.

[0009] Figure 3 This is a schematic diagram of an example of a display screen.

[0010] Figure 4 This is a schematic diagram of an example of a display screen.

[0011] Figure 5 This is a schematic diagram of an example of a display screen.

[0012] Figure 6This is a flowchart illustrating an example of the information processing flow performed by the speech processing aid device of this embodiment.

[0013] Figure 7 It is a hardware structure diagram. Detailed Implementation

[0014] The following describes in detail, with reference to the accompanying drawings, the speech processing aid device, the speech processing aid method, and the speech processing aid program.

[0015] Figure 1 This is a diagram illustrating an example of the speech processing aid 10 of this embodiment.

[0016] The speech processing auxiliary device 10 is an information processing device that assists in the processing of speech data.

[0017] The voice processing aid 10 includes a communication unit 12, a UI (user interface) unit 14, a storage unit 16, and a processing unit 20. The communication unit 12, UI unit 14, storage unit 16, and processing unit 20 are communicatively connected via a bus 18.

[0018] The communication unit 12 communicates with other external information processing devices via a network or the like. The UI unit 14 includes a display unit 14A and an input unit 14B.

[0019] Display unit 14A displays various information. Display unit 14A is, for example, an LCD (Liquid Crystal Display), an organic EL (Electro-Luminescence) display, a projection device, etc.

[0020] The input unit 14B accepts user input. The input unit 14B can be, for example, a pointing device such as a digital pen, mouse, or trackball, or an input device such as a keyboard. Alternatively, it can be configured as a touch panel that integrates at least a portion of the display unit 14A and the input unit 14B. In this embodiment, the input unit 14B includes a first operation unit 14B1 and a second operation unit 14B2.

[0021] The first operation unit 14B1 is an input device for operation direction and operation amount. For example, the first operation unit 14B1 is an input device capable of inputting operation direction and operation amount based on tilt direction and tilt angle by tilting a joystick (rocker). Such a first operation unit 14B1 is sometimes referred to as a joystick. Alternatively, the first operation unit 14B1 may also be, for example, a touchpad. In this embodiment, the first operation unit 14B1 will be described as a joystick as an example.

[0022] The second operation unit 14B2 is an input device for operation quantities. The second operation unit 14B2 is, for example, a pedal-type or button-type input device that is pressed by the user's foot or the like. In this embodiment, an example of a pedal-type second operation unit 14B2 that inputs operation quantities by pressing by the user's foot or the like will be described.

[0023] The voice output unit 14C is a speaker that outputs voice.

[0024] Storage unit 16 stores various types of data. Storage unit 16 can be, for example, a semiconductor memory element such as RAM (Random Access Memory), flash memory, a hard disk, or an optical disk. Alternatively, storage unit 16 can be a storage device located external to the voice processing aid 10. Furthermore, storage unit 16 can also be a storage medium. Specifically, the storage medium can be a medium that downloads programs and various information via a LAN (Local Area Network), the Internet, etc., and stores or temporarily stores them. Additionally, storage unit 16 can be composed of multiple storage media.

[0025] Next, the processing unit 20 will be described. The processing unit 20 performs various information processing tasks. The processing unit 20 includes a display control unit 20A, an input receiving unit 20B, a setting unit 20C, an acquisition unit 20D, a playback unit 20E, and a recording unit 20F.

[0026] The display control unit 20A, input receiving unit 20B, setting unit 20C, acquisition unit 20D, playback unit 20E, and recording unit 20F are implemented by one or more processors, for example. For instance, each of these units can also be implemented in software by having a processor such as a CPU (Central Processing Unit) execute a program. Alternatively, each of these units can be implemented in hardware, such as a dedicated IC (Integrated Circuit). Each of these units can also be implemented using both software and hardware. When using multiple processors, each processor can implement one or more of these units.

[0027] In addition, at least one of the above-mentioned units and at least a portion of the information stored in the storage unit 16 may also be mounted on a cloud server or the like for processing in the cloud.

[0028] The display control unit 20A displays various screens on the UI unit 14. Details of the display screens will be described later.

[0029] The input receiving unit 20B accepts user input for operations on the UI unit 14.

[0030] For example, imagine a scenario where a user wants to use certain voice data to generate synthesized voice data. More specifically, imagine a user wants to process and edit the voice data to create synthesized voice data that contains a specified emotion along a timeline.

[0031] In this case, the user operates the input unit 14B to set voice dictionary data corresponding to the type of emotion used in setting the voice data. In this embodiment, the input receiving unit 20B receives user input via a display screen shown on the display unit 14A.

[0032] Figure 2 This is a schematic diagram of an example of display screen 30A. Display screen 30A is an example of display screen 30 displayed on display unit 14A.

[0033] The display screen 30A includes an emotion mapping chart M, an emotion setting bar 40A, and a voice dictionary data setting bar 40B.

[0034] The Emotion Map M is a mapping diagram representing the correlation between multiple types of emotions. For example, the Emotion Map M uses colors to represent different types of emotions and combines them three-dimensionally, thereby representing complex emotions such as the mixing of emotions. In the Emotion Map M, for example, the regions for each of the eight basic emotions—joy, trust, panic, surprise, sadness, disgust, anger, and anticipation—are arranged radially outwards from the region of no emotion. Furthermore, in the Emotion Map M, regions for opposite types of emotions are arranged 180° opposite each other, separated by the central region of no emotion. Additionally, in the Emotion Map M, the eight basic emotions are categorized into three groups: positive, negative, and neutral, and the regions for emotions belonging to each group are arranged in adjacent positions. Moreover, in the Emotion Map M, the regions for each type of emotion are such that the closer they are to the central region of no emotion, the weaker the emotion, and the farther away from the central region of no emotion, the stronger the emotion.

[0035] Furthermore, the sentiment mapping graph M can be any mapping graph representing the correlation between multiple types of sentiments, and is not limited to... Figure 2 As shown in the diagram.

[0036] The emotion setting bar 40A is an input bar for the type of emotion used in setting voice data. For example, by visually confirming the emotion map M on the display screen 30A while operating the input unit 14B, the user inputs the type of emotion to be attached to the voice data into the emotion setting bar 40A. Since the display screen 30A includes the emotion map M, the user can easily input the type of emotion into the emotion setting bar 40A by visually confirming the emotion map M included on the display screen 30A.

[0037] The voice dictionary data setting field 40B is an input field for setting voice dictionary data that corresponds to the type of emotion entered into the emotion setting field 40A.

[0038] Speech dictionary data is a sound model used to derive sound features from linguistic features. Speech dictionary data is prepared in advance. Linguistic features are features extracted from the text of a speaker's spoken words. Examples of linguistic features include preceding and following phonemes, pronunciation-related information, sentence ending position, text length, stressed sentence length, short syllable length, short syllable position, stress pattern, part of speech, and modification information. Linguistic features are sometimes referred to as linguistic information. Sound features are features of speech or sound extracted from speech data. Sound features can be, for example, those used in HMM (Hidden Markov Model) speech synthesis. Examples of sound features include Mel-Cepstral Coefficients (MCCs), Mel-LPC Coefficients, Mel-LSP Coefficients (representing phonology and tone), Fundamental Frequency (F0) (representing pitch), and BAP (Balanced Aperiodic Index) (representing periodicity and the proportion of non-periodic components). These coefficients are represented by speech waveforms expressed in terms of frequency, etc.

[0039] In this embodiment, a plurality of speech dictionary data are pre-stored in the storage unit 16. The plurality of speech dictionary data are used to output the voice feature quantities of one or more speakers who speak with different emotions.

[0040] The user selects the voice dictionary data for the type of emotion input into the emotion setting bar 40A by operating the input unit 14B while visually confirming the display screen 30A. Specifically, for example, the user selects voice dictionary data that corresponds to the corresponding type of emotion from multiple voice dictionary data stored in the storage unit 16 by operating the input unit 14B, and sets it in the voice dictionary data setting bar 40B.

[0041] exist Figure 2 As an example, the text shows a filename for a speech dictionary that includes the name of the corresponding emotion category. However, the filename for speech dictionary data can also be formatted without including the name of the corresponding emotion category.

[0042] In addition, speech dictionary data corresponding to the emotion category "emotionless" is pre-set. The speech dictionary data corresponding to "emotionless" is speech dictionary data used to output the vocal feature quantities of a speaker speaking without emotion.

[0043] Through these input operations performed by the user, the input receiving unit 20B of the processing unit 20 accepts the type of emotion used in the setting of the voice data and the setting of the voice dictionary data corresponding to the type of emotion.

[0044] return Figure 1 Let me continue explaining.

[0045] The setting unit 20C sets the voice dictionary data corresponding to the types of emotions received via the display screen 30A as the voice dictionary data to be used in the editing of voice data. For example, the setting unit 20C stores the voice dictionary data corresponding to the types of emotions received via the display screen 30A in specific storage areas of the storage unit 16.

[0046] The acquisition unit 20D acquires the voice data of the object to be edited. The user specifies the voice data of the object to be edited, which is stored in the storage unit 16 or an external information processing device, etc., by operating the UI unit 14. The acquisition unit 20D acquires the voice data specified by the user's operation instructions on the UI unit 14 as the voice data of the object to be edited.

[0047] In this embodiment, the input receiving unit 20B receives the specified input of the voice data of the editing object via the display screen 30 displayed on the display unit 14A.

[0048] Figure 3 This is a schematic diagram illustrating an example of display screen 30B. Display screen 30B is an example of display screen 30. Display screen 30B is displayed on the UI unit 14 when voice data is specified and parameters are input.

[0049] When the user sets the voice dictionary data through the operation instructions of the input unit 14B, the display control unit 20A displays the display screen 30B on the display unit 14A.

[0050] The display screen 30B includes a voice data file name input display bar 40C, a playback button 40D, a voice waveform display bar 40E, an emotion map M, a pointer 40F, a speech rate adjustment button 40G, a gain adjustment button 40H, an edit button 40J, a synthesized speech playback button 40K, and a save button 40L.

[0051] The voice data filename input display area 40C is used to input and display the filename of the voice data to be edited. The user operates the input unit 14B while visually confirming the display screen 30B, thereby inputting the filename of the voice data to be edited into the voice data filename input display area 40C. The acquisition unit 20D acquires the voice data of the input filename as the voice data of the edited object. Alternatively, the user can also specify the voice data of the edited object stored in the storage unit 16, etc., while visually confirming the display screen 30B and operating the input unit 14B. In this case, the acquisition unit 20D acquires the specified voice data as the voice data of the edited object.

[0052] Alternatively, the user can input the filename of the text data of the object to be edited by operating the input unit 14B while visually confirming the display screen 30B. Additionally, the user can specify the text data of the object to be edited, stored in the storage unit 16, etc., by operating the input unit 14B while visually confirming the display screen 30B.

[0053] In this case, the acquisition unit 20D obtains the speech data of the object to be edited by using text data represented by the input file name or the specified text data and speech dictionary data corresponding to the emotion type "no emotion" and generating speech data using known methods.

[0054] In this embodiment, the processing unit 20 of the speech processing aid 10 receives parameter input during the reproduction of the speech data of the editing object.

[0055] Parameters refer to the shifts in emotion over time that occur when synthesized speech data is generated from the speech data of the edited object. Specifically, parameters include at least several distinct types of emotion and the mixing ratio of these different types of emotion. Additionally, parameters may also include at least one of emotion intensity, speech rate, and sound pressure level. In this embodiment, an example will be provided where parameters include several distinct types of emotion, the mixing ratio of these different types of emotion, and the intensity, speech rate, and sound pressure level of each of the different emotions.

[0056] When a user selects voice data for an editable object, they operate the playback button 40D, which instructs the playback of that voice data. When the input receiving unit 20B receives a playback instruction signal input via the operation instruction of the playback button 40D, the playback unit 20E begins playback of that voice data. Played voice data refers to the voice represented by that voice data output from the voice output unit 14C.

[0057] If the playback of voice data begins, the display control unit 20A preferably displays the waveform representing the volume of the voice data in the voice waveform display bar 40E.

[0058] When the playback of the voice data begins and the voice of the voice data begins to be output from the voice output unit 14C, the user inputs parameters for the desired playback timing by visually confirming the display screen 30B while operating the input unit 14B. Playback timing refers to each timing point in the playback of the voice data played along the time axis. That is, the input receiving unit 20B accepts input parameters for each playback timing point during the playback of the voice data of the edited object.

[0059] In detail, the user operates the input unit 14B while looking at at least one of the pointer 40F, speech rate adjustment button 40G, and gain adjustment button 40H displayed on the emotion map M on the display screen 30B, thereby inputting the desired playback timing parameters for the reproduced speech data.

[0060] The emotion map M included in display screen 30B is the same as the emotion map M described above. A pointer 40F is shown in the emotion map M. Pointer 40F refers to a user-specified location on the emotion map M.

[0061] For example, a user adjusts the position of pointer 40F in the emotion map M by operating the first operation unit 14B1, which functions as a joystick. Specifically, for example, when the tilt direction and tilt angle of the joystick 14B1 are adjusted, the position of pointer 40F displayed on the emotion map M on the display screen 30B moves in the tilt direction of the joystick by an amount corresponding to the tilt angle. By operating the first operation unit 14B1, the user adjusts the position of pointer 40F displayed in the emotion map M to a position corresponding to the desired type of emotion, the desired mixture ratio of different types of emotion, and the desired intensity of emotion. The input receiving unit 20B receives the type of emotion, the mixture ratio of multiple types of emotion, and the intensity of emotion represented by the position of pointer 40F in the emotion map M.

[0062] The speech rate and sound pressure level are adjusted by the positions of the speech rate adjustment button 40G and the gain adjustment button 40H included in the display screen 30B. For example, the user can adjust the positions of the speech rate adjustment button 40G and the gain adjustment button 40H on the display screen 30B by operating the second pedal-type operation unit 14B2, such as by using the user's foot.

[0063] For example, the second operation unit 14B2 includes a pedal corresponding to the speech rate adjustment button 40G and a pedal corresponding to the gain adjustment button 40H.

[0064] When the user adjusts the pressure of the pedal on the second operation unit 14B2 corresponding to the speech rate adjustment button 40G, the position of the speech rate adjustment button 40G displayed on the display screen 30B moves in the direction of increasing or decreasing the speech rate. The input receiving unit 20B accepts input of speech rate corresponding to the pressure of the pedal on the second operation unit 14B2 corresponding to the speech rate adjustment button 40G.

[0065] Similarly, when the user adjusts the amount of pressure applied to the pedal of the second operation unit 14B2 corresponding to the gain adjustment button 40H, the position of the gain adjustment button 40H displayed on the display screen 30B moves in the direction of increasing or decreasing the sound pressure (gain). The input receiving unit 20B receives the sound pressure input corresponding to the amount of pressure applied to the second operation unit 14B2 corresponding to the gain adjustment button 40H.

[0066] Thus, during the reproduction of speech data, the user operates at least one of the first operation unit 14B1 and the second operation unit 14B2 at each desired reproduction timing, thereby inputting a parameter including at least one of the following: the type of emotion desired for that reproduction timing, the mixing ratio of multiple types of emotions, the intensity of the emotion, the speech rate, and the sound pressure level. Additionally, the input receiving unit 20B receives the parameters for each reproduction timing during the reproduction of the speech data.

[0067] return Figure 1 Let me continue explaining.

[0068] The recording unit 20F records the received input parameters in a way that establishes a correspondence with the playback timing of the input received in the speech data. Specifically, for example, the recording unit 20F stores the received input parameters in a way that establishes a correspondence with the timestamp in the speech data representing the playback timing of the input received. Alternatively, the recording unit 20F may also record the received input parameters in a way that establishes a correspondence with the position in the speech data corresponding to the playback timing of the input received.

[0069] The reproduction unit 20E generates synthesized speech data based on speech dictionary data corresponding to the emotion for the speech data of the editing object. The emotion corresponds to the parameters established for the reproduction timing.

[0070] Specifically, the reproduction unit 20E inputs the linguistic features (linguistic information) of the speech data of the edited object at the reproduction time into speech dictionary data corresponding to multiple types of emotions included in the parameters set for the reproduction time, thereby obtaining sound feature quantities corresponding to each type of emotion. Then, the reproduction unit 20E obtains a first mixed sound feature quantity, which is obtained by mixing the obtained sound feature quantities corresponding to each type of emotion according to the mixing ratio of the emotions included in the parameters set for the reproduction time. Next, the reproduction unit 20E inputs the linguistic features of the speech data of the edited object at the reproduction time into speech dictionary data corresponding to no emotion, thereby obtaining a second sound feature quantity corresponding to no emotion. Then, the reproduction unit 20E mixes the second sound feature quantity corresponding to no emotion and the first mixed sound feature quantity at a ratio corresponding to the intensity of the emotion included in the parameters set for the reproduction time, thereby obtaining a second mixed sound feature quantity. In detail, the reproduction unit 20E obtains the following second mixed sound feature quantity: the lower the intensity of emotion, the larger the proportion of the second sound feature quantity corresponding to no emotion; the stronger the intensity of emotion, the larger the proportion of the first mixed sound feature quantity. Then, the reproduction unit 20E generates synthesized speech of the speech waveform represented by the second mixed sound feature quantity, as the synthesized speech data for the reproduction timing.

[0071] The reproduction unit 20E generates synthesized speech for each of the multiple reproduction timings along the time axis contained in the speech data by using the above-described processing of the parameters set on the reproduction timing, thereby generating synthesized speech data obtained by synthesizing the speech data according to the parameters.

[0072] Then, the recording unit 20F establishes a corresponding record of the parameters used in the generation of the synthesized speech at each reproduction time in the generated synthesized speech data.

[0073] When the synthesized speech reproduction button 40K on the display screen 30B is activated by the user's operation instruction on the input unit 14B, the reproduction unit 20E reproduces the synthesized speech data. The reproduction unit 20E reproduces the synthesized speech data by outputting the generated synthesized speech data to the speech output unit 14C. Alternatively, the reproduction unit 20E can also generate and reproduce synthesized speech data when the synthesized speech reproduction button 40K is activated by the user's operation instruction on the input unit 14B.

[0074] Figure 4This is a schematic diagram of an example of display screen 30C. Display screen 30C is an example of display screen 30. Display screen 30C is the display screen 30 displayed on display unit 14A when synthesized speech data is reproduced. When the synthesized speech reproduction button 40K is operated by the user through the operation instruction of input unit 14B, display control unit 20A displays display screen 30C on display unit 14A.

[0075] In addition to the aforementioned display screen 30B, display screen 30C also includes a playback timing image 40I. The playback timing image 40I is an image displayed in the waveform representing the synthesized speech data on the speech waveform display bar 40E, indicating the current playback timing. Therefore, as the playback time of the synthesized speech data elapses, the display control unit 20A moves the display position of the playback timing image 40I to a position in the waveform representing the synthesized speech data corresponding to the current playback timing.

[0076] In addition, the display control unit 20A preferably adjusts the positions of the pointer 40F, the speech rate adjustment button 40G, and the gain adjustment button 40H in the display screen 30C so that they correspond to the display positions of the parameters set for each reproduction timing of the synthesized speech data.

[0077] As described above, the recording unit 20F records the parameters used in the generation of synthesized speech at each playback timing corresponding to the playback timing in the synthesized speech data. During the playback of the synthesized speech data, the display control unit 20A displays a pointer 40F in the emotion map M, indicating the type, mixing ratio, and intensity of emotion represented by the parameters recorded corresponding to the current playback timing. Furthermore, during the playback of the synthesized speech data, the display control unit 20A displays the speech rate adjustment button 40G and the gain adjustment button 40H at positions indicating the speech rate and gain represented by the parameters recorded corresponding to the current playback timing, respectively.

[0078] Users sometimes wish to edit parameters. In this case, the user instructs the edit button 40J on the display screen 30B by operating the input unit 14B. When the edit button 40J is operated, the input receiving unit 20B begins to accept parameter editing. As described above, the input unit 14B can be a pointing device such as a digital pen, mouse, or trackball, or an input device such as a keyboard. Furthermore, the input unit 14B may include a first operating unit 14B1 such as a joystick and a second operating unit 14B2 such as a pedal. Therefore, the parameter input performed by the user is not limited to the first operating unit 14B1 such as a joystick and the second operating unit 14B2 such as a pedal; input can be performed by simultaneously operating one or more of the pointing devices such as a mouse, digital pen, or trackball, or a keyboard.

[0079] In detail, the user selects an edit point in the waveform representing speech data displayed in the speech waveform display bar 40E included in the display screen 30B by operating the input unit 14B. Then, the user edits the parameters corresponding to the edit point by operating the input unit 14B while looking at at least one of the pointer 40F, speech rate adjustment button 40G, and gain adjustment button 40H displayed on the emotion map M of the display screen 30B. The operation of editing these parameters is the same as the operation of inputting parameters for speech data.

[0080] That is, the user can edit at least one parameter among the following—the type of emotion, the mixing ratio of multiple types of emotion, the intensity of emotion, the speech rate, and the sound pressure level—relative to the selected editing point by operating at least one of the first operation unit 14B1 and the second operation unit 14B2. Additionally, the input receiving unit 20B receives input for editing the parameters of the selected editing point.

[0081] The recording unit 20F appends the edited parameters to the selected edit point in the synthesized speech data and records them. Specifically, for example, the recording unit 20F stores the edited input parameters in a correspondence with a timestamp representing the selected edit point in the synthesized speech data. Alternatively, the recording unit 20F may also record the edited input parameters in a correspondence with the position in the synthesized speech data corresponding to the selected edit point.

[0082] Then, the reproduction unit 20E regenerates the synthesized speech data based on the edited parameters. The reproduction unit 20E only needs to generate synthesized speech data corresponding to the edited parameters, similar to the generation of synthesized speech data based on the speech data with set parameters. The recording unit 20F records the regenerated synthesized speech data in a correspondence with the parameters set for each reproduction timing.

[0083] When the parameter input is complete, the user operates the save button 40L. The save button 40L is a button operated by the user when instructing the storage of synthesized speech data generated from the speech data with set parameters to the storage unit 16. When the save button 40L is operated, the input receiving unit 20B receives the save instruction.

[0084] When the input receiving unit 20B accepts the save instruction, the display control unit 20A displays the display screen 30, which is used to accept text information input for the synthesized speech data, on the display unit 14A.

[0085] Figure 5 This is a schematic diagram of an example of display screen 30D. Display screen 30D is an example of display screen 30. Display screen 30D is the display screen 30 that is displayed when the user operates the save button 40L.

[0086] When the user operates the save button 40L via the input unit 14B, the display control unit 20A displays a display screen 30D on the display screen 30C, overlaid with a text information input field 40M. The text information input field 40M is an input field for text information added to the synthesized speech data. For example, the user can input text information such as descriptions of the synthesized speech data into the text information input field 40M by operating the input unit 14B.

[0087] The input receiving unit 20B receives text information for the synthesized speech data via the text information input field 40M. The recording unit 20F stores the received input text information and the synthesized speech data in the storage unit 16 in a corresponding manner.

[0088] Through these processes, a corresponding document containing descriptions related to the synthesized speech data is created for each synthesized speech data. Therefore, users of the synthesized speech data can effectively reuse the data by verifying this textual information. Furthermore, by using the synthesized speech data and the textual information added to it as learning data, a learning model can be generated that outputs the textual information as the correct label based on the synthesized speech data.

[0089] Next, an example of the information processing flow performed by the speech processing aid 10 of this embodiment will be described.

[0090] Figure 6 This is a flowchart illustrating an example of the information processing flow performed by the speech processing aid 10 of this embodiment.

[0091] The display control unit 20A displays the display screen 30A on the display unit 14A (step S100). For example, when a signal indicating the start of voice data editing is input by the user through operation instructions on the input unit 14B, the display control unit 20A displays the display screen 30A, which is used to accept the type of emotion and the setting of the voice dictionary data, on the display unit 14A.

[0092] The user operates the input unit 14B while visually confirming the display screen 30A, thereby setting the type of emotion used in the setting of voice data and the voice dictionary data used for the type of emotion. The input receiving unit 20B accepts the setting of the type of emotion used in the setting of voice data and the voice dictionary data corresponding to the type of emotion (step S102).

[0093] The setting unit 20C sets the voice dictionary data received via the display screen 30A, corresponding to the type of emotion, as the voice dictionary data to be used in the editing of voice data (step S104).

[0094] The display control unit 20A displays a display screen 30B for accepting parameter settings for voice data on the display unit 14A (step S106). The display is achieved through the processing in step S106. Figure 4 The display shown is screen 30B.

[0095] The user operates the input unit 14B while visually confirming the display screen 30B, thereby inputting the filename of the voice data of the object to be edited into the voice data filename input display field 40C. The acquisition unit 20D acquires the voice data of the input filename as the voice data of the object to be edited (step S108).

[0096] The user operates the playback button 40D to instruct the playback of voice data. When the input receiving unit 20B receives the playback instruction signal input through the operation instruction of the playback button 40D, the playback unit 20E begins to reproduce the voice data obtained in step S108 (step S110).

[0097] When the playback of the speech data begins, the user listens to the played speech while simultaneously operating a first operation unit 14B1 (such as a joystick), a second operation unit 14B2 (such as a pedal), or an input unit 14B (such as a mouse) on the emotion map M displayed on the display screen 30B. This allows the user to input parameters for the desired playback timing of the played speech data. Specifically, while listening to the played speech, the user inputs parameters such as the type of emotion, the mixing ratio of multiple emotion types, the intensity of the emotion, the speech rate, and the sound pressure level for each playback timing. Then, by performing the above operations based on the playback of the speech data, the user inputs parameters including the type of emotion, the mixing ratio of multiple emotion types, the intensity of the emotion, the speech rate, and the sound pressure level for each playback timing of the speech data.

[0098] In the reproduction of voice data, the input receiving unit 20B determines whether it has received parameter input from the input unit 14B (step S112). When a negative determination is made in step S112 (step S112: No), the process proceeds to step S116, which will be described later. When a positive determination is made in step S112 (step S112: Yes), the process proceeds to step S114.

[0099] In step S114, the recording unit 20F records the parameters that were input in step S112 in a corresponding manner with the playback timing of the input that received the parameters in the speech data (step S114).

[0100] Next, the playback unit 20E determines whether the playback of the voice data has ended (step S116). For example, the playback unit 20E determines whether the playback has ended by the final timing on the timeline of the voice data that started playback in step S110, thereby performing the determination in step S116. When a negative determination is made in step S116, the process returns to step S112. When a positive determination is made in step S116 (step S116: Yes), the process proceeds to step S118.

[0101] The reproduction unit 20E generates synthesized speech data by synthesizing speech dictionary data corresponding to the emotion obtained in step S108, and the emotion corresponds to the parameters established for the reproduction timing (step S118).

[0102] The recording unit 20F establishes a corresponding record of the parameters used in the generation of the synthesized speech at each playback timing in the generated synthesized speech data (step S120).

[0103] Furthermore, users can listen to the synthesized speech data while simultaneously setting or editing the playback timing parameters. Additionally, as described above, the processing unit 20 can also accept the selection of edit points in the waveform of the synthesized speech data displayed on the display screen 30B, and accept the editing of the parameters at those edit points, thereby establishing a correspondence between the edited parameters and the positions in the synthesized speech data corresponding to the selected edit points and recording them accordingly.

[0104] The display control unit 20A displays a screen 30D with the text input field 40M superimposed on it. The input receiving unit 20B receives text information input into the text input field 40M by the user's operation instructions on the input unit 14B (step S122).

[0105] The recording unit 20F records the text information received in step S138 and the synthesized speech data generated in step S132, establishing a correspondence (step S124). Then, the routine ends.

[0106] As explained above, the speech processing aid 10 of this embodiment includes an input receiving unit 20B and a recording unit 20F. The input receiving unit 20B receives parameters that include at least a plurality of different types of emotions and the mixing ratio of these emotions during the reproduction of the speech data of the object being edited. The recording unit 20F records the received parameters in a corresponding manner with the reproduction timing of the input received in the speech data.

[0107] In the prior art, although it is possible to set parameters such as the deformation ratio for the overall voice data, it is difficult to set parameters related to the shift of emotions over time in detail.

[0108] On the other hand, in the speech processing assist device 10 of this embodiment, during the reproduction of the speech data of the editing object, input parameters are received that include at least a plurality of different types of emotions and a mixture ratio of the plurality of types of emotions. Then, the speech processing assist device 10 records the received parameters in a corresponding manner with the reproduction timing of the input parameters in the speech data.

[0109] Therefore, in the speech processing aid 10 of this embodiment, parameters including at least the type of emotion and the mixing ratio of emotions can be set for each reproduction timing of speech data along the time axis.

[0110] Therefore, the speech processing aid 10 of this embodiment can be applied to parameters that can be set in detail in relation to the progression of emotions that change over time.

[0111] Furthermore, in the speech processing aid 10 of this embodiment, parameters can be set for speech data to achieve dynamic speech expression that changes in emotion and the intensity of emotion over time.

[0112] Furthermore, in the speech processing aid 10 of this embodiment, the parameters for each playback timing are input during the playback of speech data, thus enabling the setting of parameters for each playback timing in real time while the speech data is being played back. Additionally, in the speech processing aid 10 of this embodiment, the user can set the parameters for each playback timing via the operation input unit 14B. Therefore, the speech processing aid 10 of this embodiment reduces the burden on the user for input and editing, and allows for detailed setting of parameters that vary along the time axis.

[0113] Next, the hardware structure of the speech processing aid 10 of this embodiment will be described.

[0114] Figure 7 This is a hardware structure diagram of an example of the speech processing auxiliary device 10 of this embodiment.

[0115] The speech processing aid 10 of this embodiment includes a control device such as a CPU 10A, a storage device such as a ROM (Read Only Memory) 10B and a RAM (Random Access Memory) 10C, an HDD (Hard Disk Drive) 10D, an I / F 10E for communication via a network connection, and a bus 10F connecting each part.

[0116] The program executed in the speech processing aid 10 of this embodiment is provided by pre-loading ROM 10B or the like.

[0117] The program executed in the speech processing aid 10 of this embodiment may also be configured to be provided as a computer program product as a file recorded in a computer-readable recording medium such as CD-ROM (Compact Disk Read Only Memory), floppy disk (FD), CD-R (Compact Disk Recordable), DVD (Digital Versatile Disc) in an installable or executable form.

[0118] Furthermore, the program executed in the speech processing assist device 10 of this embodiment can also be stored on a computer connected to a network such as the Internet, and provided by downloading it via the network. Alternatively, the program executed in the speech processing assist device 10 of this embodiment can be provided or distributed via a network such as the Internet.

[0119] The program executed in the speech processing aid 10 of this embodiment enables the computer to function as a component of the speech processing aid 10. This computer is a CPU 10A capable of reading programs from a computer-readable storage medium and executing them on the main storage device.

[0120] Furthermore, in the above embodiment, the speech processing assist device 10 was described as a single unit. However, the speech processing assist device 10 may also be composed of multiple physically separate devices that are communicatively connected via a network or the like.

[0121] In addition, the speech processing aid 10 described above can also be implemented as a virtual machine operating on a cloud system.

[0122] Furthermore, the embodiments of the present invention have been described above, but these embodiments are provided as examples and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, and are included within the scope of the invention described in the patent claims and its equivalents.

[0123] Explanation of symbols 10. Speech Processing Auxiliary Device 20A Display Control Unit 20B Input Receiving Unit 20C Setting Department 20D Acquisition Department 20E Reproduction Department 20F Recording Department

Claims

1. A speech processing auxiliary device, comprising: The input receiving unit, in the reproduction of the speech data of the edited object, accepts input parameters that include at least several distinct types of emotions and the mixing ratio of these emotions; and The recording unit records the parameters that have been input and establishes a correspondence between the input parameters and the playback timing of the input parameters in the speech data.

2. The speech processing auxiliary device according to claim 1, wherein, The parameters also include at least one of the intensity of emotion, speech rate, and sound pressure level.

3. The speech processing auxiliary device according to claim 1, wherein, It includes a display control unit that displays a screen containing an emotion mapping diagram representing the correlation between multiple types of emotions. The input receiving unit accepts at least one of the following parameters: the type of emotion corresponding to a specified location of the user on the emotion map, the mixing ratio, and the intensity of the emotion.

4. The speech processing auxiliary device according to claim 1, wherein, The input receiving unit accepts settings of voice dictionary data corresponding to the type of emotion used in the settings for voice data.

5. The speech processing auxiliary device according to claim 1, wherein, The device includes a reproduction unit that reproduces synthesized speech data based on speech dictionary data corresponding to the type of emotion, wherein the type of emotion corresponds to parameters established for the reproduction timing.

6. The speech processing auxiliary device according to claim 5, wherein, The input receiving unit accepts the editing of the parameters. The recording unit stores the edited parameters in correspondence with the selected edit points in the synthesized speech data.

7. The speech processing auxiliary device according to claim 5, wherein, The input receiving unit accepts text information as input for the synthesized speech data. The recording unit records the text information in a corresponding manner with the synthesized speech data.

8. A speech processing assistance method, executed by a speech processing assistance device, comprising: The input receiving step involves receiving input parameters that include at least several different types of emotions and the mixing ratio of these emotions during the reproduction of the speech data of the edited object. as well as The recording step involves establishing a correspondence between the received input parameters and the playback timing of the input in the speech data that received the parameters, and then recording them.

9. A speech processing aid program for causing a computer to perform the following steps: The input receiving step, in the reproduction of the speech data of the edited object, accepts input parameters containing at least several distinct types of emotions and the mixing ratio of these emotions; and The recording step involves establishing a correspondence between the received input parameters and the playback timing of the input in the speech data that received the parameters, and then recording them.

Citation Information

Patent Citations

  • Voice synthetic method and device therefor

    JP1997050295A