Voice processing support system, voice processing support method, and voice processing support program
The voice processing support system allows users to easily adjust and generate processed voice data by selecting and editing character strings, addressing the lack of user-friendly mechanisms in conventional systems.
Patent Information
- Application Number
- JP2022118791
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-07-26
AI Technical Summary
Conventional voice processing systems lack user-friendly mechanisms for adjusting processed audio data, making it difficult for users to easily modify recorded voice data.
A voice processing support system and method that includes a voice processing support device and an information processing device, enabling users to select, edit, and generate processed voice data through a series of interactive screens and processes, allowing for easy adjustment of recorded voice data.
Enables users to easily adjust and generate processed voice data by selecting and modifying character strings, facilitating user-friendly voice data processing.
Smart Images

Figure 0007788963000001 
Figure 0007788963000002 
Figure 0007788963000003
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention is a voice processing support system ,audio processing support method, and Voice processing support program Mu Regarding. [Background technology]
[0002] As a technique for processing voice data, a technique for synthesizing recorded voice data and synthetic voice data has been disclosed. For example, a conventional technique automatically extracts a substring using recorded voice and a substring using synthetic voice from an input character string, and synthesizes the recorded voice and the synthetic voice using the results of the automatic extraction.
[0003] However, in the conventional technology, processed audio data is generated by processing audio data through automatic extraction and automatic synthesis without user operation instructions, so it is difficult for the conventional technology to support easy adjustments by users regarding the processing of recorded audio data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-107454 [Patent Document 2] Japanese Patent Application Laid-Open No. 2009-20264 [Patent Document 3] Japanese Patent Application Laid-Open No. 2003-295880 Summary of the Invention [Problem to be solved by the invention]
[0005] The problem to be solved by the present invention is to provide a voice processing support system that can support a user in easily adjusting the processing of recorded voice data. system ,audio processing support method, and Voice processing support program MThe purpose is to provide. [Means for solving the problem]
[0006] In an embodiment The voice processing system includes a voice processing support device and an information processing device. The voice processing support device includes a first receiving unit, a display control unit, a second receiving unit, and a generation control unit. The first receiving unit receives selection of target recorded voice data, which is the basic recorded voice data to be processed, from one or more recorded basic recorded voice data. The display control unit converts the target recorded voice data into a basic character string and displays it. The second receiving unit receives designation of a target character string to be changed from the displayed basic character string. The generation control unit generates processed voice data according to the target recorded voice data and the target character string. The information processing device includes a receiving unit that receives the processing-related information, and a processing unit that processes the target recorded voice data based on the processing-related information to generate processed voice data. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a diagram showing a voice processing support system. [Figure 2A] FIG. 2A is a schematic diagram of the processing support screen. [Figure 2B] FIG. 2B is a schematic diagram of the processing support screen. [Figure 2C] FIG. 2C is a schematic diagram of the processing support screen. [Figure 2D] FIG. 2D is a schematic diagram of the processing support screen. [Figure 2E] FIG. 2E is a schematic diagram of the processing support screen. [Figure 2F] FIG. 2F is an explanatory diagram of an example of generation of processed voice data. [Figure 2G] FIG. 2G is a schematic diagram of the setting change screen. [Figure 2H] FIG. 2H is a schematic diagram of the detailed editing screen. [Figure 3] FIG. 3 is a flowchart showing the flow of information processing executed by the voice processing support device. [Figure 4] FIG. 4 is a flowchart showing the flow of information processing executed by the information processing device. [Figure 5]FIG. 5 is a diagram showing the hardware configuration. DETAILED DESCRIPTION OF THE INVENTION
[0008] Please refer to the attached drawing below for voice processing support. system ,audio processing support method, and Voice processing support program M Explain in detail.
[0009] FIG. 1 is a diagram showing an example of a voice processing support system 1 according to this embodiment.
[0010] The voice processing support system 1 includes a voice processing support device 10 and an information processing device 30.
[0011] The voice processing support device 10 and the information processing device 30 are configured to be able to exchange data via a network NW or the like. The voice processing support device 10 and the information processing device 30 may be configured so that various data generated by the voice processing support device 10 can be used by the information processing device 30. For this reason, the voice processing support device 10 and the information processing device 30 may be configured so that they can exchange data via various storage media such as a USB (Universal Serial Bus) memory.
[0012] The voice processing support device 10 is an information processing device for supporting the processing of recorded voice data.
[0013] The voice processing support device 10 includes a memory unit 12, an output unit 14, an input unit 16, a communication unit 18, and a processing unit 20. The memory unit 12, the output unit 14, the input unit 16, the communication unit 18, and the processing unit 20 are connected to each other via a bus 19 so as to be able to communicate with each other.
[0014] The storage unit 12 stores various types of data. The storage unit 12 is, for example, a semiconductor memory element such as a RAM (Random Access Memory), a flash memory, a hard disk, an optical disk, or the like. The storage unit 12 may be a storage device provided outside the voice processing assistance device 10. The storage unit 12 may also be a storage medium. Specifically, the storage medium may store or temporarily store programs and various types of information downloaded via a LAN (Local Area Network), the Internet, or the like. The storage unit 12 may also be composed of multiple storage media.
[0015] In this embodiment, the storage unit 12 stores one or more basic recorded audio data 70 in advance.
[0016] The basic recorded voice data 70 is recorded voice data obtained by recording the voice spoken by the user. In detail, the basic recorded voice data 70 is pre-recorded voice data that can be provided as an option for processing in the voice processing support system 1, among the recorded voice data.
[0017] For example, a user speaks lines contained in a script that is the basis for a performance, in a voice that corresponds to the scene in which the lines are spoken. Lines are words uttered by a speaker who appears in the play or creative work to be performed. A speaker is a user who is the target of speaking the lines.
[0018] For example, a user speaks lines while adjusting acoustic features such as prosody and accent. Prosody includes at least one of intonation, tone, stress, duration, and rhythm. Accent includes at least one of pitch accent and stress accent.
[0019] In the voice processing support device 10, the voice spoken by the user is collected by the microphone 16B and stored in the storage unit 12 in advance as basic recorded voice data 70.
[0020] In this embodiment, the basic recorded voice data 70 is recorded voice data recorded by a single utterance of a line by the user, as an example. However, the basic recorded voice data 70 is not limited to recorded voice data of a line uttered by the user. For example, the basic recorded voice data 70 may be recorded voice data obtained by recording a voice uttered by the user in everyday conversation, etc. Furthermore, the voice processing support device 10 may store recorded voice data recorded by another information processing device in the storage unit 12 in advance as the basic recorded voice data 70.
[0021] The output unit 14 is an output device for outputting various types of information. In this embodiment, the output unit 14 includes a display unit 14A and a speaker 14B. The display unit 14A displays various types of information. The display unit 14A is, for example, a display such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display, or a projection device. The speaker 14B outputs sound.
[0022] The input unit 16 is an input device for receiving various instructions from the user. In this embodiment, the input unit 16 includes an operation input unit 16A and a microphone 16B. The operation input unit 16A is an input device for receiving operation instructions from the user. The operation input unit 16A is, for example, a pointing device such as a digital pen, a mouse, or a trackball, or an input device such as a keyboard. The microphone 16B is an input device for inputting voice. The display unit 14A and the microphone 16B may be an integrally configured touch panel.
[0023] The communication unit 18 communicates with an external information processing device via the network NW. In this embodiment, the communication unit 18 communicates with an information processing device 30 via the network NW.
[0024] The processing unit 20 executes various types of information processing. The processing unit 20 includes a display control unit 21, a reception unit 22, a conversion unit 23, a generation control unit 24, an acquisition unit 25, a playback control unit 26, and a storage processing unit 27. The reception unit 22 includes a first reception unit 22A, a second reception unit 22B, a third reception unit 22C, a fourth reception unit 22D, a fifth reception unit 22E, and a sixth reception unit 22F.
[0025] The display control unit 21, the reception unit 22, the first reception unit 22A, the second reception unit 22B, the third reception unit 22C, the fourth reception unit 22D, the fifth reception unit 22E, the sixth reception unit 22F, the conversion unit 23, the generation control unit 24, the acquisition unit 25, the playback control unit 26, and the storage processing unit 27 are realized, for example, by one or more processors. For example, each of the above units may be realized by having a processor such as a CPU (Central Processing Unit) execute a program, i.e., by software. Each of the above units may be realized by a processor such as a dedicated IC (Integrated Circuit), i.e., by hardware. Each of the above units may be realized by a combination of software and hardware. When multiple processors are used, each processor may realize one of the units, or two or more of the units.
[0026] At least one of the above units may be installed in a cloud server that executes processing on the cloud.
[0027] The display control unit 21 displays various images on the display unit 14A.
[0028] The reception unit 22 receives operation instructions from the user and recorded audio data of the user's speech from the input unit 16. The reception unit 22 receives information representing the user's operation instructions from the operation input unit 16A. The reception unit 22 also receives recorded audio data of the user's speech from the microphone 16B.
[0029] In this embodiment, the receiving unit 22 has a first receiving unit 22A, a second receiving unit 22B, a third receiving unit 22C, a fourth receiving unit 22D, a fifth receiving unit 22E, and a sixth receiving unit 22F.
[0030] The first receiving unit 22A receives a selection of target recorded voice data from one or more recorded basic recorded voice data 70. The target recorded voice data means the basic recorded voice data 70 selected by the user as the target for processing from among the pre-recorded basic recorded voice data 70. The target for processing means the target for processing.
[0031] In this embodiment, the first accepting unit 22A accepts the selection of the target recorded voice data via the editing support screen that the display control unit 21 displays on the display unit 14A.
[0032] Fig. 2A is a schematic diagram of an example of a processing support screen 50. When a start instruction for voice processing support is input by a user operating the operation input unit 16A, for example, the display control unit 21 displays the processing support screen 50 shown in Fig. 2A.
[0033] The processing support screen 50 includes a selection field 60A. The selection field 60A is an input field for receiving a selection of basic recorded voice data 70 to be processed from one or more basic recorded voice data 70 desired by the user.
[0034] For example, consider a situation in which the selection field 60A on the processing support screen 50 is operated by a user's operation instruction via the operation input unit 16A. When the selection field 60A is operated, the display control unit 21 displays a list of one or more basic recorded voice data 70 stored in the storage unit 12 on the display unit 14A. The user operates the operation input unit 16A to select one basic recorded voice data 70 to be processed from the displayed one or more basic recorded voice data 70. The first accepting unit 22A accepts the one basic recorded voice data 70 selected via the selection field 60A as the target recorded voice data.
[0035] Fig. 2B is a schematic diagram of an example of the processing support screen 50. Fig. 2B shows an example of a scene in which the first reception unit 22A receives the selection of the target recorded voice data 72. When the first reception unit 22A receives the selection of the target recorded voice data 72, the display control unit 21 displays, in the selection field 60A of the processing support screen 50, the file name of the target recorded voice data 72, which is the basic recorded voice data 70 selected by the user as the processing target.
[0036] In this embodiment, the data format of the audio data such as the basic recorded audio data 70 and the target recorded audio data 72 is WAV (Waveform Audio File Format) as an example. However, the data format of the audio data is not limited to the WAV file format.
[0037] Returning to Figure 1, we continue the explanation.
[0038] The conversion unit 23 converts the target recorded voice data 72 for which selection has been accepted into a basic character string. A basic character string is data in which the voice represented by the target recorded voice data 72 is expressed as a character string. The conversion unit 23 can convert the target recorded voice data 72 into a basic character string using a known voice conversion technology that converts voice data into a character string. Voice data is data that represents voice. In this embodiment, voice data is used as a general term to refer to recorded voice data such as the basic recorded voice data 70 and the target recorded voice data 72, as well as synthesized voice data other than recorded voice data.
[0039] The display control unit 21 converts the selected target recorded voice data 72 into a basic character string and displays it. In this embodiment, the display control unit 21 displays the basic character string converted by the conversion unit 23 on the display unit 14A.
[0040] 2B, when the first receiving unit 22A receives the selection of the target recorded voice data 72, the display control unit 21 displays the basic character string 82 of the selected target recorded voice data 72 in the display field 60B of the editing support screen 50.
[0041] The display field 60B is a display field for the output target character string 86. The output target character string 86 is a character string of processed voice data to be output by the voice processing support device 10. The output target character string 86 displayed in the display field 60B changes depending on the input content related to processing input by the user's operation instruction (described in detail later).
[0042] When the first reception unit 22A receives the selection of the target recorded voice data 72, the display control unit 21 displays the basic character string 82 of the target recorded voice data 72 whose selection has been received in the display field 60B of the processing support screen 50.
[0043] 2B shows an example of a scene in which the basic character string 82, "I have a dear friend named Masao," is displayed in the display field 60B. By visually checking the editing assistance screen 50, the user can easily confirm the basic character string 82 of the selected target recorded voice data 72.
[0044] When the first reception unit 22A receives the selection of the target recorded voice data 72, the display control unit 21 displays a back button 60C and a play button 60F on the processing support screen 50 so that the user can select them. The back button 60C is an input button that the user operates to instruct the user to return to the previously displayed display screen. The play button 60F is an input button that the user operates to instruct the user to play the voice data.
[0045] Returning to Figure 1, we continue the explanation.
[0046] The second receiving unit 22B receives the designation of a character string to be changed from among the displayed basic character string 82. The character string to be changed is a character string to be changed from among the character strings that make up the basic character string 82. In other words, the character string to be changed is a character string that corresponds to the voice of the voice section in the target recorded voice data 72 that is to be changed. The character string to be changed may be one character or multiple characters.
[0047] This will be explained using Figure 2C, which is a schematic diagram of an example of the processing support screen 50. Figure 2C shows an example of a scene in which the second receiving unit 22B receives the designation of the character string 82A to be changed.
[0048] For example, the user operates operation input unit 16A to specify a target string 82A to be changed from among basic strings 82 displayed in display field 60B. Fig. 2C shows an example of a scene in which "Masao" from basic string 82 "I have an important friend named Masao" is specified as target string 82A.
[0049] When the user operates operation input unit 16A to specify character string 82A to be changed, second reception unit 22B receives the specification of character string 82A to be changed. In the example shown in Fig. 2C, second reception unit 22B receives character string 82A to be changed, "Masao."
[0050] The display control unit 21 displays the area of the character string 82A to be changed that has been specified by the user in a display format that differs from that of the other unspecified character strings in the basic character string 82. FIG. 2C shows an example in which the background of the area of the specified changed character string 84B is highlighted. By operating the operation input unit 16A while viewing the processing support screen 50, the user can easily specify the changed character string 84B in the basic character string 82 and confirm the specified changed character string 84B.
[0051] When the second receiving unit 22B receives the designation of the change target character string 82A, the display control unit 21 further displays a selectable save button 60G on the processing support screen 50. The save button 60G is an input button that the user operates to instruct the storage of processed voice data corresponding to the output target character string 86, which is the character string displayed in the display field 60B. Note that the display control unit 21 may display the save button 60G on the processing support screen 50 in a selectable manner before receiving the designation of the change target character string 82A. For example, the display control unit 21 may display the save button 60G on the processing support screen 50 in a selectable manner at the stage shown in FIGS. 2A and 2B.
[0052] Continuing the explanation, returning to Fig. 1, the generation control unit 24 generates processed voice data according to the target recorded voice data 72 and the target character string 82A.
[0053] 2C, a situation is assumed in which specification of a target character string 82A included in a basic character string 82 is accepted. In this case, since information regarding audio processing has not yet been input, the generation control unit 24 generates the target recorded audio data 72 as processed audio data.
[0054] Assume that a user operates the play button 60F in response to an instruction from the operation input unit 16A. When the play button 60F is operated, the sixth reception unit 22F receives an instruction to play the processed audio data. When the sixth reception unit 22F receives the instruction to play the processed audio data, the playback control unit 26 plays the processed audio data. In detail, the playback control unit 26 outputs the processed audio data recently generated by the generation control unit 24 from the microphone 16B.
[0055] The user can preview the processed audio represented by the processed audio data by operating the play button 60F using the operation input unit 16A.
[0056] Also, it is assumed that the user operates save button 60G in response to an operation instruction from operation input unit 16A. When save button 60G is operated, receiving unit 22 receives the save instruction.
[0057] Returning to Figure 1, we continue the explanation.
[0058] The storage processing unit 27 stores the processed voice data generated by the generation control unit 24 in the storage unit 12. In detail, when the reception unit 22 receives a save instruction from the operation input unit 16A, the storage processing unit 27 stores the processed voice data generated by the generation control unit 24 in the storage unit 12.
[0059] Furthermore, the storage processing unit 27 stores the processing-related information in the storage unit 12 in association with the processed voice data.
[0060] The processing-related information is information related to the processing of the processed voice data. More specifically, the processing-related information is information used to generate the processed voice data. Specifically, the processing-related information includes at least one of the target recorded voice data 72 used to generate the processed voice data and identification information of the target recorded voice data 72, a basic character string 82, and a character string to be changed 82A. Furthermore, depending on the processing content applied to the processed voice data, the processing-related information may further include at least one of a changed character string, training recorded voice data, setting change information, and detailed correction information. Details of the training recorded voice data, the changed character string, setting change information, and detailed correction information will be described later.
[0061] Assume a situation in which the user operates the operation input unit 16A to select target recorded voice data 72, specify a character string 82A to be changed in a basic character string 82 of the target recorded voice data 72, and then operate the save button 60G. Also assume a situation in which the generation control unit 24 generates the target recorded voice data 72 as processed voice data. In this case, the storage processing unit 27 stores processing-related information including the target recorded voice data 72 or identification information of the target recorded voice data 72, the basic character string 82, and the character string 82A to be changed in the storage unit 12 in association with the processed voice data.
[0062] The third receiving unit 22C receives input of a post-change string for the change target string 82A. The post-change string is a string obtained by replacing the change target string 82A with another string desired by the user. In other words, the post-change string is a string representing a phoneme or a group of phonemes after replacing the phoneme of the change target voice section corresponding to the change target string 82A in the target recorded voice data 72. The post-change string may consist of one character or multiple characters.
[0063] This will be explained using Figure 2D. Figure 2D is a schematic diagram of an example of the processing support screen 50. Figure 2D shows an example of a scene in which the third receiving unit 22C has received input of a changed character string 84B.
[0064] As described with reference to Fig. 2C, a situation is assumed in which the user operates operation input unit 16A to select "Masao" from basic character strings 82 displayed in display field 60B as character string to be changed 82A. Then, a situation is assumed in which the user further operates operation input unit 16A to input post-change character string 84B "Takumi" in place of character string to be changed 82A "Masao."
[0065] In this case, the third receiving unit 22C receives the input of the changed character string 84B "TAKUMI" from the operation input unit 16A.
[0066] The display control unit 21 replaces the change target string 82A in the basic string 82 with the changed string 84B, and displays it in the display field 60B of the processing support screen 50. Therefore, the processing support screen 50 displays, as the output target string 86, a string in which the part of the change target string 82A included in the basic string 82 has been replaced with the changed string 84B.
[0067] When the input of the post-change character string 84B is accepted, the generation control unit 24 generates processed voice data according to the target recorded voice data 72 and the post-change character string 84B for the change target character string 82A.
[0068] In detail, the generation control unit 24 generates processed audio data by combining the post-change character string audio data of the post-change character string 84B with the post-change voice section corresponding to the post-change character string 82A in the target recorded audio data 72. The post-change voice section is the voice section of the phoneme or phoneme group represented by the post-change character string 82A in the target recorded audio data 72.
[0069] Specifically, as shown in Fig. 2D, a situation is assumed in which a post-change character string 84B "Takumi" is input in place of a character string 82A to be changed "Masao" included in a basic character string 82. In this case, the generation control unit 24 identifies a target voice section to be changed, which is a voice section corresponding to the character string 82A to be changed "Masao", in the target recorded voice data 72. This identification may be performed using a known method.
[0070] The generation control unit 24 also generates post-change character string audio data, which is the audio of the post-change character string 84B "Takumi." The generation control unit 24 may generate the post-change character string audio data of the post-change character string 84B "Takumi" by using, for example, a known conversion method for converting a character string into audio data of synthetic voice.
[0071] Then, the generation control unit 24 synthesizes the generated post-change character string voice data with the specified voice section to be changed in the target recorded voice data 72, thereby generating processed voice data.
[0072] When the third reception unit 22C receives the input of the changed character string 84B, the display control unit 21 displays a detail edit button 60D, a simple teach button 60E, and a change settings button 60H in addition to the back button 60C, the play button 60F, and the save button 60G so that they can be further selected on the processing support screen 50. The detail edit button 60D, the simple teach button 60E, and the change settings button 60H will be described in detail later.
[0073] At this stage, it is assumed that the user operates the play button 60F or the save button 60G in response to an instruction from the operation input unit 16A. That is, it is assumed that the user operates the operation input unit 16A to select the target recorded voice data 72, specify the target character string 82A to be changed in the basic character string 82 of the target recorded voice data 72, and further input the changed character string 84B for the target character string 82A, and then operate the play button 60F or the save button 60G.
[0074] In this case, the playback control unit 26 and the storage processing unit 27 may each perform the same processing as described above.
[0075] More specifically, it is assumed that the user operates the operation input unit 16A to instruct the playback button 60F. In this case, the playback control unit 26 plays back the processed voice data that was generated immediately before by the generation control unit 24. Specifically, the playback control unit 26 plays back the processed voice data representing the output target string 86 "I have an important friend named Takumi" in which the change target string 82A "Masao" in the basic string 82 is changed to the changed string 84B "Takumi."
[0076] Also, assume that in this situation, the user operates the operation input unit 16A to instruct operation of the save button 60G. In this case, the storage processing unit 27 stores the processed voice data generated by the generation control unit 24 in the storage unit 12. The storage processing unit 27 also stores processing-related information including the target recorded voice data 72 or identification information of the target recorded voice data 72, the basic character string 82, the change target character string 82A, and the changed character string 84B in the storage unit 12 in association with the processed voice data.
[0077] Next, assume that the user operates the simple instruction button 60E by operating the operation input unit 16A. The simple instruction button 60E is an input button that the user operates to instruct recording of new spoken voice for the output target character string 86.
[0078] 1, when the simple instruction button 60E is operated, the acquisition unit 25 acquires, as the recorded instruction voice data 74, a voice uttered by the user for the output target string 86 obtained by converting the change target string 82A included in the basic string 82 into the changed string 84B.
[0079] This will be explained using Figure 2E. Figure 2E is a schematic diagram of an example of the processing support screen 50. Figure 2E shows, as an example, a scene in which the simple instruction button 60E is instructed to be operated after the third reception unit 22C receives input of the changed character string 84B.
[0080] When simple teaching button 60E is operated, reception unit 22 receives a recording instruction from operation input unit 16A. Upon receiving the recording instruction from reception unit 22, acquisition unit 25 starts recording the audio data collected by microphone 16B and acquires the recorded audio data as teaching recorded audio data.
[0081] For example, after operating the simple instruction button 60E, the user speaks a voice with desired acoustic features while visually checking the output target character string 86 displayed in the display field 60B. At this stage, the display field 60B displays a character string obtained by changing the change target character string 82A of the basic character string 82 to the changed character string 84B as the output target character string 86. The user speaks a voice with desired acoustic features for the output target character string 86 "I have an important friend named Takumi" obtained by changing the change target character string 82A "Masao" in the basic character string 82 to the changed character string 84B "Takumi."
[0082] For example, if the user operates the simple instruction button 60E again, the user's speech from the previous operation of the simple instruction button 60E to the current operation of the simple instruction button 60E is recorded. The acquisition unit 25 then acquires the user's speech for the output target character string 86, "I have an important friend named Takumi," as recorded speech data for instruction.
[0083] 2E, while the user is recording a speech, the display control unit 21 may change the simple instruction button 60E to a recording button 60E' displaying the text "recording." When the recording ends, the display control unit 21 displays the simple instruction button 60E in place of the recording button 60E'.
[0084] 1, when the acquisition unit 25 acquires the recorded voice data for instruction, the generation control unit 24 generates the target recorded voice data 72, a post-change character string 84B for the target character string 82A, and processed voice data corresponding to the recorded voice data for instruction.
[0085] This will be explained using Fig. 2F, which is an explanatory diagram of an example of how the processed voice data 76 is generated.
[0086] Assume that the acquisition unit 25 has acquired the instruction recorded voice data 74. As described above, the instruction recorded voice data 74 is recorded voice data in which the user speaks the output target string 86 obtained by converting the change target string 82A included in the basic string 82 into the changed string 84B.
[0087] The generation control unit 24 identifies post-change recorded audio data 74B of an audio section corresponding to the post-change character string 84B in the teaching recorded audio data 74. The generation control unit 24 also identifies a change-target audio section 72A, which is an audio section corresponding to the change-target character string 82A, in the target recorded audio data 72. The generation control unit 24 then generates processed audio data 76 by combining the change-target audio section 72A in the target recorded audio data 72 with the post-change recorded audio data 74B identified from the teaching recorded audio data 74.
[0088] Specifically, the generation control unit 24 synthesizes the post-change recorded voice data 74B corresponding to the post-change character string 84B "Takumi" included in the teaching recorded voice data 74 with the voice section 72A to be changed, which is the character string 82A "Masao" included in the basic character string 82 "I have an important friend named Masao" in the target recorded voice data 72. Through these synthesis processes, the generation control unit 24 generates processed voice data 76 corresponding to the output target character string 86.
[0089] In detail, for example, the generation control unit 24 generates the processed voice data 76 by sequentially or collectively processing the following processes.
[0090] The generation control unit 24 adjusts the pitch of the voice represented by the changed recorded voice data 74B of the voice section corresponding to the changed character string 84B identified from the training recorded voice data 74 to the pitch of the voice of the voice section 72A to be changed in the target recorded voice data 72. The pitch of the voice may be referred to as the musical interval, pitch, or key.
[0091] The generation control unit 24 also projects the prosody of the change-target voice section 72A in the target recorded voice data 72 onto the pitch-adjusted post-change recorded voice data 74B. Projecting prosody means applying prosody. In other words, the generation control unit 24 adjusts the prosody of the pitch-adjusted post-change recorded voice data 74B so that it matches the prosody of the change-target voice section 72A in the target recorded voice data 72.
[0092] Then, the generation control unit 24 generates processed voice data 76 by synthesizing the post-change recorded voice data 74B onto which the prosody has been projected with the voice section 72A to be changed in the subject recorded voice data 72.
[0093] A known method may be used to synthesize the changed recorded audio data 74B with the target recorded audio data 72. Examples of known methods used for synthesis include mixing using silent parts, cross-fading, and the like.
[0094] Specifically, the generation control unit 24 identifies the post-change recorded voice data 74B of the post-change character string 84B "Takumi" from the recorded voice training data 74 of the speech of the output target character string 86 "I have an important friend named Takumi" newly recorded by the user. Then, the generation control unit 24 adjusts the voice pitch of the post-change recorded voice data 74B of the post-change character string 84B "Takumi" to the voice pitch of the change target voice section 72A, which is the voice section of "Masao" in the target recorded voice data 72 of the basic character string 82 "I have an important friend named Masao."
[0095] In addition, the generation control unit 24 adjusts the prosody of the changed recorded voice data 74B ``Takumi'', whose pitch has been adjusted, to the prosody of the target recorded voice data 72 ``Masao'' in the target recorded voice data 72 of the basic string 82 ``I have an important friend named Masao.''
[0096] Then, the generation control unit 24 generates processed voice data 76 by synthesizing the post-change recorded voice data 74B ``Takumi'' with adjusted prosody into the voice section of the voice section 72A to be changed ``Masao'' in the target recorded voice data 72.
[0097] At this stage, it is assumed that the user operates the play button 60F or the save button 60G in response to an instruction from the operation input unit 16A. That is, it is assumed that the save button 60G is operated at the stage when the processed audio data 76 is generated using the recorded audio data 74 for training.
[0098] In this case, the playback control unit 26 and the storage processing unit 27 may each perform the same processing as described above.
[0099] Specifically, it is assumed that the user operates the operation input unit 16A to instruct the playback button 60F. In this case, the playback control unit 26 plays back the processed audio data 76 that was generated immediately before by the generation control unit 24. Specifically, the playback control unit 26 plays back the processed audio data 76 obtained by combining the target-of-change audio section 72A corresponding to the target-of-change character string 82A in the target recorded audio data 72 with the post-change recorded audio data 74B corresponding to the post-change character string 84B in the teaching recorded audio data 74.
[0100] Also, assume that in this situation, the user operates the operation input unit 16A to instruct operation of the save button 60G. In this case, the storage processing unit 27 stores the processed voice data 76 linearly generated by the generation control unit 24 in the storage unit 12. The storage processing unit 27 also stores processing-related information including the target recorded voice data 72 or identification information of the target recorded voice data 72, the basic character string 82, the change target character string 82A, the changed character string 84B, and the teaching recorded voice data 74 in the storage unit 12 in association with the processed voice data.
[0101] Next, it is assumed that the setting change button 60H is operated in response to an operation instruction from the operation input unit 16A by the user.
[0102] The setting change button 60H, which will be described with reference to FIG. 2E, is an input button that the user operates to instruct a change in the settings of at least one of the acoustic features and synthesis method of the speech of the post-change character string 84B.
[0103] When the setting change button 60H is operated by the user in response to an operation instruction from the operation input unit 16A, the display control unit 21 displays the setting change screen 52 on the display unit 14A.
[0104] FIG. 2G is a schematic diagram of an example of the setting change screen 52. The setting change screen 52 includes a setting change input field 62A and a setting change reflect button 62B. The setting change input field 62A is an input field for accepting setting changes for at least one of the acoustic features and synthesis method of the voice of the changed character string 84B. FIG. 2G shows, as an example, input fields for adjusting the pitch and gain (volume) of the voice as acoustic features. FIG. 2G also shows, as an example, an input field for adjusting crossfade as a synthesis method.
[0105] The user operates the operation input unit 16A while viewing the setting change screen 52 to input setting change information for at least one of the acoustic features and synthesis method of the speech of the post-change character string 84B. The setting change information is information that represents at least one of the acoustic features and synthesis method of the post-change character string 84B after the setting change.
[0106] Furthermore, when the setting change reflect button 62B is operated by the user in response to an operation instruction from the operation input unit 16A, the operation input unit 16A outputs the setting change information input via the setting change screen 52 to the processing unit 20.
[0107] Returning to Figure 1, we continue the explanation.
[0108] The fourth reception unit 22D receives input of setting change information from the operation input unit 16A. That is, the fourth reception unit 22D receives input of setting change information of at least one of the acoustic features and the synthesis method of the speech of the post-change character string 84B.
[0109] When the fourth receiving unit 22D receives input of setting change information, the generation control unit 24 adjusts the acoustic features of the audio data for the audio section of the post-change character string 84B to the acoustic features included in the received input setting change information. Then, the generation control unit 24 generates processed audio data 76 by synthesizing the audio data for the audio section of the post-change character string 84B, whose acoustic features have been adjusted, into the change target audio section 72A in the target recorded audio data 72 according to the synthesis method included in the setting change information.
[0110] For example, assume that the user operates the setting change button 60H in response to an instruction from the operation input unit 16A at the stage when the acquisition unit 25 acquires the recorded instruction voice data 74. Then, assume that the user operates the operation input unit 16A in response to an instruction from the operation input unit 16A to input setting change information via the setting change screen 52.
[0111] In this case, the generation control unit 24 adjusts the acoustic features of the audio represented by the changed recorded audio data 74B identified from the teaching recorded audio data 74 to the acoustic features included in the setting change information. For example, the generation control unit 24 adjusts the pitch and gain of the audio represented by the changed recorded audio data 74B identified from the teaching recorded audio data 74 to the pitch and gain of the audio included in the setting change information. Then, the generation control unit 24 synthesizes the changed recorded audio data 74B, whose acoustic features have been adjusted, with the target recorded audio data 72 in accordance with the synthesis method included in the setting change information, thereby generating processed audio data 76.
[0112] At this stage, it is assumed that the user operates the play button 60F or the save button 60G in response to an operation instruction from the operation input unit 16A. That is, it is assumed that the save button 60G is operated at the stage when the processed audio data 76 in which at least one of the acoustic features and the synthesis method has been adjusted in accordance with the setting change information has been generated.
[0113] In this case, the playback control unit 26 and the storage processing unit 27 may each perform the same processing as described above.
[0114] Specifically, it is assumed that the user operates the operation input unit 16A to instruct the playback button 60F, and in this case, the playback control unit 26 plays back the processed audio data 76 in which at least one of the acoustic features and the synthesis method has been adjusted in accordance with the setting change information.
[0115] Also, assume that in this situation, the user operates the operation input unit 16A to instruct operation of the save button 60G. In this case, the storage processing unit 27 stores the processed voice data 76 generated by the generation control unit 24 in the storage unit 12. The storage processing unit 27 also stores processing-related information, including the target recorded voice data 72 or identification information of the target recorded voice data 72, the basic character string 82, the change-target character string 82A, the changed character string 84B, the training recorded voice data 74, and setting change information, in association with the processed voice data 76 in the storage unit 12.
[0116] Next, it is assumed that the user operates the detail edit button 60D in response to an operation instruction from the operation input unit 16A.
[0117] The detailed edit button 60D is an input button that the user operates to instruct detailed editing of at least one of the acoustic features and the synthesis method of the speech of the post-change character string 84B. In other words, the detailed edit button 60D is an input button that the user operates to instruct more detailed editing than the setting change button 60H.
[0118] When the user operates the detail edit button 60D in response to an operation instruction from the operation input unit 16A, the display control unit 21 displays a detail edit screen on the display unit 14A.
[0119] 2H is a schematic diagram of an example of the detailed editing screen 54. The detailed editing screen 54 includes a detailed editing input field 64A and a detailed editing reflect button 64B. The detailed editing input field 64A is an input field for accepting input of detailed editing information for at least one of the acoustic features and synthesis method of the speech of the post-change character string 84B. For example, the detailed editing input field 64A displays a screen on which the acoustic features such as prosody and the synthesis method such as crossfade points can be set in detail.
[0120] The user operates the operation input unit 16A while viewing the detailed editing screen 54 to input detailed editing information of at least one of the acoustic features and synthesis method of the speech of the post-change character string 84B. The detailed editing information is information that represents at least one of the detailed acoustic features and synthesis method of the post-change character string 84B.
[0121] Returning to Figure 1, we continue the explanation.
[0122] The fifth receiving unit 22E receives input of detailed editing information of at least one of the acoustic feature amount and the synthesis method of the speech of the post-change character string 84B.
[0123] When the fifth receiving unit 22E receives input of detailed editing information, the generation control unit 24 adjusts the acoustic features of the audio data of the post-change character string 84B to the acoustic features included in the received input of detailed editing information. Then, the generation control unit 24 generates processed audio data 76 by synthesizing the audio data of the audio section of the post-change character string 84B, whose acoustic features have been adjusted, into the audio section 72A to be changed in the target recorded audio data 72 according to the synthesis method included in the detailed editing information.
[0124] For example, suppose that the user operates the detail editing button 60D in response to an instruction from the operation input unit 16A at the stage when the acquisition unit 25 acquires the recorded instruction voice data 74. Then, suppose that the user operates the operation input unit 16A in response to an instruction from the operation input unit 16A to input detail editing information via the detail editing screen 54.
[0125] In this case, the generation control unit 24 adjusts the acoustic features of the speech represented by the post-changed recorded speech data 74B identified from the teaching recorded speech data 74 to the acoustic features included in the detailed editing information. For example, the generation control unit 24 adjusts the acoustic features, such as prosody and accent, of the speech represented by the post-changed recorded speech data 74B identified from the teaching recorded speech data 74 to the acoustic features, such as prosody and accent, included in the detailed editing information. Then, the generation control unit 24 generates the processed speech data 76 by synthesizing the post-changed recorded speech data 74B with the adjusted acoustic features according to the synthesis method included in the detailed editing information.
[0126] At this stage, it is assumed that the user operates the play button 60F or the save button 60G in response to an instruction from the operation input unit 16A. That is, it is assumed that the save button 60G is operated at the stage when the processed audio data 76 in which at least one of the acoustic features and the synthesis method has been adjusted according to the detailed editing information has been generated.
[0127] In this case, the playback control unit 26 and the storage processing unit 27 may each perform the same processing as described above.
[0128] Specifically, it is assumed that the user operates the operation input unit 16A to instruct the playback button 60F, and in this case, the playback control unit 26 plays back the processed audio data 76 in which at least one of the acoustic features and the synthesis method has been adjusted in accordance with the detailed editing information.
[0129] Also, assume that in this situation, the user operates the operation input unit 16A to instruct operation of the save button 60G. In this case, the storage processing unit 27 stores the processed voice data 76 generated by the generation control unit 24 in the storage unit 12. The storage processing unit 27 also stores processing-related information including the target recorded voice data 72 or identification information of the target recorded voice data 72, the basic character string 82, the character string to be changed 82A, the changed character string 84B, the teaching recorded voice data 74, and detailed editing information in association with the processed voice data in the storage unit 12.
[0130] Returning to Figure 1, we continue the explanation.
[0131] Next, the information processing device 30 will be described.
[0132] The information processing device 30 processes the target recorded voice data 72 by using the processing-related information generated by the voice processing support device 10.
[0133] The information processing device 30 includes a storage unit 32, an output unit 34, an input unit 36, a communication unit 38, and a processing unit 40. The storage unit 32, the output unit 34, the input unit 36, the communication unit 38, and the processing unit 40 are connected via a bus 39 so as to be able to communicate with each other.
[0134] The storage unit 32 stores various types of data. The output unit 14 is an output device for outputting various types of information. In this embodiment, the output unit 14 includes a display unit and a speaker. The display unit and the speaker are similar to the display unit 14A and the speaker 14B of the voice processing support device 10.
[0135] The input unit 36 is an input device for receiving various instructions from the user, such as a pointing device such as a digital pen, a mouse, or a trackball, a keyboard, a microphone, or the like.
[0136] The communication unit 38 communicates with an external information processing device via the network NW. In this embodiment, the communication unit 38 communicates with the voice processing support device 10 via the network NW.
[0137] The processing unit 40 executes various types of information processing. The processing unit 40 includes a receiving unit 41 and a processing unit 42. The receiving unit 41 and the processing unit 42 are realized by, for example, one or more processors.
[0138] The reception unit 41 receives the processing-related information from the voice processing support device 10. For example, the reception unit 41 receives the processing-related information from the voice processing support device 10 via the communication unit 38, thereby receiving the processing-related information. Furthermore, for example, the reception unit 41 may store the processing-related information generated by the voice processing support device 10 in the storage unit 32 via a portable storage medium such as a USB memory, and read the processing-related information from the storage unit 32, thereby receiving the processing-related information.
[0139] As described above, the processing-related information is information relating to the processing of the processed voice data 76 .
[0140] The processing unit 42 processes the target recorded voice data 72 based on the processing-related information received by the receiving unit 41 to generate processed voice data.
[0141] For example, the processing unit 42 identifies the target recorded voice data 72 included in the processing-related information. If the processing-related information includes identification information of the target recorded voice data 72, the processing unit 42 identifies the target recorded voice data 72 identified by the identification information from the storage unit 32 or the like.
[0142] The processing unit 42 then processes the identified target recorded voice data 72 in accordance with the processing-related information to generate processed voice data. The processing unit 42 processes the identified target recorded voice data 72 in accordance with the processing-related information in the same manner as the generation control unit 24, to generate processed voice data.
[0143] For example, consider a situation in which the processing-related information received by the reception unit 41 includes the target recorded voice data 72, the basic character string 82, the character string to be changed 82A, the changed character string 84B, the instruction recorded voice data 74, and detailed correction information.
[0144] In this case, for example, the processing unit 42 replaces the character string to be changed 82A included in the basic character string 82 of the target recorded voice data 72 included in the processing-related information with the changed character string 84B. Then, the processing unit 42 adjusts the acoustic features of the changed recorded voice data 74B included in the teaching recorded voice data 74 to the acoustic features included in the detailed editing information. Then, the processing unit 42 generates the processed voice data 76 by synthesizing the changed recorded voice data 74B whose acoustic features have been adjusted according to the synthesis method included in the processing-related information.
[0145] Furthermore, the receiving unit 41 may receive, from the input unit 36, information for changing at least a part of the received processing-related information. The user inputs an instruction to change a part of the information included in the processing-related information by operating the input unit 36. In this case, the processing unit 42 processes the target recorded audio data 72 using the changed processing-related information to generate processed audio data.
[0146] For example, it is assumed that the user operates the input unit 36 to change the target character string 82A and the post-change character string 84B included in the processing-related information.
[0147] In this case, the receiving unit 41 receives the changed target character string 82A and the changed post-change character string 84B as change information. The processing unit 42 replaces the target character string 82A represented by the change information with the post-change character string 84B represented by the change information, among the basic character strings 82 of the target recorded voice data 72 included in the processing-related information.
[0148] The processing unit 42 then adjusts the acoustic features of the changed recorded audio data 74B included in the teaching recorded audio data 74 to the acoustic features included in the detailed editing information. The processing unit 42 then generates the processed audio data 76 by synthesizing the changed recorded audio data 74B, whose acoustic features have been adjusted, in accordance with the synthesis method included in the processing-related information.
[0149] As described above, the information processing device 30 of this embodiment processes the target recorded voice data 72 using the processing-related information created by the voice processing support device 10. Therefore, the information processing device 30 can easily process the target recorded voice data 72. Furthermore, the information processing device 30 of this embodiment receives, from the input unit 36, information for changing at least a part of the received processing-related information, and processes the target recorded voice data 72 using the changed processing-related information. Therefore, the information processing device 30 of this embodiment can support easy adjustment by the user regarding the processing of recorded voice data such as the target recorded voice data 72.
[0150] Next, the information processing executed by the voice processing support system 1 of this embodiment will be described.
[0151] FIG. 3 is a flowchart showing an example of the flow of information processing executed by the voice processing support device 10 of this embodiment.
[0152] The display control unit 21 of the voice processing support device 10 displays the processing support screen 50 on the display unit 14A (step S100). By the processing of step S100, for example, the processing support screen 50 shown in Fig. 2A is displayed on the display unit 14A.
[0153] The first reception unit 22A receives a selection of target recorded voice data 72 from the basic recorded voice data 70 via the processing support screen 50 displayed in step S100 (step S102). The user operates the operation input unit 16A to select one basic recorded voice data 70 to be processed from the one or more basic recorded voice data 70 displayed. The first reception unit 22A receives the one basic recorded voice data 70 selected via the selection field 60A as the target recorded voice data 72.
[0154] The conversion unit 23 converts the target recorded voice data 72 selected in step S102 into a basic character string 82 (step S104). The display control unit 21 displays the basic character string 82 converted in step S104 on the display unit 14A (step S106). By the processing of step S106, for example, the basic character string 82 is displayed in the display field 60B of the editing support screen 50, as shown in FIG. 2B.
[0155] The second receiving unit 22B receives the designation of the character string 82A to be changed from among the displayed basic character strings 82 (step S108). By the processing of step S108, for example, as shown in Fig. 2C, from the basic character string 82 "I have an important friend named Masao" displayed in the display field 60B, "Masao" is designated as the character string to be changed 82A.
[0156] Next, third reception unit 22C receives input of post-change character string 84B for change-target character string 82A received in step S108 (step S110). As shown in Fig. 2D, the user operates operation input unit 16A to input post-change character string 84B "Takumi" in place of change-target character string 82A "Masao". When post-change character string 84B is input, third reception unit 22C receives input of post-change character string 84B.
[0157] When the input of the post-change character string 84B is accepted, the generation control unit 24 generates processed audio data 76 corresponding to the target recorded audio data 72 and the post-change character string 84B for the target character string 82A (step 112). In step S122, the generation control unit 24 generates processed audio data 76 by combining the post-change recorded audio data 74B of the post-change character string 84B with the target audio section 72A corresponding to the target character string 82A in the target recorded audio data 72.
[0158] Next, the reception unit 22 determines whether or not a playback instruction has been received (step S114). The reception unit 22 makes the determination in step S114 by determining whether or not the playback button 60F has been operated in response to an operation instruction from the operation input unit 16A by the user, and whether or not a playback instruction has been received from the operation input unit 16A. If the determination in step S114 is negative (step S114: No), the process proceeds to step S118, which will be described later.
[0159] If the playback instruction is accepted (step S114: Yes), the playback control unit 26 executes a playback process to play back the processed audio data 76 that was generated immediately before by the generation control unit 24 (step S116).
[0160] Next, the receiving unit 22 determines whether or not a save instruction has been received (step S118). If a save instruction has been received (step S118: Yes), the process proceeds to step S120.
[0161] In step S120, the storage processing unit 27 executes storage processing (step S120). The storage processing unit 27 stores the most recently generated processed voice data 76 in the storage unit 12. The storage processing unit 27 also stores processing-related information related to the processing of the processed voice data 76 in association with the processed voice data 76 in the storage unit 12. Then, this routine ends.
[0162] On the other hand, if the determination in step S118 is negative (step S118: No), the process proceeds to step S122.
[0163] In step S122, the receiving unit 22 determines whether or not a simple teaching instruction has been received (step S122). The receiving unit 22 makes the determination in step S122 by determining whether or not the simple teaching button 60E has been operated by a user's operation instruction from the operation input unit 16A and a simple teaching instruction has been received from the simple teaching button 60E.
[0164] If it is determined that a simple instruction instruction has been received (step S122: Yes), the process proceeds to step S124. In step S124, the acquiring unit 25 acquires, as the recorded teaching voice data 74, the voice of the user speaking about the output target string 86 obtained by converting the change target string 82A included in the basic string 82 into the changed string 84B (step S124).
[0165] The generation control unit 24 generates processed audio data 76 based on the target recorded audio data 72, the post-change character string 84B for the change target character string 82A, and the teaching recorded audio data 74 acquired in step S124 (step S126). As shown in FIG. 2F, for example, the generation control unit 24 identifies post-change recorded audio data 74B of the audio section corresponding to the post-change character string 84B in the teaching recorded audio data 74. The generation control unit 24 also identifies the post-change audio section 72A in the target recorded audio data 72 that corresponds to the change target character string 82A. The generation control unit 24 then generates the processed audio data 76 by combining the post-change recorded audio data 74B identified from the teaching recorded audio data 74 with the post-change audio section 72A in the target recorded audio data 72. The process then proceeds to step S114.
[0166] If the determination in step S122 is negative (step S122: No), the process proceeds to step S128. In step S128, the reception unit 22 determines whether or not a setting change instruction has been received (step S128). The reception unit 22 makes the determination in step S128 by determining whether or not the setting change button 60H has been operated by a user's operation instruction from the operation input unit 16A, and whether or not a setting change instruction has been received from the setting change button 60H.
[0167] If it is determined that a setting change instruction has been received (step S128: Yes), the process proceeds to step S130. In step S130, the display control unit 21 displays a setting change screen 52 on the display unit 14A (step S130). By the processing of step S130, for example, the setting change screen 52 shown in FIG. 2G is displayed. The user operates the operation input unit 16A while visually checking the setting change screen 52 to input setting change information for at least one of the acoustic features and the synthesis method of the speech of the changed character string 84B. When the user operates the setting change reflect button 62B in response to an operation instruction from the operation input unit 16A, the fourth reception unit 22D receives the input of the setting change information (step S132).
[0168] The generation control unit 24 generates processed audio data 76 according to the setting change information received in step S132 (step S134). The generation control unit 24 adjusts the acoustic features of the audio data for the audio section of the post-change character string 84B to the acoustic features included in the setting change information received in step S132. The generation control unit 24 then generates processed audio data 76 by synthesizing the audio data for the audio section of the post-change character string 84B, whose acoustic features have been adjusted, with the change target audio section 72A in the target recorded audio data 72 in accordance with the synthesis method included in the setting change information received in step S132. The process then proceeds to step S114.
[0169] If the determination in step S128 is negative (step S128: No), the process proceeds to step S136.
[0170] In step S136, the accepting unit 22 determines whether or not a detailed editing instruction has been accepted (step S136). The accepting unit 22 makes the judgment in step S136 by determining whether or not the detailed editing button 60D has been operated in response to an operation instruction from the operation input unit 16A by the user, and whether or not a detailed editing instruction has been accepted from the detailed editing button 60D. If the judgment in step S136 is negative (step S136: No), the process proceeds to step S114. Note that if the judgment in step S136 is negative, the process may proceed to step S102.
[0171] If it is determined that a detail editing instruction has been received (step S136: Yes), the process proceeds to step S138. In step S138, the display control unit 21 displays the detail editing screen 54 on the display unit 14A (step S138). By the process of step S140, for example, the detail editing screen 54 shown in FIG. 2H is displayed on the display unit 14A.
[0172] The user operates the operation input unit 16A while viewing the detailed editing screen 54 to input detailed editing information of at least one of the acoustic features and the synthesis method of the speech of the post-change character string 84B. When the user operates the detailed edit reflect button 64B by operating the operation input unit 16A, the fifth accepting unit 22E accepts the input of the detailed editing information (step S140).
[0173] The generation control unit 24 adjusts the acoustic features of the audio data of the post-change character string 84B to the acoustic features included in the received detailed editing information. The generation control unit 24 then generates processed audio data 76 by synthesizing the audio data of the audio section of the post-change character string 84B, whose acoustic features have been adjusted, into the audio section 72A to be changed in the target recorded audio data 72 according to the synthesis method included in the detailed editing information (step S142). The process then proceeds to step 114.
[0174] Next, an example of the flow of information processing executed by the information processing device 30 of this embodiment will be described.
[0175] FIG. 4 is a flowchart showing an example of the flow of information processing executed by information processing device 30 according to this embodiment.
[0176] The receiving unit 41 receives processing-related information from the voice processing support device 10 (step S200).
[0177] The receiving unit 41 also receives information on at least part of the modification of the processing-related information received in step S200 (step S202).
[0178] The processing unit 42 generates processed voice data using the processing-related information received in step S200 and the change information received in step S202 (step S204). Then, the processing unit 42 outputs the processed voice data generated in step S204 to the output unit 34 (step S206). For example, the processing unit 42 outputs the voice of the processed voice data generated in step S204 from a microphone included in the output unit 34. The processing unit 42 may store the processed voice data generated in step S204 in the storage unit 32. Furthermore, the processing unit 42 may transmit the processed voice data generated in step S204 to another external information processing device via the communication unit 38. Then, this routine ends.
[0179] As described above, the voice processing support device 10 of this embodiment includes a first receiving unit 22A, a display control unit 21, and a second receiving unit 22B. The first receiving unit 22A receives selection of target recorded voice data 72, which is the basic recorded voice data 70 to be processed, from one or more recorded basic recorded voice data 70. The display control unit 21 converts the target recorded voice data 72 into a basic character string 82 and displays it. The second receiving unit 22B receives designation of a target character string 82A to be changed from the displayed basic character string 82. The generation control unit 24 generates processed voice data 76 corresponding to the target recorded voice data 72 and the target character string 82A.
[0180] In the conventional technology, processed audio data is generated by processing audio data through automatic extraction and synthesis without user operation instructions, making it difficult for the conventional technology to support easy user adjustments regarding the processing of recorded audio data.
[0181] Meanwhile, the voice processing support device 10 of this embodiment accepts a user's selection of target recorded voice data 72, which is the basic recorded voice data 70 to be processed, from a plurality of basic recorded voice data 70. Therefore, the user can select a desired basic recorded voice data 70 from the plurality of basic recorded voice data 70 as the target recorded voice data 72. Furthermore, the voice processing support device 10 of this embodiment displays a basic character string 82 of the target recorded voice data 72 and accepts a designation of a target character string 82A to be changed in the basic character string 82. Therefore, the user can designate a desired character string from the basic character string 82 of the target recorded voice data 72 as the target character string 82A. Then, the generation control unit 24 generates processed voice data 76 corresponding to the target recorded voice data 72 and the target character string 82A. Therefore, the generation control unit 24 can generate processed voice data 76 corresponding to the user's selection and designation.
[0182] That is, in the voice processing support device 10 of this embodiment, the user can select the desired basic recorded voice data 70 as the target recorded voice data 72 and specify the desired character string among the basic character strings 82 of the target recorded voice data 72 as the character string to be changed 82A.
[0183] Therefore, the voice processing support device 10 of this embodiment can support the user in easily adjusting the processing of recorded voice data such as the basic recorded voice data 70.
[0184] Furthermore, the voice processing support device 10 of this embodiment generates processed voice data 76 according to the target recorded voice data 72 selected from the basic recorded voice data 70 and the change target character string 82A.
[0185] Therefore, the voice processing support device 10 of this embodiment can easily generate processed voice data 76 of new lines with the same or similar voice quality as the basic recorded voice data 70, even if it is difficult to obtain a speech with the same voice quality as the basic recorded voice data 70 of the recorded speech.
[0186] The voice processing support device 10 of this embodiment may further include a third receiving unit 22C. The third receiving unit 22C receives an input of a post-change character string 84B for the change target character string 82A. The generation control unit 24 generates processed voice data 76 according to the target recorded voice data 72 and the post-change character string 84B for the change target character string 82A.
[0187] In this way, the voice processing support device 10 of this embodiment generates processed voice data 76 corresponding to the target recorded voice data 72 and the changed character string 84B. Therefore, in addition to the above-mentioned effects, the voice processing support device 10 of this embodiment can easily generate processed voice data 76 based on the pre-recorded high-quality and high-quality target recorded voice data 72 that reflects the intentions of the performer, etc.
[0188] However, in the prior art, even if there is an abundance of recorded audio, if unrecorded lines or some words need to be changed, it is necessary to re-record them.
[0189] On the other hand, the voice processing support device 10 of this embodiment generates processed voice data 76 in which the voice section 72A to be changed of the character string 82A to be changed contained in the target recorded voice data 72 selected from the plurality of basic recorded voice data 70 is replaced with the voice data of the changed character string 84B input by the user.
[0190] Therefore, in the voice processing support device 10 of this embodiment, even if a plurality of basic recorded voice data 70 has already been stored, there is no need to re-record to obtain voice data with some wording changed. Therefore, in addition to the above-mentioned effects, the voice processing support device 10 of this embodiment can reduce the burden on the user.
[0191] Furthermore, in the prior art, even when voice data of a voice spoken by the same user is synthesized, processed voice data that sounds unnatural may be generated due to changes in the physical condition of the user who speaks the lines or changes in the vocal cords, etc. Adjusting the processed voice data that sounds unnatural may require time and cost.
[0192] Furthermore, with conventional technologies, creating high-quality speech synthesis data that matches the actor's intentions requires recording a large amount of training speech data, conducting machine learning, and verifying the data. As a result, conventional technologies have been unable to meet the demand for creating performance voices for many characters in a short period of time and on a low budget. Furthermore, adjusting processed speech data by processing only the synthesized speech can sometimes be perceived as taking away work from users whose job it is to speak.
[0193] On the other hand, the voice processing support device 10 of this embodiment uses target recorded voice data 72 selected from multiple basic recorded voice data 70 as a reference, and generates processed voice data 76 in which the change target voice section 72A of the change target string 82A, which is part of the target recorded voice data 72, is replaced with voice data of the changed string 84B input by the user.
[0194] Therefore, in addition to the above-mentioned effects, the voice processing support device 10 of this embodiment can easily generate processed voice data 76 in a short period of time and with a low budget using the target recorded voice data 72 of the user's voice.
[0195] Next, the hardware configuration of the voice processing support device 10 and the information processing device 30 of this embodiment will be described.
[0196] FIG. 5 is a diagram showing an example of the hardware configuration of the voice processing support device 10 and the information processing device 30 according to the present embodiment.
[0197] The voice processing support device 10 and the information processing device 30 of this embodiment include a control device such as a CPU 10A, a storage device such as a ROM (Read Only Memory) 10B and a RAM (Random Access Memory) 10C, an HDD (Hard Disk Drive) 10D, an I / F 10E that connects to a network and communicates, and a bus 10F that connects each part.
[0198] The programs executed by the voice processing support device 10 and the information processing device 30 of this embodiment are provided in a state that they are pre-installed in the ROM 10B or the like.
[0199] The programs executed by the voice processing support device 10 and the information processing device 30 of this embodiment may be configured to be provided as a computer program product by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk).
[0200] Furthermore, the programs executed by the voice processing assistance device 10 and the information processing device 30 according to the present embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the programs executed by the voice processing assistance device 10 and the information processing device 30 according to the present embodiment may be provided or distributed via a network such as the Internet.
[0201] The programs executed by the voice processing support device 10 and the information processing device 30 of this embodiment can cause a computer to function as each unit of the above-mentioned voice processing support device 10. In this computer, the CPU 10A can read the programs from a computer-readable storage medium onto a main storage device and execute the programs.
[0202] In the above embodiment, the voice processing assistance device 10 and the information processing device 30 are described assuming that they are configured as a single device. However, the voice processing assistance device 10 and the information processing device 30 may be configured as a plurality of devices that are physically separated and connected to each other so as to be able to communicate with each other via a network or the like.
[0203] Although the embodiments of the present invention have been described above, the above embodiments are presented as examples and are not intended to limit the scope of the invention. This novel embodiment can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. This embodiment and its modifications are included within the scope and spirit of the invention, and are also included in the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0204] 1. Voice processing support system 10. Voice processing support device 21 Display control unit 22A First Reception Section 22B Second Reception Section 22C Third Reception Section 22D 4th Reception Section 22E 5th Reception Section 22F 6th Reception Desk 24 Generation control section 25 Acquisition Department 26 Playback control unit 30 Information processing equipment 41 Reception 42 Processing section
Claims
1. A voice processing support system comprising a voice processing support device and an information processing device, The voice processing support device includes: a first receiving unit that receives a selection of target recorded voice data, which is the basic recorded voice data to be processed, from one or more recorded basic recorded voice data; a display control unit that converts the target recorded voice data into a basic character string and displays it; a second receiving unit that receives a designation of a character string to be changed from among the displayed basic character strings; a generation control unit that generates processed voice data according to the target recorded voice data and the target character string to be changed; a storage control unit that stores processing-related information related to processing of the processed voice data; Equipped with The information processing device includes: a receiving unit that receives the processing-related information; a processing unit that processes the target recorded voice data based on the processing-related information to generate processed voice data; Equipped with Voice processing support system.
2. The voice processing support device comprises: a third receiving unit that receives an input of a post-change character string for the change target character string; The generation control unit generating the processed voice data according to the target recorded voice data and the post-change character string for the change target character string; The voice processing support system according to claim 1 .
3. The generation control unit generating the processed voice data by synthesizing the post-change character string voice data of the post-change character string with the voice section to be changed corresponding to the post-change character string in the target recorded voice data; The voice processing support system according to claim 2 .
4. The voice processing support device comprises: an acquisition unit that acquires, as recorded voice data for training, a speech uttered by a user of an output target string obtained by converting the target string included in the basic string into the changed string; Equipped with The generation control unit generating the processed voice data by combining the post-change recorded voice data corresponding to the post-change character string in the training recorded voice data with the voice section to be changed corresponding to the character string to be changed in the subject recorded voice data; The voice processing support system according to claim 2 .
5. The generation control unit adjusting the pitch of the voice represented by the changed recorded voice data to the pitch of the voice of the change target voice section in the target recorded voice data; projecting the prosody of the speech section to be changed in the target recorded speech data onto the post-change recorded speech data whose pitch has been converted; generating the processed speech data by synthesizing the post-change recorded speech data onto which the prosody has been projected, with the speech section to be changed in the target recorded speech data; The voice processing support system according to claim 4.
6. The voice processing support device comprises: a fourth receiving unit that receives input of setting change information for at least one of an acoustic feature of the speech of the changed character string and a synthesis method; Equipped with The generation control unit adjusting the acoustic feature of the speech data of the speech section of the changed character string to the acoustic feature included in the setting change information; generating the processed voice data by synthesizing the voice data, the acoustic features of which have been adjusted, into the voice section to be changed that corresponds to the character string to be changed in the target recorded voice data in accordance with the synthesis method included in the setting change information; The voice processing support system according to claim 2 .
7. The voice processing support device comprises: a fifth receiving unit that receives input of detailed editing information of at least one of acoustic features of the speech of the changed character string and a synthesis method; Equipped with The generation control unit adjusting the acoustic feature of the speech data of the speech section of the changed character string to the acoustic feature included in the detailed editing information; generating the processed voice data by synthesizing the voice data, the acoustic features of which have been adjusted, into the voice section to be changed corresponding to the character string to be changed in the target recorded voice data in accordance with the synthesis method included in the detailed editing information; The voice processing support system according to claim 2 .
8. The voice processing support device comprises: a sixth reception unit that receives an instruction to reproduce the processed audio data; a playback control unit that plays back the processed audio data; The voice processing support system according to claim 1 , comprising:
9. a step of receiving a selection of target recorded voice data, which is the basic recorded voice data to be processed, from one or more recorded basic recorded voice data; converting the target recorded voice data into a basic character string and displaying it; receiving a designation of a character string to be changed from among the displayed basic character strings; generating processed voice data according to the target recorded voice data and the target character string to be changed; storing processing-related information relating to processing of the processed voice data; receiving the processing-related information; generating processed voice data by processing the target recorded voice data based on the processing-related information; A voice processing support method including:
10. a step of receiving a selection of target recorded voice data, which is the basic recorded voice data to be processed, from one or more recorded basic recorded voice data; converting the target recorded voice data into a basic character string and displaying it; receiving a designation of a character string to be changed from among the displayed basic character strings; generating processed voice data according to the target recorded voice data and the target character string to be changed; storing processing-related information relating to processing of the processed voice data; receiving the processing-related information; generating processed voice data by processing the target recorded voice data based on the processing-related information; A voice processing support program that allows a computer to execute the above.
Citation Information
Patent Citations
Speech synthesis system for connecting sound-recorded speech and synthesized speech together
JP2003295880A
Voice synthesis apparatus
JP2008107454A
Voice synthesis device and voice synthesis method, and program
JP2009020264A
Voice editing composite system, voice editing composite program, and voice editing composite method
JP2009157220A
Voice data editing device
JP2011242637A