Speech synthesis device, speech conversion device, program, information processing method
The speech synthesis device allows for dynamic adjustment of speech style and characteristics by using a content encoding unit, style encoding unit, and speech decoding unit, addressing the limitation of fixed speech styles in existing models.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LY CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing speech synthesis models lack the ability for users to adjust the speech style or features of generated speech after synthesis.
A speech synthesis device and method that includes a content encoding unit, style encoding unit, and speech decoding unit, allowing for the calculation and synthesis of speech features based on user-provided prompts, enabling adjustment of speech style and characteristics.
Enables dynamic adjustment of speech style and characteristics, allowing users to modify synthesized speech to better suit their preferences.
Smart Images

Figure 2026073841000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a speech synthesis device, a speech conversion device, a program, an information processing method, etc. [Background technology]
[0002] With the rapid development and spread of generative AI based on diffusion models (DMs), high-quality speech synthesis models have been proposed. For example, Non-Patent Documents 1 and 2 disclose speech synthesis models that can flexibly control the style and characteristics of speech based on text-based prompts. However, in the methods disclosed in Non-Patent Documents 1 and 2, users of the speech synthesis model cannot adjust the speech style or features of the generated speech. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Zhifang Guo, et al. “PromptTTS: Controllable Text-to-Speech with Text Descriptions”, 2022, https: / / arxiv.org / abs / 2211.12171. [Non-Patent Document 2] Reo Shimizu, et al. “PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions”, 2023, https: / / arxiv.org / abs / 2309.08140. [Overview of the project] [Problems that the invention aims to solve]
[0004] The present invention was made against the technical background described above, and aims to provide a speech synthesis device, etc., that can synthesize a second speech by receiving modification instructions regarding the speech style and characteristics of the first speech after the first speech has been synthesized by a speech synthesis model. [Means for solving the problem]
[0005] According to a first aspect of the present invention, a speech synthesis device is provided, the control unit of the speech synthesis device comprising a content encoding unit, a style encoding unit, and a speech decoding unit, wherein the control unit acquires a content prompt and a first style prompt, the content encoding unit calculates content features based on the content prompt, the style encoding unit calculates first style features based on the first style prompt, the speech decoding unit synthesizes a first speech based on the content features and the first style features, the control unit acquires a second style prompt for the first speech, the style encoding unit calculates second style features based on the second style prompt, and the speech decoding unit synthesizes a second speech based on the content features and the internal division point between the first style features and the second style features. According to a second aspect of the present invention, a program executed by a speech synthesizer comprises a control unit of the speech synthesizer comprising a content encoding unit, a style encoding unit, and a speech decoding unit, and includes: acquiring a content prompt and a first style prompt by the control unit; calculating content features based on the content prompt by the content encoding unit; calculating first style features based on the first style prompt by the style encoding unit; synthesizing a first speech by the speech decoding unit based on the content features and the first style features; acquiring a second style prompt for the first speech by the control unit; calculating second style features based on the second style prompt by the style encoding unit; and synthesizing a second speech by the speech decoding unit based on the content features and the internal division point between the first style features and the second style features. According to a third aspect of the present invention, an information processing method performed by a speech synthesizer is provided, the control unit of the speech synthesizer comprising a content encoding unit, a style encoding unit, and a speech decoding unit, and includes: acquiring a content prompt and a first style prompt by the control unit; calculating content features based on the content prompt by the content encoding unit; calculating first style features based on the first style prompt by the style encoding unit; synthesizing a first speech by the speech decoding unit based on the content features and the first style features; acquiring a second style prompt for the first speech by the control unit; calculating second style features based on the second style prompt by the style encoding unit; and synthesizing a second speech by the speech decoding unit based on the content features and the internal division point between the first style features and the second style features. According to a fourth aspect of the present invention, a speech conversion device is provided, the control unit of the speech conversion device comprising a speech recognition unit, a content encoding unit, a style encoding unit, a reference encoding unit, and a speech decoding unit, wherein the control unit acquires a first speech and a first style prompt, the speech recognition unit infers a content prompt based on the first speech, the reference encoding unit calculates a first style feature based on the first speech, the content encoding unit calculates a content feature based on the content prompt, the style encoding unit calculates a second style feature based on the first style prompt, and the speech decoding unit synthesizes a second speech based on the content feature and the internal division point between the first style feature and the second style feature. According to a fifth aspect of the present invention, a program executed by a speech conversion device, wherein the control unit of the speech conversion device comprises a speech recognition unit, a content encoding unit, a style encoding unit, a reference encoding unit, and a speech decoding unit, and includes acquiring a first speech and a first style prompt by the control unit, inferring a content prompt based on the first speech by the speech recognition unit, calculating a first style feature quantity based on the first speech by the reference encoding unit, calculating a content feature quantity based on the content prompt by the content encoding unit, calculating a second style feature quantity based on the first style prompt by the style encoding unit, and synthesizing a second speech by the speech decoding unit based on the content feature quantity and the internal division point between the first style feature quantity and the second style feature quantity. According to a sixth aspect of the present invention, an information processing method performed by a speech conversion device, wherein the control unit of the speech conversion device comprises a speech recognition unit, a content encoding unit, a style encoding unit, a reference encoding unit, and a speech decoding unit, and includes: acquiring a first speech and a first style prompt by the control unit; inferring a content prompt based on the first speech by the speech recognition unit; calculating a first style feature quantity based on the first speech by the reference encoding unit; calculating a content feature quantity based on the content prompt by the content encoding unit; calculating a second style feature quantity based on the first style prompt by the style encoding unit; and synthesizing a second speech by the speech decoding unit based on the content feature quantity and the internal division point between the first style feature quantity and the second style feature quantity. [Brief explanation of the drawing]
[0006] [Figure 1] A diagram showing an example of the configuration of a speech synthesis device according to the first embodiment. [Figure 2] A diagram showing an example of the processing overview of the speech synthesis device according to the first embodiment. [Figure 3] A flowchart showing an example of the processing flow performed by the speech synthesis device according to the first embodiment. [Figure 4] A diagram showing an example of the configuration of a terminal equipped with a speech synthesis device according to the first embodiment. [Figure 5] A diagram showing an example of a screen displayed on the display unit of a terminal according to the first embodiment. [Figure 6] A diagram showing another example of the screen displayed on the display unit of the terminal relating to the first modified example. [Figure 7] A diagram showing another example of the screen displayed on the display unit of the terminal relating to the first modified example. [Figure 8] A diagram showing an example of the configuration of a voice conversion device according to the second embodiment. [Figure 9] A flowchart showing an example of the processing flow performed by the voice conversion device according to the second embodiment. [Figure 10]A diagram showing an example of the configuration of a terminal including the speech synthesis device according to the second embodiment. [Figure 11] A diagram showing an example of a screen displayed on the display unit of the terminal according to the second embodiment. [Figure 12] A diagram showing another example of a screen displayed on the display unit of the terminal according to the second modification. [Figure 13] A diagram showing another example of a screen displayed on the display unit of the terminal according to the second modification.
Modes for Carrying Out the Invention
[0007] <Compliance with Legal Matters> It should be noted that the disclosure described in this specification is premised on compliance with the legal matters of the country of implementation required for the implementation of the present disclosure, such as communication secrecy.
[0008] <Embodiment> In this specification, there are places described as "as an example" for easy understanding, but it should be noted that not only the relevant places but also the entire content of the embodiments described below are not limited to the described content.
[0009] Embodiments for implementing the program and the like according to the present disclosure will be described with reference to the drawings.
[0010] The production of the terminal (the terminal of the invention according to the claims of the present application) according to the claims of the present application includes, for example, a terminal owned (held) by a user, and when a program (for example, an application program) described in this specification is received (or received and stored in the terminal), a state in which the functions of the invention according to the claims of the present application can be realized (a state in which the invention according to the claims of the present application can be executed) is created in this terminal.
[0011] Furthermore, the production of the system of the claimed invention (the system of the claimed invention) may also include the concept that, for example, a state is created in which the function of the claimed system becomes achievable (a state is created in which the claimed invention becomes executable) by receiving a program described in this specification (for example, an application program) transmitted from a server included in the system to a terminal included in the system (or by the received program being stored in the terminal).
[0012] Furthermore, in this specification, a system may, for example, be configured to include multiple devices. Multiple devices may be a combination of devices of the same type, a combination of devices of different types, or a combination of devices of the same type and devices of different types. Furthermore, a system can be thought of as, for example, a system in which multiple devices work together to perform some kind of processing.
[0013] Furthermore, a system involving a client (client device) and a server can be considered, for example, as at least one of the following: (1) Terminals & Servers (2) Server (3) terminal
[0014] (1) is, for example, a system including at least one terminal and at least one server. This example is a client-server system.
[0015] The server, for example, consists of the following devices, and may be a single device or a combination of multiple devices.
[0016] Specifically, a server may be configured with, for example, at least one processor (e.g., CPU: Central Processing Unit, GPU: Graphics Processing Unit, APU: Accelerated Processing Unit, DSP: Digital Signal Processor (e.g., ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array)), a computer device (processor + memory), a control device, an arithmetic unit, a processing unit, etc., and may be configured with multiple identical units of any one device (e.g., CPU + CPU, homogeneous multicore processor, etc.), or with multiple heterogeneous units of any one device (e.g., CPU + DSP, heterogeneous multicore processor, etc.), or with a combination of multiple devices (e.g., processor + computer device, processor + arithmetic unit, multiple devices made heterogeneous, etc.). Note that the processor may be a virtual processor.
[0017] Furthermore, when a server performs some processing, if it is configured as a single device, the processing described in the embodiment will be performed by that single device. If it is configured as having multiple devices, some processing may be performed by one device and other processing may be performed by the other devices. For example, if it is configured as having a processor and an arithmetic unit, the first processing may be performed by the processor and the second processing may be performed by the arithmetic unit. Furthermore, if the system consists of multiple devices, each device may be arranged in a location that is physically separated from the others.
[0018] Furthermore, server functionality may be provided, for example, in the form of PaaS, IaaS, or SaaS in cloud computing. Some or all of the processes described herein may be implemented as programs included in an application installed from the server to the terminal. Furthermore, the manufacturer and manager of the terminal, the manufacturer and manager of the application, and the manufacturer and manager of the server may be different entities (businesses), or some or all of them may be the same entity (business).
[0019] Furthermore, the system's control unit can be at least one of either the terminal's control unit or the server's control unit. In other words, for example, the system's control unit can be any of the following: (1A) only the terminal's control unit, (1B) only the server's control unit, or (1C) both the terminal's control unit and the server's control unit.
[0020] Furthermore, the control and processing performed by the system's control unit (hereinafter collectively referred to as "control, etc.") may be (1A) performed solely by the terminal's control unit, (1B) performed solely by the server's control unit, or (1C) performed by both the terminal's control unit and the server's control unit. Furthermore, in (1C), as an example, some of the controls that the system would normally perform are to be performed by the terminal's control unit, and the remaining controls are to be performed by the server's control unit. In this case, the allocation of controls may be equal, or they may be allocated in different proportions.
[0021] Furthermore, when referring to the server's communication unit, if the server is composed of a single device, it may refer to the communication unit itself provided by that single device. Also, if the server is composed of multiple devices, the server's communication unit may include the communication units provided by each of those devices. For example, if a server comprises a first device and a second device, with the first device having a first communication unit and the second device having a second communication unit, the communication unit of the server may be considered a concept that includes both the first and second communication units.
[0022] (2) can be, for example, a system consisting of multiple servers (hereinafter referred to as the "server system"). In this case, the configuration described above can be applied similarly to the configuration of each server.
[0023] The control and other functions performed by the server system may be performed by (2A) only one server, (2B) only the other servers, or (2C) one server and the other servers. Furthermore, in (2C), as an example, one server may perform some of the controls that the server system would normally perform, while other servers perform the remaining controls. In this case, the allocation of controls may be equal, or they may be allocated in different proportions.
[0024] (3) can be, for example, a system consisting of multiple terminals. This system can be, for example, as follows: A system that assigns server functions to terminals (a distributed system). This can be achieved, for example, using blockchain technology. A system in which terminals communicate wirelessly with each other. This can be achieved, for example, by using short-range wireless communication technologies such as Bluetooth (registered trademark) to communicate in a P2P (peer-to-peer) manner.
[0025] Furthermore, the above applies not only to the control unit but also to each functional unit that can be a component of the system, such as the input / output unit, communication unit, memory unit, and clock unit.
[0026] In the following embodiments, a system including a terminal and a server (a client-server system, for example) is illustrated as an example. Furthermore, it is also possible to apply the server system described in (2) above as the server.
[0027] Furthermore, instead of a system that includes terminals and servers, it is also possible to apply a system that does not include servers, such as the system described in (3) above. In this case, the embodiment can be constructed based on the aforementioned blockchain technology. Specifically, as an example, data stored and managed on the server described in the following embodiment is stored on the blockchain. Then, when a terminal generates a transaction to the blockchain and the transaction is approved on the blockchain, the data stored on the blockchain is updated.
[0028] Furthermore, even when the term "terminal" is used, it is not limited to the meaning of a terminal as a client device in a client-server system. In other words, the term "terminal" can sometimes include the concept of a device that is not part of a client-server relationship.
[0029] Furthermore, the expression "by communication I / F" will be used as appropriate in this specification. For example, this may indicate that the device transmits and receives various information and data via a communication I / F (via a communication unit) based on the control of a control unit (processor, etc.).
[0030] Furthermore, in this specification, when terms such as "concerning" or "related" are used, "B concerning A" or "B related to A" may mean, for example, "B" which has some kind of relationship with "A". Specific examples will be given later.
[0031] Furthermore, in this specification, when a device processes two or more things, such as "transmitting A and B" or "receiving A and B," it may include both cases where "A" and "B" are performed at the same time (hereinafter referred to as "simultaneous") and cases where "A" and "B" are performed at different times (hereinafter referred to as "non-simultaneous"). For example, when referring to the transmission of first and second information, it is acceptable to include both concepts: transmitting the first and second information at the same time, and transmitting the first and second information at different times. Furthermore, taking into consideration lag (time lag), "simultaneous" may include "almost simultaneous."
[0032] Furthermore, even if "A" and "B" are performed at different times, this only requires that the processing targets "A" and "B," and their purposes do not necessarily have to be the same. For example, when transmitting the first piece of information and the second piece of information as described above, it is sufficient to simply transmit the first piece of information and the second piece of information, and this may include cases where the first piece of information and the second piece of information are transmitted for different purposes, as well as cases where they are transmitted for the same purpose.
[0033] Hereinafter, an example of an embodiment for carrying out the present invention will be described with reference to the drawings. In addition, in the descriptions of the drawings, the same reference numeral is used for identical elements, and redundant explanations may be omitted. Furthermore, the components described in this embodiment are merely illustrative and are not intended to limit the scope of the present invention to them.
[0034] <First Example> The first embodiment is an embodiment in which a speech synthesis device (which may also be called an information processing device, speech processing device, or speech generation device) receives, for example, utterance content specified by a user (referred to as "content") and specified content regarding the characteristics of the utterance and the speaker (referred to as "style"), and synthesizes a first speech. Then, it receives a request to modify the style of the first speech and synthesizes a second speech with the modified style.
[0035] The contents described in the first embodiment are equally applicable to any of the other embodiments and any of the other modifications.
[0036] Figure 1 is a block diagram showing an example of the functional configuration of a speech synthesis device 1 according to one embodiment of this model. The speech synthesis device 1 includes, for example, a content encoding unit 110, a style encoding unit 120, and a speech decoder unit 130. These are, for example, functional units (functional blocks) of the control unit (control device) (not shown) of the speech synthesis device 1. The control unit may also be called a processing unit (processing device).
[0037] The content encoding unit 110, for example, receives a content prompt regarding content via the operation unit 310 for inputting information (e.g., a string) to the speech synthesis device 1, and has the function of calculating content features (which may also be called "content vectors") based on the content prompt. The content encoding unit 110 is a functional unit that converts input text data into numerical vectors and extracts linguistic features.
[0038] The style encoding unit 120 has the function of calculating style features (which may also be called "style vectors") based on a style prompt received, for example, via the operation unit 310. The style encoding unit 120 is a functional unit for encoding a style specified by the user and using it as a control parameter for speech synthesis. For example, one could refer to Non-Patent Document 2 and introduce the concept of speaker prompts into style prompts.
[0039] The audio decoder unit 130 has the function of generating synthesized speech based, for example, on content features and style features. The audio decoder unit 130 is composed of, for example, a fusion module, an acoustic decoder, and a vocoder. The fusion module has the function of calculating comprehensive features necessary for speech generation based on features obtained from the content encoding unit 110 and the style encoding unit 120. The acoustic decoder has the ability to generate acoustic features (e.g., Mel spectrograms) based on comprehensive features in the fusion module. A vocoder has the function of converting acoustic features generated by an acoustic decoder into a time-domain waveform and calculating an audio waveform.
[0040] The control unit of the speech synthesis device 1, for example, when a speech waveform is calculated in the vocoder, causes the sound output unit 350 to output the speech waveform as sound.
[0041] The specific configurations of the content encoding unit 110, the style encoding unit 120, and the audio decoder unit 130 can be implemented by referring, for example, to Non-Patent Document 1 and Non-Patent Document 2.
[0042] Figure 2 shows an example of the processing overview of the speech synthesis device 1. For example, let's say "This is a pen." is entered as the content prompt, and "feminine, very cute" is entered as the first style prompt. In this case, the content encoding unit 110 calculates a feature representing the utterance "This is a pen." as a content feature. Furthermore, the style encoding unit 120 calculates a first style feature e to make the utterance conform to the utterance style "feminine, very cute". s This is calculated. Then, in the audio decoder unit 130, the content features and the first style features e s Based on this, the first sound is synthesized, and for example, the first sound is output from the sound output unit 350.
[0043] For example, the user listens to the first audio and inputs a second style prompt, "low pitch," to adjust (correct) the style of the first audio. Then, the style encoding unit 120 generates a second style feature e to make the speech conform to the speech style "low pitch." t This is calculated. Then, the control unit of the speech synthesis device 1 calculates a synthesized style feature amount e that reflects the first style feature amount e s and the second style feature amount e t The synthesized style feature amount e can be defined, for example, as an interpolation point between the feature vector indicated by the first style feature amount e s and the feature vector indicated by the second style feature amount e t
[0044] For example, the synthesized style feature amount e may be calculated by the following formula (1).
Equation
[0045] Note that the synthesis parameter "w" does not have to be set. For example, the synthesized style feature amount e may be the average vector between the feature vector indicated by the first style feature amount e s and the feature vector indicated by the second style feature amount e t
[0046] Also, the synthesis parameter "w" may be determined, for example, based on a modifier included in the second style prompt. For example, when "very low pitch" is input as the second style prompt, based on the modifier "very", the second style feature amount e in which "low pitch" is encoded t The value range "w>0.5" may be set to enhance the influence of . Also, if "slightly low pitch" is entered as the second style prompt, the second style feature e is encoded with "low pitch" based on the modifier "slightly". t The value may be set to a range of "w < 0.5" to minimize its influence. The composite parameter "w" may also be a value that is embedded according to the strength of the modifier, for example.
[0047] Once the synthesized style feature e is calculated, the audio decoder unit 130 synthesizes a second audio based on the content feature and the synthesized style feature e, and the second audio is output from the sound output unit 350, for example. The second audio is an audio that reflects the content of the second style prompt in relation to the style of the first audio.
[0048] Here, we consider the case where the contents of the first style prompt and the second style prompt are listed together as style prompts, and the first style feature e s and the second style feature e t This section describes the differences between calculating a composite style feature e that reflects these factors and the case where such features are calculated. When the contents of the first style prompt and the second style prompt are listed together as a style prompt, if the contents of the first style prompt and the second style prompt contain conflicting elements (for example, "low pitch" and "high pitch"), the conflicting contents will be embedded simultaneously when the prompt is embedded in the style encoding unit 120, which may hinder the calculation of appropriate style features. In contrast, the first style feature e s and the second style feature e t When calculating and separately and then calculating the combined style feature e, even if the content of the first style prompt and the second style prompt contain conflicting elements, the first style feature e s and the second style feature e tThis is calculated appropriately. Then, based on the internal division points in the feature space, the composite style feature e is calculated, and the composite style feature e is calculated as a feature that appropriately reflects the conflicting content of the first style prompt and the second style prompt.
[0049] Figure 3 is a flowchart showing an example of the processing flow performed by the speech synthesis device 1 in this embodiment. The processes described below are merely examples of processes for implementing the method of this disclosure, and are not limited to these. Additionally, you may add other steps to the process described below, or omit (delete) some steps from the process described below.
[0050] First, the control unit of the speech synthesis device 1 executes the content prompt acquisition process (S110). In the content prompt acquisition process, the control unit acquires a content prompt (e.g., a string) based on user input to the operation unit 310, for example.
[0051] Furthermore, the control unit of the speech synthesis device 1 executes the first style prompt acquisition process (S120). In the first style prompt acquisition process, the control unit acquires a first style prompt (e.g., a string) based on user input to the operation unit 310, for example.
[0052] Then, the content encoding unit 110 executes content feature calculation processing (S130). In the content feature calculation processing, the content encoding unit 110 calculates content features based on the acquired content prompt, for example.
[0053] Furthermore, the style encoding unit 120 performs a first style feature calculation process (S140). In the first style feature calculation process, the style encoding unit 120 calculates the first style features based, for example, on the acquired first style prompt.
[0054] Then, the audio decoder unit 130 performs the first audio decoding process (S150). In the first audio decoding process, the audio decoder unit 130 synthesizes the first audio based, for example, on content features and first style features.
[0055] Then, the control unit of the speech synthesis device 1 executes the first speech output process (S160). In the first speech output process, the control unit, for example, outputs the first speech from the sound output unit 350.
[0056] Subsequently, the control unit of the speech synthesis device 1 executes the process of acquiring the second style prompt (S170). In the second style prompt acquisition process, the control unit acquires a second style prompt (e.g., a string) based on user input to the operation unit 310, for example.
[0057] Furthermore, the control unit of the speech synthesis device 1 executes synthesis parameter setting processing (S180). In the synthesis parameter setting process, the control unit acquires the synthesis parameters (e.g., values) based on user input to the operation unit 310, for example.
[0058] Then, the style encoding unit 120 performs a second style feature calculation process (S190). In the second style feature calculation process, the style encoding unit 120 calculates the second style features based, for example, the acquired second style prompt.
[0059] Then, the control unit of the speech synthesis device 1 executes the synthesis style feature calculation process (S200). In the synthesis style feature calculation process, the control unit calculates the synthesis style feature based, for example, the first style feature, the second style feature, and the synthesis parameters.
[0060] Then, the audio decoder unit 130 performs a second audio decoding process (S210). In the second audio decoding process, the audio decoder unit 130 synthesizes a second audio based, for example, on content features and synthesized style features.
[0061] Then, the control unit of the speech synthesis device 1 executes the second speech output process (S220). In the second speech output process, the control unit, for example, outputs the second speech from the sound output unit 350.
[0062] Subsequently, the control unit of the speech synthesizer 1 determines, for example, whether to terminate the speech style adjustment based on user input to the operation unit 310 (S230). If it is determined that continuing the style adjustment is selected (S230: NO), the control unit of the speech synthesizer 1 stores, for example, the current synthesized style feature quantity as the first style feature quantity. Then, for example, it re-executes the steps from S170 onward.
[0063] If it is determined that style adjustment has been selected to end (S230: YES), the control unit of the speech synthesizer 1 terminates the process.
[0064] Here, "output" of sound can include not only the output of sound from the device itself (sound output), but also, for example, the output of sound to other functional parts of the device (internal output), and the output (external output) or transmission (external transmission) of sound to devices other than the device itself (external devices). Furthermore, the "acquisition" of a prompt can include not only inputting information into the device itself (operation input), but also, for example, obtaining information from other functional units of the device itself (internal acquisition), or obtaining (external acquisition) or receiving (external reception) information from devices other than the device itself (external devices).
[0065] Figure 4 shows an example of the configuration of terminal 10A, which is equipped with a speech synthesizer 1, as an example of terminal 10. Terminal 10 may be any information processing terminal capable of realizing the functions described in each embodiment. Examples of terminal 10 include smartphones, mobile phones (feature phones), computers (including, but not limited to, desktops, laptops, tablets, etc.), media computer platforms (including, but not limited to, cable, satellite set-top boxes, digital video recorders), handheld computer devices (including, but not limited to, PDAs (personal digital assistants), email clients, etc.), wearable devices (such as glasses-type devices and watch-type devices), VR (Virtual Reality) terminals, smart speakers (voice recognition devices), or other types of computers or communication platforms. Terminal 10 may also be referred to simply as an information processing terminal.
[0066] Terminal 10 includes, for example, a control unit 100 (CPU: central processing unit), a storage unit 200, an operation unit 310, a display unit 320, a communication unit 330, an audio input unit 340, and an audio output unit 350. Each component of terminal 10 is interconnected via a bus, for example. It is not necessary for terminal 10 to include all components. For example, terminal 10 may be configured to exclude individual components or multiple components, or it may not.
[0067] The operation unit 310 is implemented by any or a combination of any type of device capable of receiving user input and transmitting the information related to the input to the control unit 100. The input unit includes, for example, hardware keys such as touch panels, touch displays, and keyboards, as well as pointing devices such as mice and cameras (operation input via moving images).
[0068] The display unit 320 is implemented by any or a combination of any type of device capable of displaying data according to the display data written to the frame buffer. Examples of the display unit 320 include touch panels, touch displays, monitors (e.g., liquid crystal displays and OLEDs (organic electroluminescence displays)), head-mounted displays (HDMs), projection mapping, holograms, and devices capable of displaying images, text information, etc., in air (which may or may not be a vacuum). These display units 320 may or may not be capable of displaying display data in 3D.
[0069] The communication unit 330 transmits and receives various types of data via a network (not shown). Communication may be performed via wired or wireless connection, and any communication protocol may be used as long as communication between the units is possible. The communication unit 330 has the function of communicating with various devices, such as servers (not shown), via the network. The communication unit 330 transmits various types of data to these devices, such as servers, according to instructions from the control unit 100. The communication unit 330 also receives various types of data transmitted from these devices and transmits them to the control unit 100. The communication unit 330 may also be simply referred to as the communication unit. Furthermore, if the communication unit 330 is composed of a physically structured circuit, it may be referred to as the communication circuit.
[0070] The sound input unit 340 is used for inputting sound data (including voice data; the same applies hereinafter). The sound input unit 340 includes a microphone, etc. The sound output unit 350 is used for outputting sound data. The sound output unit 350 includes a speaker and the like.
[0071] The control unit 100 includes, for example, a central processing unit (CPU), a microprocessor, a processor core, a multiprocessor, an ASIC (application-specific integrated circuit), and an FPGA (field programmable gate array).
[0072] The storage unit 200 has the function of storing various programs and data necessary for the operation of the terminal 10. The storage unit 200 includes various storage media such as HDD (hard disk drive), SSD (solid state drive), flash memory, RAM (random access memory), and ROM (read-only memory). The storage unit 200 may or may not be referred to as memory.
[0073] Terminal 10 stores program P in the storage unit 200, and by executing program P, the control unit 100 executes the processing of each part included in the control unit 100. In other words, program P stored in the storage unit 200 enables terminal 10 to realize each function executed by the control unit 100. Furthermore, this program P may or may not be described as a program module.
[0074] The control unit 100 of terminal 10A comprises, as its main functional units, a content encoding unit 110, a style encoding unit 120, an audio decoder unit 130, and a display control unit 160. These functional units correspond to the functional units of the speech synthesis device 1 shown in Figure 1. The display control unit 160 also has a function, for example, to control the display output of the control unit 100.
[0075] The memory unit 200 stores, for example, a speech synthesis program 210 and a style feature primary memory unit 220. The speech synthesis program 210 is, for example, a program that is read by the control unit 100 and executed as a speech synthesis process. This speech synthesis process may be executed, for example, as a process based on the flowchart shown in Figure 3. The style feature primary storage unit 220 is, for example, a primary storage unit for storing first style features in speech synthesis processing.
[0076] Figure 5 shows an example of a speech synthesis application screen displayed on the display unit 320 of the terminal 10 in an application utilizing the speech synthesis device in this embodiment. This speech synthesis application screen is an example of a screen displayed when speech synthesis processing is performed by the control unit 100, for example.
[0077] In the speech synthesis application screen on the left side of Figure 5, the content prompt input area CPR for entering content prompts has text input in it, such as "This is a pen..." according to user input. Alternatively, when the voice input button to the right of the content prompt input area CPR is tapped, the user's voice may be recognized by a speech recognition unit (not shown), converted into text, and entered into the content prompt input area CPR.
[0078] The Style Prompt Input Area (SPR) has the style "feminine, very cute" entered as the first style prompt.
[0079] Furthermore, in the synthesis parameter setting area PSR, it may be possible to set the synthesis parameters, for example, by operating a slider.
[0080] For example, when a content prompt and a first style prompt are entered and the "Synthesize" button BT1 is tapped, the first voice is synthesized, and the display changes to the speech synthesis application screen shown in the center of Figure 5. On this screen, for example, the content prompt input area (CPR) is grayed out, and the content prompt can no longer be edited. Furthermore, for example, below the synthesis parameter setting area PSR, the synthesized speech confirmation area VOR is displayed for playing the first voice synthesized based on the content prompt and the first style prompt. Below the synthesized speech confirmation area VOR, the content of the first style prompt used to synthesize the first voice is displayed. For example, in the synthesized speech confirmation area (VOR), when the audio playback button is tapped, the first audio is played from the sound output unit 350.
[0081] In response to the first audio input, for example, the style prompt input area SPR is pre-filled with the style "low pitch" as the second style prompt. Then, when the "Resynthesize" button BT3 is tapped, the second voice is synthesized, and the display changes to the speech synthesis application screen shown on the right side of Figure 5. On this screen, below the synthesized speech confirmation area (VOR), the contents of the first style prompt used to synthesize the first speech and the contents of the second style prompt used to modify the style of the first speech are displayed. For example, in the synthesized speech confirmation area (VOR), when the audio playback button is tapped, the second audio is played from the sound output unit 350.
[0082] <Effects of the First Example> In this embodiment, the control unit 100 of the speech synthesizer 1 comprises a content encoding unit 110, a style encoding unit 120, and a speech decoder unit 130 (an example of a speech decoding unit). The control unit acquires a content prompt and a first style prompt. The content encoding unit calculates content features based on the content prompt. The style encoding unit calculates first style features based on the first style prompt. The speech decoding unit then synthesizes the first speech based on the content features and the first style features. The control unit also acquires a second style prompt for the first speech. The style encoding unit then calculates second style features based on the second style prompt. The speech decoding unit then synthesizes the second speech based on the content features and the internal division point between the first and second style features. According to this, the speech synthesizer can synthesize a second speech that reflects the second style prompt by using a second style prompt for a first speech synthesized based on a content prompt and a first style prompt.
[0083] Furthermore, this embodiment shows an example of a configuration in which the internal division point is determined by the composite parameter. According to this, the influence of the first style prompt and the second style prompt in the second voice can be adjusted by the synthesis parameters.
[0084] Furthermore, this embodiment demonstrates an example of a configuration in which parameters are determined based on a second style prompt. This allows users to more easily set appropriate parameters.
[0085] <First variation (1)> In the above embodiment, speech synthesis processing is performed on terminal 10A, but the invention is not limited to this. For example, terminal 10 may send the content prompt and the first style prompt to a server (not shown) upon receiving them. The server may then synthesize the first speech according to, for example, steps S130 to S150 in Figure 3 upon receiving the content prompt and the first style prompt. The server may then send the synthesized first speech to terminal 10.
[0086] Terminal 10 may output sound when it receives the first audio from the server. Terminal 10 may also send a second style prompt to the server when it receives one. When the server receives the second style prompt, it may synthesize the second audio, for example, according to steps S180 to S210 in Figure 3. The server may then send the synthesized second audio to terminal 10. Terminal 10 may also output sound when it receives the second audio from the server.
[0087] <First variation (2)> In the above embodiment, the style prompt was shown as a string, but it is not limited to this. For example, the style prompt may be selected by the user from a list of style options.
[0088] Figure 6 shows an example of the speech synthesis application screen displayed on the display unit 320 of terminal 10A in this modified example.
[0089] In the speech synthesis application screen on the left side of Figure 6, the style selection area STR is displayed instead of the style prompt input area SPR. The style selection area STR may display pre-set style tags (e.g., "feminine" or "manly"). When the user taps a style tag, for example, the style tag is highlighted and selected. Then, when the "Synthesize" button BT1 is tapped, for example, a prompt listing the selected style tags is set as the first style prompt, and the first speech is synthesized. Then, for example, the display changes to the speech synthesis application screen in the center of Figure 6.
[0090] On this screen, the style selection area STR may display style tags that are highly similar to the style tag selected as the first style prompt, for example. Alternatively, the style selection area STR may display style tags that are highly relevant to the style tag selected as the first style prompt (for example, if an aesthetic adjective such as "cute" is selected as the first style prompt, style tags that represent functional adjectives indicating properties such as "low pitch" or "faster"). Below the synthesized speech confirmation area VOR, the style tags used to synthesize the first speech are displayed.
[0091] When a style tag is selected in the style selection area STR and the "Resynthesize" button BT3 is tapped, a prompt listing the selected style tags is set as the second style prompt, and the second voice is synthesized. Then, for example, the display changes to the speech synthesis application screen shown on the right side of Figure 6. On this screen, below the synthesized speech confirmation area (VOR), the style tags used to synthesize the first voice and the style tags used to modify the style of the second voice are displayed. For example, in the synthesized speech confirmation area (VOR), when the audio playback button is tapped, the second audio is played from the sound output unit 350.
[0092] The input format for the style prompt may be selectable by the user. Figure 7 shows an example of the speech synthesis application screen displayed on the display unit 320 of terminal 10A in this case. In the speech synthesis application screen on the left side of Figure 7, for example, below the content prompt input area CPR, a style prompt input method selection menu area SMR may be displayed for selecting the input format of the style (style prompt). In the style prompt input method selection menu area SMR, for example, a "Text Input" menu for entering the style prompt as text, a "Style Selection" menu for selecting from style tags, and a "Voice Input" menu for inputting by voice input may be displayed.
[0093] When the "Voice Input" menu is selected, for example, a speech recognition unit (not shown) may acquire the style specification spoken by the user as a string. Alternatively, when the "Voice Input" menu is selected, a reference encoding unit (described later) may extract style features from the voice spoken by the user and calculate style feature quantities.
[0094] For example, when the "Text Input" menu is selected, the display changes to the speech synthesis application screen shown in the center of Figure 7. Then, when the style is entered as text in the style prompt input area SPR and the "Synthesize" button BT1 is tapped, the first voice is synthesized, and the display changes to the speech synthesis application screen shown on the right side of Figure 7.
[0095] On this screen, for example, when the audio playback button is tapped in the synthesized speech confirmation area VOR, the first audio is played from the sound output unit 350. Subsequently, in the style prompt input method selection menu area SMR, the style selection area STR is displayed because the "Style Selection" menu has been selected.
[0096] This modified example shows an example configuration in which the first style prompt and the second style prompt are selected by the user. This allows the user to configure the style prompt more easily.
[0097] Furthermore, this modified example shows an example of a configuration where the first style prompt and the second style prompt are input via voice. This allows the user to input style prompts more easily.
[0098] Furthermore, this modified example shows an example of a configuration in which the control unit controls the input format of the first style prompt and the second style prompt. With this configuration, the style prompt can be obtained according to the user's intent.
[0099] <First variation (3)> In the above embodiment, style features were calculated based on style prompts, but this is not limited to this. For example, style features could be calculated based on the style of the voice spoken by the user (referred to as "reference voice").
[0100] For example, when the speech synthesis device 1 acquires a reference speech from the sound input unit 340, it converts it into style features using the reference encoding unit. The reference encoding unit can be implemented, for example, by referring to Non-Patent Document 2.
[0101] This modified example shows a configuration in which the control unit of a speech synthesizer comprises a content encoding unit, a reference encoding unit, and a speech decoding unit. The control unit acquires a content prompt and a first speech; the content encoding unit calculates content features based on the content prompt; the reference encoding unit calculates first style features based on the first speech (first reference speech); the speech decoding unit synthesizes a second speech based on the content features and the first style features; the control unit acquires a third speech (second reference speech) for the second speech; the reference encoding unit calculates second style features based on the third speech; and the speech decoding unit synthesizes a fourth speech based on the content features and the internal division point between the first style features and the second style features. With this configuration, the speech synthesizer can synthesize a fourth speech by adjusting the second speech, which is synthesized based on the first speech, based on the third speech.
[0102] <Second Example> The second embodiment is an embodiment in which a speech conversion device (which may also be called an information processing device, a speech processing device, or a speech generation device) receives, for example, a first speech spoken by a user and a style, and synthesizes a second speech that reflects the content of the first speech and the style.
[0103] The contents described in the second embodiment are equally applicable to any of the other embodiments or other modifications.
[0104] Figure 8 is a block diagram showing an example of the functional configuration of the voice conversion device 2 according to one embodiment of this model. The voice conversion device 2 includes, for example, a content encoding unit 110, a style encoding unit 120, a voice decoder unit 130, a voice recognition unit 140, and a reference encoding unit 150. The content encoding unit 110, the style encoding unit 120, and the audio decoder unit 130 are, for example, similar to the functional units of the speech synthesis device 1.
[0105] The speech recognition unit 140, for example, receives a first speech input via the sound input unit 340, recognizes the content of the speech, and converts it into text. The speech recognition unit 140 may be configured, for example, with a speech recognition model such as Whisper.
[0106] The reference encoding unit 150, for example, receives a first speech input via the sound input unit 340, and has the function of analyzing the speech style and calculating style features. The reference encoding unit 150 may be configured, for example, by referring to Non-Patent Document 2.
[0107] Figure 9 is a flowchart showing an example of the processing flow performed by the voice conversion device 2 in this embodiment.
[0108] First, the control unit of the voice conversion device 2 executes the first voice acquisition process (S310). In the first audio acquisition process, the control unit acquires a first audio signal via the sound input unit 340 based on, for example, user input to the operation unit 310.
[0109] Then, the speech recognition unit 140 performs content prompt inference processing (S320). In the content prompt inference process, the speech recognition unit 140 performs speech recognition on the first speech, for example, and infers the content of the first speech's utterance. Then, for example, it generates a content prompt that includes the inferred content.
[0110] Once content features are calculated based on the content prompt (S130), the reference encoding unit 150 executes the first style feature calculation process (S330). In the first style feature calculation process, the reference encoding unit 150 calculates a first style feature that reflects the style of the first audio, for example, based on the first audio.
[0111] Then, the control unit of the voice conversion device 2 executes the first style prompt acquisition process (S340). In the first style prompt acquisition process, the control unit acquires a first style prompt (e.g., a string) based on user input to the operation unit 310, for example.
[0112] For example, when the synthesis parameter setting process is executed (S180), the style encoding unit 120 executes the second style feature calculation process (S350). In the second style feature calculation process, the style encoding unit 120 calculates the second style feature based, for example, the acquired first style prompt.
[0113] Figure 10 shows an example of the configuration of terminal 10B, which is equipped with a voice conversion device 2, as an example of terminal 10.
[0114] The control unit 100 of terminal 10B includes, as its main functional units, a content encoding unit 110, a style encoding unit 120, an audio decoder unit 130, an audio recognition unit 140, a reference encoding unit 150, and a display control unit 160. These functional units correspond to the functional units of the audio conversion device 2 shown in Figure 8.
[0115] The memory unit 200 stores, for example, a speech conversion program 230 and a style feature primary memory unit 220. The speech conversion program 230 is, for example, a program that is read by the control unit 100 and executed as a speech conversion process. This speech conversion process may be executed, for example, as a process based on the flowchart shown in Figure 9.
[0116] Figure 11 shows an example of a voice conversion application screen displayed on the display unit 320 of terminal 10B in an application utilizing the voice conversion device in this embodiment. This voice conversion application screen is an example of a screen displayed when, for example, the voice conversion process is executed by the control unit 100.
[0117] In the speech conversion application screen on the left side of Figure 11, the content prompt is not recognized and displayed in the content prompt input area (CPR) because the user has not yet spoken. For example, when the first style prompt is entered and the "Convert" button BT4 is tapped, the system enters a state of waiting for user input. For example, when the user inputs the first voice "This is a pen.", the display changes to the voice conversion application screen on the right side of Figure 11.
[0118] On this screen, for example, the content prompt input area (CPR) displays a content prompt based on the speech recognition results of the first voice's utterance. Furthermore, below the synthesis parameter setting area (PSR), for example, the synthesized speech confirmation area (VOR) is displayed for playing the second voice synthesized based on the first voice and the first style prompt. Below the synthesized speech confirmation area (VOR), the contents of the first style prompt used to synthesize the second voice are displayed. For example, in the synthesized speech confirmation area (VOR), when the audio playback button is tapped, the second audio is played from the sound output unit 350.
[0119] <Effects of the second example> In this embodiment, the control unit 100 of the speech conversion device 2 comprises a speech recognition unit 140, a content encoding unit 110, a style encoding unit 120, a reference encoding unit 150, and a speech decoder unit 130 (an example of a speech decoding unit). When the control unit acquires a first speech and a first style prompt, the speech recognition unit infers a content prompt based on the first speech. The reference encoding unit calculates a first style feature based on the first speech, and the content encoding unit calculates a content feature based on the content prompt. The style encoding unit also calculates a second style feature based on the first style prompt. Then, the speech decoding unit synthesizes a second speech based on the content feature and the internal division point between the first and second style features. According to this, the speech conversion device can synthesize a second speech that reflects the first style prompt by using a first style prompt for the first speech.
[0120] Furthermore, this embodiment shows an example of a configuration in which the internal division point is determined by the composite parameter. According to this, the effect that the first style prompt has on the second voice can be adjusted by the synthesis parameters.
[0121] Furthermore, this embodiment illustrates an example of a configuration in which parameters are determined based on a first style prompt. This allows users to more easily set appropriate parameters.
[0122] <Second variation (1)> In the above embodiment, the speech conversion process is performed on terminal 10B, but the invention is not limited to this. For example, terminal 10 may send the first speech and the first style prompt to a server (not shown) once it has received them. When the server receives the first speech and the first style prompt, it may synthesize a second speech according to, for example, steps S320 to S210 in Figure 9. The server may then send the synthesized second speech to terminal 10. Terminal 10 may also output sound when it receives the second audio from the server.
[0123] <Second variation (2)> In the above embodiment, the style prompt was shown as a string, but it is not limited to this. For example, the style prompt may be selected by the user from a list of style options.
[0124] Figure 12 shows an example of the voice conversion application screen displayed on the display unit 320 of terminal 10B in this modified example.
[0125] In the speech conversion application screen on the left side of Figure 12, the style selection area STR is displayed instead of the style prompt input area SPR. When a style tag is tapped by the user, for example, the style tag is highlighted and selected. Then, when the "Convert" button BT4 is tapped, for example, a prompt listing the selected style tags is set as the first style prompt, and the application enters a state of waiting for the user to speak. For example, when the user inputs the first voice "This is a pen.", the display changes to the speech conversion application screen on the right side of Figure 12.
[0126] On this screen, the style selection area STR may display style tags that are highly similar to the style tag selected as the first style prompt, for example. Alternatively, the style selection area STR may display style tags that are highly related to the style tag selected as the first style prompt, for example. Below the synthesized speech confirmation area VOR, the style tags used for synthesizing the second speech are displayed.
[0127] The input format for the style prompt may be selectable by the user. Figure 13 shows an example of the voice conversion application screen displayed on the display unit 320 of the terminal 10B in this case. In the speech conversion application screen on the left side of Figure 13, for example, below the content prompt input area CPR, a style prompt input method selection menu area SMR may be displayed for selecting the style (style prompt) input format.
[0128] When the "Voice Input" menu is selected, for example, the speech recognition unit 140 may acquire the style specification spoken by the user as a string.
[0129] For example, when the "Text Input" menu is selected, the display changes to the speech conversion application screen shown in the center of Figure 13. Then, when a style is entered as text in the style prompt input area SPR and the "Convert" button BT4 is tapped, the system enters a state of waiting for the user to speak. For example, when the user inputs the first voice "This is a pen.", the display changes to the speech conversion application screen shown on the right side of Figure 13.
[0130] On this screen, for example, when the audio playback button is tapped in the synthesized speech confirmation area VOR, the second audio is played from the sound output unit 350.
[0131] This modified example shows a configuration in which the first style prompt is selected by the user. This allows the user to configure the style prompt more easily.
[0132] Furthermore, this modified example shows a configuration where the first style prompt is entered via voice. This allows the user to input style prompts more easily.
[0133] Furthermore, this modified example shows an example of a configuration in which the control unit performs control to specify the input format of the first style prompt. With this configuration, the style prompt can be obtained according to the user's intent.
[0134] <Other>
[0135] In the above embodiment, at least some of the processing that the server was supposed to perform may be performed by the terminal 10. Conversely, in the above embodiment, at least some of the processing that the terminal 10 was supposed to perform may be performed by the server. Furthermore, a system consisting of one or more servers may be defined as a server system, and the server of the present invention may be considered a server system.
[0136] Furthermore, in the above embodiments, the server for distributing various applications (the server from which terminal 10 downloads applications) may be configured as a different server from the server that provides the corresponding services (applications). In other words, the server for distributing applications and the server that performs the application management processing described in the above embodiments may be configured as physically separate servers, or they may be configured as a single server.
[0137] Furthermore, the term "application" may include not only programs for various applications, but also, for example, programs that provide the functionality of other services as a function of the main application (for example, programs that provide the functionality of speech synthesis or speech conversion services as a function of a messaging application, or vice versa), and programs for updating the main application. It may also include data used by application programs (which may also include data for application updates).
[0138] Furthermore, as mentioned above, the contents described in each of the above embodiments, each of the modifications, other embodiments, and so on can be combined and applied accordingly. [Explanation of Symbols]
[0139] 1. Speech synthesis device 2. Voice converter 10 devices 110 Content Encoding Section 120 Style Encoding Section 130 Audio Decoder Unit 140 Voice Recognition Unit 150 Reference Encoding Section
Claims
1. A speech synthesis device, The control unit of the speech synthesis device comprises a content encoding unit, a style encoding unit, and a speech decoding unit. The control unit acquires the content prompt and the first style prompt, The content encoding unit calculates content features based on the content prompt, The style encoding unit calculates a first style feature based on the first style prompt, The speech decoding unit synthesizes a first speech based on the content features and the first style features. The control unit acquires a second style prompt for the first voice, The style encoding unit calculates a second style feature based on the second style prompt, The speech decoding unit synthesizes a second speech based on the content features and the internal division point between the first style features and the second style features. Speech synthesis device.
2. A speech synthesis device according to claim 1, The aforementioned internal division point is determined by a parameter. Speech synthesis device.
3. A speech synthesis device according to claim 2, The aforementioned parameter is determined based on the second style prompt. Speech synthesis device.
4. A speech synthesis device according to claim 1, The first style prompt and the second style prompt are selected by the user. Speech synthesis device.
5. A speech synthesis device according to claim 1, Equipped with a voice recognition unit, The first style prompt and the second style prompt are input by voice. Speech synthesis device.
6. A speech synthesis device according to claim 4 or 5, The control unit performs control to specify the input format of the first style prompt and the second style prompt. Speech synthesis device.
7. A program executed by a speech synthesis device, The control unit of the speech synthesis device comprises a content encoding unit, a style encoding unit, and a speech decoding unit. The control unit acquires the content prompt and the first style prompt, Based on the content prompt, the content features are calculated by the content encoding unit, Based on the first style prompt, the first style feature is calculated by the style encoding unit, The speech decoding unit synthesizes the first audio based on the content features and the first style features, The control unit acquires a second style prompt for the first voice, Based on the second style prompt, the second style feature is calculated by the style encoding unit, The second audio is synthesized by the speech decoding unit based on the content features and the internal division point between the first style features and the second style features. A program that includes this.
8. An information processing method performed by a speech synthesis device, The control unit of the speech synthesis device comprises a content encoding unit, a style encoding unit, and a speech decoding unit. The control unit acquires the content prompt and the first style prompt, Based on the content prompt, the content features are calculated by the content encoding unit, Based on the first style prompt, the first style feature is calculated by the style encoding unit, The speech decoding unit synthesizes the first audio based on the content features and the first style features, The control unit acquires a second style prompt for the first voice, Based on the second style prompt, the second style feature is calculated by the style encoding unit, The second audio is synthesized by the speech decoding unit based on the content features and the internal division point between the first style features and the second style features. Information processing methods, including those mentioned above.
9. It is a voice conversion device, The control unit of the aforementioned speech conversion device comprises a speech recognition unit, a content encoding unit, a style encoding unit, a reference encoding unit, and a speech decoding unit. The control unit acquires the first voice and the first style prompt, The speech recognition unit infers a content prompt based on the first speech, The reference encoding unit calculates a first style feature based on the first audio, The content encoding unit calculates content features based on the content prompt, The style encoding unit calculates a second style feature based on the first style prompt, The speech decoding unit synthesizes a second speech based on the content features and the internal division point between the first style features and the second style features. Voice conversion device.
10. The voice conversion device according to claim 9, The aforementioned internal division point is determined by a parameter. Voice conversion device.
11. A voice conversion device according to claim 10, The parameter is determined based on the first style prompt. Voice conversion device.
12. The voice conversion device according to claim 9, The first style prompt is selected by the user. Voice conversion device.
13. The voice conversion device according to claim 9, The first style prompt is input by voice. Voice conversion device.
14. A voice conversion device according to claim 12 or 13, The control unit performs control to specify the input format of the first style prompt. Voice conversion device.
15. A program executed by a speech converter, The control unit of the aforementioned speech conversion device comprises a speech recognition unit, a content encoding unit, a style encoding unit, a reference encoding unit, and a speech decoding unit. The control unit acquires the first voice and the first style prompt, Based on the first voice, the content prompt is inferred by the voice recognition unit, Based on the first audio, the first style feature is calculated by the reference encoding unit, Based on the content prompt, the content features are calculated by the content encoding unit, Based on the first style prompt, the second style feature is calculated by the style encoding unit, The second audio is synthesized by the speech decoding unit based on the content features and the internal division point between the first style features and the second style features. A program that includes this.
16. An information processing method performed by a speech conversion device, The control unit of the aforementioned speech conversion device comprises a speech recognition unit, a content encoding unit, a style encoding unit, a reference encoding unit, and a speech decoding unit. The control unit acquires the first voice and the first style prompt, Based on the first voice, the content prompt is inferred by the voice recognition unit, Based on the first audio, the first style feature is calculated by the reference encoding unit, Based on the content prompt, the content features are calculated by the content encoding unit, Based on the first style prompt, the second style feature is calculated by the style encoding unit, The second audio is synthesized by the speech decoding unit based on the content features and the internal division point between the first style features and the second style features. Information processing methods, including those mentioned above.