Electronic apparatus, terminal apparatus and controlling method thereof
Patent Information
- Application Number
- KR1020210138343
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-18
- Publication Date
- 2026-08-14
- Estimated Expiration
- 2041-10-18
Smart Images

Figure 112021118917925-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to an electronic device, a terminal device, and a method for controlling the same, and more specifically, to an electronic device, a terminal device, and a method for controlling the same that generates and outputs a voice waveform from text. Background Technology
[0002] With the advancement of speech processing technology, electronic devices that perform speech processing functions are being utilized. One of the various speech processing functions is Text-to-Speech (TTS). TTS refers to the function of converting text into speech and outputting it as a speech signal. For example, the TTS function performs speech conversion using a prosody unit and a vocoder unit. The prosody unit estimates acoustic features based on the text. In other words, the prosody unit can estimate the pronunciation and prosody of the synthesized speech. The estimated acoustic features are input to the vocoder unit. The vocoder unit estimates the speech waveform from the input acoustic features. The TTS function can be performed by outputting the speech waveform estimated by the vocoder unit through a speaker.
[0003] Generally, the prosody and vocoder sections can be trained to estimate speech waveforms from acoustic characteristics; however, since the vocoder section supports only the acoustic characteristics used for training, it can output speech waveforms at a fixed sampling rate only. Therefore, separate prosody and vocoder sections are required to output speech waveforms at various sampling rates.
[0004] In some cases, a single electronic device may be able to output voice signals at various sampling rates, and in others, different electronic devices may output voice signals at different sampling rates. Additionally, the specifications of external speakers connected to a single electronic device may also vary. The conventional method has the disadvantage of training separate prosody and vocoder units to use the trained prosody and vocoder units universally, and including multiple prosody and vocoder units in a single electronic device.
[0005] Therefore, there is a need for a technology that can output voice signals of various sampling rates with a single prosody section and vocoder section. The problem to be solved
[0006] The present disclosure is intended to solve the aforementioned problems, and the purpose of the present disclosure is to provide an electronic device including a vocoder unit that outputs voice waveforms of various sampling rates using the same acoustic characteristics estimated from a single prosody unit, and a method for controlling the same. In addition, the present disclosure is to provide an electronic device and a method for controlling the same that identify the specifications of the electronic device and output a voice signal including audio characteristics corresponding to the identified specifications. means of solving the problem
[0007] An electronic device according to one embodiment of the present disclosure includes a processor comprising an input interface, a prosody module for extracting acoustic characteristics, and a vocoder module for generating a voice waveform, wherein the processor receives text through the input interface, identifies a first acoustic characteristic from the input text using the prosody module, generates a modified acoustic characteristic having a sampling rate different from the first acoustic characteristic based on the identified first acoustic characteristic, and trains the vocoder module based on each of the first acoustic characteristic and the modified acoustic characteristic to generate a plurality of vocoder learning models.
[0008] A terminal device according to one embodiment of the present disclosure includes a processor and a speaker, wherein the processor identifies the specification of a component associated with the terminal device and selects one of the plurality of vocoder learning models based on the specification of the identified component, identifies acoustic characteristics from text using the prosody module, and generates a voice waveform corresponding to the identified acoustic characteristics using the identified vocoder learning model and outputs it through the speaker.
[0009] A control method for an electronic device according to one embodiment of the present disclosure includes the steps of receiving text input, identifying a first acoustic characteristic from the input text using a prosody module that extracts acoustic characteristics, generating a modified acoustic characteristic having a sampling rate different from the first acoustic characteristic based on the identified first acoustic characteristic, and training a vocoder module that generates a speech waveform based on each of the first acoustic characteristic and the modified acoustic characteristic to generate a plurality of vocoder learning models.
[0010] A control method for a terminal device according to one embodiment of the present disclosure includes the steps of: identifying the specification of a component associated with the terminal device; selecting one vocoder learning model among a plurality of vocoder learning models based on the specification of the identified component; identifying acoustic characteristics from text using a prosody module; and generating a voice waveform corresponding to the identified acoustic characteristics using the identified vocoder learning model and outputting it through a speaker. Brief explanation of the drawing
[0011] FIG. 1 is a drawing illustrating a system including an electronic device and a terminal device according to one embodiment of the present disclosure. FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure. FIG. 3 is a block diagram illustrating the configuration of a terminal device according to one embodiment of the present disclosure. FIG. 4 is a block diagram illustrating the specific configuration of a terminal device according to one embodiment of the present disclosure. FIG. 5 is a block diagram illustrating the configuration of a processor according to one embodiment of the present disclosure. FIGS. 6a and 6b are drawings illustrating the process of training a vocoder model according to one embodiment of the present disclosure. FIGS. 7a and 7b are drawings illustrating the process of selecting a vocoder learning model corresponding to a terminal device according to one embodiment of the present disclosure. FIG. 8 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure. FIG. 9 is a flowchart illustrating a control method for a terminal device according to one embodiment of the present disclosure. Specific details for implementing the invention
[0012] Hereinafter, various embodiments are described in more detail with reference to the attached drawings. The embodiments described in this specification may be modified in various ways. Specific embodiments may be depicted in the drawings and described in detail in the detailed description. However, specific embodiments disclosed in the attached drawings are intended only to facilitate understanding of various embodiments. Accordingly, the technical concept is not limited by specific embodiments disclosed in the attached drawings, and it should be understood that it includes all equivalents or substitutions that fall within the concept and scope of the disclosure.
[0013] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but these components are not limited by the aforementioned terms. The aforementioned terms are used solely for the purpose of distinguishing one component from another.
[0014] In this specification, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. When a component is described as being “connected” or “connected” to another component, it should be understood that it may be directly connected to or connected to that other component, or that there may be other components in between. On the other hand, when a component is described as being “directly connected” or “directly connected” to another component, it should be understood that there are no other components in between.
[0015] Meanwhile, a "module" or "part" for a component as used in this specification performs at least one function or operation. Furthermore, a "module" or "part" may perform a function or operation by hardware, software, or a combination of hardware and software. Additionally, a plurality of "modules" or a plurality of "parts," excluding a "module" or "part" that must be performed on specific hardware or on at least one processor, may be integrated into at least one module. A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0016] In describing the present disclosure, the order of each step should be understood as non-limiting, unless the preceding step must logically and temporally be performed prior to the subsequent step. That is, except for such exceptional cases, the essence of the disclosure is not affected even if the process described in the subsequent step is performed prior to the process described in the preceding step, and the scope of the rights should be defined regardless of the order of the steps. Furthermore, the designation "A or B" in this specification is defined to mean not only selectively referring to either A or B, but also including both A and B. Additionally, the term "included" in this specification has a meaning that includes additional components beyond the elements listed as included.
[0017] This specification describes only the essential components necessary for explaining the present disclosure and does not mention components unrelated to the essence of the present disclosure. Furthermore, the mentioned components should not be interpreted in an exclusive sense to include only those components, but should be interpreted in a non-exclusive sense to include other components as well.
[0018] Furthermore, in describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is abbreviated or omitted. Meanwhile, each embodiment may be implemented or operated independently, but each embodiment may also be implemented or operated in combination.
[0019] FIG. 1 is a drawing illustrating a system including an electronic device and a terminal device according to one embodiment of the present disclosure.
[0020] Referring to FIG. 1, the system may include an electronic device (100) and a terminal device (200). For example, the electronic device (100) may include a server, a cloud, etc., and the server, etc. may include a management server, a training server, etc. And, the terminal device (200) may include a smartphone, a tablet PC, a navigation device, a slate PC, a wearable device, a digital TV, a desktop computer, a laptop computer, a home appliance, an IoT device, a kiosk, etc.
[0021] The electronic device (100) includes a prosody module and a vocoder module. The prosody module may include a single prosody model, and the vocoder module may include multiple vocoder models. The prosody model and the vocoder model may be artificial intelligence neural network models. The electronic device (100) can extract acoustic characteristics from text using the prosody model. Since errors such as pronunciation errors may occur in the prosody model, the electronic device (100) can correct errors in the prosody model through an artificial intelligence learning process.
[0022] A prosody model can extract acoustic characteristics of one type of sampling rate. For example, a prosody model can extract acoustic characteristics of a sampling rate of 24 kHz. The electronic device (100) can generate modified acoustic characteristics based on the acoustic characteristics extracted from the prosody model. For example, the electronic device (100) can generate acoustic characteristics of sampling rates of 16 kHz and 8 kHz using the acoustic characteristics of a sampling rate of 24 kHz.
[0023] The electronic device (100) can train a vocoder model of a vocoder module using acoustic characteristics and modified acoustic characteristics extracted from a prosody model. The vocoder module may be one, but may include multiple training models, each trained with different acoustic characteristics. For example, the electronic device may train a first vocoder model based on acoustic characteristics of a sampling rate of 24 kHz, train a second vocoder model based on acoustic characteristics of a sampling rate of 16 kHz, and train a third vocoder model based on acoustic characteristics of a sampling rate of 8 kHz.
[0024] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0025] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform a desired characteristic (or objective) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0026] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0027] The prosody model and vocoder model learned in the electronic device (100) may be included in the terminal device (200). The terminal device (200) also includes a prosody module and a vocoder module. The electronic device (100) may transmit the prosody model and the vocoder learning model to the terminal device (200) using a wired or wireless communication method. Alternatively, the terminal device (200) may include the prosody model and the vocoder learning model during manufacturing. That is, the vocoder module of the terminal device (200) may include a plurality of vocoder learning models learned by various sampling rates. The terminal device (200) may select the optimal vocoder learning model among the plurality of vocoder learning models based on the specifications of the terminal device (200), whether streaming output is possible, sampling rate, sound quality, etc. And, the terminal device (200) may output text as a voice waveform using the selected vocoder learning model.
[0028] Up to this point, an embodiment for training a prosody model and a vocoder model in an electronic device (100) has been described. However, while the initial training process may be performed in the electronic device (100), subsequent continuous training processes for correcting errors and updating may be performed in the terminal device (200). In another embodiment, the electronic device (100) includes the trained prosody model and the vocoder training model and can generate a voice waveform from text transmitted from the terminal device (200). Then, the generated voice waveform can be transmitted to the terminal device (200). The terminal device (200) can output the voice waveform received from the electronic device (100) through a speaker.
[0029] Below, the configuration of the electronic device (100) and the terminal device (200) is described.
[0030] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0031] Referring to FIG. 2, the electronic device (100) includes an input interface (110) and a processor (120).
[0032] The input interface (110) receives text input. Alternatively, the input interface (110) may receive commands from a user. For example, the input interface (110) may include a communication interface, an input / output port, etc. The input interface (110) performs the function of receiving text input or receiving commands from a user and may be called an input unit, an input device, an input module, etc.
[0033] When the input interface (110) is implemented as a communication interface, the input interface (110) can communicate with an external device. The input interface (110) can receive text data from an external device using a wired or wireless communication method. For example, the communication interface may include a module capable of communicating via methods such as 3G, LTE (Long Term Evolution), 5G, Wi-Fi, Bluetooth, DMB (Digital Multimedia Broadcasting), ATSC (Advanced Television Systems Committee), DVB (Digital Video Broadcasting), LAN (Local Area Network), etc. The communication interface that communicates with an external device may also be referred to as a communication unit, a communication device, a communication module, a transceiver, etc.
[0034] When the input interface (110) is implemented as an input / output port, the input interface (110) can receive text data from an external device (including external memory). For example, when the input interface (110) is implemented as an input / output port, the input / output port may include ports such as HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), Thunderbolt, LAN, etc.
[0035] Meanwhile, the input interface (110) can receive control commands from the user. For example, the input interface (110) may include a keypad, a touchpad, a touchscreen, etc.
[0036] The processor (120) can control each component of the electronic device (100). For example, the processor (120) can control an input interface (110) to receive text input. The processor (120) includes a prosody module for extracting acoustic characteristics and a vocoder module for generating a voice waveform. The processor (120) can identify (extract) acoustic characteristics from the input text using the prosody module. Based on the identified acoustic characteristics, the processor (120) can generate modified acoustic characteristics with different sampling rates from the identified acoustic characteristics. For example, if the identified acoustic characteristics have a sampling rate of 24 kHz, the processor (120) can generate acoustic characteristics with a sampling rate of 16 kHz and 8 kHz based on the acoustic characteristics with a sampling rate of 24 kHz. The processor (120) can generate modified acoustic characteristics by downsampling the identified acoustic characteristics or approximating them to preset acoustic characteristics.
[0037] The processor (120) can train a vocoder model corresponding to each acoustic characteristic using each of the identified acoustic characteristic and the modified acoustic characteristic, and generate a vocoder learning model. For example, the processor (120) can generate a vocoder learning model trained with the identified acoustic characteristic. Alternatively, the processor (120) can generate a vocoder learning model trained with the downsampled modified acoustic characteristic or the modified acoustic characteristic approximated by a preset acoustic characteristic. Alternatively, the processor (120) can generate a vocoder learning model trained using both the first modified acoustic characteristic approximated by the preset acoustic characteristic and the second modified acoustic characteristic generated by downsampling the first acoustic characteristic.
[0038] The electronic device (100) can transmit the prosody model and the vocoder learning model to the terminal device (200). For example, the electronic device (100) can transmit the prosody model and the vocoder learning model to the terminal device (200) through an input / output port or a communication interface.
[0039] FIG. 3 is a block diagram illustrating the configuration of a terminal device according to one embodiment of the present disclosure.
[0040] Referring to FIG. 3, the terminal device (200) includes a processor (210) and a speaker (220).
[0041] The processor (210) can control each component of the terminal device (200). The processor (210) includes a prosody module that extracts acoustic characteristics and a vocoder module that generates a voice waveform. The prosody module of the processor (210) includes a learned prosody model, and the vocoder module may include a plurality of vocoder learning models. The plurality of vocoder learning models may be models learned at different sampling rates. The processor (210) can identify the specifications of a component related to the terminal device (200). For example, the specifications of the component may include the processor's resources, the processor's operation status, memory capacity, memory bandwidth, speaker performance, etc. The specifications of the components related to the terminal device (200) described above may be referred to as internal specifications. Meanwhile, the terminal device (200) may be connected to an external speaker. In this case, the specifications of the components related to the terminal device (200) may include various information about the external speaker.
[0042] The processor (210) can select one vocoder learning model among a plurality of vocoder learning models based on the specifications of the identified components. For example, the processor (210) can identify candidate vocoder learning models based on the specifications of the internal components of the terminal device among the components and whether streaming output of the voice waveform is possible. Then, the processor (210) can select one vocoder learning model among the candidate vocoder learning models based on the highest sampling rate and sound quality. Alternatively, the processor (210) can select one vocoder learning model based on the resources of the processor. Meanwhile, as described above, the terminal device (200) can be connected to an external speaker. In this case, the processor (210) can identify the specifications of the external speaker and select one vocoder learning model based on the identified specifications of the external speaker.
[0043] The processor (210) can identify acoustic characteristics from text using a prosody module. For example, the terminal device (200) may further include memory, and the processor (210) may identify acoustic characteristics from text data stored in the memory. Alternatively, the terminal device (200) may further include a communication interface, and the processor (210) may identify acoustic characteristics from text data received through the communication interface. Alternatively, the terminal device (200) may further include an input interface, and the processor (210) may identify acoustic characteristics from text data input through the input interface. The processor (210) may generate a speech waveform corresponding to the identified acoustic characteristics using an identified vocoder learning model.
[0044] The speaker (220) can output a generated voice waveform. Alternatively, the speaker (220) can output a user's input command, status-related information or operation-related information of the terminal device (200), etc., as voice or a notification sound.
[0045] FIG. 4 is a block diagram illustrating the specific configuration of a terminal device according to one embodiment of the present disclosure.
[0046] Referring to FIG. 4, the terminal device (200) may include a processor (210), a speaker (220), an input interface (230), a communication interface (240), a camera (250), a microphone (260), a display (270), a memory (280), and a sensor (290). Since the speaker (220) is the same as described in FIG. 3, a detailed description is omitted.
[0047] The input interface (230) can receive commands from a user. Alternatively, the input interface (230) can receive text data from a user. The input interface (230) performs the function of receiving commands or text data from an external source and may be referred to as an input unit, an input device, an input module, etc. For example, the input interface (230) may include a keypad, a touchpad, a touchscreen, etc.
[0048] The communication interface (240) can perform communication with an external device. The communication interface (240) can receive text data from an external device using a wired or wireless communication method. As an example, the text may be provided to the terminal device (200) via a web server, cloud, etc. For example, the communication interface (240) may include a module capable of performing communication in the form of 3G, LTE (Long Term Evolution), 5G, Wi-Fi, Bluetooth, DMB (Digital Multimedia Broadcasting), ATSC (Advanced Television Systems Committee), DVB (Digital Video Broadcasting), LAN (Local Area Network), etc. The communication interface (240) that performs communication with an external device may be referred to as a communication unit, a communication device, a communication module, a transceiver unit, etc.
[0049] The camera (250) can capture an external environment and receive the captured image. As an example, the camera (250) captures an image containing text, and the processor (210) can recognize the text contained in the image using an OCR (Optical Character Recognition) function. For example, the camera (250) may include a CCD sensor and a CMOS sensor.
[0050] The microphone (260) can receive an external sound signal. The processor (210) can process the input sound signal and perform a corresponding operation. For example, if the external sound signal is the user's voice, the processor (210) can recognize a control command based on the input voice and perform a control operation corresponding to the recognized control command.
[0051] The display (270) can output a video signal on which video processing has been performed. For example, the display (270) can be implemented as an LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diode), flexible display, touch screen, etc. If the display (270) is implemented as a touch screen, the terminal device (200) can receive control commands through the touch screen.
[0052] The memory (280) can store data that performs the functions of the terminal device (200), and can store programs, commands, etc. that run on the terminal device (200). For example, the memory (280) can store text data, prosody models, and multiple vocoder learning models. In addition, the prosody models and selected vocoder learning models stored in the memory (280) can be loaded into the processor (210) under the control of the processor (210) to perform operations. Programs, AI models, data, etc. stored in the memory (280) can be loaded into the processor (210) to perform operations. For example, the memory (280) can be implemented in the form of ROM, RAM, HDD, SSD, memory card, etc.
[0053] The sensor (290) can detect the user's movements, distance, location, etc. The processor (210) can recognize a control command based on the user's movements, distance, location, etc. detected by the sensor (290) and perform a control operation corresponding to the recognized control command. Alternatively, the sensor (290) can detect surrounding environment information of the terminal device (200). The processor (210) can perform a corresponding control operation based on the surrounding environment information detected by the sensor (290). For example, the sensor (290) may include an accelerometer, a gravity sensor, a gyroscope, a geomagnetic sensor, a direction sensor, a motion recognition sensor, a proximity sensor, a voltmeter, an ammeter, a barometer, a hygrometer, a thermometer, an illuminance sensor, a heat detection sensor, a touch sensor, an infrared sensor, an ultrasonic sensor, etc.
[0054] FIG. 5 is a block diagram illustrating the configuration of a processor according to one embodiment of the present disclosure.
[0055] Referring to FIG. 5, the processor (210) may include a prosody module (211) and a vocoder module (212). The prosody module (211) may include a prosody model that extracts acoustic characteristics, and the vocoder module (212) may include a vocoder learning model that generates a speech waveform from the extracted acoustic characteristics. In one embodiment, the prosody module (211) and the vocoder module (212) may be implemented in hardware or software. When the prosody module (211) and the vocoder module (212) are implemented in hardware, the prosody module (211) and the vocoder module (212) may be implemented as a component of the processor (210). If the prosody module (211) and the vocoder module (212) are implemented in software, the prosody module (211) and the vocoder module (212) are stored in memory and can be loaded from memory to the processor when the terminal device (200) executes the TTS function. Alternatively, the prosody model and the vocoder learning model are implemented in software and can be loaded from memory to the processor when the terminal device (200) executes the TTS function.
[0056] The prosody module (211) may include a prosody model that extracts acoustic characteristics from text. The vocoder module (212) may include a vocoder learning model based on the specifications of components related to the terminal device (200). Acoustic characteristics extracted from the prosody module (211) are input to the vocoder learning module (212), and the vocoder learning module (212) can generate a voice waveform corresponding to the acoustic characteristics using the selected vocoder learning model. The generated voice waveform can be output through a speaker.
[0057] Up to now, each configuration of the terminal device (200) has been described. Below, the process of training a vocoder model and selecting the optimal vocoder learning model among multiple vocoder learning models is described.
[0058] FIGS. 6a and 6b are drawings illustrating the process of training a vocoder model according to one embodiment of the present disclosure.
[0059] Referring to FIG. 6a, the process of training various vocoder models using the same prosody model is illustrated. For example, the vocoder model can be trained based on a 24 kHz speech waveform (Waveform_24). The prosody model can extract 24 kHz acoustic characteristics (feat_24) from the 24 kHz speech waveform (S110). The extracted 24 kHz acoustic characteristics can be used in the process of training a vocoder model (11) with a 24 kHz sampling rate and a vocoder model (12) with a 16 kHz sampling rate. Since the 24 kHz acoustic characteristics contain all the information of the 16 kHz speech waveform, they can also be used to train the vocoder model (12) with a 16 kHz sampling rate. The electronic device can downsample a 24 kHz voice waveform to a 16 kHz voice waveform (Waveform_16) for training a vocoder model (12) with a 16 kHz sampling rate (S120).
[0060] The extracted 24 kHz acoustic characteristics can be input into a vocoder model (11) with a 24 kHz sampling rate. Then, the vocoder model (11) with a 24 kHz sampling rate can generate a speech waveform with a 24 kHz sampling rate based on the input 24 kHz acoustic characteristics. The electronic device can identify the loss of the speech waveform based on the generated speech waveform and the 24 kHz speech waveform used for learning (S130). Based on the identified loss of the speech waveform, the electronic device can train the vocoder model (11) with a 24 kHz sampling rate to generate a vocoder learning model with a 24 kHz sampling rate.
[0061] In a similar manner, the extracted 24 kHz acoustic characteristics can be input into a vocoder model (12) with a 16 kHz sampling rate. Then, the vocoder model (12) with a 16 kHz sampling rate can generate a speech waveform with a 16 kHz sampling rate based on the input 24 kHz acoustic characteristics. The electronic device can identify the loss of the speech waveform based on the generated speech waveform and the downsampled 16 kHz speech waveform (S140). Based on the identified loss of the speech waveform, the electronic device can train the vocoder model (12) with a 16 kHz sampling rate to generate a vocoder learning model with a 16 kHz sampling rate.
[0062] Referring to FIG. 6b, a process of training various vocoder models by performing an approximation process is illustrated. For example, a prosody model can extract 24 kHz acoustic characteristics (feat_24) from a 24 kHz speech waveform (S210). The extracted 24 kHz acoustic characteristics can be used in the process of training a vocoder model (11) with a 24 kHz sampling rate and a vocoder model (12) with a 16 kHz sampling rate. The electronic device can downsample the 24 kHz speech waveform to a 16 kHz speech waveform (Waveform_16) for training the vocoder model (12) with a 16 kHz sampling rate (S220).
[0063] The extracted 24 kHz acoustic characteristics can be input into a vocoder model (11) with a 24 kHz sampling rate. Then, the vocoder model (11) with a 24 kHz sampling rate can generate a speech waveform with a 24 kHz sampling rate based on the input 24 kHz acoustic characteristics. The electronic device can identify the loss of the speech waveform based on the generated speech waveform and the 24 kHz speech waveform used for learning (S240). Based on the identified loss of the speech waveform, the electronic device can train the vocoder model (11) with a 24 kHz sampling rate to generate a vocoder learning model with a 24 kHz sampling rate.
[0064] The electronic device can approximate the 24 kHz acoustic characteristics extracted from the prosody module to acoustic characteristics of a preset sampling rate (S230). For example, the electronic device can approximate the 24 kHz acoustic characteristics to 16 kHz acoustic characteristics (feat_16). The approximated 16 kHz acoustic characteristics can be used to train a vocoder model (12) with a 16 kHz sampling rate. The approximated 16 kHz acoustic characteristics can be input to the vocoder model (12) with a 16 kHz sampling rate. Then, the vocoder model (12) with a 16 kHz sampling rate can generate a speech waveform with a 16 kHz sampling rate based on the input 16 kHz acoustic characteristics. The electronic device can identify the loss of the speech waveform based on the generated speech waveform and the downsampled 16 kHz speech waveform (S250). The electronic device can generate a vocoder learning model with a 16 kHz sampling rate by training a vocoder model (12) with a 16 kHz sampling rate based on the loss of the identified voice waveform.
[0065] FIGS. 7a and 7b are drawings illustrating the process of selecting a vocoder learning model corresponding to a terminal device according to one embodiment of the present disclosure.
[0066] Various vocoder learning models can be generated through the process described in FIGS. 6a and 6b. The generated vocoder learning models can be included in a terminal device (200).
[0067] Referring to FIG. 7a, the terminal device (200) may include a plurality of vocoder learning models (1). For example, the first vocoder learning model may include the characteristics (c1, q1, s1), and the nth vocoder learning model may include the characteristics (cn, qn, sn). c represents the complexity of the vocoder model, and the amount of computation may increase as the complexity increases. q represents the sound quality, and the higher q is, the higher the SNR (Signal-to-Noise Ratio). s represents the sampling rate.
[0068] The terminal device (200) can identify an optimal vocoder learning model based on specifications related to the terminal device (S310). For example, specifications related to the terminal device may include processor resources, whether the processor is operating, memory capacity, memory bandwidth, speaker performance, and specifications of the external speaker if an external speaker is connected. For example, specifications related to the terminal device may include internal specifications of a fixed terminal device (e.g., processor, memory, etc.) and external specifications of a variable terminal device (speaker, etc.). The terminal device (200) can identify a candidate group (2) of streaming-capable vocoder learning models based on internal specifications. Then, the terminal device (200) can select an optimal vocoder learning model (3) based on other internal specifications or external specifications. As an example of one embodiment, the candidate group (2) of the vocoder learning model may be (c1, low quality, 16 kHz), (c2, medium quality, 16 kHz), (c3, high quality, 16 kHz), and (c4, low quality, 24 kHz). When the terminal device (200) outputs to a smartphone speaker with low high-frequency expression capability, it may select the model (c3, high quality, 16 kHz), which has good sound quality even with a low sampling rate. Alternatively, when the terminal device (200) outputs to high-end headphones, it may be advantageous to provide a high bandwidth even with some noise, so it may select the model (c4, low quality, 24 kHz). Alternatively, in the case of low-cost headphones or earphones, there may be distortion and additional noise, so the terminal device (200) may select the model (c2, medium quality, 16 kHz).
[0069] The terminal device (200) can extract acoustic characteristics from text using a prosody model (31) and generate a speech waveform using the extracted acoustic characteristics and a selected vocoder learning model (3) included in the vocoder module (32). The terminal device (200) may include the same prosody model and various vocoder learning models. If the sampling rate of the acoustic characteristics extracted from the prosody model and the sampling rate of the selected vocoder learning model are different, the terminal device (200) can approximate the sampling rate of the extracted acoustic characteristics to the sampling rate of the selected vocoder learning model (33). As an example, if the sampling rate of the extracted acoustic characteristics is 24 kHz and the sampling rate of the selected vocoder learning model is 16 kHz, the terminal device (200) can approximate the sampling rate of the extracted acoustic characteristics to 16 kHz.
[0070] Referring to FIG. 7b, the terminal device (200) may include a plurality of vocoder learning models (1). The terminal device (200) may identify a candidate group (2) of vocoder learning models based on the specifications of the terminal device (S410). The terminal device (200) may identify a candidate group (2) of available vocoder learning models from all vocoder learning models (1) included in the terminal device. Additionally, the terminal device (200) may monitor the resources of the terminal device (S420). The terminal device (200) may select an optimal vocoder learning model (3) from the identified candidate group based on the resources of the terminal device, etc. (S430). As an example, when no other app or process is running on the terminal device (200), a vocoder learning model with a high sampling rate may be selected. When other apps or processes are running and resources are low, a vocoder learning model with a low sampling rate may be selected. Alternatively, if memory usage is high, the terminal device (200) may select a vocoder learning model with low complexity and a low sampling rate.
[0071] The terminal device (200) can extract acoustic characteristics from text using a prosody model (41) and generate a speech waveform using the extracted acoustic characteristics and a selected vocoder learning model (3) included in the vocoder module (42). The terminal device (200) may include the same prosody model and various vocoder learning models. If the sampling rate of the acoustic characteristics extracted from the prosody model and the sampling rate of the selected vocoder learning model are different, the terminal device (200) can approximate the sampling rate of the extracted acoustic characteristics to the sampling rate of the selected vocoder learning model (43).
[0072] So far, we have explained the process of training various vocoder learning models and selecting the optimal vocoder learning model. Below, we describe the flowcharts of the electronic devices and terminal devices.
[0073] FIG. 8 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.
[0074] Referring to FIG. 8, the electronic device receives text input (S810) and identifies a first acoustic characteristic from the input text using a prosody module that extracts acoustic characteristics (S820).
[0075] The electronic device generates a modified acoustic characteristic with a sampling rate different from the first acoustic characteristic based on the first acoustic characteristic (S830). For example, the electronic device may generate the modified acoustic characteristic by downsampling the first acoustic characteristic. Alternatively, the electronic device may generate the modified acoustic characteristic by approximating the first acoustic characteristic to a preset acoustic characteristic.
[0076] The electronic device generates a plurality of vocoder learning models by training a vocoder module that generates a speech waveform based on each of the first acoustic characteristic and the modified acoustic characteristic (S840). For example, the electronic device can train the vocoder module based on the first modified acoustic characteristic approximated by a preset acoustic characteristic and the second modified acoustic characteristic generated by downsampling the first acoustic characteristic.
[0077] FIG. 9 is a flowchart illustrating a control method for a terminal device according to one embodiment of the present disclosure.
[0078] Referring to FIG. 9, the terminal device identifies the specifications of a component associated with the terminal device (S910). For example, the specifications of the component may include the resources of the processor, whether the processor is operating, memory capacity, memory bandwidth, the performance of the speaker, etc. The specifications of the component associated with the terminal device (200) described above may be internal specifications. Meanwhile, the terminal device (200) may be connected to an external speaker. In this case, the specifications of the component associated with the terminal device (200) may include various information of the external speaker.
[0079] The terminal device selects one vocoder learning model among a plurality of vocoder learning models based on the specifications of the identified components (S920). For example, the terminal device may identify candidate vocoder learning models based on the specifications of the internal components of the terminal device and whether streaming output of the voice waveform is possible. The terminal device may select one vocoder learning model among the candidate vocoder learning models based on the highest sampling rate and sound quality. Alternatively, the terminal device may select one vocoder learning model based on the resources of the processor. If an external speaker is connected to the terminal device, the terminal device may identify the specifications of the external speaker and select one vocoder learning model based on the identified specifications of the external speaker.
[0080] The terminal device identifies acoustic characteristics from text using a prosody module (S930), and generates a voice waveform corresponding to the identified acoustic characteristics using an identified vocoder learning model and outputs it to a speaker (S940).
[0081] The control method of an electronic device and the control method of a terminal device according to the various embodiments described above may be provided as a computer program product. The computer program product may include the S / W program itself or a non-transitory computer readable medium on which the S / W program is stored.
[0082] A non-transient readable medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, the various applications or programs described above may be stored and provided on non-transient readable media such as CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.
[0083] Furthermore, although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure. Explanation of the symbols
[0084] 100: Electronic device 110: Input interface 120: Processor 200: Terminal device 210: Processor 220: Speaker
Claims
Claim 1 An electronic device comprising: an input interface; and a processor including a prosody module for extracting acoustic characteristics and a vocoder module for generating a voice waveform; wherein the processor receives text through the input interface, identifies a first acoustic characteristic from the input text using the prosody module, generates a modified acoustic characteristic having a sampling rate different from the first acoustic characteristic based on the identified first acoustic characteristic, and generates a plurality of vocoder learning models by training the vocoder module based on each of the first acoustic characteristic and the modified acoustic characteristic, and wherein the processor generates the modified acoustic characteristic by approximating the first acoustic characteristic to a preset acoustic characteristic. Claim 2 In claim 1, the processor is an electronic device that generates the modified acoustic characteristic by downsampling the first acoustic characteristic. Claim 3 delete Claim 4 An electronic device according to claim 1, wherein the processor trains the vocoder module based on a first modified acoustic characteristic approximated to the preset acoustic characteristic and a second modified acoustic characteristic generated by downsampling the first acoustic characteristic. Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 A control method for an electronic device comprising: a step of receiving text input; a step of identifying a first acoustic characteristic from the input text using a prosody module that extracts acoustic characteristics; a step of generating a modified acoustic characteristic having a sampling rate different from the first acoustic characteristic based on the identified first acoustic characteristic; and a step of generating a plurality of vocoder learning models by training a vocoder module that generates a speech waveform based on each of the first acoustic characteristic and the modified acoustic characteristic, wherein the step of generating the modified acoustic characteristic approximates the first acoustic characteristic to a preset acoustic characteristic to generate the modified acoustic characteristic. Claim 12 In claim 11, the step of generating the modified acoustic characteristic is a method for controlling an electronic device, wherein the modified acoustic characteristic is generated by downsampling the first acoustic characteristic. Claim 13 delete Claim 14 In claim 11, the step of generating the plurality of vocoder learning models is a method for controlling an electronic device, wherein the vocoder module is trained based on a first modified acoustic characteristic approximated by the preset acoustic characteristic and a second modified acoustic characteristic generated by downsampling the first acoustic characteristic. Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete