Electronic device and control method thereof
By generating and recognizing candidate reference vectors for user speech, the complexity and performance uncertainty of personalized TTS services in existing technologies are resolved, enabling the provision of high-quality personalized TTS services without the need for extensive recording.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for providing personalized text-to-speech (TTS) services require the target speaker to record their speech for an extended period of time, and the performance of the reference vectors is uncertain, resulting in complex model training and poor performance.
By using the voice of a regular user, multiple candidate reference vectors are generated. The voice is synthesized using a TTS model, and the best synthesized voice is identified based on similarity and features and stored as the user's reference vector. This reduces the number of registered voice sentences and provides personalized TTS services.
It enables personalized TTS services without retraining the model for each user, reducing voice registration time and model training complexity, and improving service performance and user experience.
Smart Images

Figure CN116457872B_ABST
Abstract
Description
Technical Field
[0001] The apparatus and methods consistent with this disclosure relate to an electronic device and a method for controlling the same, and more specifically, to an electronic device and a method for controlling the same for providing text-to-speech (TTS) services. Background Technology
[0002] Text-to-speech (TTS) refers to speech synthesis technology that uses machines to synthesize (or convert) text into human speech.
[0003] To provide speech with a style similar to that of the target speaker (e.g., tone, accent, delivery speed, pronunciation, intonation, and speaking habits) via TTS services, a process is needed to pre-record the speech of the target speaker and process the recorded speech data. To achieve natural speech with a style similar to that of the target speaker for various texts, model learning methods such as model adaptation and delivery learning based on the target speaker's spoken speech of two hundred or more sentences (or one hour or more of scripts) are required.
[0004] There are difficulties in providing personalized TTS services, where the voice of a general user is used as the TTS service's audio. This is because the target speaker needs to speak a large number of sentences with precise pronunciation over a long period to register the target speaker's voice as the audio for the TTS service as described above. Meanwhile, there is a method to obtain reference vectors from the target speaker's voice and input the text and reference vectors into a TTS model to obtain synthesized voice with the target speaker's speech characteristics to provide personalized TTS services. In this case, there is the advantage that the model may not need to be trained (zero / few training iterations), but there is the disadvantage that it may not know whether the reference vectors have optimal performance (voice quality / prosody / pronunciation / speaker similarity, etc.). Summary of the Invention
[0005] This disclosure provides an electronic device and a control method thereof for providing text-to-speech (TTS) services using the voice of a general user.
[0006] According to embodiments of this disclosure, an electronic device includes a microphone, a memory storing a TTS model and a plurality of evaluation texts, and a processor. The processor is configured to: when receiving user speech through the microphone, obtain a reference vector of the user's spoken speech; generate a plurality of candidate reference vectors based on the reference vectors; obtain a plurality of synthesized voices by inputting the plurality of candidate reference vectors and the plurality of evaluation texts into the TTS model; identify at least one of the plurality of synthesized voices based on the similarity between the plurality of synthesized voices and the user's speech, and the features of the plurality of synthesized voices; and store the reference vector of the at least one synthesized voice in the memory as a reference vector of the TTS model corresponding to the user.
[0007] According to another embodiment of this disclosure, a control method for an electronic device including a memory stores a TTS model and multiple evaluation texts in the memory. The control method includes: when receiving user speech through a microphone, obtaining a reference vector of the user's spoken speech; generating multiple candidate reference vectors based on the reference vectors; obtaining multiple synthesized voices by inputting the multiple candidate reference vectors and multiple evaluation texts into the TTS model; identifying at least one synthesized voice among the multiple synthesized voices based on the similarity between the multiple synthesized voices and the user speech and the features of the multiple synthesized voices; and storing the reference vector of the at least one synthesized voice in the memory as a reference vector for the TTS model corresponding to the user.
[0008] According to different embodiments of this disclosure, an electronic device and its control method for providing TTS services using the voice of a general user can be provided.
[0009] Furthermore, according to embodiments of this disclosure, the system can be configured to minimize the number of sentences spoken during the request to register for the TTS service, thus achieving optimal performance. Additionally, according to embodiments of this disclosure, personalized TTS services can be provided via a user's voice without retraining the TTS model for each user. Attached Figure Description
[0010] Figure 1 These are diagrams used to describe electronic devices according to embodiments of the present disclosure;
[0011] Figure 2 It is a block diagram used to describe the operation of each component of an electronic device according to embodiments of the present disclosure;
[0012] Figure 3 It is a block diagram used to describe the operation of each component of an electronic device according to embodiments of the present disclosure;
[0013] Figure 4It is a block diagram used to describe the operation of each component of an electronic device according to embodiments of the present disclosure;
[0014] Figure 5 This is a diagram illustrating a method for obtaining a reference vector according to embodiments of the present disclosure;
[0015] Figure 6A is a diagram illustrating a method for generating candidate reference vectors according to an embodiment of the present disclosure;
[0016] Figure 6B is a diagram illustrating a method for generating candidate reference vectors according to an embodiment of the present disclosure;
[0017] Figure 6C is a diagram illustrating a method for generating candidate reference vectors according to an embodiment of the present disclosure;
[0018] Figure 7 This is a diagram illustrating a text-to-speech (TTS) model according to embodiments of the present disclosure;
[0019] Figure 8A is a diagram illustrating a method for recognizing synthesized sound according to an embodiment of the present disclosure;
[0020] Figure 8B is a diagram illustrating a method for recognizing synthesized sound according to an embodiment of the present disclosure;
[0021] Figure 8C is a diagram illustrating a method for recognizing synthesized sound according to an embodiment of the present disclosure;
[0022] Figure 8D is a diagram illustrating a method for recognizing synthesized sound according to an embodiment of the present disclosure;
[0023] Figure 9A is a diagram illustrating the hardware components of an electronic device according to an embodiment of the present disclosure;
[0024] Figure 9B is a diagram illustrating additional hardware components of an electronic device according to an embodiment of the present disclosure; and
[0025] Figure 10 This is a flowchart illustrating a control method for an electronic device according to an embodiment of the present invention. Detailed Implementation
[0026] In describing this disclosure, detailed descriptions of known functions or configurations related to this disclosure will be omitted where it is determined that such detailed descriptions might unnecessarily obscure the essential points of this disclosure. Furthermore, the following embodiments can be modified in many different forms, and the scope and spirit of this disclosure are not limited to the following embodiments. Rather, these embodiments make this disclosure thorough and complete, and are provided to fully transfer the technical spirit of this disclosure to those skilled in the art.
[0027] It should be understood that the techniques mentioned in this disclosure are not limited to the specific embodiments, but include all modifications, equivalents, and / or substitutions of the embodiments according to this disclosure. In all the drawings, similar components will be indicated by similar reference numerals.
[0028] The terms “first” or “second” used in this disclosure may refer to various components, regardless of the order and / or importance of the components, and are only used to distinguish one component from others, and do not limit the scope of these components.
[0029] In this disclosure, the expressions “A or B”, “at least one of A and / or B”, or “one or more of A and / or B” can include all possible combinations of the items listed together. For example, “A or B”, “at least one of A and B”, or “at least one of A or B” can mean all of the following: 1) including at least one A, 2) including at least one B, or 3) including both at least one A and at least one B.
[0030] In this disclosure, the singular form includes the plural form unless the context clearly indicates otherwise. It should be understood that the terms "comprising" or "forming of" as used in the specification specify the presence of a feature, number, step, operation, part, component, or combination thereof mentioned in the specification, but do not exclude the presence or addition of one or more other features, numbers, steps, operations, parts, components, or combinations thereof.
[0031] When it is said that any component (e.g., a first component) is (operably or communicatively) coupled to or connected to another component (e.g., a second component), it should be understood that any component is directly coupled to the other component or can be coupled to the other component through the other component (e.g., a third component). On the other hand, when it is said that any component (e.g., a first component) is "directly coupled" or "directly connected" to another component (e.g., a second component), it should be understood that the other component (e.g., a third component) does not exist between any component and the other component.
[0032] The expression “configured (or set) to” as used in this disclosure may be replaced by “suitable,” “capable,” “designed to,” “appropriate,” “made,” or “able” depending on the context. The term “configured (or set) to” may not necessarily mean “specifically designed for” in hardware. Rather, in some cases, the expression “configured to” may mean that the device can “do” with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may refer to a dedicated processor (e.g., an embedded processor) for performing these operations, or a general-purpose processor (e.g., a central processing unit (CPU) or application processor) that can perform these operations by executing one or more software programs stored in a memory device.
[0033] Figure 1 This is a diagram used to describe an electronic device according to embodiments of the present disclosure.
[0034] Reference Figure 1 The electronic device 100 according to embodiments of the present disclosure can be implemented as an interactive system.
[0035] Here, electronic device 100 may include at least one of the following: smartphone, tablet PC, mobile phone, video phone, e-book reader, desktop PC, laptop PC, network book computer, workstation, server, personal digital assistant (PDA), portable multimedia player (PMP), MP3 player, mobile medical device, camera, wearable device, or robot. According to different embodiments, the wearable device may be implemented as an accessory-type wearable device (e.g., watch, ring, bracelet, anklet, necklace, glasses, contact lens, or head-mounted device (HMD)), a textile or clothing-integrated wearable device (e.g., electronic clothing), a body-attached wearable device (e.g., skin pad or tattoo), or a bio-implantable wearable device (e.g., implantable circuitry). However, this is merely an example, and electronic device 100 is not limited thereto, and may be implemented as electronic devices with various shapes and purposes.
[0036] An interactive system is a system that can interact with users through dialogue, such as capturing the user's intention through voice and outputting a response corresponding to the user's intention.
[0037] As a specific embodiment, the electronic device 100 may include an automatic speech recognition (ASR) module 1, a natural language processing (NLP) module 2, and a text-to-speech (TTS) module 3. Furthermore, the electronic device 100 may include a microphone 110 for receiving user speech and an output interface 140 for outputting information in response to user speech. For example, the output interface 140 may include a speaker for outputting sound.
[0038] ASR module 1 can use a language model and an acoustic model to convert speech signals (i.e., user speech) received through microphone 110 into text (strings), such as words or phoneme sequences. The language model can be a model that assigns probabilities to words or phoneme sequences, and the acoustic model can be a model that indicates the relationship between the speech signal and the text containing the speech signal. These models can be configured based on probabilistic statistics or artificial neural networks.
[0039] NLP Module 2 can use various analysis methods (e.g., morphological analysis, syntactic analysis, and semantic analysis of text corresponding to user speech) to identify the meaning of words or sentences in the text corresponding to user speech relative to the text corresponding to user speech. Based on the identified meaning, it can grasp the user's intent and obtain response information corresponding to the user's intent. In this case, the response information can be in the form of text.
[0040] The TTS module 3 can convert text into speech signals and output the speech signals through the output interface 140 (e.g., a speaker). In other words, the response information obtained by the NLP module 2 can be converted from text to speech signals through the TTS module 3.
[0041] Meanwhile, the electronic device 100 according to embodiments of this disclosure can provide personalized TTS services. Personalized TTS services refer to the service of converting (or synthesizing) text into speech signals using the voice of a single user (or another user) through the TTS module 3. For this purpose, it may be necessary to pre-register the user's voice in the electronic device 100. The electronic device 100 according to this disclosure can minimize the number of sentences used to request the user to register the TTS service. Furthermore, according to embodiments of this disclosure, personalized TTS services can be provided using the user's voice without retraining the TTS model for each user. Specific details will be described with reference to the accompanying drawings.
[0042] Figure 2 and Figure 3 This is a block diagram used to describe the operation of each component of an electronic device according to embodiments of the present disclosure. Figure 3 The operation is shown when a reference vector that satisfies at least one evaluation criterion for the evaluation text does not exist.
[0043] Reference Figure 2 An electronic device 100 according to an embodiment of the present disclosure may include a microphone 110, a memory 120, and a processor 130.
[0044] Microphone 110 can receive user voice.
[0045] The memory 120 can store multiple evaluation texts. For example, multiple evaluation texts can be stored in an evaluation text database 123 within the memory 120. The unit of evaluation text can be a sentence, but this is just an example, and the unit of evaluation text can be modified differently.
[0046] Furthermore, memory 120 can store reference vectors of users registered as speakers in TTS module 30. The reference vectors of registered users can be stored in reference vector storage module 124 within memory 120. The reference vector of a registered user can indicate the reference vector that best matches the unique speech feature of the registered user.
[0047] The processor 130 can extract the optimal reference vector from the user's speech and register the extracted optimal reference vector as the user's reference vector to synthesize text into the user's speech.
[0048] To this end, processor 130 can execute instructions included in each of the speaker encoder module 10, candidate reference vector generation module 20, TTS module 30, and synthesized voice evaluation module 40 to perform operations corresponding to each instruction. Here, speaker encoder module 10, candidate reference vector generation module 20, TTS module 30, and synthesized voice evaluation module 40 can be stored in memory 120 or in memory within processor 130.
[0049] Specifically, when user A's voice is received through microphone 110, processor 130 can obtain a reference vector of the user's voice from the user's voice received through speaker encoder module 10.
[0050] For example, when a request for user registration with TTS module 30 is received from user A (e.g., in the form of touch input, voice command, etc.), processor 130 can provide reference text (r) set to be spoken by user A. Subsequently, when user speech spoken by user A is received via microphone 110, processor 130 can obtain a reference vector from the user speech received via speaker encoder module 10. However, this is merely an example, and natural language can also be recognized as reference text when user A speaks natural language without providing the set reference text.
[0051] Here, the reference vector (RV)(s) r,A ) is defined as the user speech (x) that appears when user A (speaker A) has spoken the reference text r. r,AThe reference text is a multidimensional (i.e., two-dimensional or more) vector (or column of vectors) of speech features. Each dimension (or column) of the reference vector can indicate speech features such as prosody, pronunciation, frequency band, speaker's age, and speaker's gender. The reference text refers to the sentence (or word, etc.) spoken by the user and can be assigned a domain based on the speaking style (e.g., reading style, conversational style, or news style).
[0052] Reference Figure 5 The analog sound signal received from microphone 110 can be converted into a digital sound signal using an analog-to-digital converter (ADC). Here, the sound signal may include user speech (x) of user A who has spoken the reference text (r). r,A Meanwhile, the ADC can be implemented by embedding it into the microphone 110 or the processor 130, or it can be implemented as a standalone device. That is to say, various modifications can be made to the ADC.
[0053] In this case, the processor 130 can identify the user's speech (x) from the acoustic signal based on the energy level. r ,A ( ) part of the acoustic signal.
[0054] In addition, the processor 130 can output the corresponding user voice (x r,A The acoustic signal is divided into frames (e.g., 20ms and 40ms), and a Fourier transform is applied to each frame to calculate the spectrum. Here, the acoustic signal can represent a waveform in the time domain, such as amplitude (or sound pressure) varying over time, and the spectrum can represent a waveform in the frequency domain, such as amplitude (or sound pressure) according to frequency. For example, the acoustic signal can represent a waveform with the horizontal axis being time and the vertical axis being amplitude, and the spectrum can represent a waveform with the horizontal axis being frequency and the vertical axis being amplitude. In this case, the spectrum can be the spectrum in the general frequency domain, or it can be various types of spectra, such as the mel spectrum obtained by applying a filter bank based on a mel scale, which indicates the relationship between frequencies that are perceptibly perceived by humans and a spectrogram (SPG) having a relationship between the frequency axis and the amplitude axis. Furthermore, the spectrum can be a cepstrum or mel cepstrum from the spectrum transform, and can include pitch lag or pitch correlation with pitch / harmonic information. However, this is only one example, and the spectrum can be various acoustic feature vectors with speech characteristics.
[0055] Furthermore, processor 130 can obtain a reference vector from the spectrum. As an example, processor 130 can obtain the reference vector (s) by inputting the mel spectrum into the speaker recognizer. r,AHowever, this is merely an example, and the processor 130 can use one of various algorithms (or neural networks) such as cepstral calculations, linear prediction coefficients (LPC) and filter bank energy, Wav2Vec, SincNet, and PASE to obtain the reference vector. r,A In this case, the obtained reference vector (s) r,A () can be various types of vectors, such as vector i, vector d, and vector x.
[0056] Meanwhile, we will refer to later Figure 5 Describes a specific method for obtaining a reference vector through the speaker encoder module 10.
[0057] Furthermore, the processor 130 can base its work on the reference vectors generated by the candidate reference vector generation module 20. r,A To generate multiple candidate reference vectors
[0058] Here, the multiple candidate reference vectors may include one of the following: as a first embodiment, a reference vector randomly selected based on a reference vector; as a second embodiment, a reference vector generated based on a reference vector and a reference vector used to train the TTS module 30; and as a third embodiment, a reference vector generated by applying a masking vector to a reference vector; or a combination thereof. Specific details will be described with reference to Figures 6A to 6C.
[0059] Figures 6A to 6C are diagrams illustrating a method for generating candidate reference vectors according to embodiments of the present disclosure. Figures 6A to 6C represent reference vectors on planes 610, 620, and 630, and imply that the closer the reference vectors are to each other on planes 610, 620, and 630, the more similar their characteristics are.
[0060] Referring to Figure 6A, as a first embodiment of this disclosure, a plurality of candidate reference vectors 612(S) r,A ) can include reference vector 611 (S r,A At least one reference vector is randomly selected.
[0061] For example, processor 130 can add noise to reference vector 611(s) according to the following formula (1). r,A To generate at least one candidate reference vector 612
[0062]
[0063] Here, noise can be random values following a normal distribution, a uniform distribution, or various probability distributions, and can be a reference vector (s). r,A The value of at least one of the dimensions of ).
[0064] Furthermore, the noise can have values within a predetermined range. In this case, as shown in Figure 6A, the candidate reference vector 612 It can be based on reference vector 611(s) r,A A vector that exists within a predetermined radius.
[0065] Meanwhile, referring to FIG6B, as a second embodiment of this disclosure, a plurality of candidate reference vectors 625(S) r,A ) can include reference vector 621 (S r,A ) and at least one reference vector generated from the reference vector used to train the TTS module 30.
[0066] For example, processor 130 can compare reference vectors (s r,A ) and reference vectors of multiple speakers used for mutual training of the TTS module 30 (e.g., reference vectors of speaker B). B The reference vector s of speaker C C The distance between reference vectors of multiple speakers (e.g., reference vector s of speaker B) is used to identify the reference vector that has the closest distance to the reference vector among the reference vectors of multiple speakers. B ).
[0067] Furthermore, according to the following formula (2), the processor 130 can adjust the reference vector (s) r,A ) and reference vector 623 (e.g., speaker B's reference vector s) B Interpolation is applied to generate at least one candidate reference vector. The reference vector 623 has the closest distance among the reference vectors used to train the TTS module 30:
[0068]
[0069] Here, w a and w b This indicates that the candidate reference vector Positioning at the intersection of two reference vectors (s) r,A and s B The arbitrary coefficients of a function (e.g., a linear or quadratic function). For example, in the case of a linear function, w a It can be 0.9 and w b It can be 0.1, w a It can be 0.8 and w b It can be 0.2, or w a It can be 0.7 and w b It can be 0.3.
[0070] Meanwhile, as a third embodiment of this disclosure, multiple candidate reference vectors (S) r,A This can include applying the masking vector to the reference vector (s) r,A At least one reference vector is generated.
[0071] For example, according to formula (3), processor 130 can apply the masking vector W to the reference vector (s) r,A To generate at least one candidate reference vector
[0072]
[0073] Here, W is the masking vector, and can be configured as a matrix with values of 0 or 1, or values between 0 and 1, or various values.
[0074] Meanwhile, according to embodiments of this disclosure, multiple candidate reference vectors (S) r,A The candidate reference vector (S) can be at least one combination of the first to third embodiments. That is, multiple candidate reference vectors (S) r,A ) can include reference vectors (s) r,A At least one reference vector (s) is randomly selected. r,A ), based on reference vector (s) r,A ) and at least one reference vector generated from the reference vector used to train the TTS module 30, or by applying a masking vector to the reference vector (s r,A At least one of the at least one reference vectors generated. For example, a combination of the first and second embodiments can appear as shown in FIG6C.
[0075] Furthermore, the processor 130 can input multiple candidate reference vectors stored in the memory 120 into the TTS module 30. and multiple evaluation texts (T = t1, ..., t) M To obtain multiple synthesized sounds
[0076] Specifically, based on multiple candidate reference vectors For each of the multiple candidate reference vectors, processor 130 can... and multiple evaluation texts (T = t1, ..., t) M The input is fed into the TTS module 30 to obtain multiple evaluation texts (T = t1, ..., t2). M Each of the multiple synthesized sounds generated in ) Here, candidate reference vectors are synthesized. and evaluation text (t) m To generate the synthesized sound In this case, when the number of candidate reference vectors is N and the number of evaluation texts is M, the number of synthesized voices can be N×M.
[0077] Here, multiple evaluation texts (T = t1, ..., t) are used. M The evaluation text can include at least one evaluation text belonging to each of multiple domains (e.g., reading style, conversational style, and news style). That is, a domain can be assigned to multiple evaluation texts (T = t1, ..., t2). M Each of them.
[0078] For example, depending on the style of the text, a domain can include types such as reading style, conversational style, interrogative sentences, and exclamatory sentences, and can include types such as chatbots based on text content, natural language generation (NLG), news, dictionaries, wikis, music (song titles, artists, and lyrics), home appliances (e.g., air conditioners), big data (web crawling content), children's books, and novels. However, this is just an example; domains are not limited to this and can be modified in various ways.
[0079] As one embodiment, the TTS module 30 may include an acoustic model (AM) and a speech decoder (speech encoder). Reference will be made later. Figure 7 A detailed description of TTS module 30.
[0080] Acoustic models can use at least one of various algorithms, such as Tacotron, Tacotron 2, Transformer, text2mel, and Deep Convolutional TTS (DCTTS), to convert text and reference vectors into acoustic features. In this case, the resulting acoustic features can have reference vectors, i.e., features of the corresponding speaker (e.g., pitch, intonation, intensity, and pronunciation). Here, acoustic features can indicate unique characteristics (e.g., pitch, intonation, intensity, and pronunciation) of sound within a speech segment (e.g., frame unit or sentence unit). For example, acoustic features can be implemented as one or a combination of spectrograms, mel spectrograms, cepstrum, pitch lag, pitch correlation, mel frequency cepstrum modulation energy (MCME), mel frequency cepstrum coefficients (MFCC), etc., where waveforms and spectra are combined with each other.
[0081] Voice encoders can generate synthetic sound by synthesizing reference vectors and acoustic features using various algorithms such as WaveNet, Parallel WaveNet, WaveGlow, WaveRNN, and LPCNet. For example, a voice encoder can be a neural network-based artificial intelligence model that is learned to output synthetic sound when given acoustic features such as mel spectrum and reference vectors.
[0082] In addition, based on multiple synthesized sounds Similarity to user speech, and multiple synthesized voices The processor 130 can identify multiple synthesized sounds through the synthesized sound evaluation module 40, based on the characteristics of the synthesized sound. At least one of them.
[0083] Specifically, the processor 130 can identify candidate synthesized sounds, which are compared with multiple synthesized sounds. User voice (x) r,A The similarity (i.e., speaker similarity) is a predetermined threshold or greater. The specific implementation of this will be described with reference to Figures 8A and 8B. Furthermore, the processor 130 can identify at least one candidate synthesized voice based on at least one of the prosody, pronunciation, or voice quality of each candidate synthesized voice. The specific implementation of this will be described below with reference to Figures 8C and 8D.
[0084] Figures 8A to 8D are diagrams illustrating a method for recognizing synthesized sounds according to embodiments of the present disclosure.
[0085] Referring to Figure 8A, as an example, when multiple synthesized sounds are combined... When each input is fed into the speaker encoder module 10, the processor 130 can obtain multiple synthesized sounds for output from the speaker encoder module 10. reference vector In addition, the processor 130 can synthesize multiple sounds Each reference vector With user voice (x) r,A The reference vector (s) r,A The similarity is determined by comparing user voice (x) with other voices. Here, when user voice (x) is compared with other voices, the similarity is determined by comparing user voice (x) with other voices. r,A When the signal is input to the speaker encoder module 10, it is used for the user's speech (x) r,A The reference vector (s) r,A The output is obtained from the speaker encoder module 10.
[0086] Here, similarity can be calculated using various methods such as Euclidean distance and cosine similarity. Furthermore, the similarity between reference vectors can be considered as the speaker similarity between synthesized voices. That is, the processor 130 can identify synthesized voices with reference vectors, which have characteristics for multiple synthesized voices. reference vector The similarity is at or above a predetermined threshold.
[0087] At the same time, multiple synthesized sounds can be used. Classified as a tool for generating multiple synthesized sounds Candidate reference vector The unit group. For example, through the first candidate reference vector. and the first to the Mth evaluation texts (T = t1, ..., t M The synthesized sound generated They can be classified into the same group. That is, synthesized sounds generated from a candidate reference vector and M evaluation texts can be classified into the same group.
[0088] In this scenario, processor 130 can identify multiple synthesized sounds in groups. reference vector The processor 130 can identify the reference vector of the group with the smallest deviation. In this case, the processor 130 can identify the synthesized sound synthesized from the reference vector of the group with the smallest deviation as a candidate synthesized sound.
[0089] For example, as shown in plane 810 of Figure 8A, when passing through the first candidate reference vector and the first to the Mth evaluation texts (T = t1, ..., t M The synthesized sound generated Classified as Group 1 Through the second candidate reference vector and the first to the Mth evaluation texts (T = t1, ..., t M The synthesized sound generated Classified as Group 2 and through the third candidate reference vector and the first to the Mth evaluation texts (T = t1, ..., t M The synthesized sound generated Classified as Group 3 At that time, we can assume the location of the first to third groups. In this case, the processor 130 can base its location on the user's voice (x r,A The reference vector (s) r,A To identify the third group with the smallest bias reference vector In this case, processor 130 can be connected via the third group reference vector Synthesized sounds were identified as candidate synthesized sounds.
[0090] Meanwhile, as shown in FIG8B, the processor 130 according to an embodiment of the present disclosure may use a plurality of speaker encoder modules 10-1 to 10-3 to identify candidate synthesized voices. For example, the first to third speaker encoder modules 10-1 to 10-3 may be modeled to output different types of reference vectors (e.g., i vector, d vector, x vector, etc.).
[0091] When multiple synthesized sounds and user voice (x r,A When input is given to the first speaker encoder module 10-1, the processor 130 can obtain multiple synthesized sounds for output from the first speaker encoder module 10-1. First reference vector and for user voice (x r,A The first reference vector (s) r ,A ), and the first reference vector and the first reference vector (s) r,A They are compared with each other to determine the first similarity.
[0092] Furthermore, when multiple synthesized sounds are combined and user voice (x r,A When input is given to the second speaker encoder module 10-2, the processor 130 can obtain multiple synthesized sounds for output from the second speaker encoder module 10-2. Second reference vector and for user voice (x r,A The second reference vector (i) r,A ), and the second reference vector Second reference vector (i) r,A They are compared to each other to determine the second similarity.
[0093] Furthermore, when multiple synthesized sounds are combined and user voice (x r,A When input is given to the third speaker encoder module 10-3, the processor 130 can obtain multiple synthesized sounds for output from the third speaker encoder module 10-3. The third reference vector and for user voice (x r,A The third reference vector (t) r,A ), and the third reference vector With the third reference vector (t) r,A They are compared with each other to determine the third similarity.
[0094] Furthermore, the processor 130 can identify the reference vector that has the best performance among the first to third similarity levels, and can identify the synthesized sound synthesized using the identified reference vector as a candidate synthesized sound. The reference vector with the best performance can be the vector with the smallest deviation value or the vector that exceeds a predetermined threshold for each of the first to third similarity levels.
[0095] Furthermore, the processor 130 can identify at least one of the candidate synthesized sounds based on at least one of the prosody, pronunciation, or sound quality of each candidate synthesized sound. That is, the processor 130 can identify a synthesized sound by one or a combination of the prosody, pronunciation, and sound quality of each candidate synthesized sound.
[0096] Specifically, the processor 130 can calculate the prosody score, pronunciation score, and sound quality score of each candidate synthesized voice, and identify at least one synthesized voice from the candidate synthesized voices, whose prosody score, pronunciation score, and sound quality score are at or above a predetermined threshold.
[0097] As an example, processor 130 can calculate a prosodic score for each candidate synthesized voice and identify at least one synthesized voice whose prosodic score is at or above a predetermined threshold. For instance, processor 130 can calculate the prosodic score by comparing length, speaking speed, and pitch rise / fall curves, where the pitch rise / fall curve represents the evaluation text t. m Each phoneme set in the text is evaluated based on its pitch over time, targeting the text t. m Corresponding candidate synthesized voices For each phoneme, evaluate the text t m It features length, speaking speed, and pitch curves.
[0098] As an example, processor 130 can calculate the pronunciation score of each candidate synthesized voice and identify at least one synthesized voice whose pronunciation score is at or above a predetermined threshold.
[0099] As shown in Figure 8C(1), the processor 130 according to an embodiment of the present disclosure can synthesize sound. Input is given to ASR module 1 to obtain text, and the obtained text is then compared with the corresponding synthesized sound. Evaluation text t m The comparisons are used to calculate pronunciation scores. As mentioned above, ASR module 1 can use various algorithms to analyze speech and convert the speech content into text format.
[0100] As shown in Figure 8C(2), the processor 130 according to an embodiment of the present disclosure can synthesize sound. The input is fed into the forced alignment module 45 to identify the boundaries between speech-phonemes and speech-words, and the identified boundaries are matched with the corresponding synthesized sounds. Evaluation text t m The boundaries between phonemes and phonemes in the text are compared to calculate the likelihood. This likelihood may then be used as a pronunciation score.
[0101] As an example, processor 130 can calculate the sound quality score for each candidate synthesized sound and identify at least one synthesized sound whose sound quality score is at or above a predetermined threshold.
[0102] For example, processor 130 can calculate synthesized sound using various algorithms, such as signal-to-noise ratio (SNR), harmonic noise ratio (HNR), and sound quality assessment through spatial characteristic estimation. The sound quality score.
[0103] Simultaneously, the processor 130 can segment multiple candidate synthesized sounds based on the domain to which each of the multiple candidate synthesized sounds belongs (i.e., the domain of the evaluation text used to generate the corresponding synthesized sound). The processor 130 can identify at least one synthesized sound in each domain based on at least one of the similarity, prosody, pronunciation, or sound quality of each of the one or more candidate synthesized sounds belonging to each domain.
[0104] Specifically, the synthesized sound evaluation module 40, based on multiple synthesized sounds, will be described with reference to Figure 8. Similarity to user speech and multiple synthesized voices The implementation of selecting the optimal reference vector is based on the characteristics of the vector.
[0105] In tables (1) to (4) of Figure 8D, each row indicates the evaluation text (t1, t2, t3, ...) assigned to each domain, and each column indicates the candidate reference vector. The combination of each row and column indicates the score of the synthesized voice generated based on the combination of the evaluation text and candidate reference vectors (speaker similarity, prosody score, pronunciation score, voice quality score, etc.).
[0106] As an example, as shown in (1) of FIG8D, when the speaker similarity (i.e., the value in the same column) of all synthesized voices of one candidate reference vector and multiple evaluation texts combined with each other is a predetermined value (e.g., 60) or higher, the corresponding candidate reference vector can be determined to meet the evaluation criteria of speaker similarity, and when at least one speaker similarity is less than the predetermined value (e.g., 60), the corresponding candidate reference vector can be determined not to meet the evaluation criteria of speaker similarity.
[0107] As an example, as shown in (2) of FIG8D, when at least one of the prosodic scores of a plurality of synthesized voices in which one of the candidate reference vectors and the plurality of evaluation texts are combined is a predetermined value (e.g., 80) or greater, the corresponding candidate reference vector can be determined to meet the prosodic score evaluation criteria, and when all the prosodic scores of the plurality of synthesized voices are less than the predetermined value (e.g., 80), the corresponding candidate reference vector can be determined not to meet the prosodic score evaluation criteria.
[0108] As an example, as shown in (3) of FIG8D, when at least one of the pronunciation scores of a plurality of synthesized voices that are a combination of one of the candidate reference vectors and a plurality of evaluation texts is a predetermined value (e.g., 90) or greater, the corresponding candidate reference vector can be determined to meet the evaluation criteria for pronunciation scores, and when all the pronunciation scores of the plurality of synthesized voices are less than the predetermined value (e.g., 90), the corresponding candidate reference vector can be determined to not meet the evaluation criteria for pronunciation scores.
[0109] As an example, as shown in (4) of Figure 8D, when all the sound quality scores (i.e., values in the same column) of one of the candidate reference vectors and multiple synthesized sounds combined with multiple evaluation texts are all predetermined values (e.g., 80) or greater, the corresponding candidate reference vector can be determined to meet the sound quality score evaluation criteria, and when at least one sound quality score is less than the predetermined value (e.g., 80), the corresponding candidate reference vector can be determined to not meet the sound quality score evaluation criteria.
[0110] Furthermore, the processor 130 can store the reference vector of at least one synthesized sound identified in the memory 120. As a reference vector corresponding to user A in TTS module 30 In other words, a reference vector that meets the evaluation criteria among multiple candidate reference vectors can be registered as a reference vector of user A and stored in the reference vector storage module 124 of memory 120.
[0111] As described above, the electronic device 100 according to embodiments of this disclosure can obtain reference vectors optimized for various texts by utilizing the fact that the reference vectors of the same user have a distribution within a predetermined range, even if the user speaks a very small amount of text (e.g., 1 to 5 texts). That is, unlike the prior art, the electronic device 100 can ensure good performance through synthesized voice evaluation, and can obtain multiple reference vectors from the spoken text even if the user speaks the text only once. Because the number of texts used to provide personalized TTS services is very small, the convenience for users to register for personalized TTS services can be improved.
[0112] Meanwhile, the electronic device 100 according to embodiments of the present disclosure can provide feedback to user A when the user's voice spoken by user A during the process of registering the reference vector as user A's reference vector is insufficient to provide personalized TTS service.
[0113] Taking Figure 8D as an example, for all evaluation texts, the reference vectors of candidate synthesized voices that meet a predetermined value (e.g., 60) or more of speaker similarity can be identified as... For at least one evaluation text, the reference vector of a candidate synthesized voice that satisfies a predetermined value (e.g., 80) or more of a prosodic score can be identified as: For at least one evaluation text, the reference vector of a candidate synthesized voice that meets a predetermined value (e.g., 90) or more of the pronunciation score can be identified as: And for all evaluation texts, candidate reference vectors for candidate synthesized sounds that meet a predetermined sound quality score (e.g., 80) or higher can be identified as...
[0114]
[0115] In this case, through the synthesized sound evaluation module 40, the processor 130 can store the reference vectors of the identified candidate synthesized sounds that meet all evaluation criteria in the reference vector storage module 124 of the memory 120. As a reference vector corresponding to user A.
[0116] Furthermore, the processor 130 can divide multiple candidate synthesized sounds based on the domain to which each candidate synthesized sound belongs. Here, candidate synthesized sounds Reference vector It is generated by combining multiple evaluation texts (t1, t2, t3, ...) and candidate synthesized voices. The domain to which it belongs can be the domain assigned to the evaluation text (t1, t2, t3, ...) used to generate candidate synthesized sounds.
[0117] Furthermore, the processor 130 can identify at least one synthesized voice for each domain based on at least one of speaker similarity, prosody, pronunciation, or voice quality of each of one or more candidate synthesized voices belonging to each domain. Additionally, a reference vector of the identified at least one synthesized voice can be stored in memory 120 according to the domain to which each evaluated text belongs.
[0118] Specifically, processor 130 can determine whether a synthesized voice exists for a particular domain that meets at least one of the evaluation criteria (e.g., speaker similarity, prosody, pronunciation, or voice quality).
[0119] For example, as shown in Figure 8D, the processor 130 can evaluate the text t1 and the reference vector. The candidate synthesized voice generated by the combination of these factors is identified as a candidate synthesized voice whose prosodic score and pronunciation score meet predetermined values. Furthermore, the processor 130 can further identify candidate synthesized voices based on the evaluation text t2 and the reference vector. The candidate synthesized voice generated by the combination is identified as a candidate synthesized voice whose prosodic score and pronunciation score meet predetermined values. In this case, the processor 130 can use the reference vector that satisfies the prosodic score and pronunciation score of the evaluated text t1. The evaluation (selection) is a reference vector that can cover the domain of the evaluated text t1. Furthermore, the processor 130 can use reference vectors that satisfy the prosodic score and pronunciation score of the evaluated text t2. The evaluation (selection) is a reference vector that can cover the domain of the evaluation text t2.
[0120] Reference Figure 3 The processor 130 can control the output interface 140 (see FIG. 9B) so that when at least one synthesized voice that meets the evaluation criteria (e.g., at least one of speaker similarity, prosody, pronunciation or voice quality) does not exist in a particular domain, the synthesized voice evaluation module 40 outputs information requesting the utterance of a sentence (r') belonging to the particular domain.
[0121] For example, as shown in Figure 8D, when a synthesized sound (or reference vector) satisfying the prosody score and pronunciation score of the evaluation text t3 does not exist, the processor 130 can provide feedback to the user with a sentence (r') belonging to the domain assigned to the evaluation text t3. Here, the feedback sentence (r') can include sentences, words, etc., that encourage the user to speak to cover the domain of the evaluation text t3. For example, if the evaluation text t3 is a news domain, the feedback sentence (r') can be either t3 or the news domain text.
[0122] For example, processor 130 can determine at least one candidate synthesized sound belonging to a domain in which at least one synthesized sound is absent in a plurality of domains, and determine features of the synthesized sound, wherein relatively low scores for prosody, pronunciation, and sound quality are calculated based on prosodic scores, pronunciation scores, and sound quality scores calculated for the determined candidate synthesized sound. Processor 130 can output speech of a sentence generated based on the determined features, requested to be spoken, via speaker 141.
[0123] As described above, the electronic device 100 according to this disclosure can perform evaluations based on various evaluation criteria during the process of registering a user's voice as the voice of the TTS module 30. Therefore, the reference vector with optimal performance can be determined as the user's reference vector. Furthermore, when using only the user's spoken voice is insufficient to provide a personalized TTS service, a reference vector that can cover various types of text can be obtained by providing feedback to the user.
[0124] Simultaneously, the electronic device 100 can synthesize a speech signal using the registered user voice after the user's voice is registered in the TTS module 30. This will refer to... Figure 4 Detailed description.
[0125] Figure 4 This is a block diagram used to describe the operation of each component of an electronic device according to embodiments of the present disclosure. Figure 4 The process of synthesizing a speech signal using the user's speech after registering the user's speech in the TTS module 30 is shown.
[0126] Reference Figure 4 Suppose that input data 15 (e.g., text t) is provided to processor 130. Input data 15 may be text obtained as a result of performing speech recognition on subsequent user speech. Alternatively, input data 15 may be text t input via an input device (e.g., a keyboard, etc.).
[0127] For example, when subsequent user voice is received via microphone 110, processor 130 can obtain text t for responding to the subsequent user voice. In this case, text t can be text obtained through ASR module 1 and NLP module 2.
[0128] Furthermore, one or more reference vectors S corresponding to user A stored in reference vector storage module 124 of memory 120. A In this context, the processor 130 can select a reference vector belonging to the domain of the text t through the reference vector selection module 25.
[0129] Here, when multiple reference vectors belonging to the domain of text t are selected, processor 130 can obtain a reference vector for synthesized speech whose score (e.g., prosody score or pronunciation score) calculated based on the characteristics of synthesized speech synthesized from the evaluation text belonging to the domain of text t is the highest among the multiple reference vectors. Here, during the registration of user A's user speech, the score calculated based on the characteristics of synthesized speech synthesized from the evaluation text belonging to the domain of text t can be stored in memory 120.
[0130] For example, suppose the reference vector of the synthesized voice with the highest score among synthesized voices synthesized from evaluation texts belonging to reading style is... Furthermore, the reference vector of the synthesized voice with the highest score among those synthesized using evaluation text belonging to the dialogue style is... When the domain of the text t used as input data 15 is reading style, one or more stored reference vectors S corresponding to user A can be selected. A reference vector As a reference vector belonging to the domain of text t Simultaneously, the processor 130 can also utilize any statistical model (DNN, HMM, GMM, etc.) to select the optimal performance for a given text t.
[0131] Furthermore, the processor 130 can take the text t as input data 15 and the selected reference vector. Input to TTS module 30 to obtain reference vector-based... The speech generated for text t.
[0132] In this case, the processor 130 can control the speaker 141 (see Figure 9B) to output the acquired speech.
[0133] Figure 5 This is a diagram illustrating a method for obtaining a reference vector according to embodiments of the present disclosure.
[0134] The speaker encoder module 10 can obtain reference vectors from user speech. Here, the speaker encoder module 10 can include various types of modules, such as reference encoders, global style tokens (GST), variable autoencoders (VAE), I-vectors, and neural network modules.
[0135] As an example, refer to Figure 5 The speaker encoder module 10 may include an acoustic feature extractor 11 and recurrent neural network (RNN) modules 13-1 to 13-T.
[0136] The acoustic feature extractor 11 can extract acoustic features on a frame-by-frame basis. In this case, the scale of the acoustic features can be represented as (T×D). For example, when a frame is 10ms and 80-dimensional acoustic features are extracted, if a 3-second speech waveform is input, then T is 300 and D is 80, thus outputting (300×80) acoustic features. Typically, the acoustic features are fixed when designing the TTS module 30, therefore, D can have a fixed value regardless of the speech input.
[0137] RNN modules 13-1 to 13-T can output vectors of fixed dimensions, regardless of T. For example, assuming the reference vector is 256-dimensional, RNN modules 13-1 to 13-T can always output vectors of 256 dimensions, regardless of T and D. The reference vector can be output in a state that compresses prosodic or pitch information (global information) included in the corresponding speech rather than phoneme information (local information). In this case, the final state of RNN modules 13-1 to 13-T can be used as the reference vector of this disclosure.
[0138] Figure 7This is a diagram used to describe a TTS model according to embodiments of the present disclosure.
[0139] Reference Figure 7 According to embodiments of the present disclosure, the TTS module 30 can perform preprocessing on the text and speech waveforms through the language processor 31 and the acoustic feature extractor 33 to extract phonemes and acoustic features, and use the preprocessed phonemes and acoustic features as learning data to learn a neural network-based acoustic model (AM) 35 and a voice encoder 37.
[0140] Subsequently, the TTS module 30 can extract phonemes from the text through the language processor 31, input the extracted phonemes into the learned AM 35 to obtain the desired acoustic features as output, and input the obtained acoustic features into the learned sound encoder 37 to obtain the synthesized sound as output.
[0141] However, the above embodiments are merely examples, and this disclosure is not limited thereto, and various modifications can be made.
[0142] Figure 9A is a diagram illustrating the hardware components of an electronic device according to an embodiment of the present invention.
[0143] Referring to FIG9A, an electronic device 100 according to an embodiment of the present disclosure may include a microphone 110, a memory 120, and a processor 130.
[0144] Microphone 110 is a component for receiving analog sound signals. Microphone 110 can receive sound signals including user voice. The sound signal can indicate a sound wave having information such as frequency and amplitude.
[0145] Memory 120 is a component used to store an operating system (OS) that controls the general operation of components of electronic device 100 and various data related to the components of electronic device 100. Memory 120 can store information in various ways, such as electrical or magnetic. Data stored in memory 120 can be accessed by processor 130, and reading, writing, correcting, deleting, updating, etc., of data in memory 120 can be performed by processor 130.
[0146] Therefore, memory 120 can be configured by hardware for temporary or permanent storage of data or information. For example, memory 120 can be implemented as at least one of non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), solid-state drive (SDD), random access memory (RAM), or read-only memory (ROM).
[0147] The processor 130 may be implemented as a general-purpose processor, such as a central processing unit (CPU) or application processor (AP), a graphics-specific processor (such as a graphics processing unit (GPU) or a vision processing unit (VPU)), or an artificial intelligence-specific processor (such as a neural processing unit (NPU)). Furthermore, the processor 130 may include volatile memory for loading at least one instruction or module.
[0148] Figure 9B is a diagram illustrating additional hardware components of an electronic device according to an embodiment of the present invention.
[0149] Referring to FIG9B, in addition to microphone 110, memory 120 and processor 130, electronic device 100 according to embodiments of the present disclosure may include at least one of output interface 140, input interface 150, communication interface 160, sensor 170 or power supply 180.
[0150] Output interface 140 is a component capable of outputting information. Furthermore, output interface 140 may include at least one of speaker 141 or display 143. Speaker 141 can directly output various alarm or audio messages and various audio data, which are then processed by an audio processor (not shown) such as decoding, amplification, and noise filtering. Display 143 can output information or data in a visual form. Display 143 can display image frames that can be driven into pixels on one or all areas of the display. For this purpose, display 143 can be implemented as a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a micro LED display, a quantum dot LED (QLED) display, etc. Furthermore, at least a portion of display 143 can be implemented as a flexible display, and a flexible display can be deformed, bent, or rolled up from a thin, flexible substrate (e.g., paper) without damage.
[0151] Input interface 150 can receive various user commands and transmit the received user commands to processor 130. That is, processor 130 can recognize user commands input by the user through input interface 150. Here, user commands can be implemented in various ways, such as user touch input (touch panel), key (keyboard), or button (physical button or mouse) input.
[0152] The communication interface 160 can send and receive various types of data by performing communication with various types of external devices according to various communication methods. The communication interface 160 may include at least one of the following as circuitry for various types of wireless communication: Bluetooth module, Wi-Fi module, wireless communication module (cellular, such as 3G, 4G or 5G), Near Field Communication (NFC) module, Infrared (IR) module, Zigbee module, and Ultrasonic module, or Ethernet module, Universal Serial Bus (USB) module, High Definition Multimedia Interface (HDMI), DisplayPort (DP), D-SUB, Digital Vision Interface (DVI), Thunderbolt interface, and components for wired communication.
[0153] Sensor 170 can be implemented as various sensors, such as a camera, proximity sensor, illuminance sensor, motion sensor, time-of-flight (ToF) sensor, and global positioning system (GPS) sensor. For example, a camera can divide light into pixel units, sense the intensity of red (R), green (G), and blue (B) light for each pixel, and convert the light intensity into electrical signals to obtain data representing the color, shape, and contrast of an object. In this case, the data type can be an image with R, G, and B color values for each of the multiple pixels. A proximity sensor can sense the presence of surrounding objects and obtain data about whether the surrounding objects exist or are approaching the electronic device. An illuminance sensor can sense the amount of light (or brightness) in the ambient light surrounding the electronic device 100 to obtain data about illuminance. A motion sensor can sense the distance, direction, gradient, etc., of the electronic device 100. For this purpose, a motion sensor can be implemented using a combination of accelerometers, gyroscopes, geomagnetic sensors, etc. ToF sensors can sense the time of flight of various electromagnetic waves (e.g., ultrasonic, infrared, laser beams, and ultra-wideband (UWB) waves) at specific speeds until they return to their original positions to obtain data related to the distance to a target (or the target's location). GPS sensors can receive radio signals from multiple satellites, calculate the distance to each satellite using the transmission time of the received signals, and obtain data about the current location of electronic device 100 using triangulation over the calculated distances. However, the implementation of sensor 170 described above is merely an example; sensor 170 is not limited to this and can be implemented as various types of sensors.
[0154] Power supply 180 can supply power to electronic device 100. For example, power supply 180 can supply power to each component of electronic device 100 via an external commercial power supply or battery.
[0155] Figure 10 This is a flowchart illustrating a control method for an electronic device according to an embodiment of the present invention.
[0156] Reference Figure 10 The control method of the electronic device 100 may include: when a user's voice is received through the microphone 110, obtaining a reference vector of the user's voice (S1010); generating multiple candidate reference vectors based on the reference vector (S1020); obtaining multiple synthesized voices by inputting the multiple candidate reference vectors and multiple evaluation texts into a TTS model (S1030); identifying at least one synthesized voice among the multiple synthesized voices based on the similarity between the multiple synthesized voices and the user's voice and the features of the multiple synthesized voices (S1040); and storing the reference vector of the at least one synthesized voice in the memory 120 as a reference vector of the TTS model corresponding to the user (S1050).
[0157] Specifically, in the control method of the electronic device 100 according to the present disclosure, when the user's voice is received by the microphone 110, a reference vector of the user's voice can be obtained (S1010).
[0158] In addition, multiple candidate reference vectors can be generated based on the reference vector (S1020).
[0159] Here, multiple candidate reference vectors may include: at least one reference vector randomly selected based on a reference vector; at least one reference vector generated based on the reference vector and a reference vector used to train the TTS model; and at least one reference vector generated by applying a masking vector to the reference vector.
[0160] In addition, multiple synthesized voices can be obtained by inputting multiple candidate reference vectors and multiple evaluation texts into the TTS model (S1030).
[0161] As a specific embodiment, by inputting multiple candidate reference vectors and multiple evaluation texts into a TTS model, multiple synthesized voices can be obtained based on each of the multiple candidate reference vectors for each of the multiple evaluation texts.
[0162] Furthermore, at least one of the multiple synthesized voices can be identified based on the similarity between the multiple synthesized voices and the user's speech and the characteristics of the multiple synthesized voices (S1040).
[0163] As a specific embodiment, candidate synthesized sounds whose similarity to user speech is at or above a predetermined threshold among a plurality of synthesized sounds can be identified. Furthermore, at least one synthesized sound among the candidate synthesized sounds can be identified based on at least one of prosody, pronunciation, or sound quality of each candidate synthesized sound.
[0164] Specifically, a prosodic score, pronunciation score, and sound quality score can be calculated for each candidate synthesized voice. Furthermore, at least one synthesized voice among the candidate synthesized voices can be identified whose prosodic score, pronunciation score, and sound quality score are all at or above a predetermined threshold.
[0165] Additionally, the plurality of evaluation texts may include at least one evaluation text belonging to each of the plurality of domains.
[0166] In this context, during the identification of at least one synthesized sound, multiple candidate synthesized sounds can be segmented based on multiple domains, according to the domain to which each of the multiple candidate synthesized sounds belongs. Furthermore, at least one synthesized sound for each domain can be identified based on at least one of the prosody, pronunciation, or sound quality of each of the one or more candidate synthesized sounds belonging to each domain.
[0167] Furthermore, at least one reference vector of the synthesized voice can be stored in memory 120 as a reference vector corresponding to the user of the TTS model (S1050).
[0168] Additionally, the electronic device 100 according to embodiments of the present disclosure may further include an output interface 140, which includes at least one of a speaker 141 or a display 143.
[0169] In this case, in the control method of the electronic device 100, it can be determined that at least one domain of synthesized sound does not exist among multiple domains. Furthermore, when it is determined that no domain of synthesized sound exists, the output interface 140 can be controlled to output information requesting the utterance of a sentence belonging to the determined domain.
[0170] Specifically, at least one candidate synthesized sound can be identified belonging to a domain in which at least one synthesized sound is absent across multiple domains. Furthermore, when a domain in which no synthesized sound is identified, features of synthesized sounds with relatively low prosody, pronunciation, and sound quality scores can be determined based on prosody scores, pronunciation scores, and sound quality scores calculated for the identified candidate synthesized sounds. Additionally, the output interface 140 can be controlled to output information requesting the spoken sentence generated based on the identified features.
[0171] Additionally, the electronic device 100 according to embodiments of the present disclosure may include a speaker 141.
[0172] In this case, in the control method of the electronic device 100, when a user's subsequent voice is received through the microphone 110, text for responding to the subsequent voice can be obtained.
[0173] Furthermore, speech generated from text based on reference vectors can be obtained by inputting the obtained text and one or more reference vectors corresponding to the user stored in memory 120 into the TTS model.
[0174] To this end, a reference vector of synthesized sound can be obtained, wherein the score of the synthesized sound calculated based on the characteristics of the synthesized sound is the highest among one or more reference vectors corresponding to the user stored in memory 120.
[0175] In addition, the speaker 141 can be controlled to output the acquired voice.
[0176] According to the various embodiments of the present disclosure described above, an electronic device and its control method for providing TTS services using the voice of a general user can be provided. Furthermore, according to embodiments of the present disclosure, the number of sentences in which a request to register for a TTS service is made can be minimized. Moreover, according to embodiments of the present disclosure, personalized TTS services can be provided using a user's voice without retraining the TTS model for each user.
[0177] Various embodiments of this disclosure can be implemented by software including instructions stored in a machine-readable storage medium (e.g., a computer-readable storage medium). The machine can be a device that invokes the stored instructions from the storage medium and can operate according to the invoked instructions, and may include electronic devices (e.g., electronic device 100) according to the disclosed embodiments. Where the command is executed by a processor, the processor can directly execute the function corresponding to the command, or other components can execute the function corresponding to the command under the control of the processor. The command may include code created or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory" means that the storage medium is tangible without excluding signals, and does not distinguish whether the data is semi-permanent or temporarily stored in the storage medium.
[0178] Methods according to various embodiments may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a storage medium (e.g., an optical disc read-only memory (CD-ROM)) that can be read by a machine or stored via an application (e.g., the Play Store). TM Online access. In the case of online distribution, at least a portion of the computer program product may be temporarily stored in a storage medium, such as the memory of the manufacturer's server, an application storage server, or a relay server, or may be temporarily created.
[0179] Each component (e.g., module or program) according to different embodiments may include a single entity or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in different embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, and the functions performed by the corresponding components may be executed prior to integration in the same or similar manner. Operations performed by modules, programs, or other components according to different embodiments may be performed sequentially, in parallel, iteratively, or probingly; at least some operations may be performed in a different order, omitted, or additional operations may be added.
Claims
1. An electronic device, comprising: microphone; The memory stores a text-to-speech (TTS) model and multiple evaluation texts; as well as The processor is configured to: When the user's voice is received through the microphone, the reference vector of the user's voice is obtained; Multiple candidate reference vectors are generated based on the aforementioned reference vector. Multiple synthesized voices are obtained by inputting the multiple candidate reference vectors and the multiple evaluation texts into the TTS model; Based on the similarity between the plurality of synthesized voices and the user's speech, and the features of the plurality of synthesized voices, at least one synthesized voice among the plurality of synthesized voices is identified, and The reference vector of the at least one synthesized voice is stored in the memory as a reference vector corresponding to the user in the TTS model.
2. The electronic device as claimed in claim 1, wherein, The plurality of candidate reference vectors includes: at least one reference vector randomly selected based on the reference vector, at least one reference vector generated based on the reference vector and a reference vector used to train the TTS model, and at least one reference vector generated by applying a masking vector to the reference vector.
3. The electronic device as claimed in claim 1, wherein, The processor is configured to obtain the plurality of synthesized voices generated based on each of the plurality of candidate reference vectors being each of the plurality of evaluation texts by inputting the plurality of candidate reference vectors and the plurality of evaluation texts into the TTS model.
4. The electronic device as claimed in claim 1, wherein, The processor is configured to: Among the plurality of synthesized voices, candidate synthesized voices with a similarity to the user's speech that is at or above a predetermined threshold are identified, and At least one synthesized sound of the candidate synthesized sounds is identified based on at least one of the prosody, pronunciation, or sound quality of each of the candidate synthesized sounds.
5. The electronic device as claimed in claim 4, wherein, The processor is configured to: Calculate the prosodic score, pronunciation score, and sound quality score for each of the candidate synthesized sounds. In the candidate synthesized voices, each of the prosody score, the pronunciation score, and the voice quality score is identified as at least one synthesized voice that is at or above a predetermined threshold.
6. The electronic device as claimed in claim 4, wherein, The plurality of evaluation texts includes at least one evaluation text belonging to each of the plurality of domains. The processor is configured as follows: Based on the multiple domains, the multiple candidate synthesized sounds are divided according to the domain to which each of the multiple candidate synthesized sounds belongs. The at least one synthesized sound in each domain is identified based on at least one of the prosody, pronunciation, or sound quality of each of one or more candidate synthesized sounds belonging to each domain.
7. The electronic device as claimed in claim 6, wherein, Based on the domain to which each evaluation text belongs, the reference vector of at least one identified synthetic voice is stored in the memory.
8. The electronic device of claim 7, further comprising an output interface, said output interface including at least one of a speaker or a display. in, The processor is configured to: Among the plurality of domains, a domain in which the at least one synthesized voice does not exist is determined, and features of a synthesized voice with a lower prosodic score, pronunciation score, and voice quality score are determined based on prosodic scores, pronunciation scores, and voice quality scores calculated for at least one candidate synthesized voice belonging to the determined domain. as well as The speaker outputs a voice requesting the speaker to speak a sentence generated based on determined features.
9. The electronic device of claim 1, further comprising a speaker, in, The processor is configured to: When subsequent user voice is received through the microphone, obtain the text of the response to the subsequent user voice. By inputting the obtained text and one or more reference vectors corresponding to the user stored in the memory into the TTS model, speech generated for the text based on the reference vectors is obtained. Control the speaker to output the obtained voice.
10. The electronic device of claim 9, wherein, The processor is configured to: obtain, from one or more reference vectors stored in the memory corresponding to the user, the reference vector with the highest score calculated based on the features of the text to be synthesized.
11. A control method for an electronic device, the electronic device including a memory storing a TTS model and a plurality of evaluation texts, the control method comprising: When a user's voice is received through a microphone, a reference vector of the user's voice is obtained; Multiple candidate reference vectors are generated based on the aforementioned reference vector; Multiple synthesized voices are obtained by inputting the multiple candidate reference vectors and the multiple evaluation texts into the TTS model; Based on the similarity between the multiple synthesized voices and the user's speech, and the features of the multiple synthesized voices, at least one synthesized voice among the multiple synthesized voices is identified; as well as The reference vector of the at least one synthesized voice is stored in the memory as a reference vector corresponding to the user in the TTS model.
12. The control method as described in claim 11, wherein, The plurality of candidate reference vectors includes: at least one reference vector randomly selected based on the reference vector, at least one reference vector generated based on the reference vector and a reference vector used to train the TTS model, and at least one reference vector generated by applying a masking vector to the reference vector.
13. The control method as described in claim 11, wherein, In the process of obtaining the plurality of synthesized voices, the plurality of synthesized voices are obtained by inputting the plurality of candidate reference vectors and the plurality of evaluation texts into the TTS model, based on each of the plurality of candidate reference vectors for each of the plurality of evaluation texts.
14. The control method as described in claim 11, wherein, Identifying the at least one synthesized sound includes: Among the plurality of synthesized voices, candidate synthesized voices with a similarity to the user's speech that is at or above a predetermined threshold are identified, and At least one synthesized sound of the candidate synthesized sounds is identified based on at least one of the prosody, pronunciation, or sound quality of each of the candidate synthesized sounds.
15. The control method as described in claim 14, wherein, Identifying the at least one synthesized sound includes: Calculate the prosodic score, pronunciation score, and sound quality score for each of the candidate synthesized sounds. Among the candidate synthesized voices, each of the prosody score, the pronunciation score, and the voice quality score is identified as at least one synthesized voice that is at or above a predetermined threshold.
Citation Information
Patent Citations
Method and device for personalized voice synthesis
CN111369966A
Method of producing individual characteristic speech sound from text
CN1379391A