Electronic device using synthesized speech acquired from text and control method therefor

By using synthesized voice and embedding models to generate personalized trigger words, the electronic device addresses the challenge of accurate voice recognition activation in noisy environments, enhancing the reliability of voice command execution.

WO2025159341A1PCT designated stage Publication Date: 2025-07-31SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020432
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2024-12-16
Publication Date
2025-07-31

Smart Images

  • Figure KR2024020432_31072025_PF_FP_ABST
    Figure KR2024020432_31072025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device comprises: a memory which stores at least one instruction; and one or more processors which are connected to the memory and control the electronic device, wherein the one or more processors: input text corresponding to a user input to a speech synthesis model to acquire a synthesized sound corresponding to the text; input the synthesized sound to a first embedding model to acquire first embedding, and store the first embedding in the memory; when an utterance sound of a user is received, input the utterance sound to a second embedding model to acquire second embedding; and identify whether the utterance sound corresponds to the text on the basis of a first similarity between the first embedding and the second embedding, and perform an operation corresponding to the text if the utterance sound corresponds to the text.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device using synthesized speech obtained from text and method for controlling the same

[0001] The present invention relates to an electronic device and a control method thereof, and more particularly, to an electronic device that uses a synthesized voice obtained from text as a wake word and a control method thereof.

[0002] Recently, many electronic devices have been equipped with voice recognition capabilities. Users can easily activate voice recognition by uttering a designated trigger word.

[0003] When the electronic device determines that the user has uttered a trigger word, it can activate a voice recognition mode to identify the intent contained in the user's voice command and perform a corresponding action.

[0004] Meanwhile, as multiple home appliances having the same trigger word are installed in the home, when a user utters the trigger word with the intention of controlling a specific home appliance among the multiple home appliances by voice, problems such as the voice recognition mode of a home appliance that does not match the user's intention (e.g., the home appliance closest to the user) being activated or a specific home appliance not recognizing the trigger word due to surrounding noise frequently occur.

[0005] Accordingly, there has been a demand for a method and technology that allows a user to register different trigger words for each of multiple home appliances and smoothly activate the voice recognition mode of a specific home appliance that the user wishes to control by voice among the multiple home appliances.

[0006] According to one embodiment of the present disclosure, an electronic device includes a memory storing at least one instruction and one or more processors connected to the memory to control the electronic device, wherein the one or more processors input text corresponding to a user input into a speech synthesis model to obtain a synthesized sound corresponding to the text, input the synthesized sound into a first embedding model to obtain a first embedding, and store the first embedding in the memory, and when a user's speech sound is received, input the speech sound into a second embedding model to obtain a second embedding, and identify whether the speech sound corresponds to the text based on a first similarity between the first embedding and the second embedding, and perform an operation corresponding to the text if the speech sound corresponds to the text.

[0007] A method for controlling an electronic device according to an embodiment of the present disclosure includes the steps of: inputting text corresponding to a user input into a speech synthesis model to obtain a synthesized sound corresponding to the text; inputting the synthesized sound into a first embedding model to register a first embedding obtained; when a user's speech sound is received, inputting the speech sound into a second embedding model to obtain a second embedding; identifying whether the speech sound corresponds to the text based on a first similarity between the first embedding and the second embedding; and performing an operation corresponding to the text if the speech sound corresponds to the text.

[0008] According to one embodiment of the present disclosure for achieving the above-described object, a computer-readable recording medium including a program for executing a method for controlling an electronic device includes the steps of: inputting text corresponding to a user input into a speech synthesis model to obtain a synthesized sound corresponding to the text; inputting the synthesized sound into a first embedding model to register a first embedding obtained; when a user's speech sound is received, inputting the speech sound into a second embedding model to obtain a second embedding; identifying whether the speech sound corresponds to the text based on a first similarity between the first embedding and the second embedding; and performing an operation corresponding to the text if the speech sound corresponds to the text.

[0009] FIG. 1 is a diagram illustrating a trigger word for calling an electronic device according to an embodiment of the present disclosure.

[0010] FIG. 2 is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure.

[0011] FIG. 3 is a drawing for explaining a synthetic sound corresponding to text according to an embodiment of the present disclosure.

[0012] FIG. 4 is a diagram illustrating an electronic device for comparing a spoken sound and a synthesized sound according to an embodiment of the present disclosure.

[0013] FIG. 5 is a diagram illustrating an electronic device for training a first embedding model according to an embodiment of the present disclosure.

[0014] FIG. 6 is a diagram for explaining text embedding corresponding to text according to an embodiment of the present disclosure.

[0015] FIG. 7 is a diagram for explaining a voice enhancement model according to an embodiment of the present disclosure.

[0016] FIG. 8 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.

[0017] Hereinafter, the present disclosure will be described in detail with reference to the attached drawings.

[0018] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of this disclosure.

[0019] In this specification, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.

[0020] The expression "at least one of A and / or B" should be understood to mean either "A" or "B" or "A and B".

[0021] As used herein, the expressions “first,” “second,” “first,” or “second,” etc., may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.

[0022] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).

[0023] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0024] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented by one or more processors (not shown), excluding any "modules" or "parts" that need to be implemented in specific hardware.

[0025] In this specification, the term user may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).

[0026] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.

[0027] FIG. 1 is a diagram illustrating a trigger word for calling an electronic device according to an embodiment of the present disclosure.

[0028] Referring to FIG. 1, the electronic device (100) can be controlled based on the user's spoken voice. For example, the electronic device (100) is equipped with a microphone and can receive the user's spoken voice through the microphone. Additionally, the electronic device (100) can also receive the user's spoken voice from a remote control device equipped with a microphone.

[0029] Next, the electronic device (100) can enter voice recognition mode when the spoken sound corresponds to the trigger word.

[0030] According to an embodiment, the trigger word may be a preset three- or four-syllable word that activates the voice recognition mode of the electronic device (100) or calls the electronic device (100).

[0031] Depending on the embodiment, the trigger word may be called a wake-up word, a wake-up word, etc. Hereinafter, for convenience of explanation, it will be collectively referred to as a trigger word. The trigger word may be preset during the manufacturing stage of the electronic device (100), may be added, changed, or deleted through firmware updates, etc., and may be edited, such as added, changed, or deleted, according to the user's settings.

[0032] Referring to FIG. 1, when an electronic device (100) receives a user's spoken sound 'Hi Galaxy', and the electronic device (100) receives the received spoken sound 'Hi Galaxy' and corresponds to the trigger word 'Hi Galaxy', the electronic device (100) can activate the voice recognition mode.

[0033] For example, the electronic device (100) may store an embedding (hereinafter, referred to as a first embedding) corresponding to a trigger word and compare the first embedding with an embedding (hereinafter, referred to as a second embedding) corresponding to the user's speech. According to an embodiment, if the first similarity between the first embedding and the second embedding is greater than or equal to a threshold value, the electronic device (100) may identify the user's speech as corresponding to the trigger word and activate a voice recognition mode.

[0034] For example, the electronic device (100) can recognize the user's speech to obtain a voice command, and activate a voice recognition mode for controlling the electronic device (100) according to the voice command (e.g., switching a component within the electronic device (100) for recognizing the user's speech from a standby mode to a normal mode, or supplying power to a component within the electronic device (100) for recognizing the user's speech, etc.).

[0035] As another example, if the first similarity between the first embedding and the second embedding is less than a threshold value, the electronic device (100) may identify that the user's speech does not correspond to the trigger word and may not activate the voice recognition mode (e.g., maintain a component within the electronic device (100) for recognizing the user's speech in standby mode).

[0036] In an embodiment, if the trigger word corresponding to each of a plurality of electronic devices (100) located in a space (e.g., inside a house) is the same, there is a problem in that it is difficult for a user to activate only one of the plurality of electronic devices (100) by uttering the trigger word. For example, if a user utters the trigger word, a problem may occur in that a plurality of electronic devices (100) are simultaneously activated in response to the user's utterance.

[0037] In addition to the above-described circumstances, there has been a demand for a method of setting personalized or customized trigger words, as the trigger words set at the manufacturing stage do not reflect the user's tastes and preferences.

[0038] According to an embodiment, an electronic device (100) provides a method for setting a personalized trigger word, and in particular, the electronic device (100) can provide a method for setting a personalized trigger word without reflecting the speech characteristics of a user setting the trigger word (or a user of the electronic device (100)) and the surrounding environment at the time of setting the trigger word.

[0039] For example, the electronic device (100) can set a personalized trigger word based on text corresponding to user input so that the user's unique frequency, intensity, pronunciation, prosody, speech rate, etc. that sets the trigger word are not reflected when setting the trigger word.

[0040] For example, when a user's voice is received through a microphone and a word included in the voice is set as a trigger word, the electronic device (100) can set a personalized trigger word based on text corresponding to the user input so that noise or interference generated around the electronic device (100) at the time of setting the trigger word (or at the time of receiving the user's voice) is not reflected in the setting of the trigger word.

[0041] For example, when text is input through an input device (e.g., a keyboard, mouse, touchpad, microphone, etc.), the electronic device (100) can obtain a synthetic voice (or synthetic speech) corresponding to the text. For example, the electronic device (100) can obtain a synthetic voice corresponding to the text by inputting the text into a text-to-speech (TTS) model.

[0042] According to an embodiment, the electronic device (100) sets a text as a trigger word for calling the electronic device (100) and can identify whether to activate the voice recognition mode using a first embedding corresponding to the synthesized sound.

[0043] FIG. 2 is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure.

[0044] Referring to FIG. 2, the electronic device (100) includes a memory (110) and one or more processors (120).

[0045] The memory (110) can store a computer program including at least one instruction or instructions for controlling the electronic device (100).

[0046] In the case of memory (110) embedded in the electronic device (100), it may be implemented in the form of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).

[0047] In the case of a memory (110) that can be attached or detached to an electronic device (100), it can be implemented in the form of a memory card (e.g., CF (compact flash), SD (secure digital), Micro-SD (micro secure digital), Mini-SD (mini secure digital), xD (extreme digital), MMC (multi-media card), etc.), an external memory that can be connected to a USB port (e.g., USB memory), a cartridge, etc.

[0048] One or more processors (120) may be implemented as a digital signal processor (DSP), a microprocessor, or a timing controller (TCON) that processes digital signals. However, the present invention is not limited thereto, and may include one or more of a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, or an artificial intelligence (AI) processor, or may be defined by the relevant terms. In addition, one or more processors (120) may be implemented as a system on chip (SoC) or large scale integration (LSI) having a processing algorithm built in, or may be implemented in the form of a field programmable gate array (FPGA). One or more processors (120) may perform various functions by executing computer executable instructions stored in a memory.

[0049] The one or more processors (120) may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors (120) may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing. The one or more processors (120) may execute one or more programs or instructions stored in a memory. For example, the one or more processors (120) may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in a memory.

[0050] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-dedicated processor).

[0051] One or more processors (120) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When one or more processors (120) are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include an internal processor memory, such as a cache memory or an on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.

[0052] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among a plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.

[0053] In embodiments of the present disclosure, a processor may mean a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto.

[0054] According to an embodiment, one or more processors (120) may input text corresponding to a user input into a speech synthesis model (e.g., TTS) to obtain a synthesized sound corresponding to the text.

[0055] One or more processors (120) can input a synthesized sound into a first embedding model to obtain a first embedding and store the first embedding in a memory (110). When a user's speech sound is received, one or more processors (120) can input the speech sound into a second embedding model to obtain a second embedding.

[0056] According to an embodiment, one or more processors (120) may perform an action corresponding to the text if the speech corresponds to the text based on the first similarity between the first embedding and the second embedding.

[0057] According to an embodiment, the action corresponding to the text is not limited to the action of entering the voice recognition mode, and may include various actions that the electronic device (100) can perform (e.g., turning the display on / off, activating some components among the components included in the electronic device (100), etc.).

[0058] FIG. 3 is a drawing for explaining a synthetic sound corresponding to text according to an embodiment of the present disclosure.

[0059] Referring to FIG. 3, when a user inputs text (20) to be set as a trigger word using an input device or the like, one or more processors (120) can input the text (20) into a voice synthesis model (1) to obtain a synthesized sound (30) corresponding to the text (20).

[0060] According to an embodiment, when text (20) is input, the voice synthesis model (1) can obtain tokens by dividing the text (20) into word (or syllable) units (e.g., tokenization).

[0061] The speech synthesis model (1) can embed tokens and generate speech parameters for the embedded tokens. Here, the speech parameters can represent fixed (or preset) speech characteristics of a speaker (e.g., frequency, intensity, pronunciation, prosody, speech rate, etc.).

[0062] According to an embodiment, the voice synthesis model (1) can convert voice parameters into a synthesized sound (30) and output it.

[0063] According to an embodiment, one or more processors (120) may obtain a synthesized sound (30) that reflects the speech characteristics of a fixed (or preset) speaker, rather than the speech characteristics of a user of the electronic device (100) (or a user who wishes to set a trigger word).

[0064] According to an embodiment, one or more processors (120) may input a synthesized sound (30) into a first embedding model to obtain a first embedding, and store the first embedding in a memory (110) (or register the first embedding).

[0065] FIG. 4 is a diagram illustrating an electronic device for comparing a spoken sound and a synthesized sound according to an embodiment of the present disclosure.

[0066] Referring to FIG. 4, one or more processors (120) can input a synthetic sound (30) into a first embedding model (2) to obtain a first embedding (50). Depending on the embodiment, the first embedding (50) may be called a synthetic sound embedding or a synthetic sound embedding vector, but for the convenience of explanation, it will be collectively referred to as the first embedding (50).

[0067] According to an embodiment, one or more processors (120) input a synthetic sound (30) into a first embedding model (2), and the first embedding model (2) can output a first embedding (50) that expresses the synthetic sound (30) in vector form.

[0068] The first embedding model (2) may also be called the SWE (Synthetic Word Embedding) Model, but for convenience of explanation, it is referred to as the first embedding model (2).

[0069]

[0070] According to an embodiment, after storing (or registering) the first embedding (50), when a user's speech sound (40) is received, one or more processors (120) may input the speech sound (40) into the second embedding model (3) to obtain a second embedding (60). Depending on the embodiment, the second embedding (60) may be called a speech sound embedding or a speech sound embedding vector, but for the convenience of explanation, it will be collectively referred to as the second embedding (60).

[0071] The second embedding model (3) may also be called the AWE (Acoustic Word Embedding) Model, but for convenience of explanation, it will be referred to as the second embedding model (3).

[0072] According to an embodiment, one or more processors (120) input a speech sound (40) into a second embedding model (3), and the second embedding model (3) can output a second embedding (60) that expresses the speech sound (40) in vector form.

[0073] According to an embodiment, one or more processors (120) measure the distance (d) between the first embedding (50) and the second embedding (60). SA ) can identify the similarity (or similarity) between the first embedding (50) and the second embedding (60).

[0074] For example, one or more processors (120) may determine the distance (d) between the first embedding (50) and the second embedding (60). SA ) is longer, the more dissimilar the first embedding (50) and the second embedding (60) are, and the distance (d) between the first embedding (50) and the second embedding (60) SA ) is shorter, the first embedding (50) and the second embedding (60) can be identified as similar.

[0075] One or more processors (120) are configured to determine whether the first similarity between the first embedding (50) and the second embedding (60) is greater than or equal to a threshold value (e.g., the distance (d) between the first embedding (50) and the second embedding (60) SA ) is less than the threshold distance), the user's speech (40) can be identified as corresponding to the text (20).

[0076] According to an embodiment, one or more processors (120) may perform an action corresponding to the text when the user's speech (40) corresponds to the text (20).

[0077] For example, the text is a customized trigger word, and if the first similarity between the first embedding (50) and the second embedding (60) is greater than a threshold value, one or more processors (120) can identify the user's speech (40) as having uttered the trigger word and enter voice recognition mode.

[0078]

[0079] According to an embodiment, one or more processors (120) may input text (20) into a speech synthesis model (1) to obtain a synthesized sound (30) and characteristic information related to the synthesized sound (30). Here, the characteristic information related to the synthesized sound (30) may include at least one of length (e.g., duration) information or prosody information of the synthesized sound (30).

[0080] According to an embodiment, one or more processors (120) can input a synthetic sound (30) and characteristic information related to the synthetic sound (30) into a first embedding model (2) to obtain a first embedding (50).

[0081] For example, since a trigger word is a word of a preset length of 3 or 4 syllables, in order to more accurately identify whether the user's speech (40) corresponds to the text (20), one or more processors (120) input the synthesized sound (30) and the characteristic information related to the synthesized sound (30) into the first embedding model (2) to obtain the first embedding (50), and compare the first embedding (50) with the second embedding (60).

[0082] According to an embodiment, since the characteristic information related to the synthetic sound (30) is an intermediate output rather than a final output (e.g., the synthetic sound (30)) of the voice synthesis model (1), one or more processors (120) input the characteristic information, which is an intermediate output of the synthetic sound (30) and the voice synthesis model (1), into the first embedding model (2), so that the synthetic sound (30) and the spoken sound (40) can be compared more accurately.

[0083] FIG. 5 is a diagram illustrating an electronic device for training a first embedding model according to an embodiment of the present disclosure.

[0084] Referring to FIG. 5, the plurality of sample speech sounds (70-1, ..., 70-n) may include a plurality of first sample speech sounds corresponding to the text (20) and a plurality of second sample speech sounds not corresponding to the text (20).

[0085] For example, one or more processors (120) can input a plurality of first sample utterances into a second embedding model (3) to obtain a plurality of first sample embeddings.

[0086] One or more processors (120) can input a plurality of second sample utterances into a second embedding model (3) to obtain a plurality of second sample embeddings.

[0087] According to an embodiment, one or more processors (120) are configured to ensure that the similarity between each of the plurality of first sample embeddings and the first embedding (50) is greater than or equal to a threshold value (or, the distance (d) between each of the plurality of first sample embeddings and the first embedding (50) SA ) can be trained to shorten the first embedding model (2).

[0088] Additionally, one or more processors (120) are configured to ensure that the similarity between each of the plurality of second sample embeddings and the first embedding (50) is less than a threshold value (or, the distance (d) between each of the plurality of second sample embeddings and the first embedding (50) SA ) can be trained to make the first embedding model (2) longer.

[0089] The artificial intelligence-related function according to the present disclosure is operated through one or more processors (120) and memory (110) of an electronic device (100).

[0090] One or more processors (120) may include at least one of a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), and an NPU (Neural Processing Unit), but are not limited to the examples of the processors described above.

[0091] CPUs are general-purpose processors capable of performing not only general calculations but also artificial intelligence calculations. Their multi-layered cache structure allows for the efficient execution of complex programs. CPUs are advantageous for serial processing, enabling organic linking of previous and subsequent calculation results through sequential calculations. General-purpose processors are not limited to the examples described above, except where specifically identified as CPUs.

[0092] A GPU is a processor designed for large-scale computations, such as floating-point operations used in graphics processing. It integrates a large number of cores to perform large-scale computations in parallel. In particular, GPUs may be advantageous over CPUs in parallel processing methods, such as convolution operations. Furthermore, GPUs can be used as coprocessors to supplement the functions of CPUs. Processors for large-scale computations are not limited to the examples described above, except in cases where they are specifically referred to as GPUs.

[0093] An NPU is a processor specialized in artificial intelligence computation using artificial neural networks, and each layer of the artificial neural network can be implemented in hardware (e.g., silicon). Since an NPU is designed specifically according to the company's specifications, it has less freedom than a CPU or GPU, but can efficiently process the AI ​​computations requested by the company. Meanwhile, as a processor specialized in artificial intelligence computation, an NPU can be implemented in various forms, such as a Tensor Processing Unit (TPU), an Intelligence Processing Unit (IPU), or a Vision Processing Unit (VPU). Except as specifically stated as an NPU, an AI processor is not limited to the examples described above.

[0094] Additionally, one or more processors (120) may be implemented as a System on Chip (SoC). In this case, the SoC may further include, in addition to one or more processors, a memory, and a network interface such as a bus for data communication between the processor and the memory.

[0095] When a plurality of processors are included in a SoC (System on Chip) included in an electronic device (100), the electronic device (100) may perform operations related to artificial intelligence (e.g., operations related to learning or inference of an artificial intelligence model) by using some of the plurality of processors. For example, the electronic device (100) may perform operations related to artificial intelligence by using at least one of a GPU, an NPU, a VPU, a TPU, and a hardware accelerator specialized in artificial intelligence operations such as convolution operations and matrix multiplication operations among the plurality of processors. However, this is merely an example, and it is of course possible to process operations related to artificial intelligence by using a CPU or a general-purpose processor.

[0096] Additionally, the electronic device (100) can perform operations related to functions related to artificial intelligence by utilizing multiple cores (e.g., dual cores, quad cores, etc.) included in a single processor. In particular, the electronic device (100) can perform artificial intelligence operations, such as convolution operations and matrix multiplication operations, in parallel by utilizing multiple cores included in the processor.

[0097] One or more processors are controlled to process input data according to predefined operation rules or artificial intelligence models stored in memory (110). The predefined operation rules or artificial intelligence models are characterized by being created through learning.

[0098] Here, "created through learning" means that a predefined set of behavioral rules or an AI model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the AI ​​according to the present disclosure is implemented, or through a separate server / system.

[0099] An artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), a neural recurrent network (NeRF), a deep Q-network (Deep Q-Network), and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.

[0100] A learning algorithm is a method for training a target device (e.g., a robot) using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.

[0101] FIG. 6 is a diagram for explaining text embedding corresponding to text according to an embodiment of the present disclosure.

[0102] Referring to FIG. 6, one or more processors (120) can input text (20) into a third embedding model (4) to obtain a third embedding (80).

[0103] According to an embodiment, the third embedding model (4) may also be called a TWE (Text Word Embedding) Model.

[0104] For example, when text (20) is input, the third embedding model (4) can map the text (20) to a high-dimensional vector space and output a third embedding (or third embedding vector) (80) corresponding to the text (20).

[0105] According to an embodiment, one or more processors (120) may compare the first embedding (50) and the second embedding (60) to identify a first similarity between the first embedding (50) and the second embedding (60).

[0106] Additionally, one or more processors (120) can compare the second embedding (60) and the third embedding (80) to identify a second similarity between the second embedding (60) and the third embedding (80). For example, the one or more processors (120) can determine the distance (d) between the second embedding (60) and the third embedding (80). TA ) can be identified.

[0107] According to an embodiment, one or more processors (120) may identify the speech sound (40) as corresponding to the text (20) if at least one of the first similarity or the second similarity is greater than or equal to a threshold value.

[0108] According to an embodiment, one or more processors (120) may input a speech sound (40) into a second embedding model (3) to obtain a second embedding (60) and input text (20) into a third embedding model (4) to obtain a third embedding (80). If the second similarity between the second embedding (60) and the third embedding (80) is greater than a threshold value, the one or more processors (120) may identify that the user has uttered a trigger word and enter a voice recognition mode.

[0109] For example, one or more processors (120) may identify that the user has uttered a trigger word and enter a voice recognition mode if the second similarity between the second embedding (60) and the third embedding (80) is greater than or equal to the threshold value, even if the first similarity between the first embedding (50) and the second embedding (60) is less than or equal to the threshold value.

[0110] Not limited thereto, one or more processors (120) may identify that the speech sound (40) corresponds to the text (20) if each of the first similarity and the second similarity is greater than or equal to a threshold value. For example, one or more processors (120) may identify that the user has not spoken the trigger word and may not enter the voice recognition mode if at least one of the first similarity and the second similarity is less than the threshold value.

[0111]

[0112] As illustrated in FIG. 5, the plurality of sample speech sounds (70-1, ..., 70-n) may include a plurality of first sample speech sounds corresponding to the text (20) and a plurality of second sample speech sounds not corresponding to the text (20).

[0113] According to an embodiment, one or more processors (120) can input a plurality of first sample speech sounds into a second embedding model (3) to obtain a plurality of first sample embeddings, and can input a plurality of second sample speech sounds into a second embedding model (3) to obtain a plurality of second sample embeddings.

[0114] According to an embodiment, one or more processors (120) are configured to ensure that the similarity between each of the plurality of first sample embeddings and the third embedding (80) is greater than or equal to a threshold value (or, the distance (d) between each of the plurality of first sample embeddings and the third embedding (80) TA ) can be trained to shorten the third embedding model (4).

[0115] Additionally, one or more processors (120) may be configured to ensure that the similarity between each of the plurality of second sample embeddings and the third embedding (80) is less than a threshold value (or, the distance (d) between each of the plurality of second sample embeddings and the third embedding (80) TA ) can be trained to make the third embedding model (4) longer.

[0116] FIG. 7 is a diagram for explaining a voice enhancement model according to an embodiment of the present disclosure.

[0117] According to an embodiment of the present disclosure, one or more processors (120) may apply a speech sound (40) to a speech enhancement (SE) model (5) and then input it into a second embedding model (3).

[0118] For example, one or more processors (120) can apply the speech sound (40) to a voice enhancement model (5) to remove noise contained in the speech sound (40), amplify the speech sound, and then input the amplified speech sound into a second embedding model (3) to obtain a second embedding (60).

[0119] For example, in order to more accurately compare a synthesized sound (30) and a spoken sound (40), a voice enhancement model (5) receives a user's voice through a microphone to obtain a spoken sound (40), and removes noise and interference included in the spoken sound (40) that occur around the electronic device (100) (or microphone) at the time of receiving the user's voice, and can amplify the spoken sound (40) from which the noise has been removed so that it is greater than a preset size (e.g., SPL, Loudness, dB, etc.).

[0120] For example, the process of removing noise and noise from a speech sound (40) and amplifying it using a speech enhancement (SE) model (5) may be called a preprocessing process.

[0121] According to an embodiment, one or more processors (120) may train a voice enhancement model based on a plurality of third sample speech sounds and transcript information corresponding to each of the plurality of third sample speech sounds.

[0122] According to an embodiment, when text (20) is changed, one or more processors (120) can input the changed text (20') into a voice synthesis model (1) to obtain a new synthesized sound (30') corresponding to the changed text (20').

[0123] One or more processors (120) can input a new synthetic sound (30') into the first embedding model (2) to obtain and store (or register) a new first embedding (50').

[0124] According to various embodiments of the present disclosure, an electronic device can be provided that can easily change a trigger word by generating a synthesized sound (30) that reflects the speech characteristics of a fixed (or preset) speaker without reflecting the speech characteristics of a user setting a trigger word (or a user of an electronic device (100)) and the surrounding environment at the time of setting the trigger word (i.e., minimizing variation factors).

[0125] FIG. 8 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.

[0126] A method for controlling an electronic device according to an embodiment of the present disclosure inputs text corresponding to a user input into a voice synthesis model to obtain a synthesized sound corresponding to the text (S810).

[0127] The first embedding obtained by inputting the synthesized sound into the first embedding model is registered (S820).

[0128] When the user's speech sound is received, the speech sound is input into the second embedding model to obtain the second embedding (S830).

[0129] Based on the first similarity between the first embedding and the second embedding, it is identified whether the utterance corresponds to the text (S840).

[0130] If the speech sound corresponds to text, an action corresponding to the text is performed (S850).

[0131] A control method according to an embodiment of the present disclosure further includes a step of inputting text into a speech synthesis model to obtain characteristic information related to a synthesized sound, and the step S820 of registering a first embedding includes a step of inputting a synthesized sound and characteristic information into a first embedding model to obtain a first embedding, and the characteristic information may include length information of the synthesized sound and prosody information of the synthesized sound.

[0132] A control method according to an embodiment of the present disclosure may further include a step of obtaining a plurality of first sample embeddings by inputting a plurality of first sample speech sounds corresponding to text into a second embedding model, a step of obtaining a plurality of second sample embeddings by inputting a plurality of second sample speech sounds not corresponding to text into the second embedding model, and a step of training the first embedding model such that the similarity between each of the plurality of first sample embeddings and the first embedding is greater than or equal to a threshold value, and the similarity between each of the plurality of second sample embeddings and the first embedding is less than or equal to the threshold value.

[0133] A control method according to an embodiment of the present disclosure further includes a step of inputting a text into a third embedding model to obtain a third embedding, and the step S850 of performing an operation may include a step of identifying a speech sound as corresponding to a text and performing an operation corresponding to the text if at least one of a first similarity between the first embedding and the second embedding or a second similarity between the first embedding and the third embedding is greater than or equal to a threshold value.

[0134] A control method according to an embodiment of the present disclosure may further include a step of obtaining a plurality of first sample embeddings by inputting a plurality of first sample utterance sounds corresponding to text into a second embedding model, a step of obtaining a plurality of second sample embeddings by inputting a plurality of second learning sounds not corresponding to text into the second embedding model, and a step of training a third embedding model such that the similarity between each of the plurality of first sample embeddings and the third embedding is greater than or equal to a threshold value, and the similarity between each of the plurality of second sample embeddings and the third embedding is less than or equal to the threshold value.

[0135] A control method according to an embodiment of the present disclosure may further include a step of applying a speech sound to a speech enhancement model to remove noise contained in the speech sound and amplify the same, and then inputting the same into a second embedding model to obtain a second embedding.

[0136] A control method according to an embodiment of the present disclosure may further include a step of training a voice enhancement model based on a plurality of third sample speech sounds and transcript information corresponding to each of the plurality of third sample speech sounds.

[0137] According to an embodiment of the present disclosure, the text is a trigger word for activating a voice recognition mode of an electronic device, and the step S850 of performing the operation may further include a step of activating the voice recognition mode by identifying the spoken sound as corresponding to the text if the first similarity between the first embedding and the second embedding is greater than or equal to a threshold value.

[0138] A control method according to an embodiment of the present disclosure may further include, when text is changed, a step of inputting the changed text into a speech synthesis model to obtain a new synthesized sound corresponding to the changed text, and a step of inputting the new synthesized sound into a first embedding model to register the obtained new first embedding.

[0139] The speech synthesis model according to an embodiment of the present disclosure may be a text-to-speech model that converts text into synthesized speech.

[0140] However, it goes without saying that the various embodiments of the present disclosure can be applied to various types of electronic devices.

[0141] Meanwhile, the various embodiments described above may be implemented in a computer-readable recording medium or similar device using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented by the processor itself. In a software implementation, embodiments, such as the procedures and functions described herein, may be implemented as separate software modules. Each of the software modules may perform one or more functions and operations described herein.

[0142] Meanwhile, computer instructions for performing processing operations of an electronic device according to various embodiments of the present disclosure described above may be stored in a non-transitory computer-readable medium. When the computer instructions stored in such a non-transitory computer-readable medium are executed by a processor of a specific device, the computer instructions cause the specific device to perform processing operations in the electronic device according to various embodiments described above.

[0143] A non-transitory computer-readable medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.

[0144] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, memory that stores at least one instruction; and one or more processors connected to the memory and controlling the electronic device; One or more of the above processors, Inputting text corresponding to user input into a speech synthesis model to obtain a synthesized sound corresponding to the text, Inputting the above synthetic sound into a first embedding model to obtain a first embedding and storing the first embedding in the memory, When the user's speech sound is received, the speech sound is input into the second embedding model to obtain the second embedding, Identifying whether the utterance corresponds to the text based on the first similarity between the first embedding and the second embedding, An electronic device that performs an action corresponding to the text when the above-mentioned speech sound corresponds to the above-mentioned text.

2. In paragraph 1, One or more of the above processors, By inputting the above text into the above speech synthesis model, characteristic information related to the synthesized sound is obtained, The first embedding is obtained by inputting the above synthetic sound and the above characteristic information into the first embedding model, The above characteristic information is, An electronic device including length information of the synthesized sound and prosody information of the synthesized sound.

3. In paragraph 1, One or more of the above processors, A plurality of first sample utterances corresponding to the above text are input into the second embedding model to obtain a plurality of first sample embeddings, A plurality of second sample utterances that do not correspond to the above text are input into the second embedding model to obtain a plurality of second sample embeddings, An electronic device that trains the first embedding model so that the similarity between each of the plurality of first sample embeddings and the first embedding is greater than or equal to a threshold value, and the similarity between each of the plurality of second sample embeddings and the first embedding is less than the threshold value.

4. In paragraph 1, One or more of the above processors, Input the above text into the third embedding model to obtain the third embedding, An electronic device that identifies the speech sound as corresponding to the text and performs an action corresponding to the text if at least one of the first similarity between the first embedding and the second embedding or the second similarity between the first embedding and the third embedding is greater than or equal to a threshold value.

5. In paragraph 4, One or more of the above processors, A plurality of first sample utterances corresponding to the above text are input into the second embedding model to obtain a plurality of first sample embeddings, A plurality of second learning sounds that do not correspond to the above text are input into the second embedding model to obtain a plurality of second sample embeddings, An electronic device that trains the third embedding model so that the similarity between each of the plurality of first sample embeddings and the third embedding is greater than or equal to the threshold value, and the similarity between each of the plurality of second sample embeddings and the third embedding is less than the threshold value.

6. In paragraph 1, One or more of the above processors, An electronic device that applies the above-mentioned speech sound to a speech enhancement model to remove noise contained in the above-mentioned speech sound and amplify the same, and then inputs the above-mentioned speech sound to the second embedding model to obtain the above-mentioned second embedding.

7. In paragraph 6, One or more of the above processors, An electronic device that trains the speech enhancement model based on a plurality of third sample speech sounds and transcript information corresponding to each of the plurality of third sample speech sounds.

8. In paragraph 1, The above text is a trigger word for activating the voice recognition mode of the electronic device, One or more of the above processors, An electronic device that identifies the speech sound as corresponding to the text and activates the voice recognition mode when the first similarity between the first embedding and the second embedding is greater than or equal to a threshold value.

9. In paragraph 1, One or more of the above processors, When the above text is changed, the changed text is input into the speech synthesis model to obtain a new synthesized sound corresponding to the changed text, An electronic device that inputs the new synthetic sound into the first embedding model to obtain and store a new first embedding.

10. In paragraph 1, The above speech synthesis model is, An electronic device, which is a text-to-speech model that converts the above text into the above synthesized sound.

11. In a method for controlling an electronic device, A step of inputting text corresponding to a user input into a speech synthesis model to obtain a synthesized sound corresponding to the text; A step of registering a first embedding obtained by inputting the above-mentioned synthetic sound into a first embedding model; When a user's speech sound is received, a step of inputting the speech sound into a second embedding model to obtain a second embedding; A step of identifying whether the utterance corresponds to the text based on the first similarity between the first embedding and the second embedding; and A control method comprising: a step of performing an action corresponding to the text when the above-mentioned speech sound corresponds to the above-mentioned text.

12. In paragraph 11, The above control method is, It further includes a step of obtaining characteristic information related to the synthesized sound by inputting the text into the speech synthesis model; The step of registering the above first embedding is: A step of obtaining the first embedding by inputting the synthetic sound and the characteristic information into the first embedding model; The above characteristic information is, A control method including length information of the above-mentioned synthetic sound and prosody information of the above-mentioned synthetic sound.

13. In paragraph 12, The above control method is, A step of obtaining a plurality of first sample embeddings by inputting a plurality of first sample utterances corresponding to the above text into the second embedding model; A step of obtaining a plurality of second sample embeddings by inputting a plurality of second sample utterances that do not correspond to the above text into the second embedding model; and A control method further comprising: a step of training the first embedding model so that the similarity between each of the plurality of first sample embeddings and the first embedding is greater than or equal to a threshold value, and the similarity between each of the plurality of second sample embeddings and the first embedding is less than the threshold value.

14. In paragraph 11, The above control method is, Further comprising a step of obtaining a third embedding by inputting the above text into a third embedding model; The steps for performing the above actions are: A control method comprising: a step of identifying the speech sound as corresponding to the text and performing an action corresponding to the text if at least one of the first similarity between the first embedding and the second embedding or the second similarity between the first embedding and the third embedding is greater than or equal to a threshold value; 15. In paragraph 14, The above control method is, A step of obtaining a plurality of first sample embeddings by inputting a plurality of first sample utterances corresponding to the above text into the second embedding model; A step of obtaining a plurality of second sample embeddings by inputting a plurality of second learning sounds that do not correspond to the above text into the second embedding model; and A control method further comprising: a step of training the third embedding model so that the similarity between each of the plurality of first sample embeddings and the third embedding is greater than or equal to the threshold value, and the similarity between each of the plurality of second sample embeddings and the third embedding is less than the threshold value.

Citation Information

Patent Citations

  • Individual wheel type pallet transfer equipment

    KR1020210146574A

  • A shampoo pump with the ability to attach and remove the shampoo head to the shampoo bottle or pump by itself

    KR1020250021244A

  • A cartridge of minimizing discharge of residual contents and a caulking gun using the same

    KR102669285B1

  • Electronic device, method and computer program

    US20200312322A1

  • Passive and continuous multi-speaker voice biometrics

    US20210326421A1