Electronic device and method for speech recognition
The method and device improve keyword detection accuracy and power efficiency in voice recognition systems by employing a keyword-adaptive detection model and threshold determination model, addressing the challenges of noisy environments and power management.
Patent Information
- Application Number
- PCT/KR2024/012936
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2024-08-29
- Publication Date
- 2025-08-28
AI Technical Summary
Existing voice recognition systems face challenges in accurately identifying keywords in noisy environments and optimizing power consumption, particularly in systems utilizing Wake-On-Voice technology.
A method and device utilizing a keyword-adaptive detection model and a threshold value determination model to determine the presence of a keyword in a speech signal, enabling accurate keyword identification and efficient power management.
Enhances the accuracy of keyword detection in noisy conditions while optimizing power consumption, ensuring reliable voice recognition performance.
Smart Images

Figure KR2024012936_28082025_PF_FP_ABST
Abstract
Description
Electronic device and method for speech recognition
[0001] The present disclosure relates to a device and method for performing speech recognition. More specifically, the present disclosure relates to a device and method for determining whether a speech signal contains a keyword.
[0002] Recently, electronic devices equipped with voice recognition have been released to enhance the controllability and operability of various functions. Voice recognition offers the advantage of allowing easy control of the device by recognizing the user's voice without the need for separate button operations or touch modules.
[0003] For example, this voice recognition feature allows you to make calls or write text messages without pressing separate buttons on a variety of electronic devices, including but not limited to portable devices like smartphones and home appliances like TVs and refrigerators. This allows you to easily configure various functions, such as directions, internet searches, and alarm settings.
[0004] To be controlled by the user's voice at a considerable distance from the voice recognition device, the device must be able to perform reliably even in noisy environments. To ensure stable performance, Wake-On-Voice (WoV) technology can be utilized, allowing the user to indicate to the voice recognition device when to initiate voice recognition. To wake up the voice recognition device, the user utters a predetermined keyword (or wake word) before the main command. Since WoV technology represents the first step in voice control, it requires high accuracy.
[0005] Recently, in the field of voice recognition, various technologies for recognizing the user's voice are being studied, and in particular, as a technology to solve the problems of power consumption and malfunction due to the constant operation of the system for voice recognition service, Wake On Voice technology for starting the voice recognition service is being actively studied.
[0006] The above information is provided solely as background information to aid in understanding the present disclosure and, therefore, the present disclosure should not be considered limited to the aforementioned information.
[0007] In one embodiment of the present disclosure, a method for speech recognition is provided. The method may include obtaining a text input including a keyword. The method may include obtaining a speech signal corresponding to a user's utterance. The method may include obtaining a probability value for the keyword using a keyword-adaptive detection model. The method may include obtaining a threshold value for the keyword using a threshold value determination model. The method may include determining whether the acquired speech signal includes the keyword based on the probability value and the threshold value.
[0008] In one embodiment of the present disclosure, a computer-readable recording medium having a program recorded thereon is provided. The program may include a program for causing a computer to perform a method comprising any of the steps described above.
[0009] In one embodiment of the present disclosure, an electronic device for speech recognition is provided. The electronic device may include at least one processor including a processing circuit, and a memory including one or more storage media storing at least one instruction. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a text input including a keyword. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a voice signal corresponding to a user's utterance. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a probability value for the keyword using a keyword-adaptive detection model. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a threshold value for the keyword using a threshold value determination model. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to determine whether the acquired voice signal includes the keyword based on the probability value and the threshold value.
[0010] FIG. 1 is a schematic diagram illustrating a method for speech recognition according to one embodiment of the present disclosure.
[0011] FIG. 2 is a flowchart of a method for speech recognition according to one embodiment of the present disclosure.
[0012] FIG. 3 is a block diagram of an electronic device that determines whether a voice signal contains a keyword according to one embodiment of the present disclosure.
[0013] FIG. 4 is a diagram illustrating a process for training models of an electronic device according to one embodiment of the present disclosure.
[0014] FIG. 5 is a flowchart of an operation for obtaining a probability value for a keyword according to one embodiment of the present disclosure.
[0015] FIG. 6 is a diagram illustrating a process of training at least one model of an electronic device using user voice data according to one embodiment of the present disclosure.
[0016] FIG. 7 is a diagram for explaining a process for determining whether a voice signal according to one embodiment of the present disclosure is a voice signal obtained by a user's speech.
[0017] FIG. 8A is a diagram illustrating a user interface for enrolling or registering a keyword according to one embodiment of the present disclosure.
[0018] FIG. 8b is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0019] FIG. 9A is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0020] FIG. 9b is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0021] FIG. 9c is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0022] FIG. 10 is a diagram of a system in which voice recognition is performed using a registered keyword according to one embodiment of the present disclosure.
[0023] FIG. 11 is a block diagram of an electronic device for voice recognition according to one embodiment of the present disclosure.
[0024] FIG. 12 is a block diagram of an electronic device for voice recognition according to one embodiment of the present disclosure.
[0025] Hereinafter, the present disclosure will be described in detail by describing an embodiment of the present disclosure with reference to the attached drawings.
[0026] The detailed description below is provided to help the reader gain a comprehensive understanding of the methods, devices, and / or systems described in this disclosure. However, various modifications, alterations, and equivalents to the methods, devices, and / or systems described in this disclosure will become apparent upon understanding this disclosure. For example, the operational sequences described in this disclosure are merely exemplary and are not limited to those set forth in this disclosure, and may be modified to a degree apparent upon understanding the disclosure, except for operations that necessarily occur in a specific order. Furthermore, descriptions of features that are known upon understanding this disclosure may be omitted for clarity and conciseness.
[0027] In this disclosure, the expression “at least one of a, b or c” may refer to “a”, “b”, “c”, “a and b”, “a and c”, “b and c”, “all of a, b and c”, or variations thereof.
[0028] The terms used in this disclosure are selected from widely used, common terms, taking into account the functions of the disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings can be understood through the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of the disclosure.
[0029] In this disclosure, singular expressions may include plural expressions unless the context clearly dictates otherwise. For example, the description "a constituent surface" may also refer to one or more of such surfaces. Terms containing ordinal numbers, such as "first" or "second," used in this disclosure may be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another.
[0030] When a part of this disclosure is said to "include" a component, this does not exclude other components, but rather may include other components, unless otherwise specifically stated. In this disclosure, terms such as "part" and "module" refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.
[0031] The expression "configured to" as used herein can be used interchangeably with, for example, "suitable for", "having the capacity to", "designed to", "adapted to", "made to", or "capable of", depending on the context. The term "configured to" does not necessarily mean something is "specifically designed to" in hardware. Alternatively, in some contexts, the expression "a system configured to" can include that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can include a dedicated processor for performing the operations (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in a memory.
[0032] When a component is referred to as being "connected" or "connected" to another component in this disclosure, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless otherwise specifically stated.
[0033] In describing the present disclosure, descriptions of technical details that are well known in the technical field to which the present disclosure pertains and are not directly related to the present disclosure may be omitted. This is to convey the gist of the present disclosure more clearly without obscuring unnecessary explanation. In the drawings, parts irrelevant to the description are omitted for clarity in describing the present disclosure, and similar parts are designated with similar reference numerals throughout the specification. The size of each component does not entirely reflect the actual size. The same or corresponding components in each drawing are given the same reference numerals.
[0034] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure is complete and to fully inform those skilled in the art of the disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims.
[0035] In the present disclosure, combinations of each block in the flowchart and the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing equipment, and the instructions executed by the processor of the computer or other programmable data processing equipment can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to perform a function in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing equipment. It should be understood that each combination of blocks in the flowchart and the flowchart diagrams can be performed by one or more computer programs that include computer-executable instructions. One or more computer programs may be stored entirely in a single memory, or may be split across multiple different memories.
[0036] All functions or operations described in the present disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).
[0037] At least one processor according to an embodiment of the present invention may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits comprising at least one processor, one or more of which are configured to individually and / or collectively perform the various functions described herein in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms may include, for example, without limitation, a single processor performing some of the recited functions, other processor(s) performing other of the recited functions, and still other situations where a single processor may perform all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform the various functions.
[0038] In the present disclosure, each block in the flowchart may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.
[0039] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.
[0040] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts that are not related to the description are omitted in order to clearly describe the present disclosure, and similar parts are designated with similar reference numerals throughout the specification. In addition, the reference numerals used in each drawing are only for the purpose of describing each drawing, and different reference numerals used in different drawings do not indicate different elements. The present disclosure will be described in detail below with reference to the attached drawings.
[0041] In one embodiment of the present disclosure, a voice recognition assistant service may include a service that identifies a request contained in a voice signal (or audio signal) and processes the identified request. The request may be a user request. The voice recognition assistant service may be implemented using an artificial intelligence model. In one embodiment of the present disclosure, the voice recognition assistant service may be performed using an artificial intelligence model that infers the content of a voice signal by inputting a voice signal. The voice recognition assistant service may be referred to as, but is not limited to, a voice assistant, a virtual assistant, or a voice control system.
[0042] In one embodiment of the present disclosure, a keyword may include a syllable or word for initiating or executing a specific function. For example, the keyword may include a word for a voice recognition assistant service. In one embodiment of the present disclosure, the keyword may include a predetermined syllable or word, or may be arbitrarily determined by the user. The keyword may be referred to as, but is not limited to, a wake-up word, a wake word, a wake phrase, a call word, an activation word, or a trigger word.
[0043] FIG. 1 is a schematic diagram illustrating a method for speech recognition according to one embodiment of the present disclosure.
[0044] Referring to FIG. 1, an electronic device (100) can recognize a user's (110) voice. The electronic device (100) can determine or identify whether a keyword is included in a voice signal corresponding to the user's (110) utterance. The electronic device (100) can initiate or execute a specific function based on whether the keyword is included in the voice signal. For example, if the keyword is included in the voice signal, the electronic device (100) can execute a voice recognition assistant service. According to one embodiment, the electronic device (100) can recognize content uttered in association with a keyword and perform a specific function based on the content of the user's (110) utterance. For example, the electronic device (100) can recognize content included in an utterance of the user (110) input after a keyword is uttered and perform a specific function based on the content of the user's utterance. For example, the electronic device (100) may recognize one or more commands included after a keyword in a voice signal corresponding to a user's (110) speech, and perform a specific function or task according to one or more commands in the user's (110) speech. However, the present disclosure is not limited thereto, and according to one embodiment, one or more commands may be spoken before the keyword. In one embodiment of the present disclosure, the electronic device (100) may provide at least one of an image, text, or sound based on whether a keyword is included in the acquired voice signal. For example, the electronic device (100) may provide an image or text through a display, or provide sound through a speaker.
[0045] In one embodiment of the present disclosure, the electronic device (100) can change or add a keyword for performing a specific function. In one embodiment of the present disclosure, the keyword can include a preset word. For example, if no separate setting is changed, the electronic device (100) can determine whether “Hi Bixby” is included in a voice signal by using “Hi Bixby” as a keyword. In one embodiment of the present disclosure, the electronic device (100) can perform voice recognition using a keyword input by the user (110). For example, the electronic device (100) can perform voice recognition using a keyword input by the user (110) (“Hi Galaxy”) instead of the preset keyword (“Hi Bixby”). However, the present disclosure is not limited thereto, and according to one embodiment, the electronic device (100) can perform voice recognition using the preset keyword (“Hi Bixby”) and / or the keyword input by the user (110) (“Hi Galaxy”). For example, the electronic device (100) can perform voice recognition using a pre-set keyword (e.g., “Hello, Bixby”) or a keyword entered by the user (110) (e.g., “Hello, Galaxy”).
[0046] In one embodiment of the present disclosure, an electronic device (100) can obtain a text input regarding a new keyword. The electronic device (100) can enroll (or register) the keyword based on the text input. For example, the electronic device (100) can enroll the keyword without test voice data of the user (110) regarding the new keyword. In one embodiment of the present disclosure, the electronic device (100) can obtain test voice data by inducing the user (110) to speak regarding the keyword. This method can improve the accuracy of voice recognition. However, the present disclosure is not limited thereto, and the electronic device (100) can improve the user experience by registering a new keyword using only text input, without requiring the user (110) to speak regarding the new keyword. The electronic device (100) according to one embodiment of the present disclosure includes obtaining the user's test voice data to register the keyword, and optionally can obtain the test voice data to improve the accuracy of voice recognition.
[0047] FIG. 2 is a flowchart of a method for speech recognition according to one embodiment of the present disclosure.
[0048] In one embodiment of the present disclosure, a method for voice recognition may be performed by an electronic device (100). For example, the electronic device (100) may perform each operation of the method for voice recognition by having a processor of the electronic device (100) execute at least one instruction contained in a memory.
[0049] In operation S210, the method may include an operation of obtaining a text input including a keyword. For example, the electronic device (100) may obtain a text input including a keyword. For example, the electronic device (100) may obtain text type information including words or syllables representing the keyword.
[0050] In one embodiment of the present disclosure, the electronic device (100) can obtain a text input representing a keyword through an input interface. For example, during the process of registering a new keyword, the electronic device (100) can display text, voice, or an image that prompts a text input for the keyword. The electronic device (100) can obtain text-type information representing the new keyword through an input interface (e.g., a touchscreen, a touchpad, a keypad).
[0051] In one embodiment of the present disclosure, the electronic device (100) can identify a keyword stored in a memory. The electronic device (100) can store a text input representing the keyword in the memory and identify the keyword stored in the memory.
[0052] In operation S220, the method may include an operation of acquiring a voice signal. For example, the electronic device (100) may acquire a voice signal corresponding to the user's speech. In one embodiment of the present disclosure, the voice signal may include a signal acquired from a wave generated by the user's speech. The voice signal may include an analog signal identified from the wave and / or a digital signal digitized from an analog signal.
[0053] In one embodiment of the present disclosure, the electronic device (100) can acquire a voice signal through an input interface. For example, the electronic device (100) can acquire the voice signal in analog form using a microphone. The electronic device (100) can convert the acquired analog signal into a digital signal. The electronic device (100) can store the voice signal in an analog or digital signal format.
[0054] In one embodiment of the present disclosure, the voice signal acquired by the electronic device (100) may include various sounds, including a voice signal generated by the user's speech. For example, the electronic device (100) may additionally acquire ambient noise (or voice) that is not generated by the user, in addition to the voice signal generated by the user's speech. In operation S220, the electronic device (100) acquiring a voice signal generated by the user's speech does not exclude an operation of acquiring other voice signals.
[0055] In one embodiment of the present disclosure, the electronic device (100) can obtain a voice signal in response to satisfying a condition. For example, the electronic device (100) can obtain a voice signal that satisfies a condition based on the intensity of a sound obtained through an input interface. For example, the electronic device (100) can obtain a voice signal in response to the intensity of a sound obtained through the input interface being greater than or equal to a predetermined level. For example, the electronic device (100) can obtain a voice signal in response to the intensity of the obtained sound being greater than or equal to 60 dB.
[0056] The electronic device (100) can acquire a voice signal in response to a change in the intensity of a sound acquired through an input interface being greater than or equal to a predetermined amount. For example, the electronic device (100) can acquire a voice signal in response to a 20 dB increase in the intensity of the acquired sound from 50 dB to 70 dB.
[0057] In operation S230, the method may include an operation of obtaining a probability value regarding a keyword. The electronic device (100) may use a keyword adaptive detection model to determine a probability value regarding whether the acquired voice signal includes a keyword.
[0058] In one embodiment of the present disclosure, the probability value may include a confidence value or score indicating whether a keyword is included in the acquired speech signal. For example, the probability value may be expressed as a confidence value of a keyword-adaptive detection model indicating whether a keyword is included in the speech signal, and may be expressed as a value between 0 and 1. However, the present disclosure is not limited thereto, and according to one embodiment, the probability value may be expressed as a score indicating whether a keyword is included in the speech signal, and the score may be determined without an upper bound and / or a lower bound.
[0059] In one embodiment of the present disclosure, the keyword adaptive detection model may include an artificial intelligence model trained to input a speech signal and output a probability value. The keyword adaptive detection model may be trained using a speech training data set. The keyword adaptive detection model may include a Large Language Model (LLM).
[0060] In operation S240, the method may include an operation of obtaining a threshold value corresponding to a keyword. For example, the electronic device (100) may obtain a threshold value for the keyword using a threshold value determination model.
[0061] In one embodiment of the present disclosure, a threshold value may include a reference value that is compared with a probability value to determine whether a keyword is included in a speech signal. The threshold value may be determined differently depending on the keyword. For example, a first keyword may have a first threshold value, and a second keyword may have a second threshold value. For example, the first keyword may have a first threshold value based on the first syllable of the first keyword, and the second keyword may have a second threshold value based on the second syllable of the second keyword. For example, the threshold value may be determined differently if some syllables of the keyword are changed. Similarly, the probability value may also be determined differently depending on the keyword.
[0062] In one embodiment of the present disclosure, the threshold determination model may include an artificial intelligence model trained to input text input and output a threshold value. The threshold determination model may be trained using a text training data set corresponding to a speech training data set of a keyword adaptive detection model.
[0063] In operation S250, the method may include an operation of determining whether a keyword is included in the acquired voice signal. The electronic device (100) may determine whether a keyword is included in the acquired voice signal based on a probability value and a threshold value.
[0064] In one embodiment of the present disclosure, the electronic device (100) can compare a probability value and a threshold value. For example, the electronic device (100) can identify whether the probability value is greater than the threshold value. However, the present disclosure is not limited thereto, and according to one embodiment, the electronic device (100) can identify a difference (or ratio) between the probability value and the threshold value. For example, the electronic device (100) can identify whether the difference (or ratio) between the probability value and the threshold value satisfies a criterion.
[0065] The electronic device (100) can determine whether a keyword is included in a voice signal based on the comparison result of a probability value and a threshold value. The electronic device (100) can determine whether a keyword is included in a voice signal in response to a probability value being greater than the threshold value. The electronic device (100) can determine whether a keyword is included in a voice signal when a difference (or ratio) between the probability value and the threshold value is greater than a predetermined value.
[0066] Operations S210 to S250 described in FIG. 2 describe an exemplary method for voice recognition. The electronic device (100) may omit at least some of operations S210 to S250 or additionally perform other operations.
[0067] FIG. 3 is a block diagram of an electronic device that determines whether a voice signal contains a keyword according to one embodiment of the present disclosure.
[0068] Referring to FIG. 3, the electronic device (100) may include a keyword adaptive detection model (310), a threshold value determination model (320), and a comparison unit (330).
[0069] The keyword adaptive detection model (310) can acquire a voice signal. In one embodiment of the present disclosure, the voice signal may include a voice signal acquired through an input interface of the electronic device (100). However, the present disclosure is not limited thereto, and according to one embodiment, the voice signal may include a voice signal stored in the memory of the electronic device (100).
[0070] The voice signal may include a voice signal generated by the user's speech. However, the voice signal is not limited thereto and may also include background sounds. The voice signal may also include background sounds or noises not generated by the user's speech.
[0071] A speech signal may comprise a continuous audio stream. For example, the speech signal may comprise a speech signal for sounds acquired in real time through an input interface. In one embodiment of the present disclosure, the continuous speech signal may be divided into multiple units and processed.
[0072] A speech signal can be acquired discretely. For example, a speech signal can be acquired based on satisfying certain conditions. For example, a speech signal can be acquired when the intensity of an ambient sound exceeds a predetermined level.
[0073] The keyword adaptive detection model (310) can determine a probability value indicating whether a keyword is included in an input speech signal. For example, the probability value may be determined differently depending on the training data of the keyword adaptive detection model (310). For example, the probability value may be determined to be higher when the keyword adaptive detection model (310) includes a large number of keywords or words similar to keywords.
[0074] The keyword adaptive detection model (310) can recognize content included in a speech signal. In one embodiment of the present disclosure, the keyword adaptive detection model (310) can extract features from an input speech signal. For example, the keyword adaptive detection model (310) can determine a feature vector for the speech signal. The keyword adaptive detection model (310) can recognize content included in the speech signal based on the features. For example, the keyword adaptive detection model (310) can recognize syllables, words, or sentences corresponding to the feature vector. The keyword adaptive detection model (310) can recognize content included in the speech signal based on stored feature vector data. However, the present disclosure is not limited thereto, and according to one embodiment, the keyword adaptive detection model (310) can recognize content included in the speech signal based on an acoustic model trained using a plurality of feature vectors. In one embodiment of the present disclosure, the keyword adaptive detection model (310) can recognize content included in the speech signal using a language model. For example, a language model can determine the probability of content inferred from an acoustic model. If the probability of content inferred from the acoustic model is low, the language model can modify some of the content.
[0075] The keyword adaptive detection model (310) is trained together with the threshold determination model (320), so that the probability value determined by the keyword adaptive detection model (310) may be correlated with the threshold value determined by the threshold determination model (320).
[0076] In one embodiment of the present disclosure, the keyword adaptive detection model (310) may include an artificial intelligence model trained to input a speech signal and output a probability value. The process of training the keyword adaptive detection model (310) according to one embodiment of the present disclosure will be described with reference to FIG. 4.
[0077] The threshold value determination model (320) can obtain keyword text. The threshold value determination model (320) can obtain a text input representing a keyword. The threshold value determination model (320) can obtain the text input from an input interface or a memory. The memory may store keywords set in advance by the user. The keywords may include keywords set by the user rather than keywords set in advance by the manufacturer of the electronic device (100). For example, the manufacturer of the electronic device (100) may set "Hi Bixby" as the initial keyword, and the user may set the keyword as "Hi Galaxy."
[0078] In one embodiment of the present disclosure, the threshold value determination model (320) can determine a threshold value for a keyword. The threshold value can include a reference value that is compared with a probability value to determine whether a keyword is included in a speech signal.
[0079] In one embodiment of the present disclosure, the threshold value determination model (320) may include an artificial intelligence model trained to input text input and output a threshold value. The process of training the threshold value determination model (320) according to one embodiment of the present disclosure will be described with reference to FIG. 4.
[0080] In one embodiment of the present disclosure, if a user voice signal speaking a keyword is stored, the electronic device (100) may compare the stored user voice signal with an acquired voice signal to determine whether the keyword is included in the voice signal. The electronic device (100) may not store a user voice signal speaking a keyword. For example, when registering a new keyword, the electronic device (100) may not acquire test voice data of a user speaking the keyword. If the electronic device (100) does not store a user voice signal for a keyword, the electronic device (100) may determine whether the keyword is included in the input voice signal based on a text input regarding the keyword. The electronic device (100) may determine whether the keyword is included in the input voice signal using a probability value of a keyword adaptive detection model (310) and a threshold value of a threshold value determination model (320).
[0081] The comparison unit (330) can obtain a probability value and a threshold value. Based on the probability value and the threshold value, the comparison unit (330) can determine whether a keyword is included in the voice signal. The comparison unit (330) can compare the probability value and the threshold value. Based on the comparison result, the comparison unit (330) can determine whether a keyword is included in the voice signal.
[0082] In one embodiment of the present disclosure, the comparison unit (330) can determine whether a probability value is greater than or equal to a threshold value. If the probability value is greater than or equal to the threshold value, the comparison unit (330) can determine that a keyword is included in the speech signal. If the probability value is greater than or equal to the threshold value, it can be understood that a keyword is included in the speech signal.
[0083] The electronic device (100) can determine whether to execute a specific function based on a probability value and a threshold value. Based on the comparison unit (330) determining that the voice signal includes a keyword, the electronic device (100) can execute a predetermined function. In response to the comparison unit (330) determining that the voice signal includes a keyword, the electronic device (100) can decide to execute (or accept) the predetermined function. Additionally, in response to the comparison unit (330) determining that the voice signal does not include a keyword, the electronic device (100) can decide not to execute (or reject) the predetermined function. The electronic device (100) can execute a voice recognition assistant service if the voice signal includes a keyword. The function executed by the electronic device (100) can be set differently depending on the keyword.
[0084] Although FIG. 3 is described with three models, a keyword adaptive detection model (310), a threshold value determination model (320), and a comparison unit (330), the present invention is not limited thereto, and at least some of the keyword adaptive detection model (310), the threshold value determination model (320), and the comparison unit (330) may be merged or subdivided into one model.
[0085] FIG. 4 is a diagram illustrating a process for training models of an electronic device according to one embodiment of the present disclosure.
[0086] Referring to FIG. 4, a keyword adaptive detection model (310) and a threshold value determination model (320) for speech recognition may be trained together. In one embodiment of the present disclosure, the keyword adaptive detection model (310) may be a model pre-trained to identify words included in an input speech signal.
[0087] In one embodiment of the present disclosure, the keyword adaptive detection model (310) may be trained using a speech training data set. The threshold determination model (320) may be trained using a text training data set corresponding to the speech training data set of the keyword adaptive detection model (310). For example, if the speech training data set includes speech data saying "Samsung," the text training data set may include textual data of "Samsung."
[0088] In one embodiment of the present disclosure, the probability value for a keyword of the keyword adaptive detection model (310) may be determined differently depending on the voice training data set. For example, if the voice training data set includes many words that are at least partially identical to the keyword, the keyword adaptive detection model (310) may determine a high probability value. For example, if the registered keyword is "Hi Galaxy," the more identical words there are in the voice training data set, the higher the probability value may be determined by the keyword adaptive detection model (310). For example, the keyword adaptive detection model (310) may determine a high probability value when the voice training data set includes a phrase that is completely identical to the keyword, such as "Hello, Galaxy," and / or multiple words that are partially identical to the keyword, such as "Hello" or "Galaxy." Similarly, if the voice training data set includes few words that are at least partially identical to the keyword, the keyword adaptive detection model (310) may determine a low probability value for the keyword.
[0089] Since the keyword adaptive detection model (310) determines probability values differently depending on the training data and keywords, the threshold value must be determined differently depending on the keywords. For example, if the keyword adaptive detection model (310) indicates a large probability value for a keyword, the threshold value corresponding to the keyword also needs to be large. For example, if the keyword adaptive detection model (310) indicates a small probability value for a keyword, the threshold value corresponding to the keyword also needs to be small. In an example where the threshold value is not adaptively determined depending on the keyword but is constant, a problem may occur in which a keyword may be determined to be included in a speech signal even though it is included, or a keyword may be determined to be included in a speech signal even though it is not included.
[0090] Similar to the keyword adaptive detection model (310), the threshold of the threshold determination model (320) may be determined differently depending on the text training data set. The threshold determination model (320) may include a language model. The language model may perform inference by dividing text into minimal units (e.g., syllables or tokens). The language model may predict a subsequent unit based on the current unit. For example, the language model may predict what the subsequent unit will be. Additionally, the language model may determine the probability of the subsequent unit. For example, if the current unit is "hai" or "ee," the probability that the subsequent unit will be "gall" or "galaxy" may be determined. In one embodiment of the present disclosure, the threshold may be determined as a confidence value for the keyword. The confidence value may be determined using the probability determined when a keyword is entered using the language model. For example, the threshold may be determined as the sum, weighted sum, average, or product of the probabilities determined for each unit when a keyword is entered.
[0091] In one embodiment of the present disclosure, the keyword adaptive detection model (310) may include a language model. The keyword adaptive detection model (310) may include the same language model as the threshold determination model (320), but is not limited thereto, and may include a different language model. For example, the language model of the keyword adaptive detection model (310) may be trained with the same data set as the language model of the threshold determination model (320).
[0092] In one embodiment of the present disclosure, the keyword adaptive detection model (310) can infer text included in a speech signal and determine a probability value of the inferred text using a language model.
[0093] In one embodiment of the present disclosure, the keyword adaptive detection model (310) can determine probability values for keywords using a speech recognition model and a language model.
[0094] FIG. 5 is a flowchart of an operation for obtaining a probability value for a voice signal according to one embodiment of the present disclosure.
[0095] Referring to FIG. 5, operation S230 may include operations S510 and S520. In one embodiment of the present disclosure, operations S510 and S520 may be performed by the electronic device (100). For example, the electronic device (100) may perform each of operations S510 and S520 by having the processor of the electronic device (100) execute at least one instruction contained in a memory.
[0096] In operation S510, the method may include an operation of dividing the voice signal into multiple units. For example, the electronic device (100) may divide the voice signal into multiple units. In one embodiment of the present disclosure, the electronic device (100) may divide the voice signal into predetermined time interval units. For example, the electronic device (100) may divide the voice signal into 10 ms units.
[0097] In operation S520, the method may include an operation of sequentially inputting a plurality of segmented units into a keyword adaptive detection model to obtain a probability value regarding whether the input unit contains a keyword. For example, the electronic device (100) may sequentially input a plurality of segmented units into a keyword adaptive detection model to determine a probability value regarding whether the input unit contains a keyword.
[0098] In one embodiment of the present disclosure, the electronic device (100) can extract features for an input unit. The electronic device (100) can recognize content (e.g., syllables, words, or sentences) corresponding to the extracted features.
[0099] In one embodiment of the present disclosure, the electronic device (100) can determine a probability value regarding whether a keyword is included in an input unit based on recognized content. The electronic device (100) can determine a probability value regarding whether at least a portion of the keyword is included in the recognized content. For example, if the keyword is "Hi Galaxy," the recognized content may be "Hi," which is part of the keyword, "Hi Galaxy," which is identical to the keyword, "Hey Hi Galaxy," which includes the keyword, etc. The electronic device (100) can identify whether the recognized content is part of the keyword and determine a probability value using the content of sequentially input units corresponding to being part of the keyword. For example, if the recognized content is "Hi," which is part of the keyword, the electronic device (100) can determine a probability value by referring to the recognized content according to a subsequent unit. In one embodiment of the present disclosure, the process of the electronic device (100) determining the probability value is omitted since it has been previously described.
[0100] Operations S510 to S520 described in FIG. 5 describe exemplary detailed operations of operation S230. The electronic device (100) may omit at least some of operations S510 to S520 or additionally perform other operations.
[0101] FIG. 6 is a diagram illustrating a process of training at least one model of an electronic device using user voice data according to one embodiment of the present disclosure.
[0102] Referring to FIG. 6, the electronic device (100) may include a user voice database (Data Base; DB, 610). The keyword adaptive detection model (310), the threshold value determination model (320), and the comparison unit (330) have been described in detail with reference to FIGS. 3 and 4, and thus, any duplicate content will be omitted.
[0103] The comparison unit (330) can determine whether a keyword is included in the voice signal. The electronic device (100) can store the voice signal based on whether the keyword is included in the acquired voice signal. In an exemplary case where the keyword is not included in the acquired voice signal, the electronic device (100) can remove the voice signal. For example, the electronic device (100) may not store the voice signal based on a determination that the keyword is not included in the acquired voice signal. The electronic device (100) can store a voice signal that includes a keyword (or in which execution of a specific function is accepted) in the user voice DB (610).
[0104] In one embodiment of the present disclosure, the electronic device (100) can train a keyword adaptive detection model using a stored voice signal. The electronic device (100) can update the keyword adaptive detection model (310) using voice data included in the user voice DB (610). The electronic device (100) can train (or retrain) the keyword adaptive detection model (310) using the voice data included in the user voice DB (610). Since the voice data included in the user voice DB (610) is a voice signal including keywords, the prediction accuracy for keywords can be improved by training the keyword adaptive detection model (310) using voice data including keywords.
[0105] In one embodiment of the present disclosure, the electronic device (100) can update the threshold value determination model (320) or update the threshold value using voice data included in the voice DB (610). In the exemplary case where the keyword adaptive detection model (310) is trained using the user voice DB (610), the output probability value will change (increase), so the threshold value also needs to be changed. In one embodiment of the present disclosure, the electronic device (100) can update the threshold value based on the probability value corresponding to the stored voice signal in response to training the keyword adaptive detection model (310). In one embodiment of the present disclosure, the electronic device (100) can train the threshold value determination model using the stored voice signal based on training the keyword adaptive detection model (310).
[0106] In one embodiment of the present disclosure, the electronic device (100) may determine a threshold value based on a probability value for a voice signal included in the user voice DB (610). Since the probability value of the voice signal included in the user voice DB (610) will be higher than the threshold value prior to the update, the threshold value may be updated using the probability value of the voice signal included in the user voice DB (610). For example, the threshold value may be determined as an average of the probability values of the voice signals included in the user voice DB (610).
[0107] In one embodiment of the present disclosure, the electronic device (100) can train (or retrain) a threshold value determination model (320) using voice data included in a user voice DB (610). Depending on the training, the threshold value determination model (320) can determine a different threshold value for the same keyword than before.
[0108] In one embodiment of the present disclosure, the electronic device (100) can periodically update models and / or threshold values using the user voice DB (610). For example, when a predetermined number of voice data is stored in the user voice DB (610), the electronic device (100) can update models and / or threshold values. However, the present disclosure is not limited thereto, and according to one embodiment, the electronic device (100) can update models and / or threshold values at predetermined intervals (e.g., a predetermined period of time such as a week or a month).
[0109] FIG. 7 is a diagram for explaining a process for determining whether a voice signal according to one embodiment of the present disclosure is a voice signal obtained by a user's speech.
[0110] Referring to FIG. 7, the electronic device (100) may include a user embedding model (710). The keyword adaptive detection model (310), the threshold value determination model (320), the comparison unit (330), and the user voice DB (610) have been described in detail with reference to FIGS. 3, 4, and 6, and thus, any duplicated content will be omitted.
[0111] In one embodiment of the present disclosure, the electronic device (100) can recognize the speaker of the acquired voice signal. The electronic device (100) can determine whether the speaker of the acquired voice signal is an authorized user and / or an existing user. The electronic device (100) can execute a specific function based on the voice signal of the authorized user and / or the existing user.
[0112] In one embodiment of the present disclosure, the electronic device (100) can determine the similarity between a stored voice signal and an acquired voice signal. The electronic device (100) can determine the similarity between voice data stored in the user voice DB (610) and the acquired voice signal. The electronic device (100) can determine the similarity by comparing the characteristics of the stored voice signal and the acquired voice signal.
[0113] In one embodiment of the present disclosure, the electronic device (100) can input the acquired voice signal and / or voice data stored in the user voice DB (610) into a user embedding model (710). The user embedding model (710) can embed the input voice signal. The user embedding model (710) can extract features of the input voice signal. For example, the user embedding model (710) can determine a feature vector for the input voice signal. The feature vector can include distinguishing feature information for the voice signals.
[0114] In one embodiment of the present disclosure, the electronic device (100) can determine the similarity between voice data stored in the user voice DB (610) and the acquired voice signal. The electronic device (100) can determine the similarity between embedded voice data and the embedded voice signal. The electronic device (100) can determine the similarity using a feature vector determined by the user embedding model (710). For example, the electronic device (100) can determine the similarity by calculating the cosine similarity between the feature vectors determined by the user embedding model (710).
[0115] In one embodiment of the present disclosure, the electronic device (100) can determine whether the acquired voice signal is from a registered user based on similarity. In an exemplary case where the similarity is greater than or equal to a predetermined value, the electronic device (100) can determine that the acquired voice signal is from a registered user.
[0116] In one embodiment of the present disclosure, registered users may include users who have previously used the voice recognition function. In an exemplary case where the electronic device (100) does not initially perform separate user registration, the electronic device (100) stores the voice signal of a user uttering a keyword in the user voice database (610) and recognizes the speaker using the user voice database (610). Accordingly, the electronic device (100) can recognize a user who has previously performed voice recognition using a keyword. Even if the electronic device (100) does not initially acquire voice data from the user, it can identify the speaker using the user voice data acquired while using voice recognition.
[0117] In one embodiment of the present disclosure, the electronic device (100) can recognize a speaker based on a confidence value for recognizing the speaker being greater than or equal to a predetermined value. For example, the electronic device (100) can perform speaker recognition only when a significant number of voice data is stored in the user voice DB (610). Speaker recognition includes identifying the speaker of the acquired voice signal by the electronic device (100). For example, speaker recognition may include determining whether the acquired voice signal corresponds to a person stored in the electronic device (100) or determining which of the stored persons it represents.
[0118] FIG. 7 illustrates an example of an electronic device (100) performing voice recognition and then speaker recognition according to an embodiment of the present disclosure. However, the electronic device is not limited thereto, and speaker recognition may be performed first and then voice recognition may be performed. For example, the electronic device (100) may compare the acquired voice signal with voice data contained in the user voice DB (610) to determine whether they are similar, and if so, perform voice recognition.
[0119] In one embodiment of the present disclosure, a keyword can be registered in an electronic device (100) with reference to FIGS. 8a, 8b, 9a, 9b, and 9c. In one embodiment of the present disclosure, the user interface (UI) described in FIGS. 8a, 8b, 9a, 9b, and 9c can be output through a display of the electronic device (100).
[0120] FIG. 8A is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0121] In one embodiment of the present disclosure, as illustrated in FIG. 8A, the first user interface (810) may include a UI component that controls general settings for voice calls (or, a voice assistant service, a function for determining whether a voice signal contains a keyword). The first user interface (810) may include a UI component configured to control whether to use voice calls. For example, the UI component that controls whether to use voice calls may include a toggle, a sliding bar, a button, and / or a tab.
[0122] The first user interface (810) may include a UI component (814) for setting keywords. The UI component (814) for setting keywords may be a user interface configured to select from among preset keywords or to select a new keyword. For example, the UI component (814) for setting keywords may include a user interface configured to select from among a user interface for displaying preset keywords or a UI component (816) for registering a new keyword.
[0123] FIG. 8b is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0124] As illustrated in FIG. 8B , in one embodiment of the present disclosure, the second user interface (820) may include a UI component configured to control operations related to registering or registering a new keyword. For example, the electronic device (100) may proceed to the second user interface (820) in response to a UI component (816) for registering a new keyword.
[0125] In one embodiment of the present disclosure, the second user interface (820) may include a user interface that provides instructions on how to register a keyword. For example, the second user interface (820) may provide an example keyword such as "Hi Galaxy" and an explanation such as "Enter this." The second user interface (820) may include a UI component (825) for proceeding to the next user interface. In one embodiment of the present disclosure, the UI component (825) may lead to the third user interface (910) of FIG. 9 .
[0126] FIG. 9A is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0127] In one embodiment of the present disclosure, as illustrated in FIG. 9A, the third user interface (910) may include a user interface for receiving keyword input. For example, the third user interface (910) may provide instructions prompting the user to input a keyword.
[0128] The third user interface (910) may include a UI component (912) for entering keywords. For example, the UI component (912) may include a UI component for entering characters via a keyboard or a user's touch input. The electronic device (100) may obtain keywords entered into the UI component (912).
[0129] The third user interface (910) may include a UI component (914) that displays keywords. The UI component (914) may display keywords entered through the UI component (912).
[0130] The third user interface (910) may include a UI component (916) that displays progress. The UI component (916) may be a UI component that indicates whether a keyword is being entered, before being entered, or has been completed.
[0131] In one embodiment of the present disclosure, if there is no keyword input in the third user interface (910) for a predetermined period of time or information corresponding to completion is input, the process may proceed to the fourth user interface (920).
[0132] FIG. 9b is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0133] In one embodiment of the present disclosure, as illustrated in FIG. 9B , the fourth user interface (920) may include a user interface for confirming keywords. The fourth user interface (920) may include a UI component (922) for displaying keywords. The UI component (922) may display keywords entered in the third user interface (910).
[0134] The fourth user interface (920) may include a UI component (924) that displays progress. The UI component (924) may be a UI component that indicates whether input has been completed. For example, the UI component (924) may be expressed differently from the UI component (916) that indicates whether a keyword is being input or before being input.
[0135] In one embodiment of the present disclosure, the fourth user interface (920) may include a UI component (926) that proceeds to the next user interface.
[0136] FIG. 9c is a diagram illustrating a user interface for registering a keyword according to one embodiment of the present disclosure.
[0137] In one embodiment of the present disclosure, as illustrated in FIG. 9c, the fifth user interface (930) may include a user interface that is output after keyword registration is completed.
[0138] The fifth user interface (930) may include a UI component (932) that displays progress. The UI component (932) may be a UI component indicating that input has been completed. For example, the UI component (932) may be expressed in the same manner as the UI component (924).
[0139] In one embodiment of the present disclosure, the fifth user interface (930) may include a UI component (934) that terminates keyword registration.
[0140] In one embodiment of the present disclosure, the first user interface (810) to the fifth user interface (930) of FIGS. 8 and 9 are exemplary, and the present disclosure is not limited thereto. In one embodiment of the present disclosure, some interfaces of the first user interface (810) to the fifth user interface (930) may be omitted, or additional interface configurations may be provided. Furthermore, some UI components included in each user interface may be omitted or added.
[0141] FIG. 10 is a diagram of a system in which voice recognition is performed using a registered keyword according to one embodiment of the present disclosure.
[0142] In one embodiment of the present disclosure, the electronic device (100) may be, but is not limited to, a smartphone, a tablet PC, a PC, a smart TV, a mobile phone, a personal digital assistant (PDA), a laptop, a media player, a server, a micro server, a global positioning system (GPS) device, an e-book reader, a digital broadcasting terminal, a navigation device, a kiosk, an MP3 player, a digital camera, a speaker, or any other mobile or non-mobile computing device having a voice recognition function.
[0143] Referring to FIG. 10, a user (110) can make a voice call in a system including a smartphone (1010), a speaker (1020), and a tablet PC (1030). The user (110) can call the electronic device (100) by saying a keyword. In an exemplary case where the keywords of the smartphone (1010), the speaker (1020), and the tablet PC (1030) are the same, the smartphone (1010), the speaker (1020), and the tablet PC (1030) can all respond to the user's (110) utterance of the keyword. For example, the smartphone (1010), the speaker (1020), and the tablet PC (1030) can all respond to the user's (110) keyword "Hi Bixby." In this case, the smartphone (1010), the speaker (1020), and the tablet PC (1030) can each execute a specific function in response to the keyword. For example, each electronic device can provide at least one of an image, text, or sound.
[0144] In one embodiment of the present disclosure, the user (110) can register different keywords for at least some of the smartphone (1010), the speaker (1020), and the tablet PC (1030). For example, the user (110) can set "Hi Galaxy" for the smartphone (1010). The smartphone (1010) may not respond to the user's (110) utterance of the existing keyword ("Hi Bixby"), but may respond to the new keyword ("Hi Galaxy"). Accordingly, when calling multiple electronic devices with voice recognition functions, the user (110) can distinguish and call the electronic devices by setting different keywords for each.
[0145] In one embodiment of the present disclosure, an electronic device (100) including a smartphone (1010) can register keywords using only text input, without a user's voice input. This can enhance user convenience (110).
[0146] FIG. 11 is a block diagram of an electronic device for voice recognition according to one embodiment of the present disclosure.
[0147] In one embodiment of the present disclosure, an electronic device (100) may include a processor (1110) and a memory (1120).
[0148] The processor (1110) can control the overall operations of the electronic device (100). For example, the processor (1110) can control the overall operations of the electronic device (100) for performing voice recognition by executing one or more instructions of a program stored in the memory (1120).
[0149] In one embodiment of the present disclosure, the processor (1110) may include a configuration that controls a series of processes so that the electronic device (100) operates according to the embodiments described in the present disclosure. The processor (1110) may be composed of one or more processors. The one or more processors included in the processor (1110) may include circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), and the like. The processor (1110) may be one or more processors including, but not limited to, a central processing unit, a microprocessor unit, an application processor, a digital signal processor (DSP), a graphic processing unit, a vision processing unit (VPU), an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a neural processing unit, a communication processor, and / or an artificial intelligence processor designed with a hardware structure specialized for processing an artificial intelligence model.
[0150] Meanwhile, although not illustrated in FIG. 11, the electronic device (100) may further include additional components to perform the operations described in the aforementioned embodiments. For example, the electronic device (100) may further include a display, a camera, a microphone, a speaker, an input / output interface, and the like.
[0151] In an exemplary case where a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first and second operations may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an AI-specific processor). Here, an AI-specific processor, which is an example of the second processor, may perform operations for training / inference of an AI model. However, the embodiments of the present disclosure are not limited thereto.
[0152] One or more processors (1110) according to the present disclosure may be implemented as a single-core processor or as a multi-core processor.
[0153] When an exemplary method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one core or may be performed by multiple cores included in one or more processors.
[0154] At least one processor (1110) according to an embodiment of the present invention may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits including at least one processor, one or more of which are configured to individually and / or collectively perform the various functions described herein in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms encompass, for example, without limitation, a single processor performing some of the recited functions, other processor(s) performing other of the recited functions, and even a single processor performing all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform the various functions.
[0155] The memory (1120) may store instructions, data structures, and program codes that can be read by the processor (1110). Operations performed by the processor (1110) may be implemented by executing instructions or codes of a program stored in the memory (1120).
[0156] The memory (1120) may include at least one of volatile memory and non-volatile memory. For example, the memory (1120) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, an optical disk, a RAM (Random Access Memory), or a SRAM (Static Random Access Memory).
[0157] The memory (1120) may store one or more instructions and / or programs that cause the electronic device (100) to perform voice recognition. For example, the memory (1120) may store instructions and / or programs for implementing operations for performing voice recognition. The processor (1110) may write data to the memory (1120) or read data stored in the memory (1120). The processor (1110) may process data according to predefined operation rules or artificial intelligence models by executing the program or at least one instruction stored in the memory (1120). The processor (1110) may perform operations described in the embodiments of the present disclosure. Optionally, operations described as being performed by the electronic device (100) or detailed components included in the electronic device (100) in the embodiments of the present disclosure may be performed by the processor (1110).
[0158] Meanwhile, the memory (1120) may further store instructions and / or programs for implementing functions of an automatic speech recognition module (not shown).
[0159] In one embodiment of the present disclosure, the processor (1110) may be configured to obtain an input image by performing one or more instructions included in the memory (1120). The processor (1110) may obtain a text input for a keyword by performing one or more instructions included in the memory (1120). The processor (1110) may obtain a voice signal corresponding to a user's speech by performing one or more instructions included in the memory (1120). The processor (1110) may determine a probability value regarding whether a keyword is included in the obtained voice signal by using a keyword-adaptive detection model by performing one or more instructions included in the memory (1120). The processor (1110) may obtain a threshold value for a keyword by using a threshold value determination model by performing one or more instructions included in the memory (1120). The processor (1110) can determine whether a keyword is included in an acquired voice signal based on a probability value and a threshold value by performing one or more instructions included in the memory (1120).
[0160] In one embodiment of the present disclosure, the electronic device (100) may include other components in addition to the processor (1110) and the memory (1120). Components that the electronic device (100) may include in one embodiment of the present disclosure are described in detail with reference to FIG. 12.
[0161] FIG. 12 is a block diagram of an electronic device for voice recognition according to one embodiment of the present disclosure.
[0162] As illustrated in FIG. 12, an electronic device (100) according to one embodiment of the present disclosure may further include a communication interface (1210) and / or a user interface (1220) in addition to a processor (1110) and a memory (1120).
[0163] The processor (1110) controls the operation of the electronic device (100). The processor (1110) can control the communication interface (1210), and / or the user interface (1220), and the memory (1120) by executing programs stored in the memory (1120).
[0164] The memory (1120) may store programs for processing and controlling the processor (1110), and may also store input / output data (e.g., user voice data, etc.). The memory (1120) may also store an artificial intelligence model. For example, the memory (1120) may store an automatic speech recognition (ASR) model, a natural language understanding (NLU) model, and / or a text-to-speech (TTS) model.
[0165] The communication interface (1210) may include one or more components that enable communication between an electronic device (100) and a server device (not shown), or an electronic device (100) and a mobile terminal (not shown). For example, the communication interface (1210) may include a short-range communication unit (1212), a long-range communication unit (1214), etc.
[0166] The short-range communication unit (1212) may include, but is not limited to, a Bluetooth communication unit, a BLE (Bluetooth Low Energy) communication unit, a near field communication unit (NFC, Near Field Communication unit), a WLAN (Wi-Fi) communication unit, a Zigbee communication unit, an infrared (IrDA, infrared Data Association) communication unit, a WFD (Wi-Fi Direct) communication unit, a UWB (ultra wideband) communication unit, an Ant+ communication unit, etc.
[0167] The remote communication unit (1214) may include the Internet, a computer network (e.g., a LAN or WAN), and a mobile communication unit. The mobile communication unit transmits and receives a wireless signal with at least one of a base station, an external terminal, and a server on the mobile communication network. Here, the wireless signal may include various types of data according to a voice call signal, a video call call signal, or a text / multimedia message transmission and reception. The mobile communication unit may include, but is not limited to, a 3G module, a 4G module, a 5G module, an LTE module, an NB-IoT module, an LTE-M module, etc.
[0168] The user interface (1220) may include an output interface (1222) and an input interface (1224). The output interface (1222) is for outputting an audio signal or a video signal and may include a display and / or an audio output unit.
[0169] In one embodiment of the present disclosure, the display may be configured as a touch screen by forming a layer structure with a touchpad. When the display and the touchpad are configured as a touch screen by forming a layer structure, the display may be used as an input interface (1224) in addition to an output interface (1222). The display may include at least one of a liquid crystal display, a thin film transistor-liquid crystal display, a light-emitting diode (LED), an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display. In addition, depending on the implementation form of the electronic device (100), the electronic device (100) may include two or more displays. For example, the electronic device (100) may include a front-facing display and a rear-facing display opposite to the front-facing display.
[0170] According to one embodiment of the present disclosure, the display can display and output information processed in the electronic device (100). For example, the display can display an image stored in the memory (1120) of the electronic device (100). The display can output an interface for controlling the electronic device (100), an interface for displaying the status of the electronic device (100), and the like.
[0171] The audio output unit may output audio data received from the communication interface (1210) or stored in the memory (1120). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the audio output unit may output an audio signal related to a function performed in the electronic device (100). The audio output unit may include a speaker, a buzzer, or the like. For example, a speaker or a buzzer may output a signal related to a function performed in the electronic device (100) (e.g., a call signal reception sound, a message reception sound, a notification sound) as sound.
[0172] The input interface (1224) can receive input from a user. The input interface (1224) can include, but is not limited to, at least one of a key pad, a dome switch, a touch pad (contact electrostatic capacitance type, pressure resistive film type, infrared detection type, surface ultrasonic conduction type, integral tension measurement type, piezo effect type, etc.), a jog wheel, a jog switch, and a microphone.
[0173] The microphone can receive audio signals. For example, the microphone can receive a voice signal corresponding to the user's speech. In addition to the user's voice, the microphone can also receive an audio signal that includes noise signals generated from multiple sound sources. The microphone can transmit the acquired audio signal to the processor (1110), thereby enabling a voice recognition service to be performed.
[0174] In one embodiment of the present disclosure, a method for speech recognition is provided. The method may include obtaining a text input including a keyword. The method may include obtaining a speech signal corresponding to a user's utterance. The method may include obtaining a probability value for the keyword using a keyword-adaptive detection model. The method may include obtaining a threshold value for the keyword using a threshold value determination model. The method may include determining whether the acquired speech signal includes the keyword based on the probability value and the threshold value.
[0175] In one embodiment of the present disclosure, a keyword adaptive detection model may be trained using a speech training data set. A threshold determination model may be trained using a text training data set corresponding to the speech training data set.
[0176] In one embodiment of the present disclosure, the keyword adaptive detection model may include an artificial intelligence model trained to output a probability value based on an input of an acquired speech signal.
[0177] A method wherein the threshold determination model comprises an artificial intelligence model trained to output a threshold value based on a text input.
[0178] In one embodiment of the present disclosure, the method may include a step of storing a speech signal based on whether the acquired speech signal contains a keyword. The method may include a step of training a keyword-adaptive detection model using the stored speech signal.
[0179] In one embodiment of the present disclosure, the method may include updating a threshold value based on a probability value corresponding to a stored speech signal, based on training a keyword adaptive detection model.
[0180] In one embodiment of the present disclosure, the method may include training a threshold determination model using a stored speech signal based on training a keyword adaptive detection model.
[0181] In one embodiment of the present disclosure, the method may include determining a similarity between a stored voice signal and an acquired voice signal. The method may also include determining whether the acquired voice signal is from a registered user based on the similarity.
[0182] In one embodiment of the present disclosure, the step of obtaining a probability value may include the step of segmenting the acquired speech signal into a plurality of units each containing one or more syllables. The step of obtaining a probability value may include the step of sequentially inputting the segmented plurality of units into a keyword-adaptive detection model, thereby acquiring a probability value regarding whether each segmented unit contains a keyword.
[0183] In one embodiment of the present disclosure, the threshold value determination model can obtain the threshold value without the user's voice information.
[0184] In one embodiment of the present disclosure, the method can provide at least one of an image, text, or sound based on whether a keyword is included in the acquired voice signal.
[0185] In one embodiment of the present disclosure, a computer-readable recording medium having a program recorded thereon is provided. The program may include a program for causing a computer to perform a method comprising any of the steps described above.
[0186] In one embodiment of the present disclosure, an electronic device for speech recognition is provided. The electronic device may include at least one processor including a processing circuit, and a memory including one or more storage media storing at least one instruction. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a text input including a keyword. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a voice signal corresponding to a user's utterance. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a probability value for the keyword using a keyword-adaptive detection model. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to obtain a threshold value for the keyword using a threshold value determination model. The at least one instruction may be individually or collectively executed by the at least one processor, thereby enabling the electronic device to determine whether the acquired voice signal includes the keyword based on the probability value and the threshold value.
[0187] In one embodiment of the present disclosure, a keyword adaptive detection model may be trained using a speech training data set. A threshold determination model may be trained using a text training data set corresponding to the speech training data set.
[0188] In one embodiment of the present disclosure, the keyword adaptive detection model may include an artificial intelligence model trained to output a probability value based on an input of an acquired speech signal. The threshold determination model may include an artificial intelligence model trained to output a threshold value based on a text input.
[0189] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor, thereby allowing an electronic device to store a speech signal based on whether the acquired speech signal contains a keyword. At least one instruction may be individually or collectively executed by at least one processor, thereby allowing the electronic device to train a keyword-adaptive detection model using the stored speech signal.
[0190] In one embodiment of the present disclosure, at least one instruction is individually or collectively executed by at least one processor so that the electronic device can update a threshold value based on a probability value corresponding to a stored speech signal based on training a keyword adaptive detection model.
[0191] In one embodiment of the present disclosure, at least one instruction is individually or collectively executed by at least one processor to enable the electronic device to train a threshold determination model using a stored speech signal based on which the electronic device trains a keyword adaptive detection model.
[0192] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor, thereby enabling an electronic device to determine a similarity between a stored voice signal and an acquired voice signal. At least one instruction may be individually or collectively executed by at least one processor, thereby enabling the electronic device to determine whether the acquired voice signal is from a registered user based on the similarity.
[0193] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor, thereby allowing an electronic device to segment an acquired speech signal into a plurality of units each containing one or more syllables. At least one instruction may be individually or collectively executed by at least one processor, thereby allowing the electronic device to sequentially input the segmented units into a keyword-adaptive detection model, thereby obtaining a probability value regarding whether each segmented unit contains a keyword.
[0194] In one embodiment of the present disclosure, the threshold value determination model can obtain the threshold value without the user's voice information.
[0195] In one embodiment of the present disclosure, the method can provide at least one of an image, text, or sound based on whether a keyword is included in the acquired voice signal.
[0196] According to one embodiment of the present disclosure, functions related to artificial intelligence are operated via a processor and memory. The processor may be comprised of one or more processors. In these examples, the one or more processors may be general-purpose processors such as a CPU, an AP, a Digital Signal Processor (DSP), a graphics-only processor such as a GPU or a Vision Processing Unit (VPU), or an artificial intelligence-only processor such as an NPU. The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence models stored in memory. In the exemplary case where one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0197] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is trained using a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0198] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0199] In a method for voice recognition of an electronic device according to the present disclosure, a method for recognizing a user's voice and interpreting the user's intent to determine whether a keyword is included in the voice signal includes receiving a voice signal, which is an analog signal, through an input / output device (e.g., a microphone), and converting the voice portion into computer-readable text using an Automatic Speech Recognition (ASR) model. The converted text can be interpreted using a Natural Language Understanding (NLU) model to obtain the user's utterance intent. Here, the ASR model or the NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by an artificial intelligence-dedicated processor designed with a hardware structure specialized for processing artificial intelligence models. The artificial intelligence model may be created through learning. Here, being created through learning means that a basic artificial intelligence model is learned using a plurality of learning data by a learning algorithm, thereby creating a predefined operation rule or artificial intelligence model set to perform a desired characteristic (or purpose). The artificial intelligence model may be composed of a plurality of neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weight values.
[0200] Linguistic understanding is the technology of recognizing, applying, and processing human language / characters, including natural language processing, machine translation, dialog systems, question answering, and speech recognition / synthesis.
[0201] According to one embodiment of the present disclosure, the computer-executable instructions, such as program modules executed by a computer, may also be implemented in the form of a recording medium. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.
[0202] A computer-readable storage medium according to one embodiment of the present disclosure may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is stored semi-permanently in the storage medium and cases where data is stored temporarily. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0203] A method according to one embodiment of the present disclosure may be provided as a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0204] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.
[0205] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. In a method for voice recognition, Step (S210) of obtaining text input including keywords; Step (S220) of acquiring a voice signal corresponding to the user's utterance; A step (S230) of obtaining a probability value for the keyword using a keyword adaptive detection model; A step (S240) of obtaining a threshold value for the keyword using a threshold determining model; and A method comprising a step (S250) of determining whether the keyword is included in the acquired voice signal based on the probability value and the threshold value.
2. In paragraph 1, The above keyword adaptive detection model is trained using a voice training data set, A method, characterized in that the threshold value determination model is trained using a text training data set corresponding to the voice training data set.
3. In any one of paragraphs 1 and 2, The above keyword adaptive detection model includes an artificial intelligence model trained to output the probability value based on the input of the acquired speech signal, A method wherein the threshold value determination model comprises an artificial intelligence model trained to output the threshold value based on the text input.
4. In any one of paragraphs 1 to 3, A step of storing the voice signal based on the keyword being included in the acquired voice signal; and A method comprising the step of training the keyword adaptive detection model using the stored speech signal.
5. In paragraph 4, A method comprising the step of updating the threshold value based on a probability value corresponding to the stored speech signal, based on training the keyword adaptive detection model.
6. In any one of paragraphs 4 to 5, A method comprising the step of training the threshold value determination model using the stored speech signal based on training the keyword adaptive detection model.
7. In any one of paragraphs 4 to 6, A step of determining the similarity between the stored voice signal and the acquired voice signal; and A method comprising a step of determining whether the acquired voice signal is from a registered user based on the similarity.
8. In any one of paragraphs 1 to 7, The step (S230) of obtaining the above probability value is: A step of dividing the acquired voice signal into a plurality of units including one or more syllables; and A method comprising the step of sequentially inputting the divided plurality of units into the keyword adaptive detection model and obtaining the probability value regarding whether the keyword is included in each of the divided units.
9. In any one of paragraphs 1 to 8, A method characterized in that the threshold value determination model obtains the threshold value without the user's voice information.
10. In any one of paragraphs 1 to 9, A method comprising the step of providing at least one of an image, text, or sound based on the keyword being included in the acquired voice signal.
11. In an electronic device for voice recognition, At least one processor comprising a processing circuit; and An electronic device comprising a memory including one or more storage media storing at least one instruction, wherein the at least one instruction is individually or collectively executed by the at least one processor, Obtain a text input containing keywords, Obtain a voice signal corresponding to the user's utterance, Using a keyword adaptive detection model, a probability value for the keyword is obtained, Using the threshold value determination model, obtain a threshold value for the above keyword, An electronic device that determines whether the keyword is included in the acquired voice signal based on the probability value and the threshold value.
12. In paragraph 11, The above keyword adaptive detection model is trained using a voice training data set, An electronic device, characterized in that the threshold value determination model is trained using a text training data set corresponding to the voice training data set.
13. In any one of paragraphs 11 to 12, The above keyword adaptive detection model includes an artificial intelligence model trained to output the probability value based on the input of the acquired speech signal, An electronic device, wherein the threshold value determination model includes an artificial intelligence model trained to output the threshold value based on the text input.
14. In any one of paragraphs 11 to 13, The electronic device, wherein the at least one instruction is individually or collectively executed by the at least one processor, Based on the keyword being included in the acquired voice signal, the voice signal is stored, An electronic device that trains the keyword adaptive detection model using the stored voice signal.
15. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 10 on a computer.
Citation Information
Patent Citations
Semiconductor memory device, method for manufacturing the same, and electronic system including the same
KR1020240082820A
Cookie with tteok-core and fruit chips and method of manufacturing the cookie
KR1020240114052A
Electrical connection member and electronic device comprising same
KR1020250107655A
Voice control system and wake-up method thereof, wake-up device and home appliance, coprocessor
KR102335717B1
User interfacing device and method for setting wake-up word activating speech recognition
KR102392992B1